Introducing Gemini 3 FlashGoogleReleased December 17, 2025

Gemini 3 Flash

The third-generation speed-optimized foundation model combining sub-second latency with near-Pro intelligence.

multimodal-realtimeProprietary API$0.15 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1360

Technical Specifications

Architecture Type
Gemini 3 Fast Multimodal Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
16.384K (16,384 tokens)
Knowledge Cutoff
September 2025
Supported Modalities
text, code, vision, audio, video
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Math500
84.2
Mmlu Pro
74.8
Humaneval
91.5
Chatbot Arena ELO
1360

Deep Architectural Overview

Gemini 3 Flash introduces Gemini 3 core improvements to high-speed inference. It outperforms previous Pro models on general reasoning benchmarks while operating at a fraction of the cost and compute latency.

Strengths & Considerations

Core Strengths
  • Outperforms Gemini 2.5 Pro at Flash prices
  • Sub-second response times across 1M context
  • Exceptional agentic tool calling reliability
Known Limitations
  • Complex multi-repository refactoring remains best on Pro tier

Token & API Pricing

Input Tokens (1M)$0.15
Output Tokens (1M)$0.60
Cached Input (1M)$0.0375
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)