Introducing Gemini 3.5 FlashGoogleReleased May 19, 2026

Gemini 3.5 Flash

Mid-generation multimodal speed model with configurable thinking levels, allowing developers to balance latency and reasoning depth.

multimodal-realtimeProprietary API$0.25 / 1M tok (with Thinking Levels)Context: 1.049M (1,048,576 tokens)Arena ELO: 1390

Technical Specifications

Architecture Type
Gemini 3 Reasoning Foundation
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
65.536K (65,536 tokens)
Knowledge Cutoff
February 2026
Supported Modalities
text, code, vision, audio, video
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Math500
89.1
Mmlu Pro
78.2
Humaneval
93.8
Swe Bench Verified
52.4
Chatbot Arena ELO
1390

Deep Architectural Overview

Gemini 3.5 Flash integrates thinking budget controls into the Flash architecture. Developers can tune reasoning effort from 0 (instant response) to maximum (deep analytical thinking), matching task complexity on the fly.

Strengths & Considerations

Core Strengths
  • Configurable thinking effort levels
  • Large 64k token output capacity
  • Outstanding price-to-performance ratio
Known Limitations
  • Maximum thinking levels increase latency

Token & API Pricing

Input Tokens (1M)$0.25
Output Tokens (1M)$1.00
Cached Input (1M)$0.0625
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)