Introducing Gemini 3.6 FlashGoogleReleased July 21, 2026

Gemini 3.6 Flash

Iterative performance leap in the Flash family featuring major upgrades in code refactoring and autonomous agent tool loops.

multimodal-realtimeProprietary API$0.35 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1415

Technical Specifications

Architecture Type
Gemini 3 Core Reasoning Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
65.536K (65,536 tokens)
Knowledge Cutoff
April 2026
Supported Modalities
text, code, vision, audio, video
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Math500
91.2
Mmlu Pro
80.4
Humaneval
94.6
Chatbot Arena ELO
1415

Deep Architectural Overview

Gemini 3.6 Flash bridges the gap between speed and frontier-grade capability, bringing 94.6% HumanEval performance and 64k token outputs to the high-efficiency Flash tier.

Strengths & Considerations

Core Strengths
  • 94.6% HumanEval code generation
  • 64k token single-response output
  • Excellent tool calling consistency
Known Limitations
  • Slightly higher cost than 3.5 Flash

Token & API Pricing

Input Tokens (1M)$0.35
Output Tokens (1M)$1.40
Cached Input (1M)$0.0875
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)