Introducing Gemini 3.7 FlashGoogleReleased August 13, 2026

Gemini 3.7 Flash

High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.

multimodal-realtimeProprietary API$0.50 / 1M tok (Agentic Video)Context: 1.049M (1,048,576 tokens)Arena ELO: 1445

Technical Specifications

Architecture Type
Algorithmic Reasoning Multimodal Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
65.536K (65,536 tokens)
Knowledge Cutoff
June 2026
Supported Modalities
text, code, vision, audio, video
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Mmlu Pro
82.5
Humaneval
95.1
Video Qa Long
88.9
Swe Bench Verified
65.3
Chatbot Arena ELO
1445

Deep Architectural Overview

Gemini 3.7 Flash delivers breakthrough performance across agentic video comprehension, allowing developers to query hours of surveillance, lecture, or film footage with precise timestamp-level accuracy.

Strengths & Considerations

Core Strengths
  • State-of-the-art agentic video comprehension
  • 65.3% on SWE-bench Verified in Flash tier
  • Sub-second tool invocation speed
Known Limitations
  • Video processing at full frame rate requires caching for low cost

Token & API Pricing

Input Tokens (1M)$0.50
Output Tokens (1M)$2.00
Cached Input (1M)$0.1250
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)