Introducing Gemini 3.1 Flash-LiteGoogleReleased March 3, 2026

Gemini 3.1 Flash-Lite

The Gemini 3.1 family’s budget workhorse, delivering solid reasoning and vision parsing at 6 cents per million tokens.

text-to-textProprietary API$0.06 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1290

Technical Specifications

Architecture Type
Efficient Dense Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
December 2025
Supported Modalities
text, code, vision
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Mmlu
83.2
Math500
74.8
Humaneval
86.4
Chatbot Arena ELO
1290

Deep Architectural Overview

Gemini 3.1 Flash-Lite brings third-generation architectural optimizations to budget-conscious developers, providing fast classification, structured JSON extraction, and high-volume document ingestion.

Strengths & Considerations

Core Strengths
  • Cost-effective at $0.06/1M tokens
  • 1M token context retention
  • Clean structured JSON output compliance
Known Limitations
  • No audio streaming support

Token & API Pricing

Input Tokens (1M)$0.06
Output Tokens (1M)$0.24
Cached Input (1M)$0.0150
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)