Introducing Gemini 3.8 Audio (Live, Live Extended Thinking)GoogleReleased September 15, 2026

Gemini 3.8 Audio (Live, Live Extended Thinking)

The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.

multimodal-realtimeProprietary API$0.0035 / min (Live Extended Thinking)Context: 262.144K (262,144 tokens)

Technical Specifications

Architecture Type
Live Audio Reasoning Transformer
Total Parameters
Undisclosed
Context Window
262.144K (262,144 tokens)
Max Output Tokens
16.384K (16,384 tokens)
Knowledge Cutoff
July 2026
Supported Modalities
audio, text, speech, reasoning
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Audio Math Solve
84.6
Audio Reasoning Eval
91.8
Conversational Latency Ms
195

Deep Architectural Overview

Gemini 3.8 Audio (Live, Live Extended Thinking) allows users to speak directly with an AI that reasons deeply before answering verbally. It combines sub-200ms conversational responsiveness with high-level cognitive problem solving.

Strengths & Considerations

Core Strengths
  • Live conversational audio with extended reasoning compute
  • 195ms conversational responsiveness
  • Solves spoken complex math and engineering problems without text intermediation
Known Limitations
  • Extended thinking on difficult queries introduces purposeful thinking pauses

Token & API Pricing

Input Tokens (1M)$0.40
Output Tokens (1M)$1.60
Cached Input (1M)$0.1000
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)
Google

High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.

1.049M ctx$0.50 / 1M tok (Agentic Video)