The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Gemini 3.1 Flash Audio (Flash Live, TTS)
Dedicated conversational voice model providing native bidirectional speech streaming with 220ms end-to-end latency.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini 3.1 Flash Audio eliminates traditional text-to-speech transcoding pipelines. By processing and synthesizing acoustic tokens natively, it conveys subtle emotional inflection, natural pauses, and human-like interruption handling.
Strengths & Considerations
- 220ms glass-to-glass audio latency
- Native emotional prosody and accent synthesis
- Seamless interruption handling
- Context limited to 128k audio tokens
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.