The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Gemini Omni Flash
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini Omni Flash breaks down modal barriers entirely. It can take a voice command and video clip, edit the video, generate sound effects, narrate an explanation, and produce code in a single coherent inference stream.
Strengths & Considerations
- Universal any-to-any multimodal translation
- Synchronized video and audio generation
- Single unified API endpoint for all media modalities
- Video generation token consumption is higher than text queries
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.
High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.