Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.
Qwen3.8-Omni-Flash
Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.
Qwen3.8-Omni-Flash Quick Facts & Executive Summary
Qwen3.8-Omni-Flash by Alibaba Cloud (September 18, 2026) is a Native Omnimodal End-to-End Foundation Transformer model featuring a 1M (1,000,000 tokens) context window and USD 0.15 / 1M input tok; USD 0.60 / 1M output tok. Key highlights include Native omnimodal processing across text, images, video, and audio streams.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Qwen3.8-Omni-Flash processes text, vision, audio, and video end-to-end within a single unified model. Featuring a 1-million-token context window and 98% reduced audio input pricing, it enables real-time conversational agents, GUI automation, and multimodal tool use at low latency.
Strengths & Considerations
- Native omnimodal processing across text, images, video, and audio streams
- Sub-200ms real-time conversational audio latency (180ms)
- 1,000,000 token context window for lengthy video and document inputs
- 98% lower audio stream pricing ($0.003/min) with ultra-low $0.15/1M token base
- High-speed video streaming requires stable network bandwidth
- Cloud API only (not open weights)
Token & API Pricing
Frequently Asked Questions About Qwen3.8-Omni-Flash
Direct answers and verified technical specifications for software developers and AI evaluation engines.
What is Qwen3.8-Omni-Flash and who created it?
Qwen3.8-Omni-Flash is an advanced AI model developed by Alibaba Cloud. Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction. It operates in the omnimodal-realtime category, supporting text, code, vision, audio, video modalities with an architecture based on Native Omnimodal End-to-End Foundation Transformer.
How much does Qwen3.8-Omni-Flash cost per 1 million tokens?
Qwen3.8-Omni-Flash is priced at $0.15 per 1M input tokens and $0.60 per 1M output tokens. Cached input prompt tokens are discounted at $0.0300 per 1M tokens. Summary badge: USD 0.15 / 1M input tok; USD 0.60 / 1M output tok.
What is the context window and output token capacity of Qwen3.8-Omni-Flash?
Qwen3.8-Omni-Flash provides a context window of 1M (1,000,000 tokens) and supports a maximum output generation limit of 65.536K (65,536 tokens). Knowledge cutoff is August 2026.
Is Qwen3.8-Omni-Flash open weights or proprietary?
Qwen3.8-Omni-Flash is a proprietary closed-weight model accessible via official cloud APIs under Alibaba Cloud API Terms of Service.
What are the key benchmark scores for Qwen3.8-Omni-Flash?
Qwen3.8-Omni-Flash reported evaluations: Video Mme: 84.2%, Mmmu Reasoning: 78.4%, Gui Agent Osworld: 74.5%, Voice Interaction Latency Ms: 180. Chatbot Arena ELO is 1485.
What are the primary strengths and limitations of Qwen3.8-Omni-Flash?
Key strengths: Native omnimodal processing across text, images, video, and audio streams; Sub-200ms real-time conversational audio latency (180ms); 1,000,000 token context window for lengthy video and document inputs; 98% lower audio stream pricing ($0.003/min) with ultra-low $0.15/1M token base. Considerations: High-speed video streaming requires stable network bandwidth; Cloud API only (not open weights).
Similar & Alternative Models
Explore other frontier models from Alibaba Cloud and comparable reasoning engines.
Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.