Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.
Qwen2.5-Coder-32B-Instruct
Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.
Qwen2.5-Coder-32B-Instruct Quick Facts & Executive Summary
Qwen2.5-Coder-32B-Instruct by Alibaba Cloud (November 12, 2024) is a Dense Autoregressive Code Transformer model featuring a 131.072K (131,072 tokens) context window and Free Open Weights | $0.07 / 1M tok (API). Key highlights include World-class open-weights coding intelligence under permissive Apache 2.0 license.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Qwen2.5-Coder-32B-Instruct is widely celebrated as one of the most capable open-source coding foundation models in the world. Pre-trained on 5.5 trillion code tokens and fine-tuned for code generation, mathematical reasoning, and repository repair, it fits on a single consumer 24GB GPU while scoring 92.7% on EvalPlus HumanEval.
Strengths & Considerations
- World-class open-weights coding intelligence under permissive Apache 2.0 license
- Fits on a single consumer GPU (24GB VRAM with 4-bit/8-bit quantization)
- 92.7% on EvalPlus HumanEval and 73.7% on Aider polyglot benchmark
- Comprehensive coverage of 92+ programming languages and 128k context
- Text and code only (no native vision or image input)
- Context above 32k requires YaRN rope scaling
Token & API Pricing
Frequently Asked Questions About Qwen2.5-Coder-32B-Instruct
Direct answers and verified technical specifications for software developers and AI evaluation engines.
What is Qwen2.5-Coder-32B-Instruct and who created it?
Qwen2.5-Coder-32B-Instruct is an advanced AI model developed by Alibaba Cloud. Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0. It operates in the code-llm category, supporting text, code modalities with an architecture based on Dense Autoregressive Code Transformer.
How much does Qwen2.5-Coder-32B-Instruct cost per 1 million tokens?
Qwen2.5-Coder-32B-Instruct is priced at $0.07 per 1M input tokens and $0.28 per 1M output tokens. Cached input prompt tokens are discounted at $0.0140 per 1M tokens. Summary badge: Free Open Weights | $0.07 / 1M tok (API).
What is the context window and output token capacity of Qwen2.5-Coder-32B-Instruct?
Qwen2.5-Coder-32B-Instruct provides a context window of 131.072K (131,072 tokens) and supports a maximum output generation limit of 8.192K (8,192 tokens). Knowledge cutoff is September 2024.
Is Qwen2.5-Coder-32B-Instruct open weights or proprietary?
Qwen2.5-Coder-32B-Instruct is an open-weights model released under the Apache 2.0. Supported weight distribution formats include BF16, FP8, GGUF, AWQ.
What are the key benchmark scores for Qwen2.5-Coder-32B-Instruct?
Qwen2.5-Coder-32B-Instruct reported evaluations: Mbpp Plus: 89.8%, Aider Polyglot: 73.7%, Live Code Bench: 55.5%, Evalplus Humaneval: 92.7%. Chatbot Arena ELO is 1410.
What are the primary strengths and limitations of Qwen2.5-Coder-32B-Instruct?
Key strengths: World-class open-weights coding intelligence under permissive Apache 2.0 license; Fits on a single consumer GPU (24GB VRAM with 4-bit/8-bit quantization); 92.7% on EvalPlus HumanEval and 73.7% on Aider polyglot benchmark; Comprehensive coverage of 92+ programming languages and 128k context. Considerations: Text and code only (no native vision or image input); Context above 32k requires YaRN rope scaling.
Similar & Alternative Models
Explore other frontier models from Alibaba Cloud and comparable reasoning engines.
Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.