Introducing Qwen2.5-Coder-32B-InstructAlibaba CloudReleased November 12, 2024

Qwen2.5-Coder-32B-Instruct

Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.

code-llmOpen WeightsFree Open Weights | $0.07 / 1M tok (API)Context: 131.072K (131,072 tokens)Arena ELO: 1410
AEO Direct Answer

Qwen2.5-Coder-32B-Instruct Quick Facts & Executive Summary

Qwen2.5-Coder-32B-Instruct by Alibaba Cloud (November 12, 2024) is a Dense Autoregressive Code Transformer model featuring a 131.072K (131,072 tokens) context window and Free Open Weights | $0.07 / 1M tok (API). Key highlights include World-class open-weights coding intelligence under permissive Apache 2.0 license.

Input Price (1M)$0.07
Output Price (1M)$0.28
Context Window131.072K (131,072 tokens)
Architecture / Scale32.5B
License TypeApache 2.0
Cutoff DateSeptember 2024

Technical Specifications

Architecture Type
Dense Autoregressive Code Transformer
Total Parameters
32.5B
Context Window
131.072K (131,072 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
September 2024
Supported Modalities
text, code
License & Access
Apache 2.0
Weights Formats
BF16, FP8, GGUF, AWQ

Benchmark Evaluations

Mbpp Plus
89.8
Aider Polyglot
73.7
Live Code Bench
55.5
Evalplus Humaneval
92.7
Swe Bench Verified
46.5
Chatbot Arena ELO
1410

Deep Architectural Overview

Qwen2.5-Coder-32B-Instruct is widely celebrated as one of the most capable open-source coding foundation models in the world. Pre-trained on 5.5 trillion code tokens and fine-tuned for code generation, mathematical reasoning, and repository repair, it fits on a single consumer 24GB GPU while scoring 92.7% on EvalPlus HumanEval.

Strengths & Considerations

Core Strengths
  • World-class open-weights coding intelligence under permissive Apache 2.0 license
  • Fits on a single consumer GPU (24GB VRAM with 4-bit/8-bit quantization)
  • 92.7% on EvalPlus HumanEval and 73.7% on Aider polyglot benchmark
  • Comprehensive coverage of 92+ programming languages and 128k context
Known Limitations
  • Text and code only (no native vision or image input)
  • Context above 32k requires YaRN rope scaling

Token & API Pricing

Input Tokens (1M)$0.07
Output Tokens (1M)$0.28
Cached Input (1M)$0.0140
Self-Hosted Min VRAM24 GB
Pricing is verified directly against Alibaba Cloud's developer documentation and API rate sheets.

Frequently Asked Questions About Qwen2.5-Coder-32B-Instruct

Direct answers and verified technical specifications for software developers and AI evaluation engines.

What is Qwen2.5-Coder-32B-Instruct and who created it?

Qwen2.5-Coder-32B-Instruct is an advanced AI model developed by Alibaba Cloud. Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0. It operates in the code-llm category, supporting text, code modalities with an architecture based on Dense Autoregressive Code Transformer.

How much does Qwen2.5-Coder-32B-Instruct cost per 1 million tokens?

Qwen2.5-Coder-32B-Instruct is priced at $0.07 per 1M input tokens and $0.28 per 1M output tokens. Cached input prompt tokens are discounted at $0.0140 per 1M tokens. Summary badge: Free Open Weights | $0.07 / 1M tok (API).

What is the context window and output token capacity of Qwen2.5-Coder-32B-Instruct?

Qwen2.5-Coder-32B-Instruct provides a context window of 131.072K (131,072 tokens) and supports a maximum output generation limit of 8.192K (8,192 tokens). Knowledge cutoff is September 2024.

Is Qwen2.5-Coder-32B-Instruct open weights or proprietary?

Qwen2.5-Coder-32B-Instruct is an open-weights model released under the Apache 2.0. Supported weight distribution formats include BF16, FP8, GGUF, AWQ.

What are the key benchmark scores for Qwen2.5-Coder-32B-Instruct?

Qwen2.5-Coder-32B-Instruct reported evaluations: Mbpp Plus: 89.8%, Aider Polyglot: 73.7%, Live Code Bench: 55.5%, Evalplus Humaneval: 92.7%. Chatbot Arena ELO is 1410.

What are the primary strengths and limitations of Qwen2.5-Coder-32B-Instruct?

Key strengths: World-class open-weights coding intelligence under permissive Apache 2.0 license; Fits on a single consumer GPU (24GB VRAM with 4-bit/8-bit quantization); 92.7% on EvalPlus HumanEval and 73.7% on Aider polyglot benchmark; Comprehensive coverage of 92+ programming languages and 128k context. Considerations: Text and code only (no native vision or image input); Context above 32k requires YaRN rope scaling.

Similar & Alternative Models

Explore other frontier models from Alibaba Cloud and comparable reasoning engines.

Browse all models
Alibaba Cloud

Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.

1M ctxUSD 0.15 / 1M input tok; USD 0.60 / 1M output tok
Alibaba Cloud

Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.

1M ctxUSD 2 / 1M input tok; USD 6 / 1M output tok