Introducing Kimi K3Moonshot AIReleased July 20, 2026

Kimi K3

Moonshot AI’s premier flagship model with 2.8 trillion parameters, Kimi Delta Attention, 1M context, and open weights.

reasoning-llmOpen Weights$3.00 / 1M tok (2.8T Open Weights)Context: 1.049M (1,048,576 tokens)Arena ELO: 1550

Technical Specifications

Architecture Type
Kimi Delta Attention (KDA) Hybrid Linear Transformer
Total Parameters
2.8T
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
128K (128,000 tokens)
Knowledge Cutoff
June 2026
Supported Modalities
text, code, vision, reasoning
License & Access
Open License
Weights Formats
safetensors, fp8

Benchmark Evaluations

Math500
99.2
Mmlu Pro
92.8
Gpqa Diamond
87.2
Swe Bench Verified
84.8
Chatbot Arena ELO
1550

Deep Architectural Overview

Kimi K3 is the world’s first open-source model in the 3-trillion parameter class. Powered by Kimi Delta Attention (KDA)—a hybrid linear attention mechanism—and Attention Residuals, it combines native visual understanding, 1M context retention, and 128k output tokens to set new frontier benchmarks across software engineering, mathematical proof, and deep autonomous knowledge work.

Strengths & Considerations

Core Strengths
  • World’s first 2.8 Trillion parameter open model
  • Kimi Delta Attention (KDA) cuts compute while keeping 1M lossless context
  • 84.8% SWE-bench Verified coding capability
  • 128,000 token output buffer with $0.30 cache hits
Known Limitations
  • Massive parameter footprint requires high compute clustering for self-hosting

Token & API Pricing

Input Tokens (1M)$3.00
Output Tokens (1M)$15.00
Cached Input (1M)$0.3000
Self-Hosted Min VRAMVaries by quantization
Pricing is verified directly against Moonshot AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Moonshot AI and comparable reasoning engines.

Browse all models
Moonshot AI

High-speed variant of Kimi K2.7 Code with generation speeds of 180 to 260 tokens per second for instantaneous coding feedback.

262.144K ctx$1.90 / 1M tok (180-260 tok/s)
Moonshot AI

Dedicated coding model delivering higher task success rates, tighter instruction following, and a 30% reduction in overthinking.

262.144K ctx$0.95 / 1M tok (Dedicated Coding)
Moonshot AI

General-purpose model supporting text, image, and video inputs, switchable thinking modes, and autonomous agent tasks over 256k context.

262.144K ctx$0.95 / 1M tok (Vision & Agentic)