Introducing GLM-5.3-FlashZ.AIReleased August 26, 2026

GLM-5.3-Flash

The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.

multimodal-realtimeProprietary API$0.15 / 1M tok (320B MoE Visual Coding)Context: 200K (200,000 tokens)Arena ELO: 1470

Technical Specifications

Architecture Type
Hybrid Sparse + Linear Attention MoE
Total Parameters
320B
Active Parameters (MoE)
18B
Context Window
200K (200,000 tokens)
Max Output Tokens
64K (64,000 tokens)
Knowledge Cutoff
July 2026
Supported Modalities
text, code, vision, reasoning
License & Access
Z.AI Terms of Service

Benchmark Evaluations

Mmbench
88.6
Mmlu Pro
86.4
Humaneval
94.8
Swe Bench Verified
75.8
Chatbot Arena ELO
1470

Deep Architectural Overview

GLM-5.3-Flash weaves visual perception directly into the software development loop. It observes rendered web interfaces, browser developer tools, and interaction feedback to autonomously test and refine full-stack apps at just $0.15/$0.50 per million tokens.

Strengths & Considerations

Core Strengths
  • Hybrid linear + sparse attention cuts KV cache by 4.44x
  • Native visual coding: observes UI renders and browser feedback
  • Extremely economical at $0.15 input / $0.03 cache reads
Known Limitations
  • 200k context window compared to 1M context on GLM-5.3

Token & API Pricing

Input Tokens (1M)$0.15
Output Tokens (1M)$0.50
Cached Input (1M)$0.0300
Pricing is verified directly against Z.AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Z.AI and comparable reasoning engines.

Browse all models
Z.AI

Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.

200K ctx$0.37 / 1M tok (Ultra-Low Latency MoE)
Z.AI

Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.

1M ctx$1.40 / 1M tok (Flagship Coding & Cyber)
Z.AI

Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.

1M ctx$1.40 / 1M tok (1M Lossless Context)
Z.AI

Engineered for long-horizon tasks, able to work independently for up to 8 hours in a single run, aligned with Claude Opus 4.6.

200K ctx$1.40 / 1M tok (8h Autonomous Agent)