Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.
GLM-ASR-2512
High-accuracy automatic speech recognition model delivering a Character Error Rate as low as 0.0717 with user-defined vocabulary support.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
GLM-ASR-2512 is Z.AI’s production speech-to-text model. Engineered for complex acoustic environments, regional accents, and specialized domain terminology, it bills at roughly $0.0024 per minute of audio.
Strengths & Considerations
- Ultra-low CER of 0.0717
- User-defined custom vocabularies and terminology dictionaries
- Extremely low cost (~$0.0024/minute)
- Audio input only; text output transcription only
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Z.AI and comparable reasoning engines.
The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.
Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.
Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.