Zhipu AI achieves record-breaking 400 tokens/s inference speed

💡Zhipu AI hits 400 tokens/s, setting a new speed record for top-tier LLMs that could disrupt your current API stack.
⚡ 30-Second TL;DR
What Changed
Achieved a peak inference speed of 400 tokens/s
Why It Matters
This speed improvement drastically reduces latency for real-time conversational agents and high-throughput coding assistants. It sets a new competitive benchmark for inference efficiency in the Chinese AI market.
What To Do Next
Benchmark your current latency-sensitive applications against Zhipu's API to see if this speed boost can replace more expensive or slower model providers.
Key Points
- •Achieved a peak inference speed of 400 tokens/s
- •Significant performance leap compared to existing top-tier models
- •Enhanced efficiency for real-time AI application development
🧠 Deep Insight
Web-grounded analysis with 23 cited sources.
🔑 Enhanced Key Takeaways
- •The record-breaking 400 tokens/s inference speed is specifically achieved by Zhipu AI's GLM-5.1 high-speed API, which is currently available to select enterprise clients.
- •This performance represents a 'dual breakthrough in flagship capability and ultra-low latency,' challenging the industry's previous assumption that high speed necessitates a compromise in model quality.
- •The optimization was a collaborative effort between Zhipu AI's GLM and TileRT teams, involving extensive enhancements across the inference engine, scheduling system, and underlying infrastructure.
- •Key technical improvements include rewriting the core inference pathway for increased per-card throughput, implementing dynamic batching and KV cache scheduling to reduce tail latency, and co-optimizing cluster and network performance to sustain the 400 tokens/s output.
- •The enhanced speed makes Zhipu AI's models particularly well-suited for latency-sensitive applications such as AI coding, real-time interactive agents, and real-time voice processing.
📊 Competitor Analysis▸ Show
| Model/Feature | Inference Speed (tokens/s) | Parameters (Total/Active) | Context Window (tokens) | Pricing (Input/Output per 1M tokens) |
|---|---|---|---|---|
| Zhipu AI GLM-5.1 (High-Speed API) | 400 | 754B / 40B (MoE) | 200K | $1.40 / $4.40 (GLM-5.1 API) |
| Mercury 2 | 640 (fastest output) | N/A | N/A | N/A |
| Xiaomi MiMo-V2-Flash | 150 | 309B | 256K | N/A |
| MiniMax M2 | 150 | 230B / 10B (active) | N/A | N/A |
| Zhipu AI GLM-4.7 | 55 | 400B | 200K | $0.60 / $2.20 |
| Zhipu AI GLM-4.5 | 42 | 355B / 32B (MoE) | 128K | $0.60 / $2.20 (GLM-4.5 API) |
| Cerebras (inference provider) | ~2,600 (ultra-fast) | N/A | N/A | Free tier (1M tokens/day cap) |
| GPT-5.4 (OpenAI) | N/A | N/A | N/A | $2.50 / $15 |
| Claude Opus 4.6 (Anthropic) | N/A | N/A | N/A | $15 / $75 |
| Gemini 3.1 Pro (Google) | N/A | N/A | N/A | $2.33 (Input) |
| Kimi K2.6 (Moonshot AI) | N/A | 1T | 262.1K | $0.95 / 1M tokens (cheapest in top 10) |
🛠️ Technical Deep Dive
- Zhipu AI's GLM (General Language Model) family, developed since 2019, forms the foundation of their models.
- GLM-5, a flagship model, employs a decoder-only Transformer architecture with a Mixture-of-Experts (MoE) design, featuring 744 billion total parameters and approximately 40 billion active parameters per token.
- The GLM-5 architecture incorporates DeepSeek Sparse Attention (DSA) for efficient long-context processing, RoPE positional encoding, SwiGLU activations, and post-Layer Normalization.
- GLM-4.5 also utilizes a Mixture-of-Experts (MoE) architecture, optimized for dual modes: a 'thinking' mode for complex reasoning and tool use, and a 'non-thinking' mode for faster responses.
- GLM-4.5's architecture prioritizes depth over width, featuring 96 attention heads per layer, and integrates QK-Norm, Grouped Query Attention, Multi-Token Prediction, and the Muon optimizer for improved performance.
- The 400 tokens/s speed for GLM-5.1 is powered by a high-performance inference engine, a joint development by Zhipu AI and the TileRT team.
- The TileRT engine optimizes GPU operation scheduling by restructuring it into a persistent Engine Kernel that resides on the GPU, which significantly reduces kernel startup and memory read/write delays inherent in traditional inference pipelines.
- For multi-card setups, TileRT specializes GPU nodes within an 8-card NVL topology into distinct functional Workers to enhance attention layer computation and cross-card communication efficiency.
- Zhipu AI plans further optimizations, including FP8 inference and extended context capabilities, to support even lower-latency scenarios.
- The multimodal model GLM-5V-Turbo uses a proprietary vision encoder called CogViT and employs multi-token prediction during inference to accelerate output generation.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (23)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗