Zhipu MaaS Revenue Surges 400%

💡Zhipu's API usage and paid users are surging, offering a signal for enterprise AI demand and pricing power.
⚡ 30-Second TL;DR
What Changed
Open platform and API revenue reached RMB 825 million, accounting for 86.5% of total revenue.
Why It Matters
The results suggest that demand for enterprise AI agents and model APIs is rapidly shifting toward usage-based MaaS platforms. However, the large R&D investment and adjusted loss indicate that scaling model capabilities remains capital intensive.
What To Do Next
Benchmark Zhipu's MaaS API against your current model provider using your highest-volume production prompts, comparing latency, token cost, and output quality.
Key Points
- •Open platform and API revenue reached RMB 825 million, accounting for 86.5% of total revenue.
- •Token usage increased more than 40 times from the start of the year, while paid daily active users grew 603%.
- •Zhipu reported average API prices up approximately 101% and top-ten customer daily usage up 98%.
- •The company invested RMB 2.13 billion in R&D and recorded an adjusted net loss of RMB 1.964 billion.
🧠 Deep Insight
Background and context from public sources — not the original article. 16 sources cited.
🔑 Enhanced Key Takeaways
- •Zhipu's open platform and API business gross margin turned positive, rising from -0.4% in H1 2025 to 24.6% in H1 2026.
- •The company successfully scaled inference operations across a cluster of over 100,000 domestic AI chips, driving an 80% reduction in unit token costs.
- •Zhipu's revenue mix has undergone a radical shift, with API services growing from 15.2% of total revenue in H1 2025 to 86.5% in H1 2026.
- •The company was identified as the developer behind the 'Ox Alpha' anonymous model, which was later rebranded as the GLM-5.3-Flash model.
- •Zhipu AI completed an IPO on the Hong Kong Stock Exchange (HKEX: 2513), becoming one of the first major Chinese LLM developers to go public.
📊 Competitor Analysis▸ Show
| Feature | Zhipu AI | Alibaba (Qwen) | DeepSeek | MiniMax |
|---|---|---|---|---|
| Primary Model | GLM-5.3-Flash | Qwen-2.5 | DeepSeek-V3 | abab 7.0 |
| Market Focus | Enterprise MaaS | Cloud Ecosystem | Open Source/API | Consumer/Agent |
| Hardware Strategy | Domestic Chip Cluster | Proprietary/NVIDIA | Optimized Inference | Hybrid Cloud |
🛠️ Technical Deep Dive
- Model Architecture: Utilizes the GLM (General Language Model) framework, specifically the GLM-5.3-Flash iteration optimized for low-latency inference.
- Infrastructure: Implements large-scale distributed inference across a heterogeneous cluster of over 100,000 domestic AI accelerators.
- Cost Optimization: Achieved 80% reduction in unit token inference costs through custom kernel optimization and memory management on domestic hardware.
- API Integration: Supports high-concurrency RESTful and gRPC interfaces tailored for enterprise-grade RAG (Retrieval-Augmented Generation) pipelines.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (16)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
