智譜致歉 GLM-5 灰度與定價問題

💡Zhipu GLM-5 pricing fixes + refunds: optimize your China LLM costs now
⚡ 30-Second TL;DR
有什麼變化
GLM-5 非高峰 2 倍、高峰 3 倍消耗,因規模達 Claude Opus 級別
為什麼重要
回應中國領先 LLM 定價爭議,穩定用戶信任於競爭中。補償助留存,但凸顯前沿模型擴容挑戰。
下一步行動
Check Zhipu dashboard and apply for GLM-5 Pro/Lite refund if usage surged unexpectedly.
關鍵要點
- •GLM-5 非高峰 2 倍、高峰 3 倍消耗,因規模達 Claude Opus 級別
- •看板刷新從 1 小時優化至 10 分鐘,購買頁全面顯示規則
- •Lite/Pro 用戶自 2026 年 1 月 1 日起全額退款,2/12-16 誤升級一鍵回滾
- •Max 全開,Pro 高峰限流,Lite 節後灰度
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 7 個來源。
🔑 增強重點摘要
- •Zhipu AI apologized on February 21, 2026, for GLM Coding Plan issues including lack of transparency, slow GLM-5 rollout due to traffic surge, and flawed upgrade mechanisms for old users[1][2].
- •GLM-5 rollout is phased: Max tier fully open, Pro tier with peak-hour limits due to high cluster load, Lite tier post-holiday grayscale; refunds offered to affected Lite/Pro users since Jan 1[1][2].
- •GLM-5 is 2x larger than GLM-4.7 with 744B total parameters (40B active) in MoE architecture using DeepSeek Sparse Attention, trained on 28.5T tokens, targeting Claude Opus-level coding and agentic performance[3][5][6].
- •Token costs increased 2-3x due to model scale; dashboard improvements reduced refresh from 1hr to 10min with rules now on purchase page; one-click rollback for Feb 12-16 mis-upgrades[1].
- •Optimized for domestic chips like Huawei Ascend, Moore Threads; compute constraints caused serving delays and pricing hikes amid 10x traffic increase[2][4][6].
📊 競品分析▸ Show
| Model | Parameters | Key Benchmarks | Pricing Notes |
|---|---|---|---|
| Zhipu GLM-5 | 744B total (40B active, MoE) | Leads open models in coding/agentic; surpasses Gemini 3 Pro, lags Claude Opus | 2-3x GLM-4.7 tokens; 30% coding plan hike [3][5] |
| DeepSeek (recent) | N/A | Sparse Attention pioneer; 10x context expansion | Efficiency-focused [4] |
| Anthropic Claude Opus | Proprietary | Top coding benchmark | N/A [3] |
| Kimi K2.5 | N/A | Below GLM-5 on GDPVal-AA | Cheap metering [2][5] |
🛠️ 技術深入
• GLM-5: 744 billion total parameters, 40 billion active parameters in Mixture-of-Experts (MoE) architecture; doubled from GLM-4.7's 355B[3][5][6]. • Trained on 28.5 trillion tokens; adopts DeepSeek Sparse Attention for computational efficiency[3][4]. • Supports deployment on non-NVIDIA chips: Huawei Ascend, Moore Threads, Cambricon, Kunlunxin, MetaX via kernel optimization and quantization[4][6]. • Serving challenges: MLA models with one KV head cause tensor parallelism KV cache waste; mitigations like SGLang's DP Attention (DPA) for zero KV redundancy and +92% throughput[2]. • Pivot to 'agentic engineering' from 'vibe coding' for scaled AI-automated coding[3].
🔮 前景展望AI analysis grounded in cited sources
Zhipu's GLM-5 launch and apology highlight compute bottlenecks in China's AI race, signaling shift to agentic/coding models amid GPU shortages; pricing hikes buck price wars, while domestic chip optimization reduces NVIDIA reliance, potentially accelerating open-weight SOTA competition with global leaders like Claude Opus[2][3][4][5].
⏳ 時間線
📎 來源 (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- kucoin.com — Zhipu AI Apologizes for Glm Coding Plan Issues and Announces Compensation
- latent.space — Ainews Zai Glm 5 New Sota Open Weights
- scmp.com — Chinas Zhipu AI Launches New Major Model Glm 5 Challenge Its Rivals
- trendforce.com — News Deepseek Expands Context Tenfold As Zhipu Rolls Out New Model in Chinas AI Race
- chinatalk.media — Chinese AI Rings in the Year of the
- jessleao.substack.com — Something Big Is Definitely Happening
- businesstimes.com.sg — Chinas Zhipu Unveils New AI Model Jolting Race Deepseek
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: IT之家 ↗
每週 AI 簡報
每週一封,可隨時退訂。
