🏠較早收集於 3h

智譜致歉 GLM-5 灰度與定價問題

智譜致歉 GLM-5 灰度與定價問題
PostLinkedIn
🏠閱讀原文: IT之家
#token-pricing#grayscale-rollout#user-compensationglm-5

💡Zhipu GLM-5 pricing fixes + refunds: optimize your China LLM costs now

⚡ 30-Second TL;DR

有什麼變化

GLM-5 非高峰 2 倍、高峰 3 倍消耗,因規模達 Claude Opus 級別

為什麼重要

回應中國領先 LLM 定價爭議,穩定用戶信任於競爭中。補償助留存,但凸顯前沿模型擴容挑戰。

下一步行動

Check Zhipu dashboard and apply for GLM-5 Pro/Lite refund if usage surged unexpectedly.

誰應關注:Developers & AI Engineers

關鍵要點

  • GLM-5 非高峰 2 倍、高峰 3 倍消耗,因規模達 Claude Opus 級別
  • 看板刷新從 1 小時優化至 10 分鐘,購買頁全面顯示規則
  • Lite/Pro 用戶自 2026 年 1 月 1 日起全額退款,2/12-16 誤升級一鍵回滾
  • Max 全開,Pro 高峰限流,Lite 節後灰度

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 7 個來源。

🔑 增強重點摘要

  • Zhipu AI apologized on February 21, 2026, for GLM Coding Plan issues including lack of transparency, slow GLM-5 rollout due to traffic surge, and flawed upgrade mechanisms for old users[1][2].
  • GLM-5 rollout is phased: Max tier fully open, Pro tier with peak-hour limits due to high cluster load, Lite tier post-holiday grayscale; refunds offered to affected Lite/Pro users since Jan 1[1][2].
  • GLM-5 is 2x larger than GLM-4.7 with 744B total parameters (40B active) in MoE architecture using DeepSeek Sparse Attention, trained on 28.5T tokens, targeting Claude Opus-level coding and agentic performance[3][5][6].
  • Token costs increased 2-3x due to model scale; dashboard improvements reduced refresh from 1hr to 10min with rules now on purchase page; one-click rollback for Feb 12-16 mis-upgrades[1].
  • Optimized for domestic chips like Huawei Ascend, Moore Threads; compute constraints caused serving delays and pricing hikes amid 10x traffic increase[2][4][6].
📊 競品分析▸ Show
ModelParametersKey BenchmarksPricing Notes
Zhipu GLM-5744B total (40B active, MoE)Leads open models in coding/agentic; surpasses Gemini 3 Pro, lags Claude Opus2-3x GLM-4.7 tokens; 30% coding plan hike [3][5]
DeepSeek (recent)N/ASparse Attention pioneer; 10x context expansionEfficiency-focused [4]
Anthropic Claude OpusProprietaryTop coding benchmarkN/A [3]
Kimi K2.5N/ABelow GLM-5 on GDPVal-AACheap metering [2][5]

🛠️ 技術深入

• GLM-5: 744 billion total parameters, 40 billion active parameters in Mixture-of-Experts (MoE) architecture; doubled from GLM-4.7's 355B[3][5][6]. • Trained on 28.5 trillion tokens; adopts DeepSeek Sparse Attention for computational efficiency[3][4]. • Supports deployment on non-NVIDIA chips: Huawei Ascend, Moore Threads, Cambricon, Kunlunxin, MetaX via kernel optimization and quantization[4][6]. • Serving challenges: MLA models with one KV head cause tensor parallelism KV cache waste; mitigations like SGLang's DP Attention (DPA) for zero KV redundancy and +92% throughput[2]. • Pivot to 'agentic engineering' from 'vibe coding' for scaled AI-automated coding[3].

🔮 前景展望AI analysis grounded in cited sources

Zhipu's GLM-5 launch and apology highlight compute bottlenecks in China's AI race, signaling shift to agentic/coding models amid GPU shortages; pricing hikes buck price wars, while domestic chip optimization reduces NVIDIA reliance, potentially accelerating open-weight SOTA competition with global leaders like Claude Opus[2][3][4][5].

時間線

2025-12
Zhipu releases GLM-4.7, marketed as coding partner with subscription pivot to coding plans
2026-02-12
GLM-5 launched alongside 30% coding plan price hike; rollout begins with upgrade issues
2026-02-21
Zhipu issues public apology for GLM Coding Plan transparency, rollout delays, and upgrades; announces refunds and improvements
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: IT之家

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。