🔥較早收集於 25h

xAI 的 Grok V9-Medium 模型訓練完成

xAI 的 Grok V9-Medium 模型訓練完成
PostLinkedIn
🔥閱讀原文: 36氪

💡xAI 即將推出整合大量編碼數據的 1.5T 基礎模型,預計數週內發布。

⚡ 30-Second TL;DR

有什麼變化

Grok V9-Medium (1.5T) 基礎模型已完成訓練。

為什麼重要

整合 Cursor 數據顯示其對編碼能力的重視,使 Grok 有望成為對開發者更具競爭力的工具。

下一步行動

準備在 Grok V9-Medium 發布後,將其與現有的編碼助手進行基準測試,以評估其在實際開發任務中的表現。

誰應關注:Developers & AI Engineers

關鍵要點

  • Grok V9-Medium (1.5T) 基礎模型已完成訓練。
  • 訓練數據包含了大量來自 Cursor 的編碼數據。
  • 預計在接下來的強化學習階段後,於 2 至 3 週內正式發布。

🧠 深度解析

Web-grounded analysis with 17 cited sources.

🔑 增強重點摘要

  • The Grok V9-Medium model, with 1.5 trillion parameters, is three times larger than its predecessor, the v8-small model (0.5T parameters), which currently handles all Grok production traffic.
  • Elon Musk previously acknowledged that the v8-small model suffered from deficiencies in training data quality, comprehensiveness, and balance, which the V9-Medium aims to rectify.
  • Grok V9-Medium has been specifically optimized for NVIDIA Blackwell architecture GPUs, indicating a focus on leveraging cutting-edge hardware for performance.
  • The substantial incorporation of Cursor-sourced coding data is intended to significantly enhance Grok's capabilities in complex programming tasks, positioning it as an 'AI engineer' within the developer ecosystem.
  • Following the completion of training, the model is currently undergoing supervised fine-tuning, with reinforcement learning scheduled to commence in the coming days before its public release.
📊 競品分析▸ Show
Feature/ModelGrok V9-Medium (Expected)Grok 4.3 (Current)Claude Opus 4.7 (April 2026)GPT-5.5 (April 2026)Gemini 2.5 Pro (2026)
Parameter Count1.5TN/A (Grok 4 family)N/AN/AN/A
Primary FocusComplex Programming, AI EngineerReal-time research, social media, X integrationCoding, long-horizon agent workAgentic workflows, multimodal tasksGoogle Workspace, Android, multimodal
SWE-bench VerifiedExpected significant improvementN/A (Grok 4 at 75%)87.6%74.9% (GPT-5)N/A
Context WindowN/A1M tokens (Grok 4.3)1M tokens1M tokens1M tokens
GPU OptimizationNVIDIA Blackwell architectureN/AN/AN/AN/A
Pricing (API/Subscription)N/A (Grok 4.3 Beta: $300/month SuperGrok Heavy tier)$30/mo SuperGrok (Grok 4.3 Beta)$20/mo ProFree (ads) / $20/mo PlusFree / $20/mo Advanced

🛠️ 技術深入

  • Parameter Scale: Grok V9-Medium is a 1.5 trillion parameter foundation model.
  • GPU Optimization: The model has been specifically optimized for NVIDIA Blackwell architecture GPUs.
  • Training Data Focus: Incorporates a large volume of Cursor-sourced coding data, aiming for enhanced performance in complex programming tasks.
  • Development Stages: Currently in the supervised fine-tuning phase, with reinforcement learning to follow before public release.
  • Predecessor Comparison: It is three times the size of the previous v8-small model (0.5T parameters) that currently serves Grok's production traffic.
  • Training Infrastructure (Historical): Earlier Grok models like Grok-3 were trained on the Colossus supercomputer cluster, which xAI stated utilized over 100,000 NVIDIA H100 GPUs.
  • Architectural Basis (Historical): Grok-1, an earlier model, was a 314-billion-parameter Mixture-of-Experts (MoE) Transformer with 64 layers, 48 attention heads, an embedding dimension of 6,144, and an 8,192-token context window.
  • Software Stack (Historical): xAI's training stack includes a JAX-based modeling and training layer, a Rust control plane for orchestration, and a Kubernetes substrate for scheduling.

🔮 前景展望AI analysis grounded in cited sources

Grok V9-Medium will significantly elevate xAI's standing in the AI coding assistant market.
The explicit focus on integrating extensive Cursor coding data and Elon Musk's statements about 'much better coding' capabilities suggest a direct challenge to established coding AI leaders.
The optimization for NVIDIA Blackwell GPUs indicates xAI's commitment to high-performance, next-generation AI infrastructure.
Early adoption and optimization for advanced GPU architectures like Blackwell suggest xAI is preparing for increasingly demanding AI workloads and larger models.
The planned open-sourcing of the 0.5T parameter model (v8-small) could accelerate broader developer engagement with xAI's ecosystem.
Making a capable, albeit smaller, model openly available can foster community contributions, drive innovation, and increase adoption of xAI's technology.

時間線

2023-11
Grok-0 (33B parameters) launched as xAI's prototype language model.
2024-03
Grok-1 (314B parameters, Mixture-of-Experts architecture) released and open-sourced.
2024-08
Grok-2 and Grok-2 mini introduced, featuring enhanced reasoning and image generation capabilities.
2025-02
Grok 3 released, reportedly trained on the Colossus supercomputer cluster.
2025-07
Grok 4, a further iteration in the model family, was released.
2026-05-25
Grok V9-Medium (1.5T) foundation model training completed, with public release expected in 2-3 weeks.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 36氪