🔥Stalecollected in 25h

xAI's Grok V9-Medium Model Training Completed

xAI's Grok V9-Medium Model Training Completed
PostLinkedIn
🔥Read original on 36氪

💡New 1.5T foundation model from xAI with heavy coding data integration arriving in weeks.

⚡ 30-Second TL;DR

What Changed

Grok V9-Medium (1.5T) foundation model training is complete.

Why It Matters

The integration of Cursor data suggests a strong focus on coding capabilities, positioning Grok as a more competitive tool for developers.

What To Do Next

Prepare to benchmark Grok V9-Medium against current coding assistants once released to evaluate its performance on real-world development tasks.

Who should care:Developers & AI Engineers

Key Points

  • Grok V9-Medium (1.5T) foundation model training is complete.
  • Training data includes a large volume of Cursor-sourced coding data.
  • Official release is scheduled within 2 to 3 weeks following upcoming RL phases.

🧠 Deep Insight

Web-grounded analysis with 17 cited sources.

🔑 Enhanced Key Takeaways

  • The Grok V9-Medium model, with 1.5 trillion parameters, is three times larger than its predecessor, the v8-small model (0.5T parameters), which currently handles all Grok production traffic.
  • Elon Musk previously acknowledged that the v8-small model suffered from deficiencies in training data quality, comprehensiveness, and balance, which the V9-Medium aims to rectify.
  • Grok V9-Medium has been specifically optimized for NVIDIA Blackwell architecture GPUs, indicating a focus on leveraging cutting-edge hardware for performance.
  • The substantial incorporation of Cursor-sourced coding data is intended to significantly enhance Grok's capabilities in complex programming tasks, positioning it as an 'AI engineer' within the developer ecosystem.
  • Following the completion of training, the model is currently undergoing supervised fine-tuning, with reinforcement learning scheduled to commence in the coming days before its public release.
📊 Competitor Analysis▸ Show
Feature/ModelGrok V9-Medium (Expected)Grok 4.3 (Current)Claude Opus 4.7 (April 2026)GPT-5.5 (April 2026)Gemini 2.5 Pro (2026)
Parameter Count1.5TN/A (Grok 4 family)N/AN/AN/A
Primary FocusComplex Programming, AI EngineerReal-time research, social media, X integrationCoding, long-horizon agent workAgentic workflows, multimodal tasksGoogle Workspace, Android, multimodal
SWE-bench VerifiedExpected significant improvementN/A (Grok 4 at 75%)87.6%74.9% (GPT-5)N/A
Context WindowN/A1M tokens (Grok 4.3)1M tokens1M tokens1M tokens
GPU OptimizationNVIDIA Blackwell architectureN/AN/AN/AN/A
Pricing (API/Subscription)N/A (Grok 4.3 Beta: $300/month SuperGrok Heavy tier)$30/mo SuperGrok (Grok 4.3 Beta)$20/mo ProFree (ads) / $20/mo PlusFree / $20/mo Advanced

🛠️ Technical Deep Dive

  • Parameter Scale: Grok V9-Medium is a 1.5 trillion parameter foundation model.
  • GPU Optimization: The model has been specifically optimized for NVIDIA Blackwell architecture GPUs.
  • Training Data Focus: Incorporates a large volume of Cursor-sourced coding data, aiming for enhanced performance in complex programming tasks.
  • Development Stages: Currently in the supervised fine-tuning phase, with reinforcement learning to follow before public release.
  • Predecessor Comparison: It is three times the size of the previous v8-small model (0.5T parameters) that currently serves Grok's production traffic.
  • Training Infrastructure (Historical): Earlier Grok models like Grok-3 were trained on the Colossus supercomputer cluster, which xAI stated utilized over 100,000 NVIDIA H100 GPUs.
  • Architectural Basis (Historical): Grok-1, an earlier model, was a 314-billion-parameter Mixture-of-Experts (MoE) Transformer with 64 layers, 48 attention heads, an embedding dimension of 6,144, and an 8,192-token context window.
  • Software Stack (Historical): xAI's training stack includes a JAX-based modeling and training layer, a Rust control plane for orchestration, and a Kubernetes substrate for scheduling.

🔮 Future ImplicationsAI analysis grounded in cited sources

Grok V9-Medium will significantly elevate xAI's standing in the AI coding assistant market.
The explicit focus on integrating extensive Cursor coding data and Elon Musk's statements about 'much better coding' capabilities suggest a direct challenge to established coding AI leaders.
The optimization for NVIDIA Blackwell GPUs indicates xAI's commitment to high-performance, next-generation AI infrastructure.
Early adoption and optimization for advanced GPU architectures like Blackwell suggest xAI is preparing for increasingly demanding AI workloads and larger models.
The planned open-sourcing of the 0.5T parameter model (v8-small) could accelerate broader developer engagement with xAI's ecosystem.
Making a capable, albeit smaller, model openly available can foster community contributions, drive innovation, and increase adoption of xAI's technology.

Timeline

2023-11
Grok-0 (33B parameters) launched as xAI's prototype language model.
2024-03
Grok-1 (314B parameters, Mixture-of-Experts architecture) released and open-sourced.
2024-08
Grok-2 and Grok-2 mini introduced, featuring enhanced reasoning and image generation capabilities.
2025-02
Grok 3 released, reportedly trained on the Colossus supercomputer cluster.
2025-07
Grok 4, a further iteration in the model family, was released.
2026-05-25
Grok V9-Medium (1.5T) foundation model training completed, with public release expected in 2-3 weeks.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪