xAI's Grok V9-Medium Model Training Completed

💡New 1.5T foundation model from xAI with heavy coding data integration arriving in weeks.
⚡ 30-Second TL;DR
What Changed
Grok V9-Medium (1.5T) foundation model training is complete.
Why It Matters
The integration of Cursor data suggests a strong focus on coding capabilities, positioning Grok as a more competitive tool for developers.
What To Do Next
Prepare to benchmark Grok V9-Medium against current coding assistants once released to evaluate its performance on real-world development tasks.
Key Points
- •Grok V9-Medium (1.5T) foundation model training is complete.
- •Training data includes a large volume of Cursor-sourced coding data.
- •Official release is scheduled within 2 to 3 weeks following upcoming RL phases.
🧠 Deep Insight
Web-grounded analysis with 17 cited sources.
🔑 Enhanced Key Takeaways
- •The Grok V9-Medium model, with 1.5 trillion parameters, is three times larger than its predecessor, the v8-small model (0.5T parameters), which currently handles all Grok production traffic.
- •Elon Musk previously acknowledged that the v8-small model suffered from deficiencies in training data quality, comprehensiveness, and balance, which the V9-Medium aims to rectify.
- •Grok V9-Medium has been specifically optimized for NVIDIA Blackwell architecture GPUs, indicating a focus on leveraging cutting-edge hardware for performance.
- •The substantial incorporation of Cursor-sourced coding data is intended to significantly enhance Grok's capabilities in complex programming tasks, positioning it as an 'AI engineer' within the developer ecosystem.
- •Following the completion of training, the model is currently undergoing supervised fine-tuning, with reinforcement learning scheduled to commence in the coming days before its public release.
📊 Competitor Analysis▸ Show
| Feature/Model | Grok V9-Medium (Expected) | Grok 4.3 (Current) | Claude Opus 4.7 (April 2026) | GPT-5.5 (April 2026) | Gemini 2.5 Pro (2026) |
|---|---|---|---|---|---|
| Parameter Count | 1.5T | N/A (Grok 4 family) | N/A | N/A | N/A |
| Primary Focus | Complex Programming, AI Engineer | Real-time research, social media, X integration | Coding, long-horizon agent work | Agentic workflows, multimodal tasks | Google Workspace, Android, multimodal |
| SWE-bench Verified | Expected significant improvement | N/A (Grok 4 at 75%) | 87.6% | 74.9% (GPT-5) | N/A |
| Context Window | N/A | 1M tokens (Grok 4.3) | 1M tokens | 1M tokens | 1M tokens |
| GPU Optimization | NVIDIA Blackwell architecture | N/A | N/A | N/A | N/A |
| Pricing (API/Subscription) | N/A (Grok 4.3 Beta: $300/month SuperGrok Heavy tier) | $30/mo SuperGrok (Grok 4.3 Beta) | $20/mo Pro | Free (ads) / $20/mo Plus | Free / $20/mo Advanced |
🛠️ Technical Deep Dive
- Parameter Scale: Grok V9-Medium is a 1.5 trillion parameter foundation model.
- GPU Optimization: The model has been specifically optimized for NVIDIA Blackwell architecture GPUs.
- Training Data Focus: Incorporates a large volume of Cursor-sourced coding data, aiming for enhanced performance in complex programming tasks.
- Development Stages: Currently in the supervised fine-tuning phase, with reinforcement learning to follow before public release.
- Predecessor Comparison: It is three times the size of the previous v8-small model (0.5T parameters) that currently serves Grok's production traffic.
- Training Infrastructure (Historical): Earlier Grok models like Grok-3 were trained on the Colossus supercomputer cluster, which xAI stated utilized over 100,000 NVIDIA H100 GPUs.
- Architectural Basis (Historical): Grok-1, an earlier model, was a 314-billion-parameter Mixture-of-Experts (MoE) Transformer with 64 layers, 48 attention heads, an embedding dimension of 6,144, and an 8,192-token context window.
- Software Stack (Historical): xAI's training stack includes a JAX-based modeling and training layer, a Rust control plane for orchestration, and a Kubernetes substrate for scheduling.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (17)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗