xAI Completes Training of Grok V9-Medium Model
💡New 1.5T parameter model from xAI with heavy focus on coding data; release expected in weeks.
⚡ 30-Second TL;DR
What Changed
Grok V9-Medium (1.5T) base model training is complete.
Why It Matters
The integration of Cursor data suggests a strong focus on coding capabilities, potentially positioning Grok as a top-tier assistant for software development.
What To Do Next
Monitor xAI's official release notes to evaluate the model's coding performance against current SOTA models like Claude 3.5 Sonnet.
Key Points
- •Grok V9-Medium (1.5T) base model training is complete.
- •Training data includes significant amounts of Cursor-related data.
- •Fine-tuning is underway with reinforcement learning starting in days.
- •Official release is scheduled for 2-3 weeks from now.
🧠 Deep Insight
Web-grounded analysis with 27 cited sources.
🔑 Enhanced Key Takeaways
- •The Grok V9-Medium model, with 1.5 trillion parameters, represents a threefold increase in scale compared to its predecessor, the Grok V8 model (0.5 trillion parameters), which currently handles all Grok production traffic.
- •The substantial incorporation of Cursor-related data during supplementary training indicates a strategic focus by xAI on enhancing Grok's capabilities for developer and complex coding use cases.
- •Grok V9-Medium has been specifically optimized for the Blackwell architecture GPUs, suggesting a forward-looking hardware strategy and potential performance advantages on next-generation NVIDIA infrastructure.
📊 Competitor Analysis▸ Show
Grok AI vs. Key Competitors (as of May 2026)
| Feature/Metric | xAI Grok (V4.x / V9-Medium) | OpenAI GPT-4o / GPT-5 | Anthropic Claude 3.7 Sonnet / Opus 4.7 | Google Gemini 2.5 Pro |
|---|---|---|---|---|
| Model Parameters | V9-Medium: 1.5 Trillion (training complete). Grok-1: 314B MoE. | GPT-4: Not fully disclosed. GPT-5: Larger, specific details not public. | Not fully disclosed. | Not fully disclosed. |
| Context Window | Grok 4: 256K tokens. Grok 4.1 Fast: Up to 2M tokens. | GPT-4: Up to 128K tokens. GPT-5: Extended thinking trades latency for higher accuracy. | Claude Opus 4.7: 1M tokens. | Gemini 2.5 Pro: 1M tokens. |
| Real-time Data | Real-time access to X (Twitter) firehose and live web search. | GPT-4o: Bing integration. GPT-5: Likely similar. | Limited direct real-time access. | Integrates real-time data from search engines. |
| Content Guardrails | Looser content moderation, answers questions other models often refuse. | More restrictive, emphasizes safety and helpfulness. | Emphasizes safety, helpfulness, and ethical AI development. | Emphasizes safety and responsible AI. |
| Coding Benchmarks (SWE-Bench Verified) | Grok 4: ~72-75%. Grok Build 0.1: 70.8%. | GPT-4: ~65-70%. GPT-5.5: 88.7%. | Claude Opus 4.7: 87.6%. | Competitive, specific scores vary. |
| Reasoning Benchmarks (ARC-AGI v2) | Grok 4: 15.9% (nearly double Claude 4's 8.6%). | Competitive, specific scores vary. | Claude 4: 8.6%. | Competitive, specific scores vary. |
| Multimodal Capabilities | Grok 2: Image generation. Grok 3: Web browsing, image/file analysis. Grok 4: Text, image, video capable, live browsing. Grok Imagine (text-to-image/video), Grok Voice. | GPT-4o: Multimodal. GPT-5: Advanced multimodal. | Claude 3.7 Sonnet / Opus 4.7: Competitive on image inputs. | Native multimodal handling (text, image, audio, video, code). |
| API Pricing (per 1M tokens) | Grok 4: $3 input, $15 output (doubled above 128K tokens). Grok Build 0.1: $1 input, $2 output. | Varies by model and tier. | Varies by model and tier. | Varies by model and tier. |
| Subscription Access | X Premium, standalone Grok web/mobile app. SuperGrok Heavy: $300/month ($99 intro for Grok Build). | ChatGPT Plus, Enterprise tiers. | Claude Pro, Enterprise tiers. | Gemini Advanced, Enterprise tiers. |
🛠️ Technical Deep Dive
- Architecture: Grok models, starting with Grok-1, utilize a decoder-only Transformer architecture with a Mixture-of-Experts (MoE) design. Grok-1 specifically had 314 billion parameters with 2 out of 8 experts active per token.
- Training Infrastructure: xAI developed a custom training and inference stack built on Kubernetes, Rust, and JAX. Grok 3 was notably trained on the Colossus supercomputer cluster, which reportedly includes over 100,000 NVIDIA H100 GPUs.
- Hardware Optimization: The Grok V9-Medium model is specifically optimized to run on Blackwell architecture GPUs, indicating a focus on leveraging next-generation hardware for performance.
- Data Pipeline: A key differentiator is Grok's real-time access to the X (Twitter) public stream, which serves as both a training and retrieval source, enabling the model to handle recency, slang, and internet culture effectively.
- Coding Agent Design: The recently launched Grok Build CLI, powered by
grok-code-fast-1, features Git-worktree isolation for parallel sub-agents, allowing them to experiment in isolated branches before merging, which is considered an innovation in agentic-CLI design.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (27)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- basenor.com
- digg.com
- kucoin.com
- kucoin.com
- phemex.com
- digg.com
- aicerts.ai
- guptadeepak.com
- datastudios.org
- business-standard.com
- medium.com
- medium.com
- futureagi.com
- medium.com
- guptadeepak.com
- x.ai
- x.ai
- sintra.ai
- digitalapplied.com
- deeplearning.ai
- datacamp.com
- wikipedia.org
- timesofai.com
- openrouter.ai
- x.ai
- mindstudio.ai
- issarice.com
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗