SpaceX leases Colossus 1 to Anthropic due to latency

💡Learn why SpaceX's massive AI cluster failed to support Grok and why infrastructure location is a hidden model killer.
⚡ 30-Second TL;DR
What Changed
SpaceX failed to optimize Colossus 1 for Grok AI workloads
Why It Matters
This highlights the critical importance of network topology and physical proximity in large-scale cluster deployments. It serves as a warning for infrastructure teams regarding the complexity of multi-site AI training.
What To Do Next
Audit your distributed training architecture for network latency bottlenecks before committing to multi-site data center leases.
Key Points
- •SpaceX failed to optimize Colossus 1 for Grok AI workloads
- •Significant latency issues occurred between Memphis and other data centers
- •Anthropic is now utilizing the facility instead of SpaceX
- •Infrastructure challenges remain a bottleneck for large-scale AI training
🧠 Deep Insight
Background and context from public sources — not the original article. 26 sources cited.
🔑 Enhanced Key Takeaways
- •Colossus 1, originally built by xAI in 122 days in 2024 at a former Electrolux site in Memphis, Tennessee, houses over 220,000 Nvidia GPUs (H100, H200, and next-generation GB200 accelerators) and draws over 300 megawatts of power.
- •The latency issues that led to the lease stemmed from Colossus 1's heterogeneous GPU architecture, which made distributed training for Grok inefficient, prompting xAI to shift its core training workloads to the more uniformly built Colossus 2.
- •Anthropic's lease for Colossus 1, signed on May 6, 2026, is for $1.25 billion per month, but Elon Musk clarified on May 28, 2026, that it is structured as a 180-day term with rolling 90-day cancellation rights, not a multi-year commitment.
- •SpaceX, having merged with xAI in February 2026, is aggressively pursuing orbital AI computing, aiming for initial demonstrations by late 2027 and seeking regulatory approval to deploy up to 1 million space-based data center satellites.
- •Anthropic is simultaneously pursuing a multi-cloud infrastructure strategy, including partnerships with AWS Project Rainier, Google Cloud TPUs, a $30 billion Microsoft Azure commitment, and a $50 billion data center partnership with Fluidstack.
📊 Competitor Analysis▸ Show
| Feature/Strategy | Anthropic | OpenAI | SpaceX/xAI |
|---|---|---|---|
| Terrestrial Data Centers | Multi-cloud (AWS, Google, Azure, Fluidstack, direct leases including Colossus 1) | Stargate project (single $500B JV with SoftBank, Oracle, MGX targeting 10 GW by 2029) | Colossus 1 (leased to Anthropic), Colossus 2 (Blackwell-based, leased partly to Google) |
| Orbital AI Compute | Expressed interest in partnering with SpaceX | Not publicly detailed | Aggressively pursuing, targeting initial tests by late 2027, filed for 1M space-based data center satellites |
| Compute Capacity (Terrestrial) | Access to 220,000+ Nvidia GPUs (Colossus 1), 500K-1M Trainium2 chips (AWS), up to 1M TPU chips (Google Cloud), Fluidstack facilities | ~250,000 GPUs (CoreWeave deal) | Colossus 1 (220,000+ GPUs), Colossus 2 (initial 550,000 NVIDIA GPUs, scaling to 1M) |
| Lease Cost (Annualized) | ~$15 billion/year for Colossus 1 (short-term) | ~$2.38 billion/year for CoreWeave (5-year deal) | N/A (provider) |
| AI Model Focus | Claude (known for safety and enterprise applications) | GPT series (leading general-purpose LLMs) | Grok (consumer-focused, integrated with X, Tesla, Starlink) |
🛠️ Technical Deep Dive
- Colossus 1 Architecture: Built with over 220,000 Nvidia GPUs, including a mix of H100, H200, and early Blackwell architectures. It is capable of drawing over 300 megawatts of power and utilizes advanced custom liquid-cooling loops. Initially, 100,000 Nvidia H100 GPUs were connected via a single RDMA fabric.
- Latency Issues: The primary technical challenge with Colossus 1 for Grok's training was its heterogeneous GPU configuration. Distributed AI training requires all GPUs in a cluster to complete each computational step simultaneously, and a mixed architecture with different generations of Nvidia silicon (H100, H200, GB200) created significant efficiency problems and latency.
- Colossus 2 Design: In contrast, Colossus 2, located in Southaven, Mississippi, is a unified cluster built entirely on cutting-edge Blackwell chips, designed for more efficient and scalable AI training. It hosts an initial batch of 550,000 NVIDIA GPUs with a roadmap to scale to 1 million interconnected GPUs.
- Orbital AI Computing: SpaceX's proposed orbital AI infrastructure, exemplified by the AI1 satellite design, incorporates advanced thermal management and power systems for high-density computing in space. These satellites would likely connect to ground stations via laser links for low-latency data transmission. For Grok, edge computing is planned, deploying quantized AI models directly onto spacecraft to process data and make localized decisions instantly, mitigating space communication latency.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (26)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- capacityglobal.com
- wikipedia.org
- medium.com
- datacentermap.com
- moomoo.com
- tomshardware.com
- forbes.com
- qz.com
- actuia.com
- stocktwits.com
- indiatimes.com
- sentisight.ai
- techjacksolutions.com
- cleanroomtechnology.com
- introl.com
- medium.com
- constellationr.com
- yellow.com
- localmemphis.com
- aibusiness.com
- datacenterknowledge.com
- xsoneconsultants.com
- ca.gov
- morningstar.com
- anthropic.com
- anthropic.com
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.

