Google to pay SpaceX $920M monthly for AI compute

๐กGoogle's $920M/month compute deal reveals the massive scale of infrastructure needed for modern AI.
โก 30-Second TL;DR
What Changed
Contract value: $920 million per month from Oct 2026 to June 2029.
Why It Matters
This deal underscores the extreme scarcity of high-end AI compute and the massive capital expenditure required by big tech to maintain competitive AI infrastructure.
What To Do Next
Evaluate your infrastructure scaling strategy, as compute costs remain a critical bottleneck for large-scale AI deployment.
Key Points
- โขContract value: $920 million per month from Oct 2026 to June 2029.
- โขHardware: Access to approximately 110,000 Nvidia GPUs, CPUs, and memory.
- โขInfrastructure: Utilization of a data center originally built for Grok.
- โขStrategic partnership between Google and SpaceX for AI compute scaling.
๐ง Deep Insight
Web-grounded analysis with 21 cited sources.
๐ Enhanced Key Takeaways
- โขThe data center Google is leveraging, Colossus 1 in Memphis, Tennessee, was originally built for xAI's Grok but reportedly proved difficult for xAI to effectively train its models on due to a 'mish-mash architecture' of H100, H200, and GB200 GPUs, leading xAI to shift its training to Colossus 2.
- โขThis deal, along with a similar prior agreement with Anthropic for 220,000 GPUs at $1.25 billion per month, is part of SpaceX's broader strategy to monetize its underutilized AI compute resources from Colossus 1 and enhance its financial standing ahead of an anticipated initial public offering (IPO).
- โขGoogle's primary motivation for this partnership is to secure 'bridge capacity' to address unexpectedly high customer demand for its agentic AI platform, Gemini Enterprise.
- โขThe agreement includes specific termination clauses, allowing Google to cancel if SpaceX fails to deliver the committed GPU capacity by September 30, 2026, and either party can terminate with 90 days' notice after December 31, 2026.
- โขSpaceX's AI division recently reported an operating loss of $2.5 billion in the last quarter, despite generating $818 million in revenue, highlighting the significant capital expenditures of $7.7 billion allocated to AI infrastructure development.
๐ Competitor Analysisโธ Show
| Feature / Provider | Google Cloud | AWS | Azure | GMI Cloud (Specialized) |
|---|---|---|---|---|
| GPU Offerings | Nvidia H100 (A3 Mega), upcoming Vera Rubin NVL72, proprietary TPUs (v5p, v5e) | Nvidia Blackwell, Rubin, LPUs (planned 1M+ GPUs) | Nvidia H100, H200, upcoming GB200 NVL72 (with InfiniBand) | Nvidia H100, H200, Blackwell (GB200 NVL72, GB200 NVL4, HGX B300) |
| AI Platform | Vertex AI, Gemini integration, GKE | Amazon Bedrock | Azure AI Foundry, exclusive OpenAI models (GPT-4o, o3) | Inference Engine, specialized GPU compute |
| Pricing Model | On-demand, committed use, spot VMs (up to 91% off) | Flexible, various instance types | Flexible, various instance types | On-demand, reserved (e.g., H200 for $3.35/GPU-hour) |
| Specialization | General cloud with strong AI focus, proprietary TPUs | General cloud, broad ecosystem | General cloud, strong enterprise & OpenAI integration | AI-native, high-performance GPU infrastructure |
| Energy Efficiency | Committed to 100% carbon-free data centers by 2030, 6x more computing power per unit of electricity than 5 years ago | Focus on sustainability | Focus on sustainability | SOC 2 certified, focus on performance/cost efficiency |
๐ ๏ธ Technical Deep Dive
- GPU Mix in Colossus 1: The SpaceX data center (Colossus 1) being leased to Google contains an 'eclectic mix' of Nvidia H100, H200, and GB200 GPUs.
- Nvidia H100 Specifications: Each Nvidia H100 accelerator provides approximately 4 petaflops of FP8 compute (with sparsity) and 80 GB of HBM2e memory, offering 2 TB/s bandwidth.
- Colossus 1 Power and Networking: For Grok 3's training, the Colossus supercomputer, which included 200,000 H100 GPUs, consumed an estimated 250 megawatts of power (140 MW for GPUs alone). It utilized Nvidia's Spectrum-X Ethernet platform, featuring 800Gb/s switches and BlueField-3 SuperNICs, to achieve high throughput and low latency comparable to InfiniBand solutions.
- xAI's Training Stack: xAI's infrastructure for Grok training is built on a JAX-based modeling and training layer for composable parallelism, a Rust control plane for orchestrating training jobs and managing failures, and a Kubernetes substrate for scheduling workers and abstracting GPU clusters.
- Google Cloud's Internal AI Infrastructure: Google's A3 Mega supercomputers integrate Nvidia H100 GPUs with Google's proprietary Titanium offload architecture, enabling high-throughput distributed training with standard Ethernet networking up to 800 Gbps. Google also develops its own Tensor Processing Units (TPUs), such as the v5p and v5e, specifically optimized for large language models.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (21)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) โ

