๐Ÿ“ŠStalecollected in 16m

Google Researchers Compete for Internal AI Compute Resources

PostLinkedIn
๐Ÿ“ŠRead original on Bloomberg Technology

๐Ÿ’กUnderstand how compute scarcity at Google is shaping the future of AI development and infrastructure access.

โšก 30-Second TL;DR

What Changed

Google is balancing internal AI research needs with external cloud commitments.

Why It Matters

This resource scarcity highlights the bottleneck of compute in large-scale AI development. It suggests that even tech giants are struggling to scale infrastructure fast enough to meet the demands of modern LLM training.

What To Do Next

Diversify your cloud infrastructure strategy by testing multi-cloud deployments to avoid dependency on a single provider's constrained compute.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขGoogle is balancing internal AI research needs with external cloud commitments.
  • โ€ขThe company produces its own custom AI chips to power its infrastructure.
  • โ€ขStrategic partnerships with Anthropic and Meta are straining available compute capacity.

๐Ÿง  Deep Insight

Web-grounded analysis with 31 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขGoogle established an internal 'compute allocation council' comprising the Cloud CEO, DeepMind head, Search/Ads head, and CFO to manage the scarcity of AI compute resources, with the CFO disclosing that Google Cloud receives approximately half of the company's total capacity.
  • โ€ขThe demand for Google's custom Tensor Processing Units (TPUs) has led to multi-billion dollar deals, including a six-year agreement worth over $10 billion with Meta Platforms for AI infrastructure, and a separate agreement with Anthropic for access to up to 1 million TPUs and over a gigawatt of capacity.
  • โ€ขGoogle's latest TPU generation, the 8th generation (TPU 8t and 8i), bifurcates the architecture into specialized chips for training and inference, respectively, marking a significant design shift from previous generations that handled both workloads.
  • โ€ขGoogle's TPU fleet grew 11.5x in seven quarters, and by Q4 2025, its TPU fleet alone (3.81 million H100-equivalents) exceeded Microsoft's entire AI compute position (3.42 million H100-equivalents), drawing more power than Microsoft's total estimated AI compute stack.
  • โ€ขThe initial development of TPUs in 2013 was driven by an internal projection that if a small percentage of Google users adopted voice search for a few minutes daily, it would require doubling Google's entire compute capacity at the time, highlighting the early recognition of AI's immense compute demands.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/ProviderGoogle Cloud (TPUs)AWS (Trainium/GPUs)Microsoft Azure (GPUs)
Custom AI ChipsTensor Processing Units (TPUs) - v1 to v8t/8i, specialized for AI workloads.AWS Trainium (for ML training), AWS Inferentia (for inference).No custom chips widely offered to public; relies heavily on NVIDIA GPUs.
Primary HardwareCustom TPUs (e.g., Ironwood, TPU 8t/8i) with some NVIDIA GPUs.NVIDIA GPUs (e.g., H100/H200 series) and AWS Trainium/Inferentia.NVIDIA GPUs (e.g., H100/H200 series) with strong integration for OpenAI.
Network FabricOptical Circuit Switching, Virgo Network, high-speed Inter-Chip Interconnect (ICI) up to 19.2 Tb/s.Significant investments in networking innovations, 10p10u infrastructure (10s of petabits bandwidth, <10 microseconds latency).Accelerated networking, highly reliable clusters, extreme throughput for GPU VMs.
ScalabilitySuperpods scaling up to 9,600 chips (TPU 8t) and 1.77 petabytes of shared HBM (Ironwood).EC2 UltraClusters, SageMaker HyperPod for large-scale model development.Scalable clusters with robust checkpointing, GPU-powered VMs.
Key PartnershipsAnthropic, Meta, OpenAI (earlier in 2025).Anthropic (multi-platform approach, primary cloud provider).Exclusive cloud provider for OpenAI.
Market Position (AI Accelerators)Owns about one-quarter of global AI compute capacity, largely with TPUs, less reliant on NVIDIA than competitors.Significant player, offers custom chips and NVIDIA GPUs.Second in compute capacity after Google, heavily reliant on NVIDIA.
Pricing/Cost EfficiencyTPU v5e offers 2.5x better price-performance than v4 for inference; TPU 8i offers 80% better performance-per-dollar for inference.AWS Trainium aims to make AI more affordable. EC2 Capacity Blocks for ML for predictable access.Intel's Gaudi chips aim to be 50% cheaper than NVIDIA H100 (relevant for Azure customers).

๐Ÿ› ๏ธ Technical Deep Dive

  • TPU Architecture Evolution: Google's TPUs are Application-Specific Integrated Circuits (ASICs) optimized for neural network machine learning, specifically matrix multiplication and tensor operations.
  • Systolic Arrays: The core compute engine, the Matrix Multiply Unit (MXU), uses systolic arrays. TPU v1 had a 256x256 INT8 array. TPU v2 introduced dual 128x128 bfloat16 multiply-accumulators per core. Starting with Trillium (v6e), the MXU expanded to 256x256, quadrupling operations per cycle.
  • Memory: High Bandwidth Memory (HBM) capacity and bandwidth are critical. TPU v2 had 16GB HBM per chip (600 GB/s bandwidth). TPU v5p features 95 GB HBM per chip (2,765 GB/s bandwidth). Ironwood (v7) increased HBM to 192 GB per chip, and TPU 8i boasts 288 GB HBM3e with 8,601 GB/s bandwidth and 384 MB of on-chip SRAM.
  • Inter-Chip Interconnect (ICI): Connects chips within a pod. TPU v4 introduced optical circuit switches. TPU v5p scales to 8,960 chips per pod with 4,800 Gbps ICI bandwidth per chip. TPU 8i doubled ICI bandwidth to 19.2 Tb/s and introduced a 'Boardfly' network topology to reduce network diameter.
  • SparseCores: Introduced in TPU v4 and enhanced in v5p and Ironwood, these are designed for efficiently handling large embeddings, crucial for recommendation systems and Large Language Models (LLMs).
  • Precision: TPU v1 was inference-only using 8-bit integer precision. TPU v2 introduced bfloat16 format, enabling both training and inference.
  • Cooling: Liquid cooling was added with TPU v3 to address efficiency needs. TPU 8t and 8i utilize 4th-generation liquid cooling.
  • Host CPUs: For the first time, TPU 8t and 8i run on Google's custom Arm-based Axion CPUs, allowing for full system optimization.
  • Specialized Design (TPU 8th Gen): TPU 8t is optimized for large-scale pre-training and embedding-heavy workloads, scaling to 9,600 chips in a single superpod (121 FP4 ExaFLOPs). TPU 8i is designed for high-speed inference, AI agents, and long-context reasoning, delivering 10.1 FP4 PFLOPs and featuring a Collectives Acceleration Engine (CAE) to reduce synchronization latency by 5x.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Google's internal compute allocation challenges will intensify, potentially impacting the speed of its proprietary AI model development.
With increasing external commitments to partners like Meta and Anthropic, and continued internal demand from various Google products, the competition for finite TPU resources will likely grow, requiring more stringent allocation strategies.
The bifurcation of Google's TPU architecture into specialized training (8t) and inference (8i) chips will lead to greater efficiency and performance gains for specific AI workloads.
By optimizing chips for distinct tasks, Google can achieve higher performance-per-dollar and better resource utilization for both large-scale model training and high-speed, low-latency inference for AI agents.
Google Cloud's strategic partnerships and custom silicon offerings will further solidify its position as a major AI infrastructure provider, challenging NVIDIA's dominance.
Securing multi-billion dollar deals with major AI players like Meta and Anthropic for TPU access validates the commercial viability and performance of Google's custom chips, offering a strong alternative to NVIDIA's GPUs.

โณ Timeline

2013-XX
Google Brain team identifies that AI compute demand will outpace existing infrastructure, initiating custom silicon development.
2015-XX
First-generation Tensor Processing Unit (TPU v1) deployed internally at Google for inference workloads.
2017-05
Second-generation TPU (TPU v2) launched, capable of both training and inference, and made available on Google Compute Engine.
2018-05
Third-generation TPU (TPU v3) announced, doubling power and deployed in pods with four times as many chips as v2.
2021-XX
Fourth-generation TPU (TPU v4) unveiled, introducing optical circuit switches and SparseCores.
2023-12
Fifth-generation TPUs (TPU v5e and v5p) launched, offering cost-efficiency (v5e) and performance (v5p) optimizations.
2025-08
Meta Platforms signs a six-year, $10+ billion deal with Google Cloud for AI infrastructure, including TPU access.
2025-10
Anthropic expands its partnership with Google Cloud, securing access to up to 1 million TPUs and over a gigawatt of capacity.
2026-04
Google announces its eighth-generation TPUs (TPU 8t and 8i), specialized for training and inference respectively, running on custom Axion Arm-based CPUs.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Bloomberg Technology โ†—