🐯Stalecollected in 14m

NVIDIA Enters CPU Market with Vera

PostLinkedIn
🐯Read original on 虎嗅

💡NVIDIA's pivot to CPUs could redefine data center infrastructure and supply chain priorities for AI developers.

⚡ 30-Second TL;DR

What Changed

NVIDIA's Vera CPU is built on Arm architecture for AI inference and agentic workloads.

Why It Matters

NVIDIA's move into CPUs could disrupt the data center market and pressure traditional x86 providers while straining supply chains.

What To Do Next

Evaluate the performance benchmarks of Vera-based systems versus traditional x86 servers for your AI inference deployment.

Who should care:Developers & AI Engineers

Key Points

  • NVIDIA's Vera CPU is built on Arm architecture for AI inference and agentic workloads.
  • NVIDIA projects $20 billion in CPU-related revenue visibility for the current year.
  • Analysts suggest the $20B figure includes bundled system solutions rather than just standalone chips.
  • The entry into the CPU market intensifies competition for wafer and memory capacity.

🧠 Deep Insight

Web-grounded analysis with 18 cited sources.

🔑 Enhanced Key Takeaways

  • NVIDIA's Vera CPU is specifically engineered for reinforcement learning (RL) and agentic AI, designed to manage complex code, tools, and data workflows that extend beyond traditional model inference.
  • The projected $20 billion in CPU-related revenue for the current fiscal year (FY2027) includes sales of standalone Vera CPU servers and Vera CPUs integrated within NVIDIA's Grace Blackwell and Vera Rubin superchips.
  • The Vera CPU features 88 custom-designed 'Olympus cores' based on Armv9.2 architecture, marking NVIDIA's first in-house CPU core design since the Denver cores nearly a decade ago, rather than licensing an off-the-shelf Arm design.
  • Vera CPUs offer up to 1.2 terabytes per second (TB/s) of LPDDR5X memory bandwidth and support up to 1.5 TB of memory per socket, which is crucial for memory-intensive agentic AI and analytics workloads and significantly surpasses traditional CPU memory capabilities.
📊 Competitor Analysis▸ Show

Competitor Analysis: NVIDIA Vera/Grace vs. Intel Xeon & AMD EPYC

Feature/MetricNVIDIA Vera CPU (Olympus Cores)NVIDIA Grace CPU (Neoverse V2 Cores)AMD EPYC 9005 Turin (Zen 5c)Intel Xeon 6 Granite Rapids
ArchitectureCustom Armv9.2 (Olympus)Arm Neoverse V2x86 (Zen 5c)x86
Core Count (per chip/socket)88 cores (176 threads with Spatial Multithreading)72 cores (single chip), 144 cores (Superchip)Up to 192 cores (Zen 5c)Up to 136 cores (estimated for top SKU)
Memory TypeLPDDR5X (SOCAMM modules)LPDDR5X with ECCDDR5DDR5 / MRDIMM
Memory BandwidthUp to 1.2 TB/sUp to 546 GB/s (single Grace), 1 TB/s (Grace Superchip)Up to ~614 GB/s (DDR5-6400)Up to ~845 GB/s (MRDIMM-8800)
Memory Capacity (per socket)Up to 1.5 TBUp to 480GB (single Grace), 960GB (Grace Superchip)Up to 3 TBHigh capacity, specific details vary
InterconnectNVLink-C2C (1.8 TB/s), PCIe Gen6, CXL 3.1NVLink-C2C (900 GB/s), PCIe Gen5PCIe Gen5 (up to 128 lanes)PCIe Gen5 (up to 136 lanes)
AI FeaturesFP8 precision support (first CPU), agentic AI optimizationOptimized for AI/HPC, CUDA integrationAI inference acceleration, AVX-512AMX acceleration for CPU-based inference
Performance Claims2x performance of predecessor, 50% faster agentic sandbox, 4x sandbox density, 2x perf/watt over x86 racks2x performance per watt, 2x packaging density vs. leading serversUp to 2.75x better power efficiency vs. Grace (dual-socket), 2.17x higher database perf vs. Grace5.5x ResNet50 performance advantage (AMX)
PricingNot specified for standalone chip; part of rack-scale solutionsNot specified for standalone chipEPYC 9965: ~$14,813Not specified

Note: Performance and pricing can vary significantly based on specific configurations, workloads, and market conditions. "Vera" is a newer product, so some comparisons draw on "Grace" where direct Vera benchmarks against competitors are still emerging or less detailed. Vera is positioned as an evolution of Grace, with significant improvements.

🛠️ Technical Deep Dive

  • CPU Cores: Features 88 NVIDIA-designed 'Olympus cores' with full Armv9.2 compatibility.
  • Multithreading: Implements NVIDIA Spatial Multithreading, enabling 176 total threads by physically partitioning core resources for optimized performance or density.
  • Precision Support: First CPU to support FP8 precision, crucial for reinforcement learning and agentic AI workloads.
  • Memory Subsystem: Utilizes LPDDR5X memory delivered through SOCAMM modules, offering up to 1.5 TB capacity and 1.2 TB/s bandwidth per socket. This provides approximately 13.6-14 GB/s of memory bandwidth per core.
  • Architecture: Built on a single monolithic compute die with adjacent dielets for memory and I/O subsystems, preserving a uniform compute topology and avoiding NUMA domains.
  • Coherency Fabric: Incorporates the second-generation NVIDIA Scalable Coherency Fabric (SCF) with 3.4 TB/s on-chip bisection bandwidth for efficient data flow.
  • Interconnect: Features NVIDIA NVLink™ Chip-to-Chip (C2C) connectivity with 1.8 TB/s of coherent bandwidth for high-speed CPU-to-CPU and CPU-to-GPU communication.
  • I/O: Supports PCIe Gen6 and CXL 3.1 for external connectivity.
  • Confidential Computing: Supports full confidential computing across the CPU-GPU domain for enhanced security.
  • Rack-Scale Integration: Designed for rack-scale deployment, such as the NVIDIA Vera CPU Rack (up to 256 Vera CPUs) and the NVIDIA Vera Rubin NVL72 platform (72 Rubin GPUs, 36 Vera CPUs).

🔮 Future ImplicationsAI analysis grounded in cited sources

NVIDIA's Vera CPU will significantly accelerate the adoption of agentic AI.
Vera is purpose-built for reinforcement learning and agentic AI, offering specialized features like FP8 precision and high memory bandwidth crucial for these workloads, enabling faster iterations and efficient KV-cache management.
The entry of Vera will intensify the shift towards Arm-based CPUs in data centers for AI workloads.
Vera's performance and efficiency claims, coupled with NVIDIA's established AI ecosystem, will drive further adoption of Arm in a market traditionally dominated by x86, especially for AI-specific tasks.
NVIDIA's strategy of offering integrated CPU-GPU platforms (Vera Rubin) will increase its ecosystem lock-in and make it harder for competitors to offer comparable full-stack solutions.
The Vera CPU is designed to pair seamlessly with NVIDIA GPUs (Rubin GPUs) via NVLink-C2C, creating a unified memory pool and optimized data movement, which is a key differentiator for AI factories.

Timeline

2020-09
NVIDIA announces intent to acquire Arm Holdings.
2022-02
NVIDIA's $40 billion acquisition of Arm collapses due to significant regulatory challenges.
2022-03
NVIDIA unveils the Grace CPU Superchip, its first Arm-based data center CPU.
2026-01
NVIDIA Vera Rubin architecture, including the Vera CPU, is detailed as a comprehensive AI supercomputing platform.
2026-03
NVIDIA officially launches the Vera CPU, purpose-built for agentic AI and reinforcement learning.
2026-H2
Commercial availability of Vera CPU from major OEMs and the Vera Rubin platform is expected.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅