NVIDIA Enters CPU Market with Vera
💡NVIDIA's pivot to CPUs could redefine data center infrastructure and supply chain priorities for AI developers.
⚡ 30-Second TL;DR
What Changed
NVIDIA's Vera CPU is built on Arm architecture for AI inference and agentic workloads.
Why It Matters
NVIDIA's move into CPUs could disrupt the data center market and pressure traditional x86 providers while straining supply chains.
What To Do Next
Evaluate the performance benchmarks of Vera-based systems versus traditional x86 servers for your AI inference deployment.
Key Points
- •NVIDIA's Vera CPU is built on Arm architecture for AI inference and agentic workloads.
- •NVIDIA projects $20 billion in CPU-related revenue visibility for the current year.
- •Analysts suggest the $20B figure includes bundled system solutions rather than just standalone chips.
- •The entry into the CPU market intensifies competition for wafer and memory capacity.
🧠 Deep Insight
Web-grounded analysis with 18 cited sources.
🔑 Enhanced Key Takeaways
- •NVIDIA's Vera CPU is specifically engineered for reinforcement learning (RL) and agentic AI, designed to manage complex code, tools, and data workflows that extend beyond traditional model inference.
- •The projected $20 billion in CPU-related revenue for the current fiscal year (FY2027) includes sales of standalone Vera CPU servers and Vera CPUs integrated within NVIDIA's Grace Blackwell and Vera Rubin superchips.
- •The Vera CPU features 88 custom-designed 'Olympus cores' based on Armv9.2 architecture, marking NVIDIA's first in-house CPU core design since the Denver cores nearly a decade ago, rather than licensing an off-the-shelf Arm design.
- •Vera CPUs offer up to 1.2 terabytes per second (TB/s) of LPDDR5X memory bandwidth and support up to 1.5 TB of memory per socket, which is crucial for memory-intensive agentic AI and analytics workloads and significantly surpasses traditional CPU memory capabilities.
📊 Competitor Analysis▸ Show
Competitor Analysis: NVIDIA Vera/Grace vs. Intel Xeon & AMD EPYC
| Feature/Metric | NVIDIA Vera CPU (Olympus Cores) | NVIDIA Grace CPU (Neoverse V2 Cores) | AMD EPYC 9005 Turin (Zen 5c) | Intel Xeon 6 Granite Rapids |
|---|---|---|---|---|
| Architecture | Custom Armv9.2 (Olympus) | Arm Neoverse V2 | x86 (Zen 5c) | x86 |
| Core Count (per chip/socket) | 88 cores (176 threads with Spatial Multithreading) | 72 cores (single chip), 144 cores (Superchip) | Up to 192 cores (Zen 5c) | Up to 136 cores (estimated for top SKU) |
| Memory Type | LPDDR5X (SOCAMM modules) | LPDDR5X with ECC | DDR5 | DDR5 / MRDIMM |
| Memory Bandwidth | Up to 1.2 TB/s | Up to 546 GB/s (single Grace), 1 TB/s (Grace Superchip) | Up to ~614 GB/s (DDR5-6400) | Up to ~845 GB/s (MRDIMM-8800) |
| Memory Capacity (per socket) | Up to 1.5 TB | Up to 480GB (single Grace), 960GB (Grace Superchip) | Up to 3 TB | High capacity, specific details vary |
| Interconnect | NVLink-C2C (1.8 TB/s), PCIe Gen6, CXL 3.1 | NVLink-C2C (900 GB/s), PCIe Gen5 | PCIe Gen5 (up to 128 lanes) | PCIe Gen5 (up to 136 lanes) |
| AI Features | FP8 precision support (first CPU), agentic AI optimization | Optimized for AI/HPC, CUDA integration | AI inference acceleration, AVX-512 | AMX acceleration for CPU-based inference |
| Performance Claims | 2x performance of predecessor, 50% faster agentic sandbox, 4x sandbox density, 2x perf/watt over x86 racks | 2x performance per watt, 2x packaging density vs. leading servers | Up to 2.75x better power efficiency vs. Grace (dual-socket), 2.17x higher database perf vs. Grace | 5.5x ResNet50 performance advantage (AMX) |
| Pricing | Not specified for standalone chip; part of rack-scale solutions | Not specified for standalone chip | EPYC 9965: ~$14,813 | Not specified |
Note: Performance and pricing can vary significantly based on specific configurations, workloads, and market conditions. "Vera" is a newer product, so some comparisons draw on "Grace" where direct Vera benchmarks against competitors are still emerging or less detailed. Vera is positioned as an evolution of Grace, with significant improvements.
🛠️ Technical Deep Dive
- CPU Cores: Features 88 NVIDIA-designed 'Olympus cores' with full Armv9.2 compatibility.
- Multithreading: Implements NVIDIA Spatial Multithreading, enabling 176 total threads by physically partitioning core resources for optimized performance or density.
- Precision Support: First CPU to support FP8 precision, crucial for reinforcement learning and agentic AI workloads.
- Memory Subsystem: Utilizes LPDDR5X memory delivered through SOCAMM modules, offering up to 1.5 TB capacity and 1.2 TB/s bandwidth per socket. This provides approximately 13.6-14 GB/s of memory bandwidth per core.
- Architecture: Built on a single monolithic compute die with adjacent dielets for memory and I/O subsystems, preserving a uniform compute topology and avoiding NUMA domains.
- Coherency Fabric: Incorporates the second-generation NVIDIA Scalable Coherency Fabric (SCF) with 3.4 TB/s on-chip bisection bandwidth for efficient data flow.
- Interconnect: Features NVIDIA NVLink™ Chip-to-Chip (C2C) connectivity with 1.8 TB/s of coherent bandwidth for high-speed CPU-to-CPU and CPU-to-GPU communication.
- I/O: Supports PCIe Gen6 and CXL 3.1 for external connectivity.
- Confidential Computing: Supports full confidential computing across the CPU-GPU domain for enhanced security.
- Rack-Scale Integration: Designed for rack-scale deployment, such as the NVIDIA Vera CPU Rack (up to 256 Vera CPUs) and the NVIDIA Vera Rubin NVL72 platform (72 Rubin GPUs, 36 Vera CPUs).
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (18)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗


