Nvidia Tests Smaller Rubin Ultra Memory Configurations

Rubin Ultra may ship with far less memory, changing how large AI workloads are planned and budgeted.
30-Second TL;DR
What Changed
At least three Rubin Ultra memory configurations are reportedly being tested.
Why It Matters
Reduced memory capacity could limit the model size, batch size, or parallel workloads supported by Rubin Ultra systems. For AI infrastructure planners, the change could also affect server density, software partitioning, and procurement assumptions.
What To Do Next
Add 192 GB HBM4 and 1 TB HBM4E scenarios to your capacity-planning spreadsheet, then recalculate maximum model size and batch size for each configuration.
Key Points
- •At least three Rubin Ultra memory configurations are reportedly being tested.
- •The tested designs may include as little as 192 GB of memory.
- •The configurations would reduce capacity from the originally announced 1 TB.
- •Some designs reportedly step back from HBM4E to HBM4.
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •The shift toward lower memory capacities is driven by the extreme manufacturing complexity and yield challenges associated with HBM4E stacks, which require advanced 12-high or 16-high stacking processes.
- •Nvidia's Rubin architecture is designed to utilize a modular chiplet-based approach, allowing for greater flexibility in mixing and matching memory controllers and HBM stacks based on specific customer workload requirements.
- •Industry analysts suggest that the 192 GB configurations are targeted at inference-heavy data center deployments where memory bandwidth is critical, but total capacity requirements are lower than those needed for massive model training.
- •The transition between HBM4 and HBM4E involves a move to a 2048-bit interface per stack, which significantly increases the physical footprint on the interposer, complicating the design of high-capacity modules.
- •Supply chain reports indicate that Nvidia is working closely with SK Hynix and Samsung to optimize the thermal management of these smaller memory configurations to maintain performance parity with higher-capacity variants.
Competitor Analysis
- Nvidia Rubin Ultra (192GB)
- HBM4 / HBM4E
- AMD Instinct MI400 Series
- HBM4
- Intel Gaudi 4
- HBM3E / HBM4
- Nvidia Rubin Ultra (192GB)
- Blackwell Successor
- AMD Instinct MI400 Series
- CDNA 4
- Intel Gaudi 4
- Falcon Shores
- Nvidia Rubin Ultra (192GB)
- AI Training/Inference
- AMD Instinct MI400 Series
- AI Training
- Intel Gaudi 4
- AI Inference
- Nvidia Rubin Ultra (192GB)
- Testing
- AMD Instinct MI400 Series
- Development
- Intel Gaudi 4
- Development
| Feature | Nvidia Rubin Ultra (192GB) | AMD Instinct MI400 Series | Intel Gaudi 4 |
|---|---|---|---|
| Memory Type | HBM4 / HBM4E | HBM4 | HBM3E / HBM4 |
| Architecture | Blackwell Successor | CDNA 4 | Falcon Shores |
| Target Market | AI Training/Inference | AI Training | AI Inference |
| Status | Testing | Development | Development |
Technical Deep Dive
- Rubin Ultra utilizes a 2048-bit memory interface per HBM stack, doubling the width compared to previous HBM3E generations.
- The 192 GB configuration likely employs a reduced number of HBM4 stacks or lower-density dies to manage power consumption and thermal output.
- Implementation relies on advanced CoWoS-L (Chip-on-Wafer-on-Substrate) packaging to integrate the GPU compute die with the memory stacks.
- The architecture supports dynamic memory allocation, allowing the GPU to treat the 192 GB pool as a unified high-speed cache for large language model weights.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2024-06Nvidia announces the Rubin architecture roadmap at Computex.
- 2025-03Initial specifications for Rubin Ultra featuring 1 TB HBM4E are leaked.
- 2026-02Nvidia begins internal validation of early Rubin silicon samples.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Tom's Hardware ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.