SourceStalecollected in 36m

Nvidia Tests Smaller Rubin Ultra Memory Configurations

Read original on Tom's Hardware
#memory-supply#accelerator-design#data-center

Rubin Ultra may ship with far less memory, changing how large AI workloads are planned and budgeted.

30-Second TL;DR

What Changed

At least three Rubin Ultra memory configurations are reportedly being tested.

Why It Matters

Reduced memory capacity could limit the model size, batch size, or parallel workloads supported by Rubin Ultra systems. For AI infrastructure planners, the change could also affect server density, software partitioning, and procurement assumptions.

What To Do Next

Add 192 GB HBM4 and 1 TB HBM4E scenarios to your capacity-planning spreadsheet, then recalculate maximum model size and batch size for each configuration.

Who should care:Researchers & Academics

Key Points

  • •At least three Rubin Ultra memory configurations are reportedly being tested.
  • •The tested designs may include as little as 192 GB of memory.
  • •The configurations would reduce capacity from the originally announced 1 TB.
  • •Some designs reportedly step back from HBM4E to HBM4.

Deep Insight

AI-generated analysis for this event — not the original article.

Enhanced Key Takeaways

  • •The shift toward lower memory capacities is driven by the extreme manufacturing complexity and yield challenges associated with HBM4E stacks, which require advanced 12-high or 16-high stacking processes.
  • •Nvidia's Rubin architecture is designed to utilize a modular chiplet-based approach, allowing for greater flexibility in mixing and matching memory controllers and HBM stacks based on specific customer workload requirements.
  • •Industry analysts suggest that the 192 GB configurations are targeted at inference-heavy data center deployments where memory bandwidth is critical, but total capacity requirements are lower than those needed for massive model training.
  • •The transition between HBM4 and HBM4E involves a move to a 2048-bit interface per stack, which significantly increases the physical footprint on the interposer, complicating the design of high-capacity modules.
  • •Supply chain reports indicate that Nvidia is working closely with SK Hynix and Samsung to optimize the thermal management of these smaller memory configurations to maintain performance parity with higher-capacity variants.

Competitor Analysis

Memory Type
Nvidia Rubin Ultra (192GB)
HBM4 / HBM4E
AMD Instinct MI400 Series
HBM4
Intel Gaudi 4
HBM3E / HBM4
Architecture
Nvidia Rubin Ultra (192GB)
Blackwell Successor
AMD Instinct MI400 Series
CDNA 4
Intel Gaudi 4
Falcon Shores
Target Market
Nvidia Rubin Ultra (192GB)
AI Training/Inference
AMD Instinct MI400 Series
AI Training
Intel Gaudi 4
AI Inference
Status
Nvidia Rubin Ultra (192GB)
Testing
AMD Instinct MI400 Series
Development
Intel Gaudi 4
Development

Technical Deep Dive

  • Rubin Ultra utilizes a 2048-bit memory interface per HBM stack, doubling the width compared to previous HBM3E generations.
  • The 192 GB configuration likely employs a reduced number of HBM4 stacks or lower-density dies to manage power consumption and thermal output.
  • Implementation relies on advanced CoWoS-L (Chip-on-Wafer-on-Substrate) packaging to integrate the GPU compute die with the memory stacks.
  • The architecture supports dynamic memory allocation, allowing the GPU to treat the 192 GB pool as a unified high-speed cache for large language model weights.

Future ImplicationsAI analysis grounded in cited sources

Nvidia will prioritize modularity over maximum capacity in the Rubin generation.
The testing of multiple memory configurations indicates a strategic shift toward SKU diversification to mitigate supply chain bottlenecks.
HBM4E adoption will be slower than initially projected by market analysts.
The reported fallback to HBM4 in some Rubin Ultra designs suggests that HBM4E yields are not yet sufficient for mass production.

Timeline

2024-06
Nvidia announces the Rubin architecture roadmap at Computex.
2025-03
Initial specifications for Rubin Ultra featuring 1 TB HBM4E are leaked.
2026-02
Nvidia begins internal validation of early Rubin silicon samples.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Tom's Hardware ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.