๐Ÿ”งFreshcollected in 16h

HBF Could Add Terabytes to GPU Memory

HBF Could Add Terabytes to GPU Memory
PostLinkedIn
๐Ÿ”งRead original on Tom's Hardware

๐Ÿ’กHBF targets terabytes of extra GPU memory, potentially changing large-model serving economics.

โšก 30-Second TL;DR

What Changed

HBF is designed to extend GPU memory capacity using NAND-based stacks.

Why It Matters

If ecosystem support develops, HBF could let AI systems handle larger models and datasets without relying solely on expensive HBM. Its practical impact remains uncertain because the specification is early-stage and broad industry adoption is not yet established.

What To Do Next

Track HBF specification revisions and prototype support before designing a GPU serving platform around NAND-backed memory expansion.

Who should care:Researchers & Academics

Key Points

  • โ€ขHBF is designed to extend GPU memory capacity using NAND-based stacks.
  • โ€ขThe specification supports up to 16-Hi NAND stacks and targets up to 3 TB/s bandwidth.
  • โ€ขUCIe is part of the proposed interconnect approach, but only four companies currently show interest.

๐Ÿง  Deep Insight

AI-generated analysis for this event.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขHBF (High Bandwidth Flash) utilizes a tiered memory hierarchy that positions NAND as a capacity-extending layer behind traditional HBM3e/HBM4, effectively creating a 'near-memory' storage tier.
  • โ€ขThe architecture leverages the UCIe (Universal Chiplet Interconnect Express) standard to enable low-latency, die-to-die communication between the GPU compute die and the NAND controller.
  • โ€ขSandisk and SK hynix are positioning HBF as a solution to the 'memory wall' in Large Language Model (LLM) inference, allowing massive model weights to reside closer to the GPU than traditional NVMe SSDs.
  • โ€ขThe 16-Hi NAND stack configuration utilizes advanced hybrid bonding techniques to maintain signal integrity at the high-density levels required for terabyte-scale capacity.
  • โ€ขInitial industry feedback suggests HBF aims to reduce the total cost of ownership (TCO) for AI servers by minimizing the reliance on expensive, high-capacity HBM stacks for static model parameters.
๐Ÿ“Š Competitor Analysisโ–ธ Show
FeatureHBF (Sandisk/SK hynix)CXL-based Memory ExpansionTraditional NVMe SSDs
LatencyNear-memory (Low)Moderate (CXL overhead)High (PCIe/OS overhead)
BandwidthUp to 3 TB/sLimited by CXL 3.0/4.0Limited by PCIe Gen5/6
Primary UseGPU Model Weight StorageSystem RAM ExpansionGeneral Storage
PricingPremium (NAND-based)ModerateLow

๐Ÿ› ๏ธ Technical Deep Dive

  • Architecture: Utilizes a disaggregated memory approach where NAND is treated as a memory-mapped device rather than a block-storage device.
  • Interconnect: Employs UCIe 1.1/2.0 for chiplet-level integration, allowing for direct memory access (DMA) patterns between the GPU and the HBF stack.
  • Stack Density: 16-Hi NAND stacks utilize 3D TLC or QLC NAND, optimized for high-throughput read operations required by transformer-based AI models.
  • Controller Logic: Integrates a specialized memory controller within the HBF stack to handle wear leveling and error correction (ECC) without interrupting GPU compute cycles.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

HBF will replace traditional NVMe storage for primary AI model weight loading by 2028.
The massive bandwidth advantage of HBF over PCIe-based storage makes it the only viable solution for loading multi-terabyte models in sub-second timeframes.
GPU manufacturers will begin integrating HBF controllers directly into their silicon interposers.
To achieve the 3 TB/s target, the physical distance between the GPU die and the NAND stack must be minimized, necessitating interposer-level integration.

โณ Timeline

2025-06
SK hynix and Sandisk announce strategic partnership for next-gen memory architectures.
2026-02
Initial whitepaper on HBF (High Bandwidth Flash) concept presented at industry memory summit.
2026-08
Formal introduction of the HBF specification and UCIe integration roadmap.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Tom's Hardware โ†—