🏠較早收集於 10m

NVIDIA Vera Rubin:每瓦效能躍升10倍

NVIDIA Vera Rubin:每瓦效能躍升10倍
PostLinkedIn
🏠閱讀原文: IT之家
#gpu#liquid-cooling#ai-rack#energy-efficiencynvidia-vera-rubinnvidiavera-rubingrace-blackwellrubin-gpuvera-cpu

💡NVIDIA's 10x efficient AI rack slashes datacenter energy costs amid power shortages.

⚡ 30-Second TL;DR

有什麼變化

每瓦效能較 Grace Blackwell 提升 10 倍,儘管功耗加倍

為什麼重要

此效能突破解決 AI 資料中心能源危機,讓訓練規模擴大而不成比例增加功耗。推動液冷成為標準,減少用水並降低超大規模業者成本。

下一步行動

Benchmark Vera Rubin's specs against Blackwell for your next AI training cluster power budget.

誰應關注:Enterprise & Security Teams

關鍵要點

  • 每瓦效能較 Grace Blackwell 提升 10 倍,儘管功耗加倍
  • 72 顆 Rubin GPU + 36 顆 Vera CPU,由台積電生產,來自 130 萬全球零件
  • 首款 100% 液冷機架,具 260TB/s NVLink 及 5000 根銅纜
  • Kyber 機架原型:288 顆 GPU,Vera Rubin Ultra 明年上市,重量僅增 50%

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 7 個來源。

🔑 增強重點摘要

  • Each Rubin GPU features 336 billion transistors on TSMC N3 process and integrates 288GB HBM4 memory with 22 TB/s bandwidth.[1]
  • Vera CPU includes 88 custom Olympus ARMv9.2-compatible cores, 1.5 TB LPDDR5X memory, and 1.2 TB/s memory bandwidth with confidential computing support.[2][3]
  • Rubin platform introduces third-generation Transformer Engine, NVLink 6 with 3.6 TB/s per GPU, ConnectX-9 SuperNIC, and BlueField-4 DPU for enhanced networking.[2][4]
  • Vera Rubin NVL72 delivers 3.6 EFLOPS NVFP4 inference (5x Blackwell) and supports rack-scale confidential computing across CPU, GPU, and NVLink.[2][4]
  • Rubin GPUs include second-generation RAS engine for proactive maintenance and modular cable-free trays enabling 18x faster assembly than Blackwell.[3]

🛠️ 技術深入

  • Rubin GPU: 336B transistors (TSMC N3), 288GB HBM4 (22 TB/s bandwidth), 50 PFLOPS FP4 inference (2.5x Blackwell), 3.6 TB/s NVLink 6 per GPU.[1][2]
  • Vera CPU: 88 Olympus cores (176 threads with spatial multi-threading), 1.5 TB LPDDR5X (1.2 TB/s bandwidth), 1.8 TB/s NVLink-C2C, 256 PCIe Gen6 lanes, full confidential computing.[1][2][3]
  • Vera Rubin NVL72 rack: 72 Rubin GPUs (20.7 TB total HBM4), 36 Vera CPUs (54 TB total LPDDR5X, 3,168 cores), 260 TB/s NVLink, 3.6 EFLOPS NVFP4 inference, 2.5 EFLOPS training.[1][2][5]
  • Additional components: ConnectX-9 SuperNIC, BlueField-4 DPU, third-generation Transformer Engine with adaptive compression, SHARP for 50% reduced network congestion.[2][3][4]

🔮 前景展望AI analysis grounded in cited sources

Rubin enables training of MoE models with 4x fewer GPUs than Blackwell
The platform's enhanced NVLink 6 and Transformer Engine optimizations reduce GPU requirements for massive-scale mixture-of-experts models.[4]
Vera Rubin NVL72 cuts AI inference cost per million tokens by 10x versus Blackwell
Combined advancements in NVFP4 compute, HBM4 memory, and efficiency deliver superior performance at lower operational costs.[4][5]
Rack-scale confidential computing protects proprietary trillion-parameter models
Third-generation confidential computing spans CPU, GPU, and NVLink domains, enabling secure training and inference without data exposure.[2][4]

時間線

2025-03
NVIDIA announces Blackwell platform as predecessor to Rubin.
2026-01
NVIDIA unveils Rubin platform, Vera CPU, and NVL72 specs at CES 2026 with full production plans.[1][7]
2026-02
Detailed Vera Rubin NVL72 specifications released, highlighting 10x perf/watt over Grace Blackwell.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: IT之家

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。