🟩較早收集於 2m

NVIDIA Run:ai GPU 分割提升 Token 吞吐量

NVIDIA Run:ai GPU 分割提升 Token 吞吐量
PostLinkedIn
🟩閱讀原文: NVIDIA Developer Blog
#gpu-fractioning#token-throughput#ai-schedulingnvidia-run:ai

💡Unlock massive AI token throughput via GPU fractioning in any environment with NVIDIA Run:ai.

⚡ 30-Second TL;DR

有什麼變化

引入 GPU 分割以實現 AI 工作負載的智慧排程

為什麼重要

使 AI 團隊最大化 GPU 使用率,降低成本並改善大型推論與訓練的 SLA。在多樣環境中民主化高性能 AI 運算存取。將 NVIDIA Run:ai 定位為企業 AI 基礎設施的關鍵工具。

下一步行動

Deploy GPU fractioning in your NVIDIA Run:ai cluster to test token throughput improvements on current AI workloads.

誰應關注:Developers & AI Engineers

關鍵要點

  • 引入 GPU 分割以實現 AI 工作負載的智慧排程
  • 實現大量 Token 吞吐量提升
  • 無縫支援雲端、NCP 和內部部署環境
  • NVIDIA 與 AI 合作夥伴聯合基準測試驗證效能
  • 解決擴展挑戰如延遲與效率

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • NVIDIA Run:ai enables GPU fractioning to allocate portions of GPU memory (1-100% or in MB/GB units) per device for workloads, supporting requests and optional limits for efficient resource provisioning across pods[1].
  • Improved fractional GPU support in v2.24 extends to multi-container pods, allowing explicit specification of containers via annotations, beyond just the first container[2].
  • GPU fractioning delivers high token throughput, with joint benchmarking showing massive gains in AI workloads across cloud, NCP, and on-premises environments[article].
  • Feature integrates with autoscaling for services like NIM, enabling dynamic scaling, partial GPU usage, and multi-node deployments for better efficiency[2].
  • Supports intelligent scheduling to address scaling challenges like latency, efficiency, and resource usage in shared GPU clusters[1][2][article].
📊 競品分析▸ Show
FeatureNVIDIA Run:ai GPU FractioningClarifai GPU Fractioning
Memory Allocation% of device, MB/GB per device; requests/limits per pod [1]Smart autoscaling with fractioning on GH200 [4]
Throughput GainsMassive token throughput; validated benchmarks [article]7.6× higher throughput vs H100 [4]
EnvironmentsCloud, NCP, on-premises [article]Cross-cloud orchestration [4]
Pricing/BenchmarksNot specified8× lower cost per token vs H100 [4]

🛠️ 技術深入

• Enable GPU fractioning in compute resources to set GPU devices per pod and memory per device (1-100%, MB, GB); request is minimum provisioned, limit is maximum (limit ≥ request to avoid OOM kills)[1]. • In v2.24, fractional GPUs assignable to specific containers in multi-container pods via annotations; default to first container[2]. • Works with DynamoGraphDeployment for inference workloads and NIM services supporting autoscaling, fractional GPUs, multi-node[2]. • Complements time-based fairshare scheduling for balanced GPU allocation over time windows[3].

🔮 前景展望AI analysis grounded in cited sources

NVIDIA Run:ai's GPU fractioning enhances AI workload scaling by improving GPU utilization, reducing waste in shared clusters, and enabling predictable performance for inference and training, potentially lowering costs and accelerating adoption in multi-tenant environments amid growing AI demands.

時間線

2023-12
Run:ai v2.20 introduces core workload scheduling and orchestration with initial GPU optimization[5]
2024-10
Run:ai v2.23 adds improved fractional GPU support for multi-container pods (beta) and Dynamo/NIM integration[2]
2025-01
Run:ai v2.24 releases with enhanced fractional GPU features, time-based fairshare, and global replica scaling[2][3]
2026-02
NVIDIA Developer Blog announces Run:ai GPU fractioning with token throughput boosts and cross-environment support[article]
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: NVIDIA Developer Blog

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。