🦙Reddit r/LocalLLaMA•較早收集於 6h
RTX PRO 6000 MoE 基準測試:最高 50.5 tok/s
#moe#blackwell#benchmarkqwen3.5-397bqwen3.5-397brtx-pro-6000marlincutlass
💡397B MoE 實測 50.5 tok/s + NVIDIA SM120 錯誤曝光 (22字)
⚡ 30-Second TL;DR
有什麼變化
Marlin TP=4 持續達 50.5 tok/s—SM120 硬體最佳。
為什麼重要
凸顯 NVIDIA 在工作站 Blackwell 的驗證缺口,限制 FP4 吞吐量。引導推論最佳化偏向 Marlin,而非故障 CUTLASS 用於消費 GPU。
下一步行動
在你的 SM120 設定基準測試 Marlin TP=4,再試 CUTLASS 或 MTP。
誰應關注:Developers & AI Engineers
關鍵要點
- •Marlin TP=4 持續達 50.5 tok/s—SM120 硬體最佳。
- •CUTLASS grouped GEMM 在 RTX PRO 6000 (SM120) 失敗,提交 bug #3096。
- •MTP=2 使 Marlin 從 50.5 降至 39.6 tok/s,因去量化不符。
- •無 NVLink PCIe 限制專家並行僅 1.4-2.6 tok/s。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 9 個來源。
🔑 增強重點摘要
🛠️ 技術深入
🔮 前景展望AI analysis grounded in cited sources
⏳ 時間線
2025-03
NVIDIA 發布 Blackwell 架構,RTX PRO 6000 亮相為工作站頂級 GPU
2025-10
CloudRift 發布 RTX PRO 6000 單 GPU 30B AWQ 基準,達 8,400 tok/s
2026-01
GamersNexus 與 Trooper.ai 比較 RTX PRO 6000 對 RTX 5090/4090 LLM 效能
2026-02
CloudRift 基準顯示 RTX PRO 6000 單 GPU 超越 H100 推論吞吐量
2026-03
Reddit r/LocalLLaMA 發布 Qwen3.5-397B 在 4x RTX PRO 6000 上 50.5 tok/s MoE 測試
📎 來源 (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- spheron.network — Rent Nvidia Rtx Pro 6000
- gamersnexus.net — Nvidia Rtx Pro 6000 Blackwell Benchmarks Tear Down Thermals Gaming LLM Acoustic Tests
- cloudrift.ai — Benchmarking Rtx6000 vs Datacenter Gpus
- trooper.ai — Benchmarks Compare Rtx 5090 with Rtx Pro 6000 Blackwell
- acecloud.ai — Rtx Pro 6000 Blackwell Rendering Workstation Builds
- youtube.com — Watch
- vast.ai — Which Nvidia Rtx 6000 Is Right for You
- yottalabs.ai — Which Nvidia Rtx 6000 GPU Is Right for You in 2026
- databasemart.com — Vllm GPU Benchmark Pro6000
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA ↗
每週 AI 簡報
每週一封,可隨時退訂。