Search

Tag: #nvidia21 results

DGX Spark 用於 vLLM 本地推論設定

DGX Spark 用於 vLLM 本地推論設定

一位使用者拆箱 NVIDIA DGX Spark,用於教育應用程式的本地 LLM 推論,搭配 vLLM、PyTorch 和 Hugging Face 模型。他們尋求最佳模型、統一記憶體 vLLM 調校建議,以及實際吞吐量表現。這是從雲端 GPU 轉向本地設定的首次嘗試。

Reddit r/LocalLLaMACommunityApr 15#unified-memory#local-inference#nvidia
DeepGEMM nv_dev_c491439 開發版本發布

DeepGEMM nv_dev_c491439 開發版本發布

DeepSeek 在 GitHub 上發布 DeepGEMM 的 nv_dev_c491439 開發更新。此版本無詳細發布說明或變更日誌。可能針對 NVIDIA GPU 的矩陣乘法進行 AI 工作負載優化。

DeepSeek (GitHub Releases: DeepGEMM)MediaApr 22#nvidia#gpu#gemm
Nvidia Rubin Speeds MoE Inference 10x Cheaper

Nvidia Rubin Speeds MoE Inference 10x Cheaper

Nvidia's Rubin platform features advanced NVLink interconnects to accelerate agentic AI, reasoning, and massive-scale MoE model inference at up to 10x lower cost per token. The article analogizes tech growth to a pyramid's limestone blocks, highlighting shifts from CPUs to GPUs and now efficient architectures. Groq complements this with ultra-fast inference to solve latency issues in real-time AI.

VentureBeatMediaFeb 15#launch#nvidia#rubin
Nvidia's DMS Slashes LLM Costs 8x

Nvidia's DMS Slashes LLM Costs 8x

Nvidia's DMS compresses LLM KV cache up to 8x, reducing memory costs without accuracy loss. Enables longer chain-of-thought reasoning and more parallel paths. Outperforms heuristic eviction and paging methods.

VentureBeatMediaFeb 12#research#nvidia#dms
Page 2 of 3