Search

Tag: #cuda17 results

RotorQuant:比 TurboQuant 快 10-19 倍

RotorQuant:比 TurboQuant 快 10-19 倍

RotorQuant 利用 Clifford 轉子進行向量量化,比 TurboQuant 快 10-19 倍,參數僅 1/44。它在 Qwen2.5-3B KV 快取上餘弦相似度近乎相同(0.990),並在融合 CUDA/Metal 核心中表現出色。程式碼與論文現已開放。

Reddit r/LocalLLaMACommunityMar 26#quantization#kv-cache#cuda
在 NVIDIA cuTile 中調優 Flash Attention 達峰值效能

在 NVIDIA cuTile 中調優 Flash Attention 達峰值效能

NVIDIA 開發者部落格詳細說明使用 cuTile Python 實作 Flash Attention,這是現代 AI 的關鍵工作負載,以達峰值效能。它涵蓋透過 quickstart 文件的環境設定,以及 transformer 模型核心的 attention 機制。從業者可學習優化 token 處理的技巧。

NVIDIA Developer BlogOfficialMar 4#transformers#cuda
Page 1 of 2