Search

Tag: #optimization77 results

8GB VRAM 本地微調 Gemma 4

8GB VRAM 本地微調 Gemma 4

Unsloth 支援在 8GB VRAM 上本地微調 Gemma 4 E2B/E4B,比 FlashAttention-2 快 1.5 倍且 VRAM 少 60%。修復梯度累積損失爆炸、26B/31B 索引錯誤、use_cache 問題及 float16 音頻溢位等 bug。提供文字、視覺、音頻的免費 Colab 筆記本及 Unsloth Studio UI。

Reddit r/LocalLLaMACommunityApr 7#fine-tuning#optimization#notebooks
Google TurboQuant 加速 AI 推論

Google TurboQuant 加速 AI 推論

Google 的 TurboQuant 壓縮 LLM 推論中的 KV 快取及向量搜尋,在 Nvidia H100 上對 Gemma/Mistral 模型實現 6 倍記憶體節省及 8 倍加速。它解決關鍵瓶頸,提升硬體效率而不損準確度。分析師強調擴展潛力,但指出可能增加使用而非節省。

ComputerworldMediaMar 26#kv-cache#inference#optimization
Page 2 of 8