Search

Tag: #speculative-decoding32 results

Luce DFlash:在 RTX 3090 上 Qwen3.6 速度翻倍

Luce DFlash:在 RTX 3090 上 Qwen3.6 速度翻倍

Luce DFlash 是 DFlash 推測解碼的 GGUF 移植版,針對 Qwen3.6-27B,在單張 RTX 3090 上實現高達 2 倍吞吐量,使用基於 ggml 的獨立 C++/CUDA 堆疊。支援 256K 上下文,具 KV 快取壓縮與滑動視窗注意力,在程式碼/數學任務上基準測試達 1.98 倍平均加速。部署簡單,只需 CMake 建置與 Hugging Face 下載,並提供 OpenAI 相容伺服。

Reddit r/LocalLLaMACommunityApr 27#speculative-decoding#local-llm#cuda-optimization
⚙️

推測解碼儲存庫上線

新 GitHub 儲存庫從頭實作推測解碼方法:EAGLE-3、Medusa-1、PARD、草稿模型、n-gram、尾綴解碼。針對 Qwen2.5-7B-Instruct 提供共享訓練/推論,基準測試澄清提案者品質對驗證者成本及吞吐量細微差異。定位為演算法-系統教育資源。

Reddit r/MachineLearningCommunityApr 26#speculative-decoding#inference#benchmarks
Page 2 of 4