Search

Few direct matches — filled in with the latest updates.

Tag: #cuda-optimization2 results

Luce DFlash Doubles Qwen3.6 Speed on 3090

Luce DFlash Doubles Qwen3.6 Speed on 3090

Luce DFlash is a new GGUF port of DFlash speculative decoding for Qwen3.6-27B, delivering up to 2x throughput on a single RTX 3090 using a standalone C++/CUDA stack on ggml. It supports 256K context with KV cache compression and sliding-window attention, benchmarked at 1.98x mean speedup on coding/math tasks. Deployment is simple via CMake build and Hugging Face downloads, with OpenAI-compatible serving.

Reddit r/LocalLLaMACommunityApr 27#speculative-decoding#local-llm#cuda-optimization