SourceStalecollected in 27m

TurboQuant Skips 90% KV Dequant for +22.8% Speed

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#kv-cache#apple-siliconturboquantllama.cppturboquantqwen3.5-35b

💡22.8% faster 32K decode via KV sparsity – must-try for local inference tuning

⚡ 30-Second TL;DR

What Changed

+22.8% decode speed at 32K on Qwen3.5-35B-A3B (M5 Max)

Why It Matters

Dramatically improves long-context inference efficiency on Apple Silicon, benefiting local LLM deployments.

What To Do Next

Clone github.com/TheTom/turboquant_plus and test sparse V dequant on your 32K setups.

Who should care:Developers & AI Engineers

Key Points

  • +22.8% decode speed at 32K on Qwen3.5-35B-A3B (M5 Max)
  • Skips 90% V dequant using attention sparsity
  • Repo: github.com/TheTom/turboquant_plus
  • Works on standard q8_0 KV (+5% speed)
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.