SourceReddit r/LocalLLaMA•Stalecollected in 28m
DFlash achieves 85 tok/s on Apple Silicon

#speculative-decoding#apple-mlx#benchmarks#quantizationdflash-mlxdflashmlxqwen3.5apple-siliconm5-max
💡3.3x LLM speedup on M5 Max via new MLX DFlash
⚡ 30-Second TL;DR
What Changed
85 tok/s on Qwen3.5-9B, 3.3x vs baseline on M5 Max
Why It Matters
Dramatically boosts on-device LLM inference speeds on Apple hardware. Enables practical long-context generation locally. Shifts optimization focus for bandwidth-limited devices.
What To Do Next
Install MLX on Apple Silicon and monitor DFlash repo for open-source release.
Who should care:Researchers & Academics
Key Points
- •85 tok/s on Qwen3.5-9B, 3.3x vs baseline on M5 Max
- •Optimizations: head_dim=256 patch, sync elision, packed QKV
- •80-87% acceptance rate; better on 8bit than 4bit quantized
- •Bandwidth-bound insights for Apple unified memory
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.