SourceStalecollected in 28m

DFlash achieves 85 tok/s on Apple Silicon

DFlash achieves 85 tok/s on Apple Silicon
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#speculative-decoding#apple-mlx#benchmarks#quantizationdflash-mlxdflashmlxqwen3.5apple-siliconm5-max

💡3.3x LLM speedup on M5 Max via new MLX DFlash

⚡ 30-Second TL;DR

What Changed

85 tok/s on Qwen3.5-9B, 3.3x vs baseline on M5 Max

Why It Matters

Dramatically boosts on-device LLM inference speeds on Apple hardware. Enables practical long-context generation locally. Shifts optimization focus for bandwidth-limited devices.

What To Do Next

Install MLX on Apple Silicon and monitor DFlash repo for open-source release.

Who should care:Researchers & Academics

Key Points

  • 85 tok/s on Qwen3.5-9B, 3.3x vs baseline on M5 Max
  • Optimizations: head_dim=256 patch, sync elision, packed QKV
  • 80-87% acceptance rate; better on 8bit than 4bit quantized
  • Bandwidth-bound insights for Apple unified memory
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.