π¦Reddit r/LocalLLaMAβ’Stalecollected in 6h
DFlash Doubles Qwen3.5-27B Speed on M5 Max

#apple-silicon#speculative-decoding#mlxomlxomlxdfashqwen3.5-27bm5-max
π‘2x LLM speed on M5 Max: oMLX DFlash for Qwen3.5-27B BF16
β‘ 30-Second TL;DR
What Changed
DFlash in oMLX 0.3.5 RC1: 9β22 t/s for Qwen3.5-27B BF16
Why It Matters
Makes high-quality 27B models viable at interactive speeds on Apple hardware, boosting local deployment for Mac users.
What To Do Next
Download oMLX 0.3.5 RC1 from omlx.ai and test DFlash with Qwen3.5-27B on Mac.
Who should care:Developers & AI Engineers
Key Points
- β’DFlash in oMLX 0.3.5 RC1: 9β22 t/s for Qwen3.5-27B BF16
- β’Tested on M5 Max 128GB Apple Silicon
- β’Models: MLX-Qwopus3.5-27B-v3-bf16 + Qwen3.5-27B-DFlash
- β’GitHub: github.com/bstnxbt/dflash-mlx
π°
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.