πŸ¦™Stalecollected in 6h

DFlash Doubles Qwen3.5-27B Speed on M5 Max

DFlash Doubles Qwen3.5-27B Speed on M5 Max
PostLinkedIn
πŸ¦™Read original on Reddit r/LocalLLaMA
#apple-silicon#speculative-decoding#mlxomlxomlxdfashqwen3.5-27bm5-max

πŸ’‘2x LLM speed on M5 Max: oMLX DFlash for Qwen3.5-27B BF16

⚑ 30-Second TL;DR

What Changed

DFlash in oMLX 0.3.5 RC1: 9β†’22 t/s for Qwen3.5-27B BF16

Why It Matters

Makes high-quality 27B models viable at interactive speeds on Apple hardware, boosting local deployment for Mac users.

What To Do Next

Download oMLX 0.3.5 RC1 from omlx.ai and test DFlash with Qwen3.5-27B on Mac.

Who should care:Developers & AI Engineers

Key Points

  • β€’DFlash in oMLX 0.3.5 RC1: 9β†’22 t/s for Qwen3.5-27B BF16
  • β€’Tested on M5 Max 128GB Apple Silicon
  • β€’Models: MLX-Qwopus3.5-27B-v3-bf16 + Qwen3.5-27B-DFlash
  • β€’GitHub: github.com/bstnxbt/dflash-mlx
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.