🦙Freshcollected in 3h

llama.cpp May Close MLX’s M5 Prefill Advantage

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#apple-silicon#metal-acceleration#inference-benchmarkllama.cpp-metalllama.cppmlxqwen3.8apple m5

💡See whether llama.cpp now matches MLX for Qwen3.8 prefill on Apple M5.

⚡ 30-Second TL;DR

What Changed

The reported setup reaches roughly 300–350 tokens per second for prefill with Unsloth Q8 GGUF.

Why It Matters

If independently confirmed, the change could simplify Mac inference stacks by making GGUF viable for both prompt processing and generation. However, practitioners should treat the report as an observation rather than a universal benchmark result because it focuses on one model, quantization range, and Apple chip generation.

What To Do Next

Run identical prompt-length and generation-length tests for llama.cpp Metal GGUF and MLX on your M5 hardware before removing MLX from your deployment stack.

Who should care:Developers & AI Engineers

Key Points

  • The reported setup reaches roughly 300–350 tokens per second for prefill with Unsloth Q8 GGUF.
  • Generation performance is approximately 19 tokens per second on an M5 Pro with Qwen3.8 27B.
  • llama.cpp Metal appears to benefit from M5 matmul or neural accelerator support.
  • The comparison remains workload- and model-specific, and mainstream MLX MTP support is still an open question.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.