SourceStalecollected in 4h

Gemma 4 Excels Over Qwen in Local Tests

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#local-inference#benchmark#quantizationgemma-4gemma-4qwen-3.5llama.cppmlx-vlm

💡Gemma 4 crushes Qwen locally: speed + smarts for your Mac Studio setup

⚡ 30-Second TL;DR

What Changed

Gemma 26b a4b: ~1000pp, ~60tg at 20k context on Mac Studio

Why It Matters

Positions Gemma 4 as top open-weight option for local inference, potentially drawing users from Qwen due to better usability and coherence. KV cache issues may limit long-context apps until fixes.

What To Do Next

Benchmark Gemma 4 26b Q4_K_XL vs Qwen3.5 on Mac Studio using llama.cpp at 20k context.

Who should care:Developers & AI Engineers

Key Points

  • Gemma 26b a4b: ~1000pp, ~60tg at 20k context on Mac Studio
  • Concise, coherent CoT vs Qwen's looping and gaslighting
  • Strong visual understanding and multilingual performance
  • Large KV cache without optimizations; mlx-vlm prompt caching fails for Qwen
  • Censorship heavy in e4b variant

📰 Event Coverage

📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.