๐Ÿฆ™Stalecollected in 56m

Mac Mini M4: 34 tok/s on 20B LLM

PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#local-llm#apple-silicon#benchmarksopenclawopenclawlm-studiounslothgpt-oss-20b

๐Ÿ’กMac Mini M4 benchmark: 34 t/s on 20B Q4 w/ 26k ctx via OpenClaw/LM Studio

โšก 30-Second TL;DR

What Changed

Model: unsloth gpt-oss-20b-Q4_K_S.gguf, context 26035 tokens.

Why It Matters

Proves Mac Mini M4 viable for fast local inference on 20B models, aiding desktop AI devs. Highlights OpenClaw/LM Studio combo for high perf without discrete GPU.

What To Do Next

Benchmark OpenClaw on your Mac Mini M4 with gpt-oss-20b-Q4_K_S.gguf.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขModel: unsloth gpt-oss-20b-Q4_K_S.gguf, context 26035 tokens.
  • โ€ขPerformance: 34 tok/s decode, 0.7s TTFT post-first prompt.
  • โ€ขSetup: OpenClaw 2026.3.8, LM Studio 0.4.6+1, GPU offload=18, flash attention=on.

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 5 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขMac Mini M4 base model (10 CPU cores, 10 GPU cores) achieves Geekbench 6 single-core score of 3788 and multi-core of 14696, positioning it strongly among 2024-2025 Macs[3].
  • โ€ขM4 Neural Engine delivers 38 trillion operations per second, a 3x improvement over M1 and outperforming Intel's 14th-gen i9 in AI tasks[2].
  • โ€ขMac Mini M4 shows significant performance degradation with model sizes: 77.1 tok/s on 1B models dropping to 17.7 tok/s on 8B and 9.6 tok/s on 14B models[1].
  • โ€ขM4 memory bandwidth reaches 120 GB/s, a 20% boost over M2, aiding heavy data applications like video editing and AI[2].
๐Ÿ“Š Competitor Analysisโ–ธ Show
DeviceModel SizeGen Speed (tok/s)TTFT (ms)Prompt Proc (tok/s)
Mac Studio1B1782035719
Mac Mini M41B77.111801111
Mac Studio8B62.710601119
Mac Mini M48B17.76850186
Mac Studio14B35.82040583
Mac Mini M414B9.61330096

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขM4 chip: 10 CPU cores at 4.4 GHz, 10 GPU cores; Neural Engine at 38 TOPS (trillion operations per second)[2][3].
  • โ€ขMemory bandwidth: 120 GB/s, 20% higher than M2, optimized for AI and matrix multiplication tasks[2].
  • โ€ขGeekbench 6 scores: Single-core 3788, Multi-core 14696 for base M4 Mac Mini[3].
  • โ€ขPerformance scales poorly with model size due to unified memory limits in base 32GB config, evident in 77% drop from 1B to 8B models[1].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

M5 Mac Mini will offer 15% single-core and 25% multi-core Geekbench gains over M4
Leaked benchmarks indicate M5 single-core over 4300 vs M4's 3800, with multi-core jumping significantly for AI workloads[4].
M4 Mac Mini enables smooth 4K video editing and local AI on base model
Reviews confirm lag-free 4K editing and ML tasks on base M4 Mac Mini with 32GB RAM[5].

โณ Timeline

2024-10
Apple releases Mac Mini M4 with 10-core CPU/GPU and 38 TOPS Neural Engine
2024-11
Early benchmarks compare Mac Mini M4 vs Mac Studio local AI performance
2025-11
Mac Studio vs Mac Mini M4 AI benchmarks highlight scaling issues on larger models
2026-03
Reddit benchmarks show 34 tok/s on 20B LLM with OpenClaw on Mac Mini M4 32GB
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.