Mac Mini M4: 34 tok/s on 20B LLM
๐กMac Mini M4 benchmark: 34 t/s on 20B Q4 w/ 26k ctx via OpenClaw/LM Studio
โก 30-Second TL;DR
What Changed
Model: unsloth gpt-oss-20b-Q4_K_S.gguf, context 26035 tokens.
Why It Matters
Proves Mac Mini M4 viable for fast local inference on 20B models, aiding desktop AI devs. Highlights OpenClaw/LM Studio combo for high perf without discrete GPU.
What To Do Next
Benchmark OpenClaw on your Mac Mini M4 with gpt-oss-20b-Q4_K_S.gguf.
Key Points
- โขModel: unsloth gpt-oss-20b-Q4_K_S.gguf, context 26035 tokens.
- โขPerformance: 34 tok/s decode, 0.7s TTFT post-first prompt.
- โขSetup: OpenClaw 2026.3.8, LM Studio 0.4.6+1, GPU offload=18, flash attention=on.
๐ง Deep Insight
Background and context from public sources โ not the original article. 5 sources cited.
๐ Enhanced Key Takeaways
- โขMac Mini M4 base model (10 CPU cores, 10 GPU cores) achieves Geekbench 6 single-core score of 3788 and multi-core of 14696, positioning it strongly among 2024-2025 Macs[3].
- โขM4 Neural Engine delivers 38 trillion operations per second, a 3x improvement over M1 and outperforming Intel's 14th-gen i9 in AI tasks[2].
- โขMac Mini M4 shows significant performance degradation with model sizes: 77.1 tok/s on 1B models dropping to 17.7 tok/s on 8B and 9.6 tok/s on 14B models[1].
- โขM4 memory bandwidth reaches 120 GB/s, a 20% boost over M2, aiding heavy data applications like video editing and AI[2].
๐ Competitor Analysisโธ Show
| Device | Model Size | Gen Speed (tok/s) | TTFT (ms) | Prompt Proc (tok/s) |
|---|---|---|---|---|
| Mac Studio | 1B | 178 | 203 | 5719 |
| Mac Mini M4 | 1B | 77.1 | 1180 | 1111 |
| Mac Studio | 8B | 62.7 | 1060 | 1119 |
| Mac Mini M4 | 8B | 17.7 | 6850 | 186 |
| Mac Studio | 14B | 35.8 | 2040 | 583 |
| Mac Mini M4 | 14B | 9.6 | 13300 | 96 |
๐ ๏ธ Technical Deep Dive
- โขM4 chip: 10 CPU cores at 4.4 GHz, 10 GPU cores; Neural Engine at 38 TOPS (trillion operations per second)[2][3].
- โขMemory bandwidth: 120 GB/s, 20% higher than M2, optimized for AI and matrix multiplication tasks[2].
- โขGeekbench 6 scores: Single-core 3788, Multi-core 14696 for base M4 Mac Mini[3].
- โขPerformance scales poorly with model size due to unified memory limits in base 32GB config, evident in 77% drop from 1B to 8B models[1].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.