🦙Reddit r/LocalLLaMA•較早收集於 56m
Mac Mini M4 32GB:20B LLM 達 34 tok/s
#local-llm#apple-silicon#benchmarksopenclawopenclawlm-studiounslothgpt-oss-20b
💡Mac Mini M4 基準:OpenClaw/LM Studio 於 20B Q4 達 34 t/s、26k 脈絡
⚡ 30-Second TL;DR
有什麼變化
模型:unsloth gpt-oss-20b-Q4_K_S.gguf,脈絡 26035 令牌。
為什麼重要
證明 Mac Mini M4 適合 20B 模型快速本地推論,助桌面 AI 開發者。凸顯 OpenClaw/LM Studio 組合無需獨立 GPU 即高性能。
下一步行動
在你的 Mac Mini M4 上以 gpt-oss-20b-Q4_K_S.gguf 基準測試 OpenClaw。
誰應關注:Developers & AI Engineers
關鍵要點
- •模型:unsloth gpt-oss-20b-Q4_K_S.gguf,脈絡 26035 令牌。
- •效能:34 tok/s 解碼,首提示後 0.7s TTFT。
- •設定:OpenClaw 2026.3.8、LM Studio 0.4.6+1、GPU 卸載=18、flash attention=開。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 9 個來源。
🔑 增強重點摘要
🛠️ 技術深入
- •模型採用原生 MXFP4 量化於 MoE 層,使 20B 版本可在 16GB 記憶體內運行,並具備代理功能如函數呼叫、網頁瀏覽與 Python 程式碼執行[4]。
- •推理等級設定:低等級優化快速回應、一般對話;中等級平衡效能;高等級提供深度分析但延遲較高,可經系統提示如 'Reasoning: high' 指定[3][4]。
- •Unsloth Dynamic 2.0 GGUF 量化優化,Q4_K_S 適合消費者硬體,檔案大小 11.6 GB,推薦總可用記憶體(統一記憶體 + VRAM + RAM)超過模型大小以達最快速度[3][6]。
- •微調支援:QLoRA 訓練全 20B 模型僅需 14GB VRAM,比傳統方法快 1.5 倍且 VRAM 消耗少 70%,但 GGUF 匯出需 LoRA bf16 權重導致更高 VRAM 需求[1]。
🔮 前景展望AI analysis grounded in cited sources
⏳ 時間線
2026-03
unsloth 發布 gpt-oss-20b-GGUF 量化模型系列
2026-03
Mac Mini M4 32GB 基準測試達 20B LLM 34 tok/s
📎 來源 (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- skywork.ai — Unsloth Gpt Oss 20b Gguf Free Chat Online
- artificialanalysis.ai — Gpt Oss 20b vs Grok 4
- unsloth.ai — Gpt Oss How to Run and Fine Tune
- Hugging Face — Gpt Oss 20b Gguf
- latent.space — Ainews the High Return Activity of
- Hugging Face — Gpt Oss 20b Q4 K S
- datarobot.com — Testing Gpt Oss Models
- unsloth.ai — Qwen3
- GitHub — 15396
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA ↗
每週 AI 簡報
每週一封,可隨時退訂。