Search

Tag: #moe72 results

🤖

家用實驗室整合至 122B Qwen MoE

使用者在 Strix Halo 上基準測試並將家用實驗室從三個模型整合至單一 122B Qwen3.5 MoE,達到 27.4 tok/s 並獲最高分數。IQ3 量化在半 VRAM 下性能匹配 Q4_K_M。同時處理電子郵件、財務和攝像頭等多個應用。

Reddit r/LocalLLaMACommunityMar 27#homelab#moe#quantization
⚙️

M5 Max 128GB LLM 基準測試 v2

Apple M5 Max 128GB 的後續基準測試顯示優異提示處理速度,特別是 MoE 模型如 Qwen3.5-35B-A3B 達 2,845 tok/s。Token 生成排名強調 MoE 優於密集模型。測試使用 llama-bench 並比較 llama.cpp 與 MLX。

Reddit r/LocalLLaMACommunityMar 22#benchmarks#moe#apple-silicon
Page 6 of 8