
Gemma 4 124B MoE 開放發佈傳聞
Reddit 上根據 Jeff Dean 刪除的推文猜測 Gemma 4 124B MoE 可能超越 Gemini 3 Flash-Lite。社群期待 Google 開放此更大模型變體。最近發佈後,更多驚喜預期。
Tag: #moe72 results

Reddit 上根據 Jeff Dean 刪除的推文猜測 Gemma 4 124B MoE 可能超越 Gemini 3 Flash-Lite。社群期待 Google 開放此更大模型變體。最近發佈後,更多驚喜預期。
用戶報告 Qwen3.5-122B-A10B 在 Q4 表現良好,但 Q3_K_M 或 UD_Q2_K_XL 急劇退化,毀壞程式碼庫,儘管速度快。它處理工具呼叫和語法但決策失敗。分享警示他人量化 >100B 大模型。
使用者偏好 35B A3B MoE 在 16GB VRAM/64GB RAM 上達 49-55 t/s,勝過 9B 的 23 t/s。對 35B 效能愛不釋手。提及 9B 傳聞勝過 120B OSS 模型。

ArtificialAnalysis.ai 在智慧指數、程式設計指數與代理指數上,將 Qwen3.5-27B 評為高於 Qwen3.5-122B-A10B 與 Qwen3.5-35B-A3B。小型模型在所有類別勝過更大 MoE 兄弟模型。

r/LocalLLaMA 的 Reddit 貼文推測 Opus 模型擁有 0.5T 參數 × 10,等於約 5T 總參數。未提供額外細節或證據。引發社群對模型規模的興趣。
Reddit 貼文質疑在 MoE 模型如 Qwen3-30B-A3B 中將專家擴增至 A6B 是否顯著提升效能。A6B 配置過去短暫熱門,現已鮮少實驗。可輕鬆在 Llama.cpp 中測試。
Reddit 用戶因辦公室頻寬限制,尋求 Qwen 3.5 MoE 35B 不含推理鏈的指令模式測試結果。對 Qwen 在 2507 發布後重返混合推理模型感到驚訝。社群討論持續中。

Nvidia's Rubin platform features advanced NVLink interconnects to accelerate agentic AI, reasoning, and massive-scale MoE model inference at up to 10x lower cost per token. The article analogizes tech growth to a pyramid's limestone blocks, highlighting shifts from CPUs to GPUs and now efficient architectures. Groq complements this with ultra-fast inference to solve latency issues in real-time AI.

JD.com open-sourced JoyAI-LLM-Flash, a 48B total parameter MoE model with 3B active params, pre-trained on 20T tokens. It excels in knowledge, reasoning, coding, and agents using FiberPO framework and Muon optimizer. Features 1.3x-1.7x throughput gains via dense MTP.