Search

Tag: #moe72 results

Zyphra 推出高效 ZAYA1-8B MoE 模型

Zyphra 推出高效 ZAYA1-8B MoE 模型

Zyphra 發布 ZAYA1-8B,這是一個開放的 8B 參數混合專家推理模型,在 AMD Instinct MI300 GPU 上訓練,在基準測試中與 GPT-5-High 和 DeepSeek-V3.2 競爭。Apache 2.0 授權可在 Hugging Face 免費下載並自訂。它具備創新的 MoE++ 架構,提供優異效率。

VentureBeatMediaMay 7#moe#amd-gpu#reasoning-model
Qwen3.6-35B-A3B 開源 MoE 模型發布

Qwen3.6-35B-A3B 開源 MoE 模型發布

Qwen3.6-35B-A3B 是全新開源稀疏 MoE 模型,總參數 35B,活躍參數 3B,採用 Apache 2.0 許可。代理編碼能力媲美活躍參數大 10 倍的模型,並具強大多模態感知與思考/非思考模式。可於 HuggingFace、ModelScope 及 Qwen Studio 下載。

Reddit r/LocalLLaMACommunityApr 16#moe#multimodal#agentic-coding
🤖

GigaChat 3.1 Ultra 702B 與 Lightning 10B 發布

GigaChat 在 Hugging Face 以 MIT 授權發布 GigaChat-3.1-Ultra(702B MoE)與 GigaChat-3.1-Lightning(10B MoE)開放權重。從頭預訓練,在基準測試超越 DeepSeek-V3 與 Qwen3,優化英文/俄文及工具呼叫。Ultra 適合高資源環境;Lightning 擅長本地推理,支援 256k 上下文。

Reddit r/LocalLLaMACommunityMar 24#open-weights#moe#multilingual
Page 2 of 8