🦙Reddit r/LocalLLaMA•較早收集於 59m
Nemotron 3 Super自身推理循環
#model-bug#reasoning-loop#llama-cppnemotron-3-supernemotron-3-superllama-serveraider
💡修復Nemotron缺陷:自身推理變無限循環!
⚡ 30-Second TL;DR
有什麼變化
模型將自身推理鏈視為新用戶訊息
為什麼重要
暴露進階推理模型在本地推論的潛在缺陷。可能影響使用Nemotron的代理工作流程可靠性。
下一步行動
在llama-server執行Nemotron 3 Super,無--special旗標隔離推理循環。
誰應關注:Developers & AI Engineers
關鍵要點
- •模型將自身推理鏈視為新用戶訊息
- •後端:llama-server旗標--special及--jinja
- •客戶端:Aider;產生8k token循環元分析
- •其他模型未見的獨特問題
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 6 個來源。
🔑 增強重點摘要
🛠️ 技術深入
- •架構:混合 Mamba-Transformer 主幹,採用 LatentMoE(令牌壓縮至潛在維度路由至專家),總參數 120.6B,活躍參數 12.7B[2][3][4]。
- •多令牌預測 (MTP):單次前向傳遞預測多個未來令牌,提升生成速度並內建推測解碼[2][3][4]。
- •NVFP4 預訓練:使用 NVIDIA Blackwell 優化,低精度訓練減少記憶體需求,在 B200 GPU 上比 H100 FP8 快 4×[2][4]。
- •推理配置:最大模型長度 65,536,支援 FLASH_ATTN 後端,編譯器通過 fuse_allreduce_rms 和 eliminate_noops[1]。
- •後訓練:25 兆令牌預訓練 + SFT + 多環境 RL(21 配置,超過 120 萬 rollout)[2][3]。
🔮 前景展望AI analysis grounded in cited sources
⏳ 時間線
2025-12
發布 Nemotron 3 Nano,為 Super 模型家族先驅
2026-03
發布 Nemotron 3 Super,引入 LatentMoE 與 MTP 創新架構
📎 來源 (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- artificialanalysis.ai — Nvidia Nemotron 3 Super the New Leader in Open Efficient Intelligence
- developer.nvidia.com — Introducing Nemotron 3 Super an Open Hybrid Mamba Transformer Moe for Agentic Reasoning
- research.nvidia.com — Nvidia Nemotron 3 Super Technical Report
- Hugging Face — Nvidia Nemotron 3 Super 120b A12b Bf16
- greptile.com — Nvidia Nemotron Super in Code Review
- aiagentsdirectory.com — Nvidia Nemotron 3 Super Everything You Need to Know
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA ↗
每週 AI 簡報
每週一封,可隨時退訂。
