Search

直接匹配不多,已補上最新動態。

Tag: #error-analysis2 results

大型語言模型除錯系統化方法

大型語言模型除錯系統化方法

這篇論文提出一種系統化、模型無關的大型語言模型(LLM)除錯方法,將其視為可觀測系統。提供從問題偵測到模型精煉的結構化方法,統一評估、可解釋性和錯誤分析。即使缺乏標準化基準,此方法仍能實現弱點的迭代診斷、提示精煉及資料調整。

AI Fails Basic Arithmetic Despite Advanced Math Wins

AI Fails Basic Arithmetic Despite Advanced Math Wins

Frontier AI models excel in advanced math but consistently fail at multi-digit integer addition. Errors primarily stem from operand misalignment or carry failures, explaining most mistakes in top models like Claude, GPT, and Gemini. These issues link to tokenization and random carrying failures.

ArXiv AIResearchFeb 12#research#ai-rithmetic#v1
Ox Alpha:神秘模型公開亮相

Ox Alpha:神秘模型公開亮相

Ox Alpha 是一款透過 OpenRouter 與 OpenCode 提供的匿名模型,具備 100 萬 token 上下文、圖片與影片輸入、工具呼叫及免費使用等能力。早期測試顯示其推理與程式設計表現強勁,但開發者、參數量、訓練資料與正式基準排名仍未獲確認。