Search

Few direct matches — filled in with the latest updates.

Tag: #error-analysis2 results

Systematic LLM Debugging Approach

Systematic LLM Debugging Approach

This paper introduces a systematic, model-agnostic approach for debugging large language models (LLMs), treating them as observable systems. It provides structured methods from issue detection to model refinement, unifying evaluation, interpretability, and error analysis. The methodology enables iterative diagnosis of weaknesses, prompt refinement, and data adaptation, even without standardized benchmarks.

AI Fails Basic Arithmetic Despite Advanced Math Wins

AI Fails Basic Arithmetic Despite Advanced Math Wins

Frontier AI models excel in advanced math but consistently fail at multi-digit integer addition. Errors primarily stem from operand misalignment or carry failures, explaining most mistakes in top models like Claude, GPT, and Gemini. These issues link to tokenization and random carrying failures.

ArXiv AIResearchFeb 12#research#ai-rithmetic#v1