Latest AI Benchmarks News & Updates

Arena rankings, eval controversies and what benchmarks actually measure. Read the fine print behind every "SOTA" claim.

135 articles

💰
钛媒体6h ago

Tool-Free Tests Expose the Opus 5–GPT-5.6 Gap

A tool-free evaluation compares Opus 5 with GPT-5.6 and argues that their capability gap becomes clearer without external tools. The article also suggests that stronger native reasoning could make traditional prompt engineering less important.

📄
ArXiv AI11h ago

A New Map for Diagnosing Agent Failures

This research introduces an interaction-centric taxonomy that localizes agent failures to interactions between models, harnesses, users, tools, memory, and environments. Its 41 failure modes indicate whether remediation belongs in model post-training, harness engineering, environment redesign, or benchmark repair.

📄
ArXiv AIYesterday

Benchmarking Autonomous AI Scientists with Multi-Model Peer Review

This study introduces a rigorous benchmarking protocol for autonomous AI research systems using multi-model LLM peer review. It evaluates four frameworks against commercial benchmarks, revealing significant performance gaps and validating the reliability of automated evaluation.

📄
ArXiv AIYesterday

LSR-Synth Tests Whether AI Discovers or Recalls Equations

This study evaluates whether LSR-Synth can distinguish genuine symbolic discovery from equation memorization and conventional operator search. It finds that a fixed, semantics-free vocabulary solves most current tasks, while language-model-generated candidates add substantial value mainly when parts of that vocabulary are deliberately removed.

🤖
Reddit r/MachineLearning18d ago

Schema harness achieves 99% on ARC-AGI-3 benchmark

The new 'Schema' harness improves performance on the ARC-AGI-3 benchmark by optimizing the interaction process between models and game environments. It uses Claude Opus 4.8 and Fable 5 to achieve 99% accuracy without modifying model weights.

🐯
虎嗅23d ago

GPT-5.6 outperforms competitors in complex 3D generation tasks

A comparative test of 63 complex prompts shows GPT-5.6 Sol Ultra significantly outperforming other models in generating detailed, interactive 3D web environments.

🌍
The Next Web (TNW)27d ago

AI Hacking Capabilities Outpace Current Safety Benchmarks

Current safety benchmarks are failing to accurately measure the evolving hacking capabilities of frontier AI models. This leaves security teams and regulators without reliable tools to assess the risks posed by advanced systems.

🐯
虎嗅29d ago

New benchmark MemoBench exposes world model memory gaps

MemoBench is a new diagnostic benchmark designed to test world models' ability to maintain object permanence and state evolution in dynamic environments. It reveals that current video generation models struggle to accurately track and update object states during occlusion.

🇭🇰
SCMP Technology36d ago

Zhipu AI’s GLM-5.2 model excels in cybersecurity tasks

Beijing-based Zhipu AI has released GLM-5.2, a model that demonstrates competitive performance against Anthropic’s Claude Opus 4.8 in cybersecurity bug-hunting tasks. This development highlights the narrowing gap between Chinese AI models and top-tier US counterparts.

🐯
虎嗅41d ago

Humans outperform AI in rigorous mathematical research testing

The First Proof project tested four AI systems against original, unpublished research-level math problems, finding that while AI shows promise, it still lags behind top human researchers.

💼
VentureBeat48d ago

Weibo's 3B Model Challenges AI Scaling Laws on Benchmarks

Sina Weibo researchers introduced VibeThinker-3B, a 3-billion parameter model that allegedly matches or exceeds the reasoning performance of massive flagship models. The release has sparked intense debate within the AI community regarding the validity of current benchmarks.

💼
VentureBeat54d ago

GPT-5.5 Outperforms Claude Fable 5 on New ALE Benchmark

UC Berkeley researchers launched the Agents’ Last Exam (ALE), a rigorous benchmark designed to evaluate AI agents on real-world, long-horizon professional workflows. GPT-5.5 achieved the top score, outperforming Anthropic's Claude Fable 5 in a test that emphasizes practical labor impact over static coding puzzles.

🌍
The Next Web (TNW)60d ago

Spirit AI beats Nvidia on robotics benchmark

Chinese startup Spirit AI has surpassed Nvidia's robotics models on the RoboArena leaderboard. Their Spirit v1.6 model achieved a score of 1,924, outperforming Nvidia's Cosmos3-Nano-Policy.

💼
VentureBeat69d ago

DeepSWE Benchmark Challenges AI Coding Leaderboards and GPT-5.5 Supremacy

Datacurve launched DeepSWE, a new coding benchmark that reveals significant performance gaps between frontier models, crowning GPT-5.5 as the leader. The report also exposes a 32% error rate in existing SWE-Bench Pro evaluations, questioning the reliability of current industry standards.

⚛️
量子位74d ago

Fei-Fei Li unveils ImageNet for spatial intelligence

AI pioneer Fei-Fei Li has introduced a new benchmark designed to evaluate embodied spatial intelligence. This initiative aims to provide a standardized dataset and evaluation framework for the next generation of robotic and spatial AI systems.

📄
ArXiv AI82d ago

BenchJack: Automating Red-Teaming to Expose AI Benchmark Flaws

BenchJack is an automated red-teaming system designed to identify reward-hacking exploits in AI agent benchmarks. By applying it to 10 major benchmarks, researchers discovered 219 flaws, demonstrating that many current evaluation metrics are easily gamed without actual task completion.

🦙
Reddit r/LocalLLaMA141d ago

Hard-to-Fake Coding Benchmark Exposes LLM Limits

Researchers developed EsoLang-Bench using esoteric languages like Brainfuck and Befunge-98 to test genuine coding reasoning beyond pattern matching. Top models including GPT-5.2, O4-mini, and Gemini achieved a best score of 11.2%, with medium/hard problems at 0%. The benchmark highlights the need for ungameable evaluations and includes a website and paper.

🦙
Reddit r/LocalLLaMA18d ago

Kimi K3 Ranks 3rd on ArtificialAnalysis, Surpassing Claude Opus

The Kimi K3 model has achieved a 3rd place ranking on the ArtificialAnalysis leaderboard. It notably outperformed the Claude Opus 4.8 model in this benchmark assessment.

🐼
Pandaily19d ago

Moonshot AI's Kimi K3 Challenges Fable 5 in Video Generation

Moonshot AI's new Kimi K3 model has appeared on benchmark platforms, demonstrating video generation capabilities that rival Fable 5. Early social media tests highlight significant improvements in animation smoothness and visual detail.

🤖
Reddit r/MachineLearning19d ago

Papers with Code Launches Dedicated Robotics Benchmark Page

Papers with Code has introduced a dedicated Robotics section to track major benchmarks, trending papers, and open-source artifacts. It currently features over 110 entries across benchmarks like LIBERO and SimplerEnv.

🤖
Reddit r/MachineLearning20d ago

New Benchmark for Open-Ended Multi-Agent LLM Coordination

研究人員發布了 ALEM 基準測試,旨在評估 LLM 代理在複雜、開放式環境中的協作能力。測試結果顯示,儘管大多數模型表現不佳,但 Gemini 3.1 Pro 在高難度任務中展現了與專業訓練模型相當的潛力。

📄
ArXiv AI22d ago

New Benchmark Tests AI Agents on Long-Horizon Terminal Tasks

Long-Horizon-Terminal-Bench introduces a new evaluation framework for AI agents using 46 complex, multi-step tasks that require long-term planning and iterative debugging. The benchmark provides dense intermediate rewards to better measure progress in open-ended workflows compared to traditional outcome-only evaluations.

Vercel News22d ago

AI Gateway leaderboards now feature open data and charts

Vercel's AI Gateway leaderboards now offer shareable charts and open access to production traffic data. Users can analyze model, lab, app, and provider performance via downloadable CSVs or a programmatic export endpoint.

📄
ArXiv AI25d ago

Survey on LLMs for Medical Reasoning and Clinical Needs

This survey evaluates LLMs in healthcare by mapping clinical competency levels to computational reasoning patterns. It introduces a new benchmark across five reasoning levels and compares 18 state-of-the-art models to identify performance gaps in clinical practice.

📄
ArXiv AI26d ago

SageMath-Augmented Agents Boost Mathematical Problem Solving

Researchers introduced a ReAct-style agentic framework that integrates LLM reasoning with SageMath for verifiable mathematical computation. The study demonstrates significant performance gains across various models, narrowing the gap between open-weight and closed-source systems.

📄
ArXiv AI26d ago

Measuring AI Intelligence Beyond Human Capability

This research proposes a new paradigm for evaluating AI models that have surpassed human capability by using relative measurement instead of static benchmarks. The framework utilizes model-generated challenges to create an adversarial psychometric rating system that scales alongside agent capabilities.

📄
ArXiv AI26d ago

AgentLens: A New Benchmark for Evaluating Coding Agent Trajectories

AgentLens is a new open-source benchmark that evaluates coding agents based on their entire interaction trajectory rather than just binary pass/fail results. It combines formal verification with LLM-generated reviews to provide actionable insights into agent behavior and performance.

⚛️
量子位27d ago

RoboDojo: A New Benchmark for Embodied AI

RoboDojo has been introduced as a challenging new benchmark for embodied AI, highlighting the significant performance gap between humans and current AI models. The benchmark reveals that even the most advanced models struggle to reach human-level proficiency in physical tasks.

🤖
OpenAI News27d ago

OpenAI identifies reliability issues in SWE-Bench Pro coding benchmark

OpenAI released an analysis highlighting significant flaws in the SWE-Bench Pro benchmark. The report questions the accuracy and reliability of current methods used to evaluate AI coding capabilities.

🦙
Reddit r/LocalLLaMA27d ago

RAG is essential for accurate local LLM technical answers

A developer experiment shows that local LLMs struggle with technical accuracy without RAG, but perform exceptionally well when provided with a knowledge base. The study tested various models including Apple Intelligence and Qwen, highlighting the effectiveness of context injection.