OpenAI’s AGI Score Collapses Under the Official Harness

💡A 99.9% AGI claim became 62.7% under the benchmark’s own software.
⚡ 30-Second TL;DR
What Changed
OpenAI cited a 99.9% benchmark score in declaring the AGI era.
Why It Matters
The discrepancy could undermine confidence in headline AGI claims and make independent replication more important. AI teams may need to treat benchmark scores as implementation-dependent results rather than universal measures of model capability.
What To Do Next
Re-run GPT-6 Astra or comparable models with ARC Prize’s official harness and document every evaluator configuration before citing the benchmark.
Key Points
- •OpenAI cited a 99.9% benchmark score in declaring the AGI era.
- •Running the same model through the benchmark’s own harness produced a 62.7% score.
- •ARC Prize published both results, highlighting a substantial discrepancy caused by evaluation software.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


