BDH Crushes Extreme Sudoku Benchmark at 97.4%
💡BDH beats LLMs 97-0 on clean constraint benchmark—exposes transformer reasoning flaws
⚡ 30-Second TL;DR
What Changed
250,000 'Sudoku Extreme' hard instances as benchmark
Why It Matters
Challenges over-reliance on CoT scaling for reasoning; pushes for internal search architectures. May shift focus from verbalization to native constraint solving in AI research.
What To Do Next
Read Pathway's blog and paper linked in Reddit comments to replicate BDH on Sudoku.
Key Points
- •250,000 'Sudoku Extreme' hard instances as benchmark
- •LLMs (O3-mini, DeepSeek R1, Claude 3.7) at 0% accuracy
- •BDH achieves 97.4% natively, no CoT/tools/backtracking
- •Critiques transformers' poor fit for search-heavy tasks
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.