New AI IQ site scores frontier models on human scale

๐กA controversial new way to visualize LLM performance that is currently dividing the AI research community.
โก 30-Second TL;DR
What Changed
Maps 50+ frontier LLMs onto a standard human IQ bell curve.
Why It Matters
The project provides a simplified visualization for enterprise stakeholders to compare model performance, though researchers warn it may oversimplify the 'jagged' nature of AI intelligence.
What To Do Next
Visit aiiq.org to see how your preferred models rank and evaluate if this methodology aligns with your specific use-case requirements.
Key Points
- โขMaps 50+ frontier LLMs onto a standard human IQ bell curve.
- โขCalculates composite IQ using 12 benchmarks across four reasoning dimensions.
- โขUses hand-calibrated difficulty curves to prevent score inflation from data contamination.
- โขSparks debate over whether reducing complex AI capabilities to a single number is meaningful.
๐ง Deep Insight
Web-grounded analysis with 7 cited sources.
๐ Enhanced Key Takeaways
- โขThe AI IQ platform was created by Ryan Shea, an engineer, entrepreneur, and angel investor known for co-founding the blockchain platform Stacks, among other ventures.
- โขBeyond a single IQ score, the platform offers interactive visualizations that track how frontier AI intelligence changes over time, compares models on both IQ and EQ (emotional intelligence), and plots intelligence against operational cost.
- โขThe platform estimates emotional intelligence (EQ) using data from EQ-Bench 3 and Arena signals, mapping these scores onto a normalized scale for direct comparison with IQ.
- โขThe introduction of AI IQ has generated significant debate within the tech community, with some enterprise technologists praising its ability to make a complex market more legible, while researchers and commentators criticize the reduction of complex AI capabilities to a single, potentially misleading, number.
- โขAs of mid-May 2026, OpenAI's GPT-5.5 leads the AI IQ bell curve with an estimated IQ near 136, closely followed by Anthropic's Opus 4.7.
๐ Competitor Analysisโธ Show
Competitor Analysis: AI IQ vs. Other LLM Benchmarking Platforms
| Feature / Platform | AI IQ | LLM Stats Leaderboard | Artificial Analysis Intelligence Index | LiveBench |
|---|---|---|---|---|
| Primary Metric | Composite IQ score (human IQ scale) | Composite LLM Stats Score (aggregates GPQA, SWE-Bench, coding-arena, pricing) | Intelligence Index v4.0 (ELO ratings, real-world tasks) | Global Average (across 23 diverse tasks) |
| Key Visualizations | IQ bell curve, IQ over time, IQ vs. EQ, Intelligence vs. Cost | Ranked leaderboard, performance index | ELO ratings, accuracy, hallucination rates | Leaderboard with subcategory averages |
| Reasoning Dimensions | Abstract, Mathematical, Programmatic, Academic | GPQA (reasoning), SWE-Bench (coding), coding-arena | Agents, coding, scientific reasoning, general knowledge | Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, IF |
| Contamination Prevention | Hand-calibrated difficulty curves, compressed ceilings for easier/gameable benchmarks | Continuously updated from public benchmarks and live API metrics; focuses on non-saturated benchmarks | Hand-curated problems, guess-resistant, machine-verifiable answers; removed saturated benchmarks | New questions regularly, complete refresh every 6 months, verifiable objective ground-truth answers |
| Unique Features | EQ estimates, direct IQ/EQ comparison, founder Ryan Shea (Stacks) | Ranks 300+ models, continuously updated pricing and speed data | Focus on 'economically useful action,' agentic harness ('Stirrup'), blind pairwise comparisons | Limits potential contamination by releasing new questions regularly, no LLM judge needed |
| Pricing Information | Not explicitly stated for platform access | Includes cost per million tokens in ranking | Not explicitly stated for platform access | Not explicitly stated for platform access |
| Update Frequency | Not explicitly stated, but tracks 'frontier changes over time' | Continuously updated, weekly for benchmark scores, hourly for pricing | ELO ratings frozen at evaluation time for stability; pricing data hourly | Questions updated regularly, benchmark refreshes every 6 months |
| Criticism/Debate | Debate over reducing complex AI to a single number, scientific validity | Benchmark saturation issues acknowledged | Addresses benchmark saturation by making curve harder to climb | Designed with test set contamination and objective evaluation in mind |
๐ ๏ธ Technical Deep Dive
- โขThe AI IQ platform utilizes a composite score derived from 12 benchmarks, categorized into four primary reasoning dimensions: Abstract, Mathematical, Programmatic, and Academic Reasoning.
- โขEach raw benchmark score is translated into an implied IQ through a system of 'hand-calibrated difficulty curves.'
- โขTo prevent score inflation from data contamination, benchmarks considered easier or more susceptible to 'gaming' have compressed IQ ceilings, limiting their influence above an IQ of 100.
- โขConversely, harder and less 'gameable' benchmarks are designed to retain high IQ ceilings, allowing for greater differentiation among top-performing models.
- โขFor a model to receive a derived IQ, it must have coverage in at least two of the four reasoning dimensions.
- โขMissing benchmark data and dimensions are conservatively imputed within the scoring pipeline. This imputation prioritizes direct predecessor lineage when explicit, otherwise using a matched lower-quartile cap based on models with similar capabilities across other dimensions.
- โขThe Mathematical Reasoning dimension specifically includes benchmarks such as FrontierMath Tier 1-3 and ProofBench.
- โขEmotional intelligence (EQ) is estimated using signals from EQ-Bench 3 and Arena, which are then mapped onto a normalized scale comparable to the IQ scores.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat โ
