New AI IQ site scores frontier models on human scale

A controversial new way to visualize LLM performance that is currently dividing the AI research community.
30-Second TL;DR
What Changed
Maps 50+ frontier LLMs onto a standard human IQ bell curve.
Why It Matters
The project provides a simplified visualization for enterprise stakeholders to compare model performance, though researchers warn it may oversimplify the 'jagged' nature of AI intelligence.
What To Do Next
Visit aiiq.org to see how your preferred models rank and evaluate if this methodology aligns with your specific use-case requirements.
Key Points
- •Maps 50+ frontier LLMs onto a standard human IQ bell curve.
- •Calculates composite IQ using 12 benchmarks across four reasoning dimensions.
- •Uses hand-calibrated difficulty curves to prevent score inflation from data contamination.
- •Sparks debate over whether reducing complex AI capabilities to a single number is meaningful.
Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
Enhanced Key Takeaways
- •The AI IQ platform was created by Ryan Shea, an engineer, entrepreneur, and angel investor known for co-founding the blockchain platform Stacks, among other ventures.
- •Beyond a single IQ score, the platform offers interactive visualizations that track how frontier AI intelligence changes over time, compares models on both IQ and EQ (emotional intelligence), and plots intelligence against operational cost.
- •The platform estimates emotional intelligence (EQ) using data from EQ-Bench 3 and Arena signals, mapping these scores onto a normalized scale for direct comparison with IQ.
- •The introduction of AI IQ has generated significant debate within the tech community, with some enterprise technologists praising its ability to make a complex market more legible, while researchers and commentators criticize the reduction of complex AI capabilities to a single, potentially misleading, number.
- •As of mid-May 2026, OpenAI's GPT-5.5 leads the AI IQ bell curve with an estimated IQ near 136, closely followed by Anthropic's Opus 4.7.
Competitor Analysis
- AI IQ
- Composite IQ score (human IQ scale)
- LLM Stats Leaderboard
- Composite LLM Stats Score (aggregates GPQA, SWE-Bench, coding-arena, pricing)
- Artificial Analysis Intelligence Index
- Intelligence Index v4.0 (ELO ratings, real-world tasks)
- LiveBench
- Global Average (across 23 diverse tasks)
- AI IQ
- IQ bell curve, IQ over time, IQ vs. EQ, Intelligence vs. Cost
- LLM Stats Leaderboard
- Ranked leaderboard, performance index
- Artificial Analysis Intelligence Index
- ELO ratings, accuracy, hallucination rates
- LiveBench
- Leaderboard with subcategory averages
- AI IQ
- Abstract, Mathematical, Programmatic, Academic
- LLM Stats Leaderboard
- GPQA (reasoning), SWE-Bench (coding), coding-arena
- Artificial Analysis Intelligence Index
- Agents, coding, scientific reasoning, general knowledge
- LiveBench
- Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, IF
- AI IQ
- Hand-calibrated difficulty curves, compressed ceilings for easier/gameable benchmarks
- LLM Stats Leaderboard
- Continuously updated from public benchmarks and live API metrics; focuses on non-saturated benchmarks
- Artificial Analysis Intelligence Index
- Hand-curated problems, guess-resistant, machine-verifiable answers; removed saturated benchmarks
- LiveBench
- New questions regularly, complete refresh every 6 months, verifiable objective ground-truth answers
- AI IQ
- EQ estimates, direct IQ/EQ comparison, founder Ryan Shea (Stacks)
- LLM Stats Leaderboard
- Ranks 300+ models, continuously updated pricing and speed data
- Artificial Analysis Intelligence Index
- Focus on 'economically useful action,' agentic harness ('Stirrup'), blind pairwise comparisons
- LiveBench
- Limits potential contamination by releasing new questions regularly, no LLM judge needed
- AI IQ
- Not explicitly stated for platform access
- LLM Stats Leaderboard
- Includes cost per million tokens in ranking
- Artificial Analysis Intelligence Index
- Not explicitly stated for platform access
- LiveBench
- Not explicitly stated for platform access
- AI IQ
- Not explicitly stated, but tracks 'frontier changes over time'
- LLM Stats Leaderboard
- Continuously updated, weekly for benchmark scores, hourly for pricing
- Artificial Analysis Intelligence Index
- ELO ratings frozen at evaluation time for stability; pricing data hourly
- LiveBench
- Questions updated regularly, benchmark refreshes every 6 months
- AI IQ
- Debate over reducing complex AI to a single number, scientific validity
- LLM Stats Leaderboard
- Benchmark saturation issues acknowledged
- Artificial Analysis Intelligence Index
- Addresses benchmark saturation by making curve harder to climb
- LiveBench
- Designed with test set contamination and objective evaluation in mind
| Feature / Platform | AI IQ | LLM Stats Leaderboard | Artificial Analysis Intelligence Index | LiveBench |
|---|---|---|---|---|
| Primary Metric | Composite IQ score (human IQ scale) | Composite LLM Stats Score (aggregates GPQA, SWE-Bench, coding-arena, pricing) | Intelligence Index v4.0 (ELO ratings, real-world tasks) | Global Average (across 23 diverse tasks) |
| Key Visualizations | IQ bell curve, IQ over time, IQ vs. EQ, Intelligence vs. Cost | Ranked leaderboard, performance index | ELO ratings, accuracy, hallucination rates | Leaderboard with subcategory averages |
| Reasoning Dimensions | Abstract, Mathematical, Programmatic, Academic | GPQA (reasoning), SWE-Bench (coding), coding-arena | Agents, coding, scientific reasoning, general knowledge | Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, IF |
| Contamination Prevention | Hand-calibrated difficulty curves, compressed ceilings for easier/gameable benchmarks | Continuously updated from public benchmarks and live API metrics; focuses on non-saturated benchmarks | Hand-curated problems, guess-resistant, machine-verifiable answers; removed saturated benchmarks | New questions regularly, complete refresh every 6 months, verifiable objective ground-truth answers |
| Unique Features | EQ estimates, direct IQ/EQ comparison, founder Ryan Shea (Stacks) | Ranks 300+ models, continuously updated pricing and speed data | Focus on 'economically useful action,' agentic harness ('Stirrup'), blind pairwise comparisons | Limits potential contamination by releasing new questions regularly, no LLM judge needed |
| Pricing Information | Not explicitly stated for platform access | Includes cost per million tokens in ranking | Not explicitly stated for platform access | Not explicitly stated for platform access |
| Update Frequency | Not explicitly stated, but tracks 'frontier changes over time' | Continuously updated, weekly for benchmark scores, hourly for pricing | ELO ratings frozen at evaluation time for stability; pricing data hourly | Questions updated regularly, benchmark refreshes every 6 months |
| Criticism/Debate | Debate over reducing complex AI to a single number, scientific validity | Benchmark saturation issues acknowledged | Addresses benchmark saturation by making curve harder to climb | Designed with test set contamination and objective evaluation in mind |
Technical Deep Dive
- •The AI IQ platform utilizes a composite score derived from 12 benchmarks, categorized into four primary reasoning dimensions: Abstract, Mathematical, Programmatic, and Academic Reasoning.
- •Each raw benchmark score is translated into an implied IQ through a system of 'hand-calibrated difficulty curves.'
- •To prevent score inflation from data contamination, benchmarks considered easier or more susceptible to 'gaming' have compressed IQ ceilings, limiting their influence above an IQ of 100.
- •Conversely, harder and less 'gameable' benchmarks are designed to retain high IQ ceilings, allowing for greater differentiation among top-performing models.
- •For a model to receive a derived IQ, it must have coverage in at least two of the four reasoning dimensions.
- •Missing benchmark data and dimensions are conservatively imputed within the scoring pipeline. This imputation prioritizes direct predecessor lineage when explicit, otherwise using a matched lower-quartile cap based on models with similar capabilities across other dimensions.
- •The Mathematical Reasoning dimension specifically includes benchmarks such as FrontierMath Tier 1-3 and ProofBench.
- •Emotional intelligence (EQ) is estimated using signals from EQ-Bench 3 and Arena, which are then mapped onto a normalized scale comparable to the IQ scores.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-05AI IQ platform launched by Ryan Shea, mapping over 50 frontier language models onto a human IQ bell curve.
Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.