SourceStalecollected in 13m

New AI IQ site scores frontier models on human scale

Read original on VentureBeat
#benchmarking#llm-evaluation#model-comparison

A controversial new way to visualize LLM performance that is currently dividing the AI research community.

30-Second TL;DR

What Changed

Maps 50+ frontier LLMs onto a standard human IQ bell curve.

Why It Matters

The project provides a simplified visualization for enterprise stakeholders to compare model performance, though researchers warn it may oversimplify the 'jagged' nature of AI intelligence.

What To Do Next

Visit aiiq.org to see how your preferred models rank and evaluate if this methodology aligns with your specific use-case requirements.

Who should care:Researchers & Academics

Key Points

  • Maps 50+ frontier LLMs onto a standard human IQ bell curve.
  • Calculates composite IQ using 12 benchmarks across four reasoning dimensions.
  • Uses hand-calibrated difficulty curves to prevent score inflation from data contamination.
  • Sparks debate over whether reducing complex AI capabilities to a single number is meaningful.

Deep Insight

Background and context from public sources — not the original article. 7 sources cited.

Enhanced Key Takeaways

  • The AI IQ platform was created by Ryan Shea, an engineer, entrepreneur, and angel investor known for co-founding the blockchain platform Stacks, among other ventures.
  • Beyond a single IQ score, the platform offers interactive visualizations that track how frontier AI intelligence changes over time, compares models on both IQ and EQ (emotional intelligence), and plots intelligence against operational cost.
  • The platform estimates emotional intelligence (EQ) using data from EQ-Bench 3 and Arena signals, mapping these scores onto a normalized scale for direct comparison with IQ.
  • The introduction of AI IQ has generated significant debate within the tech community, with some enterprise technologists praising its ability to make a complex market more legible, while researchers and commentators criticize the reduction of complex AI capabilities to a single, potentially misleading, number.
  • As of mid-May 2026, OpenAI's GPT-5.5 leads the AI IQ bell curve with an estimated IQ near 136, closely followed by Anthropic's Opus 4.7.

Competitor Analysis

Primary Metric
AI IQ
Composite IQ score (human IQ scale)
LLM Stats Leaderboard
Composite LLM Stats Score (aggregates GPQA, SWE-Bench, coding-arena, pricing)
Artificial Analysis Intelligence Index
Intelligence Index v4.0 (ELO ratings, real-world tasks)
LiveBench
Global Average (across 23 diverse tasks)
Key Visualizations
AI IQ
IQ bell curve, IQ over time, IQ vs. EQ, Intelligence vs. Cost
LLM Stats Leaderboard
Ranked leaderboard, performance index
Artificial Analysis Intelligence Index
ELO ratings, accuracy, hallucination rates
LiveBench
Leaderboard with subcategory averages
Reasoning Dimensions
AI IQ
Abstract, Mathematical, Programmatic, Academic
LLM Stats Leaderboard
GPQA (reasoning), SWE-Bench (coding), coding-arena
Artificial Analysis Intelligence Index
Agents, coding, scientific reasoning, general knowledge
LiveBench
Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, IF
Contamination Prevention
AI IQ
Hand-calibrated difficulty curves, compressed ceilings for easier/gameable benchmarks
LLM Stats Leaderboard
Continuously updated from public benchmarks and live API metrics; focuses on non-saturated benchmarks
Artificial Analysis Intelligence Index
Hand-curated problems, guess-resistant, machine-verifiable answers; removed saturated benchmarks
LiveBench
New questions regularly, complete refresh every 6 months, verifiable objective ground-truth answers
Unique Features
AI IQ
EQ estimates, direct IQ/EQ comparison, founder Ryan Shea (Stacks)
LLM Stats Leaderboard
Ranks 300+ models, continuously updated pricing and speed data
Artificial Analysis Intelligence Index
Focus on 'economically useful action,' agentic harness ('Stirrup'), blind pairwise comparisons
LiveBench
Limits potential contamination by releasing new questions regularly, no LLM judge needed
Pricing Information
AI IQ
Not explicitly stated for platform access
LLM Stats Leaderboard
Includes cost per million tokens in ranking
Artificial Analysis Intelligence Index
Not explicitly stated for platform access
LiveBench
Not explicitly stated for platform access
Update Frequency
AI IQ
Not explicitly stated, but tracks 'frontier changes over time'
LLM Stats Leaderboard
Continuously updated, weekly for benchmark scores, hourly for pricing
Artificial Analysis Intelligence Index
ELO ratings frozen at evaluation time for stability; pricing data hourly
LiveBench
Questions updated regularly, benchmark refreshes every 6 months
Criticism/Debate
AI IQ
Debate over reducing complex AI to a single number, scientific validity
LLM Stats Leaderboard
Benchmark saturation issues acknowledged
Artificial Analysis Intelligence Index
Addresses benchmark saturation by making curve harder to climb
LiveBench
Designed with test set contamination and objective evaluation in mind

Technical Deep Dive

  • The AI IQ platform utilizes a composite score derived from 12 benchmarks, categorized into four primary reasoning dimensions: Abstract, Mathematical, Programmatic, and Academic Reasoning.
  • Each raw benchmark score is translated into an implied IQ through a system of 'hand-calibrated difficulty curves.'
  • To prevent score inflation from data contamination, benchmarks considered easier or more susceptible to 'gaming' have compressed IQ ceilings, limiting their influence above an IQ of 100.
  • Conversely, harder and less 'gameable' benchmarks are designed to retain high IQ ceilings, allowing for greater differentiation among top-performing models.
  • For a model to receive a derived IQ, it must have coverage in at least two of the four reasoning dimensions.
  • Missing benchmark data and dimensions are conservatively imputed within the scoring pipeline. This imputation prioritizes direct predecessor lineage when explicit, otherwise using a matched lower-quartile cap based on models with similar capabilities across other dimensions.
  • The Mathematical Reasoning dimension specifically includes benchmarks such as FrontierMath Tier 1-3 and ProofBench.
  • Emotional intelligence (EQ) is estimated using signals from EQ-Bench 3 and Arena, which are then mapped onto a normalized scale comparable to the IQ scores.

Future ImplicationsAI analysis grounded in cited sources

The 'IQ' metaphor for AI will persist and gain wider adoption for communicating AI capabilities to a general audience.
Despite academic criticism regarding oversimplification, the intuitive nature of an IQ score makes complex AI performance more accessible and understandable for non-expert users and enterprise decision-makers, driving its continued use.
AI development will increasingly focus on optimizing for a broader range of 'intelligence' metrics, including emotional intelligence (EQ), due to platforms like AI IQ.
By providing comparable scores for both IQ and EQ, AI IQ encourages model developers to consider and improve aspects of AI beyond purely cognitive reasoning, potentially leading to more holistically capable AI systems.
Methodologies for preventing data contamination and ensuring robust, 'ungameable' benchmarks, such as hand-calibrated difficulty curves, will become standard practice in AI evaluation.
As frontier models become more sophisticated and prone to memorization or 'gaming' benchmarks, rigorous and adaptive evaluation techniques are essential for meaningful differentiation and accurate assessment of true AI progress.

Timeline

2026-05
AI IQ platform launched by Ryan Shea, mapping over 50 frontier language models onto a human IQ bell curve.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.