๐Ÿ’ผStalecollected in 13m

New AI IQ site scores frontier models on human scale

New AI IQ site scores frontier models on human scale
PostLinkedIn
๐Ÿ’ผRead original on VentureBeat

๐Ÿ’กA controversial new way to visualize LLM performance that is currently dividing the AI research community.

โšก 30-Second TL;DR

What Changed

Maps 50+ frontier LLMs onto a standard human IQ bell curve.

Why It Matters

The project provides a simplified visualization for enterprise stakeholders to compare model performance, though researchers warn it may oversimplify the 'jagged' nature of AI intelligence.

What To Do Next

Visit aiiq.org to see how your preferred models rank and evaluate if this methodology aligns with your specific use-case requirements.

Who should care:Researchers & Academics

Key Points

  • โ€ขMaps 50+ frontier LLMs onto a standard human IQ bell curve.
  • โ€ขCalculates composite IQ using 12 benchmarks across four reasoning dimensions.
  • โ€ขUses hand-calibrated difficulty curves to prevent score inflation from data contamination.
  • โ€ขSparks debate over whether reducing complex AI capabilities to a single number is meaningful.

๐Ÿง  Deep Insight

Web-grounded analysis with 7 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe AI IQ platform was created by Ryan Shea, an engineer, entrepreneur, and angel investor known for co-founding the blockchain platform Stacks, among other ventures.
  • โ€ขBeyond a single IQ score, the platform offers interactive visualizations that track how frontier AI intelligence changes over time, compares models on both IQ and EQ (emotional intelligence), and plots intelligence against operational cost.
  • โ€ขThe platform estimates emotional intelligence (EQ) using data from EQ-Bench 3 and Arena signals, mapping these scores onto a normalized scale for direct comparison with IQ.
  • โ€ขThe introduction of AI IQ has generated significant debate within the tech community, with some enterprise technologists praising its ability to make a complex market more legible, while researchers and commentators criticize the reduction of complex AI capabilities to a single, potentially misleading, number.
  • โ€ขAs of mid-May 2026, OpenAI's GPT-5.5 leads the AI IQ bell curve with an estimated IQ near 136, closely followed by Anthropic's Opus 4.7.
๐Ÿ“Š Competitor Analysisโ–ธ Show

Competitor Analysis: AI IQ vs. Other LLM Benchmarking Platforms

Feature / PlatformAI IQLLM Stats LeaderboardArtificial Analysis Intelligence IndexLiveBench
Primary MetricComposite IQ score (human IQ scale)Composite LLM Stats Score (aggregates GPQA, SWE-Bench, coding-arena, pricing)Intelligence Index v4.0 (ELO ratings, real-world tasks)Global Average (across 23 diverse tasks)
Key VisualizationsIQ bell curve, IQ over time, IQ vs. EQ, Intelligence vs. CostRanked leaderboard, performance indexELO ratings, accuracy, hallucination ratesLeaderboard with subcategory averages
Reasoning DimensionsAbstract, Mathematical, Programmatic, AcademicGPQA (reasoning), SWE-Bench (coding), coding-arenaAgents, coding, scientific reasoning, general knowledgeReasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, IF
Contamination PreventionHand-calibrated difficulty curves, compressed ceilings for easier/gameable benchmarksContinuously updated from public benchmarks and live API metrics; focuses on non-saturated benchmarksHand-curated problems, guess-resistant, machine-verifiable answers; removed saturated benchmarksNew questions regularly, complete refresh every 6 months, verifiable objective ground-truth answers
Unique FeaturesEQ estimates, direct IQ/EQ comparison, founder Ryan Shea (Stacks)Ranks 300+ models, continuously updated pricing and speed dataFocus on 'economically useful action,' agentic harness ('Stirrup'), blind pairwise comparisonsLimits potential contamination by releasing new questions regularly, no LLM judge needed
Pricing InformationNot explicitly stated for platform accessIncludes cost per million tokens in rankingNot explicitly stated for platform accessNot explicitly stated for platform access
Update FrequencyNot explicitly stated, but tracks 'frontier changes over time'Continuously updated, weekly for benchmark scores, hourly for pricingELO ratings frozen at evaluation time for stability; pricing data hourlyQuestions updated regularly, benchmark refreshes every 6 months
Criticism/DebateDebate over reducing complex AI to a single number, scientific validityBenchmark saturation issues acknowledgedAddresses benchmark saturation by making curve harder to climbDesigned with test set contamination and objective evaluation in mind

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขThe AI IQ platform utilizes a composite score derived from 12 benchmarks, categorized into four primary reasoning dimensions: Abstract, Mathematical, Programmatic, and Academic Reasoning.
  • โ€ขEach raw benchmark score is translated into an implied IQ through a system of 'hand-calibrated difficulty curves.'
  • โ€ขTo prevent score inflation from data contamination, benchmarks considered easier or more susceptible to 'gaming' have compressed IQ ceilings, limiting their influence above an IQ of 100.
  • โ€ขConversely, harder and less 'gameable' benchmarks are designed to retain high IQ ceilings, allowing for greater differentiation among top-performing models.
  • โ€ขFor a model to receive a derived IQ, it must have coverage in at least two of the four reasoning dimensions.
  • โ€ขMissing benchmark data and dimensions are conservatively imputed within the scoring pipeline. This imputation prioritizes direct predecessor lineage when explicit, otherwise using a matched lower-quartile cap based on models with similar capabilities across other dimensions.
  • โ€ขThe Mathematical Reasoning dimension specifically includes benchmarks such as FrontierMath Tier 1-3 and ProofBench.
  • โ€ขEmotional intelligence (EQ) is estimated using signals from EQ-Bench 3 and Arena, which are then mapped onto a normalized scale comparable to the IQ scores.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

The 'IQ' metaphor for AI will persist and gain wider adoption for communicating AI capabilities to a general audience.
Despite academic criticism regarding oversimplification, the intuitive nature of an IQ score makes complex AI performance more accessible and understandable for non-expert users and enterprise decision-makers, driving its continued use.
AI development will increasingly focus on optimizing for a broader range of 'intelligence' metrics, including emotional intelligence (EQ), due to platforms like AI IQ.
By providing comparable scores for both IQ and EQ, AI IQ encourages model developers to consider and improve aspects of AI beyond purely cognitive reasoning, potentially leading to more holistically capable AI systems.
Methodologies for preventing data contamination and ensuring robust, 'ungameable' benchmarks, such as hand-calibrated difficulty curves, will become standard practice in AI evaluation.
As frontier models become more sophisticated and prone to memorization or 'gaming' benchmarks, rigorous and adaptive evaluation techniques are essential for meaningful differentiation and accurate assessment of true AI progress.

โณ Timeline

2026-05
AI IQ platform launched by Ryan Shea, mapping over 50 frontier language models onto a human IQ bell curve.

๐Ÿ“Ž Sources (7)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. Google Search Source
  2. Google Search Source
  3. Google Search Source
  4. Google Search Source
  5. Google Search Source
  6. Google Search Source
  7. Google Search Source
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat โ†—