A new startup project called AI IQ has released interactive charts that try to make frontier language-model progress easier to scan. The site, aiiq.org, assigns an estimated “intelligence quotient” to more than 50 top language models and plots them on a bell curve.
The results spread fast. VentureBeat says the charts “ricocheted across social media” over the past week, with enterprise technologists praising the visualization and researchers pushing back on the underlying idea. Technology commentator Thibaut Mélen called it “super useful” for turning progress into something “legible.” Business strategist Brian Vellmure said it “tracks” with his personal experience. The criticism landed just as quickly, with AI Deeply posting that “AI is far too jagged” and that “the map is not the territory.”
How AI IQ builds an “IQ” number
AI IQ’s methodology takes a set of 12 benchmarks and groups them into four reasoning dimensions. The dimensions are described as abstract, mathematical, programmatic, and academic.
The site then computes a composite score as a straight average of the four dimension scores.
The component benchmarks match this structure, per the VentureBeat write-up.
- Abstract reasoning: ARC-AGI-1 and ARC-AGI-2
- Mathematical reasoning: FrontierMath (Tiers 1–3 and Tier 4), AIME, ProofBench
- Programmatic reasoning: Terminal-Bench 2.0, SWE-Bench Verified, SciCode
- Academic reasoning: Humanity’s Last Exam, CritPt, GPQA Diamond
Each benchmark raw score gets mapped to an implied IQ using what the site describes as “hand-calibrated difficulty curves.” VentureBeat also reports two design choices meant to prevent inflated totals:
- Ceiling compression for benchmarks the site considers easier or more vulnerable to data contamination. Easier tasks cap out sooner, harder tasks retain higher ceilings.
- Conservative handling of missing data. VentureBeat says models must have results on at least two of the four dimensions to get a derived IQ. When benchmarks are absent, the pipeline pulls scores down rather than up, and the site reports that derived IQ “averages all four dimensions,” so omission cannot make a model look better by leaving it out.
Why it matters
The core dispute is not about whether charts are helpful. It is about whether compressing uneven capabilities into one controversial number creates an illusion of precision.
VentureBeat frames the pushback as a methodological warning. AI Deeply’s comment captures the concern researchers keep returning to. If model performance is jagged across tasks and benchmarks, a single scalar can look more stable than it is.
On the other side, supporters like Mélen and Vellmure argue the opposite: that a single view makes a complicated landscape easier to interpret than endless leaderboard tables.
AI IQ tries to thread that needle with ceiling compression and missing-data rules. But the backlash shows how quickly “legible” can turn into “oversimplified,” depending on how readers treat the number.
Market impact
This is where the story crosses into crypto infrastructure culture. VentureBeat tags the story under exchanges and NFTs, and that matters because public model-ranking narratives often spill into broader tech speculation.
Even if AI IQ is purely informational, a single bell-curve chart tends to travel like a headline. VentureBeat reports that it is already being shared widely. Once something gets shared in a scalar form, people use it as shorthand, even when the underlying benchmark coverage is incomplete or the mapping is contested.
What to watch next
The most important next step is not a redesign of the bell curve. It is how AI IQ handles ongoing changes in benchmarks, scoring updates, and model release cycles.
VentureBeat says AI IQ tracks convergence at the top of the frontier and broader diversity below. If the charts keep showing tight clustering among top models, critics will likely push harder on whether “difference compression” is a property of the models, the benchmarks, or the mapping.
Here is the early top cluster described by VentureBeat, based on mid-May 2026 AI IQ charts.
| Model (provider) | Approx. AI IQ (per AI IQ charts) |
|---|---|
| GPT-5.5 (OpenAI) | ~136 |
| Opus 4.7 (Anthropic) | ~132 |
| GPT-5.4 (OpenAI) | ~131 |
| Gemini 3.1 Pro (Google) | ~131 |
| Opus 4.6 (Anthropic) | ~129 |
reality check AI IQ gives a clean visualization of messy benchmark results. It may help people compare progress at a glance. But VentureBeat’s reporting also captures the reason it drew pushback. A single “IQ” number can mask task-by-task jaggedness. If you use the chart, treat the score as an estimate tied to specific benchmarks, not a measurement of intelligence in any human sense.