AI IQ

The AI record book

Instead of a leaderboard we keep a record book: five disciplines, one named test each. You can see who holds the best score, since when, who held it before, and how quickly records fall. Every value comes from an independent evaluator, and what was never measured stays a gap.

5disciplines
22record breaks in 12 months
31days since the last record
167models in the sources

As of 15 August 2026

The current records

A record is the best published run on the discipline’s one test, entered on the day the model appeared. No average across exams of different difficulty, no picking the best score out of several tests.

The record changed hands most often in Mathematics recently. It has held longest in Reliability.

Mathematics

FrontierMath Tier 4 (v2) ↗

Purpose-written, unpublished problems at research level. The only maths test in the sources that still separates frontier models.

Claude Fable 587.8 %

Record since 9 June 2026 · held for 67 days

before that GPT-5.5 Pro at 78.0 %

Chasing todayGPT-5.6 Sol · 82.9 %GPT-5.5 Pro · 78.0 %
46 models measuredevaluated by Epoch AI1 without a release date left off

Science

GPQA Diamond ↗

Biology, physics and chemistry questions that experts with a search engine mostly fail.

GPT-5.4 Pro94.6 %

Record since 5 March 2026 · held for 163 days

before that Gemini 3.1 Pro at 94.4 %

Chasing todayGemini 3.1 Pro · 94.4 %Gemini 3.6 Flash · 94.1 %
162 models measuredevaluated by Epoch AI2 without a release date left off

Software

SWE-bench Verified ↗

Real, documented bugs in real open-source projects. It only counts if the test suite passes afterwards.

Claude Opus 4.783.5 %

Record since 16 April 2026 · held for 121 days

before that Claude Opus 4.6 at 78.7 %

Chasing todayGPT-5.5 · 80.6 %Gemini 3.5 Flash · 79.3 %
32 models measuredevaluated by Epoch AI1 without a release date left off

Reliability

SimpleQA Verified ↗

Short factual questions with exactly one right answer. A fabricated answer scores worse than admitting ignorance.

Gemini 3.1 Pro77.3 %

Record since 19 February 2026 · held for 177 days

before that Gemini 3 Pro at 72.9 %

Chasing todayGemini 3 Pro · 72.9 %GPT-5.6 Sol · 71.6 %
66 models measuredevaluated by Epoch AI1 without a release date left off

Abstraction

ARC-AGI-2 ↗

Novel visual patterns that were not in training. Humans solve them reliably; models so far barely do.

Inkling36.5 %

Record since 15 July 2026 · held for 31 days

before that Gemini 3 Pro at 31.1 %

Chasing todayGemini 3 Pro · 31.1 %GLM-5.2 · 22.8 %
20 models measuredevaluated by the benchmark authors

Every dot a model, the staircase the record

Horizontally the release date, vertically the score on the chosen discipline’s test. The line rises exactly when a new model beats the best score. Hover over a dot to see the model and its value.

0255075100202420252026o3-mini · 0.0 % · 31 January 2025o4-mini · 4.9 % · 16 April 2025Gemini 2.5 Pro (Jun 2025) · 0.0 % · 17 June 2025Claude Opus 4.1 · 2.4 % · 5 August 2025GPT-5 mini · 12.2 % · 7 August 2025GPT-5 nano · 2.4 % · 7 August 2025Claude Sonnet 4.5 · 2.4 % · 29 September 2025GPT-5 Pro · 19.5 % · 7 October 2025Claude Opus 4.5 · 4.9 % · 24 November 2025GPT-5.2 · 31.7 % · 11 December 2025GPT-5.2 Pro · 46.0 % · 11 December 2025Gemini 3 Flash · 17.1 % · 17 December 2025Claude Opus 4.6 · 26.8 % · 5 February 2026Grok 4.20 · 17.1 % · 17 February 2026Gemini 3.1 Pro · 26.8 % · 19 February 2026GPT-5.4 · 49.0 % · 5 March 2026GPT-5.4 Pro · 58.5 % · 5 March 2026GPT-5.4 Mini · 9.8 % · 17 March 2026GPT-5.4 Nano · 12.2 % · 17 March 2026Claude Opus 4.7 · 31.7 % · 16 April 2026Grok 4.3 Beta · 14.6 % · 17 April 2026Kimi K2.6 · 25.6 % · 20 April 2026GPT-5.5 · 72.5 % · 23 April 2026GPT-5.5 Pro · 78.0 % · 23 April 2026DeepSeek-V4-Pro · 2.4 % · 24 April 2026GPT-5.5 Instant · 2.4 % · 5 May 2026AI Co-Mathematician · 75.6 % · 8 May 2026Gemini 3.5 Flash · 26.8 % · 19 May 2026Qwen3.7-Max · 34.1 % · 19 May 2026Claude Opus 4.8 · 56.1 % · 28 May 2026Claude Fable 5 · 87.8 % · 9 June 2026Kimi K2.7 Code · 12.2 % · 12 June 2026GLM-5.2 · 29.3 % · 16 June 2026Claude Sonnet 5 · 29.3 % · 30 June 2026Grok 4.5 · 24.4 % · 8 July 2026GPT-5.6 Luna · 61.0 % · 9 July 2026GPT-5.6 Sol · 82.9 % · 9 July 2026GPT-5.6 Terra · 70.7 % · 9 July 2026Inkling · 4.9 % · 15 July 2026Kimi K3 · 39.0 % · 16 July 2026Gemini 3.5 Flash-Lite · 0.0 % · 21 July 2026Gemini 3.6 Flash · 22.0 % · 21 July 2026Claude Opus 5 · 73.2 % · 24 July 2026DeepSeek V4 Flash 0731 · 24.4 % · 31 July 2026Qwen 3.8 Max · 46.3 % · 2 August 2026Claude Fable 5 · 87.8 %

measured modelrecord on releaserecord progression

The easy test and the hard test

Both are mathematics, both from the same model, both evaluated by Epoch AI. On the left, competition problems of the kind that appear in school olympiads. On the right, purpose-written, unpublished problems at research level. When you read somewhere that an AI solves mathematics almost perfectly, the left-hand test is almost always the one meant.

Competition problems OTIS Mock AIME 2024-2025 ↗Research problems FrontierMath Tier 4 (v2) ↗
  • Qwen 3.8 MaxAlibaba99.4 %46.3 %53.1Gap
  • DeepSeek V4 Flash 0731DeepSeek94.4 %24.4 %70.0Gap
  • Claude Opus 5Anthropic98.9 %73.2 %25.7Gap
  • Gemini 3.5 Flash-LiteGoogle DeepMind71.1 %0.0 %71.1Gap
  • Gemini 3.6 FlashGoogle DeepMind94.2 %22.0 %72.2Gap
  • Kimi K3Moonshot97.2 %39.0 %58.2Gap
  • InklingThinking Machines88.9 %4.9 %84.0Gap
  • GPT-5.6 LunaOpenAI98.3 %61.0 %37.3Gap

That is why the record book scores mathematics on the hard test. The easy one no longer separates the frontier: 36 of the 139 models measured on it sit at 90 per cent or above. On the hard test no model has managed that yet.

Two models, head to head

Pick two models and see them side by side, discipline by discipline. Only what both were measured on counts. Hatching means: no run there.

GPT-5.6 SolClaude Fable 5
82.9 %Mathematics87.8 %
93.5 %Science85.9 %
not measuredSoftwarenot measured
71.6 %Reliability68.3 %
not measuredAbstractionnot measured

GPT-5.6 Sol leads in 2 of 3 jointly measured disciplines, Claude Fable 5 in 1.

All current models

Ordered by the Epoch AI Capabilities Index, the one overall index we carry and attribute; models without an index follow at the end, newest first. The bar strip shows the five disciplines in a fixed order. A hatched cell means this model has never been run on that test.

Vendor

1Mathematics2Science3Software4Reliability5Abstraction

Select a row for the test, the evaluator and the number of runs

112 further models, measured too thinly for a profile

These models are in the sources, but on fewer than 3 of the five tests. A profile built from one or two scores says nothing about a model’s shape, so this lists only what they were measured on.

  • Grok 4.6Science, Reliability
  • Qwen3.7 FlashScience
  • Gemini 3.5 Flash-LiteMathematics, Science
  • Qwen3.7-PlusScience
  • MiniMax-M3Science
  • AI Co-MathematicianMathematics
  • Qwen 3.6 FlashScience, Reliability
  • Qwen3.6 27BScience
  • Qwen 3.6 35B-A3BScience
  • Muse SparkScience, Reliability
  • Gemma 4 31B ITScience, Reliability
  • Gemini 3.1 Flash-LiteScience
  • Qwen 3.5 Flash (hosted 35B-A3B)Science, Reliability
  • Qwen 3.5 Plus (hosted 397B-A17B)Science, Reliability
  • Qwen3.5 397B-A17BScience
  • GPT-5.3 CodexSoftware
  • GLM-4.7Science, Reliability
  • GPT-5.2 ProMathematics
  • DeepSeek-V3.2Science, Reliability
  • Kimi K2 ThinkingScience, Reliability
  • GPT-5 ProMathematics, Abstraction
  • Qwen3-MaxScience, Reliability
  • Qwen3-235B-A22B (Jul 2025)Science, Reliability
  • gpt-oss-120bScience, Reliability
  • gpt-oss-20bScience
  • Grok 4Science, Reliability
  • Magistral Small 1.0Science
  • DeepSeek-R1 (May 2025)Science, Reliability
  • Claude Sonnet 4Science, Abstraction
  • Mistral Medium 3Science
  • Gemini 2.5 Pro (May 2025)Science
  • Qwen3-235B-A22BScience
  • GPT-4.1 miniScience
  • GPT-4.1 nanoScience
  • Grok 3Science, Abstraction
  • Grok-3 miniScience, Reliability
  • QWQ-PlusScience
  • Llama 4 MaverickScience, Abstraction
  • Llama 4 ScoutScience, Abstraction
  • Gemini 2.5 Pro (Mar 2025)Science
  • DeepSeek-V3 (Mar 2025)Science
  • Mistral Small 3.1Science
  • Gemma 3 27BScience
  • GPT-4.5Science, Abstraction
  • Claude 3.7 SonnetScience, Software
  • Gemini 2.0 FlashScience, Abstraction
  • Gemini 2.0 ProScience
  • o3-miniMathematics, Science
  • Qwen2.5-MaxScience
  • Mistral Small 3Science
  • Qwen PlusScience
  • Gemini 2.0 Flash ThinkingScience
  • DeepSeek-R1Science
  • DeepSeek-R1-Distill-Llama-70BScience
  • DeepSeek-R1-Distill-Qwen-14BScience
  • Eurus-2-7B-PRIMEScience
  • DeepSeek-V3Science
  • o1Science
  • Grok-2Science
  • Phi-4Science
  • Llama 3.3 70BScience
  • Tulu 3 (Tülu 3) 70BScience
  • Qwen-TurboScience
  • Claude 3.5 HaikuScience, Reliability
  • Claude 3.5 Sonnet (October 2024)Science
  • Ministral 3BScience
  • Ministral 8BScience
  • Gemini 1.5 Flash 8BScience
  • Llama 3.2 90BScience
  • Qwen2.5-72BScience
  • Qwen2.5-32BScience
  • o1-miniScience, Abstraction
  • o1-previewScience
  • Mistral Large 2Science
  • Llama 3.1-405BScience
  • Llama 3.1-70BScience
  • Llama 3.1-8BScience
  • GPT-4o miniScience
  • Mistral NeMoScience
  • Gemma 2 27BScience
  • Claude 3.5 SonnetScience
  • Hermes 2 Theta Llama-3 70BScience
  • Qwen2-72BScience
  • Gemini 1.5 FlashScience
  • Yi-1.5-34BScience
  • phi-3-medium 14BScience
  • Llama 3-70BScience
  • Llama 3-8BScience
  • Mixtral 8x22BScience
  • WizardLM-2 8x22BScience
  • GPT-4 Turbo (Apr 2024)Science
  • DBRXScience
  • Claude 3 HaikuScience
  • Claude 3 OpusScience
  • Claude 3 SonnetScience
  • Mistral LargeScience
  • Gemini 1.5 ProScience, Abstraction
  • Qwen1.5-32BScience
  • Qwen1.5-72BScience
  • Gemini 1.0 ProScience
  • Mixtral 8x7BScience
  • DeepSeek LLM 67BScience
  • Claude 2.1Science
  • GPT-4 Turbo (Nov 2023)Science
  • Yi-34BScience
  • Mistral 7BScience
  • Llama 2-70BScience
  • Claude 2Science
  • GPT-3.5 TurboScience
  • GPT-4 (Jun 2023)Science
  • GPT-4 (Mar 2023)Science
  • Gemma 2 9BScience

How this is built, and what it does not claim

  • One discipline, one test

    The sources carry four mathematics tests of very different difficulty. Averaging them, or taking the best score among them, turns a solved research problem and a solved competition problem into the same figure. We bind each discipline to exactly one test and show the others separately, as what they are.

  • What a record is here

    A model’s best published run on the discipline’s test, entered on the model’s release day. Only dated models enter the series; the rest are counted. If an older model is later re-measured with a better result, the series can change retroactively. That is why the observation time is stated at the top.

  • No score of our own

    The index shown is the Epoch AI Capabilities Index with its confidence interval. We carry it through and attribute it rather than inventing a competing number. We crown no overall winner across disciplines.

  • No imputed values

    A discipline without a successful public run stays empty and lowers the published coverage. Nothing is estimated to fill a row.

  • Who measured it is recorded

    Every test names its evaluator. Epoch AI distinguishes its own evaluation, evaluation by the benchmark authors, and self-report by the model vendor. Scores of the last kind are flagged; on this page there are currently none.

  • Names are never merged

    A result from a second source attaches only on an exact model-name match. Anything without an exact counterpart is left out and counted, rather than guessed into place. That is why the abstraction discipline is thinly populated.

A benchmark score is a snapshot under one test setup. It does not establish that a model is intelligent, safe, or suitable for a given task.

Frequently asked questions

Is this an IQ test?
No. A language model has no IQ in the human sense, and we do not calculate one. What we keep is a record book across five named tests. Where Epoch AI publishes its Capabilities Index we carry it through with its confidence interval and attribute it.
What counts as a record here?
A model’s best published run on the discipline’s one test, entered on the model’s release day. Only a strictly better score breaks a record; a tie does not. If an older model is later re-measured with a better result, the series can change retroactively, which is why the page states its observation time.
Why are your mathematics scores lower than elsewhere?
Because we show the hard test. The sources carry four mathematics tests of very different difficulty. On the competition problems several models reach 100 per cent; on the research problems no model has ever reached 90. The easy test no longer separates the frontier, so the record book scores on the hard one. You will find the easy score in each model’s detail and in the direct comparison.
Why are some cells hatched?
Because there is no successful public run there. The gap is shown as a gap and lowers the published coverage. Values we filled in ourselves used to sit in those places; we removed them.
Why do you crown no overall winner?
Because each discipline has its own record holder and the five tests measure different things. Any total would have to weight the disciplines, and that weighting would be our invention. Within one discipline a rank does hold, because every model there was measured on the same test.
Why are models I have heard of missing?
A model needs scores on at least three of the five tests to appear in the model list. Everything else is named below it, together with what it was measured on. Models that appear in neither source do not show up at all.
Where do the scores come from?
From Epoch AI’s benchmarking hub and the ARC Prize data files. Both are official, machine-readable, and permit retrieval. Every cell names its test and who evaluated it.
How current is this?
The time of the last collection is stated at the top. The data is gathered afresh on every publication of the page; we claim no fixed cadence that the pipeline does not keep.

Sources & status

As of 15 August 2026

Benchmark scores are snapshots under one test setup. They do not establish that a model is intelligent, safe, or suitable for a given task.

Which model fits your use case?

A record does not settle that. We put the right model into your process and tell you what we based that on.