Instead of a leaderboard we keep a record book: five disciplines, one named test each. You can see who holds the best score, since when, who held it before, and how quickly records fall. Every value comes from an independent evaluator, and what was never measured stays a gap.
5disciplines
22record breaks in 12 months
31days since the last record
167models in the sources
As of 15 August 2026
The current records
A record is the best published run on the discipline’s one test, entered on the day the model appeared. No average across exams of different difficulty, no picking the best score out of several tests.
The record changed hands most often in Mathematics recently. It has held longest in Reliability.
Novel visual patterns that were not in training. Humans solve them reliably; models so far barely do.
Inkling36.5 %
Record since 15 July 2026 · held for 31 days
before that Gemini 3 Pro at 31.1 %
Chasing todayGemini 3 Pro · 31.1 %GLM-5.2 · 22.8 %
Every dot a model, the staircase the record
Horizontally the release date, vertically the score on the chosen discipline’s test. The line rises exactly when a new model beats the best score. Hover over a dot to see the model and its value.
measured modelrecord on releaserecord progression
The easy test and the hard test
Both are mathematics, both from the same model, both evaluated by Epoch AI. On the left, competition problems of the kind that appear in school olympiads. On the right, purpose-written, unpublished problems at research level. When you read somewhere that an AI solves mathematics almost perfectly, the left-hand test is almost always the one meant.
That is why the record book scores mathematics on the hard test. The easy one no longer separates the frontier: 36 of the 139 models measured on it sit at 90 per cent or above. On the hard test no model has managed that yet.
Two models, head to head
Pick two models and see them side by side, discipline by discipline. Only what both were measured on counts. Hatching means: no run there.
GPT-5.6 SolClaude Fable 5
82.9 %Mathematics87.8 %
93.5 %Science85.9 %
not measuredSoftwarenot measured
71.6 %Reliability68.3 %
not measuredAbstractionnot measured
GPT-5.6 Sol leads in 2 of 3 jointly measured disciplines, Claude Fable 5 in 1.
All current models
Ordered by the Epoch AI Capabilities Index, the one overall index we carry and attribute; models without an index follow at the end, newest first. The bar strip shows the five disciplines in a fixed order. A hatched cell means this model has never been run on that test.
Select a row for the test, the evaluator and the number of runs
112 further models, measured too thinly for a profile
These models are in the sources, but on fewer than 3 of the five tests. A profile built from one or two scores says nothing about a model’s shape, so this lists only what they were measured on.
Qwen 3.5 Plus (hosted 397B-A17B)Science, Reliability
Qwen3.5 397B-A17BScience
GPT-5.3 CodexSoftware
GLM-4.7Science, Reliability
GPT-5.2 ProMathematics
DeepSeek-V3.2Science, Reliability
Kimi K2 ThinkingScience, Reliability
GPT-5 ProMathematics, Abstraction
Qwen3-MaxScience, Reliability
Qwen3-235B-A22B (Jul 2025)Science, Reliability
gpt-oss-120bScience, Reliability
gpt-oss-20bScience
Grok 4Science, Reliability
Magistral Small 1.0Science
DeepSeek-R1 (May 2025)Science, Reliability
Claude Sonnet 4Science, Abstraction
Mistral Medium 3Science
Gemini 2.5 Pro (May 2025)Science
Qwen3-235B-A22BScience
GPT-4.1 miniScience
GPT-4.1 nanoScience
Grok 3Science, Abstraction
Grok-3 miniScience, Reliability
QWQ-PlusScience
Llama 4 MaverickScience, Abstraction
Llama 4 ScoutScience, Abstraction
Gemini 2.5 Pro (Mar 2025)Science
DeepSeek-V3 (Mar 2025)Science
Mistral Small 3.1Science
Gemma 3 27BScience
GPT-4.5Science, Abstraction
Claude 3.7 SonnetScience, Software
Gemini 2.0 FlashScience, Abstraction
Gemini 2.0 ProScience
o3-miniMathematics, Science
Qwen2.5-MaxScience
Mistral Small 3Science
Qwen PlusScience
Gemini 2.0 Flash ThinkingScience
DeepSeek-R1Science
DeepSeek-R1-Distill-Llama-70BScience
DeepSeek-R1-Distill-Qwen-14BScience
Eurus-2-7B-PRIMEScience
DeepSeek-V3Science
o1Science
Grok-2Science
Phi-4Science
Llama 3.3 70BScience
Tulu 3 (Tülu 3) 70BScience
Qwen-TurboScience
Claude 3.5 HaikuScience, Reliability
Claude 3.5 Sonnet (October 2024)Science
Ministral 3BScience
Ministral 8BScience
Gemini 1.5 Flash 8BScience
Llama 3.2 90BScience
Qwen2.5-72BScience
Qwen2.5-32BScience
o1-miniScience, Abstraction
o1-previewScience
Mistral Large 2Science
Llama 3.1-405BScience
Llama 3.1-70BScience
Llama 3.1-8BScience
GPT-4o miniScience
Mistral NeMoScience
Gemma 2 27BScience
Claude 3.5 SonnetScience
Hermes 2 Theta Llama-3 70BScience
Qwen2-72BScience
Gemini 1.5 FlashScience
Yi-1.5-34BScience
phi-3-medium 14BScience
Llama 3-70BScience
Llama 3-8BScience
Mixtral 8x22BScience
WizardLM-2 8x22BScience
GPT-4 Turbo (Apr 2024)Science
DBRXScience
Claude 3 HaikuScience
Claude 3 OpusScience
Claude 3 SonnetScience
Mistral LargeScience
Gemini 1.5 ProScience, Abstraction
Qwen1.5-32BScience
Qwen1.5-72BScience
Gemini 1.0 ProScience
Mixtral 8x7BScience
DeepSeek LLM 67BScience
Claude 2.1Science
GPT-4 Turbo (Nov 2023)Science
Yi-34BScience
Mistral 7BScience
Llama 2-70BScience
Claude 2Science
GPT-3.5 TurboScience
GPT-4 (Jun 2023)Science
GPT-4 (Mar 2023)Science
Gemma 2 9BScience
How this is built, and what it does not claim
One discipline, one test
The sources carry four mathematics tests of very different difficulty. Averaging them, or taking the best score among them, turns a solved research problem and a solved competition problem into the same figure. We bind each discipline to exactly one test and show the others separately, as what they are.
What a record is here
A model’s best published run on the discipline’s test, entered on the model’s release day. Only dated models enter the series; the rest are counted. If an older model is later re-measured with a better result, the series can change retroactively. That is why the observation time is stated at the top.
No score of our own
The index shown is the Epoch AI Capabilities Index with its confidence interval. We carry it through and attribute it rather than inventing a competing number. We crown no overall winner across disciplines.
No imputed values
A discipline without a successful public run stays empty and lowers the published coverage. Nothing is estimated to fill a row.
Who measured it is recorded
Every test names its evaluator. Epoch AI distinguishes its own evaluation, evaluation by the benchmark authors, and self-report by the model vendor. Scores of the last kind are flagged; on this page there are currently none.
Names are never merged
A result from a second source attaches only on an exact model-name match. Anything without an exact counterpart is left out and counted, rather than guessed into place. That is why the abstraction discipline is thinly populated.
A benchmark score is a snapshot under one test setup. It does not establish that a model is intelligent, safe, or suitable for a given task.
Frequently asked questions
+Is this an IQ test?
No. A language model has no IQ in the human sense, and we do not calculate one. What we keep is a record book across five named tests. Where Epoch AI publishes its Capabilities Index we carry it through with its confidence interval and attribute it.
+What counts as a record here?
A model’s best published run on the discipline’s one test, entered on the model’s release day. Only a strictly better score breaks a record; a tie does not. If an older model is later re-measured with a better result, the series can change retroactively, which is why the page states its observation time.
+Why are your mathematics scores lower than elsewhere?
Because we show the hard test. The sources carry four mathematics tests of very different difficulty. On the competition problems several models reach 100 per cent; on the research problems no model has ever reached 90. The easy test no longer separates the frontier, so the record book scores on the hard one. You will find the easy score in each model’s detail and in the direct comparison.
+Why are some cells hatched?
Because there is no successful public run there. The gap is shown as a gap and lowers the published coverage. Values we filled in ourselves used to sit in those places; we removed them.
+Why do you crown no overall winner?
Because each discipline has its own record holder and the five tests measure different things. Any total would have to weight the disciplines, and that weighting would be our invention. Within one discipline a rank does hold, because every model there was measured on the same test.
+Why are models I have heard of missing?
A model needs scores on at least three of the five tests to appear in the model list. Everything else is named below it, together with what it was measured on. Models that appear in neither source do not show up at all.
+Where do the scores come from?
From Epoch AI’s benchmarking hub and the ARC Prize data files. Both are official, machine-readable, and permit retrieval. Every cell names its test and who evaluated it.
+How current is this?
The time of the last collection is stated at the top. The data is gathered afresh on every publication of the page; we claim no fixed cadence that the pipeline does not keep.
Sources & status
As of 15 August 2026
Benchmark scores are snapshots under one test setup. They do not establish that a model is intelligent, safe, or suitable for a given task.
Which model fits your use case?
A record does not settle that. We put the right model into your process and tell you what we based that on.