Skip to content

Frontier Models In Mathematics Sampled Or Sealed

rg78803 edited this page Oct 10, 2026 · 1 revision

Frontier models in mathematics β€” a share over sampled runs, and one sealed reading of every key and every answer

Why this page exists

Frontier labs print mathematics figures built from sampled runs β€” one sample, or 4, 8, 64 or 1,000 β€” a checker the lab chose marks the answers, and an average, a vote or another model's score makes the printed figure (Β§2). Run it again and the answers can change. On the 30 AIME 2025 problems, 465 of 870 model–problem pairs returned more than one answer across four runs (counted from MathArena's published rows, theirs; Β§1). One prompt sent 1,000 times at temperature 0 came back as 80 different completions in 1,000 runs (Thinking Machines Lab, theirs). They train their models and sample them.

Affine.Earth trains nothing and samples nothing. One law of exact whole numbers reads the bytes: 12 Swift files, 0 imports, no float. Study 56 puts that law to MATH, a public benchmark of 12,500 competition problems, and to OpenAI's prm800k release, 815,632 scored attempts at 500 of them. It answers what a sample cannot: which answers equal their key exactly. Whether the nine Affine.Earth cells read them the same way is not known: no cell runs Study 56 yet (Β§5). This page sets the two side by side, every figure of theirs labelled theirs.

What Study 56 proved

  1. One verdict for every answer the law can hold. All 12,500 keys and 815,612 of OpenAI's 815,632 attempts have one verdict each; the other 20 attempts are longer than the law's room of 10,000 digits, and they are named and not read. The law was designed and sealed without the marks, and committed before the marks were read again (an earlier reading had read them once, on 2026-10-07). Run again after the seal, all 815,632 attempts reproduced their sealed outcome. Run twice on one Mac, the law wrote byte-identical files (Β§4). The Affine IDE's agent swarm seals one set of bytes with 1, 2 or 5 agents.
  2. 1,967 of MATH's 12,500 answer keys are not a whole number or a fraction of whole numbers, the only values the law holds, and the law names why for each. 595 hold a root Β· 383 hold letters Β· 265 hold pi Β· 256 are written with a decimal point Β· 165 are words Β· 99 hold the imaginary unit Β· 91 hold infinity Β· 71 hold a container of containers Β· 19 use a form outside the law's grammar Β· 11 are unions of sets Β· 6 hold a trigonometric function Β· 4 hold no answer Β· 1 holds a logarithm Β· 1 has a comma between digit groups. Of OpenAI's 500 problems, 83: 24 roots, 20 letters, 10 pi, 9 decimals, 7 words, 6 the imaginary unit, 4 infinity, 2 unions of sets, 1 container of containers.
  3. The law decides none of the 134,995 attempts on those 83 problems, and their mark reads correct on 73,200 of them (theirs). 47,027 of the 134,995 copy the key byte for byte, and a copy of a text the law cannot read is still not read to a value. How many of the 73,200 are such copies is not known.
  4. Every time the law reads an answer unequal to the key, their mark also reads false: 243,578 of 243,578, and true on none.
  5. 4,825 answers equal to the key are marked false (theirs). Where the law reads an answer equal to the key, their mark is_correct reads false on 4,825 attempts in 49 problems (theirs). 1,793 attempts answer 864 and are marked false; their copy of the key reads 864\mbox{inches}^2, and MATH's key is 864. \frac{2187}{5625} is marked false; it reads to 243/625, and the key is 243/625.
  6. Their grader normalises answers with sympy, with no sympy version pinned and no timeout (their grader.py, theirs). No file of theirs states how is_correct was produced.
  7. Side by side, never subtracted. Among OpenAI's own attempts at each problem, the answer most of them hold, counted by exact value, equals the key on 293 of the 417 problems whose key the law reads to a value. The other 83 have no key value to equal. Theirs (Let's Verify Step by Step, Figure 3, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it counts problems; their eval.py prints their two reward-model picks as means over 400 trials, not as counts. Their printed figures stand beside the law's as published; the two cannot be subtracted (Β§4).

Theirs is a sample. This reading came out byte-identical every time it ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.

Status: PUBLISHED FIGURES SET BESIDE STUDY 56 β€” 2026-10-09 β€” what the labs that build frontier models print about their mathematics, what independent evaluators counted run by run, and what Affine.Earth's substrate computes from sha-pinned bytes, side by side. Every figure of theirs is labelled theirs, with its source, its date and the conditions it was printed under; where no date was found on a source, its row says so. Every figure of the substrate's names its commit in Β§4: the blind seal, made before OpenAI's marks were read again, and the comparison, made after; a count added up on 2026-10-09 from a committed table says so beside it. Integers only: a published share is written as a whole count k of N where the source states N and k is whole; otherwise the page says what kind of number it is and does not print it. The sources of the labs, the evaluators and the run-to-run studies were read on 2026-10-09; the paper, the repository and the scored file's headers, the sources of the attempts the substrate read, were read on 2026-10-08. Study 56 β€” MATH, read exactly Β· Program index Β· Study board Β· Study 55 β€” IceCube: the light in the ice


For every reader

When a lab says its model scores so much on a mathematics test, each row below says how that figure was built, as far as its source says: how many runs, which checker marked them (a script, another model, or people), and under which conditions the lab set β€” which tools, how much thinking time, how many samples. Run again, the answers can change, and the evaluators in Β§1 counted how often they do.

Affine.Earth does a different thing, and it solves no problem. Study 56 reads the answer key that MATH's authors wrote for each of the 12,500 problems in MATH, a public set of competition problems, and decides whether that key closes to an exact integer or fraction, or cannot close, with the reason named β€” a root, a letter, pi, a decimal point, a word. It does not check that a key is right. It then reads the 815,632 answers OpenAI published in 2023 with its paper Let's Verify Step by Step, for 500 of those problems. They were written by the paper's large-scale generator, a model it finetunes from the base model it names GPT-4, as 1,860 solutions per problem, drawn so that a reward model or a vote could pick one. The law decides each answer against the key: equal, unequal, or cannot close, with its reason. Anyone can run the law on the same committed bytes. The seals came out byte-identical every time it ran, twice on one Mac and with 1, 2 or 5 agents; whether the nine Affine.Earth cells write the same bytes is not known until they answer Study 56. The substrate has read none of the models named in Β§2 and Β§3. The two sides are set next to each other here, with nothing added.


1 Β· The same question, different answers

What the evaluators measured, run to run (theirs)

who (theirs) what was run what they counted, as integers source
Thinking Machines Lab one prompt, sent 1,000 times at temperature 0 to Qwen/Qwen3-235B-A22B-Instruct-2507 in non-thinking mode, 1,000 tokens per completion; the prompt is about Richard Feynman, not a mathematics question 80 different completions in 1,000 runs; the most common appeared 78 times; all 1,000 agreed for the first 102 tokens, and at token 103, 992 went one way and 8 another. With their batch-invariant kernels: 1 completion in 1,000 runs. The cause they name: kernels whose order of addition changes with batch size Defeating Nondeterminism in LLM Inference, 2025-09-10
Atil and others the MMLU college-mathematics questions, 10 runs per model, temperature 0, top-p 1, a fixed seed, the same infrastructure, 5-shot GPT-4o: 88 of 100 in its best run, 44 of 100 in its worst; 50 questions got one parsed answer in all 10 runs; 0 questions got one raw output in all 10 runs. Llama-3-70B: 85 best, 22 worst, 0 with one raw output. Mixtral 8x7B: 75 best, 3 worst, 0 with one raw output. The paper prints shares; 100 is the row count of the public test split (cais/mmlu), read 2026-10-09 Non-Determinism of "Deterministic" LLM Settings, v1 2024-08-06; figures from v5, 2025-04-02
Atil and others 5 models, 8 tasks, zero-shot and few-shot: 80 conditions, 10 runs each their reading: no model gave the same accuracy on every task, and identical output strings were rarer still; the spreads are printed as shares over task sets of different sizes, not counts of problems same paper
Yuan and others DeepSeek-R1-Distill-Qwen-7B, greedy decoding, the same seed, the same prompts, run under 12 configurations (2 GPU types Γ— 2 GPU counts Γ— 3 batch sizes) on the 30 AIME'24 problems the accuracy moved between configurations (printed as a share, not a count); response length differed by up to 9,000 tokens. The cause they name: floating-point addition is not associative at limited precision Give Me FP32 or Give Me Death?, v1 2025-06-11
Hochlehnert and others the same weights and prompts on 5 compute clusters, and under 2 evaluation frameworks AIME'24 scores moved with the cluster and with the framework; printed as shares (means over seeds), which their text says can change model rankings; not counts A Sober Look at Progress in Language Model Reasoning, v1 2025-04-09
MathArena how it scores every model each model answers each question 4 times; the score is the average of the 4, with no majority vote; final answers are parsed into sympy and also checked by an LLM judge, and every disagreement between the two is looked at by hand MathArena: Evaluating LLMs on Uncontaminated Math Competitions, v1 2025-05-29

The same AIME 2025 problems, 4 runs each β€” counted for this page from MathArena's published rows

MathArena publishes every run's parsed answer and its grade. Counting those rows for AIME 2025 I and II (30 problems), for 29 of the 31 models in the files, 4 runs each:

count
model–problem pairs 870
runs 3,480
runs graded correct (MathArena's grading) 1,979
pairs graded correct in all 4 runs 375
pairs graded correct in none of the 4 269
pairs graded correct in 1, 2 or 3 of the 4 226
pairs where the 4 runs did not all return one answer 465 of 870

Some of the answers, as returned:

model (theirs) runs graded correct problems with more than one answer across 4 runs one problem, its 4 answers
DeepSeek-R1 84 of 120 14 of 30 AIME 2025 I problem 11, key 259: 259, 4423, 758, 0 Β· I problem 13, key 204: 829, 829, 622, 754 Β· II problem 7, key 237: 237, 237, 60671, 60671
o3 (high) 107 of 120 6 of 30 AIME 2025 I problem 15, key 735: 147, 423, 441, 343
o4-mini (high) 110 of 120 5 of 30 β€”
o1 (medium) 96 of 120 11 of 30 β€”

Datasets MathArena/aime_2025_I_outputs (revision 05b2f50ac7cb76ff) and MathArena/aime_2025_II_outputs (revision 9809200ca4e719df), 1,860 rows each; created 2025-04-16, last modified 2026-05-15, read 2026-10-09; each model at the settings MathArena lists for it. A run with no parsed answer counts as one answer value. The counts are this page's, made from MathArena's rows; MathArena does not print them.

The substrate's side of this question

The substrate's answer to "run it again" is in Β§4, under The same bytes, the same seals: its law, run twice as separate processes on one Mac, wrote byte-identical files. Two of the studies above changed the hardware; on more than one machine the substrate's result is not known until the nine Affine.Earth cells answer Study 56.


2 Β· What the labs publish (theirs)

Every row in this section is the lab's own figure about its own model, from its own page. Where the source prints a share and states no problem count, the share is not turned into a count and is not printed; the cell says what kind of number it is. Grouped by test.

The words in these rows, in plain words

  • pass@1 (DeepSeek writes Pass@1): the share of problems marked correct when one answer per problem is marked; where several samples are drawn per problem, it is that share averaged over them.
  • consensus of 64, cons@64, majority vote over 64 samples: the answer given most often across 64 samples is the one marked. Consensus of 8 is the same over 8.
  • re-ranking of 1,000 samples, best-of-N: a scoring model picks one of N samples, and only that one is marked.
  • em_maj1@1 (Meta): exact match of the answer, majority vote over 1 sample β€” that is, one sample, marked by exact match.
  • 4 shots, 5-shot: the prompt carries 4 or 5 worked examples before the question.
  • temperature: how widely the model samples among its possible next tokens; at temperature 0 it is asked to take the most likely token every time, which is also what greedy decoding means.
  • top-p: the model samples only among its most likely next tokens that together hold a share p of the probability; top-p 1 keeps all of them.
  • bf16: bfloat16, a 16-bit floating-point format in which the model's weights are held.
  • 128K context, thinking capped at 96k tokens: how many tokens (pieces of words) the model may read, or may write before it answers.
  • reasoning effort, thinking, Think Max, heavy mode: settings on the lab's own service for how long the model may work, or how many attempts it runs side by side, before it answers.
  • Pass@8192, 1,024 samples per problem (formal proofs): a problem counts as solved if any one of that many attempts is a proof the checker accepts.

MATH β€” the problem set Study 56 reads

lab model their figure, as integers (theirs) conditions (theirs) date source
DeepSeek DeepSeek-R1, MATH-500 printed as a mean over k samples per question (k typically 4 to 64); not a count of problems up to 32,768 generated tokens; temperature 6/10, top-p 95/100 2025-01-22 arXiv paper
Meta Llama 4 Maverick, Llama 4 Scout (pre-trained); Llama 3.1 405B and 70B beside them printed as shares; problem count not stated; not converted base models; 4 shots; metric em_maj1@1; bf16 weights 2025-04-05 model card

AIME 2024

lab model their figure, as integers (theirs) conditions (theirs) date source
OpenAI o1, o1-preview, GPT-4o printed as an average score out of 15 per exam, for one sample, a consensus of 64 samples and a re-ranking of 1,000 samples; not a whole count of problems o1 at its maximal test-time compute; re-ranking uses a learned scoring function; tools not mentioned; grader not stated 2024-09-12 Learning to reason with LLMs
DeepSeek DeepSeek-R1 printed as a mean over k samples per question; not a count of problems as for MATH-500 above 2025-01-22 arXiv paper
DeepSeek DeepSeek-R1-Zero pass@1 printed as a mean over samples (16 responses per question in the training curve); cons@64 printed as a share; not converted cons@64 is a majority vote over 64 samples 2025-01-22 arXiv paper
xAI Grok 3 mini printed as a share; not converted conditions not stated in the sentence 2025-02-19 Grok 3
Mistral AI Magistral Medium, Magistral Small each printed as two shares, the second with a majority vote over 64 samples; not converted conditions of the first figure not stated in the sentence 2025-06-10 Magistral

AIME 2025

lab model their figure, as integers (theirs) conditions (theirs) date source
OpenAI o4-mini, o3 consensus of 8: every problem, for both; pass@1 printed as shares a Python interpreter; high reasoning effort; OpenAI writes that tool results should not be compared with no-tool results 2025-04-16 Introducing o3 and o4-mini
OpenAI GPT-5 pro, GPT-5, o3 GPT-5 pro with Python: every problem at pass@1; every other bar printed as a pass@1 share; not converted pass@1; high effort; with and without Python; with and without thinking 2025-08-07 Introducing GPT-5
OpenAI GPT-5.2 Thinking, GPT-5.2 Pro, GPT-5.1 Thinking every problem for GPT-5.2 Thinking and GPT-5.2 Pro; GPT-5.1 Thinking printed as a share no tools; maximum reasoning effort in the API; research environment 2025-12-11 Introducing GPT-5.2
Google DeepMind Gemini 3 Pro, Gemini 2.5 Pro Gemini 3 Pro with code execution: every problem; the no-tool figures are pass@1 shares averaged over an unstated number of trials; not converted pass@1, a single attempt, no majority vote; default sampling; results as of November 2025 2025-11 Gemini 3 Pro evaluation
xAI Grok 3 (Think) printed as a share; not converted majority vote over 64 samples; the exam was released 7 days before the post; the model still in training 2025-02-19 Grok 3
xAI Grok 4, Grok 4 Heavy Grok 4 Heavy: every problem; the other figures printed as shares the tool is Python; Heavy uses parallel test-time compute; sample counts not stated 2025-07-09 Grok 4
xAI Grok 4 Fast, Grok 4 printed as shares; not converted pass@1; no tools 2025-09-19 Grok 4 Fast
DeepSeek DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking printed as Pass@1 shares; sample count not stated; not converted temperature 1; 128K context 2025-12-01 paper
Moonshot AI Kimi K2 Thinking no tools: printed as a mean over 32 runs; with Python: a mean over 16 runs; not counts of problems. Heavy mode: every problem temperature 1; thinking capped at 96k tokens; heavy mode rolls out 8 trajectories in parallel and aggregates them; heavy run count not stated 2025-11-04 model card

AIME 2026

lab model their figure, as integers (theirs) conditions (theirs) date source
Alibaba (Qwen team) Qwen3.5-397B-A17B, Qwen3-Max-Thinking printed as shares that are not whole counts of a 30-problem set; run count not stated; not converted sampling for the table not stated 2026-02-16 model card
Moonshot AI Kimi K2.6, Kimi K2.5 printed as shares; run count not stated; not converted thinking mode; temperature 1, top-p 1; generation capped at 98,304 tokens 2026-04-14 model card

HMMT

lab model their figure, as integers (theirs) conditions (theirs) date source
OpenAI GPT-5.2 Pro, GPT-5.2 Thinking, GPT-5.1 Thinking β€” February 2025 GPT-5.2 Pro: every problem; the other two printed as shares no tools; maximum effort 2025-12-11 Introducing GPT-5.2
xAI Grok 4, Grok 4 Heavy β€” 2025 printed as shares; not converted with and without Python; Heavy in parallel 2025-07-09 Grok 4
xAI Grok 4 Fast, Grok 4 β€” 2025 printed as shares; not converted pass@1; no tools 2025-09-19 Grok 4 Fast
DeepSeek DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking β€” February and November 2025 printed as Pass@1 shares; not converted temperature 1; 128K context 2025-12-01 paper
DeepSeek DeepSeek-V4-Pro-Max β€” February 2026 printed as a Pass@1 share; not converted Think Max mode, with a special system prompt 2026-04-22 model card
Alibaba (Qwen team) Qwen3.5-397B-A17B, Qwen3-Max-Thinking β€” February and November 2025 printed as shares; not converted sampling not stated 2026-02-16 model card
Moonshot AI Kimi K2 Thinking β€” 2025 printed as means over 32 runs (no tools) and 16 runs (Python); heavy mode printed as a mean, not a count as above 2025-11-04 model card
Moonshot AI Kimi K2.6, Kimi K2.5 β€” February 2026 printed as shares; not converted thinking mode 2026-04-14 model card

Olympiad proofs β€” IMO, USAMO, CMO, Putnam

lab model their figure, as integers (theirs) conditions (theirs) date source
Google DeepMind AlphaProof with AlphaGeometry 2 β€” IMO 2024 28 of 42 points; 4 of 6 problems (two combinatorics problems unsolved) problems translated by hand into formal language first; one solved in minutes, others took up to three days; scored by Timothy Gowers and Joseph Myers; that year's gold began at 29 points, reached by 58 of 609 contestants 2024-07-25 AI solves IMO problems at silver-medal level
OpenAI an unnamed experimental reasoning LLM β€” IMO 2025 35 of 42 points; 5 of 6 problems; no solution for problem 6 two 270-minute sessions; no tools or internet; each proof graded by three former IMO medalists, by unanimous consensus; grading by IMO coordinators not claimed 2025-07-19 thread by Alexander Wei
Google DeepMind Gemini Deep Think, advanced version β€” IMO 2025 35 of 42 points; 5 of 6 problems graded and certified by IMO coordinators with the student criteria; natural language from the official statements; within the 270-minute limit 2025-07-21 Gemini Deep Think at the IMO
Harmonic Aristotle β€” IMO 2025, formal 5 of 6 problems, as formal solutions Lean 4 with Mathlib; geometry problems handled by a solver working outside Lean 2025-10-01 arXiv paper
DeepSeek DeepSeekMath-V2 β€” IMO 2025, CMO 2024, Putnam 2024 IMO 2025: 5 of 6 problems fully solved (point share not converted). CMO 2024: 4 of 6 fully solved, partial credit on 1. Putnam 2024: 118 of 120 points, 11 of 12 problems fully solved 64 proof samples, each checked by 64 verification analyses, up to 16 refinement rounds; DeepSeek's own experts assessed the top proofs; the paper gives the highest human Putnam 2024 score as 90 2025-11-27 repository
DeepSeek DeepSeek-V3.2-Speciale β€” IMO 2025, CMO 2025 IMO 2025: 35 of 42 points. CMO 2025: 102 of 126 points up to 128k generated tokens; no tools or internet; a generate–verify–refine loop; the grader is not named 2025-12-01 paper
dots team (Xiaohongshu, RedNote) dots-note-3.0, internal version β€” IMO 2026 42 of 42 points; 6 of 6 problems invited by the IMO Organizing Committee and marked through its official marking; reads the original LaTeX; uses Python; time limit not stated; 7 of 666 human contestants also scored 42; the affiliation is from press reports 2026-07 dots at IMO 2026
Huawei Celia (Xiaoyi Agent system) β€” IMO 2026 42 of 42 points as reported by IT Home from Huawei's announcement; per-problem points and grading channel not given 2026-07-22 IT Home
xAI Grok 4, Grok 4 Heavy β€” USAMO 2025 printed as shares of the score; point total and run count not stated; not converted Heavy with Python and parallel test-time compute; who graded the proofs is not stated 2025-07-09 Grok 4

Answer and proof benchmarks built from olympiads

lab model their figure, as integers (theirs) conditions (theirs) date source
Google DeepMind Gemini 3 Pro, Gemini 2.5 Pro β€” MathArena Apex printed as shares; not converted Google states these were reported by matharena.ai 2025-11 Gemini 3 Pro evaluation
Google DeepMind Gemini Deep Think, January 2026 version β€” IMO-ProofBench Advanced printed as a share, rising as inference-time compute scales; not converted graded by human experts 2026-02-11 Accelerating discovery with Gemini Deep Think
DeepSeek DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking β€” IMOAnswerBench printed as Pass@1 shares; not converted temperature 1 2025-12-01 paper
DeepSeek DeepSeek-V4-Pro-Max β€” IMOAnswerBench, Apex, Apex Shortlist printed as Pass@1 shares; not converted Think Max mode 2026-04-22 model card
Alibaba (Qwen team) Qwen3.5-397B-A17B, Qwen3-Max-Thinking β€” IMOAnswerBench printed as shares; not converted sampling not stated 2026-02-16 model card
Moonshot AI Kimi K2 Thinking β€” IMO-AnswerBench printed as a mean over 8 runs; not a count of problems thinking capped at 128k tokens 2025-11-04 model card
Moonshot AI Kimi K2.6, Kimi K2.5 β€” IMO-AnswerBench printed as shares; not converted thinking mode 2026-04-14 model card

FrontierMath

lab model their figure, as integers (theirs) conditions (theirs) date source
OpenAI o3-mini (high) printed as lower bounds on a share; problem count not stated; not converted; OpenAI calls the numbers provisional Python tool; first attempt 2025-01-31 OpenAI o3-mini
OpenAI GPT-5.2 Thinking, GPT-5.1 Thinking β€” Tiers 1–3 and Tier 4 printed as shares; problem count not stated; not converted with Python; maximum effort; research environment 2025-12-11 Introducing GPT-5.2
OpenAI GPT-5.4, GPT-5.4 Pro, GPT-5.2, GPT-5.2 Pro printed as shares; not converted. The GPT-5.2 figures on this page differ from those on the GPT-5.2 page, and GPT-5.2 Pro carries a Tier 4 figure here that the GPT-5.2 page does not print xhigh effort unless stated; tool setting not stated 2026-03-05 Introducing GPT-5.4
OpenAI GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5 β€” v2 printed as shares; not converted effort and tools not stated 2026-07-09 GPT-5.6
OpenAI GPT-6 Astra, GPT-5.6 Sol β€” Tier 4 v2 printed as shares; not converted the maximum at any effort; the page prints no date (third-party reports give 2026-09-03); as first read on 2026-10-09 β€” a second read that day returned HTTP 403, so this row was not re-verified 2026-09-03 GPT-6 Astra

Formal proof sets

lab model their figure, as integers (theirs) conditions (theirs) date source
DeepSeek DeepSeek-Prover-V2-671B MiniF2F-test: 217 of 244 problems at Pass@8192 (the paper's Table 2). PutnamBench: the repository prints 49 of 658 and states no sample budget. The paper, in its second version, ran the model on the 649 of the 658 problems that work with Lean 4.9.0; its first run solved 49 at 1,024 samples per problem; after the proofs went to the benchmark's maintainers, 2 problems were excluded for misformulated statements, and its Table 4 prints 47 of 658 Lean 4 proofs checked by the Lean proof assistant repository 2025-04-30; paper v2 2025-07-18 repository Β· paper
Google DeepMind AlphaGeometry 2 β€” IMO geometry problems of the past 25 years printed as a share; problem count not stated; not converted formal geometry system; measured before IMO 2024 2024-07-25 AI solves IMO problems at silver-medal level

Research problems

lab model their figure, as integers (theirs) conditions (theirs) date source
OpenAI GPT-6 Astra β€” bounded gaps between primes helped establish a bound of 186 for infinitely many prime pairs; earlier bounds 246, then 240 (Julia Stadlmann) OpenAI states it shares the proofs, an abridged chain of thought and verification materials; who verified the proof is not known from this page; as first read on 2026-10-09 β€” a second read that day returned HTTP 403, so this row was not re-verified 2026-09-03 GPT-6 Astra
Google DeepMind Aletheia, built on Gemini Deep Think β€” ErdΕ‘s problems 4 autonomous solutions among 700 open problems evaluated uses a natural-language verifier, Google Search and web browsing 2026-02-11 Accelerating discovery with Gemini Deep Think

3 Β· What independent evaluators measured (theirs)

These are figures counted by someone other than the model's maker. Where a source prints one figure that is a mean over runs, the per-run counts it also publishes are given instead. No date was found on MathArena's competition tables as read, so their rows give the read date and, where the source names it, the competition's month.

How the evaluators run the models

evaluator how what it means in counts source
MathArena (SRI Lab, ETH Zurich, with INSAIT) every model 4 times on every problem; the score is the average of the 4; a mark beside a model means it was released after the competition 4 runs per problem per model matharena.ai, no date found on the page; read 2026-10-09
Vals AI AIME 2024 and 2025, 60 questions, 8 runs each; the boxed answer compared with the reference 480 answers per model: Gemini 3.1 Pro Preview (02/26) 471 of 480; GPT 5.2 465 of 480. The page no longer runs new models vals.ai, updated 2026-04-16
Artificial Analysis AIME 2025 as pass@1 with 10 repeats; MATH-500, a 500-problem subset of MATH; both now retired AIME 2025: 30 questions Γ— 10 repeats = 300 answers per model; graded by script with SymPy and an LLM equality checker as backup methodology, no date found on the page; read 2026-10-09

IMO 2025 and IMO 2024

who model per-run or whole counts conditions source
MathArena GPT-5 (high) runs 1 to 4: 19, 15, 14, 16 of 42 β€” 64 of 168. By problem, runs 1 to 4: P1 7 1 1 0 Β· P2 0 0 0 0 Β· P3 2 2 1 2 Β· P4 4 5 5 7 Β· P5 6 7 7 7 Β· P6 0 0 0 0 marked as released after the competition IMO 2025 table, no date found on the table; competition July 2025; read 2026-10-09
MathArena Gemini 2.5 Pro runs 1 to 4: 17, 10, 11, 15 of 42 β€” 53 of 168; both judges gave the same grade on every run. By problem: P1 1 2 0 1 Β· P2 0 0 0 0 Β· P3 7 2 4 7 Β· P4 5 2 3 3 Β· P5 4 4 4 4 Β· P6 0 0 0 0 best of 32, chosen in a bracket tournament where the model judged its own answers; each answer graded by 2 of 4 human judges IMO 2025 table, as above; method: matharena.ai/imo, 2025-07-17
MathArena (blog) the July 2025 post bronze needed 19 of 42; the top model's figure is printed as a mean over 4 runs, not a count β€” its per-run counts are 17, 10, 11 and 15 of 42 4 judges with IMO-level background, grading in pairs matharena.ai/imo, 2025-07-17
UK Mathematics Trust the IMO 2025 boundaries the students were cut at gold from 35 of 42, silver from 28, bronze from 19; 110 countries Sunshine Coast, Australia UKMT news, 2025-07-18
Google DeepMind Gemini Deep Think, advanced version 35 of 42 points; 5 of 6 problems graded and certified by IMO coordinators; the IMO President is quoted confirming the 35 DeepMind blog, 2025-07-21
OpenAI, as quoted by Simon Willison the experimental reasoning LLM 35 of 42 points; 5 of 6 problems; 3 graders per proof three former IMO medalists by unanimous agreement; whether IMO coordinators graded it is not stated simonwillison.net, 2025-07-19
OpenAI, on GitHub the same model's proofs 5 proof files; 0 for problem 6 first commit 2025-07-18 aw31/openai-imo-2025-proofs
Google DeepMind AlphaProof with AlphaGeometry 2 β€” IMO 2024 28 of 42; 4 of 6; AlphaGeometry 2 proved Problem 4 within 19 seconds of receiving it formalized; gold began at 29 points, 58 of 609 contestants problems formalized by hand first; scored by Gowers and Myers DeepMind blog, 2024-07-25

USAMO 2025 and 2026 β€” proofs, 6 problems of 7 points, 4 runs each

who model points per run, of 42 conditions source
Petrov, Dekoninck and others ("Proof or Bluff?"); the grades as MathArena stores them DeepSeek-R1 0, 7, 0, 1 β€” 8 of 168, the same from both judges graded within hours of the problems' release; four expert judges, two per problem per-run grades: USAMO 2025 table, its stored trace grades, no date found on the table, read 2026-10-09; the study: arXiv v1, 2025-03-27
MathArena Gemini 2.5 Pro judge 1: 14, 7, 6, 14; judge 2: 14, 7, 8, 12 β€” each judge 41 of 168 released after the competition USAMO 2025 table, no date found on the table; read 2026-10-09
MathArena DeepSeek-R1-0528 judge 1: 14, 7, 15, 15 (51 of 168); judge 2: 14, 7, 14, 15 (50 of 168); one split, problem 6 run 3, 1 and 0 released after the competition same
MathArena o4-mini (high) 7, 14, 4, 7 β€” 32 of 168, both judges released after the competition same
MathArena GPT-5.4 (xhigh) β€” 2026 42, 42, 41, 35 β€” 160 of 168 graded by an LLM jury of three models that includes GPT-5.4, the lowest of the three scores kept, then every solution reviewed by hand USAMO 2026 table, no date found on the table; competition March 2026; read 2026-10-09; jury: Proof, Not Bluff, 2026-03-28
MathArena GPT-5.5 (xhigh) β€” 2026 41, 41, 42, 41 β€” 165 of 168 released after the competition same
MathArena Gemini 3.1 Pro Preview β€” 2026 33, 35, 32, 25 β€” 125 of 168 Gemini-3.1-Pro also sits on the grading jury same
MathArena DeepSeek-v4-Pro (Max) β€” 2026 28, 25, 30, 19 β€” 102 of 168 released after the competition; open weights same

Putnam 2025 β€” 12 proof problems of 10 points, graded by the Putnam graders

system points of 120 conditions source
DeepSeek-v3.2-Speciale Agent 103 (by problem: 10 10 10 9 2 9 10 10 10 10 10 3); within the top 3 of 4,329 human participants one submission, chosen before the contest; the graders knew the answers were AI-written, not which model; up to 1,000 model queries per problem (cut from 65,000) matharena.ai/putnam, 2026-02-26
Gemini Deep Think 3 (12/25) 100 one submission each, graded in the Putnam graders' second round Putnam 2025 table, no date found on the table, read 2026-10-09; with the post of 2026-02-26
GPT-5-Pro 93 same same
Gemini 3 Pro (preview) 91 same same
DeepSeek-v3.2-Speciale 71 same same
GPT-5.1 (high) 58 same same

The evaluation was partially supported by an evaluation grant from Google DeepMind, which also makes two of the six systems graded (matharena.ai/putnam, 2026-02-26).

Final-answer contests, 2026, 4 runs each

who model correct answers per run conditions source
MathArena GPT-5.2 (high) β€” AIME 2026, 30 problems 30, 29, 30, 29 of 30 β€” 118 of 120 answers checked against the official integers AIME 2026 table, no date found on the table; competition February 2026; read 2026-10-09
MathArena GPT-5.2 (high) β€” HMMT February 2026, 33 problems 31, 32, 33, 32 of 33 β€” 128 of 132 not marked as released after the competition HMMT Feb 2026 table, no date found on the table; competition February 2026; read 2026-10-09
MathArena (Sun, Dekoninck, Vechev) Gemini 3.1 Pro, 4 attempts on each of 176 problems from contests since August 2025 162 solved in all 4 attempts; 14 solved at least once; 0 unsolved in every attempt β€” so no problem met the Apex rule. MathArena concludes recent public contest problems no longer serve as hard frontier benchmarks the Apex rule keeps a problem only if it is unsolved in every attempt matharena.ai/no_final_answer, 2026-05-12

BrokenArXiv β€” 31 subtly false statements built from arXiv papers

model points per run how graded source
GPT-5.4 (xhigh) 28, 21, 23 of 62 and 22 of 60 (problem 21 has 3 runs) β€” 94 points over 123 graded answers 0 points = proves the false statement as given; 1 = silently repairs it; 2 = any other response, such as saying it is false; judged by Gemini-3.1-Pro, an LLM; the prompt asks the model to try to prove the statement BrokenArXiv, February, no date found on the table, read 2026-10-09; method, 2026-03-13
Gemini 3.1 Pro Preview 12, 10, 10, 14 of 62 β€” 46 of 248 same same

FrontierMath (Epoch AI)

what counts source
who commissioned it OpenAI commissioned 300 core problems and 50 Tier 4 problems: 350. 53 core solutions are held out from OpenAI; 20 of the 50 Tier 4 problems are held out; the work is supported by OpenAI About FrontierMath, no publication date; its latest update note is dated 2026-09-22; read 2026-10-09
the 2026-06-12 correction 123 problems corrected in Tiers 1–3 and 12 in Tier 4; 5 and 7 removed β€” 147 of 350 in all, added here; now 338 problems (295 + 43), 12 of them public (10 from Tiers 1–3, 2 from Tier 4) FrontierMath Tier 4 v2, 2026-06-12
o3 and GPT-5, run by Epoch on the private core set of 2025-02-28 o3-2025-04-16_high and gpt-5-2025-08-07_high: each printed as a mean score; the file states no problem count, so neither is converted. OpenAI's December 2024 claim was on an earlier version, FrontierMath_11-26-24, which held 180 questions Epoch benchmark data, file frontiermath.csv, runs started 2025-11-16 and 2025-11-13; the 180: About FrontierMath, as above
the human baseline at MIT 8 teams of four or five, 23 problems, 270 minutes, internet allowed: all teams together, 8 of 23; the average team is printed as a mean over 8 teams, not a count About FrontierMath, as above

MATH Level 5 and HLE

who what counts source
Epoch AI MATH Level 5: the 1,324 hardest test problems of the same MATH set Study 56 reads gpt-5-2025-08-07_high: printed as a mean score; samples per problem not stated; not a count of problems, and not converted. Three scorers; the plotted one asks gemini-1.5-flash-002 whether the answer matches the reference MATH Level 5, no date found on the page, read 2026-10-09; the run started 2025-10-29 (benchmark data, file math_level_5.csv)
Center for AI Safety and Scale AI HLE: how often experts disagreed with the reference answers 2 audit rounds of 200 sampled questions each; the disagreement rate is an estimate from samples, not a count of questions HLE paper, v11 2026-07-28

4 Β· What the substrate computes

Study 56 reads 8 published files with Affine's own Swift, each pinned by its byte count and sha256: MATH, and OpenAI's prm800k release β€” the repository at commit 7ecc7947 (2023-06-01) and its scored file of 2,159,374,091 bytes (last modified 2023-05-25), which the repository's README says holds the large-scale samples behind Figure 3 of Let's Verify Step by Step (arXiv 2305.20050, v1 2023-05-31). Every value is an integer or a fraction of integers; a decimal point refuses its text. Each count below is the law's own print, committed in the folder and at the commit named, unless the row says it was added up here. Study 56 gives every table in full.

One verdict per problem, and per attempt

what the law computes count commit
MATH problems whose key closes to an exact Affine value 10,533 of 12,500 f2d890e48
MATH problems whose key cannot close, each with its named reason 1,967 of 12,500 f2d890e48
MATH problems not read 0 f2d890e48
the 500 problems OpenAI scored: key closes / cannot close 417 / 83 f2d890e48
OpenAI's 815,632 attempts, each decided against the MATH key: EQUAL 362,184 f2d890e48
UNEQUAL 243,578 f2d890e48
CANNOT_CLOSE, each with its reason 209,850 f2d890e48
NOT_READ_ROOM (an answer past 10,000 digits) 20 f2d890e48
run again after the seal: attempts that reproduce their sealed outcome, the 20 not read among them all 815,632; 0 left out of the join to the sealed records 3c25e8cf9

Why the 83 keys cannot close, by problems: a root 24 Β· letters 20 Β· pi 10 Β· a written decimal 9 Β· words 7 Β· the imaginary unit 6 Β· infinity 4 Β· a union of sets 2 Β· a container of containers 1. No attempt on these 83 can close, whatever it writes; 134,995 attempts sit on them (3c25e8cf9).

How many different values the attempts wrote

These attempts were drawn for selection. Their paper says its generator wrote 1,860 solutions for each test problem (Figure 3 caption), for a reward model or a majority vote to pick one, and their README says the file holds the samples used to evaluate their reward models: up to 1,860 per test problem, fewer where a solution reached no answer within 1,024 tokens. The paper does not state the sampling temperature of these solutions. The law counted between 152 and 1,860 attempts per problem in the file. A count of different values over samples drawn this way can only grow as more samples are drawn.

what the law counted count commit
attempts whose answer reads to an exact Affine value 618,869 of 815,632 f2d890e48
attempts whose answer reads to no Affine value 196,743 f2d890e48
attempts past the room (an answer longer than 10,000 digits) 20 f2d890e48
different exact values written, counted once per problem and added over the 500 33,586 added up 2026-10-09 from the 3c25e8cf9 table
problems on which the attempts wrote more than one value 471 of 500 (414 of the 417 whose key closes) counted 2026-10-09 from the 3c25e8cf9 table
problems on which they wrote exactly one value 15 of 500 (3 of the 417) counted 2026-10-09 from the 3c25e8cf9 table
problems on which no attempt read to a value 14 of 500, all 14 among the 83 whose key cannot close counted 2026-10-09 from the 3c25e8cf9 table
the middle problem, by number of different values 32 values counted 2026-10-09 from the 3c25e8cf9 table
the most on one problem: test/algebra/2626 1,149 different values in 1,854 attempts; the key is 32,348; the value held most often, 32,372, by 78 attempts; 3 attempts equal the key 3c25e8cf9; the row found 2026-10-09
one problem where they gathered: test/algebra/1004 1,860 attempts, 6 different values; 1,821 hold 4, which is the key 3c25e8cf9
problems where two values share the top count 2 on keys that close: 1791 (key βˆ’3/8; the values 0 and 3/4 hold 104 attempts each) and 989 (key 12; 0 and 2 hold 6 each) 3c25e8cf9

Counted from evidence/study-56-v3-compare-20261008/majority/scored.compare-v3.majority.tsv, column values (law), rows 4 to 503 β€” sha256 6bc0023748dbb486… at 3c25e8cf9. The file is the law's sealed per-problem table, written after OpenAI's marks had been read; the sum and the counts of problems were made from it on 2026-10-09 and are not seals themselves.

Beside OpenAI's own mark and figures

the law's outcome their mark is_correct true (theirs) their mark false (theirs) attempts
EQUAL β€” the answer and the key hold one exact value 357,359 4,825, across 49 problems 362,184
UNEQUAL 0 243,578 243,578
CANNOT_CLOSE, with its reason β€” the law decides nothing 75,708 134,142 209,850
NOT_READ_ROOM β€” the law decides nothing 0 20 20
all 433,067 382,565 815,632

Every cell is the comparison census at 3c25e8cf9; the cannot-close row adds its 40 reason lines, added up here on 2026-10-09. Of the 75,708 attempts their mark reads true where the law decides nothing, 73,200 sit on the 83 problems whose key cannot close (census, 3c25e8cf9).

On the 4,825. Their published grader, grader.py, compares two fractions as written strings only, and its comment says an unreduced fraction must not be marked correct (rule 11a, L275-278; theirs, not run here). No file of theirs states how is_correct was produced. The law compares values; that published rule compares the written fractions. One example: on test/algebra/1072 the MATH key is written \frac{243}{625}; the answers \frac{2187}{5625}, \frac{54675}{140625} and \frac{25 \cdot 3^6}{3 \cdot 5^6} each read to 243/625, which is the key, and their mark reads false (theirs). On that problem 1,011 attempts equal to the key carry their mark true and 178 carry it false.

a vote over each problem's attempts count commit
the law's majority: the exact value held by the most attempts equals the MATH key 293 of the 417 whose key closes; on the other 83 no majority is decided 3c25e8cf9
… is unequal to the key / two values tie / both sides cannot close / the key cannot close 119 / 2 / 3 / 83 3c25e8cf9
their majority voting, best-of-1,860, in Let's Verify Step by Step (theirs) a share equal to 348/500 (theirs); the paper does not say whether it is a count of problems or a rounded mean; majority voting is not in their eval.py; how their vote matches two answers or breaks a tie is not published transcribed at 3c25e8cf9 from the paper's Figure 3 table, page 7
their process and outcome reward models, best-of-1,860 (theirs) their eval.py, which makes these two selections, prints a mean over 400 trials; not a count of problems read at 3c25e8cf9

The law's 293 counts problems among the 417 whose key closes, and it decides no majority on the other 83. Their share, 348/500, is printed over all 500, and it is not stated to be a count. The law's majority is taken over OpenAI's own attempts; it solves no problem. The two figures stand side by side and are not subtracted: the law decides no majority on 83 of the 500, so these figures cannot show where any difference lies.

The same bytes, the same seals

what was run twice, or checked again result commit
the blind seal, run A and run B on one Mac: separate processes, separate HTTP streams of OpenAI's 2,159,374,091-byte scored file 2 of 2 streams with the same bytes and the same sha256 46a01478dceefe81…; 26 of 26 output files byte-identical; 50 of 50 seals recompute (48 + 2); 14 of 14 local runs exit 0 f2d890e48
the comparison, run A and run B on one Mac 13 of 13 output files byte-identical; 26 of 26 seals recompute; 0 output files hold a decimal 3c25e8cf9
every HTTP stream of their scored file in the comparison rounds 7 of 8 complete (one was stopped), all 7 with the same bytes and the same sha256 3c25e8cf9
every committed file of both Study 56 evidence folders against its own SHA256SUMS, checked again on 2026-10-09 41 of 41 and 217 of 217 match 3c25e8cf9
the Affine IDE agent swarm, the Study 56 goal, with 1, 2 and 5 agents one goal seal, c6949876a8639991, for every agent count; census and index byte-equal to the committed record 8eb8f4ca9
history: the same goal on an earlier reading that turned decimals into fractions, with 1, 2, 4 and 5 agents one goal seal, 38c8079f18ed23d3, for every agent count af49986a8
the nine Affine.Earth cells not known: they do not yet answer Study 56 β€”

5 Β· What this measures, and what it does not

  • It solves no problem. The substrate reads the key MATH's authors wrote and the answers OpenAI's model wrote, and decides whether each reads to an exact integer or fraction and whether the two are equal. It does not check that a key is right: an answer equal to a key holds the key's value, whatever the key is.
  • The substrate has read MATH, and only MATH. It has not read AIME, HMMT, USAMO, the IMO, CMO, Putnam, FrontierMath, MathArena's sets or HLE. Every figure in Β§2 and Β§3 about those tests is theirs, and the substrate makes no figure of its own beside them.
  • It has read none of the models named in Β§2 and Β§3. The 815,632 attempts it read were written by the large-scale generator of Let's Verify Step by Step, which the paper finetunes from the base model it names GPT-4, and published in 2023; the figures in Β§2 and Β§3 were published from 2024 to 2026.
  • It decides answers; it does not write proofs, and it does not grade them. The olympiad and Putnam rows are proofs graded by people or by models; nothing on this page grades a proof.
  • It does not write the attempts. The sampling settings that produced them are not stated in their paper beyond the 1,860 per problem and the 1,024-token cut-off. The substrate counts what was written and decides each answer against the key.
  • Where it decides nothing, it says why. 1,967 MATH keys cannot close, 83 of them among the 500; 209,850 attempts cannot close, each with its reason, and their mark reads true on 75,708 of them; 20 lie past the room. The law does not take roots, evaluate functions, or read pi, i, infinity, letters, words or decimals as values.
  • Not known: how OpenAI produced is_correct; whether their majority-voting figure is a whole count or a rounded mean; how their majority vote matches two answers or breaks a tie. Known: their eval.py prints its reward-model selections as a mean over 400 trials.
  • The run-to-run studies in Β§1 were measured on other models and other sets β€” one of them on a prompt that is not mathematics. None was measured on OpenAI's 815,632 attempts. The count of different values per problem in Β§4 is the substrate's own, over those attempts, which were drawn as 1,860 samples per problem.
  • On which machines. The seals came out byte-identical every time the study ran: twice on one Mac (Β§4), and with 1, 2 or 5 agents. No other machine has run it yet. The nine Affine.Earth cells do not yet answer Study 56: no court on them reads it, so whether they write the same bytes is not known. The precedent is Study 55, which the cells answer with one digest on 9 of 9 (043de8021).
  • What is tabled: 29 of the 31 models in MathArena's AIME 2025 files are tabled. The tables of the labs and the evaluators are what was read on 2026-10-09; they are not every figure every lab has printed.

6 Β· Sources

The sources of the labs, the evaluators and the run-to-run studies were read on 2026-10-09; the three sources of the attempts the substrate read were read on 2026-10-08, and their entries say so. Where no date was found on a page, the entry says so.

Labs (theirs)

The attempts the substrate read (theirs)

  • OpenAI β€” Lightman, Kosaraju, Burda and others, Let's Verify Step by Step, arXiv 2305.20050, v1 2023-05-31, the only version; the abstract page and the PDF read 2026-10-08
  • OpenAI β€” prm800k repository, commit 7ecc7947, 2023-06-01, MIT licence; read at that commit 2026-10-08
  • OpenAI β€” the scored file, scored-test-samples.jsonl, 2,159,374,091 bytes, last modified 2023-05-25; its headers read 2026-10-08

Independent evaluators and studies of run-to-run variation (theirs)

The substrate (Affine.Earth, committed)

  • evidence/study-56-v3-blind-seal-20261008/ β€” the blind seal, commit f2d890e48
  • evidence/study-56-v3-compare-20261008/ β€” the comparison, commit 3c25e8cf9
  • Study 56 β€” MATH, read exactly

🧬 CURES β€” read in this order

Each step is the reason the next one exists. Nothing here is medical advice, and no page calls any medicine safe or unsafe.

1 Β· Why an exact safety screen at all

2 Β· The three libraries, which grow rather than close

3 Β· The maps β€” every place a molecule could act, counted

4 Β· One medicine at a time

  • Zilganersen β€” the first treatment for Alexander disease, screened on the real approved sequence
  • A drug an AI designed β€” rentosertib for pulmonary fibrosis, and exactly what our instruments reach
  • CAR-T, halted β€” the verdict a regulator could re-derive
  • N-of-1 antisense β€” the only safety net at a population of one
  • VERVE-102 β€” the off-target lattice a stranger can re-derive
  • PM359 β€” prime editing, certified before anyone is dosed
  • Del-Zota β€” the one safety question that can be made exact

5 Β· What keeps a disease alive, and what moves it

βš–οΈ How to read any page here

πŸ”¬ The method β€” exact against float, domain by domain

The same move every time: take a domain where a floating-point model is the accepted instrument, compute the same quantity in exact integers, and seal the cases where the two render opposite verdicts. The subject under grading is always the instrument, never the phenomenon.

⚑ Fusion β€” the energy case

🌍 The planet, and the sky

πŸ› Markets, money and risk

βš›οΈ Run a court yourself

πŸ“’ Program ledger β€” every study by lifecycle

A study appears here under the state its evidence has earned, and above under the question it answers. The two are different filings of the same work, on purpose.

βœ… LAW FROZEN Β· DATA SEALED

πŸ”΄ LIVE CLAIM β€” standing, not sealed

🌊 CHARTER Β· OPEN β€” the findings, published either way

β˜€οΈπŸŒ‘ Eclipse 2026 β€” Study 01, DATA SEALED

πŸ”¬ Discoveries and flows

Clone this wiki locally