Repository navigation
Frontier Models In Mathematics Sampled Or Sealed
Frontier models in mathematics β a share over sampled runs, and one sealed reading of every key and every answer
Frontier labs print mathematics figures built from sampled runs β one sample, or 4, 8, 64 or 1,000 β a checker the lab chose marks the answers, and an average, a vote or another model's score makes the printed figure (Β§2). Run it again and the answers can change. On the 30 AIME 2025 problems, 465 of 870 modelβproblem pairs returned more than one answer across four runs (counted from MathArena's published rows, theirs; Β§1). One prompt sent 1,000 times at temperature 0 came back as 80 different completions in 1,000 runs (Thinking Machines Lab, theirs). They train their models and sample them.
Affine.Earth trains nothing and samples nothing. One law of exact whole numbers reads the bytes: 12 Swift files, 0 imports, no float. Study 56 puts that law to MATH, a public benchmark of 12,500 competition problems, and to OpenAI's prm800k release, 815,632 scored attempts at 500 of them. It answers what a sample cannot: which answers equal their key exactly. Whether the nine Affine.Earth cells read them the same way is not known: no cell runs Study 56 yet (Β§5). This page sets the two side by side, every figure of theirs labelled theirs.
- One verdict for every answer the law can hold. All 12,500 keys and 815,612 of OpenAI's 815,632 attempts have one verdict each; the other 20 attempts are longer than the law's room of 10,000 digits, and they are named and not read. The law was designed and sealed without the marks, and committed before the marks were read again (an earlier reading had read them once, on 2026-10-07). Run again after the seal, all 815,632 attempts reproduced their sealed outcome. Run twice on one Mac, the law wrote byte-identical files (Β§4). The Affine IDE's agent swarm seals one set of bytes with 1, 2 or 5 agents.
- 1,967 of MATH's 12,500 answer keys are not a whole number or a fraction of whole numbers, the only values the law holds, and the law names why for each. 595 hold a root Β· 383 hold letters Β· 265 hold pi Β· 256 are written with a decimal point Β· 165 are words Β· 99 hold the imaginary unit Β· 91 hold infinity Β· 71 hold a container of containers Β· 19 use a form outside the law's grammar Β· 11 are unions of sets Β· 6 hold a trigonometric function Β· 4 hold no answer Β· 1 holds a logarithm Β· 1 has a comma between digit groups. Of OpenAI's 500 problems, 83: 24 roots, 20 letters, 10 pi, 9 decimals, 7 words, 6 the imaginary unit, 4 infinity, 2 unions of sets, 1 container of containers.
- The law decides none of the 134,995 attempts on those 83 problems, and their mark reads correct on 73,200 of them (theirs). 47,027 of the 134,995 copy the key byte for byte, and a copy of a text the law cannot read is still not read to a value. How many of the 73,200 are such copies is not known.
- Every time the law reads an answer unequal to the key, their mark also reads false: 243,578 of 243,578, and true on none.
-
4,825 answers equal to the key are marked false (theirs). Where the law reads an answer equal to the key, their mark
is_correctreads false on 4,825 attempts in 49 problems (theirs). 1,793 attempts answer 864 and are marked false; their copy of the key reads864\mbox{inches}^2, and MATH's key is 864.\frac{2187}{5625}is marked false; it reads to 243/625, and the key is 243/625. -
Their grader normalises answers with sympy, with no sympy version pinned and no timeout (their
grader.py, theirs). No file of theirs states howis_correctwas produced. -
Side by side, never subtracted. Among OpenAI's own attempts at each problem, the answer most of them hold, counted by exact value, equals the key on 293 of the 417 problems whose key the law reads to a value. The other 83 have no key value to equal. Theirs (Let's Verify Step by Step, Figure 3, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it counts problems; their
eval.pyprints their two reward-model picks as means over 400 trials, not as counts. Their printed figures stand beside the law's as published; the two cannot be subtracted (Β§4).
Theirs is a sample. This reading came out byte-identical every time it ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.
Status: PUBLISHED FIGURES SET BESIDE STUDY 56 β 2026-10-09 β what the labs that build frontier models print about their mathematics, what independent evaluators counted run by run, and what Affine.Earth's substrate computes from sha-pinned bytes, side by side. Every figure of theirs is labelled theirs, with its source, its date and the conditions it was printed under; where no date was found on a source, its row says so. Every figure of the substrate's names its commit in Β§4: the blind seal, made before OpenAI's marks were read again, and the comparison, made after; a count added up on 2026-10-09 from a committed table says so beside it. Integers only: a published share is written as a whole count k of N where the source states N and k is whole; otherwise the page says what kind of number it is and does not print it. The sources of the labs, the evaluators and the run-to-run studies were read on 2026-10-09; the paper, the repository and the scored file's headers, the sources of the attempts the substrate read, were read on 2026-10-08. Study 56 β MATH, read exactly Β· Program index Β· Study board Β· Study 55 β IceCube: the light in the ice
When a lab says its model scores so much on a mathematics test, each row below says how that figure was built, as far as its source says: how many runs, which checker marked them (a script, another model, or people), and under which conditions the lab set β which tools, how much thinking time, how many samples. Run again, the answers can change, and the evaluators in Β§1 counted how often they do.
Affine.Earth does a different thing, and it solves no problem. Study 56 reads the answer key that MATH's authors wrote for each of the 12,500 problems in MATH, a public set of competition problems, and decides whether that key closes to an exact integer or fraction, or cannot close, with the reason named β a root, a letter, pi, a decimal point, a word. It does not check that a key is right. It then reads the 815,632 answers OpenAI published in 2023 with its paper Let's Verify Step by Step, for 500 of those problems. They were written by the paper's large-scale generator, a model it finetunes from the base model it names GPT-4, as 1,860 solutions per problem, drawn so that a reward model or a vote could pick one. The law decides each answer against the key: equal, unequal, or cannot close, with its reason. Anyone can run the law on the same committed bytes. The seals came out byte-identical every time it ran, twice on one Mac and with 1, 2 or 5 agents; whether the nine Affine.Earth cells write the same bytes is not known until they answer Study 56. The substrate has read none of the models named in Β§2 and Β§3. The two sides are set next to each other here, with nothing added.
| who (theirs) | what was run | what they counted, as integers | source |
|---|---|---|---|
| Thinking Machines Lab | one prompt, sent 1,000 times at temperature 0 to Qwen/Qwen3-235B-A22B-Instruct-2507 in non-thinking mode, 1,000 tokens per completion; the prompt is about Richard Feynman, not a mathematics question | 80 different completions in 1,000 runs; the most common appeared 78 times; all 1,000 agreed for the first 102 tokens, and at token 103, 992 went one way and 8 another. With their batch-invariant kernels: 1 completion in 1,000 runs. The cause they name: kernels whose order of addition changes with batch size | Defeating Nondeterminism in LLM Inference, 2025-09-10 |
| Atil and others | the MMLU college-mathematics questions, 10 runs per model, temperature 0, top-p 1, a fixed seed, the same infrastructure, 5-shot | GPT-4o: 88 of 100 in its best run, 44 of 100 in its worst; 50 questions got one parsed answer in all 10 runs; 0 questions got one raw output in all 10 runs. Llama-3-70B: 85 best, 22 worst, 0 with one raw output. Mixtral 8x7B: 75 best, 3 worst, 0 with one raw output. The paper prints shares; 100 is the row count of the public test split (cais/mmlu), read 2026-10-09 |
Non-Determinism of "Deterministic" LLM Settings, v1 2024-08-06; figures from v5, 2025-04-02 |
| Atil and others | 5 models, 8 tasks, zero-shot and few-shot: 80 conditions, 10 runs each | their reading: no model gave the same accuracy on every task, and identical output strings were rarer still; the spreads are printed as shares over task sets of different sizes, not counts of problems | same paper |
| Yuan and others | DeepSeek-R1-Distill-Qwen-7B, greedy decoding, the same seed, the same prompts, run under 12 configurations (2 GPU types Γ 2 GPU counts Γ 3 batch sizes) | on the 30 AIME'24 problems the accuracy moved between configurations (printed as a share, not a count); response length differed by up to 9,000 tokens. The cause they name: floating-point addition is not associative at limited precision | Give Me FP32 or Give Me Death?, v1 2025-06-11 |
| Hochlehnert and others | the same weights and prompts on 5 compute clusters, and under 2 evaluation frameworks | AIME'24 scores moved with the cluster and with the framework; printed as shares (means over seeds), which their text says can change model rankings; not counts | A Sober Look at Progress in Language Model Reasoning, v1 2025-04-09 |
| MathArena | how it scores every model | each model answers each question 4 times; the score is the average of the 4, with no majority vote; final answers are parsed into sympy and also checked by an LLM judge, and every disagreement between the two is looked at by hand | MathArena: Evaluating LLMs on Uncontaminated Math Competitions, v1 2025-05-29 |
MathArena publishes every run's parsed answer and its grade. Counting those rows for AIME 2025 I and II (30 problems), for 29 of the 31 models in the files, 4 runs each:
| count | |
|---|---|
| modelβproblem pairs | 870 |
| runs | 3,480 |
| runs graded correct (MathArena's grading) | 1,979 |
| pairs graded correct in all 4 runs | 375 |
| pairs graded correct in none of the 4 | 269 |
| pairs graded correct in 1, 2 or 3 of the 4 | 226 |
| pairs where the 4 runs did not all return one answer | 465 of 870 |
Some of the answers, as returned:
| model (theirs) | runs graded correct | problems with more than one answer across 4 runs | one problem, its 4 answers |
|---|---|---|---|
| DeepSeek-R1 | 84 of 120 | 14 of 30 | AIME 2025 I problem 11, key 259: 259, 4423, 758, 0 Β· I problem 13, key 204: 829, 829, 622, 754 Β· II problem 7, key 237: 237, 237, 60671, 60671 |
| o3 (high) | 107 of 120 | 6 of 30 | AIME 2025 I problem 15, key 735: 147, 423, 441, 343 |
| o4-mini (high) | 110 of 120 | 5 of 30 | β |
| o1 (medium) | 96 of 120 | 11 of 30 | β |
Datasets MathArena/aime_2025_I_outputs (revision 05b2f50ac7cb76ff) and MathArena/aime_2025_II_outputs (revision 9809200ca4e719df), 1,860 rows each; created 2025-04-16, last modified 2026-05-15, read 2026-10-09; each model at the settings MathArena lists for it. A run with no parsed answer counts as one answer value. The counts are this page's, made from MathArena's rows; MathArena does not print them.
The substrate's answer to "run it again" is in Β§4, under The same bytes, the same seals: its law, run twice as separate processes on one Mac, wrote byte-identical files. Two of the studies above changed the hardware; on more than one machine the substrate's result is not known until the nine Affine.Earth cells answer Study 56.
Every row in this section is the lab's own figure about its own model, from its own page. Where the source prints a share and states no problem count, the share is not turned into a count and is not printed; the cell says what kind of number it is. Grouped by test.
- pass@1 (DeepSeek writes Pass@1): the share of problems marked correct when one answer per problem is marked; where several samples are drawn per problem, it is that share averaged over them.
- consensus of 64, cons@64, majority vote over 64 samples: the answer given most often across 64 samples is the one marked. Consensus of 8 is the same over 8.
- re-ranking of 1,000 samples, best-of-N: a scoring model picks one of N samples, and only that one is marked.
- em_maj1@1 (Meta): exact match of the answer, majority vote over 1 sample β that is, one sample, marked by exact match.
- 4 shots, 5-shot: the prompt carries 4 or 5 worked examples before the question.
- temperature: how widely the model samples among its possible next tokens; at temperature 0 it is asked to take the most likely token every time, which is also what greedy decoding means.
- top-p: the model samples only among its most likely next tokens that together hold a share p of the probability; top-p 1 keeps all of them.
- bf16: bfloat16, a 16-bit floating-point format in which the model's weights are held.
- 128K context, thinking capped at 96k tokens: how many tokens (pieces of words) the model may read, or may write before it answers.
- reasoning effort, thinking, Think Max, heavy mode: settings on the lab's own service for how long the model may work, or how many attempts it runs side by side, before it answers.
- Pass@8192, 1,024 samples per problem (formal proofs): a problem counts as solved if any one of that many attempts is a proof the checker accepts.
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| DeepSeek | DeepSeek-R1, MATH-500 | printed as a mean over k samples per question (k typically 4 to 64); not a count of problems | up to 32,768 generated tokens; temperature 6/10, top-p 95/100 | 2025-01-22 | arXiv paper |
| Meta | Llama 4 Maverick, Llama 4 Scout (pre-trained); Llama 3.1 405B and 70B beside them | printed as shares; problem count not stated; not converted | base models; 4 shots; metric em_maj1@1; bf16 weights | 2025-04-05 | model card |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| OpenAI | o1, o1-preview, GPT-4o | printed as an average score out of 15 per exam, for one sample, a consensus of 64 samples and a re-ranking of 1,000 samples; not a whole count of problems | o1 at its maximal test-time compute; re-ranking uses a learned scoring function; tools not mentioned; grader not stated | 2024-09-12 | Learning to reason with LLMs |
| DeepSeek | DeepSeek-R1 | printed as a mean over k samples per question; not a count of problems | as for MATH-500 above | 2025-01-22 | arXiv paper |
| DeepSeek | DeepSeek-R1-Zero | pass@1 printed as a mean over samples (16 responses per question in the training curve); cons@64 printed as a share; not converted | cons@64 is a majority vote over 64 samples | 2025-01-22 | arXiv paper |
| xAI | Grok 3 mini | printed as a share; not converted | conditions not stated in the sentence | 2025-02-19 | Grok 3 |
| Mistral AI | Magistral Medium, Magistral Small | each printed as two shares, the second with a majority vote over 64 samples; not converted | conditions of the first figure not stated in the sentence | 2025-06-10 | Magistral |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| OpenAI | o4-mini, o3 | consensus of 8: every problem, for both; pass@1 printed as shares | a Python interpreter; high reasoning effort; OpenAI writes that tool results should not be compared with no-tool results | 2025-04-16 | Introducing o3 and o4-mini |
| OpenAI | GPT-5 pro, GPT-5, o3 | GPT-5 pro with Python: every problem at pass@1; every other bar printed as a pass@1 share; not converted | pass@1; high effort; with and without Python; with and without thinking | 2025-08-07 | Introducing GPT-5 |
| OpenAI | GPT-5.2 Thinking, GPT-5.2 Pro, GPT-5.1 Thinking | every problem for GPT-5.2 Thinking and GPT-5.2 Pro; GPT-5.1 Thinking printed as a share | no tools; maximum reasoning effort in the API; research environment | 2025-12-11 | Introducing GPT-5.2 |
| Google DeepMind | Gemini 3 Pro, Gemini 2.5 Pro | Gemini 3 Pro with code execution: every problem; the no-tool figures are pass@1 shares averaged over an unstated number of trials; not converted | pass@1, a single attempt, no majority vote; default sampling; results as of November 2025 | 2025-11 | Gemini 3 Pro evaluation |
| xAI | Grok 3 (Think) | printed as a share; not converted | majority vote over 64 samples; the exam was released 7 days before the post; the model still in training | 2025-02-19 | Grok 3 |
| xAI | Grok 4, Grok 4 Heavy | Grok 4 Heavy: every problem; the other figures printed as shares | the tool is Python; Heavy uses parallel test-time compute; sample counts not stated | 2025-07-09 | Grok 4 |
| xAI | Grok 4 Fast, Grok 4 | printed as shares; not converted | pass@1; no tools | 2025-09-19 | Grok 4 Fast |
| DeepSeek | DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking | printed as Pass@1 shares; sample count not stated; not converted | temperature 1; 128K context | 2025-12-01 | paper |
| Moonshot AI | Kimi K2 Thinking | no tools: printed as a mean over 32 runs; with Python: a mean over 16 runs; not counts of problems. Heavy mode: every problem | temperature 1; thinking capped at 96k tokens; heavy mode rolls out 8 trajectories in parallel and aggregates them; heavy run count not stated | 2025-11-04 | model card |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| Alibaba (Qwen team) | Qwen3.5-397B-A17B, Qwen3-Max-Thinking | printed as shares that are not whole counts of a 30-problem set; run count not stated; not converted | sampling for the table not stated | 2026-02-16 | model card |
| Moonshot AI | Kimi K2.6, Kimi K2.5 | printed as shares; run count not stated; not converted | thinking mode; temperature 1, top-p 1; generation capped at 98,304 tokens | 2026-04-14 | model card |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| OpenAI | GPT-5.2 Pro, GPT-5.2 Thinking, GPT-5.1 Thinking β February 2025 | GPT-5.2 Pro: every problem; the other two printed as shares | no tools; maximum effort | 2025-12-11 | Introducing GPT-5.2 |
| xAI | Grok 4, Grok 4 Heavy β 2025 | printed as shares; not converted | with and without Python; Heavy in parallel | 2025-07-09 | Grok 4 |
| xAI | Grok 4 Fast, Grok 4 β 2025 | printed as shares; not converted | pass@1; no tools | 2025-09-19 | Grok 4 Fast |
| DeepSeek | DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking β February and November 2025 | printed as Pass@1 shares; not converted | temperature 1; 128K context | 2025-12-01 | paper |
| DeepSeek | DeepSeek-V4-Pro-Max β February 2026 | printed as a Pass@1 share; not converted | Think Max mode, with a special system prompt | 2026-04-22 | model card |
| Alibaba (Qwen team) | Qwen3.5-397B-A17B, Qwen3-Max-Thinking β February and November 2025 | printed as shares; not converted | sampling not stated | 2026-02-16 | model card |
| Moonshot AI | Kimi K2 Thinking β 2025 | printed as means over 32 runs (no tools) and 16 runs (Python); heavy mode printed as a mean, not a count | as above | 2025-11-04 | model card |
| Moonshot AI | Kimi K2.6, Kimi K2.5 β February 2026 | printed as shares; not converted | thinking mode | 2026-04-14 | model card |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| Google DeepMind | AlphaProof with AlphaGeometry 2 β IMO 2024 | 28 of 42 points; 4 of 6 problems (two combinatorics problems unsolved) | problems translated by hand into formal language first; one solved in minutes, others took up to three days; scored by Timothy Gowers and Joseph Myers; that year's gold began at 29 points, reached by 58 of 609 contestants | 2024-07-25 | AI solves IMO problems at silver-medal level |
| OpenAI | an unnamed experimental reasoning LLM β IMO 2025 | 35 of 42 points; 5 of 6 problems; no solution for problem 6 | two 270-minute sessions; no tools or internet; each proof graded by three former IMO medalists, by unanimous consensus; grading by IMO coordinators not claimed | 2025-07-19 | thread by Alexander Wei |
| Google DeepMind | Gemini Deep Think, advanced version β IMO 2025 | 35 of 42 points; 5 of 6 problems | graded and certified by IMO coordinators with the student criteria; natural language from the official statements; within the 270-minute limit | 2025-07-21 | Gemini Deep Think at the IMO |
| Harmonic | Aristotle β IMO 2025, formal | 5 of 6 problems, as formal solutions | Lean 4 with Mathlib; geometry problems handled by a solver working outside Lean | 2025-10-01 | arXiv paper |
| DeepSeek | DeepSeekMath-V2 β IMO 2025, CMO 2024, Putnam 2024 | IMO 2025: 5 of 6 problems fully solved (point share not converted). CMO 2024: 4 of 6 fully solved, partial credit on 1. Putnam 2024: 118 of 120 points, 11 of 12 problems fully solved | 64 proof samples, each checked by 64 verification analyses, up to 16 refinement rounds; DeepSeek's own experts assessed the top proofs; the paper gives the highest human Putnam 2024 score as 90 | 2025-11-27 | repository |
| DeepSeek | DeepSeek-V3.2-Speciale β IMO 2025, CMO 2025 | IMO 2025: 35 of 42 points. CMO 2025: 102 of 126 points | up to 128k generated tokens; no tools or internet; a generateβverifyβrefine loop; the grader is not named | 2025-12-01 | paper |
| dots team (Xiaohongshu, RedNote) | dots-note-3.0, internal version β IMO 2026 | 42 of 42 points; 6 of 6 problems | invited by the IMO Organizing Committee and marked through its official marking; reads the original LaTeX; uses Python; time limit not stated; 7 of 666 human contestants also scored 42; the affiliation is from press reports | 2026-07 | dots at IMO 2026 |
| Huawei | Celia (Xiaoyi Agent system) β IMO 2026 | 42 of 42 points | as reported by IT Home from Huawei's announcement; per-problem points and grading channel not given | 2026-07-22 | IT Home |
| xAI | Grok 4, Grok 4 Heavy β USAMO 2025 | printed as shares of the score; point total and run count not stated; not converted | Heavy with Python and parallel test-time compute; who graded the proofs is not stated | 2025-07-09 | Grok 4 |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| Google DeepMind | Gemini 3 Pro, Gemini 2.5 Pro β MathArena Apex | printed as shares; not converted | Google states these were reported by matharena.ai | 2025-11 | Gemini 3 Pro evaluation |
| Google DeepMind | Gemini Deep Think, January 2026 version β IMO-ProofBench Advanced | printed as a share, rising as inference-time compute scales; not converted | graded by human experts | 2026-02-11 | Accelerating discovery with Gemini Deep Think |
| DeepSeek | DeepSeek-V3.2-Speciale, DeepSeek-V3.2-Thinking β IMOAnswerBench | printed as Pass@1 shares; not converted | temperature 1 | 2025-12-01 | paper |
| DeepSeek | DeepSeek-V4-Pro-Max β IMOAnswerBench, Apex, Apex Shortlist | printed as Pass@1 shares; not converted | Think Max mode | 2026-04-22 | model card |
| Alibaba (Qwen team) | Qwen3.5-397B-A17B, Qwen3-Max-Thinking β IMOAnswerBench | printed as shares; not converted | sampling not stated | 2026-02-16 | model card |
| Moonshot AI | Kimi K2 Thinking β IMO-AnswerBench | printed as a mean over 8 runs; not a count of problems | thinking capped at 128k tokens | 2025-11-04 | model card |
| Moonshot AI | Kimi K2.6, Kimi K2.5 β IMO-AnswerBench | printed as shares; not converted | thinking mode | 2026-04-14 | model card |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| OpenAI | o3-mini (high) | printed as lower bounds on a share; problem count not stated; not converted; OpenAI calls the numbers provisional | Python tool; first attempt | 2025-01-31 | OpenAI o3-mini |
| OpenAI | GPT-5.2 Thinking, GPT-5.1 Thinking β Tiers 1β3 and Tier 4 | printed as shares; problem count not stated; not converted | with Python; maximum effort; research environment | 2025-12-11 | Introducing GPT-5.2 |
| OpenAI | GPT-5.4, GPT-5.4 Pro, GPT-5.2, GPT-5.2 Pro | printed as shares; not converted. The GPT-5.2 figures on this page differ from those on the GPT-5.2 page, and GPT-5.2 Pro carries a Tier 4 figure here that the GPT-5.2 page does not print | xhigh effort unless stated; tool setting not stated | 2026-03-05 | Introducing GPT-5.4 |
| OpenAI | GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5 β v2 | printed as shares; not converted | effort and tools not stated | 2026-07-09 | GPT-5.6 |
| OpenAI | GPT-6 Astra, GPT-5.6 Sol β Tier 4 v2 | printed as shares; not converted | the maximum at any effort; the page prints no date (third-party reports give 2026-09-03); as first read on 2026-10-09 β a second read that day returned HTTP 403, so this row was not re-verified | 2026-09-03 | GPT-6 Astra |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| DeepSeek | DeepSeek-Prover-V2-671B | MiniF2F-test: 217 of 244 problems at Pass@8192 (the paper's Table 2). PutnamBench: the repository prints 49 of 658 and states no sample budget. The paper, in its second version, ran the model on the 649 of the 658 problems that work with Lean 4.9.0; its first run solved 49 at 1,024 samples per problem; after the proofs went to the benchmark's maintainers, 2 problems were excluded for misformulated statements, and its Table 4 prints 47 of 658 | Lean 4 proofs checked by the Lean proof assistant | repository 2025-04-30; paper v2 2025-07-18 | repository Β· paper |
| Google DeepMind | AlphaGeometry 2 β IMO geometry problems of the past 25 years | printed as a share; problem count not stated; not converted | formal geometry system; measured before IMO 2024 | 2024-07-25 | AI solves IMO problems at silver-medal level |
| lab | model | their figure, as integers (theirs) | conditions (theirs) | date | source |
|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra β bounded gaps between primes | helped establish a bound of 186 for infinitely many prime pairs; earlier bounds 246, then 240 (Julia Stadlmann) | OpenAI states it shares the proofs, an abridged chain of thought and verification materials; who verified the proof is not known from this page; as first read on 2026-10-09 β a second read that day returned HTTP 403, so this row was not re-verified | 2026-09-03 | GPT-6 Astra |
| Google DeepMind | Aletheia, built on Gemini Deep Think β ErdΕs problems | 4 autonomous solutions among 700 open problems evaluated | uses a natural-language verifier, Google Search and web browsing | 2026-02-11 | Accelerating discovery with Gemini Deep Think |
These are figures counted by someone other than the model's maker. Where a source prints one figure that is a mean over runs, the per-run counts it also publishes are given instead. No date was found on MathArena's competition tables as read, so their rows give the read date and, where the source names it, the competition's month.
| evaluator | how | what it means in counts | source |
|---|---|---|---|
| MathArena (SRI Lab, ETH Zurich, with INSAIT) | every model 4 times on every problem; the score is the average of the 4; a mark beside a model means it was released after the competition | 4 runs per problem per model | matharena.ai, no date found on the page; read 2026-10-09 |
| Vals AI | AIME 2024 and 2025, 60 questions, 8 runs each; the boxed answer compared with the reference | 480 answers per model: Gemini 3.1 Pro Preview (02/26) 471 of 480; GPT 5.2 465 of 480. The page no longer runs new models | vals.ai, updated 2026-04-16 |
| Artificial Analysis | AIME 2025 as pass@1 with 10 repeats; MATH-500, a 500-problem subset of MATH; both now retired | AIME 2025: 30 questions Γ 10 repeats = 300 answers per model; graded by script with SymPy and an LLM equality checker as backup | methodology, no date found on the page; read 2026-10-09 |
| who | model | per-run or whole counts | conditions | source |
|---|---|---|---|---|
| MathArena | GPT-5 (high) | runs 1 to 4: 19, 15, 14, 16 of 42 β 64 of 168. By problem, runs 1 to 4: P1 7 1 1 0 Β· P2 0 0 0 0 Β· P3 2 2 1 2 Β· P4 4 5 5 7 Β· P5 6 7 7 7 Β· P6 0 0 0 0 | marked as released after the competition | IMO 2025 table, no date found on the table; competition July 2025; read 2026-10-09 |
| MathArena | Gemini 2.5 Pro | runs 1 to 4: 17, 10, 11, 15 of 42 β 53 of 168; both judges gave the same grade on every run. By problem: P1 1 2 0 1 Β· P2 0 0 0 0 Β· P3 7 2 4 7 Β· P4 5 2 3 3 Β· P5 4 4 4 4 Β· P6 0 0 0 0 | best of 32, chosen in a bracket tournament where the model judged its own answers; each answer graded by 2 of 4 human judges | IMO 2025 table, as above; method: matharena.ai/imo, 2025-07-17 |
| MathArena (blog) | the July 2025 post | bronze needed 19 of 42; the top model's figure is printed as a mean over 4 runs, not a count β its per-run counts are 17, 10, 11 and 15 of 42 | 4 judges with IMO-level background, grading in pairs | matharena.ai/imo, 2025-07-17 |
| UK Mathematics Trust | the IMO 2025 boundaries the students were cut at | gold from 35 of 42, silver from 28, bronze from 19; 110 countries | Sunshine Coast, Australia | UKMT news, 2025-07-18 |
| Google DeepMind | Gemini Deep Think, advanced version | 35 of 42 points; 5 of 6 problems | graded and certified by IMO coordinators; the IMO President is quoted confirming the 35 | DeepMind blog, 2025-07-21 |
| OpenAI, as quoted by Simon Willison | the experimental reasoning LLM | 35 of 42 points; 5 of 6 problems; 3 graders per proof | three former IMO medalists by unanimous agreement; whether IMO coordinators graded it is not stated | simonwillison.net, 2025-07-19 |
| OpenAI, on GitHub | the same model's proofs | 5 proof files; 0 for problem 6 | first commit 2025-07-18 | aw31/openai-imo-2025-proofs |
| Google DeepMind | AlphaProof with AlphaGeometry 2 β IMO 2024 | 28 of 42; 4 of 6; AlphaGeometry 2 proved Problem 4 within 19 seconds of receiving it formalized; gold began at 29 points, 58 of 609 contestants | problems formalized by hand first; scored by Gowers and Myers | DeepMind blog, 2024-07-25 |
| who | model | points per run, of 42 | conditions | source |
|---|---|---|---|---|
| Petrov, Dekoninck and others ("Proof or Bluff?"); the grades as MathArena stores them | DeepSeek-R1 | 0, 7, 0, 1 β 8 of 168, the same from both judges | graded within hours of the problems' release; four expert judges, two per problem | per-run grades: USAMO 2025 table, its stored trace grades, no date found on the table, read 2026-10-09; the study: arXiv v1, 2025-03-27 |
| MathArena | Gemini 2.5 Pro | judge 1: 14, 7, 6, 14; judge 2: 14, 7, 8, 12 β each judge 41 of 168 | released after the competition | USAMO 2025 table, no date found on the table; read 2026-10-09 |
| MathArena | DeepSeek-R1-0528 | judge 1: 14, 7, 15, 15 (51 of 168); judge 2: 14, 7, 14, 15 (50 of 168); one split, problem 6 run 3, 1 and 0 | released after the competition | same |
| MathArena | o4-mini (high) | 7, 14, 4, 7 β 32 of 168, both judges | released after the competition | same |
| MathArena | GPT-5.4 (xhigh) β 2026 | 42, 42, 41, 35 β 160 of 168 | graded by an LLM jury of three models that includes GPT-5.4, the lowest of the three scores kept, then every solution reviewed by hand | USAMO 2026 table, no date found on the table; competition March 2026; read 2026-10-09; jury: Proof, Not Bluff, 2026-03-28 |
| MathArena | GPT-5.5 (xhigh) β 2026 | 41, 41, 42, 41 β 165 of 168 | released after the competition | same |
| MathArena | Gemini 3.1 Pro Preview β 2026 | 33, 35, 32, 25 β 125 of 168 | Gemini-3.1-Pro also sits on the grading jury | same |
| MathArena | DeepSeek-v4-Pro (Max) β 2026 | 28, 25, 30, 19 β 102 of 168 | released after the competition; open weights | same |
| system | points of 120 | conditions | source |
|---|---|---|---|
| DeepSeek-v3.2-Speciale Agent | 103 (by problem: 10 10 10 9 2 9 10 10 10 10 10 3); within the top 3 of 4,329 human participants | one submission, chosen before the contest; the graders knew the answers were AI-written, not which model; up to 1,000 model queries per problem (cut from 65,000) | matharena.ai/putnam, 2026-02-26 |
| Gemini Deep Think 3 (12/25) | 100 | one submission each, graded in the Putnam graders' second round | Putnam 2025 table, no date found on the table, read 2026-10-09; with the post of 2026-02-26 |
| GPT-5-Pro | 93 | same | same |
| Gemini 3 Pro (preview) | 91 | same | same |
| DeepSeek-v3.2-Speciale | 71 | same | same |
| GPT-5.1 (high) | 58 | same | same |
The evaluation was partially supported by an evaluation grant from Google DeepMind, which also makes two of the six systems graded (matharena.ai/putnam, 2026-02-26).
| who | model | correct answers per run | conditions | source |
|---|---|---|---|---|
| MathArena | GPT-5.2 (high) β AIME 2026, 30 problems | 30, 29, 30, 29 of 30 β 118 of 120 | answers checked against the official integers | AIME 2026 table, no date found on the table; competition February 2026; read 2026-10-09 |
| MathArena | GPT-5.2 (high) β HMMT February 2026, 33 problems | 31, 32, 33, 32 of 33 β 128 of 132 | not marked as released after the competition | HMMT Feb 2026 table, no date found on the table; competition February 2026; read 2026-10-09 |
| MathArena (Sun, Dekoninck, Vechev) | Gemini 3.1 Pro, 4 attempts on each of 176 problems from contests since August 2025 | 162 solved in all 4 attempts; 14 solved at least once; 0 unsolved in every attempt β so no problem met the Apex rule. MathArena concludes recent public contest problems no longer serve as hard frontier benchmarks | the Apex rule keeps a problem only if it is unsolved in every attempt | matharena.ai/no_final_answer, 2026-05-12 |
| model | points per run | how graded | source |
|---|---|---|---|
| GPT-5.4 (xhigh) | 28, 21, 23 of 62 and 22 of 60 (problem 21 has 3 runs) β 94 points over 123 graded answers | 0 points = proves the false statement as given; 1 = silently repairs it; 2 = any other response, such as saying it is false; judged by Gemini-3.1-Pro, an LLM; the prompt asks the model to try to prove the statement | BrokenArXiv, February, no date found on the table, read 2026-10-09; method, 2026-03-13 |
| Gemini 3.1 Pro Preview | 12, 10, 10, 14 of 62 β 46 of 248 | same | same |
| what | counts | source |
|---|---|---|
| who commissioned it | OpenAI commissioned 300 core problems and 50 Tier 4 problems: 350. 53 core solutions are held out from OpenAI; 20 of the 50 Tier 4 problems are held out; the work is supported by OpenAI | About FrontierMath, no publication date; its latest update note is dated 2026-09-22; read 2026-10-09 |
| the 2026-06-12 correction | 123 problems corrected in Tiers 1β3 and 12 in Tier 4; 5 and 7 removed β 147 of 350 in all, added here; now 338 problems (295 + 43), 12 of them public (10 from Tiers 1β3, 2 from Tier 4) | FrontierMath Tier 4 v2, 2026-06-12 |
| o3 and GPT-5, run by Epoch on the private core set of 2025-02-28 | o3-2025-04-16_high and gpt-5-2025-08-07_high: each printed as a mean score; the file states no problem count, so neither is converted. OpenAI's December 2024 claim was on an earlier version, FrontierMath_11-26-24, which held 180 questions |
Epoch benchmark data, file frontiermath.csv, runs started 2025-11-16 and 2025-11-13; the 180: About FrontierMath, as above |
| the human baseline at MIT | 8 teams of four or five, 23 problems, 270 minutes, internet allowed: all teams together, 8 of 23; the average team is printed as a mean over 8 teams, not a count | About FrontierMath, as above |
| who | what | counts | source |
|---|---|---|---|
| Epoch AI | MATH Level 5: the 1,324 hardest test problems of the same MATH set Study 56 reads | gpt-5-2025-08-07_high: printed as a mean score; samples per problem not stated; not a count of problems, and not converted. Three scorers; the plotted one asks gemini-1.5-flash-002 whether the answer matches the reference |
MATH Level 5, no date found on the page, read 2026-10-09; the run started 2025-10-29 (benchmark data, file math_level_5.csv) |
| Center for AI Safety and Scale AI | HLE: how often experts disagreed with the reference answers | 2 audit rounds of 200 sampled questions each; the disagreement rate is an estimate from samples, not a count of questions | HLE paper, v11 2026-07-28 |
Study 56 reads 8 published files with Affine's own Swift, each pinned by its byte count and sha256: MATH, and OpenAI's prm800k release β the repository at commit 7ecc7947 (2023-06-01) and its scored file of 2,159,374,091 bytes (last modified 2023-05-25), which the repository's README says holds the large-scale samples behind Figure 3 of Let's Verify Step by Step (arXiv 2305.20050, v1 2023-05-31). Every value is an integer or a fraction of integers; a decimal point refuses its text. Each count below is the law's own print, committed in the folder and at the commit named, unless the row says it was added up here. Study 56 gives every table in full.
| what the law computes | count | commit |
|---|---|---|
| MATH problems whose key closes to an exact Affine value | 10,533 of 12,500 | f2d890e48 |
| MATH problems whose key cannot close, each with its named reason | 1,967 of 12,500 | f2d890e48 |
| MATH problems not read | 0 | f2d890e48 |
| the 500 problems OpenAI scored: key closes / cannot close | 417 / 83 | f2d890e48 |
| OpenAI's 815,632 attempts, each decided against the MATH key: EQUAL | 362,184 | f2d890e48 |
| UNEQUAL | 243,578 | f2d890e48 |
| CANNOT_CLOSE, each with its reason | 209,850 | f2d890e48 |
| NOT_READ_ROOM (an answer past 10,000 digits) | 20 | f2d890e48 |
| run again after the seal: attempts that reproduce their sealed outcome, the 20 not read among them | all 815,632; 0 left out of the join to the sealed records | 3c25e8cf9 |
Why the 83 keys cannot close, by problems: a root 24 Β· letters 20 Β· pi 10 Β· a written decimal 9 Β· words 7 Β· the imaginary unit 6 Β· infinity 4 Β· a union of sets 2 Β· a container of containers 1. No attempt on these 83 can close, whatever it writes; 134,995 attempts sit on them (3c25e8cf9).
These attempts were drawn for selection. Their paper says its generator wrote 1,860 solutions for each test problem (Figure 3 caption), for a reward model or a majority vote to pick one, and their README says the file holds the samples used to evaluate their reward models: up to 1,860 per test problem, fewer where a solution reached no answer within 1,024 tokens. The paper does not state the sampling temperature of these solutions. The law counted between 152 and 1,860 attempts per problem in the file. A count of different values over samples drawn this way can only grow as more samples are drawn.
| what the law counted | count | commit |
|---|---|---|
| attempts whose answer reads to an exact Affine value | 618,869 of 815,632 | f2d890e48 |
| attempts whose answer reads to no Affine value | 196,743 | f2d890e48 |
| attempts past the room (an answer longer than 10,000 digits) | 20 | f2d890e48 |
| different exact values written, counted once per problem and added over the 500 | 33,586 | added up 2026-10-09 from the 3c25e8cf9 table |
| problems on which the attempts wrote more than one value | 471 of 500 (414 of the 417 whose key closes) | counted 2026-10-09 from the 3c25e8cf9 table |
| problems on which they wrote exactly one value | 15 of 500 (3 of the 417) | counted 2026-10-09 from the 3c25e8cf9 table |
| problems on which no attempt read to a value | 14 of 500, all 14 among the 83 whose key cannot close | counted 2026-10-09 from the 3c25e8cf9 table |
| the middle problem, by number of different values | 32 values | counted 2026-10-09 from the 3c25e8cf9 table |
the most on one problem: test/algebra/2626
|
1,149 different values in 1,854 attempts; the key is 32,348; the value held most often, 32,372, by 78 attempts; 3 attempts equal the key |
3c25e8cf9; the row found 2026-10-09 |
one problem where they gathered: test/algebra/1004
|
1,860 attempts, 6 different values; 1,821 hold 4, which is the key | 3c25e8cf9 |
| problems where two values share the top count | 2 on keys that close: 1791 (key β3/8; the values 0 and 3/4 hold 104 attempts each) and 989 (key 12; 0 and 2 hold 6 each) |
3c25e8cf9 |
Counted from evidence/study-56-v3-compare-20261008/majority/scored.compare-v3.majority.tsv, column values (law), rows 4 to 503 β sha256 6bc0023748dbb486β¦ at 3c25e8cf9. The file is the law's sealed per-problem table, written after OpenAI's marks had been read; the sum and the counts of problems were made from it on 2026-10-09 and are not seals themselves.
| the law's outcome | their mark is_correct true (theirs) |
their mark false (theirs) | attempts |
|---|---|---|---|
| EQUAL β the answer and the key hold one exact value | 357,359 | 4,825, across 49 problems | 362,184 |
| UNEQUAL | 0 | 243,578 | 243,578 |
| CANNOT_CLOSE, with its reason β the law decides nothing | 75,708 | 134,142 | 209,850 |
| NOT_READ_ROOM β the law decides nothing | 0 | 20 | 20 |
| all | 433,067 | 382,565 | 815,632 |
Every cell is the comparison census at 3c25e8cf9; the cannot-close row adds its 40 reason lines, added up here on 2026-10-09. Of the 75,708 attempts their mark reads true where the law decides nothing, 73,200 sit on the 83 problems whose key cannot close (census, 3c25e8cf9).
On the 4,825. Their published grader, grader.py, compares two fractions as written strings only, and its comment says an unreduced fraction must not be marked correct (rule 11a, L275-278; theirs, not run here). No file of theirs states how is_correct was produced. The law compares values; that published rule compares the written fractions. One example: on test/algebra/1072 the MATH key is written \frac{243}{625}; the answers \frac{2187}{5625}, \frac{54675}{140625} and \frac{25 \cdot 3^6}{3 \cdot 5^6} each read to 243/625, which is the key, and their mark reads false (theirs). On that problem 1,011 attempts equal to the key carry their mark true and 178 carry it false.
| a vote over each problem's attempts | count | commit |
|---|---|---|
| the law's majority: the exact value held by the most attempts equals the MATH key | 293 of the 417 whose key closes; on the other 83 no majority is decided | 3c25e8cf9 |
| β¦ is unequal to the key / two values tie / both sides cannot close / the key cannot close | 119 / 2 / 3 / 83 | 3c25e8cf9 |
| their majority voting, best-of-1,860, in Let's Verify Step by Step (theirs) | a share equal to 348/500 (theirs); the paper does not say whether it is a count of problems or a rounded mean; majority voting is not in their eval.py; how their vote matches two answers or breaks a tie is not published |
transcribed at 3c25e8cf9 from the paper's Figure 3 table, page 7 |
| their process and outcome reward models, best-of-1,860 (theirs) | their eval.py, which makes these two selections, prints a mean over 400 trials; not a count of problems |
read at 3c25e8cf9
|
The law's 293 counts problems among the 417 whose key closes, and it decides no majority on the other 83. Their share, 348/500, is printed over all 500, and it is not stated to be a count. The law's majority is taken over OpenAI's own attempts; it solves no problem. The two figures stand side by side and are not subtracted: the law decides no majority on 83 of the 500, so these figures cannot show where any difference lies.
| what was run twice, or checked again | result | commit |
|---|---|---|
| the blind seal, run A and run B on one Mac: separate processes, separate HTTP streams of OpenAI's 2,159,374,091-byte scored file | 2 of 2 streams with the same bytes and the same sha256 46a01478dceefe81β¦; 26 of 26 output files byte-identical; 50 of 50 seals recompute (48 + 2); 14 of 14 local runs exit 0 |
f2d890e48 |
| the comparison, run A and run B on one Mac | 13 of 13 output files byte-identical; 26 of 26 seals recompute; 0 output files hold a decimal | 3c25e8cf9 |
| every HTTP stream of their scored file in the comparison rounds | 7 of 8 complete (one was stopped), all 7 with the same bytes and the same sha256 | 3c25e8cf9 |
| every committed file of both Study 56 evidence folders against its own SHA256SUMS, checked again on 2026-10-09 | 41 of 41 and 217 of 217 match | 3c25e8cf9 |
| the Affine IDE agent swarm, the Study 56 goal, with 1, 2 and 5 agents | one goal seal, c6949876a8639991, for every agent count; census and index byte-equal to the committed record |
8eb8f4ca9 |
| history: the same goal on an earlier reading that turned decimals into fractions, with 1, 2, 4 and 5 agents | one goal seal, 38c8079f18ed23d3, for every agent count |
af49986a8 |
| the nine Affine.Earth cells | not known: they do not yet answer Study 56 | β |
- It solves no problem. The substrate reads the key MATH's authors wrote and the answers OpenAI's model wrote, and decides whether each reads to an exact integer or fraction and whether the two are equal. It does not check that a key is right: an answer equal to a key holds the key's value, whatever the key is.
- The substrate has read MATH, and only MATH. It has not read AIME, HMMT, USAMO, the IMO, CMO, Putnam, FrontierMath, MathArena's sets or HLE. Every figure in Β§2 and Β§3 about those tests is theirs, and the substrate makes no figure of its own beside them.
- It has read none of the models named in Β§2 and Β§3. The 815,632 attempts it read were written by the large-scale generator of Let's Verify Step by Step, which the paper finetunes from the base model it names GPT-4, and published in 2023; the figures in Β§2 and Β§3 were published from 2024 to 2026.
- It decides answers; it does not write proofs, and it does not grade them. The olympiad and Putnam rows are proofs graded by people or by models; nothing on this page grades a proof.
- It does not write the attempts. The sampling settings that produced them are not stated in their paper beyond the 1,860 per problem and the 1,024-token cut-off. The substrate counts what was written and decides each answer against the key.
- Where it decides nothing, it says why. 1,967 MATH keys cannot close, 83 of them among the 500; 209,850 attempts cannot close, each with its reason, and their mark reads true on 75,708 of them; 20 lie past the room. The law does not take roots, evaluate functions, or read pi, i, infinity, letters, words or decimals as values.
-
Not known: how OpenAI produced
is_correct; whether their majority-voting figure is a whole count or a rounded mean; how their majority vote matches two answers or breaks a tie. Known: theireval.pyprints its reward-model selections as a mean over 400 trials. - The run-to-run studies in Β§1 were measured on other models and other sets β one of them on a prompt that is not mathematics. None was measured on OpenAI's 815,632 attempts. The count of different values per problem in Β§4 is the substrate's own, over those attempts, which were drawn as 1,860 samples per problem.
-
On which machines. The seals came out byte-identical every time the study ran: twice on one Mac (Β§4), and with 1, 2 or 5 agents. No other machine has run it yet. The nine Affine.Earth cells do not yet answer Study 56: no court on them reads it, so whether they write the same bytes is not known. The precedent is Study 55, which the cells answer with one digest on 9 of 9 (
043de8021). - What is tabled: 29 of the 31 models in MathArena's AIME 2025 files are tabled. The tables of the labs and the evaluators are what was read on 2026-10-09; they are not every figure every lab has printed.
The sources of the labs, the evaluators and the run-to-run studies were read on 2026-10-09; the three sources of the attempts the substrate read were read on 2026-10-08, and their entries say so. Where no date was found on a page, the entry says so.
Labs (theirs)
- OpenAI β Learning to reason with LLMs, 2024-09-12
- OpenAI β OpenAI o3-mini, 2025-01-31
- OpenAI β Introducing o3 and o4-mini, 2025-04-16
- OpenAI β Introducing GPT-5, 2025-08-07
- OpenAI β Introducing GPT-5.2, 2025-12-11
- OpenAI β Introducing GPT-5.4, 2026-03-05
- OpenAI β GPT-5.6, 2026-07-09
- OpenAI β GPT-6 Astra, 2026-09-03 (no date on the page; the date is from third-party reports); a second read on 2026-10-09 returned HTTP 403, so its two rows are as first read that day
- OpenAI β Alexander Wei, IMO 2025 thread, 2025-07-19; proofs
- Google DeepMind β AI solves IMO problems at silver-medal level, 2024-07-25
- Google DeepMind β Gemini Deep Think at the IMO, 2025-07-21
- Google DeepMind β Gemini 3 Pro model evaluation, November 2025
- Google DeepMind β Accelerating mathematical and scientific discovery with Gemini Deep Think, 2026-02-11
- xAI β Grok 3, 2025-02-19 Β· Grok 4, 2025-07-09 Β· Grok 4 Fast, 2025-09-19
- DeepSeek β DeepSeek-R1, 2025-01-22 Β· DeepSeek-Prover-V2 repository, 2025-04-30 Β· DeepSeek-Prover-V2 paper, v2 2025-07-18 Β· DeepSeekMath-V2, 2025-11-27 Β· DeepSeek-V3.2 paper, 2025-12-01 Β· DeepSeek-V4-Pro, 2026-04-22
- Alibaba (Qwen team) β Qwen3.5-397B-A17B, 2026-02-16
- Moonshot AI β Kimi K2 Thinking, 2025-11-04 Β· Kimi K2.6, 2026-04-14
- Mistral AI β Magistral, 2025-06-10
- Meta β Llama 4 model card, 2025-04-05
- Harmonic β Aristotle, 2025-10-01
- dots team β dots at IMO 2026, July 2026
- Huawei, as reported by IT Home β article, 2026-07-22
The attempts the substrate read (theirs)
- OpenAI β Lightman, Kosaraju, Burda and others, Let's Verify Step by Step, arXiv 2305.20050, v1 2023-05-31, the only version; the abstract page and the PDF read 2026-10-08
- OpenAI β prm800k repository, commit
7ecc7947, 2023-06-01, MIT licence; read at that commit 2026-10-08 - OpenAI β the scored file, scored-test-samples.jsonl, 2,159,374,091 bytes, last modified 2023-05-25; its headers read 2026-10-08
Independent evaluators and studies of run-to-run variation (theirs)
- MathArena β home, no date found on the page Β· IMO 2025 post, 2025-07-17 Β· USAMO 2026 post, Proof, Not Bluff, 2026-03-28 Β· Putnam 2025 post, 2026-02-26 Β· no final answer, 2026-05-12 Β· BrokenArXiv, 2026-03-13
- MathArena β competition tables, no date found on any of them, each read 2026-10-09: IMO 2025, USAMO 2025, USAMO 2026, Putnam 2025, AIME 2026, HMMT February 2026, BrokenArXiv February
- MathArena β published AIME 2025 I outputs and II outputs, created 2025-04-16, last modified 2026-05-15
- BalunoviΔ, Dekoninck, Petrov, JovanoviΔ, Vechev β MathArena: Evaluating LLMs on Uncontaminated Math Competitions, 2025-05-29
- Petrov, Dekoninck and others β Proof or Bluff?, 2025-03-27
- UK Mathematics Trust β IMO 2025 results, 2025-07-18
- Simon Willison β OpenAI's IMO announcement, 2025-07-19
- Epoch AI β About FrontierMath, no publication date, latest update note 2026-09-22 Β· FrontierMath Tier 4 v2, 2026-06-12 Β· MATH Level 5, no date found on the page Β· benchmark data, no date on the file; the runs cited started 2025-10-29 to 2025-11-16
- Center for AI Safety and Scale AI β Humanity's Last Exam, v11 2026-07-28
- Vals AI β AIME, updated 2026-04-16
- Artificial Analysis β intelligence benchmarking methodology, no date found on the page
- Thinking Machines Lab β Defeating Nondeterminism in LLM Inference, 2025-09-10
- Atil and others β Non-Determinism of "Deterministic" LLM Settings, 2024-08-06
- Yuan and others β Give Me FP32 or Give Me Death?, 2025-06-11
- Hochlehnert and others β A Sober Look at Progress in Language Model Reasoning, 2025-04-09
The substrate (Affine.Earth, committed)
-
evidence/study-56-v3-blind-seal-20261008/β the blind seal, commitf2d890e48 -
evidence/study-56-v3-compare-20261008/β the comparison, commit3c25e8cf9 - Study 56 β MATH, read exactly
Rights β source-available, all rights reserved. This wiki and its repository are published for public inspection and to let anyone re-derive the figures. They carry no LICENSE; under default copyright, all rights are reserved. No right is given or intended to use, run, or deploy it for any purpose other than re-deriving the published figures, nor to modify or build on it β any other use requires a written licensing agreement with the authors. Β· Affine.Earth Β· zero float Β· zero shear
Each step is the reason the next one exists. Nothing here is medical advice, and no page calls any medicine safe or unsafe.
1 Β· Why an exact safety screen at all
- Cures Without the Gatekeeper β the medicine front door: six real written medicines, one screen anyone can re-run
- The library admission law β what may enter, and the 71 arms that prove it refuses. The primary artefact.
2 Β· The three libraries, which grow rather than close
- The Library of Compound Cures β exact off-target maps for the medicines the registry publishes
- The Library of Proteins β 80,080 generated sequences, novel chemical matter, graded honestly
- The Library of Material Systems β what a system is, what was measured, where the law lives. C-007 absolute: no recipes
3 Β· The maps β every place a molecule could act, counted
- The off-target atlas β every nucleic-acid medicine the registry publishes a sequence for: WHERE it can pair
- The order of the bases β WHETHER THAT BURDEN IS UNUSUAL: 472 strands ranked against sixteen rearrangements of their own bases
- Where else could this guide cut? β the whole human genome, counted
- Designed, or forced by its own bases? β every clinical CRISPR guide, with its own composition as the control
- What a public genome deposit will tell you β and four ways it will mislead a health tool first
- Study 45 β which of nine billion answers a laboratory can act on β a safety review of AlphaGenome Atlas, measured live on 1,200 real variants at two genes. The headline score separates every one. The detailed tracks do not: splice-site usage hands back 950 of every 1,000 values shared with another variant at HBB and 998 at CFTR, and the shared values pile up in the quiet band where a bench clears a variant
4 Β· One medicine at a time
- Zilganersen β the first treatment for Alexander disease, screened on the real approved sequence
- A drug an AI designed β rentosertib for pulmonary fibrosis, and exactly what our instruments reach
- CAR-T, halted β the verdict a regulator could re-derive
- N-of-1 antisense β the only safety net at a population of one
- VERVE-102 β the off-target lattice a stranger can re-derive
- PM359 β prime editing, certified before anyone is dosed
- Del-Zota β the one safety question that can be made exact
5 Β· What keeps a disease alive, and what moves it
- Study 26 β master regulator bonds β 17 tumour types, 7,673 tumours; eleven compound pairs where no single agent among 20,308 cleared any
- Study 20 β Rife frequency β light and frequency, measured rather than dismissed
- Study 37 β five molecules β 37,910 "validated discoveries", 5 distinct molecules; why per-item validation cannot see a corpus-level defect
- Are the generated cures new? β 80,080 peptides against the human proteome
- Study 16 β disease type Β· Study 17 β chemistry InChIKey Β· Study 14 β protein lattice
- No language model in this stack β what the answers here are made of: measured 2026-09-12, no cell runs a model process, opens a model port or holds an unmasked model unit, and a gate refuses their return
- Run any study in your browser β all ninety programs open on your own device, forty-nine run there, and the run tells you whether it printed the sealed bytes
- The ontology β grades, terminals, controls, and what each page may say
- Zero Float Β· Zero Shear β the method in one page
- Ask someone you trust to check this β what to hand a sceptic
- Readersβ guide Β· Program index β all 42 studies Β· White paper Β· Roadmap
- The full-grade replacement β 49 retired instruments, 4 verticals
- The exactness seam β the business case
- Build a study β Falcon walkthrough β how to add one yourself
The same move every time: take a domain where a floating-point model is the accepted instrument, compute the same quantity in exact integers, and seal the cases where the two render opposite verdicts. The subject under grading is always the instrument, never the phenomenon.
- Study 48 β the atom already has an address β silicon dimers 3.840 Γ apart, the smallest commanded scale on the board: a length carried in single precision mis-addresses its first atom at step 8,783; an address cannot
- Study 49 β the phase code never needs Ο β a phase-only modulator takes 256 codes per pixel; the code is a ratio of integers
- Study 50 β CMS raw data from the LHC, read exactly β CMS's 2011 collision bytes streamed from CERN Open Data into the Affine IDE and read in exact integers, every collision a hologram you can turn: 138 of 3,564 bunch slots carry 93,110 of 120,742 collisions, and in 3,854 the event record reads its slot exactly 3 lower than the pixel boards Β· public release
- Study 55 β IceCube: the light in the ice, hit by hit β IceCube's calibrated hits read byte for byte: 4 published files, 9,749 events, 2,289,821 hits, a census seal per file
- Study 47 β translation shear: the meaning that survives a language β LAW FROZEN Β· LIVE CLAIM, measured 2026-09-11 and again fleet-wide 2026-09-12: translation as an exact coordinate transform, charts derived in memory at every start from the raw rows of a pinned public weight file and never written down; one lattice digest on 9/9 cells, zero drift, every refusal named. The generative comparison arm is ABSENT β there is no generative translator in the stack
- Study 34 β the observer-invariant verdict β why a safety verdict needs an exact law, not a bigger computer
- Study 35 β the safety brain that forgets β deaf in 8.4 seconds, forgets across machines, disagrees with itself
- Study 36 β the language game of Fermat's Last Theorem β guess and shear, or project
- Study 40 β the number the simulation throws away β their ICO result computed as a fraction; in float the effect returns 0 at every width, and an effect returned as zero cannot be searched for
- Study 41 β fifty years of solving the wrong problem β the ordering was never about time, it was about arithmetic; 177Γ the work and 2,400Γ the wrong guesses to return the answer the machine already had
- Study 42 β The Exact Contract β 2.7M flood settlements in Int128 cents; the step exists and the rigidity does not
- Study 29 β continuous-model shear
- The lattice holds Β· Impact study β continuum dead Β· Death of continuous shear
- Fourier Phantom β Anima FNO vs 11+12+13 Β· Stellar dynamo kill shot
- QCD: freedom is dilation Β· UUM-8D vs IUT β WIN
- Peer-review bundle Β· Conjecture alignment
- We need fusion β the verdict every machine can check
- Affine Fusion Control β the local exact-integer court Β· public release
- Fusion researcher's guide
- Study 33 β the fusion control verdict court
- Every season, fifty tonnes β the biosphere-safety case
- The forcing nobody measures Β· Impact study β the SpaceX trajectory
- Study 31 β the biosphere joint ledger β LIVE on the court, 9/9 cells
- Study 28 β the wet-bulb threshold court β Act 1 sealed
- Study 32 β the taxi-out floor court
- Where humans actually yield β the fatigue curves, and where the rules already agree
- Study 30 β sovereign edge pod Β· Manufacture contracts
- The detector that flags the whole market β a manipulation geometry in exact integers, and the regulator's own indicator scored against a legitimate quoter
- Study 43 β almost every order is cancelled, and that is normal β nine sessions, three operators, two continents: 935 to 998 of every 1,000 orders that ended, ended without trading. A check that flags almost everything is a denominator, not a detector β and the stock you pick moves it further than the exchange does
- Study 44 β nine billion answers, four billion ways to say them β AlphaGenome Atlas ships 9 billion predictions in single-precision floats, which hold 4.28 billion distinct values: 52 of every 100 variants MUST share a score with another. Agreement and exhaustion look identical on the wire
- Study 38 β the loss-reserve triangle β a reserve is an exact rational; 481 of 482 verdicts identical in both arithmetics; the sixteen-billion figure comes from an unchecked premise
- Study 39 β the actuarial domain β life, pensions, multi-state and aggregation; the margin is 8 significant digits at its tightest
- Run any study in your browser β the βΆ badge beside a program name opens it in the Studio, already built and carrying its inputs, and runs it on your machine with nothing sent back
- Explore the live courts
- MCP user guide β all 51 tools Β· Deterministic no-float courts for LLMs
- Court Client β generic wasm IDE for every court Β· Court-client checkpoint
- Coding Court β the verdict IS the artifact
-
Zed β the coding agent, for developers β set Zed 1.20.2 up on
https://affine.earth/v1, no language model anywhere; what a turn does, the wire, the autonomous closure -
Zed β Minecraft comes to life β the two-person interaction, sealed: it asks, cites, clones a sibling with a value you supply, verifies by replay; the court flips
REFUSED_UNKNOWN_BUDGET β WIN - Zed β the agent that teaches the whole domain β architecture, protocols, server management and git, each answered from lines it read and instruments it ran; five closures PROVEN, and the cattle question answered with a counter the fleet did not have
- Math Court on Glama Β· Math Court user guide Β· Example app β entire court
- Quantum algorithms inventory Β· Shor witness certifier
- MCP clients (public)
- Glama connector
- Look in the UI (no visitor data)
A study appears here under the state its evidence has earned, and above under the question it answers. The two are different filings of the same work, on purpose.
β LAW FROZEN Β· DATA SEALED
- Study 06 β explosion vs earthquake Β· Study 07 β Sgr A* raw visibilities
- Study 11 β Ehrhart volume Β· Study 12 β parallel repetition Β· Study 13 β Connes rigidity
- Study 14 β protein lattice Β· Study 16 β disease type Β· Study 17 β chemistry InChIKey
- Study 18 β material STD Β· Study 19 β Go First dice
- Study 26 β master regulator bonds β 17 tumour types, every finding published
- Study 56 β MATH, read exactly β 12,500 problems read blind: 10,533 keys close, 1,967 cannot, each with the law's reason; OpenAI's marks and printed figures set beside it afterwards, labelled theirs
- Frontier models in mathematics β beside Study 56 β they train their models and sample them; Affine.Earth trains nothing: each lab's printed figure, labelled theirs with its runs and its grader, beside one sealed reading of every key and every answer
π΄ LIVE CLAIM β standing, not sealed
- Study 02 β launch ionospheric holes Β· Study 02 β regulatory alarm
- Study 09 β global convective bond Β· Study 20 β Rife frequency Β· Study 21 β stellar dynamo
- Study 22 β 2-local Hamiltonian Β· Study 23 β spin glass Β· Study 24 β N-representability Β· Study 25 β exact permanent
π CHARTER Β· OPEN β the findings, published either way
- Study 03 β flare SIDs β archive went dead Β· predictions and validations
- Study 04 β tsunami vs surge β partial seal Β· Study 05 β Forbush decreases
- Study 08 β Gaia BH1 β no corpus until DR4 Β· Study 10 β Fermi / dark matter β does not disprove DM
- Study 15 β Skala DFT shear Β· Study 27 β exact nuclear scattering
- Overview Β· First 27 days Β· Success criteria
- The science, and what history says Β· Blind spots β five stories magnitude models miss
- Historical corpus Β· Data archives β every source, exactly how to reach it
- Model shear Β· Benchmark results Β· Prediction registry
- Substrate architecture β how a shadow becomes a geometry
- Operations runbook Β· Satellite & aviation advisory