Repository navigation
Study 56 MATH Read Exactly
Study 56 β MATH, read exactly: 12,500 keys and 815,632 graded attempts, one verdict for every answer the law can hold
Frontier labs print mathematics figures built from sampled runs: OpenAI printed o1's AIME 2024 result for one sample, a consensus of 64 samples and a re-ranking of 1,000 (theirs). Run it again and the answers can change: on the 30 AIME 2025 problems, 465 of 870 modelβproblem pairs returned more than one answer across four runs (counted from MathArena's published rows, theirs). They train their models and sample them.
Affine.Earth trains nothing and samples nothing. One law of exact whole numbers reads the bytes: 12 Swift files, 0 imports, no float. Study 56 puts that law to MATH, a public benchmark of 12,500 competition problems, and to OpenAI's prm800k release, 815,632 scored attempts at 500 of them. It answers what a sample cannot: which answers equal their key exactly. Whether the nine Affine.Earth cells read them the same way is not known: no cell runs Study 56 yet.
MATH and PRM800K were built to train models. MATH's own train folder holds 7,500 problems. PRM800K's 800,000 step labels trained OpenAI's process reward model, and 4,500 MATH test problems were moved into its training set (theirs, their README). Affine.Earth trained on none of it: it read every one of the 12,500 problems, train and test alike, as a claim to decide.
Answers change even at temperature 0, where nothing is sampled. Thinking Machines Lab names the cause: kernels whose order of addition changes with batch size (theirs). Yuan and others, under greedy decoding: floating-point addition is not associative at limited precision (theirs; both in Frontier models in mathematics). The substrate adds whole numbers: split across 1, 2 or 5 agents and merged in forward, reverse or interleaved order, the Study 56 goal seals one set of bytes, and all 815,632 attempts reproduce their sealed outcome, the 20 not read among them.
- One verdict for every answer the law can hold. All 12,500 keys and 815,612 of OpenAI's 815,632 attempts have one verdict each. A key closes or cannot close, with its reason; an attempt is equal or unequal to the key, or cannot close, with its reason. The other 20 attempts are longer than the law's room of 10,000 digits; they are named and not read. The law was designed and sealed without the marks, and committed before the marks were read again (an earlier reading had read them once, on 2026-10-07). Run again after the seal, all 815,632 attempts reproduced their sealed outcome. Run twice on one Mac, the law wrote byte-identical files: 26 of 26 for the seal, 13 of 13 for the comparison. The Affine IDE's agent swarm seals one set of bytes with 1, 2 or 5 agents.
- 1,967 of MATH's 12,500 answer keys are not a whole number or a fraction of whole numbers, the only values the law holds, and the law names why for each. 595 hold a root Β· 383 hold letters Β· 265 hold pi Β· 256 are written with a decimal point Β· 165 are words Β· 99 hold the imaginary unit Β· 91 hold infinity Β· 71 hold a container of containers Β· 19 use a form outside the law's grammar Β· 11 are unions of sets Β· 6 hold a trigonometric function Β· 4 hold no answer Β· 1 holds a logarithm Β· 1 has a comma between digit groups. Of OpenAI's 500 problems, 83: 24 roots, 20 letters, 10 pi, 9 decimals, 7 words, 6 the imaginary unit, 4 infinity, 2 unions of sets, 1 container of containers.
- The law decides none of the 134,995 attempts on those 83 problems, and their mark reads correct on 73,200 of them (theirs). 47,027 of the 134,995 copy the key byte for byte, and a copy of a text the law cannot read is still not read to a value. How many of the 73,200 are such copies is not known.
- Every time the law reads an answer unequal to the key, their mark also reads false: 243,578 of 243,578, and true on none.
-
4,825 answers equal to the key are marked false (theirs). Where the law reads an answer equal to the key, their mark
is_correctreads false on 4,825 attempts in 49 problems (theirs). 1,793 attempts answer 864 and are marked false; their copy of the key reads864\mbox{inches}^2, and MATH's key is 864.\frac{2187}{5625}is marked false; it reads to 243/625, and the key is 243/625. -
Their grader normalises answers with sympy, with no sympy version pinned and no timeout (their
grader.py, theirs). It marks two answers alike when their normalised strings match or when sympy simplifies their difference to zero. No file of theirs states howis_correctwas produced. -
Side by side, never subtracted. Among OpenAI's own attempts at each problem, the answer most of them hold, counted by exact value, equals the key on 293 of the 417 problems whose key the law reads to a value. The other 83 have no key value to equal. Theirs (Let's Verify Step by Step, Figure 3, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it counts problems; their
eval.pyprints their two reward-model picks as means over 400 trials, not as counts. Their printed figures stand beside the law's as published; the two cannot be subtracted.
Theirs is a sample. This reading came out byte-identical every time it ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.
MATH is a public set of 12,500 competition problems in seven subjects, from algebra to precalculus, each given a level from 1 to 5, except two Geometry problems whose level is written "?". Every problem comes with a worked solution, and its answer β the key β is the last \boxed text of the solution. OpenAI's later release, prm800k, is built on it. It holds 815,632 attempts written by a model at 500 of the test problems, each with a mark called is_correct, and four files of model solutions whose steps people rated one by one. Study 56 reads all of it with Affine's own Swift and answers every problem and every attempt only as it relates to the Affine substrate. A text closes when the law reads it to an Affine value: an integer, an exact fraction of two integers, or a base-N integer. Any other text cannot close, and the law says why in a sentence of its own β a root is written; pi is not a fraction of integers; letters make an expression, not a value; a decimal point is a float's spelling. A text is not valid just because it is written. The law does not check that a key is right: an answer equal to a key holds the key's value, whatever the key is.
The law did all of this blind: its reader passes over the marks by name, and its results were sealed and committed before the marks were read again. Only then were the marks set beside it, labelled theirs. Every count is sealed with SHA-256, and the seals came out byte-identical every time the study ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.
What the labs that build frontier models print about their own mathematics, and what independent evaluators counted run by run, is set beside this reading, each figure labelled theirs with its runs and its grader, in Frontier models in mathematics β a share over sampled runs, and one sealed reading of every key and every answer.
Status: SEALED BLIND, THEN SET SIDE BY SIDE β 2026-10-08 β 8 published files opened by the law on one Mac: all 12,500 MATH problems, test and train, and every attempt that OpenAI's prm800k release holds for them. All of it sealed and committed before the marks were read again; OpenAI's marks and published figures, labelled theirs, set beside it afterwards. The commits are in Β§10. In Affine IDE 0.2.6.6 Ξ² the stage, the Mathematics folder and the swarm read this sealed record (Β§9).
Program index Β· Study board Β· Study 55 β IceCube: the light in the ice Β· Study 50 β CMS raw data from the LHC
Eight files. Each is pinned in the law by its byte count and its sha256, and the law checks both before it reads a record: the bytes first, then the digest. All eight pins read VERIFIED.
| file | what it holds | bytes | records | distinct problems | sha256, first 16 |
|---|---|---|---|---|---|
MATH.tar |
MATH: one JSON file per problem, with its text, subject, level and worked solution | 20,327,936 | 12,500 | 12,500 | 0fbe4fad0df66942 |
math_splits/test.jsonl |
prm800k (theirs): the 500 test problems and their keys | 446,564 | 500 | 500 | 35dc41080a368085 |
math_splits/train.jsonl |
prm800k (theirs): the training problems and their keys | 10,896,985 | 12,000 | 11,999 | 90d96daeac3fe343 |
scored-test-samples.jsonl |
the scored attempts at the 500 test problems, with is_correct (theirs) |
2,159,374,091 | 815,632 | 500 | 46a01478dceefe81 |
phase1_test.jsonl |
phase one: solutions whose steps people rated (theirs) | 829,105 | 106 | 101 | f4b3bc5b095e45c8 |
phase1_train.jsonl |
phase one, training | 7,900,236 | 949 | 903 | e9da6a73f827ffb9 |
phase2_test.jsonl |
phase two: model solutions, the model's own steps and answer, step ratings (theirs) | 12,240,719 | 2,762 | 458 | 6b172efa884ac834 |
phase2_train.jsonl |
phase two, training | 456,135,365 | 97,782 | 10,828 | 1110237feeb51d1b |
-
MATH.tar β the Wayback capture 20231017015604 of
people.eecs.berkeley.edu/~hendrycks/MATH.tar: 12,518 members, 17 directories and 12,501 files; 12,500 problems, test 5,000 and train 7,500. MIT is the hendrycks/math repository licence (GitHub API); the tar's own README states none. -
math_splits and the phase files β the openai/prm800k repository at commit
7ecc794703b2877f(theirs), MIT licence; LICENSE, 1,062 bytes, in the repository tree. -
The scored file β
https://openaipublic.blob.core.windows.net/process-supervision/scored-test-samples.jsonl. No licence is stated on the blob or in its HEAD response. It is streamed, never stored: 0 files of exactly 2,159,374,091 bytes were found after the runs. It holds between 152 and 1,860 attempts per problem, as the law counted them; neither of OpenAI's sources states the smallest. -
Where the Affine IDE fetches two of them β the IDE fetches
MATH.tarandphase2_test.jsonlfromhttps://affine.earth/language-game/studies/math56/and keeps each only when its byte count and then its sha256 equal the pin. The scored file is never stored and never served.
MATH.tar is the spine, one row per problem; every other source joins to it. No phase-file join is ambiguous. 3,854 phase2_train records name 10 problems that join no math_splits record; they are decided against their own source's key and credited to no row.
An Affine value is an integer of any width, carried as a decimal-string integer by the mesh's own arithmetic (ArbitraryStringMath.swift, vendored byte-identical from its one home, 6,379 bytes, commit 73c986c9e); an exact fraction of two such integers in lowest terms; or a base-N integer read to its integer. The law is 12 files with 0 import lines, no float type and no float literal. It writes no integer longer than 10,000 decimal digits; a text that would pass that room is not read (NOT_READ_ROOM), a state that decides nothing and says nothing of the text's meaning.
No decimals. A decimal point is how a float is written, so the law refuses the text and reads no value from it. The law's sentence:
WRITTEN_WITH_A_DECIMAL a decimal point is a float's spelling; Affine refuses the text and reads no value from it
The substrate holds integers and fractions of integers, so that every cell reads the same value from the same text. The law does not turn a decimal into a fraction. The rule is decided first, at every level: key, answer, option, container member, the right side of an assignment, a side of a step. The code that read decimals into fractions is deleted. 256 MATH keys are written with a decimal (test 99, train 157); 9 of them are among the 500.
Invent nothing. Everything this law adds to the earlier reading (Β§8) is a composition of a law the substrate already has:
| what it reads | counts | |
|---|---|---|
| 2a | wide integers, by the mesh's decimal-string arithmetic | 6 MATH keys; answers scored 860, phase2_test 4, phase2_train 108 |
| 2b | identity: byte equality after the earlier reading's own unwrapping (Β§8) | step pairs scored 2,710, phase2_test 18, phase2_train 376, MATH solutions 4 |
| 2c | containers β tuples, sets, intervals β member by member | 523 MATH keys; answers scored 29,126 |
| 2d |
x = value, read as its value |
32 MATH keys; answers scored 39 |
| 2e | a choice letter, read as that option of the problem's own text | 788 option texts resolved: 321 to a value, 467 cannot close; 68 letter keys, 33 resolved, none to a value |
| 2f | a word, by its exact bytes inside \text{β¦}
|
1 MATH key |
Not added: no number fields, no radical canonical forms, no polynomial algebra, no transcendental rings, no tolerance, no refinement and no bound.
Cannot close β why a text lacks shared meaning. When the law cannot read a text to an Affine value, it names the reason in its own words. The first reason in this order is the one recorded.
| reason | the law's own sentence | MATH keys | keys of the 500 |
|---|---|---|---|
NO_ANSWER |
there is no text to read: the field is absent or null, empty, the string None, or holds no braced boxed answer | 4 | 0 |
WRITTEN_WITH_A_DECIMAL |
a decimal point is a float's spelling; Affine refuses the text and reads no value from it | 256 | 9 |
PROSE_OR_WORD |
words carry no number | 165 | 7 |
HOLDS_LETTERS |
letters standing for unknowns make an expression, not a value | 383 | 20 |
HOLDS_IMAGINARY_UNIT |
i squared is minus one, and no integer and no fraction of integers squares to minus one | 99 | 6 |
HOLDS_TRIG |
a trigonometric or hyperbolic function is written; the law does not evaluate functions | 6 | 0 |
HOLDS_LOG |
a logarithm is written; the law does not evaluate functions | 1 | 0 |
HOLDS_PI |
pi is not an integer and not a fraction of integers | 265 | 10 |
HOLDS_RADICAL |
a root is written; the law reads integers and fractions joined by plus, minus, times, divide, integer powers and factorials, and does not take roots | 595 | 24 |
HOLDS_INFINITY |
infinity is not an integer and not a fraction of integers | 91 | 4 |
UNION_OF_SETS |
a union, intersection or difference of sets is set algebra; the law compares one container of values at a time | 11 | 2 |
CONTAINER_OF_CONTAINERS |
a member that is itself a container (a matrix row, a list of pairs) is not a value | 71 | 1 |
COMMA_BETWEEN_DIGIT_GROUPS |
a comma between digit groups, in a text that is not wholly one integer written in comma groups, is written both as a thousands separator and as a list separator, and the text does not say which | 1 | 0 |
OUTSIDE_THE_GRAMMAR |
the text uses a symbol or a form the law does not read β among them a floor, a ceiling, a binomial, a modulus, an absolute value, a ratio, plus-or-minus, dots, a function name, an environment, a text-family command other than text, an unclosed bracket, a list on one side of an equals sign | 19 | 0 |
| all | 1,967 | 83 |
Five more reasons fall on no MATH key: digits written against the letter e, a float's exponent spelling (WRITTEN_WITH_A_FLOAT_EXPONENT); a unit written in words or a degree, percent or dollar mark, on a step side only (DECORATED); an equation or a comparison, a statement about values (RELATION_NOT_A_VALUE); the letter e (HOLDS_E); and a division by zero (DIVIDES_BY_ZERO). Two pair reasons fall where both sides read to values and still share no form: a container against a single value, or two kinds of container (KINDS_DIFFER); and two unbracketed lists in another order, since an unbracketed list does not say whether order counts (ORDER_NOT_WRITTEN).
Three levels. A problem: its key alone decides. An attempt: if the key cannot close, every attempt on it cannot close, with the key's reason; if the key closes, two values are equal or unequal; otherwise the attempt cannot close, on the answer's side, or because the two sides read to values that share no form. A step: there is no key; two values decide by value, and two byte-identical sides that lack a value and hold no decimal close by identity.
Identical spellings. An attempt that copies a key the law cannot read is not read to a value by being copied: being written does not make a text valid, and a text the law cannot read to a value is not read to one by being repeated. The identity is printed beside each such attempt and never counted as closed: scored 47,027, phase2_test 66, phase2_train 506 (505 on problems that join a MATH row, and 1 on a problem that joins none).
The reader compares every field name, byte for byte, with the names and fragments it withholds, and never decodes a withheld value:
withheld by name: is_correct rating corrected_rating finish_reason flagged prm_score orm_score pre_generated_verifier_score rating_probs chosen_completion human_completion label
withheld by any field containing: correct rating score grade verdict finish flag label_
Flip every withheld member and the seals do not move: BLIND_HOLDS, 5 of 5 sources. A review read back every recorded read and command of the runs: none returned a line from a withheld path. The law was designed and run without the marks; they had been read once before, on 2026-10-07, by the earlier reading (Β§8).
On 2026-10-08, UTC: at 22:03:15 every blind output and law file was verified against the git objects of the blind seal's commit (Β§10), 63 checks; at 22:35:07 the phase files' rating and finish_reason (theirs) were first read; at 22:35:14 is_correct (theirs) was first read. prm_score, orm_score, rating_probs and pre_generated_verifier_score were never read by any run: they are another model's decimal scores.
| problems | rows | key closes | key cannot close |
|---|---|---|---|
| all MATH problems | 12,500 | 10,533 | 1,967 |
| test split | 5,000 | 4,215 | 785 |
| train split | 7,500 | 6,318 | 1,182 |
| the 500 scored test problems | 500 | 417 | 83 |
The 10,533 closing keys read to: an integer 8,252 Β· a fraction 1,657 Β· a container 523 Β· a base-N integer 62 Β· an assignment 32 Β· a wide integer 6 Β· a word 1.
| subject (MATH's own) | problems | key closes | key cannot close |
|---|---|---|---|
| Algebra | 2,931 | 2,580 | 351 |
| Counting & Probability | 1,245 | 1,205 | 40 |
| Geometry | 1,349 | 976 | 373 |
| Intermediate Algebra | 2,198 | 1,689 | 509 |
| Number Theory | 1,409 | 1,369 | 40 |
| Prealgebra | 2,076 | 1,854 | 222 |
| Precalculus | 1,292 | 860 | 432 |
MATH's own worked solutions. The last \boxed text of a solution is the key, by MATH's own definition, so read against the key it is the key read against itself: 10,533 equal, 0 unequal, every other row cannot close with the key's reason. The steps of the worked solutions hold 63,448 = pairs: 11,146 equal, 4 by identity, 67 unequal, 52,226 cannot close, 5 past the room. 18,194 equals signs sit outside every span the step law opens; they are counted and not read.
Every attempt in every source.
| source | records | equal | unequal | cannot close | not read |
|---|---|---|---|---|---|
| scored attempts, the 500 test problems | 815,632 | 362,184 | 243,578 | 209,850 | 20 (room) |
| phase2_test, the model's own final answer | 2,762 | 558 | 1,415 | 789 | 0 |
| phase2_train, the model's own final answer | 97,782 | 7,061 | 60,305 | 30,409 | 7 (room) |
| phase1_test, the final answer | 106 | β | β | β | 106 (withheld by the blind seal; its rated steps read after it, Β§7) |
| phase1_train, the final answer | 949 | β | β | β | 949 (withheld by the blind seal; its rated steps read after it, Β§7) |
Each row adds to its records. Credited to the 12,500 problems, 912,322 attempts were read (the census of every problem counts them as attempts read, ATTEMPTS_READ): 675,101 closed, 237,194 cannot close, and 20 scored and 7 phase2_train answers past the room. The other 3,854 phase2_train records sit on 10 problems that join no MATH row; they were read too, against their source's own key, and all 3,854 cannot close: 3,853 on 9 problems whose source key holds no answer (NO_ANSWER), and 1 on 1 problem whose source key is a container of containers (CONTAINER_OF_CONTAINERS). They are inside phase2_train's 30,409. Of the 3,885 phase2_train attempts whose key holds no answer, those 3,853 are the unjoined records and the other 32 are credited to problem rows. The final answers of the 106 phase1_test and 949 phase1_train attempts are not read: the blind seal withheld them with the label tree. After the seal, the comparison joined all 106 and all 949 to their sealed records, with 0 left out of the join, and read their rated steps with the same step law: 3,119 pairs in the rated completions and 135 in the steps the labellers wrote in phase1_test, 31,620 and 1,316 in phase1_train (Β§7). The blind seal counts 1,082 attempts not read (ATTEMPTS_NOT_READ): the 1,055 phase-one attempts and the same 27 answers past the room. 1,175 problems have no attempt in any source.
All 815,632 attempts were joined to their sealed records and run again after the seal: all 815,632 reproduced their sealed outcome. Their mark is_correct (theirs) reads true on 433,067 and false on 382,565.
| the law's outcome | is_correct true (theirs) | is_correct false (theirs) | attempts |
|---|---|---|---|
| equal | 357,359 | 4,825 | 362,184 |
| unequal | 0 | 243,578 | 243,578 |
| cannot close Β· the key lacks shared meaning | 73,200 | 61,795 | 134,995 |
| cannot close Β· the answer lacks shared meaning | 2,507 | 67,627 | 70,134 |
cannot close Β· both sides read to values: a container against a single value, or two kinds of container (KINDS_DIFFER) |
1 | 4,143 | 4,144 |
cannot close Β· both sides read to values: two lists in another order (ORDER_NOT_WRITTEN) |
0 | 577 | 577 |
not read (NOT_READ_ROOM) Β· an answer past 10,000 digits |
0 | 20 | 20 |
| all | 433,067 | 382,565 | 815,632 |
Unequal, with their mark true: 0. Every attempt the law reads as unequal to the key, their mark also reads false.
Equal, with their mark false: 4,825 attempts, across 49 problems. In each, the answer and the MATH key hold one exact value. This page prints both columns and does not rule between them. Examples, as the law read them:
-
test/algebra/1072β\frac{2187}{5625},\frac{54675}{140625}and\frac{25 \cdot 3^6}{3 \cdot 5^6}each read to 243/625; the MATH key is 243/625; their mark, false (theirs). -
test/algebra/722β485149/49and99^2+99+1each read to 9,901, against the key 9,901. -
test/geometry/473β 1,793 attempts answer 864. The scored file's own copy of the key (theirs) is864\mbox{inches}^2; the MATH key's value is 864. -
test/prealgebra/1114β 550 attempts answer 15; the key copy (theirs) is15\mbox{cm}^2; the MATH key's value is 15.
The 83 problems whose key cannot close. No attempt on them can close, whatever it writes, because the key it would be decided against has no Affine value.
| why the key lacks shared meaning | problems | attempts | is_correct true (theirs) | is_correct false (theirs) |
|---|---|---|---|---|
HOLDS_RADICAL β a root is written |
24 | 37,862 | 19,015 | 18,847 |
HOLDS_LETTERS β an expression, not a value |
20 | 31,834 | 15,423 | 16,411 |
HOLDS_PI β pi is not a fraction of integers |
10 | 16,155 | 5,828 | 10,327 |
WRITTEN_WITH_A_DECIMAL β a float's spelling |
9 | 16,674 | 12,781 | 3,893 |
PROSE_OR_WORD β words carry no number |
7 | 12,722 | 9,942 | 2,780 |
HOLDS_IMAGINARY_UNIT β no fraction of integers squares to minus one |
6 | 9,556 | 5,352 | 4,204 |
HOLDS_INFINITY β infinity is not a fraction of integers |
4 | 6,326 | 4,107 | 2,219 |
UNION_OF_SETS β set algebra, not one container |
2 | 3,561 | 752 | 2,809 |
CONTAINER_OF_CONTAINERS β a member that is a container |
1 | 305 | 0 | 305 |
| all | 83 | 134,995 | 73,200 | 61,795 |
What they built (theirs): the paper "Let's Verify Step by Step" (OpenAI), version one. 1,860 solutions were generated per test problem (Figure 3 caption, page 7), and one is chosen by a process reward model, an outcome reward model or majority voting. A process reward model scores a solution as the product of its per-step probabilities. 4,500 MATH test problems went into training, and evaluation uses the remaining 500, selected uniformly at random (Appendix C). prm800k holds 800,000 step-level correctness labels (README L5).
How they graded (theirs): the solution ranked highest is graded automatically on its final answer, and the fraction correct is reported; the paper gives no implementation of the grader. In the repository, grader.py (8,101 bytes) counts two ways to be correct β the two strings normalise to the same string, or sympy simplifies their difference to zero (L236-240) β and eval.py (3,153 bytes) selects by the stored score and counts the stored is_correct; no grading happens in eval.py, and majority voting is not in it. No repository file states how is_correct was produced. The substrate did not run their grader; is_correct is printed exactly as stored.
The substrate's majority, by exact Affine value, over OpenAI's own attempts. It solves no problem: it counts the answers OpenAI's model wrote at each problem. Only attempts whose answer reads to an Affine value vote β 618,869 of 815,632. One value, one vote class. A tie is named and counted, and left a tie. The majority value is decided against the MATH key by the law's own decide.
| the substrate's majority, by exact Affine value | problems of 500 |
|---|---|
| the Affine majority value equals the MATH key | 293 (293 of the 417 whose key can close) |
| it is unequal to the key | 119 |
| a tie: two values hold the top count | 2 |
| the majority value and the key read to values that share no form (a container against a single value, or two kinds of container: 2; two lists in another order: 1) | 3 |
| the key cannot close | 83 |
Their figures, as published (theirs; Let's Verify Step by Step, Figure 3 inset table, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it is a count of problems or a rounded mean. Their eval.py, which makes the two reward-model selections, prints each as a mean over 400 trials, not as a count of problems, so neither is written here as k of 500.
The law's 293 counts problems among the 417 whose key closes; their share, 348/500, is printed over all 500 and is not stated to be a count. The two stand side by side and are not subtracted: the law decides no majority on 83 of the 500, so these figures cannot show where any difference lies.
The two ties: intermediate_algebra/1791 (key β3/8; the values 0 and 3/4 each hold 104 attempts) and precalculus/989 (key 12; the values 0 and 2 each hold 6 attempts).
Not computed, and why. Their two reward-model selections choose one attempt per problem by a model's decimal score (prm_score, orm_score); a decimal score is a float, so those fields were not read. Their step-score strategies (rating_probs) and the verifier score are decimal scores, not read. Their grader's normalisation is theirs, not run. Their majority vote's grouping is not stated in their paper. Anyone may read those scores on their own machine; no figure on this page comes from a decimal score.
What their sources state, recorded without comment (theirs): the paper does not state how majority voting decides two answers are the same, how it breaks a tie, or which solution it returns; their README states up to 1,860 scored samples per test problem, and the problem with the fewest holds 152 attempts as the law counted them.
| step pairs read | pairs | equal | by identity | unequal | cannot close | past the room |
|---|---|---|---|---|---|---|
| MATH worked solutions | 63,448 | 11,146 | 4 | 67 | 52,226 | 5 |
| scored attempts | 5,396,497 | 728,819 | 2,710 | 60,672 | 4,602,962 | 1,334 |
| phase2_test, the model's own steps | 18,476 | 2,248 | 18 | 300 | 15,910 | 0 |
| phase2_train, the model's own steps | 750,346 | 93,415 | 376 | 9,228 | 647,279 | 48 |
| phase1_test, the rated completions | 3,119 | 275 | 1 | 45 | 2,798 | 0 |
| phase1_test, the steps the labellers wrote | 135 | 12 | 1 | 0 | 122 | 0 |
| phase1_train, the rated completions | 31,620 | 3,132 | 74 | 925 | 27,489 | 0 |
| phase1_train, the steps the labellers wrote | 1,316 | 113 | 0 | 6 | 1,197 | 0 |
The phase-one pairs were read by the comparison after the seal, with the same step law; they are not part of the seal. The law reads no rating. Every step pair the phase files hold is printed beside its rating (theirs) in the comparison's censuses; the full tables are in the report and in evidence/study-56-v3-compare-20261008/.
An earlier reading turned a decimal into a fraction; this law refuses it: a decimal point is how a float is written, so the law refuses the text and reads no value from it. Only the decimal rule moves a value: no earlier equal became unequal and no earlier unequal became equal. In the scored file, 14,289 earlier equal and 21,986 earlier unequal attempts now cannot close, each for a decimal, and 12,487 equal and 10,176 unequal are newly decided out of the earlier reading's not-decided attempts. Of the MATH keys, 247 that the earlier reading read to a value now cannot close, each written with a decimal, and 562 that it read no value from now close. The earlier reading is history, not deleted: commit c56e3dbba holds it, its comparison is at 20af6f29e, and anyone may run it on their own machine. Its swarm goal, also history, sealed 38c8079f18ed23d3 with 1, 2, 4 and 5 agents (af49986a8).
In Affine IDE 0.2.6.6 Ξ² the stage, the Mathematics folder and the swarm read the sealed record of Β§3 to Β§6.
-
The Study 56 stage: three holograms β Key, the 12,500 MATH problems as a lattice by subject, level and split, the 500 test problems ringed; Samples, the 500 test problems as columns of their 815,632 attempts, the sealed outcome beside
is_correct(theirs) from chapter 3 on; Steps, one phase-two record as a tower of its steps and its=pairs. Six chapters in the chat, every figure read from the sealed record: 10,533 keys close, 1,967 cannot, and the 500 test problems are ringed. - The Mathematics folder: seven type folders in MATH's own order β Algebra 2,931 Β· Counting & Probability 1,245 Β· Geometry 1,349 Β· Intermediate Algebra 2,198 Β· Number Theory 1,409 Β· Prealgebra 2,076 Β· Precalculus 1,292 β each by level, 20 problems to a page. Selecting a problem opens Study 56 on it: its cube ringed on the Key picture, its text, its key, MATH's own answer (theirs), and how the substrate answered it: it closes, with its value, or it cannot close, with its reason in the law's own sentence.
-
The internal swarm: deterministic substrate workers running the substrate's own laws, never a language model, each owning a space and a goal in its own sandbox, sized to the machine as measured. The Study 56 goal runs this law and seals one set of bytes,
c6949876a8639991, with 1, 2 or 5 agents; its census and index are byte-equal to the committed record.
| commit | when | what |
|---|---|---|
c56e3dbba |
2026-10-07 20:32:30 -0400 | the earlier reading's blind seal β history (Β§8) |
20af6f29e |
2026-10-07 22:00:55 -0400 | the earlier reading's comparison β history (Β§8) |
f2d890e48 |
2026-10-08 18:01:19 -0400 | the blind seal |
3c25e8cf9 |
2026-10-08 21:30:24 -0400 | the comparison |
8eb8f4ca9 |
2026-10-09 10:44:53 -0400 | the Affine IDE reads the sealed record: the stage, the Mathematics folder and the swarm |
The census seals, first 16 hex digits, at f2d890e48:
| source | census seal |
|---|---|
| MATH.tar | f14929f70bd49d42 |
| math_splits_test | 62cae89ea1c30246 |
| math_splits_train | f98a49208cc1e0b7 |
| scored | 03b5a3a6ff1ba4f1 |
| phase1_test | 964e7d1395b06617 |
| phase1_train | cc1c9367f11f70a6 |
| phase2_test | e02306d6d9fbb45a |
| phase2_train | aed6415d3ff09154 |
| the 12,500 problems | 69e4a013c2fb79a3 |
The comparison's scored census at 3c25e8cf9: 1f482961312e517d. The swarm's Study 56 goal at 8eb8f4ca9: c6949876a8639991. Repeat runs: 26 of 26 blind outputs and 13 of 13 comparison outputs byte-identical between two runs, each from its own stream of the scored file. Every number on this page can be read from git at its commit, for example git show f2d890e48:evidence/study-56-v3-blind-seal-20261008/census/scored.v3.census.txt.
| item | status |
|---|---|
| the 1,967 keys that cannot close, and every attempt on them | no Affine value; each reason is named in Β§2 |
| answers and step pairs past the room (answers 20 scored and 7 phase2_train; step pairs 1,334 scored, 48 phase2_train, 5 MATH solutions) | not read; decides nothing |
| the 106 and 949 phase-one attempts | their final answers are not read: the blind seal withheld them; after the seal the comparison read their rated steps, 3,119 and 135 pairs in phase1_test and 31,620 and 1,316 in phase1_train |
| their reward-model selections, step-score strategies and verifier score | not computed: decimal scores |
how is_correct was produced |
not stated in any repository file (theirs) |
| whether their majority-voting figure is a whole count or a rounded mean | not stated (theirs) |
| how their majority voting matches answers or breaks ties | not stated (theirs) |
| how many of the 73,200 attempts marked correct on the 83 keys that cannot close copy the key byte for byte | not known: not counted |
| whether the nine Affine.Earth cells read Study 56 the same way | not known: no cell runs Study 56 yet |
Edge. Where a key lacks shared meaning, the substrate decides nothing about it. The law states that it cannot close, and why.
Rights β source-available, all rights reserved. This wiki and its repository are published for public inspection and to let anyone re-derive the figures. They carry no LICENSE; under default copyright, all rights are reserved. No right is given or intended to use, run, or deploy it for any purpose other than re-deriving the published figures, nor to modify or build on it β any other use requires a written licensing agreement with the authors. Β· Affine.Earth Β· zero float Β· zero shear
Each step is the reason the next one exists. Nothing here is medical advice, and no page calls any medicine safe or unsafe.
1 Β· Why an exact safety screen at all
- Cures Without the Gatekeeper β the medicine front door: six real written medicines, one screen anyone can re-run
- The library admission law β what may enter, and the 71 arms that prove it refuses. The primary artefact.
2 Β· The three libraries, which grow rather than close
- The Library of Compound Cures β exact off-target maps for the medicines the registry publishes
- The Library of Proteins β 80,080 generated sequences, novel chemical matter, graded honestly
- The Library of Material Systems β what a system is, what was measured, where the law lives. C-007 absolute: no recipes
3 Β· The maps β every place a molecule could act, counted
- The off-target atlas β every nucleic-acid medicine the registry publishes a sequence for: WHERE it can pair
- The order of the bases β WHETHER THAT BURDEN IS UNUSUAL: 472 strands ranked against sixteen rearrangements of their own bases
- Where else could this guide cut? β the whole human genome, counted
- Designed, or forced by its own bases? β every clinical CRISPR guide, with its own composition as the control
- What a public genome deposit will tell you β and four ways it will mislead a health tool first
- Study 45 β which of nine billion answers a laboratory can act on β a safety review of AlphaGenome Atlas, measured live on 1,200 real variants at two genes. The headline score separates every one. The detailed tracks do not: splice-site usage hands back 950 of every 1,000 values shared with another variant at HBB and 998 at CFTR, and the shared values pile up in the quiet band where a bench clears a variant
4 Β· One medicine at a time
- Zilganersen β the first treatment for Alexander disease, screened on the real approved sequence
- A drug an AI designed β rentosertib for pulmonary fibrosis, and exactly what our instruments reach
- CAR-T, halted β the verdict a regulator could re-derive
- N-of-1 antisense β the only safety net at a population of one
- VERVE-102 β the off-target lattice a stranger can re-derive
- PM359 β prime editing, certified before anyone is dosed
- Del-Zota β the one safety question that can be made exact
5 Β· What keeps a disease alive, and what moves it
- Study 26 β master regulator bonds β 17 tumour types, 7,673 tumours; eleven compound pairs where no single agent among 20,308 cleared any
- Study 20 β Rife frequency β light and frequency, measured rather than dismissed
- Study 37 β five molecules β 37,910 "validated discoveries", 5 distinct molecules; why per-item validation cannot see a corpus-level defect
- Are the generated cures new? β 80,080 peptides against the human proteome
- Study 16 β disease type Β· Study 17 β chemistry InChIKey Β· Study 14 β protein lattice
- No language model in this stack β what the answers here are made of: measured 2026-09-12, no cell runs a model process, opens a model port or holds an unmasked model unit, and a gate refuses their return
- Run any study in your browser β all ninety programs open on your own device, forty-nine run there, and the run tells you whether it printed the sealed bytes
- The ontology β grades, terminals, controls, and what each page may say
- Zero Float Β· Zero Shear β the method in one page
- Ask someone you trust to check this β what to hand a sceptic
- Readersβ guide Β· Program index β all 42 studies Β· White paper Β· Roadmap
- The full-grade replacement β 49 retired instruments, 4 verticals
- The exactness seam β the business case
- Build a study β Falcon walkthrough β how to add one yourself
The same move every time: take a domain where a floating-point model is the accepted instrument, compute the same quantity in exact integers, and seal the cases where the two render opposite verdicts. The subject under grading is always the instrument, never the phenomenon.
- Study 48 β the atom already has an address β silicon dimers 3.840 Γ apart, the smallest commanded scale on the board: a length carried in single precision mis-addresses its first atom at step 8,783; an address cannot
- Study 49 β the phase code never needs Ο β a phase-only modulator takes 256 codes per pixel; the code is a ratio of integers
- Study 50 β CMS raw data from the LHC, read exactly β CMS's 2011 collision bytes streamed from CERN Open Data into the Affine IDE and read in exact integers, every collision a hologram you can turn: 138 of 3,564 bunch slots carry 93,110 of 120,742 collisions, and in 3,854 the event record reads its slot exactly 3 lower than the pixel boards Β· public release
- Study 55 β IceCube: the light in the ice, hit by hit β IceCube's calibrated hits read byte for byte: 4 published files, 9,749 events, 2,289,821 hits, a census seal per file
- Study 47 β translation shear: the meaning that survives a language β LAW FROZEN Β· LIVE CLAIM, measured 2026-09-11 and again fleet-wide 2026-09-12: translation as an exact coordinate transform, charts derived in memory at every start from the raw rows of a pinned public weight file and never written down; one lattice digest on 9/9 cells, zero drift, every refusal named. The generative comparison arm is ABSENT β there is no generative translator in the stack
- Study 34 β the observer-invariant verdict β why a safety verdict needs an exact law, not a bigger computer
- Study 35 β the safety brain that forgets β deaf in 8.4 seconds, forgets across machines, disagrees with itself
- Study 36 β the language game of Fermat's Last Theorem β guess and shear, or project
- Study 40 β the number the simulation throws away β their ICO result computed as a fraction; in float the effect returns 0 at every width, and an effect returned as zero cannot be searched for
- Study 41 β fifty years of solving the wrong problem β the ordering was never about time, it was about arithmetic; 177Γ the work and 2,400Γ the wrong guesses to return the answer the machine already had
- Study 42 β The Exact Contract β 2.7M flood settlements in Int128 cents; the step exists and the rigidity does not
- Study 29 β continuous-model shear
- The lattice holds Β· Impact study β continuum dead Β· Death of continuous shear
- Fourier Phantom β Anima FNO vs 11+12+13 Β· Stellar dynamo kill shot
- QCD: freedom is dilation Β· UUM-8D vs IUT β WIN
- Peer-review bundle Β· Conjecture alignment
- We need fusion β the verdict every machine can check
- Affine Fusion Control β the local exact-integer court Β· public release
- Fusion researcher's guide
- Study 33 β the fusion control verdict court
- Every season, fifty tonnes β the biosphere-safety case
- The forcing nobody measures Β· Impact study β the SpaceX trajectory
- Study 31 β the biosphere joint ledger β LIVE on the court, 9/9 cells
- Study 28 β the wet-bulb threshold court β Act 1 sealed
- Study 32 β the taxi-out floor court
- Where humans actually yield β the fatigue curves, and where the rules already agree
- Study 30 β sovereign edge pod Β· Manufacture contracts
- The detector that flags the whole market β a manipulation geometry in exact integers, and the regulator's own indicator scored against a legitimate quoter
- Study 43 β almost every order is cancelled, and that is normal β nine sessions, three operators, two continents: 935 to 998 of every 1,000 orders that ended, ended without trading. A check that flags almost everything is a denominator, not a detector β and the stock you pick moves it further than the exchange does
- Study 44 β nine billion answers, four billion ways to say them β AlphaGenome Atlas ships 9 billion predictions in single-precision floats, which hold 4.28 billion distinct values: 52 of every 100 variants MUST share a score with another. Agreement and exhaustion look identical on the wire
- Study 38 β the loss-reserve triangle β a reserve is an exact rational; 481 of 482 verdicts identical in both arithmetics; the sixteen-billion figure comes from an unchecked premise
- Study 39 β the actuarial domain β life, pensions, multi-state and aggregation; the margin is 8 significant digits at its tightest
- Run any study in your browser β the βΆ badge beside a program name opens it in the Studio, already built and carrying its inputs, and runs it on your machine with nothing sent back
- Explore the live courts
- MCP user guide β all 51 tools Β· Deterministic no-float courts for LLMs
- Court Client β generic wasm IDE for every court Β· Court-client checkpoint
- Coding Court β the verdict IS the artifact
-
Zed β the coding agent, for developers β set Zed 1.20.2 up on
https://affine.earth/v1, no language model anywhere; what a turn does, the wire, the autonomous closure -
Zed β Minecraft comes to life β the two-person interaction, sealed: it asks, cites, clones a sibling with a value you supply, verifies by replay; the court flips
REFUSED_UNKNOWN_BUDGET β WIN - Zed β the agent that teaches the whole domain β architecture, protocols, server management and git, each answered from lines it read and instruments it ran; five closures PROVEN, and the cattle question answered with a counter the fleet did not have
- Math Court on Glama Β· Math Court user guide Β· Example app β entire court
- Quantum algorithms inventory Β· Shor witness certifier
- MCP clients (public)
- Glama connector
- Look in the UI (no visitor data)
A study appears here under the state its evidence has earned, and above under the question it answers. The two are different filings of the same work, on purpose.
β LAW FROZEN Β· DATA SEALED
- Study 06 β explosion vs earthquake Β· Study 07 β Sgr A* raw visibilities
- Study 11 β Ehrhart volume Β· Study 12 β parallel repetition Β· Study 13 β Connes rigidity
- Study 14 β protein lattice Β· Study 16 β disease type Β· Study 17 β chemistry InChIKey
- Study 18 β material STD Β· Study 19 β Go First dice
- Study 26 β master regulator bonds β 17 tumour types, every finding published
- Study 56 β MATH, read exactly β 12,500 problems read blind: 10,533 keys close, 1,967 cannot, each with the law's reason; OpenAI's marks and printed figures set beside it afterwards, labelled theirs
- Frontier models in mathematics β beside Study 56 β they train their models and sample them; Affine.Earth trains nothing: each lab's printed figure, labelled theirs with its runs and its grader, beside one sealed reading of every key and every answer
π΄ LIVE CLAIM β standing, not sealed
- Study 02 β launch ionospheric holes Β· Study 02 β regulatory alarm
- Study 09 β global convective bond Β· Study 20 β Rife frequency Β· Study 21 β stellar dynamo
- Study 22 β 2-local Hamiltonian Β· Study 23 β spin glass Β· Study 24 β N-representability Β· Study 25 β exact permanent
π CHARTER Β· OPEN β the findings, published either way
- Study 03 β flare SIDs β archive went dead Β· predictions and validations
- Study 04 β tsunami vs surge β partial seal Β· Study 05 β Forbush decreases
- Study 08 β Gaia BH1 β no corpus until DR4 Β· Study 10 β Fermi / dark matter β does not disprove DM
- Study 15 β Skala DFT shear Β· Study 27 β exact nuclear scattering
- Overview Β· First 27 days Β· Success criteria
- The science, and what history says Β· Blind spots β five stories magnitude models miss
- Historical corpus Β· Data archives β every source, exactly how to reach it
- Model shear Β· Benchmark results Β· Prediction registry
- Substrate architecture β how a shadow becomes a geometry
- Operations runbook Β· Satellite & aviation advisory