Skip to content

Study 56 MATH Read Exactly

rg78803 edited this page Oct 10, 2026 · 1 revision

Study 56 β€” MATH, read exactly: 12,500 keys and 815,632 graded attempts, one verdict for every answer the law can hold

Why this study exists

Frontier labs print mathematics figures built from sampled runs: OpenAI printed o1's AIME 2024 result for one sample, a consensus of 64 samples and a re-ranking of 1,000 (theirs). Run it again and the answers can change: on the 30 AIME 2025 problems, 465 of 870 model–problem pairs returned more than one answer across four runs (counted from MathArena's published rows, theirs). They train their models and sample them.

Affine.Earth trains nothing and samples nothing. One law of exact whole numbers reads the bytes: 12 Swift files, 0 imports, no float. Study 56 puts that law to MATH, a public benchmark of 12,500 competition problems, and to OpenAI's prm800k release, 815,632 scored attempts at 500 of them. It answers what a sample cannot: which answers equal their key exactly. Whether the nine Affine.Earth cells read them the same way is not known: no cell runs Study 56 yet.

MATH and PRM800K were built to train models. MATH's own train folder holds 7,500 problems. PRM800K's 800,000 step labels trained OpenAI's process reward model, and 4,500 MATH test problems were moved into its training set (theirs, their README). Affine.Earth trained on none of it: it read every one of the 12,500 problems, train and test alike, as a claim to decide.

Answers change even at temperature 0, where nothing is sampled. Thinking Machines Lab names the cause: kernels whose order of addition changes with batch size (theirs). Yuan and others, under greedy decoding: floating-point addition is not associative at limited precision (theirs; both in Frontier models in mathematics). The substrate adds whole numbers: split across 1, 2 or 5 agents and merged in forward, reverse or interleaved order, the Study 56 goal seals one set of bytes, and all 815,632 attempts reproduce their sealed outcome, the 20 not read among them.

What it proved

  1. One verdict for every answer the law can hold. All 12,500 keys and 815,612 of OpenAI's 815,632 attempts have one verdict each. A key closes or cannot close, with its reason; an attempt is equal or unequal to the key, or cannot close, with its reason. The other 20 attempts are longer than the law's room of 10,000 digits; they are named and not read. The law was designed and sealed without the marks, and committed before the marks were read again (an earlier reading had read them once, on 2026-10-07). Run again after the seal, all 815,632 attempts reproduced their sealed outcome. Run twice on one Mac, the law wrote byte-identical files: 26 of 26 for the seal, 13 of 13 for the comparison. The Affine IDE's agent swarm seals one set of bytes with 1, 2 or 5 agents.
  2. 1,967 of MATH's 12,500 answer keys are not a whole number or a fraction of whole numbers, the only values the law holds, and the law names why for each. 595 hold a root Β· 383 hold letters Β· 265 hold pi Β· 256 are written with a decimal point Β· 165 are words Β· 99 hold the imaginary unit Β· 91 hold infinity Β· 71 hold a container of containers Β· 19 use a form outside the law's grammar Β· 11 are unions of sets Β· 6 hold a trigonometric function Β· 4 hold no answer Β· 1 holds a logarithm Β· 1 has a comma between digit groups. Of OpenAI's 500 problems, 83: 24 roots, 20 letters, 10 pi, 9 decimals, 7 words, 6 the imaginary unit, 4 infinity, 2 unions of sets, 1 container of containers.
  3. The law decides none of the 134,995 attempts on those 83 problems, and their mark reads correct on 73,200 of them (theirs). 47,027 of the 134,995 copy the key byte for byte, and a copy of a text the law cannot read is still not read to a value. How many of the 73,200 are such copies is not known.
  4. Every time the law reads an answer unequal to the key, their mark also reads false: 243,578 of 243,578, and true on none.
  5. 4,825 answers equal to the key are marked false (theirs). Where the law reads an answer equal to the key, their mark is_correct reads false on 4,825 attempts in 49 problems (theirs). 1,793 attempts answer 864 and are marked false; their copy of the key reads 864\mbox{inches}^2, and MATH's key is 864. \frac{2187}{5625} is marked false; it reads to 243/625, and the key is 243/625.
  6. Their grader normalises answers with sympy, with no sympy version pinned and no timeout (their grader.py, theirs). It marks two answers alike when their normalised strings match or when sympy simplifies their difference to zero. No file of theirs states how is_correct was produced.
  7. Side by side, never subtracted. Among OpenAI's own attempts at each problem, the answer most of them hold, counted by exact value, equals the key on 293 of the 417 problems whose key the law reads to a value. The other 83 have no key value to equal. Theirs (Let's Verify Step by Step, Figure 3, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it counts problems; their eval.py prints their two reward-model picks as means over 400 trials, not as counts. Their printed figures stand beside the law's as published; the two cannot be subtracted.

Theirs is a sample. This reading came out byte-identical every time it ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.

For every reader

MATH is a public set of 12,500 competition problems in seven subjects, from algebra to precalculus, each given a level from 1 to 5, except two Geometry problems whose level is written "?". Every problem comes with a worked solution, and its answer β€” the key β€” is the last \boxed text of the solution. OpenAI's later release, prm800k, is built on it. It holds 815,632 attempts written by a model at 500 of the test problems, each with a mark called is_correct, and four files of model solutions whose steps people rated one by one. Study 56 reads all of it with Affine's own Swift and answers every problem and every attempt only as it relates to the Affine substrate. A text closes when the law reads it to an Affine value: an integer, an exact fraction of two integers, or a base-N integer. Any other text cannot close, and the law says why in a sentence of its own β€” a root is written; pi is not a fraction of integers; letters make an expression, not a value; a decimal point is a float's spelling. A text is not valid just because it is written. The law does not check that a key is right: an answer equal to a key holds the key's value, whatever the key is.

The law did all of this blind: its reader passes over the marks by name, and its results were sealed and committed before the marks were read again. Only then were the marks set beside it, labelled theirs. Every count is sealed with SHA-256, and the seals came out byte-identical every time the study ran: twice on one Mac, and with 1, 2 or 5 agents. No other machine has run it yet.

What the labs that build frontier models print about their own mathematics, and what independent evaluators counted run by run, is set beside this reading, each figure labelled theirs with its runs and its grader, in Frontier models in mathematics β€” a share over sampled runs, and one sealed reading of every key and every answer.

Status: SEALED BLIND, THEN SET SIDE BY SIDE β€” 2026-10-08 β€” 8 published files opened by the law on one Mac: all 12,500 MATH problems, test and train, and every attempt that OpenAI's prm800k release holds for them. All of it sealed and committed before the marks were read again; OpenAI's marks and published figures, labelled theirs, set beside it afterwards. The commits are in Β§10. In Affine IDE 0.2.6.6 Ξ² the stage, the Mathematics folder and the swarm read this sealed record (Β§9).
Program index Β· Study board Β· Study 55 β€” IceCube: the light in the ice Β· Study 50 β€” CMS raw data from the LHC


1 Β· The files

Eight files. Each is pinned in the law by its byte count and its sha256, and the law checks both before it reads a record: the bytes first, then the digest. All eight pins read VERIFIED.

file what it holds bytes records distinct problems sha256, first 16
MATH.tar MATH: one JSON file per problem, with its text, subject, level and worked solution 20,327,936 12,500 12,500 0fbe4fad0df66942
math_splits/test.jsonl prm800k (theirs): the 500 test problems and their keys 446,564 500 500 35dc41080a368085
math_splits/train.jsonl prm800k (theirs): the training problems and their keys 10,896,985 12,000 11,999 90d96daeac3fe343
scored-test-samples.jsonl the scored attempts at the 500 test problems, with is_correct (theirs) 2,159,374,091 815,632 500 46a01478dceefe81
phase1_test.jsonl phase one: solutions whose steps people rated (theirs) 829,105 106 101 f4b3bc5b095e45c8
phase1_train.jsonl phase one, training 7,900,236 949 903 e9da6a73f827ffb9
phase2_test.jsonl phase two: model solutions, the model's own steps and answer, step ratings (theirs) 12,240,719 2,762 458 6b172efa884ac834
phase2_train.jsonl phase two, training 456,135,365 97,782 10,828 1110237feeb51d1b
  • MATH.tar β€” the Wayback capture 20231017015604 of people.eecs.berkeley.edu/~hendrycks/MATH.tar: 12,518 members, 17 directories and 12,501 files; 12,500 problems, test 5,000 and train 7,500. MIT is the hendrycks/math repository licence (GitHub API); the tar's own README states none.
  • math_splits and the phase files β€” the openai/prm800k repository at commit 7ecc794703b2877f (theirs), MIT licence; LICENSE, 1,062 bytes, in the repository tree.
  • The scored file β€” https://openaipublic.blob.core.windows.net/process-supervision/scored-test-samples.jsonl. No licence is stated on the blob or in its HEAD response. It is streamed, never stored: 0 files of exactly 2,159,374,091 bytes were found after the runs. It holds between 152 and 1,860 attempts per problem, as the law counted them; neither of OpenAI's sources states the smallest.
  • Where the Affine IDE fetches two of them β€” the IDE fetches MATH.tar and phase2_test.jsonl from https://affine.earth/language-game/studies/math56/ and keeps each only when its byte count and then its sha256 equal the pin. The scored file is never stored and never served.

MATH.tar is the spine, one row per problem; every other source joins to it. No phase-file join is ambiguous. 3,854 phase2_train records name 10 problems that join no math_splits record; they are decided against their own source's key and credited to no row.


2 Β· How the substrate answers

An Affine value is an integer of any width, carried as a decimal-string integer by the mesh's own arithmetic (ArbitraryStringMath.swift, vendored byte-identical from its one home, 6,379 bytes, commit 73c986c9e); an exact fraction of two such integers in lowest terms; or a base-N integer read to its integer. The law is 12 files with 0 import lines, no float type and no float literal. It writes no integer longer than 10,000 decimal digits; a text that would pass that room is not read (NOT_READ_ROOM), a state that decides nothing and says nothing of the text's meaning.

No decimals. A decimal point is how a float is written, so the law refuses the text and reads no value from it. The law's sentence:

WRITTEN_WITH_A_DECIMAL a decimal point is a float's spelling; Affine refuses the text and reads no value from it

The substrate holds integers and fractions of integers, so that every cell reads the same value from the same text. The law does not turn a decimal into a fraction. The rule is decided first, at every level: key, answer, option, container member, the right side of an assignment, a side of a step. The code that read decimals into fractions is deleted. 256 MATH keys are written with a decimal (test 99, train 157); 9 of them are among the 500.

Invent nothing. Everything this law adds to the earlier reading (Β§8) is a composition of a law the substrate already has:

what it reads counts
2a wide integers, by the mesh's decimal-string arithmetic 6 MATH keys; answers scored 860, phase2_test 4, phase2_train 108
2b identity: byte equality after the earlier reading's own unwrapping (Β§8) step pairs scored 2,710, phase2_test 18, phase2_train 376, MATH solutions 4
2c containers β€” tuples, sets, intervals β€” member by member 523 MATH keys; answers scored 29,126
2d x = value, read as its value 32 MATH keys; answers scored 39
2e a choice letter, read as that option of the problem's own text 788 option texts resolved: 321 to a value, 467 cannot close; 68 letter keys, 33 resolved, none to a value
2f a word, by its exact bytes inside \text{…} 1 MATH key

Not added: no number fields, no radical canonical forms, no polynomial algebra, no transcendental rings, no tolerance, no refinement and no bound.

Cannot close β€” why a text lacks shared meaning. When the law cannot read a text to an Affine value, it names the reason in its own words. The first reason in this order is the one recorded.

reason the law's own sentence MATH keys keys of the 500
NO_ANSWER there is no text to read: the field is absent or null, empty, the string None, or holds no braced boxed answer 4 0
WRITTEN_WITH_A_DECIMAL a decimal point is a float's spelling; Affine refuses the text and reads no value from it 256 9
PROSE_OR_WORD words carry no number 165 7
HOLDS_LETTERS letters standing for unknowns make an expression, not a value 383 20
HOLDS_IMAGINARY_UNIT i squared is minus one, and no integer and no fraction of integers squares to minus one 99 6
HOLDS_TRIG a trigonometric or hyperbolic function is written; the law does not evaluate functions 6 0
HOLDS_LOG a logarithm is written; the law does not evaluate functions 1 0
HOLDS_PI pi is not an integer and not a fraction of integers 265 10
HOLDS_RADICAL a root is written; the law reads integers and fractions joined by plus, minus, times, divide, integer powers and factorials, and does not take roots 595 24
HOLDS_INFINITY infinity is not an integer and not a fraction of integers 91 4
UNION_OF_SETS a union, intersection or difference of sets is set algebra; the law compares one container of values at a time 11 2
CONTAINER_OF_CONTAINERS a member that is itself a container (a matrix row, a list of pairs) is not a value 71 1
COMMA_BETWEEN_DIGIT_GROUPS a comma between digit groups, in a text that is not wholly one integer written in comma groups, is written both as a thousands separator and as a list separator, and the text does not say which 1 0
OUTSIDE_THE_GRAMMAR the text uses a symbol or a form the law does not read β€” among them a floor, a ceiling, a binomial, a modulus, an absolute value, a ratio, plus-or-minus, dots, a function name, an environment, a text-family command other than text, an unclosed bracket, a list on one side of an equals sign 19 0
all 1,967 83

Five more reasons fall on no MATH key: digits written against the letter e, a float's exponent spelling (WRITTEN_WITH_A_FLOAT_EXPONENT); a unit written in words or a degree, percent or dollar mark, on a step side only (DECORATED); an equation or a comparison, a statement about values (RELATION_NOT_A_VALUE); the letter e (HOLDS_E); and a division by zero (DIVIDES_BY_ZERO). Two pair reasons fall where both sides read to values and still share no form: a container against a single value, or two kinds of container (KINDS_DIFFER); and two unbracketed lists in another order, since an unbracketed list does not say whether order counts (ORDER_NOT_WRITTEN).

Three levels. A problem: its key alone decides. An attempt: if the key cannot close, every attempt on it cannot close, with the key's reason; if the key closes, two values are equal or unequal; otherwise the attempt cannot close, on the answer's side, or because the two sides read to values that share no form. A step: there is no key; two values decide by value, and two byte-identical sides that lack a value and hold no decimal close by identity.

Identical spellings. An attempt that copies a key the law cannot read is not read to a value by being copied: being written does not make a text valid, and a text the law cannot read to a value is not read to one by being repeated. The identity is printed beside each such attempt and never counted as closed: scored 47,027, phase2_test 66, phase2_train 506 (505 on problems that join a MATH row, and 1 on a problem that joins none).


3 Β· Blind first, then compare

The reader compares every field name, byte for byte, with the names and fragments it withholds, and never decodes a withheld value:

withheld by name: is_correct rating corrected_rating finish_reason flagged prm_score orm_score pre_generated_verifier_score rating_probs chosen_completion human_completion label
withheld by any field containing: correct rating score grade verdict finish flag label_

Flip every withheld member and the seals do not move: BLIND_HOLDS, 5 of 5 sources. A review read back every recorded read and command of the runs: none returned a line from a withheld path. The law was designed and run without the marks; they had been read once before, on 2026-10-07, by the earlier reading (Β§8).

On 2026-10-08, UTC: at 22:03:15 every blind output and law file was verified against the git objects of the blind seal's commit (Β§10), 63 checks; at 22:35:07 the phase files' rating and finish_reason (theirs) were first read; at 22:35:14 is_correct (theirs) was first read. prm_score, orm_score, rating_probs and pre_generated_verifier_score were never read by any run: they are another model's decimal scores.


4 Β· All 12,500 problems

problems rows key closes key cannot close
all MATH problems 12,500 10,533 1,967
test split 5,000 4,215 785
train split 7,500 6,318 1,182
the 500 scored test problems 500 417 83

The 10,533 closing keys read to: an integer 8,252 Β· a fraction 1,657 Β· a container 523 Β· a base-N integer 62 Β· an assignment 32 Β· a wide integer 6 Β· a word 1.

subject (MATH's own) problems key closes key cannot close
Algebra 2,931 2,580 351
Counting & Probability 1,245 1,205 40
Geometry 1,349 976 373
Intermediate Algebra 2,198 1,689 509
Number Theory 1,409 1,369 40
Prealgebra 2,076 1,854 222
Precalculus 1,292 860 432

MATH's own worked solutions. The last \boxed text of a solution is the key, by MATH's own definition, so read against the key it is the key read against itself: 10,533 equal, 0 unequal, every other row cannot close with the key's reason. The steps of the worked solutions hold 63,448 = pairs: 11,146 equal, 4 by identity, 67 unequal, 52,226 cannot close, 5 past the room. 18,194 equals signs sit outside every span the step law opens; they are counted and not read.

Every attempt in every source.

source records equal unequal cannot close not read
scored attempts, the 500 test problems 815,632 362,184 243,578 209,850 20 (room)
phase2_test, the model's own final answer 2,762 558 1,415 789 0
phase2_train, the model's own final answer 97,782 7,061 60,305 30,409 7 (room)
phase1_test, the final answer 106 β€” β€” β€” 106 (withheld by the blind seal; its rated steps read after it, Β§7)
phase1_train, the final answer 949 β€” β€” β€” 949 (withheld by the blind seal; its rated steps read after it, Β§7)

Each row adds to its records. Credited to the 12,500 problems, 912,322 attempts were read (the census of every problem counts them as attempts read, ATTEMPTS_READ): 675,101 closed, 237,194 cannot close, and 20 scored and 7 phase2_train answers past the room. The other 3,854 phase2_train records sit on 10 problems that join no MATH row; they were read too, against their source's own key, and all 3,854 cannot close: 3,853 on 9 problems whose source key holds no answer (NO_ANSWER), and 1 on 1 problem whose source key is a container of containers (CONTAINER_OF_CONTAINERS). They are inside phase2_train's 30,409. Of the 3,885 phase2_train attempts whose key holds no answer, those 3,853 are the unjoined records and the other 32 are credited to problem rows. The final answers of the 106 phase1_test and 949 phase1_train attempts are not read: the blind seal withheld them with the label tree. After the seal, the comparison joined all 106 and all 949 to their sealed records, with 0 left out of the join, and read their rated steps with the same step law: 3,119 pairs in the rated completions and 135 in the steps the labellers wrote in phase1_test, 31,620 and 1,316 in phase1_train (Β§7). The blind seal counts 1,082 attempts not read (ATTEMPTS_NOT_READ): the 1,055 phase-one attempts and the same 27 answers past the room. 1,175 problems have no attempt in any source.


5 Β· The 500 and their 815,632 attempts

All 815,632 attempts were joined to their sealed records and run again after the seal: all 815,632 reproduced their sealed outcome. Their mark is_correct (theirs) reads true on 433,067 and false on 382,565.

the law's outcome is_correct true (theirs) is_correct false (theirs) attempts
equal 357,359 4,825 362,184
unequal 0 243,578 243,578
cannot close Β· the key lacks shared meaning 73,200 61,795 134,995
cannot close Β· the answer lacks shared meaning 2,507 67,627 70,134
cannot close Β· both sides read to values: a container against a single value, or two kinds of container (KINDS_DIFFER) 1 4,143 4,144
cannot close Β· both sides read to values: two lists in another order (ORDER_NOT_WRITTEN) 0 577 577
not read (NOT_READ_ROOM) Β· an answer past 10,000 digits 0 20 20
all 433,067 382,565 815,632

Unequal, with their mark true: 0. Every attempt the law reads as unequal to the key, their mark also reads false.

Equal, with their mark false: 4,825 attempts, across 49 problems. In each, the answer and the MATH key hold one exact value. This page prints both columns and does not rule between them. Examples, as the law read them:

  • test/algebra/1072 β€” \frac{2187}{5625}, \frac{54675}{140625} and \frac{25 \cdot 3^6}{3 \cdot 5^6} each read to 243/625; the MATH key is 243/625; their mark, false (theirs).
  • test/algebra/722 β€” 485149/49 and 99^2+99+1 each read to 9,901, against the key 9,901.
  • test/geometry/473 β€” 1,793 attempts answer 864. The scored file's own copy of the key (theirs) is 864\mbox{inches}^2; the MATH key's value is 864.
  • test/prealgebra/1114 β€” 550 attempts answer 15; the key copy (theirs) is 15\mbox{cm}^2; the MATH key's value is 15.

The 83 problems whose key cannot close. No attempt on them can close, whatever it writes, because the key it would be decided against has no Affine value.

why the key lacks shared meaning problems attempts is_correct true (theirs) is_correct false (theirs)
HOLDS_RADICAL β€” a root is written 24 37,862 19,015 18,847
HOLDS_LETTERS β€” an expression, not a value 20 31,834 15,423 16,411
HOLDS_PI β€” pi is not a fraction of integers 10 16,155 5,828 10,327
WRITTEN_WITH_A_DECIMAL β€” a float's spelling 9 16,674 12,781 3,893
PROSE_OR_WORD β€” words carry no number 7 12,722 9,942 2,780
HOLDS_IMAGINARY_UNIT β€” no fraction of integers squares to minus one 6 9,556 5,352 4,204
HOLDS_INFINITY β€” infinity is not a fraction of integers 4 6,326 4,107 2,219
UNION_OF_SETS β€” set algebra, not one container 2 3,561 752 2,809
CONTAINER_OF_CONTAINERS β€” a member that is a container 1 305 0 305
all 83 134,995 73,200 61,795

6 Β· What OpenAI published, set beside the substrate

What they built (theirs): the paper "Let's Verify Step by Step" (OpenAI), version one. 1,860 solutions were generated per test problem (Figure 3 caption, page 7), and one is chosen by a process reward model, an outcome reward model or majority voting. A process reward model scores a solution as the product of its per-step probabilities. 4,500 MATH test problems went into training, and evaluation uses the remaining 500, selected uniformly at random (Appendix C). prm800k holds 800,000 step-level correctness labels (README L5).

How they graded (theirs): the solution ranked highest is graded automatically on its final answer, and the fraction correct is reported; the paper gives no implementation of the grader. In the repository, grader.py (8,101 bytes) counts two ways to be correct β€” the two strings normalise to the same string, or sympy simplifies their difference to zero (L236-240) β€” and eval.py (3,153 bytes) selects by the stored score and counts the stored is_correct; no grading happens in eval.py, and majority voting is not in it. No repository file states how is_correct was produced. The substrate did not run their grader; is_correct is printed exactly as stored.

The substrate's majority, by exact Affine value, over OpenAI's own attempts. It solves no problem: it counts the answers OpenAI's model wrote at each problem. Only attempts whose answer reads to an Affine value vote β€” 618,869 of 815,632. One value, one vote class. A tie is named and counted, and left a tie. The majority value is decided against the MATH key by the law's own decide.

the substrate's majority, by exact Affine value problems of 500
the Affine majority value equals the MATH key 293 (293 of the 417 whose key can close)
it is unequal to the key 119
a tie: two values hold the top count 2
the majority value and the key read to values that share no form (a container against a single value, or two kinds of container: 2; two lists in another order: 1) 3
the key cannot close 83

Their figures, as published (theirs; Let's Verify Step by Step, Figure 3 inset table, page 7, best-of-1,860): majority voting is printed as a share equal to 348/500, and the paper does not say whether it is a count of problems or a rounded mean. Their eval.py, which makes the two reward-model selections, prints each as a mean over 400 trials, not as a count of problems, so neither is written here as k of 500.

The law's 293 counts problems among the 417 whose key closes; their share, 348/500, is printed over all 500 and is not stated to be a count. The two stand side by side and are not subtracted: the law decides no majority on 83 of the 500, so these figures cannot show where any difference lies.

The two ties: intermediate_algebra/1791 (key βˆ’3/8; the values 0 and 3/4 each hold 104 attempts) and precalculus/989 (key 12; the values 0 and 2 each hold 6 attempts).

Not computed, and why. Their two reward-model selections choose one attempt per problem by a model's decimal score (prm_score, orm_score); a decimal score is a float, so those fields were not read. Their step-score strategies (rating_probs) and the verifier score are decimal scores, not read. Their grader's normalisation is theirs, not run. Their majority vote's grouping is not stated in their paper. Anyone may read those scores on their own machine; no figure on this page comes from a decimal score.

What their sources state, recorded without comment (theirs): the paper does not state how majority voting decides two answers are the same, how it breaks a tie, or which solution it returns; their README states up to 1,860 scored samples per test problem, and the problem with the fewest holds 152 attempts as the law counted them.


7 Β· The steps and the ratings

step pairs read pairs equal by identity unequal cannot close past the room
MATH worked solutions 63,448 11,146 4 67 52,226 5
scored attempts 5,396,497 728,819 2,710 60,672 4,602,962 1,334
phase2_test, the model's own steps 18,476 2,248 18 300 15,910 0
phase2_train, the model's own steps 750,346 93,415 376 9,228 647,279 48
phase1_test, the rated completions 3,119 275 1 45 2,798 0
phase1_test, the steps the labellers wrote 135 12 1 0 122 0
phase1_train, the rated completions 31,620 3,132 74 925 27,489 0
phase1_train, the steps the labellers wrote 1,316 113 0 6 1,197 0

The phase-one pairs were read by the comparison after the seal, with the same step law; they are not part of the seal. The law reads no rating. Every step pair the phase files hold is printed beside its rating (theirs) in the comparison's censuses; the full tables are in the report and in evidence/study-56-v3-compare-20261008/.


8 Β· What changed: an earlier reading that turned decimals into fractions

An earlier reading turned a decimal into a fraction; this law refuses it: a decimal point is how a float is written, so the law refuses the text and reads no value from it. Only the decimal rule moves a value: no earlier equal became unequal and no earlier unequal became equal. In the scored file, 14,289 earlier equal and 21,986 earlier unequal attempts now cannot close, each for a decimal, and 12,487 equal and 10,176 unequal are newly decided out of the earlier reading's not-decided attempts. Of the MATH keys, 247 that the earlier reading read to a value now cannot close, each written with a decimal, and 562 that it read no value from now close. The earlier reading is history, not deleted: commit c56e3dbba holds it, its comparison is at 20af6f29e, and anyone may run it on their own machine. Its swarm goal, also history, sealed 38c8079f18ed23d3 with 1, 2, 4 and 5 agents (af49986a8).


9 Β· In the Affine IDE β€” 0.2.6.6 Ξ²

In Affine IDE 0.2.6.6 Ξ² the stage, the Mathematics folder and the swarm read the sealed record of Β§3 to Β§6.

  • The Study 56 stage: three holograms β€” Key, the 12,500 MATH problems as a lattice by subject, level and split, the 500 test problems ringed; Samples, the 500 test problems as columns of their 815,632 attempts, the sealed outcome beside is_correct (theirs) from chapter 3 on; Steps, one phase-two record as a tower of its steps and its = pairs. Six chapters in the chat, every figure read from the sealed record: 10,533 keys close, 1,967 cannot, and the 500 test problems are ringed.
  • The Mathematics folder: seven type folders in MATH's own order β€” Algebra 2,931 Β· Counting & Probability 1,245 Β· Geometry 1,349 Β· Intermediate Algebra 2,198 Β· Number Theory 1,409 Β· Prealgebra 2,076 Β· Precalculus 1,292 β€” each by level, 20 problems to a page. Selecting a problem opens Study 56 on it: its cube ringed on the Key picture, its text, its key, MATH's own answer (theirs), and how the substrate answered it: it closes, with its value, or it cannot close, with its reason in the law's own sentence.
  • The internal swarm: deterministic substrate workers running the substrate's own laws, never a language model, each owning a space and a goal in its own sandbox, sized to the machine as measured. The Study 56 goal runs this law and seals one set of bytes, c6949876a8639991, with 1, 2 or 5 agents; its census and index are byte-equal to the committed record.

10 Β· The seals

commit when what
c56e3dbba 2026-10-07 20:32:30 -0400 the earlier reading's blind seal β€” history (Β§8)
20af6f29e 2026-10-07 22:00:55 -0400 the earlier reading's comparison β€” history (Β§8)
f2d890e48 2026-10-08 18:01:19 -0400 the blind seal
3c25e8cf9 2026-10-08 21:30:24 -0400 the comparison
8eb8f4ca9 2026-10-09 10:44:53 -0400 the Affine IDE reads the sealed record: the stage, the Mathematics folder and the swarm

The census seals, first 16 hex digits, at f2d890e48:

source census seal
MATH.tar f14929f70bd49d42
math_splits_test 62cae89ea1c30246
math_splits_train f98a49208cc1e0b7
scored 03b5a3a6ff1ba4f1
phase1_test 964e7d1395b06617
phase1_train cc1c9367f11f70a6
phase2_test e02306d6d9fbb45a
phase2_train aed6415d3ff09154
the 12,500 problems 69e4a013c2fb79a3

The comparison's scored census at 3c25e8cf9: 1f482961312e517d. The swarm's Study 56 goal at 8eb8f4ca9: c6949876a8639991. Repeat runs: 26 of 26 blind outputs and 13 of 13 comparison outputs byte-identical between two runs, each from its own stream of the scored file. Every number on this page can be read from git at its commit, for example git show f2d890e48:evidence/study-56-v3-blind-seal-20261008/census/scored.v3.census.txt.


11 Β· Not decided, not known, open

item status
the 1,967 keys that cannot close, and every attempt on them no Affine value; each reason is named in Β§2
answers and step pairs past the room (answers 20 scored and 7 phase2_train; step pairs 1,334 scored, 48 phase2_train, 5 MATH solutions) not read; decides nothing
the 106 and 949 phase-one attempts their final answers are not read: the blind seal withheld them; after the seal the comparison read their rated steps, 3,119 and 135 pairs in phase1_test and 31,620 and 1,316 in phase1_train
their reward-model selections, step-score strategies and verifier score not computed: decimal scores
how is_correct was produced not stated in any repository file (theirs)
whether their majority-voting figure is a whole count or a rounded mean not stated (theirs)
how their majority voting matches answers or breaks ties not stated (theirs)
how many of the 73,200 attempts marked correct on the 83 keys that cannot close copy the key byte for byte not known: not counted
whether the nine Affine.Earth cells read Study 56 the same way not known: no cell runs Study 56 yet

Edge. Where a key lacks shared meaning, the substrate decides nothing about it. The law states that it cannot close, and why.

🧬 CURES β€” read in this order

Each step is the reason the next one exists. Nothing here is medical advice, and no page calls any medicine safe or unsafe.

1 Β· Why an exact safety screen at all

2 Β· The three libraries, which grow rather than close

3 Β· The maps β€” every place a molecule could act, counted

4 Β· One medicine at a time

  • Zilganersen β€” the first treatment for Alexander disease, screened on the real approved sequence
  • A drug an AI designed β€” rentosertib for pulmonary fibrosis, and exactly what our instruments reach
  • CAR-T, halted β€” the verdict a regulator could re-derive
  • N-of-1 antisense β€” the only safety net at a population of one
  • VERVE-102 β€” the off-target lattice a stranger can re-derive
  • PM359 β€” prime editing, certified before anyone is dosed
  • Del-Zota β€” the one safety question that can be made exact

5 Β· What keeps a disease alive, and what moves it

βš–οΈ How to read any page here

πŸ”¬ The method β€” exact against float, domain by domain

The same move every time: take a domain where a floating-point model is the accepted instrument, compute the same quantity in exact integers, and seal the cases where the two render opposite verdicts. The subject under grading is always the instrument, never the phenomenon.

⚑ Fusion β€” the energy case

🌍 The planet, and the sky

πŸ› Markets, money and risk

βš›οΈ Run a court yourself

πŸ“’ Program ledger β€” every study by lifecycle

A study appears here under the state its evidence has earned, and above under the question it answers. The two are different filings of the same work, on purpose.

βœ… LAW FROZEN Β· DATA SEALED

πŸ”΄ LIVE CLAIM β€” standing, not sealed

🌊 CHARTER Β· OPEN β€” the findings, published either way

β˜€οΈπŸŒ‘ Eclipse 2026 β€” Study 01, DATA SEALED

πŸ”¬ Discoveries and flows

Clone this wiki locally