Repository navigation
What A Gene Length Already Decides
Two instruments that answer before you ask them. The subject under grading is a mapping rule and a constraint score — never a gene, never a person, and never anyone's DNA.
Status: VERIFIED — 2026-09-08. 19,704 genes, every catchment base counted with no sampling, 16 self-test arms in both directions, 0 failed, 17 published figures recomputed and 0 disagreeing. Both results are re-derivable from a 1.5 MB public corpus in under three seconds.
Two of the most-used moves in human genetics are:
- map a variant to its nearest gene, and
- ask whether the resulting gene set is unusually constrained, usually with gnomAD's LOEUF.
Both are reasonable. Both are also instruments with a built-in answer, and the size of that built-in answer is measurable exactly. This page measures it. Nothing here says anyone's result is wrong; it says what a matched control has to hold constant before the result means what it looks like.
A long gene has a larger catchment. That is not a hypothesis, it is geometry: every base inside the gene maps to it, and so does every base in the flanks up to the point where a neighbour gets closer. So a variant with no biology attached at all, dropped uniformly at random on the genome, lands disproportionately on long genes.
Measured over the whole gene model, counting every base of every catchment — no variant list, no sampling, no binning:
| mean length of the assigned gene | against the average gene | |
|---|---|---|
| the average gene, unweighted | 66,591 bp | 1000/1000 |
| nearest by INTERVAL — the field's own rule, ≤ 100 kb | 243,881 bp | 3662/1000 |
| nearest by MIDPOINT, ≤ 100 kb | 114,478 bp | 1719/1000 |
Assignable bases: 2,070,253,229 under the interval rule (of which 1,261,815,650 lie inside a gene) and 1,628,025,115 under the midpoint rule. The two domains differ, and the reason they differ is the finding: a long gene reaches further under the interval rule.
A variant with nothing biological about it maps to a gene 3.66× the length of the average gene. The midpoint rule carries less than half that bias — 1.72× — and nobody chose it for that reason; the interval rule is the convention.
This matters because gene sets that are enriched for long genes are not rare or exotic. Brain- expressed genes are long. So an analysis that maps variants to nearest genes and then asks "is this set brain-enriched?" is asking a question its own mapping has already half-answered.
LOEUF is the upper bound of the observed/expected loss-of-function ratio. Its denominator is the expected LoF count, and the expected count is a function of how much coding sequence a gene has. So the score should fall with length even where nothing about constraint differs — and it does, monotonically, across every decile.
Median LOEUF × 10⁶ over the 19,197 genes carrying both a LOEUF and an expected-LoF count:
| decile | by gene length | by expected LoF | by CDS length |
|---|---|---|---|
| D1 | 1626000 | 1799000 | 1686000 |
| D2 | 1215000 | 1454000 | 1268000 |
| D3 | 1096000 | 1146000 | 1129000 |
| D4 | 995000 | 1013000 | 1140000 |
| D5 | 974000 | 921000 | 919000 |
| D6 | 872000 | 867000 | 867000 |
| D7 | 812000 | 765000 | 794000 |
| D8 | 720000 | 648000 | 704000 |
| D9 | 572000 | 551000 | 558000 |
| D10 | 484000 | 412000 | 412000 |
- by gene length: 1626000 → 484000, a 3359/1000 swing, strictly monotone
- by expected LoF: 1799000 → 412000, a 4366/1000 swing, strictly monotone
- by CDS length: 4092/1000, and not monotone — D4 sits above D3. Reported as it came out. A result that were fabricated, or fitted, would be monotone everywhere.
The obvious objection is that gene length and expected-LoF count are the same variable wearing two coats. So each is held inside the other's decile and the residual measured — in both directions, because running it one way only would have answered a question the page did not ask.
| held constant | moving variable | deciles where it still moves LOEUF | mean shift |
|---|---|---|---|
| expected-LoF decile | gene length | 10 of 10 | -124,600 |
| gene-length decile | expected LoF | 10 of 10 | -391,700 |
The denominator dominates by 3143/1000 — and both survive conditioning. So the honest statement is not "LOEUF is only its denominator": it is that coding length accounts for about three times as much of the LOEUF gradient as genomic length does, and neither is zero. Both halves of that are measured; neither is asserted.
Nothing here invalidates a published enrichment result. It names what the control has to hold:
- Match on gene-length decile, not only on expression level or detection rate. A background set matched on expression but not on length is matched on the wrong variable for this instrument.
- Or map by midpoint. It carries less than half the length bias, and it costs nothing.
- Match on expected-LoF count before reading LOEUF, because that is the score's own denominator.
All three are cheap. None requires new data. Each is a line of code.
We call both measurements exact and complete. Every base of every catchment on every chromosome is counted; every gene carrying the fields is scored; nothing is sampled, binned or estimated. The figures are reproducible in under three seconds from a 1.5 MB file.
We refuse to call any published enrichment result wrong. This page measures an instrument's built-in gradient. Whether a particular result survives matching on that gradient is a question for that result's own data, and we have not run it.
We refuse to say anything about any gene, any variant or any person. No individual data of any kind is in this corpus. Gene lengths and constraint scores are aggregate properties of a reference genome, published by their consortia.
We refuse to convert either result into a claim about biology. A length-weighted lottery is a statement about a mapping rule. That brain-expressed genes are long is a fact about the genome, not a verdict on any neuroscience.
Where a bench should point. Re-run one published nearest-gene enrichment with gene-length decile added to the matching, and publish both numbers side by side. That is a day of work and it would tell the field how much of its tissue-localisation literature is carrying this gradient.
git clone https://github.com/gaiaftcl-sudo/uum8dSolarResearch.git
cd uum8dSolarResearch
xcrun swiftc -O -swift-version 5 reproduce/nearest-gene-length-lottery-exact.swift -o /tmp/gl
/tmp/glSixteen self-test arms run before any corpus byte is read, in both directions — the decimal parser
must accept scientific notation and reject NA, text and a double dot; the rational comparator
must order 1/2 before 2/3, must not order 2/3 before 1/2, must leave equal rationals in
different denominators unordered, and must separate 1/3 from 33333333333333333/10¹⁷, which a
Double calls equal. One arm is a mutant that weights every gene equally and is required not
to reproduce the catchment mean — which is what shows the weighting, rather than the arithmetic, is
where 3.66× comes from.
The program takes no argument and no stdin, and prints its published reference figures as its very first action, before any file is opened, so every refusal path prints them. The corpus root is discovered by walking outward from the binary and the working directory; the seal is identical from any directory.
How you know this run computed rather than quoted. The harness runs every program with no
argument and stdin from /dev/null, and every published figure must appear in the output — so a
program that refuses still prints its figures, and grepping a transcript for a figure cannot tell
a quoted number from a computed one. This program therefore prints exactly one terminal line, last,
on every path it can take: RUN_TERMINAL COMPLETE, or RUN_TERMINAL REFUSED <reason>. The quoted
figures sit inside a fenced block a reader or a grader can exclude by structure. Read the terminal
first; if it says REFUSED, nothing inside the fences was computed on that run.
arms run = 16 failed = 0
pins disagreeing = 0
RUN_TERMINAL COMPLETE
MARKER GENE_LENGTH_INSTRUMENT__NEAREST_GENE_LOTTERY_AND_LOEUF_DENOMINATOR
sha256 c11ae542d7344b95abd87ee84d065954f322d0118d2609b236fe69ae1dda4489
Corpus and provenance: corpus/gene-length-instrument/, digest-pinned, one file, aggregate
properties of a reference genome only.
Related: The ceiling on what a genetic score can know — what the best possible genotype-only score can ever do, computed exactly.
The measurements and the law that produces them are published so anyone may check them. Re-deriving these figures from the public gene model and constraint release requires no permission and no agreement with us. Affine.Earth asserts no claim over GENCODE, Ensembl or gnomAD, which belong to their authors and to the public.
Rights — source-available, all rights reserved. This wiki and its repository are published for public inspection and to let anyone re-derive the figures. They carry no LICENSE; under default copyright, all rights are reserved. No right is given or intended to use, run, or deploy it for any purpose other than re-deriving the published figures, nor to modify or build on it — any other use requires a written licensing agreement with the authors. · Affine.Earth · zero float · zero shear
Each step is the reason the next one exists. Nothing here is medical advice, and no page calls any medicine safe or unsafe.
1 · Why an exact safety screen at all
- Cures Without the Gatekeeper — the medicine front door: six real written medicines, one screen anyone can re-run
- The library admission law — what may enter, and the 71 arms that prove it refuses. The primary artefact.
2 · The three libraries, which grow rather than close
- The Library of Compound Cures — exact off-target maps for the medicines the registry publishes
- The Library of Proteins — 80,080 generated sequences, novel chemical matter, graded honestly
- The Library of Material Systems — what a system is, what was measured, where the law lives. C-007 absolute: no recipes
3 · The maps — every place a molecule could act, counted
- The off-target atlas — every nucleic-acid medicine the registry publishes a sequence for: WHERE it can pair
- The order of the bases — WHETHER THAT BURDEN IS UNUSUAL: 472 strands ranked against sixteen rearrangements of their own bases
- Where else could this guide cut? — the whole human genome, counted
- Designed, or forced by its own bases? — every clinical CRISPR guide, with its own composition as the control
- What a public genome deposit will tell you — and four ways it will mislead a health tool first
- Study 45 — which of nine billion answers a laboratory can act on — a safety review of AlphaGenome Atlas, measured live on 1,200 real variants at two genes. The headline score separates every one. The detailed tracks do not: splice-site usage hands back 950 of every 1,000 values shared with another variant at HBB and 998 at CFTR, and the shared values pile up in the quiet band where a bench clears a variant
4 · One medicine at a time
- Zilganersen — the first treatment for Alexander disease, screened on the real approved sequence
- A drug an AI designed — rentosertib for pulmonary fibrosis, and exactly what our instruments reach
- CAR-T, halted — the verdict a regulator could re-derive
- N-of-1 antisense — the only safety net at a population of one
- VERVE-102 — the off-target lattice a stranger can re-derive
- PM359 — prime editing, certified before anyone is dosed
- Del-Zota — the one safety question that can be made exact
5 · What keeps a disease alive, and what moves it
- Study 26 — master regulator bonds — 17 tumour types, 7,673 tumours; eleven compound pairs where no single agent among 20,308 cleared any
- Study 20 — Rife frequency — light and frequency, measured rather than dismissed
- Study 37 — five molecules — 37,910 "validated discoveries", 5 distinct molecules; why per-item validation cannot see a corpus-level defect
- Are the generated cures new? — 80,080 peptides against the human proteome
- Study 16 — disease type · Study 17 — chemistry InChIKey · Study 14 — protein lattice
- No language model in this stack — what the answers here are made of: measured 2026-09-12, no cell runs a model process, opens a model port or holds an unmasked model unit, and a gate refuses their return
- Run any study in your browser — all ninety programs open on your own device, forty-nine run there, and the run tells you whether it printed the sealed bytes
- The ontology — grades, terminals, controls, and what each page may say
- Zero Float · Zero Shear — the method in one page
- Ask someone you trust to check this — what to hand a sceptic
- Readers’ guide · Program index — all 42 studies · White paper · Roadmap
- The full-grade replacement — 49 retired instruments, 4 verticals
- The exactness seam — the business case
- Build a study — Falcon walkthrough — how to add one yourself
The same move every time: take a domain where a floating-point model is the accepted instrument, compute the same quantity in exact integers, and seal the cases where the two render opposite verdicts. The subject under grading is always the instrument, never the phenomenon.
- Study 48 — the atom already has an address — silicon dimers 3.840 Å apart, the smallest commanded scale on the board: a length carried in single precision mis-addresses its first atom at step 8,783; an address cannot
- Study 49 — the phase code never needs π — a phase-only modulator takes 256 codes per pixel; the code is a ratio of integers
- Study 50 — CMS raw data from the LHC, read exactly — CMS's 2011 collision bytes streamed from CERN Open Data into the Affine IDE and read in exact integers, every collision a hologram you can turn: 138 of 3,564 bunch slots carry 93,110 of 120,742 collisions, and in 3,854 the event record reads its slot exactly 3 lower than the pixel boards · public release
- Study 55 — IceCube: the light in the ice, hit by hit — IceCube's calibrated hits read byte for byte: 4 published files, 9,749 events, 2,289,821 hits, a census seal per file
- Study 47 — translation shear: the meaning that survives a language — LAW FROZEN · LIVE CLAIM, measured 2026-09-11 and again fleet-wide 2026-09-12: translation as an exact coordinate transform, charts derived in memory at every start from the raw rows of a pinned public weight file and never written down; one lattice digest on 9/9 cells, zero drift, every refusal named. The generative comparison arm is ABSENT — there is no generative translator in the stack
- Study 34 — the observer-invariant verdict — why a safety verdict needs an exact law, not a bigger computer
- Study 35 — the safety brain that forgets — deaf in 8.4 seconds, forgets across machines, disagrees with itself
- Study 36 — the language game of Fermat's Last Theorem — guess and shear, or project
- Study 40 — the number the simulation throws away — their ICO result computed as a fraction; in float the effect returns 0 at every width, and an effect returned as zero cannot be searched for
- Study 41 — fifty years of solving the wrong problem — the ordering was never about time, it was about arithmetic; 177× the work and 2,400× the wrong guesses to return the answer the machine already had
- Study 42 — The Exact Contract — 2.7M flood settlements in Int128 cents; the step exists and the rigidity does not
- Study 29 — continuous-model shear
- The lattice holds · Impact study — continuum dead · Death of continuous shear
- Fourier Phantom — Anima FNO vs 11+12+13 · Stellar dynamo kill shot
- QCD: freedom is dilation · UUM-8D vs IUT — WIN
- Peer-review bundle · Conjecture alignment
- We need fusion — the verdict every machine can check
- Affine Fusion Control — the local exact-integer court · public release
- Fusion researcher's guide
- Study 33 — the fusion control verdict court
- Every season, fifty tonnes — the biosphere-safety case
- The forcing nobody measures · Impact study — the SpaceX trajectory
- Study 31 — the biosphere joint ledger — LIVE on the court, 9/9 cells
- Study 28 — the wet-bulb threshold court — Act 1 sealed
- Study 32 — the taxi-out floor court
- Where humans actually yield — the fatigue curves, and where the rules already agree
- Study 30 — sovereign edge pod · Manufacture contracts
- The detector that flags the whole market — a manipulation geometry in exact integers, and the regulator's own indicator scored against a legitimate quoter
- Study 43 — almost every order is cancelled, and that is normal — nine sessions, three operators, two continents: 935 to 998 of every 1,000 orders that ended, ended without trading. A check that flags almost everything is a denominator, not a detector — and the stock you pick moves it further than the exchange does
- Study 44 — nine billion answers, four billion ways to say them — AlphaGenome Atlas ships 9 billion predictions in single-precision floats, which hold 4.28 billion distinct values: 52 of every 100 variants MUST share a score with another. Agreement and exhaustion look identical on the wire
- Study 38 — the loss-reserve triangle — a reserve is an exact rational; 481 of 482 verdicts identical in both arithmetics; the sixteen-billion figure comes from an unchecked premise
- Study 39 — the actuarial domain — life, pensions, multi-state and aggregation; the margin is 8 significant digits at its tightest
- Run any study in your browser — the ▶ badge beside a program name opens it in the Studio, already built and carrying its inputs, and runs it on your machine with nothing sent back
- Explore the live courts
- MCP user guide — all 51 tools · Deterministic no-float courts for LLMs
- Court Client — generic wasm IDE for every court · Court-client checkpoint
- Coding Court — the verdict IS the artifact
-
Zed — the coding agent, for developers — set Zed 1.20.2 up on
https://affine.earth/v1, no language model anywhere; what a turn does, the wire, the autonomous closure -
Zed — Minecraft comes to life — the two-person interaction, sealed: it asks, cites, clones a sibling with a value you supply, verifies by replay; the court flips
REFUSED_UNKNOWN_BUDGET → WIN - Zed — the agent that teaches the whole domain — architecture, protocols, server management and git, each answered from lines it read and instruments it ran; five closures PROVEN, and the cattle question answered with a counter the fleet did not have
- Math Court on Glama · Math Court user guide · Example app — entire court
- Quantum algorithms inventory · Shor witness certifier
- MCP clients (public)
- Glama connector
- Look in the UI (no visitor data)
A study appears here under the state its evidence has earned, and above under the question it answers. The two are different filings of the same work, on purpose.
✅ LAW FROZEN · DATA SEALED
- Study 06 — explosion vs earthquake · Study 07 — Sgr A* raw visibilities
- Study 11 — Ehrhart volume · Study 12 — parallel repetition · Study 13 — Connes rigidity
- Study 14 — protein lattice · Study 16 — disease type · Study 17 — chemistry InChIKey
- Study 18 — material STD · Study 19 — Go First dice
- Study 26 — master regulator bonds — 17 tumour types, every finding published
🔴 LIVE CLAIM — standing, not sealed
- Study 02 — launch ionospheric holes · Study 02 — regulatory alarm
- Study 09 — global convective bond · Study 20 — Rife frequency · Study 21 — stellar dynamo
- Study 22 — 2-local Hamiltonian · Study 23 — spin glass · Study 24 — N-representability · Study 25 — exact permanent
🌊 CHARTER · OPEN — the findings, published either way
- Study 03 — flare SIDs — archive went dead · predictions and validations
- Study 04 — tsunami vs surge — partial seal · Study 05 — Forbush decreases
- Study 08 — Gaia BH1 — no corpus until DR4 · Study 10 — Fermi / dark matter — does not disprove DM
- Study 15 — Skala DFT shear · Study 27 — exact nuclear scattering
- Overview · First 27 days · Success criteria
- The science, and what history says · Blind spots — five stories magnitude models miss
- Historical corpus · Data archives — every source, exactly how to reach it
- Model shear · Benchmark results · Prediction registry
- Substrate architecture — how a shadow becomes a geometry
- Operations runbook · Satellite & aviation advisory