-
Notifications
You must be signed in to change notification settings - Fork 0
Study 26 Master Regulator Bonds
Charter published 2026-08-26. Corpus not yet ingested. Law drafted here, to be frozen verbatim at corpus stage before any scoring.
In 2021 the Califano laboratory published a pan-cancer analysis in Cell reporting that 407 master regulator (MR) proteins canalize the genetics of individual tumor samples from 20 TCGA cohorts into 112 transcriptionally distinct tumor subtypes, with the MRs organizing into 24 pan-cancer master regulator block modules (MRBs), and more than 50% of somatic alterations detected in each individual sample predicted to induce aberrant MR activity (Paull et al., Cell 184(2):334-351, 2021, PMID 33434495 — VERIFIED, abstract fetched from PubMed 2026-08-26; the abstract states no total tumor count, and this page prints none).
Set that against what the mutation counts themselves do. Across 92,439 analyzed tumor samples spanning 541 cancer types, the median tumor mutational burden was 3.6 mutations/Mb with a range of 0 to 1,241 mutations/Mb, and type medians ran from 0.8 (bone marrow myelodysplastic syndrome) to 45.2 (skin squamous cell carcinoma) — a spread of three orders of magnitude across individual tumors (Chalmers et al., Genome Medicine 2017, PMC5395719 — VERIFIED, full text fetched). Within a single type the label does not pin the number: in soft tissue angiosarcoma the median was 3.8 mutations/Mb, yet 13.4% of cases carried more than 20 (VERIFIED, same source). Across 3,083 tumor/normal pairs, the median frequency of non-synonymous mutations varied by more than 1,000-fold across cancer types (Lawrence et al., Nature 499:214-218, 2013, PMC3919509 — VERIFIED, full text fetched). And in 2026 the MSK-IMPACT 50K analysis of 54,331 tumors from 48,179 patients across 448 histological subtypes reported that one-third of all driver alterations arose in non-canonical contexts, concluding that the functional role of a driver depends on the cancer type and clinical context in which it arises (Bandlamudi, Muldoon, de Bruijn et al., Cancer Cell 44(5), 2026, PMID 41895280 — VERIFIED via the Mount Sinai institutional abstract record).
The field's own summary of this landscape is Vogelstein's: a small number of "mountains" and a much larger number of "hills," with a typical tumor carrying two to eight driver mutations drawn from roughly 140 driver genes that classify into 12 signaling pathways regulating three core cellular processes — and the concession, in the same paper, that "Methods based on mutation frequency can only prioritize genes for further analysis but cannot unambiguously identify driver genes that are mutated at relatively low frequencies" (Vogelstein et al., Science 339:1546-1558, 2013, PMC3749880 — VERIFIED, full text fetched).
So the field's published position — the "tumor checkpoint" / "oncotecture" architecture (Califano & Alvarez, Nature Reviews Cancer 17:116-130, 2017, PMID 27977008 — REPORTED from the article listing, body not fetched) — is that tumors with wildly different mutational landscapes converge onto a small, conserved set of master regulator proteins whose aberrant activity maintains the tumor state. Mutation-burden reads are magnitude science: count the lesions. The master-regulator architecture is shape science: find the load-bearing bonds. That reframing is the field's, not this program's.
The program's parallel, stated as the program's reading: a load-bearing regulator set whose intactness maintains a state regardless of surrounding noise is, structurally, what this program calls a bond — the same object graded in Study 14 — Protein lattice manifold and audited across the Eclipse 2026 model-shear lineage. The house thesis applies unchanged: an adversary can copy the SIZE of a response (a mutation count, a burden per megabase) but cannot keep the forcing APPOINTMENT (which named regulators are active, in which named disease context, inverted by which named compounds). The mapping is an analogy that motivates a decidable test. The test is what gets sealed. The analogy never is.
This study claims no cure, and exists partly to show what a claim that is not a cure claim looks like: a network-shape claim on public, already-measured bytes, graded in exact strings and exact integers, published win and miss alike.
Before any biology is graded, the instrument must be shown able to fail. A scoring rule that any random protein set can satisfy is void before the first byte is scored — the program has already retired one instrument for exactly this defect, and an instrument that cannot distinguish is a turn counter, not a court.
The score function, frozen here, verbatim, before any set is scored. It is named first because a gate that does not name its score cannot be shown to discriminate:
σ(R, context) = the integer count of regulators
r ∈ Rwhose signed integer activity, computed by the S1 aggregation rule below, falls within the top N of the context's full regulator ranking. N is frozen at the published MR set's cardinality. Ties in the ranking are broken by HGNC symbol in ascending lexicographic order.
σ reads the context's own expression and regulon bytes and the integer N. It does not reference the published MR set's identity anywhere — that is the whole point, and the reason S4 is not scored on S1's overlap-with-the-published-set. A control arm graded on the published set's identity is green whatever the biology does, because a random set trivially misses a named list.
The gate, in integers, frozen before any MR set is scored:
- Draw 1,000 size-matched random regulator sets. Each is a uniform sample, without replacement, from the regulators present in the frozen regulon file for that disease context — not from all 19,297 HGNC protein-coding symbols. A random draw from the whole protein-coding space would fill nearly every set with symbols that carry no regulon, σ would be undefined or empty by construction, K = 10 would never be approached, and the gate would be always-green: the exact defect this section exists to prevent.
-
A drawn member with no regulon in the frozen file is not scored as zero. Absence and zero are different answers. Such a member is recorded with the token
NO_REGULON, its set is discarded and redrawn, and the total number of discards is published in the seal. No set is scored with an undefined member folded in as a nil contribution. - Score every random set with the identical frozen σ applied to the published MR set (same expression signature, same regulon file, same rank quantisation, same integer threshold, same tie rule).
- PASS: at most K = 10 of the 1,000 random sets score at or above the published MR set's σ. FAIL: 11 or more do.
- K = 10 is frozen here, in the charter, before any corpus byte is read. It does not move after scoring. A FAIL is published as a FAIL.
If the gate fails, every downstream criterion in this study is void for that disease context, and the page says so. The randomised sets are the control arm; a court that is never shown the healthy population cannot tell a cause from a coincidence.
| Reading | What it holds constant | The observation that separates it | Sourcing |
|---|---|---|---|
| Mutation magnitude — the tumor is its lesion count and lesion list | The count and the list | The identical exact key gives opposite answers in different contexts: BRAF V600E responds to vemurafenib in melanoma and barely in colon cancer, via EGFR feedback the mutation list never shows (Prahallad et al., Nature 483:100-103, 2012, PMID 22281684). Naive frequency reasoning also manufactures false positives — 101 of 450 "significant" lung-squamous genes were olfactory receptors (VERIFIED, PMC3919509). | REPORTED (Prahallad, abstract summaries); VERIFIED (Lawrence) |
| Network shape — the tumor is a maintained regulatory state; the MR set is load-bearing | The identity of the active regulator set | Protein-activity inference from regulons "outperformed mutational analysis in predicting sensitivity to targeted inhibitors" (Alvarez et al., Nature Genetics 48:838-847, 2016, PMID 27322546 — VERIFIED, abstract quoted verbatim from the fetched record). Sharpest case: reversible, chromatin-mediated drug tolerance with >100-fold sensitivity change and an identical genome before, during, and after — a mutation-list method has zero signal by construction (Sharma et al., Cell 141:69-80, 2010, PMID 20371346 — REPORTED). MEK-inhibitor resistance by kinome reprogramming in hours to days, no new mutation (Duncan et al., Cell 149:307-321, 2012 — REPORTED); gefitinib resistance by MET amplification re-reaching PI3K while the EGFR key is unchanged, 4 of 18 resistant specimens (Engelman et al., Science 316:1039-1043, 2007 — REPORTED). | VERIFIED / REPORTED as marked |
| The null — MR sets are descriptive summaries of expression with no causal load | Nothing; it predicts random size-matched sets score as well | This is what Section zero exists to test, and it is not a straw man. The prospective evidence is mixed on the record: in patient-derived xenografts, OncoTreat-predicted drugs achieved 91% 30-day disease control and 15 of 18 induced the predicted MR-module inversion in vivo (Mundi et al., Cancer Discovery 13(6):1386-1407, 2023, PMID 37061969 — VERIFIED, abstract fetched; PDX disease control, not a human response rate) — while the flagship human trial of the predicted drug entinostat in neuroendocrine tumors enrolled 5 patients, terminated early, and did not meet its primary endpoint, with all 4 evaluable patients at stable disease and tumor growth rates at 17%, 20%, 33%, and 68% of pre-enrollment rates (Jamison et al., The Oncologist 29(9), 2024 — VERIFIED, PMC full text fetched). The null is retired only by the gate, never by citation. | VERIFIED as marked |
An independent benchmark also reported a GSEA-based regulator-enrichment method outperforming the VIPER algorithm in all but one of its experiments (RegEnrich, PMC8752721, 2022 — REPORTED), and the term "master regulator" itself is under published criticism as diluted ("When everything is a master regulator, nothing is," Mol Biol Cell 2025, PMC11974949 — REPORTED). Both belong on this table's null side of the ledger, and both are why the biology is graded here only as appointment-keeping, never assumed.
This study speaks only in exact string keys and exact integers, on the rails the program has already built:
-
Disease = DOID key, joined to ICD via the xrefs the ontology itself carries — the Study 16 — Disease type rails. The frozen DOID artifact (release
2026-07-31, downloaded and counted 2026-08-26 — VERIFIED) holds 12,247 active terms of 14,762 stanzas, licensed CC0 by its own header. -
Regulators = HGNC symbol sets, backed by UniProt accessions. The frozen HGNC complete-set TSV: 45,045 approved rows, of which 19,297 protein-coding (VERIFIED, downloaded and counted 2026-08-26). UniProt release 2026_02 reports 20,431 reviewed human entries, accession grammar
[OPQ][0-9][A-Z0-9]{3}[0-9]|[A-NR-Z][0-9]([A-Z][A-Z0-9]{2}[0-9]){1,2}(VERIFIED, live REST headers). -
Drugs = InChIKey, 27 uppercase characters, the Study 17 — Chemistry InChIKey rails — SMILES and float MW/logP are not the court. The LINCS metadata itself carries
inchi_keyandpubchem_cidcolumns (VERIFIED, file downloaded and header read 2026-08-26), so the chemistry rail joins to the perturbation archive with zero float arithmetic and zero name-matching. The one place a drug name enters this study is the OncoTreat comparison inside S3, and it is fenced behind a frozen, checksummed mapping rail declared below — never resolved ad hoc at scoring time. PubChem, holding 124,598,147 compound records (VERIFIED via NCBI einfo 2026-08-26), resolves any key anonymously. - Structures = 4-character PDB ids, the Study 14 — Protein lattice manifold rails (N = 258,616 holdings, law frozen and claim live).
- Activity calls are quantised to integer rank ordinals. Identity is decided by exact string equality — never by multiplying or dividing floats. Three crossings into that exactness exist, and each is declared at its own site in the frozen law below — one integer-to-integer quantisation and two genuine float-to-exact reads. None is left implied.
The court that accepts these keys already exists in public. The Affine Math Court is live on Glama — open MCP endpoint https://affine.earth/language-invariant/mcp, registry earth.affine/math-court, connector https://glama.ai/mcp/connectors/earth.affine/affine-earth-math-court-remote — with 21 public tools and court domains that already include disease, chemistry, materials, PDB, and health. The wire is decimal strings: a JSON Float64 is refused at the door, and verify_jordan_bond returns AFFINE_JZ_SHEAR_ZERO on a uniform lock. See the Math Court on Glama and the Math Court user guide.
One warning is promoted to law here rather than buried in limits: of the 12,328 genes in an L1000 profile, only 978 landmark genes are physically measured; roughly 11,350 (about 92% — arithmetic on the platform's own figures, REPORTED) are computationally inferred from the landmarks. This study's court scores only the 978 measured landmarks. The inferred genes are a model's output, not an observation, and a court whose doctrine is exact integers on public raw bytes does not seat a model's output as evidence.
Every input is retrievable without an account, anonymous, and checksummable; one carries a non-commercial license term, stated in its own row. All entries verified live on 2026-08-26 unless marked REPORTED.
| Source | URL | Format / keys | Version / cadence | Auth | Sourcing |
|---|---|---|---|---|---|
| TCGA open expression (GDC) | https://api.gdc.cancer.gov/data/<uuid> |
STAR gene counts TSV; Ensembl gene_id + HGNC gene_name; integer unstranded column is the court, TPM/FPKM floats excluded |
Data Release 46.0 (2026-08-10), API 8.5.0; md5 per file in manifests | None — 11,505 open files of this type, 0 controlled; a 4,244,281-byte file fetched anonymously | VERIFIED |
| GDC open somatic mutations (S2's mutation rail) |
https://api.gdc.cancer.gov/data/<uuid>, data_type = "Masked Somatic Mutation"
|
MAF (gzipped TSV); Hugo_Symbol + Entrez_Gene_Id are the exact-string join to HGNC; integer Start_Position / End_Position; no float column enters the court |
Data Release 46.0 (2026-08-10); md5sum returned per file by the files endpoint |
None — 24,498 open files of this type against 2,268 controlled, of which 10,640 open within the TCGA program; one 39,235-byte MAF fetched anonymously and its md5 57d08529…218106 reproduced byte-exact |
VERIFIED (live API query + download 2026-08-26) |
| LINCS L1000 Phase I | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE92nnn/GSE92742/suppl/ |
Level 5 GCTX (HDF5), 473,647 signatures x 12,328 genes; sidecar TSVs incl. pert_info (51,383 perturbagens counted: 20,413 trt_cp compounds with InChIKey column). Level 5 values are float moderated z-scores — see Crossing B |
Frozen 2017-03 vintage; GSE92742_SHA512SUMS.txt.gz for byte-exact verification |
None | VERIFIED |
| LINCS L1000 Phase II | https://ftp.ncbi.nlm.nih.gov/geo/series/GSE70nnn/GSE70138/suppl/ |
Level 5 GCTX, 118,050 x 12,328 (combined open Level 5 corpus: 591,697 signatures); float values as above | Frozen 2017-03-06 vintage; own SHA512 manifest | None (the clue.io 2020 build is registration-gated and is NOT a rail here) | VERIFIED |
| ARACNe regulons | Bioconductor aracne.networks v1.38.0 (Bioc 3.23), DOI 10.18129/B9.bioc.aracne.networks |
25 regulon objects, not 24: data/ holds 25 regulon*.rda with 25 matching man/*.Rd. 24 are TCGA cohort codes (blca…ucec) that are also GDC project ids; the 25th, regulonnet, is the GEP-NET neuroendocrine interactome (Alvarez et al. 2018 cohort) and is not a TCGA cohort code and has no GDC project — see S1. Regulator→target edge lists, Entrez ids mapped 1:1 to HGNC; write.regulon() exports flat text a court can hash. Upstream inconsistency, declared: data/datalist names a 26th object, regulonskcm, for which no .rda and no man page exist in the release — S5 names it so a byte-exact seal does not read it as corpus damage |
Package release cadence; manual dated 2026-08-25; RELEASE_3_23 cloned and counted 2026-08-26 | None to fetch, but file LICENSE is a Columbia University software evaluation license — non-commercial, non-profit/academic/educational research only, no redistribution. The caution before commercial-adjacent use stands, and it is a license term, not a suspicion |
VERIFIED (package, manual, 25-object list, license text — clone read 2026-08-26) |
| HGNC symbols | https://storage.googleapis.com/public-download-files/hgnc/tsv/tsv/hgnc_complete_set.txt |
TSV, 16,948,051 bytes; hgnc_id, symbol, ensembl_gene_id, uniprot_ids columns are the joins |
Monthly snapshots archived in the same bucket; old EBI FTP path now 404 | None | VERIFIED |
| Disease Ontology | http://purl.obolibrary.org/obo/doid.obo |
OBO, 7,192,337 bytes; DOID keys + ICD/OMIM/MeSH xrefs | Pinned in-file: releases/2026-07-31; CC0 license in header |
None | VERIFIED |
| UniProt | https://rest.uniprot.org/uniprotkb/ |
Accession strings; counts via X-Total-Results header |
Release 2026_02 (2026-06-10) via X-UniProt-Release header |
None | VERIFIED |
| PubChem | https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/<KEY>/cids/JSON |
InChIKey → CID round-trip | Live; 124,598,147 compounds | None | Count VERIFIED; PUG pattern REPORTED |
What stays controlled at GDC — raw BAM/FASTQ, germline calls, certain clinical elements under dbGaP phs000178 — is never needed by this court. Counts, masked somatic mutations, clinical, and biospecimen are open (policy quotes REPORTED from the fetched GDC access page; the open/controlled split of both expression and mutation files independently VERIFIED by live API query). The controlled-access boundary is respected by construction: the law below reads no byte a stranger cannot read.
The one hand-made rail, named because it is not a public byte. S3 compares an exact InChIKey set against OncoTreat's published rankings, and those rankings are REPORTED paper content keyed by drug names, not InChIKeys. A name is not an exact key, and resolving one at scoring time would be exactly the non-exact join this court forbids. So the mapping is pinned as a rail: a paper_drug_name → pert_iname → inchi_key table, transcribed once from the named tables of the named papers, joined only to the pert_iname and inchi_key columns of the frozen pert_info file whose SHA512 is already in the GEO manifest, then checksummed and published in the corpus before any scoring. Every row is auditable against the two public artifacts it sits between. A paper name that does not resolve to exactly one InChIKey is not guessed: it is published in an exact UNRESOLVED list, frozen with the rail, and excluded from both sides of the S3 count. This is the only artifact in the corpus this program mints rather than fetches, and it is declared here rather than discovered later.
The upstream regulon values (tfmode, likelihood) ship as floats in the R distribution (VERIFIED from the package manual). likelihood is refused outright. tfmode is not refused — its sign is read, and reading the sign of a float is a float-to-exact conversion, not a refusal of one. It is declared as Crossing C below and performed exactly once, at load time.
Drafted here; frozen verbatim, with every integer bound to a named release string, before the first byte is scored. Every criterion below is decidable by string equality and integer comparison on public bytes.
The declared crossings — three, each at its own site, each performed once. The earlier draft of this charter declared a single crossing and named the wrong one; both statements are corrected here on the page's own text:
-
Crossing A — GDC counts → rank ordinals (S1, S2). Integer
unstrandedread counts to integer per-sample gene rank ordinals. This is integer-to-integer: a quantisation, not a float crossing. Frozen tie rule: tied counts all take the minimum ordinal of their tied block (competition ranking); residual ties by HGNC symbol ascending. - Crossing B — LINCS Level 5 → rank ordinals (S3). Level 5 GCTX values are float moderated z-scores, so obtaining "measured landmark-gene rank ordinals" is a genuine float-to-exact conversion. It is declared here, performed once at read time, by a frozen rule: rank within the 978 measured landmarks of that signature only, competition ranking as above, residual ties by HGNC symbol ascending. No float survives this step into any comparison, sum, or seal.
-
Crossing C — regulon
tfmode→ sign token (S1, S2, S3).tfmodeis a float and extracting its sign is a read of that float. Declared here, performed once at load time, by a frozen rule:tfmode > 0 → "+",tfmode < 0 → "−",tfmode == 0 → "0"recorded as its own token and never folded into either sign.likelihoodis refused entirely and never read.
After A, B, and C the court is exact set arithmetic and integer comparison, end to end.
The S1 derivation rule, frozen verbatim — named here because "top-N regulators by integer rank" is not a rule until its aggregation operator, sample universe, and tie rule are written down:
Sample universe: the frozen list of GDC file UUIDs for that cohort — open STAR gene-count files, primary-tumor aliquots, one file per case selected by ascending UUID string order. The list and its md5s are named in the seal. Per-sample ordinals: Crossing A. Per-sample regulator activity: for regulator
r, the signed integer sum over its regulon targets of(+ordinal)where the Crossing C token is"+",(−ordinal)where it is"−", and no contribution where it is"0"(a"0"target is recorded, not silently dropped). Cohort aggregation operator: the integer sum of per-sample activities across the frozen sample universe. Summation, never a mean — no division, so no float can enter. Ranking: regulators ordered by that integer, descending; ties by HGNC symbol ascending.
-
S1 — Regulon recovery. For each disease context (DOID key, TCGA cohort code), the published MR set for that context must re-derive from the public regulon file plus open GDC integer counts as the top-N regulators by integer rank under the rule above, N frozen at the published set's cardinality. PASS: |published set ∩ derived top-N| ≥ M, M frozen per context at corpus stage as an exact integer, before scoring. Identity by HGNC symbol string equality only. The regulon-to-GDC-cohort join is 24-of-25, not 25-of-25:
regulonnetis a GEP-NET interactome with no GDC project id, so it carries no GDC sample universe, is excluded from S1 and S2, and appears only in S3's neuroendocrine context where its cohort is the Alvarez et al. 2018 GEP-NET cohort and not a GDC project. That exclusion is frozen here, not decided at scoring time. -
S2 — Subtype appointment. For each frozen pair of subtypes within a cohort, the MR-set identity must separate the pair where the mutation list does not: |MR(A) Δ MR(B)| ≥ m₁ (symmetric difference, exact set op on symbol strings) while the top-t mutated-gene lists of A and B overlap by ≥ m₂ of t. The mutated-gene lists are derived from the open GDC Masked Somatic Mutation MAFs named in the archive table — frozen file UUID list plus per-file md5 in the seal, genes taken by exact
Hugo_Symbolstring, ranked by integer mutated-case count with ties by symbol ascending. All of m₁, m₂, t frozen per pair before scoring. Both readings graded by the same exact set operations — the magnitude adversary is separated on SHAPE and IDENTITY, never on magnitude. -
S3 — Inversion lookup. Over the already-measured LINCS Level 5 signatures (the 591,697 frozen GEO signatures;
trt_cprows with a non-emptyinchi_keyonly), the drugs whose measured landmark-gene rank ordinals (Crossing B) invert the MR set's direction at a frozen integer ordinal threshold form an exact intersection set of InChIKeys. That set is compared against OncoTreat's published rankings as REPORTED ground truth (Mundi et al. 2023; Alvarez et al. 2018 GEP-NET: 212 tumors, 107 compounds — REPORTED), joined only through the frozen, checksummed name→InChIKey rail described above, with itsUNRESOLVEDlist excluded from both sides. PASS/FAIL is an integer overlap count against a frozen bound. This is a lookup over measured bytes. It ranks what was measured; it predicts nothing unmeasured, and it replaces no trial. - S4 — The discrimination gate of Section zero: at most K = 10 of 1,000 size-matched random sets — drawn from the regulators present in that context's frozen regulon file — may score at or above the published MR set under the frozen σ. K and σ are already frozen by this charter.
-
S5 — Byte identity. Every scored file's checksum matches the published manifest (GEO SHA512, GDC md5 for both counts and MAFs), the minted name→InChIKey rail matches its own published checksum, and every release string (GDC 46.0, GEO 2017 vintages, HGNC snapshot date, DOID
releases/2026-07-31, UniProt 2026_02, aracne.networks 1.38.0) is named in the seal. The seal records 25 regulon objects present and thedatalist-named-but-absentregulonskcmas a declared upstream inconsistency, so a stranger reproducing the corpus reads the discrepancy as upstream, not as damage. A stranger with curl reproduces the corpus byte-exact or the seal does not stand.
The honest, sharp statement of what S3 buys — and all it buys: once master-regulator sets and drug perturbation signatures are held as exact keys over already-measured public perturbation data, asking "which measured perturbations invert this regulator set" becomes an exact set-intersection lookup instead of a statistical screen — a decidable query over measured bytes. Nothing more is claimed.
Patients are failed by irreproducible computational-oncology claims, and the failure is measured, not rhetorical. The signature-inversion field's own reproducibility audit found that querying the second Connectivity Map with signatures derived from the first re-identified the queried compounds at a success rate of 17% — "Low recall is caused by low differential expression (DE) reproducibility both between CMaps and within each CMap" (Lim & Pavlidis, "Evaluation of connectivity map shows limited reproducibility in drug repositioning," Scientific Reports 11(1):17624, 2 September 2021, doi 10.1038/s41598-021-97005-z, PMID 34475469 — VERIFIED, figures and quote reproduced from the PubMed abstract; the per-cell-line 29%–58% range is REPORTED only).
The Broad's own audit of its perturbation archive found that of the 924 shRNAs targeting the same genes with significant on-target projection ranks — the subset the paper built for direct comparison against CRISPR — only 41.8% had a larger on-target than off-target component, against 97.4% on-target dominance for the CRISPR sgRNAs. The study profiled roughly 13,000 shRNAs overall; 41.8% is measured on the 924-shRNA comparison subset, not on the full archive, and this page prints the figure only against its own denominator (PMC5726721 — VERIFIED, full text fetched; denominator corrected 2026-08-26). The qualitative finding — that shRNA signatures are heavily seed-driven where CRISPR signatures are not — is what carries into the law below; the archive-wide rate is not established by that measurement and is not asserted here.
And CRISPR re-validation showed the mechanism of action of 10 drugs already in clinical trials — involving roughly 1,000 patients — was mischaracterized: the named target was not the operative target (Lin et al., Science Translational Medicine 2019, PMID 31511426 — REPORTED).
Pre-registration plus exact keys plus public corpora is the repair this program can actually contribute: a network claim becomes checkable by a stranger with a terminal. The complete pipeline a stranger runs is public end to end — curl open TCGA counts and open masked-somatic-mutation MAFs from api.gdc.cancer.gov, install the public regulon and scoring packages from Bioconductor, derive the MR calls, name every regulator by HGNC symbol, verify every input byte against a published checksum (VERIFIED for the data path; the scoring-package provenance REPORTED). No account and no privileged data anywhere. One license term, stated rather than waved past: aracne.networks carries Columbia's non-commercial software evaluation license, so the stranger installs those bytes from Bioconductor under that license rather than receiving them from us, and commercial-adjacent use is outside what the license grants.
When the claim misses, the miss is published on this wiki with the same prominence as a win — that is the program's standing practice, and it is the part of reproducibility that costs something.
The same discipline protects the commons downstream: public archives (GEO, GDC, HGNC, DOID, UniProt, PubChem) are a shared scientific resource, and a court that pins release strings and checksums exercises them as archives rather than treating them as a mutable backdrop. Nothing here consumes a wet lab, an animal, or a patient; the entire study is a re-reading of what has already been measured and published.
- Not a treatment, not medical advice, not an efficacy result. The program precedent is Study 20 — Rife frequency, wiki label verbatim: "test claim, not efficacy." Study 26 grades a network-shape claim on public data. This page does not claim a cure and grades nothing that could establish one.
- Computation replaces no laboratory and no trial. S3 ranks already-measured perturbations by exact intersection. It predicts nothing unmeasured. The one human trial of the flagship predicted drug did not meet its primary endpoint (VERIFIED above), and that fact sits in this charter's own evidence table.
- The biology is not this program's discovery. That network/protein-activity state can outperform mutation lists is already published and validated in vitro (Alvarez et al. 2016 — VERIFIED), and the move from gene-centric to context-aware framing is the field's own 2026 conclusion (Bandlamudi et al. — VERIFIED). What Study 26 contributes is the grading discipline: law frozen before scoring, identity by exact string, integers on public raw bytes, miss published beside win.
- Regulon inference is statistical upstream of our exact crossings. ARACNe (Margolin et al., BMC Bioinformatics 2006, PMID 16723010 — VERIFIED) is mutual-information network inference; it requires large cohorts, leaves orphan tissues without matched interactomes (metaVIPER exists to work around exactly that — REPORTED), and an independent benchmark reported a competing method outperforming VIPER in all but one experiment (REPORTED). The court inherits these upstream statistics and declares them, exactly as it declares Crossings A, B, and C at their own sites rather than as one implied crossing.
-
One rail is minted, not fetched. The
paper_drug_name → pert_iname → inchi_keymapping behind S3's OncoTreat comparison is transcribed by this program, not downloaded from an archive. It is frozen and checksummed before scoring, published in the corpus row by row, and itsUNRESOLVEDnames are excluded from both sides of the count — but it is the one artifact in this study a stranger audits rather than re-downloads, and it is named here for that reason. - MR causality is the field's claim. It is graded here only as appointment-keeping — does the named set re-derive, separate, and invert on schedule — never assumed. The "master regulator" term itself is under published definitional criticism (REPORTED above), which is one more reason the court grades named sets, not the word.
-
Known false-positive modes are carried in the law, not footnoted: 92% of the L1000 gene space is inferred, so the court scores landmarks only; on the 924-shRNA comparison subset 58.2% did not show a larger on-target than off-target component, so
trt_cpcompound rows with InChIKeys are the S3 universe and no shRNA row is seated as evidence — the archive-wide seed-domination rate is not established by that measurement and no criterion depends on it; drug-target annotations have a measured error rate on trial-stage drugs, so no target annotation is load-bearing anywhere in S1–S5. - What a measurement shows versus what it licenses: the long-tail sources show that the set of mutated genes differs markedly between tumors sharing a subtype label and that most driver genes sit in the low-frequency tail. They do not license the claim that most tumors lack a common driver, and this charter does not make it (guardrail VERIFIED as a negative result against PMC6336235).
- Controlled-access boundaries are respected by construction — the law reads open bytes only, expression and somatic mutations alike.
- Claim posture. Under 21 CFR 801.4, intended use is decided by objective intent as shown in written statements (VERIFIED, regulation text fetched from Cornell LII) — which means this page's own wording is the boundary. The wording is: this study scores agreement with published measurements and grades shape and identity. It does not recommend, indicate, diagnose, select a therapy, or address any patient. It is research output, not for use in diagnostic procedures, with no clinical validation performed or claimed. The current FDA clinical-decision-support guidance was reissued 2026-01-06 (VERIFIED via two independent legal analyses; the FDA primary PDF was unreachable this session), and the conservative reading — genomic-input decision software remains device-regulated — is adopted here by staying entirely outside decision support.
- Study 14 — Protein lattice manifold — the structures rail: 4-char PDB ids, N = 258,616 holdings, law frozen, claim live.
- Study 16 — Disease type — the disease rail: ICD/DOID industry keys; generated labels do not WIN; does not treat health as a pocket.
- Study 17 — Chemistry InChIKey — the drug rail: the 27-character key is the identity; SMILES and float MW/logP are not the court.
- Study 20 — Rife frequency — the medical-scoping precedent: test claim, not efficacy.
- Affine Math Court on Glama — the live court that already accepts disease, chemistry, and PDB keys on a decimal-string wire.
- Eclipse 2026 — Model shear — the shape-versus-magnitude lineage this study extends into biology.
- Shear Studies Index — the program board, Studies 01–26.
Status: OPEN — charter published, corpus not yet ingested. Nothing here is sealed until the corpus runs.
- The lattice holds
- Impact study — continuum dead
- Death of continuous shear
- Fourier Phantom — Anima FNO vs 11+12+13
- Stellar dynamo kill shot
- QCD: freedom is dilation
- Look in the UI (no visitor data)
- Explore the live courts
- MCP clients (public)
- Example app — entire court
- Math Court user guide
- Math Court on Glama
- Glama connector
- Zero Float · Zero Shear
- UUM-8D vs IUT — WIN
- Readers’ guide
- White paper
- Program index
- Language-games study board
- Known discoveries — family ledger
- Discoveries by user
- Low-friction user flows
- How the IDE hologram works
- Known molecular discoveries
- Study 14 — PDB play
- Study 11 — Ehrhart Volume
- Study 12 — Parallel Repetition
- Study 13 — Connes Rigidity
- Peer-review bundle
- Conjecture alignment
- Study 07 — Milky Way Results
- Study 06 — Explosion Results
- Study 10 — Fermi / Dark Matter
- Study 04 — Tsunami (partial)
- Study 09 — Convective bond
- Study 14 — Protein lattice
- Study 15 — Skala DFT shear
- Study 16 — Disease type
- Study 17 — Chemistry InChIKey
- Study 18 — Material STD
- Study 19 — Go First Dice
- Study 20 — Rife frequency
- Study 21 — Stellar dynamo
- Study 22 — Ground state is a coordinate
- Study 23 — Spin glass, no freezer
- Study 24 — Molecule is consistent or not
- Study 25 — Supremacy is a rounding error