Skip to content

Pipeline Fidelity Gate

ProxyPrints Docs Bot edited this page Aug 6, 2026 · 12 revisions

Pipeline-fidelity gate — canonical status page

GitHub issue #154 (internally referenced elsewhere as "task #151" — a pre-board internal task-ledger number, not a second GitHub issue; issue #154's own title cross-references both). This is the single page for this gate's status, its (now-decided) owner decisions, and any in-flight implementation gating the fire. Every underlying fact below is linked to its source, not copied — if a number here disagrees with its linked source, this page is wrong and should be fixed, not the source.

Data as of 2026-07-22T18:01Z, re-verified live against production Postgres for this page (sudo docker exec mpcautofill_django python manage.py shell, read-only queries only). §10/§11 carry a later, narrower live re-verification (2026-07-23T00:16–00:24Z) for the #340 footprint sizing and the Stage C run-identity note specifically — dated inline in those sections rather than bumping this whole-page timestamp. §12 carries a further dated update (2026-07-23T01:33Z onward) for the deploy step and the Bug-B whole-DB reparse dry-run outcome, and §9 was amended the same day per issue #347's Tron-reviewed zeroing plan. §13 carries a further dated update (2026-07-23T09:0x–09:3xZ) for the zeroing steps' actual execution (B(i) write, B(ii)+B(iii) retraction) and the §9(c) Bug-A forced-escalation sample, re-verified live at that time. §9(d) carries a further update (2026-07-23T09:14–11:12Z) recording the 4c pilot dry-run's own outcome, which was still in progress as of §13's own check. §14 records the pilot --write and the first consensus_recompute --apply (2026-07-23), after which the gate was fired end to end. §14 was extended 2026-07-24 (DB-verified live against production Postgres) with five further passes that ran the same night into the next day — a lexicon-gate retraction, a marker reparse, an artist-credit fill, a calculator re-pass, and a second consensus_recompute closer — which is the sequence's actual true end state; that update also corrects the live resolved-printing count from 4 to 3 (the closer flipped one card resolved→unresolved after §14's original snapshot was taken) and folds in a real, verified finding that all 218,345 cards remain artist_vote_status=unresolved despite the artist-fill pass (see §14's "Artist-consensus finding"). §15 (new, 2026-07-24) records a separate, later arc — the wave-1 Bug-A blank-tier-1 re-scan (10,437 cards) carried through re-scan, reparse retraction, a full-pool lands write, a wave-1-scoped Stage D write, and a closing consensus_recompute dry-run that found zero transitions — closing that slice of the Bug-A tail; the remaining ~6,535-card tail is routed to Stage E streaming's shakedown cohort instead of a further batch pass (§14's "What remains open" item 1, §15). §14's own Operational notes were also updated the same pass to record issue #402's fix (PR #412) going live via deploy-3, and "What remains open" item 2 (moderation) now reflects an owner ruling to defer package sizing pending further passes. See "Chain" below for where each number's provenance sits.

1. Gate definition

Stage D (local_calculate_verdicts.py) carries a hard precondition before any full-catalog fire: calculators must call the existing shipped identification code paths with ImageEvidence-supplied inputs, not re-derive their logic; a stratified-sample parity replay against pilot run 20260716T193408-6613a1a6's recorded outputs must show zero unexplained divergence; a full knowledge-inventory sweep (every empirically-derived constant/threshold/override/skip-reason mapped to its home in the new pipeline, or flagged missing) must be clean. Full precondition wording and the Stage D build it gates: features/catalog-completion-plan.md's "Harvest-calculate pipeline" section.

2. Current status

artifact status
Artifact 2 — knowledge-inventory sweep DONE (2026-07-22); all 3 MISSING-constant decisions now made — see §3 below. Full constant-by-constant table: reports/2026-07-22-knowledge-inventory.md.
Artifact 1 — stratified-sample parity replay DONE (2026-07-22), outcome owner-accepted. 83.2% OCR-channel agreement (28,456-card reproducible-channel subset); 373/41,586 (0.9%) unexplained divergences, 0/373 a wrong-printing vote — all conservative abstentions. Not literally "zero unexplained divergence," but ruled to satisfy the gate's soundness intent. Full outcome, methodology, and the owner-acceptance ruling: see §4 below.

Gate verdict: FIRED — §9 sequence complete end to end, true completion 2026-07-24 (the pilot --write + first consensus_recompute --apply landed 2026-07-23; five further corrective/completion passes — lexicon- gate retraction, marker reparse, artist-credit fill, calculator re-pass, and a second consensus_recompute closer, §14 — ran the same night into 2026-07-24 and are part of the same fire, not a new gate question). Every owner decision this gate needed is made (§3, §4); the deploy step is done: master a587000 (PR #345) went live in prod 2026-07-23T01:33Z — migration 0078 auto-applied, and §3 items 1–2's RESOLUTION_FLOOR_DPI/ EXCLUDED_RESOLVED_TAGS constants (PR #343) verified live in the running container (see §3, §12). Item 3 (deductive-backfill exclusion) was separately ruled NOT restored, already merged (§3). Artifact 1 (§4) is DONE and owner-accepted 2026-07-22 as satisfying the gate's soundness intent, even though the literal "zero unexplained divergence" bar wasn't hit (373 remained, all conservative abstentions; both root causes since fixed in code by merged PR #340). Artifact 1 itself is closed history, not a baseline to keep re-measuring against — see §8 for the new-data basis ratified 2026-07-23. The full §9 fire sequence (Bug-B write pass → retraction → Bug-A sample → the 4c pilot dry-run → owner sample audit → write → consensus_recompute --apply), amended 2026-07-23 per issue #347, is now COMPLETE: the Bug-B whole-DB reparse dry-run (§12), the zeroing steps (B(i) write, B(ii)+B(iii) retraction), the §9(c) Bug-A forced-escalation sample, the 4c pilot dry-run (§9(d)), the owner sample audit, the pilot --write (COMPLETE-BY-VERIFICATION, a documented execution-harness artifact — see §14), and consensus_recompute --apply (§9(e), §14) are all DONE, executed and DB-verified. The measurement of record for this fire is reports/2026-07-23-4c-pilot-dry-run.md's dry-run (pilot-dry-20260723T094518Z) plus the write's DB-verified counts (§14) — the write reproduced the dry-run's predicted counts exactly on both channels. §14 was then extended by five further passes (2026-07-23 evening into 2026-07-24) that are part of the same fire, not a new gate question: a lexicon-gate retraction, a marker reparse, an artist-credit fill, a calculator re-pass, and a second consensus_recompute closer — full per-pass table and topline end-state in §14. Two open items are carried forward past the fire: (1) the Bug-A full re-scan, deferred to post-pilot per the owner ruling in §9(c)/§13; (2) the pilot-era consensus_recompute (49,206 tag transitions, 2026-07-23T13:25:37Z) has no PilotRunLedger row — a structural, pre-ledger-convention gap, not a data-quality concern — see §14's "Pass-9 ledger gap" note.

3. Decision (a), resolved — the three MISSING constants

The knowledge-inventory sweep confirmed three pilot-era constants have no current home in Stage C/D, by direct grep/read, not inference:

  1. RESOLUTION_FLOOR_DPI = 200 — the pilot never fetched a card below this empirically-validated floor (dpi≤150 measurably degrades OCR yield). Stage C's cohort selection has no dpi condition at all. Highest-severity of the three: no downstream signal (ImageEvidence carries no dpi field) distinguishes a low-resolution extraction later.
  2. EXCLUDED_RESOLVED_TAGS = ["custom-art", "non-english"] — the pilot excluded cards already tagged custom-art/non-english (their printing-identification precondition is already falsified) from selection entirely. Stage D's _eligible_cards_queryset has no equivalent exclusion, and no other code path produces the same effect (checked directly against tag_consensus.py). Not previously tracked anywhere prior to this sweep — a genuinely new finding, not a known accepted gap.
  3. Deductive-backfill exclusion (.exclude(printing_tags__anonymous_id=DEDUCTIVE_BACKFILL_ANONYMOUS_ID)) — the pilot never re-voted a card the deductive backfill had already cast a vote for. Lower severity: only bites the narrow subset where backfill voted but the card is still UNRESOLVED (a card it fully resolved is already excluded by Stage D's own printing_tag_status=UNRESOLVED filter). Origin of this constant: DEDUCTIVE_BACKFILL_ANONYMOUS_ID = "deductive-backfill-v1" in ../MPCAutofill/cardpicker/deductive_backfill.py, run via the deductive_backfill_printing_tags management command. This is the same run that produced the 28,112 deduction-source CardPrintingTag votes in today's live pool (§6 below) — verified live: all 28,112 carry run_id=None, anonymous_id="deductive-backfill-v1", created_at between 2026-07-14T18:21:49Z and 2026-07-14T18:22:05Z. See journal/2026-07-14-deductive-printing-tag-backfill.md (gitignored, machine-local) for that run's own narrative. INTENTIONALLY NOT RESTORED (owner ruling, 2026-07-22), superseding an earlier same-day revision of this page that marked it "addressed in code": a read-only investigation of the 2026-07-14 backfill found it is pure name/metadata deduction (never phash/OCR — zero image inspection) whose votes check out sound (a 15-card sample all correct), and that excluding those cards would strand ~27,819 sound-but-UNRESOLVED cards outside Stage D for no protective benefit — re-evaluating them is safe under the human-backed consensus gate (agreement dedups, disagreement surfaces to human review). The pilot's own exclusion was a performance optimization (skip a card its weaker engines couldn't add to), not a soundness mechanism, so restoring it here would trade real coverage for a protection the vote-consensus layer already provides independently. See local_calculate_verdicts._eligible_cards_queryset's own docstring for the in-code record of this decision. These 28,112 deduction- source votes are NAME-based (the 2026-07-14 backfill, not phash/OCR- based), stay on the record as votes, and are invisible to Stage D's own calculation post-#341 — Stage D neither reads nor is blocked by them. Their consensus-layer weighting was a separate question, parked here as still-OPEN as of this section's earlier revisions — now RESOLVED (owner ruling, 2026-07-23): these votes carry ZERO consensus weight in every resolution computation, permanently, while the rows themselves stay on the record forever (raw tallies/display paths unaffected — see vote_consensus.DEDUCTIVE_BACKFILL_ANONYMOUS_ID's own docstring for the exact mechanism and the live-audit numbers behind the ruling, and theory.md's soundness section, §4, for the write-up). Re-scoped 2026-07-29 (owner clarification): that ruling zeroes THIS COHORT — the 28,112 rows this one run wrote, held out permanently as a measurement control — and does NOT disqualify name-matching deductive inference as a method, so a future run of deductive_backfill_printing_tags casts votes carrying the ordinary PRINTING_TAG_MACHINE_WEIGHT. The run_id=None fact recorded above is no longer true of these rows: migration 0097_freeze_deductive_backfill_zero_weight_cohort stamped them with vote_consensus.DEDUCTIVE_BACKFILL_ZERO_WEIGHT_RUN_ID, which is now what the zero-weight override matches on (together with source=deduction and the deductive-backfill calculator family). The created_at window above is what that migration selected on, and is not consulted at runtime by anything. This is a distinct decision from the Stage-D-exclusion question this whole numbered item is about (whether Stage D re-votes a card the backfill already touched, resolved NOT- RESTORED above) — that ruling is unchanged by this one.

None of these three are soundness violations — the human-backed consensus gate still applies to every vote Stage D casts regardless.

Owner ruling (2026-07-22T23:47Z): #1 and #2 are MUST-FIX. Sized against live data: 28 eligible cards fall below the dpi floor (0.016% of the 179,766-card eligible pool), 47 carry custom-art / 0 non-english (0.026%), zero overlap between the two, union 75 (0.042%) — zero live OCR votes rest on a sub-floor image. Fix: two one-line queryset excludes in _eligible_cards_queryset (dpi__lt=200, the custom-art/non-english tag exclusions) — merged to master in PR #343 and deployed 2026-07-23T01:33Z (§9(a) done — see §12): both constants verified live in the running container (RESOLUTION_FLOOR_DPI = 200, EXCLUDED_RESOLVED_TAGS = ['custom-art', 'non-english']). All three items above are now decided — none remain open.

Full detail, plus 3 lower grade "open items" that are separate from these 3 MISSING findings: reports/2026-07-22-knowledge-inventory.md.

4. Artifact 1 — parity-replay methodology and outcome (resolved 2026-07-22)

Artifact 1 was originally scoped as a live stratified-sample re-run against dpi=250, re-extracting evidence and comparing it to the pilot's outputs. That scoping was wrong and was replaced by this corrected methodology:

Dry-run diff of new local_calculate_verdicts verdicts vs. the pilot's recorded votes; NO Stage C re-extraction; pass bar = the DOCUMENTED "zero unexplained divergence", NOT a percentage threshold.

Concretely: run Stage D's calculator in dry-run mode against the ImageEvidence rows that already exist, diff its verdicts against the CardPrintingTag rows the pilot run (20260716T193408-6613a1a6) already cast, and account for every divergence — a match is not enough on its own; every disagreement must be explained (a known, reasoned architectural difference per the knowledge-inventory sweep) or the gate fails. No numeric pass threshold was invented or accepted anywhere in this process — "zero unexplained divergence" was the only documented bar; the outcome below did not literally hit it, which is why the owner ruling in this section exists.

Result

Baseline vs. method, stated explicitly. Pilot run 20260716T193408-6613a1a6 is the legacy multi-channel engine — OCR plus the local-phash-v1/local-fallback-v1 phash channels voting concurrently (why its 43,425 votes span only 41,586 distinct cards: some cards received votes from more than one channel). Stage D is OCR-only by design. This replay is therefore a cross-method verdict diff — the new OCR-only Stage D, computed against current ImageEvidence from the new Stage C extractions, diffed against the legacy multi-channel engine's recorded votes — not the new method validated against itself. The 83.2% figure below is OCR-channel agreement specifically; bucket d below (13,026) is exactly the phash/fallback channels Stage D doesn't run — an explained architectural difference, not a gap.

The replay ran 2026-07-22 as a read-only, full-cohort (not sampled) comparison against that pilot run's 43,425 votes / 41,586 distinct cards (source=ocr): each card's pilot vote was compared against what the current Stage D join-key calculator (calculate_join_key_verdict/_resolve_candidates_for_card, called directly) computes from its now-persisted ImageEvidence, bypassing _eligible_cards_queryset entirely. No writes.

Headline: 23,789/41,586 (57.2%) agree outright. Restricted to the 28,456 cards whose pilot vote actually used the OCR channel Stage D can reproduce (local-ocr-v1), agreement is 23,689/28,456 = 83.2%.

Divergence buckets (17,793 disagreements):

bucket count note
a. RESOLUTION_FLOOR_DPI 0 structurally 0 in this cohort — pilot's own filter already excluded these before ever voting
b. EXCLUDED_RESOLVED_TAGS 0 same — structurally pre-filtered
c. DEDUCTIVE_BACKFILL (§3 item 3) 0 same, doubly structural (deductive_backfill's own eligibility requires zero pre-existing votes)
d. pilot engine has no Stage D analogue 13,026 pilot matched via local-fallback-v1/local-phash-v1 only — both explicitly out of Stage D's scope per its own docstring/theory.md §7
e. Stage D's new veto layer (border/copyright-year mismatch) 4,394 correctly withholds matches on cards whose evidence genuinely disagrees with the real printing (spot-checked, not a classifier bug)
f. UNEXPLAINED (the gate criterion) 373 (0.9% of cohort) 2 identified root causes, 0/373 a wrong-printing vote — every one a conservative abstention. See theory.md §7c.

Buckets a/b/c reading exactly 0 is a property of this backward-looking cohort (the pilot's own eligibility filter already excluded any card that would trip them), not evidence the §3 constants don't matter going forward — that is the separate, forward-looking question §3 answers (all three items resolved 2026-07-22, per that section — items 1–2 MUST-FIX with a fix in flight, item 3 not restored).

Owner gate ruling (2026-07-22): soundness bar ACCEPTED

Not literally zero — 373 unexplained (0.9% of the 41,586-card cohort) — but every one is a conservative abstention (0/373 voted for a wrong printing). Ruling: the gate's intent (no confidently-wrong verdicts at scale) is satisfied — zero false-accept risk. The 373 were not treated as a fire blocker; both root causes were identified and fixed in the same session, code-only, in merged PR #340:

  • 155 "no-text" divergences — the 2026-07-21 OCR short-circuit skipped deeper tiers whenever both tier-1 OCR attempts were digit-free, conflating a blank/failed tier-1 read (a read failure) with a confident "no collector number here" finding. Narrowed to escalate whenever tier-1 comes back blank, exactly like a digit-bearing-but-unparseable read already did.
  • 218 is_no_match divergences (subset fixed) — a glued-token OCR parse failure: a single language-marker character glued onto the tail of a set-code token, e.g. card 41559 ("Verazol, the Split Current") parsing set_code="znre" from "znr" + an adjacent language-marker token's leading "e"; the real set is znr. Same family as PR #260's denominator/rarity-token glued-token guard.

Separately: the three §3 constants reading 0 divergence in this backward-looking replay is a structural artifact of the pilot's own pre-filtering (see "Result" above), not evidence of their forward impact — that sizing is tracked in §3, independent of this ruling.

Scope boundary — code fix vs. live cohort

PR #340's fixes are code-only. Realizing the benefit on the live 373-card cohort (or any other newly-affected cards) requires a targeted Stage C re-extraction of the affected card_ids — a gated prod write, queued behind the post-freeze deploy, and explicitly out of scope for / not run by PR #340. The 373-card cohort is bounded to the pilot-vote replay's own 41,586-card comparison set, not the full catalog — §10 sizes both root causes' actual catalog-wide footprint (17,531 cards / 284 rows respectively), which is the scope the re-extraction in §9(b) actually runs against.

Full source: issue #154's 2026-07-22 comments (the replay result and the owner-acceptance ruling) and PR #340 (the fix detail, verification, and scope-boundary note). Durable technical distillation of what the replay showed about the OCR channel and the conservative-abstention property: theory.md §7c.

5. Why Artifact 1 is a verdict-diff, not an evidence-diff

Pilot run 20260716T193408-6613a1a6 completed 2026-07-16/17 — this predates the ImageEvidence model entirely. Migration 0068 (the ImageEvidence substrate) wasn't applied to production until 2026-07-20 (see features/catalog-completion-plan.md's "Migration 0068 (only) live on production" entry). There is therefore no ImageEvidence row that existed at pilot time to diff against — Artifact 1 can only ever compare Stage D's verdicts (computed today, against today's ImageEvidence) to the pilot's recorded votes (CardPrintingTag rows, which do predate and survive the migration). This is the reason §4's methodology is a verdict-diff, not an evidence-diff, and why re-extracting Stage C evidence to match the pilot's original images would not close this gap even if it were done (the pilot's own transient-fetch images were never persisted — see CLAUDE.md's "Governing premise: we index, we do not store images").

6. Verified data snapshot (2026-07-22, live)

FIG-2 — Catalog funnel (shape only, no counts)

flowchart TD
    A["CATALOG"] --> B["PERCEPTUAL HASH<br/>extracted"]
    A -.-> A1{{"no phash yet"}}

    B --> C["IMAGE EVIDENCE EXTRACTED<br/>near-complete"]
    B -.-> B1{{"phash but no evidence yet"}}

    C --> D["CARRIES ≥1 PRINTING VOTE<br/>machine votes abundant"]
    C -.-> C1{{"evidence but no vote yet<br/>Stage D has not reached these"}}

    D --> E{"human-backed consensus gate<br/>weight ≥ 2 · share ≥ 0.6 · a human said so"}

    E -- "cleared" --> F(["RESOLVED PRINTING<br/>rare — by design"])
    E -- "held" --> G{{"UNRESOLVED<br/>machine votes alone can never clear this gate"}}
    E -- "ruled out" --> H(["NO MATCH"])

    G --> I["REVIEW QUEUE<br/>awaiting a human vote"]

    classDef stage fill:#24283b,stroke:#565f89,stroke-width:1px,color:#c0caf5
    classDef gate fill:#2f3549,stroke:#ff9e64,stroke-width:2px,color:#c0caf5
    classDef halt fill:#f7768e,stroke:#8c3d4e,stroke-width:2px,color:#1a1b26
    classDef excl fill:#7dcfff,stroke:#3f7f9c,stroke-width:2px,color:#1a1b26
    classDef done fill:#9ece6a,stroke:#5c7c3d,stroke-width:2px,color:#1a1b26

    class A,B,C,D,I stage
    class E gate
    class G halt
    class A1,B1,C1 excl
    class F,H done
Loading

Deliberately shape-only — no absolute counts in this diagram (owner ruling, 2026-07-25): every number behind each row above is real, live, and already published in the tables directly below this figure on this same page — repeating them inside the diagram itself would fork the numbers this page's own rule forbids ("don't restate gate status/ decisions elsewhere, link here") and would go stale the moment the next pass runs. What the shape says instead: extraction (phash → image evidence) is functionally complete; Stage D coverage (the "carries a vote" row) is the one place more machine capacity still helps; the human-backed consensus gate is a hard requirement, not a formality — machine votes alone can never clear it, so "held" is the gate working as designed, not a fault, and the review queue is where held cards wait for exactly the one thing that clears them: a human vote. Absolute counts return to this diagram once real user confirmations start accumulating in volume; until then, read them from the tables immediately below.

Catalog: 218,285 cards; 218,270 with a current ImageEvidence row.

Vote pool (CardPrintingTag / CardTagVote / CardArtistVote, live, grows continuously):

pool total by source
printing 101,105 ocr 72,938 / deduction 28,112 (all deductive_backfill, §3 item 3) / user 55
tag 61,334 ocr 61,294 / user 40
artist 7,137 ocr 7,131 / user 6
total 169,576

Pilot run 20260716T193408-6613a1a6 — three distinct numbers, not one flattened figure:

  • 165,980 candidates scanned (the pilot's own completion-log counter; not independently re-derivable from CardScanLog today — see reports/2026-07-22-knowledge-inventory.md's note on non-persisted run counters for why some legacy-engine in-memory stats can't be re-queried after the fact).
  • 43,425 votes cast (live CardPrintingTag.objects.filter(run_id=...) count — corrected 2026-07-22 from a previously-stated 43,426 in theory.md/catalog-completion-plan.md; the recorded PilotRunLedger ledger row still says votes_written=43426, an off-by-one against the live count with no documented retraction explaining the difference).
  • 41,586 distinct cards voted (live .values('card_id').distinct().count() — smaller than votes-cast because a card can receive votes from multiple engines, e.g. OCR and phash both voting on the same card).

Stage D run history (CardPrintingTag/PilotRunLedger, local_calculate_verdicts):

run_id votes written notes
staged-dryrun-20260721T0423Z 0 dry-run
staged-write-20260721T0434Z 0 (live, 2026-07-23) PilotRunLedger.votes_written records 8,925 — the pre-retraction figure; #258's retraction (2026-07-21) brought the live count to 8,825 (already documented in reports/2026-07-21-recovery-arc.md); the 2026-07-23 zeroing retraction (§13) then deleted 8,801 of those, leaving 24 — which B(i)'s write pass had already re-labeled off this run_id as corrected flips, so 0 remain attributed here live
staged2-0721 0 (live, 2026-07-23) was 70; all 70 deleted by the 2026-07-23 zeroing retraction (§13)
staged3-0721 0 (live, 2026-07-23) was 3,010; all 3,010 deleted by the 2026-07-23 zeroing retraction (§13)
staged4-0721 0 (live, 2026-07-23) was 999; all 999 deleted by the 2026-07-23 zeroing retraction (§13)
interim-peek-0722 0 dry-run
bugb-reparse-dry-20260723T014652Z (PilotRunLedger id 32) 0 dry-run, reparse_collector_evidence whole-DB Bug-B measurement, 197,938 considered — see §12
bugb-reparse-scoped-dry-20260723T020508Z (PilotRunLedger id 33) 0 dry-run, reparse_collector_evidence scoped to the 284-signature ID file — see §12
bugb-reparse-voted33-dry-20260723T0206Z (PilotRunLedger id 34) 0 dry-run, reparse_collector_evidence scoped to the 33 previously-voted Bug-B cards — see §12
bugb-write-dry-20260723T090258Z (PilotRunLedger id 35) 0 dry-run, pre-write confirmation for B(i) — see §13
20260723T090331-fdf5822b (PilotRunLedger id 36) dry-run, retract_stage_d_by_run_id pre-retraction preview (12,904 votes / 7,773 skips previewed) — see §13
bugb-write-20260723T0905Z (PilotRunLedger id 37) 236 B(i) live writereparse_collector_evidence --write, considered 285 / fields_fixed 285 / retracted 236 / gate_refused 0 — see §13
20260723T091446-35a1bde5 (PilotRunLedger id 38) B(ii)+B(iii) live retractionretract_stage_d_by_run_id --write: 12,880 votes deleted (+24 already flipped by B(i) = 12,904 total staged votes accounted for), 7,773 skips deleted, 20,653 cards resynced, 0 resolved-gate refusals — see §13
buga-sample-20260723T0927Z (PilotRunLedger id 39) §9(c) Bug-A samplerun_image_evidence_cohort, live, 300-card uniform sample (seed 20260723) of the 17,531-card blank-tier-1 pool, --no-shortcircuit — see §13
buga-sample-verdicts-dry-20260723T093321Z (PilotRunLedger id 40) 0 dry-run, reparse_collector_evidence verdict pass over the id-39 sample — 1 genuine match / 76 no-match / 223 skips — see §13
pilot-dry-20260723T094518Z (PilotRunLedger id 41, §9(d) dry-run) 0 dry-run, full eligible-pool Stage D pass — join-key 39,253 match/61,247 no_match, fallback 29,710 (read-only recomputation) — see reports/2026-07-23-4c-pilot-dry-run.md
pilot-write-20260723T1202Z (PilotRunLedger id 42, §9(d) write) 130,210 at write (48,151 join-key rows survive live, 2026-07-24 — 52,349 retracted by the lexicon-gate retraction, §14; fallback 29,710 unaffected) live write — join-key 100,500 (39,253 match/61,247 no_match) + fallback 29,710 match (still 29,710 live, never retracted), exact match to the dry-run's prediction; ledger row reads status=failed/votes_written=None, a documented execution-harness artifact (client timeout severed the executor's stdout mid-run, not the write) — see §14/data/2026-07-23-pilot-write-and-recompute.md
20260724T001154-d3986cfc (post-fire calculator re-pass write) 1 §14 "Five further passes" writelocal_calculate_verdicts re-run over the lexicon-gate-retracted pool; 52,348 unknown-set-code + 16 no-evidence skips, 117,442 to-review — see §14

Stage C run history (ImageEvidence.run_id, current last-writer row count per run — updated 2026-07-27 after the full-catalog re-extraction):

run_id date rows
pass-pilot-20260725 2026-07-25 100
stage-e-stream-20260725T233633221123Z 2026-07-25 22
stage-e-stream-20260725T233633687349Z 2026-07-25 23
pass-full-20260725 2026-07-25/26 194,831
pass-full-20260725-r2 2026-07-25/26 23,132
Total 218,108

The 2026-07-25 full-catalog re-extraction (pass-full-20260725 + pass-full-20260725-r2) superseded all prior Stage C last-writer rows — the original canaries (stagec-canary-20260720T1659Z, stagec-canary-decoupled-20260720T235127Z), ntx-0721, stagec-20k-20260721T0227Z, and stagec-remainder-0721 no longer appear as last-writer for any ImageEvidence row. See reports/2026-07-20-decoupled-canary-confirm.md and reports/2026-07-21-stagec-20k-extraction.md for the original canary/20k narrative detail, and reports/2026-07-26-stagec-full-catalog-completion.md for the full-catalog completion record. None of the Stage C runs have a PilotRunLedger row of their own — see §11 for what that does and doesn't affect.

ntx-0721 is NOT a pilot — stated explicitly because this number has been misremembered before: ntx-0721 is a Stage C extraction run (22,899 CardScanLog rows, a no-text cohort re-extraction — see reports/2026-07-21-recovery-arc.md). Verified live: 0 votes cast under run_id='ntx-0721' across all three vote tables (CardPrintingTag/CardTagVote/CardArtistVote). There is no 23,000-vote pilot; "23,111 votes" was a transient same-day pool snapshot at some earlier point, never a pilot run of any kind, and should not be cited as a baseline.

printing_tag_status: 218,281 unresolved, 3 resolved, 1 no_match.

consensus_impact_report dry-run (2026-07-22, --sample-limit 20, zero writes): printing 92,368 pairs checked / 0 transitions; artist 7,130 pairs / 0 transitions; tag 61,328 pairs, with 49,207 None→UNRESOLVED materializations sized as pending a separately owner-gated recompute pass. "Zero transitions" on printing/artist means the ratified resolver's answer already matched every currently-persisted status exactly, at that date. Superseded by the real apply, and then superseded again: this dry-run's tag/printing sizing is now closed history — consensus_recompute --apply ran 2026-07-23T13:25:37Z and materialized 49,206 of the sized 49,207 tag transitions (organic interim resolution accounts for the one-fewer figure) plus one additional printing resolution (3 → 4); a second consensus_recompute closer then ran 2026-07-24T01:14:48Z and flipped that same printing card back (resolved→unresolved), so the live resolved count is 3, not 4 — see §14's "Five further passes" and "Pass-9 ledger gap" for the complete before/after.

7. Chain — where each underlying fact actually lives

This page owns the gate's status and its decisions above; every fact below is the linked source, not duplicated here.

  • theory.md — the formal decoding model, the false-accept bound, and §7's stage-by-stage Stage D composition with error terms. Keeps the pilot's own per-engine breakdown and calibration numbers.
  • identification-pipeline.md — plain-language walkthrough of the same Stage C/D pipeline, stage by stage, for a reader who wants the mechanics without the formal model.
  • reports/2026-07-22-knowledge-inventory.md — Artifact 2's full constant-by-constant inventory table (SAME / CHANGED / MISSING / superseded-by-architecture / open item), the source for §3 above.
  • features/catalog-completion-plan.md — the six-part catalog-completion plan; the Stage D precondition wording (§1 above) and the Stage C/D build detail live there, not here.
  • data/2026-07-22-pipeline-snapshot.md (+ its .json sibling) — the dated raw-data record this page's §6 numbers were re-verified against; re-query before trusting any live-pool number more than an hour or two old, per that file's own provenance notes.
  • reports/2026-07-21-recovery-arc.md — the staged-write-20260721T0434Z 8,925→8,825 retraction (#258) cited in §6's Stage D table.
  • MPCAutofill/cardpicker/deductive_backfill.py — the module behind the 28,112 deduction-source printing votes and §3 item 3's MISSING exclusion.
  • GitHub issue #154 — the artifact-1 parity-replay result and the owner-acceptance ruling comments (2026-07-22), the source for §4's numbers.
  • PR #340 (merged) — the code fix for both of §4's root causes; the live 373-card cohort still needs the gated re-extraction described there.
  • PR #341 (merged) — §3 item 3's non-restoration rationale and code removal.
  • PR #343 (merged) — §3 items 1–2's RESOLUTION_FLOOR_DPI/EXCLUDED_RESOLVED_TAGS queryset-exclude fix, merged to master and deployed via PR #345.
  • PR #345 (merged) — the §9(a) deploy commit (master a587000); also makes run_image_evidence_cohort self-recording via PilotRunLedger (unrelated to §9(a) itself, ledger-tracking-only, see §11).
  • issue #347 — the Tron-reviewed pre-pilot machine-vote zeroing plan that amended §9 2026-07-23 (the retraction step, the Tron corrections, the consensus-safety verification cited in §9/§12).
  • reports/2026-07-23-4c-pilot-dry-run.md — the §9(d) measurement of record: the full eligible-pool Stage D dry-run's per-channel counters and the read-only fallback-recovery methodology, cited in §9(d)/§14.
  • data/2026-07-23-pilot-write-and-recompute.md — the §9(d)/§9(e) fire-sequence closing steps: the pilot --write's DB-verified vote counts, the status=failed execution-harness artifact explanation, and consensus_recompute --apply's outcome, the source for §14 above.
  • data/2026-07-23-bugb-reparse-dryruns.md — the §9(b)/§12 Bug-B whole-DB reparse dry-run report + resource metrics, keyed by run_id.
  • data/2026-07-23-zeroing-and-buga-sample.md — the §9 B(i)/B(ii)+B(iii)/(c)/§13 zeroing-execution and Bug-A sample report + resource metrics, keyed by run_id.
  • MPCAutofill/cardpicker/management/commands/consensus_recompute.py (PR #336, merged) — the --apply command §9(e)'s materialization step runs (STRICTLY LAST — see §9's Tron correction).

8. New-data basis (owner-ratified 2026-07-23)

The 2026-07-22 full-catalog Stage C sweep (§6's stagec-remainder-0721 main leg, plus its two preceding canary/20k runs) is the epistemic foundation for everything this gate measures going forward. Two consequences, both ratified in-session 2026-07-23:

  1. Artifact 1 (§4)'s legacy-pilot comparison is CLOSED HISTORY. Pilot run 20260716T193408-6613a1a6 served its purpose as a cross-method corroboration point (§4, §5, theory.md §7c) and stays on the record as that, permanently — but it does not get re-run or re-compared against as new data lands. It is not the baseline going forward.
  2. The new system's own pilot is DERIVED from the new data, not inherited from the old one. The upcoming full-pool Stage D dry-run (§9 step (d), the "4c pilot") IS the pilot / measurement of record from this point on — its own statistics (match/no-match/abstain counts, per-channel join-key-vs-fallback breakdown) become the cited figures for any future soundness argument about this pipeline, not a diff back against the legacy engine.

Scope note: this is a basis change, not a retraction. Every legacy (source=ocr, pilot-era) and deduction (source=deduction, deductive_backfill) vote already in the DB stays exactly as-is — no retraction, no source filter applied to either, in this ruling or any prior one. They continue to count as live, valid votes toward consensus; they simply stop being cited as the validation baseline for new soundness claims.

9. Fire sequence (owner-ratified 2026-07-23, amended 2026-07-23 per #347)

The full gated sequence this gate's verdict (§2) is blocked on, in order. Amended same-day by issue #347 (Tron-reviewed pre-pilot machine-vote zeroing plan) to insert an explicit retraction step ahead of the pilot dry-run — rationale: the ratified new-data basis (§8) requires the 4c pilot to cast 100% of fresh machine votes against the corrected evidence/parser (a new-data-basis requirement, not a data-quality complaint about the staged votes themselves), and consensus safety for doing so was independently verified (all 12,904 staged cards UNRESOLVED before and after, zero resolved-card overlap, zero user-visible delta during the no-vote window, zero ES writes from retraction).

(a) Deploy master — DONE 2026-07-23T01:33Z. PR #345 (master a587000) put PR #343's §3 items 1–2 fix (and PR #341's item 3 non-restoration) live in prod; migration 0078 auto-applied; #343's constants verified live in the running container (§3, §12).

(b) Bug-B whole-DB reparse dry-run — DONE (§12). The fixed parser applied offline across all 197,938 stored raw texts confirmed the existing 285-row prediction (§10) exactly: 284 glued-marker guard rows plus 1 unrelated improvement (card 62354). Full detail, resource metrics, and the scoped-33 sub-run: data/2026-07-23-bugb-reparse-dryruns.md.

B(i) Bug-B write pass — DONE 2026-07-23 (§13). reparse_collector_evidence --write (run bugb-write-20260723T0905Z) against the regenerated 284-signature ID file (the exact cohort §12's dry-run confirmed), NOT the broader parser-bug regex selector (a 553-card different, wider cohort — not substituted). Patch requirement (persist corrected parse fields unconditionally) verified in the outcome: considered=285, fields_fixed=285 (all 285, unconditionally — the 49-row gap is closed), retracted=236 (votes actually flipped), gate_refused=0.

B(ii)+B(iii) Retraction — one invocation, DONE 2026-07-23 (§13). The single-purpose retract_stage_d_by_run_id command (run 20260723T091446-35a1bde5), scoped to anonymous_id=stage-d-join-key-v1 AND the four staged run ids, deleted:

  • 12,880 CardPrintingTag votes (staged-write-20260721T0434Z / staged2-0721 / staged3-0721 / staged4-0721) plus the 24 votes B(i) had already flipped to genuine matches within the same target cohort = all 12,904 originally-staged votes accounted for, matching §6's table exactly (8,825 + 70 + 3,010 + 999 = 12,904; a pre-retraction dry-run preview, 20260723T090331-fdf5822b, confirmed the full 12,904 before B(i)'s write ran). This task's own independent pre-flight check (2026-07-23T09:41Z, ahead of launching (d)) confirmed the same end state: 0 CardPrintingTag rows remain with anonymous_id=stage-d-join-key-v1.
  • 7,773 non-rescannable CardScanLog stale skips from the same four runs (per-run: 7,187 / 14 / 19 / 553), exactly as predicted. (The much larger no-evidence skip population from these same runs was deliberately excluded — it is rescannable via the pilot's own native resume-filter logic, not a retraction target.)

Per-card resolve_printing() safety gate applied (same conservative card-level check as reparse_collector_evidence, §3 item 3's docstring) — 0 skipped_resolved_gate refusals, confirming zero resolved-card overlap. Verified untouched: user votes (55), deduction votes (28,112 — §3 item 3, still 28,112 live), legacy pilot votes (43,425, still 43,425 live), all tag/artist votes, the 3 resolved

  • 1 no_match cards (still 3 resolved live). 20,653 cards resynced via resolve_and_persist_printing()Tron correction still applies: that function casts no votes itself, it only recomputes/persists printing state for the card just retracted. Verified end-state (2026-07-23, live): CardPrintingTag rows with anonymous_id='stage-d-join-key-v1' = 0; non-rescannable CardScanLog skips for the 4 target runs = 0. Full run report: data/2026-07-23-zeroing-and-buga-sample.md.

(c) Bug-A forced-escalation SAMPLE — DONE 2026-07-23 (§13). 300 cards uniform-randomly sampled (seed 20260723, --no-shortcircuit) from the 17,531-card blank-tier-1 pool (§10), run buga-sample-20260723T0927Z (extraction, 85.4s) + buga-sample-verdicts-dry-20260723T093321Z (verdict dry-run). Funnel: 300 fetched → 78 non-blank OCR text (26.0%) → 78 parsed numbers → 65 set codes → 76 no-match votes → 1 genuine match (0.33%, card 122326 "Ephemerate" sketch variant → STA 68, spot-checked correct) → 223 skips. Wilson 95% extrapolation to the full 17,531-card pool: ~58 genuine matches [CI 10–327], qualitatively low-end likely (spot-checks show OCR noise dominating the non-blank yield); full re-scan cost estimated ~83–104 minutes.

Owner ruling (2026-07-23): Bug-A full re-scan DEFERRED to post-pilot, gap tracked not dropped: (a) the 17,531-card signature query regenerates on demand (fetch_ok=True, empty collector number, blank/whitespace raw text, excluding ntx-0721; 17,531 at 2026-07-23T09:19Z); (b) the pilot's own skip counters (§9(d)) will surface the blank-evidence abstentions so the gap stays visible; (c) any future re-scan MUST include a state-clear step first — this sample's no-text skips are non-rescannable (unlike no-evidence skips), so a post-pilot re-scan recipe is: re-extract with --no-shortcircuit → clear stale skip state via the reparse path → run a follow-up scoped Stage D pass to vote. Documented here as the recipe so it is not re-derived. Full run report, funnel table, and the recipe's full rationale: data/2026-07-23-zeroing-and-buga-sample.md.

(d) Stage D dry-run over the full eligible pool = THE PILOT — DONE 2026-07-23 (§8), run_id pilot-dry-20260723T094518Z, git_sha 42a09b3c794f7cf8aca5eb1ca2d4f6cdaa2895a6, eligible pool 200,366 (post- zeroing, live-verified immediately pre-run). Full statistics, the per-channel join-key-vs-fallback breakdown, and the read-only fallback recomputation methodology (the actual command reports fallback considered=0/votes=0 for a structural reason — see below — not a data gap): reports/2026-07-23-4c-pilot-dry-run.md. Structural finding: _fallback_eligible_cards_queryset sources its "join-key found no confident hit" population from PERSISTED CardPrintingTag/CardScanLog rows — which a same-invocation DRY RUN never writes — so local_calculate_verdicts (no --write) can concretely never produce a nonzero fallback count in one pass, regardless of flags; this is inherent to the current implementation, not fixable by a different invocation. The fallback channel's real first-ever numbers were recovered via a bounded, zero-write diagnostic that re-derives the join-key no-hit population in-memory using the same production functions, then calls the actual calculate_fallback_verdict per eligible card — see the linked report for the full methodology and results. Followed by an owner SAMPLE AUDIT — DONE (150 uniformly sampled verdicts, seed 20260723150) — sheet written machine-local only (not committed; see the linked report for its exact path and provenance), then --write — DONE 2026-07-23 (run pilot-write-20260723T1202Z, PilotRunLedger id 42, COMPLETE-BY-VERIFICATION — both channels' vote counts DB-verified to match the dry-run's prediction exactly, 130,210 total votes; full outcome and the status=failed execution-harness artifact explanation: data/2026-07-23-pilot-write-and-recompute.md).

(e) consensus_recompute --apply — STRICTLY LAST — DONE 2026-07-23T13:25:37Z. Materialized the pending None→UNRESOLVED tag transitions sized in §6's consensus_impact_report dry-run (49,206 of the 49,207 sized there — one fewer, organic interim resolution, not a discrepancy), via the command in consensus_recompute.py (PR #336). Exit code 0. Tron correction to an earlier "tag-orthogonal" claim: consensus_recompute recomputes printing + artist + tag state via the real resolver paths, not tag state alone — its safety in this sequence comes from idempotence plus running strictly last, not from any orthogonality to the printing-layer steps above; it was run strictly last here too. Printing resolution moved 3 → 4 resolved cards (DB-verified); artist (7,130 pairs) had zero transitions. Full outcome: data/2026-07-23-pilot-write-and-recompute.md. Note: this command predates the PilotRunLedger self-recording convention — no ledger row exists for it, a small follow-up flagged beside the already-tracked §11 Stage C run-identity gap. Re-confirmed 2026-07-24 by an independent DB-only audit (exhaustive PilotRunLedger scan, 64 rows, no gap-filling row found for this invocation): the gap is real and structural, not an omission — see §14's "Pass-9 ledger gap" note for the full corroboration attempt.

Order was load-bearing and followed exactly: (a) → (b)/B(i) → B(ii)+B(iii) → (c) → (d) → sample audit → --write → (e) consensus_recompute --apply strictly last. The full §9 sequence (as originally ratified) is COMPLETE. §14 records five further corrective/completion passes that ran after (e), extending true completion to 2026-07-24 — not a reopening of §9's own ratified order, which stands as executed.

Every run in this sequence gets its own report plus resource metrics (RSS/IO/CPU, per-card cost) committed under docs/data/, keyed by that run's run_id, and added as a new row to §6's Stage C/Stage D run history tables (not a separate table).

10. #340 root-cause footprint sizing (measured 2026-07-23T00:16–00:24Z, live, read-only)

Both root causes fixed by merged PR #340 (§4) were sized against the full catalog, not just the 373-card replay cohort (§4's cohort is bounded to the pilot-vote comparison set, 41,586 cards; this sizing covers the whole 218,285-card catalog).

Bug A (OCR short-circuit over-skip). 17,531 cards catalog-wide carry the blank-tier-1 signature — a necessary-condition proxy for the bug (matches the pattern the fix targets; not a guarantee every one flips to a match), excluding ntx-0721's cohort (already force-escalated by that run, §6). 15,948 of the 17,531 (91%) are Stage-D-eligible. Expected genuine-match conversion is LOW: ntx-0721's own base rate for this same signature is ~0.2% — the reason §9(c) samples first rather than committing to a full re-scan.

Bug B (glued language-marker set-code). 284 rows carry the signature, systemic across ≥7 distinct set codes (not one bad batch). 212 of the 284 are Stage-D-eligible; 33 already carry staged-run votes, and every one of those 33 is is_no_match=True with 0 wrong-printing votes — direct empirical confirmation, at this wider scale, of the same conservative-abstention direction §4/theory.md §7c already established at the 373-card replay scale. The whole-DB reparse (§9(b)) is pre-sized, not estimated: offline re-parsing all 197,938 stored raw texts with the fixed parser diffs exactly 284 guard rows plus 1 unrelated improvement (card 62354).

11. Stage C run-identity note (verified 2026-07-23)

Stage C runs are not tracked in PilotRunLedger — verified directly: zero rows for run_image_evidence_cohort. Completion-log counters like stagec-remainder-0721's (141,369 processed / 57,725 short-circuited / 492 fetch failures, §6) are log-only and not DB-re-derivable after the fact, the same caveat §6 already states for the older pilot's own 165,980-candidate counter. This does not affect the evidence itself: ImageEvidence rows are fully persisted regardless of ledger tracking — 197,469 current rows verified at this check, window 2026-07-21T19:29:24Z–2026-07-22T16:45:43Z (one row's drift against §6's 197,470 count for the same run is two snapshots taken at different times, not a data-integrity concern — re-query before trusting either number to the exact row). A ledger fix (Stage C runs writing their own PilotRunLedger rows, closing this gap) is in flight as a separate code PR, not yet merged as of this note.

12. Deploy + Bug-B whole-DB reparse dry-run outcome (2026-07-23)

Deploy (§9(a)): master a587000 (PR #345) merged and deployed to prod 2026-07-23T01:33Z. Migration 0078 auto-applied — verified live (manage.py showmigrations cardpicker shows [X] 0078_pilotrunledger_counters as the last applied row). §3 items 1–2's constants verified live in the running container: RESOLUTION_FLOOR_DPI = 200, EXCLUDED_RESOLVED_TAGS = ['custom-art', 'non-english'].

Bug-B whole-DB reparse dry-run (§9(b)): three reparse_collector_evidence dry-runs, PilotRunLedger ids 32–34 (§6's Stage D run history table), all dry_run=True / votes_written=0 — verified live against PilotRunLedger directly.

  • bugb-reparse-dry-20260723T014652Z (whole-DB, 197,938 candidates): the offline 285-changed-row prediction already recorded in §10 verified exactly — 284 glued-marker guard rows plus 1 unrelated improvement (card 62354). The command's own internal changed=162,866 counter is a different, broader metric (see the command's own module docstring: it compares a fresh re-parse against each card's currently-RECORDED join-key verdict, not against the specific glued-marker signature) — arithmetic cross-check: no_evidence=0 + no_prior_join_key_state=16,253 + unchanged=18,819 + changed=162,866 = 197,938, matching considered exactly. The 162,866 figure is explained as stale no-evidence skips from the 2026-07-21 staged passes predating full Stage C evidence coverage, handled natively by the pilot's own rescannable-skip resume logic — not Bug-B blast radius, and explicitly not a write target for B(i)/B(ii+iii).
  • bugb-reparse-scoped-dry-20260723T020508Z and bugb-reparse-voted33-dry-20260723T0206Z: scoped confirmation runs, completed in 2.94s and 1.69s respectively. The voted33 run found 24/33 previously-voted Bug-B cards flip false-no-match→genuine match under the fixed parser — ground-truth confirmation of those 24 is deferred to the pilot's owner sample audit (§9(d)), not asserted here.

Full run report, resource metrics (RSS/IO/CPU, per-card cost), and provenance: data/2026-07-23-bugb-reparse-dryruns.md. This also doubles as the first runtime calibration for the §9(d) 4c pilot dry-run — same verdict-computation code path.

Fallback channel (§9(d)): as of this section's original 2026-07-23T01:33Z-onward timestamp, verified live 0 CardPrintingTag rows carried anonymous_id='stage-d-fallback-v1' — the fallback channel had never cast a vote in production. §9(d)'s own entry now has the full outcome (still zero live votes — the pilot is DRY-RUN only — but the channel's own would-cast statistics are now known, via a read-only recomputation; see that entry and its linked report).

Remaining §9 steps as of this section's timestamp (2026-07-23T01:33Z onward): B(i), B(ii)+B(iii), (c), (d), sample audit, --write, (e) were all NOT YET RUN. §13 below carries a later same-day update: B(i), B(ii)+B(iii), (c), and (d) (the 4c pilot dry-run) are now all DONE (§9(d)'s own entry carries the full outcome); sample audit / --write / (e) remain NOT YET RUN.

13. Zeroing execution + Bug-A sample outcome (2026-07-23T09:0x–09:3xZ)

B(i) Bug-B write pass — DONE. Run bugb-write-20260723T0905Z (PilotRunLedger id 37, preceded by dry-run confirmation id 35): considered=285, fields_fixed=285 (all persisted unconditionally, closing the 49-row gap), retracted=236 (votes actually flipped), gate_refused=0.

B(ii)+B(iii) retraction — DONE. Run 20260723T091446-35a1bde5 (PilotRunLedger id 38, preceded by a dry-run preview id 36): 12,880 CardPrintingTag votes deleted + the 24 votes B(i) had already flipped = all 12,904 originally-staged votes accounted for exactly; 7,773 non-rescannable CardScanLog skips deleted (per-run 7,187 / 14 / 19 / 553, matching the prediction exactly); skipped_resolved_gate=0; 20,653 cards resynced via resolve_and_persist_printing(). Verified end-state (live, this section): stage-d-join-key-v1 votes = 0; non-rescannable join-key skips (4 target runs) = 0; deduction votes = 28,112 (intact); legacy pilot votes = 43,425 (intact); resolved cards = 3 (intact); eligible Stage D pool = 200,366 (relayed, not re-queried this pass to avoid a full-catalog scan against the concurrently-running §9(d) pilot).

§9(c) Bug-A forced-escalation sample — DONE. Runs buga-sample-20260723T0927Z (PilotRunLedger id 39, 300-card uniform-random extraction, seed 20260723, --no-shortcircuit, 85.4s, 7 workers) + buga-sample-verdicts-dry-20260723T093321Z (PilotRunLedger id 40, verdict dry-run). Funnel over the 300-card sample drawn from the 17,531-card blank-tier-1 pool (§10): 300 fetched → 78 non-blank OCR text (26.0%) → 78 parsed numbers → 65 set codes → 76 no-match votes → 1 genuine match (0.33%) → 223 skips. The 1 match: card 122326 ("Ephemerate", Sketch Yumiko variant) → STA 68 — spot-checked and confirmed correct. Wilson 95% extrapolation to the full pool: ~58 genuine matches [CI 10–327], qualitatively low-end likely (OCR noise dominates the non-blank yield in spot-checks); full re-scan estimated at ~83–104 minutes.

Owner ruling (2026-07-23): Bug-A full re-scan DEFERRED to post-pilot. Gap tracked, not dropped: (a) the signature query regenerates on demand (17,531 cards at 2026-07-23T09:19Z, same definition as §10); (b) the §9(d) pilot's own skip counters will surface the blank-evidence abstentions, keeping the gap visible without a separate tracker; (c) the post-pilot re-scan procedure must include a state-clear step — this sample's no-text skips are non-rescannable (unlike the no-evidence skip reason handled natively elsewhere in this sequence), so the documented recipe is: re-extract with --no-shortcircuit over the target cohort → clear stale skip state via the reparse path (reparse_collector_evidence, same mechanism B(i) used) → run a follow-up scoped Stage D pass (local_calculate_verdicts) to actually cast votes. Recorded here as the standing recipe so it is not re-derived next time.

§9(d) the 4c pilot dry-run — DONE, started immediately after (c) above completed. Full run_id, statistics, per-channel breakdown, and the read-only fallback-channel recovery methodology: see §9(d)'s own entry above and reports/2026-07-23-4c-pilot-dry-run.md.

Full run reports, resource metrics, and per-run PilotRunLedger counters for every run in this section: data/2026-07-23-zeroing-and-buga-sample.md.

B(i)/B(ii)+B(iii)/(c)/(d) are now all DONE (§9's own entries carry the live-verified detail and ledger ids); sample audit, --write, and (e) remained NOT YET RUN as of this page's edit at the time this section was written. §14 below carries the closing update: all three completed the same day, 2026-07-23 — the full §9 sequence is COMPLETE.

14. Fire sequence complete — full outcome (2026-07-23/24)

The owner sample audit, the pilot --write, and the first consensus_recompute --apply — the three steps §12/§13 left NOT YET RUN — all completed 2026-07-23. The §9 fire sequence (as originally ratified) is COMPLETE end to end and the gate is FIRED (§2). Full detail, DB-verified counters, and the status=failed artifact explanation: data/2026-07-23-pilot-write-and-recompute.md.

This section was extended 2026-07-24 with five further corrective/completion passes that ran the same night into the next day (a lexicon-gate retraction, a marker reparse, an artist-credit fill, a calculator re-pass, and a second consensus_recompute closer) — see "Five further passes" below, and the full per-pass table immediately below. All figures in this extension are DB-verified live against mpcautofill_django/Postgres, queried 2026-07-24, unless marked "docs-sourced" or "GAP."

Full per-pass table (chronological, all times UTC; condensed — dry-run rehearsals with 0 rows written are collapsed into their write row's note where the write immediately follows):

# pass run_id(s) dry/write wall time rows written notes
1 Bug-B reparse write bugb-write-20260723T0905Z (id 37) write 5.4s 236 vote-affecting rows preceded by 3 dry rehearsals (ids 33–35), see §12/§13
1r Bug-B retraction 20260723T091446-35a1bde5 (id 38) write 189s votes_deleted 12,880, cards_resynced 20,653 see §13
2 Bug-A forced-escalation sample buga-sample-20260723T0927Z write (extraction) 85.4s 300 sampled, 1 genuine match see §9(c)/§13
3 4c pilot dry-run pilot-dry-20260723T094518Z (id 41) dry 7m37s 0 see §9(d)/§14 above
4 4c pilot write pilot-write-20260723T1202Z (id 42) write 59m23s 130,210 (77,861 survive live today; 52,349 later retracted by pass 5) see §14 above
5 Lexicon-gate retraction 20260723T184334-2260859d write 675.8s retracted 52,349, unchanged 8,898 see "Five further passes" below
6 Marker reparse write 20260723T184318-6e0c73d9 write 13.4s 2,506 fields flipped (not votes) see "Five further passes" below
7 Artist fill write 20260723T213608-5479a0f0 write 51m2s 131,020 evidence fields filled see "Five further passes" below
8 Calculator re-pass write 20260724T001154-d3986cfc write 52m41s 1 vote; 52,348+16 skips, 117,442 to-review see "Five further passes" below
9 Consensus recompute #1 (pilot-era) no ledger row — GAP not verifiable 49,206 tag transitions (docs-sourced) see "Pass-9 ledger gap" below
10 Consensus recompute #2 (closer) 20260724T011448-32e08cc8 apply 383.9s 101,715 total (tag 0, artist 0, printing 1) see "Five further passes" below

Full per-pass detail (every dry-run rehearsal, exact considered/skip counts, and the reconciliation arithmetic) is DB-verified but not duplicated in full here — this table is the condensed record; the prose subsections below carry the load-bearing detail for passes 5–10.

Pilot --write — run pilot-write-20260723T1202Z, PilotRunLedger id 42, started_at 2026-07-23T12:06:53Z, finished_at 2026-07-23T13:06:16Z. DB-verified vote counts match the §9(d) dry-run's prediction exactly on both channels: join-key 100,500 (39,253 match / 61,247 no_match, same as the dry-run's would_cast figures), fallback 29,710 match (the fallback channel's first production execution — its votes did not exist before this run). Grand total 130,210 CardPrintingTag rows written.

The status=failed / votes_written=None artifact: the ledger row itself reads as a failed run with no vote count. This is NOT a failed or partial write — it is a documented execution-harness artifact. The authorization executor enforced a 1800s client-side timeout on the launching connection; that timeout severed the executor's own stdout stream at ~12:36Z, not the in-container process, which kept running to completion (finished_at 13:06:16Z, ~30 minutes past the timeout boundary) and hit a terminal-phase exception in its own summary/ledger-update code, after the last vote write, not during it. The DB-verified vote counts above — an exact match to the dry-run's prediction on both channels — are the direct evidence the writes themselves completed cleanly. Recorded here as COMPLETE-BY-VERIFICATION, not as a failed fire, so this row is never misread as one at a glance.

consensus_recompute --apply — executed 2026-07-23T13:25:37Z, exit code 0. Predates the PilotRunLedger self-recording convention (no ledger row exists for it — confirmed live, highest id remains 42) — this was flagged beside the already-tracked §11 Stage C run-identity gap as a "command predates the ledger convention" instance rather than a broken command, and is now resolved going forward: consensus_recompute gained its own self-recording ledger row (RUNNING at start, COMPLETED/FAILED at end, per-family pairs_checked/rows_written/transitions counters) as part of the command-lifecycle hardening pass, matching every other Stage C/D pilot command's lifecycle. This specific 2026-07-23T13:25:37Z run predates that fix and has no ledger row, immutably. Outcome: artist 7,130 pairs checked, 0 transitions; tag 61,329 pairs checked, 49,206 None→UNRESOLVED materializations (one fewer than the 49,207 sized in §6's earlier dry-run — organic interim resolution in the gap between measurement and apply, not a discrepancy); printing resolved 3 → 4 (DB-verified at the time: Card.printing_tag_status distribution = unresolved 218,309 / resolved 4 / no_match 1). This 4 was itself provisional — a second closer pass the following night flipped it back to 3; see "Five further passes" and "Topline end-state" below for the corrected live figure.

Pass-9 ledger gap (re-confirmed 2026-07-24): the 49,206-transition figure above remains docs-sourced, not independently re-derivable from any ledger row — an exhaustive PilotRunLedger scan (64 rows, ids 1–65 minus an unrelated gap at id 10) found no row for this invocation, and Card.tag_vote_statuses is a mutable JSONField overwritten by every later recompute, not append-only, so the exact per-run delta can't be reconstructed after the fact either. Best available corroboration (not proof): the live tag_vote_statuses aggregate is 61,332 unresolved + 2 resolved_apply = 61,334 (card, tag) pairs, and the 2026-07-24 closer pass's own dry-run rows report pairs_checked: 61,330 for the tag family with zero transitions found — consistent with a large one-time materialization having already happened before the closer ran. This is a structural, pre-ledger-convention gap (the command has since been hardened to always self-record), not a data-quality concern, and nothing further closes it.

Five further passes (2026-07-23 evening → 2026-07-24)

Five more passes ran after the original §9 sequence closed, all DB-verified live 2026-07-24. None of these were part of the ratified §9 order — each addressed a downstream correction or the pilot's own follow-on work, and none re-opens the gate question.

  • Lexicon-gate retraction — dry lexgate-dry-20260723T1835Z (18:35:54→18:38:56, 181.9s), write 20260723T184334-2260859d (18:43:34→18:54:50, 675.8s, 77.5 rows/s). Scope hash 7662d0b34e17d2a6, considered 61,247: changed 52,349, retracted 52,349, unchanged 8,898. Retracted the 52,349 stale join-key votes the pilot --write (pass above) had cast against evidence a corrected lexicon gate has since superseded, rather than recasting them itself. This is the other half of the pilot --write reconciliation math already noted in §6: live join-key rows under pilot-write-20260723T1202Z today = 48,151 (down from the 100,500 cast), and 48,151 + 52,349 = 100,500 exactly.
  • Marker reparse — two dry rehearsals (20260723T164303-f8e07e2b, superseded, considered 132,674, would-flip 2,508; markerdry-20260723T1844Z, final, considered 133,354, would-flip 2,506 — the pool grew ~680 rows between them from concurrent run_image_evidence_cohort crash-drill passes), then write 20260723T184318-6e0c73d9 (18:43:18→18:43:31, 13.4s, 186.7 flips/s): 2,506 rows flipped legal_line_proxy_marker_detected false→true. Field-level correction, casts no votes.
  • Artist fill — dry 20260723T204858-40e9408a (20:48:58→21:35:07, 46m10s), write 20260723T213608-5479a0f0 (21:36:08→22:27:10, 51m2s, 42.8 fills/s). backfill_modern_artist_names, considered 202,338: dry would_fill 131,020 / write filled 131,020 — exact match. Fills ImageEvidence.artist_ocr_name only (never overwrites a non-blank value) — evidence, not a vote table. DB cross-check: current non-blank artist_ocr_name count 144,680 = 131,020 (this run) + ~13,660 pre-existing (docs figure), consistent.
  • Calculator re-pass — two dry rehearsals (20260723T222926-022af8ca, 20260723T231506-b5c52b16, ~43.6–43.7 min each, 0 written), then write 20260724T001154-d3986cfc (00:11:54→01:04:35, 52m41s, 53.7 rows/s scan-log+vote total). local_calculate_verdicts re-run over the lexicon-gate-retracted pool: 1 CardPrintingTag vote written; DB-verified via CardScanLog, 52,348 unknown-set-code skips, 16 no-evidence skips, 117,442 to-review skips (169,806 scan-log rows total). Nearly all of the retracted pool landed in abstention/review rather than a fresh resolvable vote.
  • Consensus recompute #2 ("the closer") — dry 20260724T010750-21919b2d (139.2s), dry 20260724T011030-f216fcc1 (138.7s), apply 20260724T011448-32e08cc8 (01:14:48→01:21:12, 383.9s, 264.9 rows/s). Re-materialized tag/artist/printing consensus: tag 61,330 pairs checked, 0 transitions; artist 7,130 pairs checked, 0 transitions (every one of those 7,130 checked pairs still carries only a single vote, below the resolution threshold — see "Artist-consensus finding" below); printing 94,585 pairs checked, 1 transition (resolved→unresolved) — total_written 101,715. This is the pass that corrects the live Card.printing_tag_status resolved count from 4 back to 3 (DB-verified: unresolved 218,341 / resolved 3 / no_match 1, summing to the current 218,345-card catalog).

Artist-consensus finding (verified 2026-07-24, real, not an artifact)

Despite the artist-fill pass writing 131,020 ImageEvidence.artist_ocr_name values and the closer checking 7,130 CardArtistVote (card, votes) pairs, all 218,345 cards currently read artist_vote_status=unresolved (0 resolved, 0 unknown, 0 contested); Card.inferred_canonical_artist is non-null on 0 cards. Each of the 7,130 checked pairs carries only a single machine vote, which sits below the resolution weight/share threshold in the owner-ratified vote-weight scenario matrix — see reference/vote-weight-matrix.md (narrated in theory.md §4/§7a) for the resolution mechanics; not restated here. This is expected behavior under that ruling, not a bug: a lone machine vote never resolves anything on its own, by design.

Topline end-state (DB-verified live, queried 2026-07-24; updated 2026-07-26 post-pass — supersedes 2026-07-24 snapshot)

FIG-2 pipeline funnel

Node Value Notes
A — CATALOG (total Card rows) 218,360 +15 vs 2026-07-24 snapshot (218,345)
A1 — no phash yet 15 zero-evidence cards; equals A − B
B — phash extracted 218,345
B1 — phash but no evidence 0 clean
C — evidence extracted 218,345
C1 — evidence but no vote 121,298 46.4% of evidence-bearing cards — largest funnel gap
D — carries ≥1 vote (distinct vote-holder cards) 97,053 distinct cards with at least one vote; the 2026-07-24 figure of 166,066 was total vote rows, not distinct card count — see vote tables below
F — RESOLVED 3 unchanged
H — NO MATCH 9 was 1
G — UNRESOLVED 218,348
I — REVIEW QUEUE 3,595 most-recent-per-card CardScanLog rows with skip_reason='to-review'; the 2026-07-24 figure of 134,370 was all-time (not dedup'd to most-recent) — definition differs
  • artist_vote_status: all cards unresolved (see "Artist-consensus finding" above).

CardPrintingTag (printing-consensus votes): 169,429 total

anonymous_id rows
stage-d-join-key-v1 48,754
local-ocr-v1 41,023
stage-d-fallback-v1 29,714
deductive-backfill-v1 28,112
local-fallback-v1 11,947
local-phash-v1 8,291
lands-artist-decomp-v1 1,488
user UUIDs 100

CardArtistVote: 7,137 total (unchanged)

anonymous_id rows
residual-classify-v1 6,144
art-hash-artist-v1 987
user UUIDs 6

CardTagVote: 278,175 total (was 61,336)

anonymous_id rows
layout-class-cast-v1 216,802
local-fallback-v1 53,966
residual-classify-v1 6,144
ai-art-detector-v1 1,183
user UUIDs 80

New vote sources (2026-07-26)

Five anonymous_id sources appear in the 2026-07-26 snapshot that were absent from the 2026-07-24 topline or were previously lumped into coarser buckets:

  • lands-artist-decomp-v1 (1,488 CardPrintingTag rows) — decomposes land-card art into artist-identity signals and casts a printing vote derived from that decomposition.
  • art-hash-artist-v1 (987 CardArtistVote rows) — casts artist- identity votes by matching a perceptual hash of the card art against a known-artist hash index; previously the entire CardArtistVote non-user pool was reported as a single "ocr 7,131" bucket.
  • layout-class-cast-v1 (216,802 CardTagVote rows) — the borderless-attribute cast: classifies each card's layout and border style and casts a tag vote against the attribute-chip taxonomy seeded by seed_attribute_tags (tags 24–30); referenced as the "borderless cast" in the operational notes above.
  • residual-classify-v1 (6,144 rows in both CardArtistVote and CardTagVote) — runs a residual classifier over cards that other engines left unresolved, casting both an artist and a tag vote from the same classification pass; previously the CardArtistVote non-user pool was reported as "ocr 7,131" without this sub-source breakdown.
  • ai-art-detector-v1 (1,183 CardTagVote rows) — runs an AI-based art-style detector and casts a tag vote indicating whether the card art is AI-generated.
  • frame-style-cast-v1 (local_attribute_chip_cast, NEW 2026-07-30, 0 rows — never run) — reads stored ImageEvidence and casts the Old Border/Modern Border chip via local_fallback.classify_frame_style over collector_line_collector_number + illus_anchor_fired. Zero image fetches. Gates on BOTH collector_line_ocr and artist_ocr, because bool(None) on the nullable illus_anchor_fired would otherwise read as a real "no anchor" and classify everything modern. Reachable from the conveyor (stage_e_dispatch._run_stage_d) and from its own management command. Derivable population measured read-only 2026-07-29: 133,627 Modern Border + 9,006 Old Border.
  • bleed-edge-cast-v1 (local_attribute_chip_cast, NEW 2026-07-30, 0 rows — never run) — same module, same pass, separate identity; casts appropriate-bleed at polarity NOT_APPLICABLE for a confidently trimmed bleed_class. NEGATIVE-ONLY by design: absence of a vote is the documented convention for normal bleed, so a persistently low row count here is correct, not a coverage gap. The identity is separate from frame-style-cast-v1 precisely because of that — under one shared identity a card's frame vote would read as "already handled" and permanently strand its bleed chip. Derivable population 2026-07-29: 2,786.

Both were created because the 2026-07-29 composition audit found the only casters for these two chips inside the live-fetch pilot and inside image_evidence.extract_card_evidence, which had zero production callers — so both chips sat at zero rows with nothing able to re-derive them. See docs/features/printing-tags.md, "Who actually casts the attribute chips".

Wiring status — the streaming conveyor's _run_stage_d (2026-08-05)

The question this section answers: for EVERY identity the roster tether derives from code (19 today, 2 of them not real vote-casting calculators — evidence-transfer-v1/question-feed-hypothetical-vote, see CALCULATOR_ROSTER_ALLOWLIST in .github/scripts/docs_lint.py — leaving 17 real vote-casting identities), does a full-catalogue pass (run_pipeline, or the streaming conveyor's own stage_e_dispatch._run_stage_d) actually INVOKE it, verified by a real call site rather than a name match against this file? A prior claim in circulation held that a pass casts on "10 of roughly 28" channels — the derived roster is 17 real identities, not ~28, and as of this pass 11 of 17 are invoked by a real call site (up from 7 before this pass — see below), 6 are deliberately unwired with a stated reason.

Wired before this pass (7): stage-d-join-key-v1, stage-d-fallback-v1, stage-d-illustration-v2, stage-d-slow-path-v1 (all four local_calculate_verdicts.py, called directly by _run_stage_d), plus layout-class-cast-v1/frame-style-cast-v1/bleed-edge-cast-v1 (PR #654, _run_attribute_chip_casters). stage-d-slow-path-v1 is a router that casts 0 votes by construction (see its own entry above) but IS invoked, which is why it counts as wired here despite never appearing in a vote-count table.

Newly wired by this pass (4)ai-art-detector-v1, lands-artist-decomp-v1, residual-classify-v1, art-hash-artist-v1, via a new _run_evidence_only_calculators step in _run_stage_d. All four read only already-stored evidence (ImageEvidence OCR fields, Card.content_phash, already-resolved artist/printing chains) and fetch no image — run_lands_identify/run_frame_mismatch_recovery are called with every live-fetch budget forced to 0, which their own docstrings already documented as "the scoped, genuinely free [...] path". Per-card cost was not independently benchmarked beyond what each calculator's own docstring already measures (run_lands_identify's cached CandidateNameIndex build: 1.48s once per worker process, not per card; every other cost is a plain indexed DB read) — no full-catalogue timing run was performed as part of this pass, consistent with the constraint that no management command or backfill ran against the live, contended database while writing it. See the PR that shipped this wiring for the per-channel dispatch-level tests (each proven to fire, and proven to fail when the wiring call is removed).

Deliberately left unwired (6), with a reason and a tracked issue each:

  • art-edge-continuity-v1 — genuinely FREE, but the module's own author gated it behind a stated validation precondition (agreement/false-positive rate against Scryfall's own frame_effects ground truth) that has not run. See its own entry below and issue #721.
  • deductive-backfill-v1, local-name-frequency-v1 — read only already-stored data (no fetch/OCR/network), but NEITHER has a card_ids batch-scoping parameter: each rebuilds a whole-catalogue in-memory index or census from scratch on every call, so wiring either into a 25-card-at-a-time conveyor as-is would mean paying a full-catalogue-scale cost on every single micro-batch. Classified EXPENSIVE on IMPLEMENTATION grounds, not data-source grounds — a real finding, not an oversight. See issue #722.
  • local-ocr-v1, local-phash-v1, local-fallback-v1 — these three identities' PRIMARY caster (run_pilot/run_fallback_for_card) computes its verdict from a real image for the first time, requiring a live CDN fetch plus tesseract (OCR) or a hash computation over pixels (phash); unlike the four newly-wired channels, there is no stored-evidence reconstruction path for the PRIMARY computation (some votes under these same identities are already cast for free as a side effect of OTHER calculators' own evidence-only paths, e.g. lands-artist-decomp-v1's OCR-resolved branch — that is not this identity's own engine running for free). See issue #723.

Backfill for the 4 newly-wired channels, against the EXISTING catalogue (not future passes, which the wiring above already covers), is a SEPARATE deliverable per this project's own practice (wiring ≠ backfill) — stated in the PR that shipped this wiring, not run as part of it, since a full-catalogue pass was live and the database was contended at the time.

Calculator roster — the three identities this page omitted (2026-07-29)

Until 2026-07-29 this page enumerated eleven calculator identities. Code declared fourteen. The three below appear nowhere above, and the gate never audited them — not because it failed on them, but because it was never pointed at them. Two of the three turned out to be effectively DEAD in production, which is exactly the failure a coverage-based audit cannot surface on its own: a calculator that produces no output produces no divergence to explain, so it reads as clean by being invisible. The list is now tethered to code by .github/scripts/docs_lint.py's check_calculator_roster_tether() — every *_ANONYMOUS_ID declared under MPCAutofill/cardpicker/ must have an entry here, and CI fails if one does not. See documentation-process.md's "Roster tethers" section for the general rule.

The tether itself then had the same defect one directory down — see "Calculator roster — the identity the tether itself could not see" below.

  • stage-d-slow-path-v1 (SLOW_PATH_ANONYMOUS_ID, MPCAutofill/cardpicker/local_calculate_verdicts.py) — a ROUTER, not a voter. Working as designed. It has no printing to vote for, so it casts 0 votes of any kind by construction; its entire DB footprint is CardScanLog routing markers carrying skip_reason='to-review'135,362 of them in prod. Those markers are what populates the human review queue (the same population the "REVIEW QUEUE" funnel row above dedups to its most-recent-per-card figure). Its absence from every vote table on this page is correct behaviour and not evidence of dormancy; it is listed here so that a reader auditing vote counts does not mistake "no votes" for "not running."
  • stage-d-illustration-v2 (ILLUSTRATION_ANONYMOUS_ID, MPCAutofill/cardpicker/local_illustration.py) — was DORMANT as -v1 (3 votes in its entire existence); repaired and re-versioned 2026-07-29, not yet run in prod. It casts a printing vote from illustration identity (issue #507). The -v1 root cause: its eligibility gate read ImageEvidence.layout_class believing that field carries the card's faced-ness, when what it actually holds is a border colour (black 138,728 / borderless 72,603 / white 7,475 / '' 1,455 / silver 408), so the gate excluded 99.28% of every population handed to it — 3,409 multi-faced-v1 CardScanLog rows out of 3,426 scanned. The gate is now DELETED rather than repaired: CanonicalPrintingMetadata.face_illustrations retains every face's own illustration_id, so a back-face scan resolves to the artwork on the side actually scanned and there is no wrong-vote exposure left to guard. The -v2 rename is load-bearing, not cosmetic — multi-faced-v1 is not a rescannable skip reason and eligibility excludes cards carrying a non-rescannable scan log for the calculator's own identity, so a repaired -v1 would never re-examine the cards it wrongly skipped. Still nothing MEASURED in prod: a read-only counterfactual replay over a 30,000-card sample of the 160,585-card -v2-eligible population (2,350 of them reach the calculator with artist OCR; the rest skip as no-artist-ocr) projects ~10,277 illustration votes and ~3,233 printing votes catalog-wide — but no -v2 run has written a row. Do not read -v1's near-zero vote count as a measured statement about illustration matching's yield.
  • art-edge-continuity-v1 (ART_EDGE_ANONYMOUS_ID, MPCAutofill/cardpicker/local_art_edge.py) — DECLARED, NOT YET LIVE. Zero votes, zero CardScanLog rows, and no runner calls it — by design, not by dormancy. It is the extended-art channel: a two-sample-point pixel comparison (card edge vs. the band adjacent to the art crop) that classifies a card image as framed/extended/open and would cast the pre-existing "Extended" attribute tag. It is listed here because this roster's whole purpose is that a calculator producing no output "reads as clean by being invisible" — so the identity is declared and recorded BEFORE it can write anything, the same reasoning local_fallback's skip-reason block gives for declaring its constants ahead of first write. The gate it has not yet cleared: the classifier is validated against constructed images only, never against real card images. Before it votes, run it over the ImageEvidence rows whose confirmed printing carries Scryfall's own frame_effects extendedart (1,129 such rows in the 2026-07-28 join) and report agreement against that imported fact, plus the false-positive rate over a same-sized sample of confirmed non-extended black-bordered cards. That labelling is free and needs no human pass. Do not read its zero vote count as a measured statement about extended-art detection's yield — nothing has been measured yet. Tracked: issue #721.
  • local-name-frequency-v1 (NAME_FREQUENCY_ANONYMOUS_ID, MPCAutofill/cardpicker/local_identify_printing_tags.py) — ZERO output of any kind. Under diagnosis; may be retired. The name-frequency elimination pass deduces a match structurally, with no image fetch at all, for a name where exactly one printing is uncovered AND exactly one eligible card is unresolved. It has produced no votes and no CardScanLog rows — not even a skip log, which means it is not merely abstaining, it is not being reached. Whether it gets fixed or deleted is undecided; it is recorded here as a known hole rather than left off the page. Also missing a card_ids batch-scoping parameter (2026-08-05 finding), which is a separate, more concrete blocker than "under diagnosis" alone conveys — see the "Wiring status" section above and issue #722.

Calculator roster — the identity the tether itself could not see (2026-07-29)

The roster tether above was built to make "a vote-casting calculator nobody documented" impossible. It then failed on exactly that, one directory below where it looked: _declared_calculator_identities() scanned MPCAutofill/cardpicker/*.py with a non-recursive glob, so MPCAutofill/cardpicker/management/commands/ was never read. The non-recursion was not a scoping decision — it was there to keep cardpicker/tests/ fixture literals out of the roster, and it took the whole subtree with it as a side effect. The scan is now recursive with an explicit tests/ exclusion (.github/scripts/docs_lint.py's _roster_source_files()), which is the same exclusion stated as a decision rather than obtained as an accident.

  • scryfall-tagger-v1 (SCRYFALL_TAGGER_ANONYMOUS_ID) — RETIRED 2026-07-30, together with PrintingTagVote and its importer (management/commands/import_external_ip_tags.py, deleted). It is no longer a calculator identity, and the roster tether no longer derives it from the code, because there is no code. The entry is kept rather than deleted so that a reader meeting the string in an old report, run log or database column can find out what it was; everything below is written in the past tense and describes a design that never ran. Retirement rationale and the full behavioural record: features/printing-tags.md.

    It wrote zero rows, ever. It would have imported Scryfall Tagger's art:external-ip community art tag (Universes Beyond illustrations — Lord of the Rings, Doctor Who, Warhammer 40K) and cast machine PrintingTagVote rows against the external-ip tag, at the PRINTING level (CanonicalCard), not the catalog-image level. It was designed to write BOTH polarities: APPLY for positive Tagger matches, NOT_APPLICABLE for confirmed printings absent from the positive set, with a printing that had no data at all abstaining rather than voting. source=DEDUCTION with its own anonymous_id per the machine-caster convention — pure logical inference over already-trusted structured data, zero image inspection. Weight resolved to PRINTING_TAG_MACHINE_WEIGHT (0.5); the 2026-07-23 zero-weight override did NOT touch it (that override is scoped to source + the deductive-backfill family + one frozen run_id, all three together). Re-runs would have been idempotent via the (printing, tag, anonymous_id) uniqueness constraint, with retraction by the ordinary purge_machine_votes --run-id mechanism. What the gate could say about it was always nothing, because it produced nothing — and its absence from every vote count on this page was never evidence that it worked.

Resolved 2026-07-30, previously an open owner call: PrintingTagVote (models.py, added by PR #497) was a third vote family alongside CardPrintingTag and CardTagVote, and it appeared zero times on this page and zero times in theory.md — every vote-population figure above was silent about it. The question recorded here was whether it belonged inside this gate's scope or was deliberately outside it. It has now been answered by removal rather than by scoping: the model, its table and its only writer were retired on the owner's ruling, the table having held 0 rows on production throughout its life. The figures on this page are therefore complete as they stand, which they were not while this question was open.

Operational notes

  • Scryfall cache lost on every container rebuildissue #402, discovered 2026-07-23 after deploy-2 rebuilt the django/worker images (see docs/troubleshooting.md's "local_calculate_verdicts silently runs with an empty back-face lookup after an image rebuild" entry for the full symptom/cause writeup): MPCAutofill/scryfall_cache/default_cards.json (~2GB, later measured ~558MB compressed — see that troubleshooting entry) lived inside the django container filesystem, not on a mounted volume, so a rebuild silently deleted it; local_calculate_verdicts ran with an empty back-face lookup as a result (degraded, not crashed — easy to miss). Fix shipped: PR #412 ("Persist scryfall_cache volume and fail loud on a missing cache", merged 2026-07-24T09:07:13Z) makes scryfall_cache a named, persistent Docker volume mounted on both the django and worker services, plus a fail-loud ensure_scryfall_cache_present() guard at the start of local_calculate_verdicts.Command.handle() (raises rather than degrading silently, unless --allow-missing-scryfall-cache is passed explicitly) — went live via deploy-3 (2026-07-24), the deploy following the one (deploy-2, 2026-07-23) that originally lost the cache. This closes the gap this note previously left open; not yet independently re-verified live in this pass (no fresh cache-presence check was run while writing this section) — flagged as documentation-sourced, not re-confirmed, should that distinction matter to a future reader.
  • Border-tag seeding (7 created) — the seed_attribute_tags management command (backing the local_layout_class_cast border-vote caster, issue #369/PR #375) was run to seed the attribute-chip taxonomy it depends on: Etched, Black/White/Silver Border, Old/Modern Border, Future Frame — 7 tags, none of which existed before. Verified live: Tag ids 24–30 are exactly this set, sequential, with nothing else in that id range. Idempotent (seed_attribute_tags is safe to re-run) — mentioned here as a one-time prerequisite step, not a vote or catalog-state change. local_layout_class_cast itself went on to cast 216,811 CardTagVote rows against this taxonomy (the "borderless cast," cited in §15's pass 5 for why that pass's tag-pairs count is so much larger than this section's own 61,330/61,334 figures).

What remains open

  1. Bug-A full re-scan of the 17,531-card blank-tier-1 pool. Wave-1 (10,437 cards, the top 4 sources) is now CLOSED (2026-07-24) — see §15 for the full re-scan → reparse → lands → Stage D → closing- recompute arc, DB-verified end to end with zero consensus transitions resulting. Only the remaining ~6,535-card tail (16,972 − 10,437, issue #418) stays open, and per the owner-ratified sequencing (2026-07-24) it does not get a further batch pass — it is routed to Stage E streaming's shakedown cohort instead; see docs/proposals/stage-e-streaming.md §6/§7 (issues #153/#418).
  2. Moderation package sizing: DEFERRED (owner ruling, 2026-07-24), pending further passes. At ruling time, the raw (non-dedup'd) union of review-eligible cards across engines was 167,045 (76% of the 218,345-card catalog) — a materially larger and noisier number than this section's own dedup'd 134,370 to-review-skip figure (see "Topline end-state" above), and the owner declined to size or build the moderation package against either number while wave-1/Bug-A closure (§15), the artist-consensus gap (see above), and the other passes still in flight keep moving the underlying counts. The 134,370 dedup'd figure remains available as a working number whenever this is picked back up — not re-derived here, just not acted on yet.

15. Wave-1 Bug-A arc closure (2026-07-24)

The 10,437-card wave-1 cohort — the top 4 sources of the 16,972-card Bug-A blank-tier-1 pool (§9(c)/§13), the same cohort docs/proposals/stage-e-streaming.md §6 item 1 names as its own owner-ratified sequencing basis and §1 already measured for Stage C throughput (PilotRunLedger ids 70/75, rescan-wave1-dry-20260724/rescan-wave1b-20260724) — was carried through to a full, closed arc the same day: re-scan → reparse retraction → a full-pool lands pass → a wave-1-scoped Stage D vote → a closing consensus_recompute dry-run confirming zero transitions. Every figure below is DB-verified, keyed by run_id, in the same condensed-table style as §14's per-pass table.

# pass run_id rows notes
1 Wave-1 re-scan (Stage C) rescan-wave1b-20260724 10,437 fetched 3,282 (31%) recovered a parsed collector number, of which 2,526 further resolved a set code; 1,898 separately recovered an artist name via the artist-OCR channel. Throughput/resource detail already recorded in stage-e-streaming.md §1 (PilotRunLedger id 75, 3.351 cards/s, 3,115.1s) — this row adds the extraction-yield detail that brief didn't carry.
2 Reparse write 20260724T125238 3,243 retractions reparse_collector_evidence --write over the wave-1 cohort's newly-recovered evidence, retracting stale no-match/skip state superseded by the fresh parse — the same mechanism as the earlier Bug-B write (§9(b)/§13), scoped to wave-1 here.
3 Lands full-pool write 20260724T125355 2,701 new votes; 7,831 already-voted skips local_lands_identify --write, full eligible pool (not wave-1-scoped): 2,701 new votes cast, 7,831 honest no-op skips (PR #411's vote-collision skip-if-exists guard — a card already carrying a vote from the same anonymous_id, not a duplicate/error). Forced dry-run-before-write gate (#373) checked and passed ahead of this write.
4 Wave-1 Stage D write 20260724T133819 919 votes local_calculate_verdicts --write scoped to the wave-1 cohort: 5 genuine matches + 914 no-match — closing the loop on exactly the cohort re-scanned in pass 1.
5 Closing recompute dry-run 20260724T142039 0 transitions consensus_recompute dry-run: printing 97,212 pairs checked / 0 transitions, artist 7,130 pairs / 0 transitions, tag 226,974 pairs / 0 transitions. Tag pairs jumped from §14's 61,330/61,334 (the 2026-07-23/24 closer, still the live figure in "Topline end-state" above) to 226,974 here — consistent with the borderless-attribute cast (216,811 CardTagVote rows, verified earlier — see the "Border-tag seeding" operational note above) having landed a large new (card, tag) pair population in the interim, not a data-integrity concern. No --apply run was needed — the wave-1 votes/retractions above changed no card's resolved status on their own, consistent with the vote-weight gate's own "no machine tipping" rule (§3 item 3 / theory.md §4/§7a). Arc closed.

Wave-1 topline: of the 10,437-card cohort, 5 cards resolved a genuine new printing match via wave-1's own Stage D pass (pass 4); 2,701 additional cards (full-pool scoped, not exclusive to wave-1) received a lands-identity vote from the same re-scanned evidence (pass 3); the closing recompute (pass 5) confirms none of this moved any card's resolved/no_match status. The wave-1 slice of the Bug-A tail is closed — only the remaining ~6,535-card tail (16,972 − 10,437) is still open, and per the owner-ratified 2026-07-24 sequencing (see "What remains open" item 1 above) it does not get a further batch pass: it is routed to Stage E streaming's shakedown cohort (docs/proposals/stage-e-streaming.md §6/§7, issues #153/#418).

Clone this wiki locally