-
Notifications
You must be signed in to change notification settings - Fork 0
Pipeline Fidelity Gate
GitHub issue #154 (internally referenced elsewhere as "task #151" — a pre-board internal task-ledger number, not a second GitHub issue; issue #154's own title cross-references both). This is the single page for this gate's status, its (now-decided) owner decisions, and any in-flight implementation gating the fire. Every underlying fact below is linked to its source, not copied — if a number here disagrees with its linked source, this page is wrong and should be fixed, not the source.
Data as of 2026-07-22T18:01Z, re-verified live against production
Postgres for this page (sudo docker exec mpcautofill_django python manage.py shell, read-only queries only). §10/§11 carry a later, narrower
live re-verification (2026-07-23T00:16–00:24Z) for the #340 footprint
sizing and the Stage C run-identity note specifically — dated inline in
those sections rather than bumping this whole-page timestamp. §12
carries a further dated update (2026-07-23T01:33Z onward) for the
deploy step and the Bug-B whole-DB reparse dry-run outcome, and §9 was
amended the same day per issue #347's
Tron-reviewed zeroing plan. §13 carries a further dated update
(2026-07-23T09:0x–09:3xZ) for the zeroing steps' actual execution
(B(i) write, B(ii)+B(iii) retraction) and the §9(c) Bug-A
forced-escalation sample, re-verified live at that time. §9(d) carries
a further update (2026-07-23T09:14–11:12Z) recording the 4c pilot
dry-run's own outcome, which was still in progress as of §13's own
check. §14 records the pilot --write and the first
consensus_recompute --apply (2026-07-23), after which the gate was
fired end to end. §14 was extended 2026-07-24 (DB-verified live
against production Postgres) with five further passes that ran the
same night into the next day — a lexicon-gate retraction, a marker
reparse, an artist-credit fill, a calculator re-pass, and a second
consensus_recompute closer — which is the sequence's actual true end
state; that update also corrects the live resolved-printing count
from 4 to 3 (the closer flipped one card resolved→unresolved after
§14's original snapshot was taken) and folds in a real, verified
finding that all 218,345 cards remain artist_vote_status=unresolved
despite the artist-fill pass (see §14's "Artist-consensus finding").
§15 (new, 2026-07-24) records a separate, later arc — the wave-1
Bug-A blank-tier-1 re-scan (10,437 cards) carried through re-scan,
reparse retraction, a full-pool lands write, a wave-1-scoped Stage D
write, and a closing consensus_recompute dry-run that found zero
transitions — closing that slice of the Bug-A tail; the remaining
~6,535-card tail is routed to Stage E streaming's shakedown cohort
instead of a further batch pass (§14's "What remains open" item 1, §15).
§14's own Operational notes were also updated the same pass to record
issue #402's fix (PR #412) going live via deploy-3, and "What remains
open" item 2 (moderation) now reflects an owner ruling to defer package
sizing pending further passes. See "Chain" below for where each number's
provenance sits.
Stage D (local_calculate_verdicts.py) carries a hard precondition
before any full-catalog fire: calculators must call the existing shipped
identification code paths with ImageEvidence-supplied inputs, not
re-derive their logic; a stratified-sample parity replay against pilot
run 20260716T193408-6613a1a6's recorded outputs must show zero
unexplained divergence; a full knowledge-inventory sweep (every
empirically-derived constant/threshold/override/skip-reason mapped to
its home in the new pipeline, or flagged missing) must be clean. Full
precondition wording and the Stage D build it gates:
features/catalog-completion-plan.md's
"Harvest-calculate pipeline" section.
| artifact | status |
|---|---|
| Artifact 2 — knowledge-inventory sweep |
DONE (2026-07-22); all 3 MISSING-constant decisions now made — see §3 below. Full constant-by-constant table: reports/2026-07-22-knowledge-inventory.md. |
| Artifact 1 — stratified-sample parity replay | DONE (2026-07-22), outcome owner-accepted. 83.2% OCR-channel agreement (28,456-card reproducible-channel subset); 373/41,586 (0.9%) unexplained divergences, 0/373 a wrong-printing vote — all conservative abstentions. Not literally "zero unexplained divergence," but ruled to satisfy the gate's soundness intent. Full outcome, methodology, and the owner-acceptance ruling: see §4 below. |
Gate verdict: FIRED — §9 sequence complete end to end, true completion
2026-07-24 (the pilot --write + first consensus_recompute --apply
landed 2026-07-23; five further corrective/completion passes — lexicon-
gate retraction, marker reparse, artist-credit fill, calculator
re-pass, and a second consensus_recompute closer, §14 — ran the same
night into 2026-07-24 and are part of the same fire, not a new gate
question).
Every owner decision this gate needed is made (§3, §4); the deploy step
is done: master a587000 (PR #345) went live in prod 2026-07-23T01:33Z
— migration 0078 auto-applied, and §3 items 1–2's RESOLUTION_FLOOR_DPI/
EXCLUDED_RESOLVED_TAGS constants (PR #343) verified live in the
running container (see §3, §12). Item 3 (deductive-backfill exclusion)
was separately ruled NOT restored, already merged (§3). Artifact 1 (§4)
is DONE and owner-accepted 2026-07-22 as satisfying the gate's soundness
intent, even though the literal "zero unexplained divergence" bar
wasn't hit (373 remained, all conservative abstentions; both root
causes since fixed in code by merged PR #340). Artifact 1 itself is
closed history, not a baseline to keep re-measuring against — see
§8 for the new-data basis ratified 2026-07-23. The full §9 fire
sequence (Bug-B write pass → retraction → Bug-A sample → the 4c pilot
dry-run → owner sample audit → write → consensus_recompute --apply),
amended 2026-07-23 per
issue #347,
is now COMPLETE: the Bug-B whole-DB reparse dry-run (§12), the
zeroing steps (B(i) write, B(ii)+B(iii) retraction), the §9(c) Bug-A
forced-escalation sample, the 4c pilot dry-run (§9(d)), the owner sample
audit, the pilot --write (COMPLETE-BY-VERIFICATION, a documented
execution-harness artifact — see §14), and consensus_recompute --apply (§9(e), §14) are all DONE, executed and DB-verified. The
measurement of record for this fire is
reports/2026-07-23-4c-pilot-dry-run.md's
dry-run (pilot-dry-20260723T094518Z) plus the write's DB-verified
counts (§14) — the write reproduced the dry-run's predicted counts
exactly on both channels. §14 was then extended by five further
passes (2026-07-23 evening into 2026-07-24) that are part of the same
fire, not a new gate question: a lexicon-gate retraction, a marker
reparse, an artist-credit fill, a calculator re-pass, and a second
consensus_recompute closer — full per-pass table and topline end-state
in §14. Two open items are carried forward past the fire: (1) the Bug-A
full re-scan, deferred to post-pilot per the owner ruling in §9(c)/§13;
(2) the pilot-era consensus_recompute (49,206 tag transitions,
2026-07-23T13:25:37Z) has no PilotRunLedger row — a structural,
pre-ledger-convention gap, not a data-quality concern — see §14's
"Pass-9 ledger gap" note.
The knowledge-inventory sweep confirmed three pilot-era constants have
no current home in Stage C/D, by direct grep/read, not inference:
-
RESOLUTION_FLOOR_DPI = 200— the pilot never fetched a card below this empirically-validated floor (dpi≤150 measurably degrades OCR yield). Stage C's cohort selection has nodpicondition at all. Highest-severity of the three: no downstream signal (ImageEvidencecarries no dpi field) distinguishes a low-resolution extraction later. -
EXCLUDED_RESOLVED_TAGS = ["custom-art", "non-english"]— the pilot excluded cards already tagged custom-art/non-english (their printing-identification precondition is already falsified) from selection entirely. Stage D's_eligible_cards_querysethas no equivalent exclusion, and no other code path produces the same effect (checked directly againsttag_consensus.py). Not previously tracked anywhere prior to this sweep — a genuinely new finding, not a known accepted gap. -
Deductive-backfill exclusion (
.exclude(printing_tags__anonymous_id=DEDUCTIVE_BACKFILL_ANONYMOUS_ID)) — the pilot never re-voted a card the deductive backfill had already cast a vote for. Lower severity: only bites the narrow subset where backfill voted but the card is still UNRESOLVED (a card it fully resolved is already excluded by Stage D's ownprinting_tag_status=UNRESOLVEDfilter). Origin of this constant:DEDUCTIVE_BACKFILL_ANONYMOUS_ID = "deductive-backfill-v1"in../MPCAutofill/cardpicker/deductive_backfill.py, run via thedeductive_backfill_printing_tagsmanagement command. This is the same run that produced the 28,112deduction-sourceCardPrintingTagvotes in today's live pool (§6 below) — verified live: all 28,112 carryrun_id=None,anonymous_id="deductive-backfill-v1",created_atbetween 2026-07-14T18:21:49Z and 2026-07-14T18:22:05Z. Seejournal/2026-07-14-deductive-printing-tag-backfill.md(gitignored, machine-local) for that run's own narrative. INTENTIONALLY NOT RESTORED (owner ruling, 2026-07-22), superseding an earlier same-day revision of this page that marked it "addressed in code": a read-only investigation of the 2026-07-14 backfill found it is pure name/metadata deduction (never phash/OCR — zero image inspection) whose votes check out sound (a 15-card sample all correct), and that excluding those cards would strand ~27,819 sound-but-UNRESOLVED cards outside Stage D for no protective benefit — re-evaluating them is safe under the human-backed consensus gate (agreement dedups, disagreement surfaces to human review). The pilot's own exclusion was a performance optimization (skip a card its weaker engines couldn't add to), not a soundness mechanism, so restoring it here would trade real coverage for a protection the vote-consensus layer already provides independently. Seelocal_calculate_verdicts._eligible_cards_queryset's own docstring for the in-code record of this decision. These 28,112deduction- source votes are NAME-based (the 2026-07-14 backfill, not phash/OCR- based), stay on the record as votes, and are invisible to Stage D's own calculation post-#341 — Stage D neither reads nor is blocked by them. Their consensus-layer weighting was a separate question, parked here as still-OPEN as of this section's earlier revisions — now RESOLVED (owner ruling, 2026-07-23): these votes carry ZERO consensus weight in every resolution computation, permanently, while the rows themselves stay on the record forever (raw tallies/display paths unaffected — seevote_consensus.DEDUCTIVE_BACKFILL_ANONYMOUS_ID's own docstring for the exact mechanism and the live-audit numbers behind the ruling, andtheory.md's soundness section, §4, for the write-up). Re-scoped 2026-07-29 (owner clarification): that ruling zeroes THIS COHORT — the 28,112 rows this one run wrote, held out permanently as a measurement control — and does NOT disqualify name-matching deductive inference as a method, so a future run ofdeductive_backfill_printing_tagscasts votes carrying the ordinaryPRINTING_TAG_MACHINE_WEIGHT. Therun_id=Nonefact recorded above is no longer true of these rows: migration0097_freeze_deductive_backfill_zero_weight_cohortstamped them withvote_consensus.DEDUCTIVE_BACKFILL_ZERO_WEIGHT_RUN_ID, which is now what the zero-weight override matches on (together withsource=deductionand thedeductive-backfillcalculator family). Thecreated_atwindow above is what that migration selected on, and is not consulted at runtime by anything. This is a distinct decision from the Stage-D-exclusion question this whole numbered item is about (whether Stage D re-votes a card the backfill already touched, resolved NOT- RESTORED above) — that ruling is unchanged by this one.
None of these three are soundness violations — the human-backed consensus gate still applies to every vote Stage D casts regardless.
Owner ruling (2026-07-22T23:47Z): #1 and #2 are MUST-FIX. Sized
against live data: 28 eligible cards fall below the dpi floor (0.016%
of the 179,766-card eligible pool), 47 carry custom-art / 0
non-english (0.026%), zero overlap between the two, union 75
(0.042%) — zero live OCR votes rest on a sub-floor image. Fix: two
one-line queryset excludes in _eligible_cards_queryset
(dpi__lt=200, the custom-art/non-english tag exclusions) — merged
to master in PR #343
and deployed 2026-07-23T01:33Z (§9(a) done — see §12): both
constants verified live in the running container (RESOLUTION_FLOOR_DPI = 200, EXCLUDED_RESOLVED_TAGS = ['custom-art', 'non-english']). All
three items above are now decided — none remain open.
Full detail, plus 3 lower grade "open items" that are separate from
these 3 MISSING findings:
reports/2026-07-22-knowledge-inventory.md.
Artifact 1 was originally scoped as a live stratified-sample re-run against dpi=250, re-extracting evidence and comparing it to the pilot's outputs. That scoping was wrong and was replaced by this corrected methodology:
Dry-run diff of new
local_calculate_verdictsverdicts vs. the pilot's recorded votes; NO Stage C re-extraction; pass bar = the DOCUMENTED "zero unexplained divergence", NOT a percentage threshold.
Concretely: run Stage D's calculator in dry-run mode against the
ImageEvidence rows that already exist, diff its verdicts against the
CardPrintingTag rows the pilot run (20260716T193408-6613a1a6)
already cast, and account for every divergence — a match is not enough
on its own; every disagreement must be explained (a known, reasoned
architectural difference per the knowledge-inventory sweep) or the gate
fails. No numeric pass threshold was invented or accepted anywhere in
this process — "zero unexplained divergence" was the only documented
bar; the outcome below did not literally hit it, which is why the
owner ruling in this section exists.
Baseline vs. method, stated explicitly. Pilot run
20260716T193408-6613a1a6 is the legacy multi-channel engine — OCR
plus the local-phash-v1/local-fallback-v1 phash channels voting
concurrently (why its 43,425 votes span only 41,586 distinct cards:
some cards received votes from more than one channel). Stage D is
OCR-only by design. This replay is therefore a cross-method
verdict diff — the new OCR-only Stage D, computed against current
ImageEvidence from the new Stage C extractions, diffed against the
legacy multi-channel engine's recorded votes — not the new method
validated against itself. The 83.2% figure below is OCR-channel
agreement specifically; bucket d below (13,026) is exactly the
phash/fallback channels Stage D doesn't run — an explained
architectural difference, not a gap.
The replay ran 2026-07-22 as a read-only, full-cohort (not sampled)
comparison against that pilot run's 43,425 votes / 41,586 distinct
cards (source=ocr): each card's pilot vote was compared against what
the current Stage D join-key calculator
(calculate_join_key_verdict/_resolve_candidates_for_card, called
directly) computes from its now-persisted ImageEvidence, bypassing
_eligible_cards_queryset entirely. No writes.
Headline: 23,789/41,586 (57.2%) agree outright. Restricted to the
28,456 cards whose pilot vote actually used the OCR channel Stage D can
reproduce (local-ocr-v1), agreement is 23,689/28,456 = 83.2%.
Divergence buckets (17,793 disagreements):
| bucket | count | note |
|---|---|---|
a. RESOLUTION_FLOOR_DPI
|
0 | structurally 0 in this cohort — pilot's own filter already excluded these before ever voting |
b. EXCLUDED_RESOLVED_TAGS
|
0 | same — structurally pre-filtered |
c. DEDUCTIVE_BACKFILL (§3 item 3) |
0 | same, doubly structural (deductive_backfill's own eligibility requires zero pre-existing votes) |
| d. pilot engine has no Stage D analogue | 13,026 | pilot matched via local-fallback-v1/local-phash-v1 only — both explicitly out of Stage D's scope per its own docstring/theory.md §7 |
| e. Stage D's new veto layer (border/copyright-year mismatch) | 4,394 | correctly withholds matches on cards whose evidence genuinely disagrees with the real printing (spot-checked, not a classifier bug) |
| f. UNEXPLAINED (the gate criterion) | 373 | (0.9% of cohort) 2 identified root causes, 0/373 a wrong-printing vote — every one a conservative abstention. See theory.md §7c. |
Buckets a/b/c reading exactly 0 is a property of this backward-looking cohort (the pilot's own eligibility filter already excluded any card that would trip them), not evidence the §3 constants don't matter going forward — that is the separate, forward-looking question §3 answers (all three items resolved 2026-07-22, per that section — items 1–2 MUST-FIX with a fix in flight, item 3 not restored).
Not literally zero — 373 unexplained (0.9% of the 41,586-card cohort) — but every one is a conservative abstention (0/373 voted for a wrong printing). Ruling: the gate's intent (no confidently-wrong verdicts at scale) is satisfied — zero false-accept risk. The 373 were not treated as a fire blocker; both root causes were identified and fixed in the same session, code-only, in merged PR #340:
- 155 "no-text" divergences — the 2026-07-21 OCR short-circuit skipped deeper tiers whenever both tier-1 OCR attempts were digit-free, conflating a blank/failed tier-1 read (a read failure) with a confident "no collector number here" finding. Narrowed to escalate whenever tier-1 comes back blank, exactly like a digit-bearing-but-unparseable read already did.
-
218
is_no_matchdivergences (subset fixed) — a glued-token OCR parse failure: a single language-marker character glued onto the tail of a set-code token, e.g. card 41559 ("Verazol, the Split Current") parsingset_code="znre"from"znr"+ an adjacent language-marker token's leading"e"; the real set isznr. Same family as PR #260's denominator/rarity-token glued-token guard.
Separately: the three §3 constants reading 0 divergence in this backward-looking replay is a structural artifact of the pilot's own pre-filtering (see "Result" above), not evidence of their forward impact — that sizing is tracked in §3, independent of this ruling.
PR #340's fixes are code-only. Realizing the benefit on the live
373-card cohort (or any other newly-affected cards) requires a
targeted Stage C re-extraction of the affected card_ids — a
gated prod write, queued behind the post-freeze deploy, and explicitly
out of scope for / not run by PR #340. The 373-card cohort is bounded
to the pilot-vote replay's own 41,586-card comparison set, not the
full catalog — §10 sizes both root causes' actual catalog-wide
footprint (17,531 cards / 284 rows respectively), which is the scope
the re-extraction in §9(b) actually runs against.
Full source: issue #154's
2026-07-22 comments (the replay result and the owner-acceptance
ruling) and PR #340
(the fix detail, verification, and scope-boundary note). Durable
technical distillation of what the replay showed about the OCR channel
and the conservative-abstention property: theory.md §7c.
Pilot run 20260716T193408-6613a1a6 completed 2026-07-16/17 — this
predates the ImageEvidence model entirely. Migration 0068 (the
ImageEvidence substrate) wasn't applied to production until
2026-07-20 (see
features/catalog-completion-plan.md's
"Migration 0068 (only) live on production" entry). There is therefore
no ImageEvidence row that existed at pilot time to diff against —
Artifact 1 can only ever compare Stage D's verdicts (computed today,
against today's ImageEvidence) to the pilot's recorded votes
(CardPrintingTag rows, which do predate and survive the migration).
This is the reason §4's methodology is a verdict-diff, not an
evidence-diff, and why re-extracting Stage C evidence to match the
pilot's original images would not close this gap even if it were done
(the pilot's own transient-fetch images were never persisted — see
CLAUDE.md's "Governing premise: we index, we do not store images").
flowchart TD
A["CATALOG"] --> B["PERCEPTUAL HASH<br/>extracted"]
A -.-> A1{{"no phash yet"}}
B --> C["IMAGE EVIDENCE EXTRACTED<br/>near-complete"]
B -.-> B1{{"phash but no evidence yet"}}
C --> D["CARRIES ≥1 PRINTING VOTE<br/>machine votes abundant"]
C -.-> C1{{"evidence but no vote yet<br/>Stage D has not reached these"}}
D --> E{"human-backed consensus gate<br/>weight ≥ 2 · share ≥ 0.6 · a human said so"}
E -- "cleared" --> F(["RESOLVED PRINTING<br/>rare — by design"])
E -- "held" --> G{{"UNRESOLVED<br/>machine votes alone can never clear this gate"}}
E -- "ruled out" --> H(["NO MATCH"])
G --> I["REVIEW QUEUE<br/>awaiting a human vote"]
classDef stage fill:#24283b,stroke:#565f89,stroke-width:1px,color:#c0caf5
classDef gate fill:#2f3549,stroke:#ff9e64,stroke-width:2px,color:#c0caf5
classDef halt fill:#f7768e,stroke:#8c3d4e,stroke-width:2px,color:#1a1b26
classDef excl fill:#7dcfff,stroke:#3f7f9c,stroke-width:2px,color:#1a1b26
classDef done fill:#9ece6a,stroke:#5c7c3d,stroke-width:2px,color:#1a1b26
class A,B,C,D,I stage
class E gate
class G halt
class A1,B1,C1 excl
class F,H done
Deliberately shape-only — no absolute counts in this diagram (owner ruling, 2026-07-25): every number behind each row above is real, live, and already published in the tables directly below this figure on this same page — repeating them inside the diagram itself would fork the numbers this page's own rule forbids ("don't restate gate status/ decisions elsewhere, link here") and would go stale the moment the next pass runs. What the shape says instead: extraction (phash → image evidence) is functionally complete; Stage D coverage (the "carries a vote" row) is the one place more machine capacity still helps; the human-backed consensus gate is a hard requirement, not a formality — machine votes alone can never clear it, so "held" is the gate working as designed, not a fault, and the review queue is where held cards wait for exactly the one thing that clears them: a human vote. Absolute counts return to this diagram once real user confirmations start accumulating in volume; until then, read them from the tables immediately below.
Catalog: 218,285 cards; 218,270 with a current ImageEvidence row.
Vote pool (CardPrintingTag / CardTagVote / CardArtistVote,
live, grows continuously):
| pool | total | by source |
|---|---|---|
| printing | 101,105 | ocr 72,938 / deduction 28,112 (all deductive_backfill, §3 item 3) / user 55 |
| tag | 61,334 | ocr 61,294 / user 40 |
| artist | 7,137 | ocr 7,131 / user 6 |
| total | 169,576 |
Pilot run 20260716T193408-6613a1a6 — three distinct numbers, not
one flattened figure:
-
165,980 candidates scanned (the pilot's own completion-log
counter; not independently re-derivable from
CardScanLogtoday — seereports/2026-07-22-knowledge-inventory.md's note on non-persisted run counters for why some legacy-engine in-memory stats can't be re-queried after the fact). -
43,425 votes cast (live
CardPrintingTag.objects.filter(run_id=...)count — corrected 2026-07-22 from a previously-stated 43,426 intheory.md/catalog-completion-plan.md; the recordedPilotRunLedgerledger row still saysvotes_written=43426, an off-by-one against the live count with no documented retraction explaining the difference). -
41,586 distinct cards voted (live
.values('card_id').distinct().count()— smaller than votes-cast because a card can receive votes from multiple engines, e.g. OCR and phash both voting on the same card).
Stage D run history (CardPrintingTag/PilotRunLedger, local_calculate_verdicts):
| run_id | votes written | notes |
|---|---|---|
staged-dryrun-20260721T0423Z |
0 | dry-run |
staged-write-20260721T0434Z |
0 (live, 2026-07-23) |
PilotRunLedger.votes_written records 8,925 — the pre-retraction figure; #258's retraction (2026-07-21) brought the live count to 8,825 (already documented in reports/2026-07-21-recovery-arc.md); the 2026-07-23 zeroing retraction (§13) then deleted 8,801 of those, leaving 24 — which B(i)'s write pass had already re-labeled off this run_id as corrected flips, so 0 remain attributed here live |
staged2-0721 |
0 (live, 2026-07-23) | was 70; all 70 deleted by the 2026-07-23 zeroing retraction (§13) |
staged3-0721 |
0 (live, 2026-07-23) | was 3,010; all 3,010 deleted by the 2026-07-23 zeroing retraction (§13) |
staged4-0721 |
0 (live, 2026-07-23) | was 999; all 999 deleted by the 2026-07-23 zeroing retraction (§13) |
interim-peek-0722 |
0 | dry-run |
bugb-reparse-dry-20260723T014652Z (PilotRunLedger id 32) |
0 | dry-run, reparse_collector_evidence whole-DB Bug-B measurement, 197,938 considered — see §12 |
bugb-reparse-scoped-dry-20260723T020508Z (PilotRunLedger id 33) |
0 | dry-run, reparse_collector_evidence scoped to the 284-signature ID file — see §12 |
bugb-reparse-voted33-dry-20260723T0206Z (PilotRunLedger id 34) |
0 | dry-run, reparse_collector_evidence scoped to the 33 previously-voted Bug-B cards — see §12 |
bugb-write-dry-20260723T090258Z (PilotRunLedger id 35) |
0 | dry-run, pre-write confirmation for B(i) — see §13 |
20260723T090331-fdf5822b (PilotRunLedger id 36) |
— | dry-run, retract_stage_d_by_run_id pre-retraction preview (12,904 votes / 7,773 skips previewed) — see §13 |
bugb-write-20260723T0905Z (PilotRunLedger id 37) |
236 |
B(i) live write — reparse_collector_evidence --write, considered 285 / fields_fixed 285 / retracted 236 / gate_refused 0 — see §13 |
20260723T091446-35a1bde5 (PilotRunLedger id 38) |
— |
B(ii)+B(iii) live retraction — retract_stage_d_by_run_id --write: 12,880 votes deleted (+24 already flipped by B(i) = 12,904 total staged votes accounted for), 7,773 skips deleted, 20,653 cards resynced, 0 resolved-gate refusals — see §13 |
buga-sample-20260723T0927Z (PilotRunLedger id 39) |
— |
§9(c) Bug-A sample — run_image_evidence_cohort, live, 300-card uniform sample (seed 20260723) of the 17,531-card blank-tier-1 pool, --no-shortcircuit — see §13 |
buga-sample-verdicts-dry-20260723T093321Z (PilotRunLedger id 40) |
0 | dry-run, reparse_collector_evidence verdict pass over the id-39 sample — 1 genuine match / 76 no-match / 223 skips — see §13 |
pilot-dry-20260723T094518Z (PilotRunLedger id 41, §9(d) dry-run) |
0 | dry-run, full eligible-pool Stage D pass — join-key 39,253 match/61,247 no_match, fallback 29,710 (read-only recomputation) — see reports/2026-07-23-4c-pilot-dry-run.md
|
pilot-write-20260723T1202Z (PilotRunLedger id 42, §9(d) write) |
130,210 at write (48,151 join-key rows survive live, 2026-07-24 — 52,349 retracted by the lexicon-gate retraction, §14; fallback 29,710 unaffected) |
live write — join-key 100,500 (39,253 match/61,247 no_match) + fallback 29,710 match (still 29,710 live, never retracted), exact match to the dry-run's prediction; ledger row reads status=failed/votes_written=None, a documented execution-harness artifact (client timeout severed the executor's stdout mid-run, not the write) — see §14/data/2026-07-23-pilot-write-and-recompute.md
|
20260724T001154-d3986cfc (post-fire calculator re-pass write) |
1 |
§14 "Five further passes" write — local_calculate_verdicts re-run over the lexicon-gate-retracted pool; 52,348 unknown-set-code + 16 no-evidence skips, 117,442 to-review — see §14 |
Stage C run history (ImageEvidence.run_id, current last-writer row
count per run — updated 2026-07-27 after the full-catalog re-extraction):
| run_id | date | rows |
|---|---|---|
pass-pilot-20260725 |
2026-07-25 | 100 |
stage-e-stream-20260725T233633221123Z |
2026-07-25 | 22 |
stage-e-stream-20260725T233633687349Z |
2026-07-25 | 23 |
pass-full-20260725 |
2026-07-25/26 | 194,831 |
pass-full-20260725-r2 |
2026-07-25/26 | 23,132 |
| Total | 218,108 |
The 2026-07-25 full-catalog re-extraction (pass-full-20260725 +
pass-full-20260725-r2) superseded all prior Stage C last-writer rows —
the original canaries (stagec-canary-20260720T1659Z,
stagec-canary-decoupled-20260720T235127Z), ntx-0721,
stagec-20k-20260721T0227Z, and stagec-remainder-0721 no longer appear
as last-writer for any ImageEvidence row. See
reports/2026-07-20-decoupled-canary-confirm.md
and reports/2026-07-21-stagec-20k-extraction.md
for the original canary/20k narrative detail, and
reports/2026-07-26-stagec-full-catalog-completion.md
for the full-catalog completion record. None of the Stage C runs have a
PilotRunLedger row of their own — see §11 for what that does and
doesn't affect.
ntx-0721 is NOT a pilot — stated explicitly because this number
has been misremembered before: ntx-0721 is a Stage C extraction
run (22,899 CardScanLog rows, a no-text cohort re-extraction — see
reports/2026-07-21-recovery-arc.md).
Verified live: 0 votes cast under run_id='ntx-0721' across all
three vote tables (CardPrintingTag/CardTagVote/CardArtistVote).
There is no 23,000-vote pilot; "23,111 votes" was a transient same-day
pool snapshot at some earlier point, never a pilot run of any kind, and
should not be cited as a baseline.
printing_tag_status: 218,281 unresolved, 3 resolved, 1 no_match.
consensus_impact_report dry-run (2026-07-22, --sample-limit 20,
zero writes): printing 92,368 pairs checked / 0 transitions; artist
7,130 pairs / 0 transitions; tag 61,328 pairs, with 49,207
None→UNRESOLVED materializations sized as pending a separately
owner-gated recompute pass. "Zero transitions" on printing/artist means
the ratified resolver's answer already matched every currently-persisted
status exactly, at that date. Superseded by the real apply, and then
superseded again: this dry-run's tag/printing sizing is now closed
history — consensus_recompute --apply ran 2026-07-23T13:25:37Z and
materialized 49,206 of the sized 49,207 tag transitions (organic
interim resolution accounts for the one-fewer figure) plus one
additional printing resolution (3 → 4); a second consensus_recompute
closer then ran 2026-07-24T01:14:48Z and flipped that same printing
card back (resolved→unresolved), so the live resolved count is 3,
not 4 — see §14's "Five further passes" and "Pass-9 ledger gap" for the
complete before/after.
This page owns the gate's status and its decisions above; every fact below is the linked source, not duplicated here.
-
theory.md— the formal decoding model, the false-accept bound, and §7's stage-by-stage Stage D composition with error terms. Keeps the pilot's own per-engine breakdown and calibration numbers. -
identification-pipeline.md— plain-language walkthrough of the same Stage C/D pipeline, stage by stage, for a reader who wants the mechanics without the formal model. -
reports/2026-07-22-knowledge-inventory.md— Artifact 2's full constant-by-constant inventory table (SAME / CHANGED / MISSING / superseded-by-architecture / open item), the source for §3 above. -
features/catalog-completion-plan.md— the six-part catalog-completion plan; the Stage D precondition wording (§1 above) and the Stage C/D build detail live there, not here. -
data/2026-07-22-pipeline-snapshot.md(+ its.jsonsibling) — the dated raw-data record this page's §6 numbers were re-verified against; re-query before trusting any live-pool number more than an hour or two old, per that file's own provenance notes. -
reports/2026-07-21-recovery-arc.md— thestaged-write-20260721T0434Z8,925→8,825 retraction (#258) cited in §6's Stage D table. -
MPCAutofill/cardpicker/deductive_backfill.py— the module behind the 28,112deduction-source printing votes and §3 item 3's MISSING exclusion. - GitHub issue #154 — the artifact-1 parity-replay result and the owner-acceptance ruling comments (2026-07-22), the source for §4's numbers.
- PR #340 (merged) — the code fix for both of §4's root causes; the live 373-card cohort still needs the gated re-extraction described there.
- PR #341 (merged) — §3 item 3's non-restoration rationale and code removal.
-
PR #343
(merged) — §3 items 1–2's
RESOLUTION_FLOOR_DPI/EXCLUDED_RESOLVED_TAGSqueryset-exclude fix, merged to master and deployed via PR #345. -
PR #345
(merged) — the §9(a) deploy commit (master
a587000); also makesrun_image_evidence_cohortself-recording viaPilotRunLedger(unrelated to §9(a) itself, ledger-tracking-only, see §11). - issue #347 — the Tron-reviewed pre-pilot machine-vote zeroing plan that amended §9 2026-07-23 (the retraction step, the Tron corrections, the consensus-safety verification cited in §9/§12).
-
reports/2026-07-23-4c-pilot-dry-run.md— the §9(d) measurement of record: the full eligible-pool Stage D dry-run's per-channel counters and the read-only fallback-recovery methodology, cited in §9(d)/§14. -
data/2026-07-23-pilot-write-and-recompute.md— the §9(d)/§9(e) fire-sequence closing steps: the pilot--write's DB-verified vote counts, thestatus=failedexecution-harness artifact explanation, andconsensus_recompute --apply's outcome, the source for §14 above. -
data/2026-07-23-bugb-reparse-dryruns.md— the §9(b)/§12 Bug-B whole-DB reparse dry-run report + resource metrics, keyed byrun_id. -
data/2026-07-23-zeroing-and-buga-sample.md— the §9 B(i)/B(ii)+B(iii)/(c)/§13 zeroing-execution and Bug-A sample report + resource metrics, keyed byrun_id. -
MPCAutofill/cardpicker/management/commands/consensus_recompute.py(PR #336, merged) — the--applycommand §9(e)'s materialization step runs (STRICTLY LAST — see §9's Tron correction).
The 2026-07-22 full-catalog Stage C sweep (§6's stagec-remainder-0721
main leg, plus its two preceding canary/20k runs) is the epistemic
foundation for everything this gate measures going forward. Two
consequences, both ratified in-session 2026-07-23:
-
Artifact 1 (§4)'s legacy-pilot comparison is CLOSED HISTORY. Pilot
run
20260716T193408-6613a1a6served its purpose as a cross-method corroboration point (§4, §5,theory.md§7c) and stays on the record as that, permanently — but it does not get re-run or re-compared against as new data lands. It is not the baseline going forward. - The new system's own pilot is DERIVED from the new data, not inherited from the old one. The upcoming full-pool Stage D dry-run (§9 step (d), the "4c pilot") IS the pilot / measurement of record from this point on — its own statistics (match/no-match/abstain counts, per-channel join-key-vs-fallback breakdown) become the cited figures for any future soundness argument about this pipeline, not a diff back against the legacy engine.
Scope note: this is a basis change, not a retraction. Every legacy
(source=ocr, pilot-era) and deduction (source=deduction,
deductive_backfill) vote already in the DB stays exactly as-is — no
retraction, no source filter applied to either, in this ruling or any
prior one. They continue to count as live, valid votes toward
consensus; they simply stop being cited as the validation baseline for
new soundness claims.
The full gated sequence this gate's verdict (§2) is blocked on, in order. Amended same-day by issue #347 (Tron-reviewed pre-pilot machine-vote zeroing plan) to insert an explicit retraction step ahead of the pilot dry-run — rationale: the ratified new-data basis (§8) requires the 4c pilot to cast 100% of fresh machine votes against the corrected evidence/parser (a new-data-basis requirement, not a data-quality complaint about the staged votes themselves), and consensus safety for doing so was independently verified (all 12,904 staged cards UNRESOLVED before and after, zero resolved-card overlap, zero user-visible delta during the no-vote window, zero ES writes from retraction).
(a) Deploy master — DONE 2026-07-23T01:33Z. PR #345 (master
a587000) put PR #343's §3 items 1–2 fix (and PR #341's item 3
non-restoration) live in prod; migration 0078 auto-applied; #343's
constants verified live in the running container (§3, §12).
(b) Bug-B whole-DB reparse dry-run — DONE (§12). The fixed parser
applied offline across all 197,938 stored raw texts confirmed the
existing 285-row prediction (§10) exactly: 284 glued-marker guard
rows plus 1 unrelated improvement (card 62354). Full detail, resource
metrics, and the scoped-33 sub-run: data/2026-07-23-bugb-reparse-dryruns.md.
B(i) Bug-B write pass — DONE 2026-07-23 (§13). reparse_collector_evidence --write (run bugb-write-20260723T0905Z) against the
regenerated 284-signature ID file (the exact cohort §12's dry-run
confirmed), NOT the broader parser-bug regex selector (a 553-card
different, wider cohort — not substituted). Patch requirement (persist
corrected parse fields unconditionally) verified in the outcome:
considered=285, fields_fixed=285 (all 285, unconditionally — the
49-row gap is closed), retracted=236 (votes actually flipped),
gate_refused=0.
B(ii)+B(iii) Retraction — one invocation, DONE 2026-07-23 (§13). The
single-purpose retract_stage_d_by_run_id command (run
20260723T091446-35a1bde5), scoped to anonymous_id=stage-d-join-key-v1
AND the four staged run ids, deleted:
-
12,880
CardPrintingTagvotes (staged-write-20260721T0434Z/staged2-0721/staged3-0721/staged4-0721) plus the 24 votes B(i) had already flipped to genuine matches within the same target cohort = all 12,904 originally-staged votes accounted for, matching §6's table exactly (8,825 + 70 + 3,010 + 999 = 12,904; a pre-retraction dry-run preview,20260723T090331-fdf5822b, confirmed the full 12,904 before B(i)'s write ran). This task's own independent pre-flight check (2026-07-23T09:41Z, ahead of launching (d)) confirmed the same end state: 0CardPrintingTagrows remain withanonymous_id=stage-d-join-key-v1. -
7,773 non-rescannable
CardScanLogstale skips from the same four runs (per-run: 7,187 / 14 / 19 / 553), exactly as predicted. (The much largerno-evidenceskip population from these same runs was deliberately excluded — it is rescannable via the pilot's own native resume-filter logic, not a retraction target.)
Per-card resolve_printing() safety gate applied (same conservative
card-level check as reparse_collector_evidence, §3 item 3's
docstring) — 0 skipped_resolved_gate refusals, confirming zero
resolved-card overlap. Verified untouched: user votes (55),
deduction votes (28,112 — §3 item 3, still 28,112 live), legacy pilot
votes (43,425, still 43,425 live), all tag/artist votes, the 3 resolved
- 1 no_match cards (still 3 resolved live). 20,653 cards resynced via
resolve_and_persist_printing()— Tron correction still applies: that function casts no votes itself, it only recomputes/persists printing state for the card just retracted. Verified end-state (2026-07-23, live):CardPrintingTagrows withanonymous_id='stage-d-join-key-v1'= 0; non-rescannableCardScanLogskips for the 4 target runs = 0. Full run report:data/2026-07-23-zeroing-and-buga-sample.md.
(c) Bug-A forced-escalation SAMPLE — DONE 2026-07-23 (§13). 300
cards uniform-randomly sampled (seed 20260723, --no-shortcircuit) from
the 17,531-card blank-tier-1 pool (§10), run buga-sample-20260723T0927Z
(extraction, 85.4s) + buga-sample-verdicts-dry-20260723T093321Z
(verdict dry-run). Funnel: 300 fetched → 78 non-blank OCR text (26.0%)
→ 78 parsed numbers → 65 set codes → 76 no-match votes → 1 genuine
match (0.33%, card 122326 "Ephemerate" sketch variant → STA 68,
spot-checked correct) → 223 skips. Wilson 95% extrapolation to the full
17,531-card pool: ~58 genuine matches [CI 10–327], qualitatively
low-end likely (spot-checks show OCR noise dominating the non-blank
yield); full re-scan cost estimated ~83–104 minutes.
Owner ruling (2026-07-23): Bug-A full re-scan DEFERRED to
post-pilot, gap tracked not dropped: (a) the 17,531-card signature
query regenerates on demand (fetch_ok=True, empty collector number,
blank/whitespace raw text, excluding ntx-0721; 17,531 at
2026-07-23T09:19Z); (b) the pilot's own skip counters (§9(d)) will
surface the blank-evidence abstentions so the gap stays visible; (c)
any future re-scan MUST include a state-clear step first — this
sample's no-text skips are non-rescannable (unlike no-evidence
skips), so a post-pilot re-scan recipe is: re-extract with
--no-shortcircuit → clear stale skip state via the reparse path →
run a follow-up scoped Stage D pass to vote. Documented here as the
recipe so it is not re-derived. Full run report, funnel table, and the
recipe's full rationale:
data/2026-07-23-zeroing-and-buga-sample.md.
(d) Stage D dry-run over the full eligible pool = THE PILOT — DONE
2026-07-23 (§8), run_id pilot-dry-20260723T094518Z, git_sha
42a09b3c794f7cf8aca5eb1ca2d4f6cdaa2895a6, eligible pool 200,366 (post-
zeroing, live-verified immediately pre-run). Full statistics, the
per-channel join-key-vs-fallback breakdown, and the read-only fallback
recomputation methodology (the actual command reports fallback
considered=0/votes=0 for a structural reason — see below — not a
data gap): reports/2026-07-23-4c-pilot-dry-run.md.
Structural finding: _fallback_eligible_cards_queryset sources its
"join-key found no confident hit" population from PERSISTED
CardPrintingTag/CardScanLog rows — which a same-invocation DRY RUN
never writes — so local_calculate_verdicts (no --write) can concretely never produce a nonzero fallback count in one pass, regardless of flags;
this is inherent to the current implementation, not fixable by a
different invocation. The fallback channel's real first-ever numbers
were recovered via a bounded, zero-write diagnostic that re-derives the
join-key no-hit population in-memory using the same production
functions, then calls the actual calculate_fallback_verdict per
eligible card — see the linked report for the full methodology and
results. Followed by an owner SAMPLE AUDIT — DONE (150 uniformly
sampled verdicts, seed 20260723150) — sheet written machine-local only
(not committed; see the linked report for its exact path and
provenance), then --write — DONE 2026-07-23 (run
pilot-write-20260723T1202Z, PilotRunLedger id 42,
COMPLETE-BY-VERIFICATION — both channels' vote counts DB-verified to
match the dry-run's prediction exactly, 130,210 total votes; full
outcome and the status=failed execution-harness artifact explanation:
data/2026-07-23-pilot-write-and-recompute.md).
(e) consensus_recompute --apply — STRICTLY LAST — DONE 2026-07-23T13:25:37Z.
Materialized the pending None→UNRESOLVED tag transitions sized in
§6's consensus_impact_report dry-run (49,206 of the 49,207 sized
there — one fewer, organic interim resolution, not a discrepancy), via
the command in
consensus_recompute.py
(PR #336). Exit code 0. Tron correction to an earlier "tag-orthogonal"
claim: consensus_recompute recomputes printing + artist + tag
state via the real resolver paths, not tag state alone — its safety in
this sequence comes from idempotence plus running strictly last,
not from any orthogonality to the printing-layer steps above; it was
run strictly last here too. Printing resolution moved 3 → 4 resolved
cards (DB-verified); artist (7,130 pairs) had zero transitions. Full
outcome: data/2026-07-23-pilot-write-and-recompute.md.
Note: this command predates the PilotRunLedger self-recording
convention — no ledger row exists for it, a small follow-up flagged
beside the already-tracked §11 Stage C run-identity gap. Re-confirmed
2026-07-24 by an independent DB-only audit (exhaustive PilotRunLedger
scan, 64 rows, no gap-filling row found for this invocation): the gap is
real and structural, not an omission — see §14's "Pass-9 ledger gap"
note for the full corroboration attempt.
Order was load-bearing and followed exactly: (a) → (b)/B(i) →
B(ii)+B(iii) → (c) → (d) → sample audit → --write → (e)
consensus_recompute --apply strictly last. The full §9 sequence
(as originally ratified) is COMPLETE. §14 records five further
corrective/completion passes that ran after (e), extending true
completion to 2026-07-24 — not a reopening of §9's own ratified order,
which stands as executed.
Every run in this sequence gets its own report plus resource metrics
(RSS/IO/CPU, per-card cost) committed under docs/data/, keyed by that
run's run_id, and added as a new row to §6's Stage C/Stage D run
history tables (not a separate table).
Both root causes fixed by merged PR #340 (§4) were sized against the full catalog, not just the 373-card replay cohort (§4's cohort is bounded to the pilot-vote comparison set, 41,586 cards; this sizing covers the whole 218,285-card catalog).
Bug A (OCR short-circuit over-skip). 17,531 cards catalog-wide
carry the blank-tier-1 signature — a necessary-condition proxy for the
bug (matches the pattern the fix targets; not a guarantee every one
flips to a match), excluding ntx-0721's cohort (already
force-escalated by that run, §6). 15,948 of the 17,531 (91%) are
Stage-D-eligible. Expected genuine-match conversion is LOW:
ntx-0721's own base rate for this same signature is ~0.2% — the
reason §9(c) samples first rather than committing to a full re-scan.
Bug B (glued language-marker set-code). 284 rows carry the
signature, systemic across ≥7 distinct set codes (not one bad batch).
212 of the 284 are Stage-D-eligible; 33 already carry staged-run votes,
and every one of those 33 is is_no_match=True with 0 wrong-printing
votes — direct empirical confirmation, at this wider scale, of the
same conservative-abstention direction §4/theory.md §7c already
established at the 373-card replay scale. The whole-DB reparse (§9(b))
is pre-sized, not estimated: offline re-parsing all 197,938 stored raw
texts with the fixed parser diffs exactly 284 guard rows plus 1
unrelated improvement (card 62354).
Stage C runs are not tracked in PilotRunLedger — verified
directly: zero rows for run_image_evidence_cohort. Completion-log
counters like stagec-remainder-0721's (141,369 processed / 57,725
short-circuited / 492 fetch failures, §6) are log-only and not
DB-re-derivable after the fact, the same caveat §6 already states for
the older pilot's own 165,980-candidate counter. This does not affect
the evidence itself: ImageEvidence rows are fully persisted
regardless of ledger tracking — 197,469 current rows verified at this
check, window 2026-07-21T19:29:24Z–2026-07-22T16:45:43Z (one row's
drift against §6's 197,470 count for the same run is two snapshots
taken at different times, not a data-integrity concern — re-query
before trusting either number to the exact row). A ledger fix (Stage C
runs writing their own PilotRunLedger rows, closing this gap) is in
flight as a separate code PR, not yet merged as of this note.
Deploy (§9(a)): master a587000 (PR #345) merged and deployed to
prod 2026-07-23T01:33Z. Migration 0078 auto-applied — verified
live (manage.py showmigrations cardpicker shows [X] 0078_pilotrunledger_counters
as the last applied row). §3 items 1–2's constants verified live in the
running container: RESOLUTION_FLOOR_DPI = 200,
EXCLUDED_RESOLVED_TAGS = ['custom-art', 'non-english'].
Bug-B whole-DB reparse dry-run (§9(b)): three reparse_collector_evidence
dry-runs, PilotRunLedger ids 32–34 (§6's Stage D run history table),
all dry_run=True / votes_written=0 — verified live against
PilotRunLedger directly.
-
bugb-reparse-dry-20260723T014652Z(whole-DB, 197,938 candidates): the offline 285-changed-row prediction already recorded in §10 verified exactly — 284 glued-marker guard rows plus 1 unrelated improvement (card 62354). The command's own internalchanged=162,866counter is a different, broader metric (see the command's own module docstring: it compares a fresh re-parse against each card's currently-RECORDED join-key verdict, not against the specific glued-marker signature) — arithmetic cross-check:no_evidence=0 + no_prior_join_key_state=16,253 + unchanged=18,819 + changed=162,866 = 197,938, matchingconsideredexactly. The 162,866 figure is explained as stale no-evidence skips from the 2026-07-21 staged passes predating full Stage C evidence coverage, handled natively by the pilot's own rescannable-skip resume logic — not Bug-B blast radius, and explicitly not a write target for B(i)/B(ii+iii). -
bugb-reparse-scoped-dry-20260723T020508Zandbugb-reparse-voted33-dry-20260723T0206Z: scoped confirmation runs, completed in 2.94s and 1.69s respectively. The voted33 run found 24/33 previously-voted Bug-B cards flip false-no-match→genuine match under the fixed parser — ground-truth confirmation of those 24 is deferred to the pilot's owner sample audit (§9(d)), not asserted here.
Full run report, resource metrics (RSS/IO/CPU, per-card cost), and
provenance: data/2026-07-23-bugb-reparse-dryruns.md.
This also doubles as the first runtime calibration for the §9(d) 4c
pilot dry-run — same verdict-computation code path.
Fallback channel (§9(d)): as of this section's original
2026-07-23T01:33Z-onward timestamp, verified live 0 CardPrintingTag
rows carried anonymous_id='stage-d-fallback-v1' — the fallback
channel had never cast a vote in production. §9(d)'s own entry now has
the full outcome (still zero live votes — the pilot is DRY-RUN
only — but the channel's own would-cast statistics are now known, via
a read-only recomputation; see that entry and its linked report).
Remaining §9 steps as of this section's timestamp (2026-07-23T01:33Z
onward): B(i), B(ii)+B(iii), (c), (d), sample audit, --write, (e) were
all NOT YET RUN. §13 below carries a later same-day update: B(i),
B(ii)+B(iii), (c), and (d) (the 4c pilot dry-run) are now all DONE
(§9(d)'s own entry carries the full outcome); sample audit / --write /
(e) remain NOT YET RUN.
B(i) Bug-B write pass — DONE. Run bugb-write-20260723T0905Z
(PilotRunLedger id 37, preceded by dry-run confirmation id 35):
considered=285, fields_fixed=285 (all persisted unconditionally, closing
the 49-row gap), retracted=236 (votes actually flipped), gate_refused=0.
B(ii)+B(iii) retraction — DONE. Run 20260723T091446-35a1bde5
(PilotRunLedger id 38, preceded by a dry-run preview id 36): 12,880
CardPrintingTag votes deleted + the 24 votes B(i) had already flipped
= all 12,904 originally-staged votes accounted for exactly; 7,773
non-rescannable CardScanLog skips deleted (per-run 7,187 / 14 / 19 /
553, matching the prediction exactly); skipped_resolved_gate=0;
20,653 cards resynced via resolve_and_persist_printing(). Verified
end-state (live, this section): stage-d-join-key-v1 votes = 0;
non-rescannable join-key skips (4 target runs) = 0; deduction votes =
28,112 (intact); legacy pilot votes = 43,425 (intact); resolved cards =
3 (intact); eligible Stage D pool = 200,366 (relayed, not re-queried
this pass to avoid a full-catalog scan against the concurrently-running
§9(d) pilot).
§9(c) Bug-A forced-escalation sample — DONE. Runs
buga-sample-20260723T0927Z (PilotRunLedger id 39, 300-card
uniform-random extraction, seed 20260723, --no-shortcircuit, 85.4s,
7 workers) + buga-sample-verdicts-dry-20260723T093321Z (PilotRunLedger
id 40, verdict dry-run). Funnel over the 300-card sample drawn from the
17,531-card blank-tier-1 pool (§10): 300 fetched → 78 non-blank OCR text
(26.0%) → 78 parsed numbers → 65 set codes → 76 no-match votes → 1
genuine match (0.33%) → 223 skips. The 1 match: card 122326
("Ephemerate", Sketch Yumiko variant) → STA 68 — spot-checked and
confirmed correct. Wilson 95% extrapolation to the full pool: ~58
genuine matches [CI 10–327], qualitatively low-end likely (OCR noise
dominates the non-blank yield in spot-checks); full re-scan estimated at
~83–104 minutes.
Owner ruling (2026-07-23): Bug-A full re-scan DEFERRED to
post-pilot. Gap tracked, not dropped: (a) the signature query
regenerates on demand (17,531 cards at 2026-07-23T09:19Z, same
definition as §10); (b) the §9(d) pilot's own skip counters will surface
the blank-evidence abstentions, keeping the gap visible without a
separate tracker; (c) the post-pilot re-scan procedure must include a
state-clear step — this sample's no-text skips are non-rescannable
(unlike the no-evidence skip reason handled natively elsewhere in this
sequence), so the documented recipe is: re-extract with
--no-shortcircuit over the target cohort → clear stale skip state via
the reparse path (reparse_collector_evidence, same mechanism B(i)
used) → run a follow-up scoped Stage D pass (local_calculate_verdicts)
to actually cast votes. Recorded here as the standing recipe so it is
not re-derived next time.
§9(d) the 4c pilot dry-run — DONE, started immediately after (c)
above completed. Full run_id, statistics, per-channel breakdown, and the
read-only fallback-channel recovery methodology: see §9(d)'s own entry
above and reports/2026-07-23-4c-pilot-dry-run.md.
Full run reports, resource metrics, and per-run PilotRunLedger
counters for every run in this section:
data/2026-07-23-zeroing-and-buga-sample.md.
B(i)/B(ii)+B(iii)/(c)/(d) are now all DONE (§9's own entries carry
the live-verified detail and ledger ids); sample audit, --write, and
(e) remained NOT YET RUN as of this page's edit at the time this
section was written. §14 below carries the closing update: all three
completed the same day, 2026-07-23 — the full §9 sequence is COMPLETE.
The owner sample audit, the pilot --write, and the first
consensus_recompute --apply — the three steps §12/§13 left NOT YET
RUN — all completed 2026-07-23. The §9 fire sequence (as originally
ratified) is COMPLETE end to end and the gate is FIRED (§2). Full
detail, DB-verified counters, and the status=failed artifact
explanation:
data/2026-07-23-pilot-write-and-recompute.md.
This section was extended 2026-07-24 with five further
corrective/completion passes that ran the same night into the next day
(a lexicon-gate retraction, a marker reparse, an artist-credit fill, a
calculator re-pass, and a second consensus_recompute closer) — see
"Five further passes" below, and the full per-pass table immediately
below. All figures in this extension are DB-verified live against
mpcautofill_django/Postgres, queried 2026-07-24, unless marked
"docs-sourced" or "GAP."
Full per-pass table (chronological, all times UTC; condensed — dry-run rehearsals with 0 rows written are collapsed into their write row's note where the write immediately follows):
| # | pass | run_id(s) | dry/write | wall time | rows written | notes |
|---|---|---|---|---|---|---|
| 1 | Bug-B reparse write |
bugb-write-20260723T0905Z (id 37) |
write | 5.4s | 236 vote-affecting rows | preceded by 3 dry rehearsals (ids 33–35), see §12/§13 |
| 1r | Bug-B retraction |
20260723T091446-35a1bde5 (id 38) |
write | 189s | votes_deleted 12,880, cards_resynced 20,653 | see §13 |
| 2 | Bug-A forced-escalation sample | buga-sample-20260723T0927Z |
write (extraction) | 85.4s | 300 sampled, 1 genuine match | see §9(c)/§13 |
| 3 | 4c pilot dry-run |
pilot-dry-20260723T094518Z (id 41) |
dry | 7m37s | 0 | see §9(d)/§14 above |
| 4 | 4c pilot write |
pilot-write-20260723T1202Z (id 42) |
write | 59m23s | 130,210 (77,861 survive live today; 52,349 later retracted by pass 5) | see §14 above |
| 5 | Lexicon-gate retraction | 20260723T184334-2260859d |
write | 675.8s | retracted 52,349, unchanged 8,898 | see "Five further passes" below |
| 6 | Marker reparse write | 20260723T184318-6e0c73d9 |
write | 13.4s | 2,506 fields flipped (not votes) | see "Five further passes" below |
| 7 | Artist fill write | 20260723T213608-5479a0f0 |
write | 51m2s | 131,020 evidence fields filled | see "Five further passes" below |
| 8 | Calculator re-pass write | 20260724T001154-d3986cfc |
write | 52m41s | 1 vote; 52,348+16 skips, 117,442 to-review | see "Five further passes" below |
| 9 | Consensus recompute #1 (pilot-era) | no ledger row — GAP | — | not verifiable | 49,206 tag transitions (docs-sourced) | see "Pass-9 ledger gap" below |
| 10 | Consensus recompute #2 (closer) | 20260724T011448-32e08cc8 |
apply | 383.9s | 101,715 total (tag 0, artist 0, printing 1) | see "Five further passes" below |
Full per-pass detail (every dry-run rehearsal, exact considered/skip counts, and the reconciliation arithmetic) is DB-verified but not duplicated in full here — this table is the condensed record; the prose subsections below carry the load-bearing detail for passes 5–10.
Pilot --write — run pilot-write-20260723T1202Z, PilotRunLedger
id 42, started_at 2026-07-23T12:06:53Z, finished_at 2026-07-23T13:06:16Z.
DB-verified vote counts match the §9(d) dry-run's prediction exactly
on both channels: join-key 100,500 (39,253 match / 61,247 no_match,
same as the dry-run's would_cast figures), fallback 29,710 match (the
fallback channel's first production execution — its votes did not
exist before this run). Grand total 130,210 CardPrintingTag rows
written.
The status=failed / votes_written=None artifact: the ledger row
itself reads as a failed run with no vote count. This is NOT a
failed or partial write — it is a documented execution-harness
artifact. The authorization executor enforced a 1800s client-side
timeout on the launching connection; that timeout severed the
executor's own stdout stream at ~12:36Z, not the in-container process,
which kept running to completion (finished_at 13:06:16Z, ~30 minutes
past the timeout boundary) and hit a terminal-phase exception in its own
summary/ledger-update code, after the last vote write, not during it.
The DB-verified vote counts above — an exact match to the dry-run's
prediction on both channels — are the direct evidence the writes
themselves completed cleanly. Recorded here as
COMPLETE-BY-VERIFICATION, not as a failed fire, so this row is never
misread as one at a glance.
consensus_recompute --apply — executed 2026-07-23T13:25:37Z, exit
code 0. Predates the PilotRunLedger self-recording convention (no
ledger row exists for it — confirmed live, highest id remains 42) — this
was flagged beside the already-tracked §11 Stage C run-identity gap as a
"command predates the ledger convention" instance rather than a broken
command, and is now resolved going forward: consensus_recompute gained
its own self-recording ledger row (RUNNING at start, COMPLETED/FAILED at
end, per-family pairs_checked/rows_written/transitions counters)
as part of the command-lifecycle hardening pass, matching every other
Stage C/D pilot command's lifecycle. This specific 2026-07-23T13:25:37Z
run predates that fix and has no ledger row, immutably. Outcome: artist 7,130 pairs
checked, 0 transitions; tag 61,329 pairs checked, 49,206
None→UNRESOLVED materializations (one fewer than the 49,207 sized in
§6's earlier dry-run — organic interim resolution in the gap between
measurement and apply, not a discrepancy); printing resolved 3 → 4
(DB-verified at the time: Card.printing_tag_status distribution =
unresolved 218,309 / resolved 4 / no_match 1). This 4 was itself
provisional — a second closer pass the following night flipped it
back to 3; see "Five further passes" and "Topline end-state" below for
the corrected live figure.
Pass-9 ledger gap (re-confirmed 2026-07-24): the 49,206-transition
figure above remains docs-sourced, not independently re-derivable
from any ledger row — an exhaustive PilotRunLedger scan (64 rows, ids
1–65 minus an unrelated gap at id 10) found no row for this invocation,
and Card.tag_vote_statuses is a mutable JSONField overwritten by every
later recompute, not append-only, so the exact per-run delta can't be
reconstructed after the fact either. Best available corroboration (not
proof): the live tag_vote_statuses aggregate is 61,332 unresolved +
2 resolved_apply = 61,334 (card, tag) pairs, and the 2026-07-24 closer
pass's own dry-run rows report pairs_checked: 61,330 for the tag
family with zero transitions found — consistent with a large
one-time materialization having already happened before the closer ran.
This is a structural, pre-ledger-convention gap (the command has since
been hardened to always self-record), not a data-quality concern, and
nothing further closes it.
Five more passes ran after the original §9 sequence closed, all DB-verified live 2026-07-24. None of these were part of the ratified §9 order — each addressed a downstream correction or the pilot's own follow-on work, and none re-opens the gate question.
-
Lexicon-gate retraction — dry
lexgate-dry-20260723T1835Z(18:35:54→18:38:56, 181.9s), write20260723T184334-2260859d(18:43:34→18:54:50, 675.8s, 77.5 rows/s). Scope hash7662d0b34e17d2a6, considered 61,247: changed 52,349, retracted 52,349, unchanged 8,898. Retracted the 52,349 stale join-key votes the pilot--write(pass above) had cast against evidence a corrected lexicon gate has since superseded, rather than recasting them itself. This is the other half of the pilot--writereconciliation math already noted in §6: live join-key rows underpilot-write-20260723T1202Ztoday = 48,151 (down from the 100,500 cast), and 48,151 + 52,349 = 100,500 exactly. -
Marker reparse — two dry rehearsals (
20260723T164303-f8e07e2b, superseded, considered 132,674, would-flip 2,508;markerdry-20260723T1844Z, final, considered 133,354, would-flip 2,506 — the pool grew ~680 rows between them from concurrentrun_image_evidence_cohortcrash-drill passes), then write20260723T184318-6e0c73d9(18:43:18→18:43:31, 13.4s, 186.7 flips/s): 2,506 rows flippedlegal_line_proxy_marker_detectedfalse→true. Field-level correction, casts no votes. -
Artist fill — dry
20260723T204858-40e9408a(20:48:58→21:35:07, 46m10s), write20260723T213608-5479a0f0(21:36:08→22:27:10, 51m2s, 42.8 fills/s).backfill_modern_artist_names, considered 202,338: drywould_fill 131,020/ write filled 131,020 — exact match. FillsImageEvidence.artist_ocr_nameonly (never overwrites a non-blank value) — evidence, not a vote table. DB cross-check: current non-blankartist_ocr_namecount 144,680 = 131,020 (this run) + ~13,660 pre-existing (docs figure), consistent. -
Calculator re-pass — two dry rehearsals (
20260723T222926-022af8ca,20260723T231506-b5c52b16, ~43.6–43.7 min each, 0 written), then write20260724T001154-d3986cfc(00:11:54→01:04:35, 52m41s, 53.7 rows/s scan-log+vote total).local_calculate_verdictsre-run over the lexicon-gate-retracted pool: 1CardPrintingTagvote written; DB-verified viaCardScanLog, 52,348unknown-set-codeskips, 16no-evidenceskips, 117,442to-reviewskips (169,806 scan-log rows total). Nearly all of the retracted pool landed in abstention/review rather than a fresh resolvable vote. -
Consensus recompute #2 ("the closer") — dry
20260724T010750-21919b2d(139.2s), dry20260724T011030-f216fcc1(138.7s), apply20260724T011448-32e08cc8(01:14:48→01:21:12, 383.9s, 264.9 rows/s). Re-materialized tag/artist/printing consensus: tag 61,330 pairs checked, 0 transitions; artist 7,130 pairs checked, 0 transitions (every one of those 7,130 checked pairs still carries only a single vote, below the resolution threshold — see "Artist-consensus finding" below); printing 94,585 pairs checked, 1 transition (resolved→unresolved) — total_written 101,715. This is the pass that corrects the liveCard.printing_tag_statusresolved count from 4 back to 3 (DB-verified: unresolved 218,341 / resolved 3 / no_match 1, summing to the current 218,345-card catalog).
Despite the artist-fill pass writing 131,020 ImageEvidence.artist_ocr_name
values and the closer checking 7,130 CardArtistVote (card, votes)
pairs, all 218,345 cards currently read artist_vote_status=unresolved
(0 resolved, 0 unknown, 0 contested); Card.inferred_canonical_artist
is non-null on 0 cards. Each of the 7,130 checked pairs carries only
a single machine vote, which sits below the resolution weight/share
threshold in the owner-ratified vote-weight scenario matrix — see
reference/vote-weight-matrix.md
(narrated in theory.md §4/§7a) for the resolution
mechanics; not restated here. This is expected behavior under that
ruling, not a bug: a lone machine vote never resolves anything on its
own, by design.
Topline end-state (DB-verified live, queried 2026-07-24; updated 2026-07-26 post-pass — supersedes 2026-07-24 snapshot)
| Node | Value | Notes |
|---|---|---|
A — CATALOG (total Card rows) |
218,360 | +15 vs 2026-07-24 snapshot (218,345) |
| A1 — no phash yet | 15 | zero-evidence cards; equals A − B |
| B — phash extracted | 218,345 | |
| B1 — phash but no evidence | 0 | clean |
| C — evidence extracted | 218,345 | |
| C1 — evidence but no vote | 121,298 | 46.4% of evidence-bearing cards — largest funnel gap |
| D — carries ≥1 vote (distinct vote-holder cards) | 97,053 | distinct cards with at least one vote; the 2026-07-24 figure of 166,066 was total vote rows, not distinct card count — see vote tables below |
| F — RESOLVED | 3 | unchanged |
| H — NO MATCH | 9 | was 1 |
| G — UNRESOLVED | 218,348 | |
| I — REVIEW QUEUE | 3,595 | most-recent-per-card CardScanLog rows with skip_reason='to-review'; the 2026-07-24 figure of 134,370 was all-time (not dedup'd to most-recent) — definition differs |
-
artist_vote_status: all cardsunresolved(see "Artist-consensus finding" above).
anonymous_id |
rows |
|---|---|
stage-d-join-key-v1 |
48,754 |
local-ocr-v1 |
41,023 |
stage-d-fallback-v1 |
29,714 |
deductive-backfill-v1 |
28,112 |
local-fallback-v1 |
11,947 |
local-phash-v1 |
8,291 |
lands-artist-decomp-v1 |
1,488 |
| user UUIDs | 100 |
anonymous_id |
rows |
|---|---|
residual-classify-v1 |
6,144 |
art-hash-artist-v1 |
987 |
| user UUIDs | 6 |
anonymous_id |
rows |
|---|---|
layout-class-cast-v1 |
216,802 |
local-fallback-v1 |
53,966 |
residual-classify-v1 |
6,144 |
ai-art-detector-v1 |
1,183 |
| user UUIDs | 80 |
Five anonymous_id sources appear in the 2026-07-26 snapshot that were
absent from the 2026-07-24 topline or were previously lumped into coarser
buckets:
-
lands-artist-decomp-v1(1,488CardPrintingTagrows) — decomposes land-card art into artist-identity signals and casts a printing vote derived from that decomposition. -
art-hash-artist-v1(987CardArtistVoterows) — casts artist- identity votes by matching a perceptual hash of the card art against a known-artist hash index; previously the entireCardArtistVotenon-user pool was reported as a single "ocr 7,131" bucket. -
layout-class-cast-v1(216,802CardTagVoterows) — the borderless-attribute cast: classifies each card's layout and border style and casts a tag vote against the attribute-chip taxonomy seeded byseed_attribute_tags(tags 24–30); referenced as the "borderless cast" in the operational notes above. -
residual-classify-v1(6,144 rows in bothCardArtistVoteandCardTagVote) — runs a residual classifier over cards that other engines left unresolved, casting both an artist and a tag vote from the same classification pass; previously theCardArtistVotenon-user pool was reported as "ocr 7,131" without this sub-source breakdown. -
ai-art-detector-v1(1,183CardTagVoterows) — runs an AI-based art-style detector and casts a tag vote indicating whether the card art is AI-generated. -
frame-style-cast-v1(local_attribute_chip_cast, NEW 2026-07-30, 0 rows — never run) — reads storedImageEvidenceand casts theOld Border/Modern Borderchip vialocal_fallback.classify_frame_styleovercollector_line_collector_number+illus_anchor_fired. Zero image fetches. Gates on BOTHcollector_line_ocrandartist_ocr, becausebool(None)on the nullableillus_anchor_firedwould otherwise read as a real "no anchor" and classify everythingmodern. Reachable from the conveyor (stage_e_dispatch._run_stage_d) and from its own management command. Derivable population measured read-only 2026-07-29: 133,627Modern Border+ 9,006Old Border. -
bleed-edge-cast-v1(local_attribute_chip_cast, NEW 2026-07-30, 0 rows — never run) — same module, same pass, separate identity; castsappropriate-bleedat polarityNOT_APPLICABLEfor a confidentlytrimmedbleed_class. NEGATIVE-ONLY by design: absence of a vote is the documented convention for normal bleed, so a persistently low row count here is correct, not a coverage gap. The identity is separate fromframe-style-cast-v1precisely because of that — under one shared identity a card's frame vote would read as "already handled" and permanently strand its bleed chip. Derivable population 2026-07-29: 2,786.
Both were created because the 2026-07-29 composition audit found the only
casters for these two chips inside the live-fetch pilot and inside
image_evidence.extract_card_evidence, which had zero production callers —
so both chips sat at zero rows with nothing able to re-derive them. See
docs/features/printing-tags.md, "Who actually casts the attribute chips".
The question this section answers: for EVERY identity the roster tether
derives from code (19 today, 2 of them not real vote-casting calculators —
evidence-transfer-v1/question-feed-hypothetical-vote, see
CALCULATOR_ROSTER_ALLOWLIST in .github/scripts/docs_lint.py — leaving
17 real vote-casting identities), does a full-catalogue pass (run_pipeline,
or the streaming conveyor's own stage_e_dispatch._run_stage_d) actually
INVOKE it, verified by a real call site rather than a name match against
this file? A prior claim in circulation held that a pass casts on "10 of
roughly 28" channels — the derived roster is 17 real identities, not ~28,
and as of this pass 11 of 17 are invoked by a real call site (up from 7
before this pass — see below), 6 are deliberately unwired with a stated
reason.
Wired before this pass (7): stage-d-join-key-v1, stage-d-fallback-v1,
stage-d-illustration-v2, stage-d-slow-path-v1 (all four
local_calculate_verdicts.py, called directly by _run_stage_d), plus
layout-class-cast-v1/frame-style-cast-v1/bleed-edge-cast-v1 (PR #654,
_run_attribute_chip_casters). stage-d-slow-path-v1 is a router that
casts 0 votes by construction (see its own entry above) but IS invoked,
which is why it counts as wired here despite never appearing in a
vote-count table.
Newly wired by this pass (4) — ai-art-detector-v1,
lands-artist-decomp-v1, residual-classify-v1, art-hash-artist-v1, via
a new _run_evidence_only_calculators step in _run_stage_d. All four
read only already-stored evidence (ImageEvidence OCR fields,
Card.content_phash, already-resolved artist/printing chains) and fetch
no image — run_lands_identify/run_frame_mismatch_recovery are called
with every live-fetch budget forced to 0, which their own docstrings
already documented as "the scoped, genuinely free [...] path". Per-card
cost was not independently benchmarked beyond what each calculator's own
docstring already measures (run_lands_identify's cached
CandidateNameIndex build: 1.48s once per worker process, not per card;
every other cost is a plain indexed DB read) — no full-catalogue timing
run was performed as part of this pass, consistent with the constraint
that no management command or backfill ran against the live, contended
database while writing it. See the PR that shipped this wiring for the
per-channel dispatch-level tests (each proven to fire, and proven to fail
when the wiring call is removed).
Deliberately left unwired (6), with a reason and a tracked issue each:
-
art-edge-continuity-v1— genuinely FREE, but the module's own author gated it behind a stated validation precondition (agreement/false-positive rate against Scryfall's ownframe_effectsground truth) that has not run. See its own entry below and issue #721. -
deductive-backfill-v1,local-name-frequency-v1— read only already-stored data (no fetch/OCR/network), but NEITHER has acard_idsbatch-scoping parameter: each rebuilds a whole-catalogue in-memory index or census from scratch on every call, so wiring either into a 25-card-at-a-time conveyor as-is would mean paying a full-catalogue-scale cost on every single micro-batch. Classified EXPENSIVE on IMPLEMENTATION grounds, not data-source grounds — a real finding, not an oversight. See issue #722. -
local-ocr-v1,local-phash-v1,local-fallback-v1— these three identities' PRIMARY caster (run_pilot/run_fallback_for_card) computes its verdict from a real image for the first time, requiring a live CDN fetch plus tesseract (OCR) or a hash computation over pixels (phash); unlike the four newly-wired channels, there is no stored-evidence reconstruction path for the PRIMARY computation (some votes under these same identities are already cast for free as a side effect of OTHER calculators' own evidence-only paths, e.g.lands-artist-decomp-v1's OCR-resolved branch — that is not this identity's own engine running for free). See issue #723.
Backfill for the 4 newly-wired channels, against the EXISTING catalogue (not future passes, which the wiring above already covers), is a SEPARATE deliverable per this project's own practice (wiring ≠ backfill) — stated in the PR that shipped this wiring, not run as part of it, since a full-catalogue pass was live and the database was contended at the time.
Until 2026-07-29 this page enumerated eleven calculator identities.
Code declared fourteen. The three below appear nowhere above, and the
gate never audited them — not because it failed on them, but because it
was never pointed at them. Two of the three turned out to be effectively
DEAD in production, which is exactly the failure a coverage-based audit
cannot surface on its own: a calculator that produces no output produces
no divergence to explain, so it reads as clean by being invisible. The
list is now tethered to code by .github/scripts/docs_lint.py's
check_calculator_roster_tether() — every *_ANONYMOUS_ID declared under
MPCAutofill/cardpicker/ must have an entry here, and CI fails if one
does not. See documentation-process.md's
"Roster tethers" section for the general rule.
The tether itself then had the same defect one directory down — see "Calculator roster — the identity the tether itself could not see" below.
-
stage-d-slow-path-v1(SLOW_PATH_ANONYMOUS_ID,MPCAutofill/cardpicker/local_calculate_verdicts.py) — a ROUTER, not a voter. Working as designed. It has no printing to vote for, so it casts 0 votes of any kind by construction; its entire DB footprint isCardScanLogrouting markers carryingskip_reason='to-review'— 135,362 of them in prod. Those markers are what populates the human review queue (the same population the "REVIEW QUEUE" funnel row above dedups to its most-recent-per-card figure). Its absence from every vote table on this page is correct behaviour and not evidence of dormancy; it is listed here so that a reader auditing vote counts does not mistake "no votes" for "not running." -
stage-d-illustration-v2(ILLUSTRATION_ANONYMOUS_ID,MPCAutofill/cardpicker/local_illustration.py) — was DORMANT as-v1(3 votes in its entire existence); repaired and re-versioned 2026-07-29, not yet run in prod. It casts a printing vote from illustration identity (issue #507). The-v1root cause: its eligibility gate readImageEvidence.layout_classbelieving that field carries the card's faced-ness, when what it actually holds is a border colour (black 138,728 / borderless 72,603 / white 7,475 /''1,455 / silver 408), so the gate excluded 99.28% of every population handed to it — 3,409multi-faced-v1CardScanLogrows out of 3,426 scanned. The gate is now DELETED rather than repaired:CanonicalPrintingMetadata.face_illustrationsretains every face's ownillustration_id, so a back-face scan resolves to the artwork on the side actually scanned and there is no wrong-vote exposure left to guard. The-v2rename is load-bearing, not cosmetic —multi-faced-v1is not a rescannable skip reason and eligibility excludes cards carrying a non-rescannable scan log for the calculator's own identity, so a repaired-v1would never re-examine the cards it wrongly skipped. Still nothing MEASURED in prod: a read-only counterfactual replay over a 30,000-card sample of the 160,585-card-v2-eligible population (2,350 of them reach the calculator with artist OCR; the rest skip asno-artist-ocr) projects ~10,277 illustration votes and ~3,233 printing votes catalog-wide — but no-v2run has written a row. Do not read-v1's near-zero vote count as a measured statement about illustration matching's yield. -
art-edge-continuity-v1(ART_EDGE_ANONYMOUS_ID,MPCAutofill/cardpicker/local_art_edge.py) — DECLARED, NOT YET LIVE. Zero votes, zeroCardScanLogrows, and no runner calls it — by design, not by dormancy. It is the extended-art channel: a two-sample-point pixel comparison (card edge vs. the band adjacent to the art crop) that classifies a card image asframed/extended/openand would cast the pre-existing "Extended" attribute tag. It is listed here because this roster's whole purpose is that a calculator producing no output "reads as clean by being invisible" — so the identity is declared and recorded BEFORE it can write anything, the same reasoninglocal_fallback's skip-reason block gives for declaring its constants ahead of first write. The gate it has not yet cleared: the classifier is validated against constructed images only, never against real card images. Before it votes, run it over theImageEvidencerows whose confirmed printing carries Scryfall's ownframe_effectsextendedart(1,129 such rows in the 2026-07-28 join) and report agreement against that imported fact, plus the false-positive rate over a same-sized sample of confirmed non-extended black-bordered cards. That labelling is free and needs no human pass. Do not read its zero vote count as a measured statement about extended-art detection's yield — nothing has been measured yet. Tracked: issue #721. -
local-name-frequency-v1(NAME_FREQUENCY_ANONYMOUS_ID,MPCAutofill/cardpicker/local_identify_printing_tags.py) — ZERO output of any kind. Under diagnosis; may be retired. The name-frequency elimination pass deduces a match structurally, with no image fetch at all, for a name where exactly one printing is uncovered AND exactly one eligible card is unresolved. It has produced no votes and noCardScanLogrows — not even a skip log, which means it is not merely abstaining, it is not being reached. Whether it gets fixed or deleted is undecided; it is recorded here as a known hole rather than left off the page. Also missing acard_idsbatch-scoping parameter (2026-08-05 finding), which is a separate, more concrete blocker than "under diagnosis" alone conveys — see the "Wiring status" section above and issue #722.
The roster tether above was built to make "a vote-casting calculator
nobody documented" impossible. It then failed on exactly that, one
directory below where it looked: _declared_calculator_identities()
scanned MPCAutofill/cardpicker/*.py with a non-recursive glob, so
MPCAutofill/cardpicker/management/commands/ was never read. The
non-recursion was not a scoping decision — it was there to keep
cardpicker/tests/ fixture literals out of the roster, and it took the
whole subtree with it as a side effect. The scan is now recursive with an
explicit tests/ exclusion (.github/scripts/docs_lint.py's
_roster_source_files()), which is the same exclusion stated as a
decision rather than obtained as an accident.
-
scryfall-tagger-v1(SCRYFALL_TAGGER_ANONYMOUS_ID) — RETIRED 2026-07-30, together withPrintingTagVoteand its importer (management/commands/import_external_ip_tags.py, deleted). It is no longer a calculator identity, and the roster tether no longer derives it from the code, because there is no code. The entry is kept rather than deleted so that a reader meeting the string in an old report, run log or database column can find out what it was; everything below is written in the past tense and describes a design that never ran. Retirement rationale and the full behavioural record:features/printing-tags.md.It wrote zero rows, ever. It would have imported Scryfall Tagger's
art:external-ipcommunity art tag (Universes Beyond illustrations — Lord of the Rings, Doctor Who, Warhammer 40K) and cast machinePrintingTagVoterows against theexternal-iptag, at the PRINTING level (CanonicalCard), not the catalog-image level. It was designed to write BOTH polarities:APPLYfor positive Tagger matches,NOT_APPLICABLEfor confirmed printings absent from the positive set, with a printing that had no data at all abstaining rather than voting.source=DEDUCTIONwith its ownanonymous_idper the machine-caster convention — pure logical inference over already-trusted structured data, zero image inspection. Weight resolved toPRINTING_TAG_MACHINE_WEIGHT(0.5); the 2026-07-23 zero-weight override did NOT touch it (that override is scoped to source + thedeductive-backfillfamily + one frozenrun_id, all three together). Re-runs would have been idempotent via the(printing, tag, anonymous_id)uniqueness constraint, with retraction by the ordinarypurge_machine_votes --run-idmechanism. What the gate could say about it was always nothing, because it produced nothing — and its absence from every vote count on this page was never evidence that it worked.
Resolved 2026-07-30, previously an open owner call: PrintingTagVote
(models.py, added by PR #497) was a third vote family alongside
CardPrintingTag and CardTagVote, and it appeared zero times on
this page and zero times in theory.md — every
vote-population figure above was silent about it. The question recorded
here was whether it belonged inside this gate's scope or was deliberately
outside it. It has now been answered by removal rather than by scoping:
the model, its table and its only writer were retired on the owner's
ruling, the table having held 0 rows on production throughout its life.
The figures on this page are therefore complete as they stand, which they
were not while this question was open.
-
Scryfall cache lost on every container rebuild — issue #402,
discovered 2026-07-23 after deploy-2 rebuilt the django/worker images
(see
docs/troubleshooting.md's "local_calculate_verdictssilently runs with an empty back-face lookup after an image rebuild" entry for the full symptom/cause writeup):MPCAutofill/scryfall_cache/default_cards.json(~2GB, later measured ~558MB compressed — see that troubleshooting entry) lived inside the django container filesystem, not on a mounted volume, so a rebuild silently deleted it;local_calculate_verdictsran with an empty back-face lookup as a result (degraded, not crashed — easy to miss). Fix shipped: PR #412 ("Persist scryfall_cache volume and fail loud on a missing cache", merged 2026-07-24T09:07:13Z) makesscryfall_cachea named, persistent Docker volume mounted on both thedjangoandworkerservices, plus a fail-loudensure_scryfall_cache_present()guard at the start oflocal_calculate_verdicts.Command.handle()(raises rather than degrading silently, unless--allow-missing-scryfall-cacheis passed explicitly) — went live via deploy-3 (2026-07-24), the deploy following the one (deploy-2, 2026-07-23) that originally lost the cache. This closes the gap this note previously left open; not yet independently re-verified live in this pass (no fresh cache-presence check was run while writing this section) — flagged as documentation-sourced, not re-confirmed, should that distinction matter to a future reader. -
Border-tag seeding (7 created) — the
seed_attribute_tagsmanagement command (backing thelocal_layout_class_castborder-vote caster, issue #369/PR #375) was run to seed the attribute-chip taxonomy it depends on: Etched, Black/White/Silver Border, Old/Modern Border, Future Frame — 7 tags, none of which existed before. Verified live:Tagids 24–30 are exactly this set, sequential, with nothing else in that id range. Idempotent (seed_attribute_tagsis safe to re-run) — mentioned here as a one-time prerequisite step, not a vote or catalog-state change.local_layout_class_castitself went on to cast 216,811CardTagVoterows against this taxonomy (the "borderless cast," cited in §15's pass 5 for why that pass's tag-pairs count is so much larger than this section's own 61,330/61,334 figures).
-
Bug-A full re-scan of the 17,531-card blank-tier-1 pool. Wave-1
(10,437 cards, the top 4 sources) is now CLOSED (2026-07-24) — see
§15 for the full re-scan → reparse → lands → Stage D → closing-
recompute arc, DB-verified end to end with zero consensus transitions
resulting. Only the remaining ~6,535-card tail (16,972 − 10,437,
issue #418) stays open, and per the owner-ratified sequencing
(2026-07-24) it does not get a further batch pass — it is routed
to Stage E streaming's shakedown cohort instead; see
docs/proposals/stage-e-streaming.md§6/§7 (issues #153/#418). -
Moderation package sizing: DEFERRED (owner ruling, 2026-07-24),
pending further passes. At ruling time, the raw (non-dedup'd) union
of review-eligible cards across engines was 167,045 (76% of the
218,345-card catalog) — a materially larger and noisier number than
this section's own dedup'd 134,370
to-review-skip figure (see "Topline end-state" above), and the owner declined to size or build the moderation package against either number while wave-1/Bug-A closure (§15), the artist-consensus gap (see above), and the other passes still in flight keep moving the underlying counts. The 134,370 dedup'd figure remains available as a working number whenever this is picked back up — not re-derived here, just not acted on yet.
The 10,437-card wave-1 cohort — the top 4 sources of the 16,972-card
Bug-A blank-tier-1 pool (§9(c)/§13), the same cohort
docs/proposals/stage-e-streaming.md
§6 item 1 names as its own owner-ratified sequencing basis and §1
already measured for Stage C throughput (PilotRunLedger ids 70/75,
rescan-wave1-dry-20260724/rescan-wave1b-20260724) — was carried
through to a full, closed arc the same day: re-scan → reparse retraction
→ a full-pool lands pass → a wave-1-scoped Stage D vote → a closing
consensus_recompute dry-run confirming zero transitions. Every figure
below is DB-verified, keyed by run_id, in the same condensed-table
style as §14's per-pass table.
| # | pass | run_id | rows | notes |
|---|---|---|---|---|
| 1 | Wave-1 re-scan (Stage C) | rescan-wave1b-20260724 |
10,437 fetched | 3,282 (31%) recovered a parsed collector number, of which 2,526 further resolved a set code; 1,898 separately recovered an artist name via the artist-OCR channel. Throughput/resource detail already recorded in stage-e-streaming.md §1 (PilotRunLedger id 75, 3.351 cards/s, 3,115.1s) — this row adds the extraction-yield detail that brief didn't carry. |
| 2 | Reparse write | 20260724T125238 |
3,243 retractions |
reparse_collector_evidence --write over the wave-1 cohort's newly-recovered evidence, retracting stale no-match/skip state superseded by the fresh parse — the same mechanism as the earlier Bug-B write (§9(b)/§13), scoped to wave-1 here. |
| 3 | Lands full-pool write | 20260724T125355 |
2,701 new votes; 7,831 already-voted skips |
local_lands_identify --write, full eligible pool (not wave-1-scoped): 2,701 new votes cast, 7,831 honest no-op skips (PR #411's vote-collision skip-if-exists guard — a card already carrying a vote from the same anonymous_id, not a duplicate/error). Forced dry-run-before-write gate (#373) checked and passed ahead of this write. |
| 4 | Wave-1 Stage D write | 20260724T133819 |
919 votes |
local_calculate_verdicts --write scoped to the wave-1 cohort: 5 genuine matches + 914 no-match — closing the loop on exactly the cohort re-scanned in pass 1. |
| 5 | Closing recompute dry-run | 20260724T142039 |
0 transitions |
consensus_recompute dry-run: printing 97,212 pairs checked / 0 transitions, artist 7,130 pairs / 0 transitions, tag 226,974 pairs / 0 transitions. Tag pairs jumped from §14's 61,330/61,334 (the 2026-07-23/24 closer, still the live figure in "Topline end-state" above) to 226,974 here — consistent with the borderless-attribute cast (216,811 CardTagVote rows, verified earlier — see the "Border-tag seeding" operational note above) having landed a large new (card, tag) pair population in the interim, not a data-integrity concern. No --apply run was needed — the wave-1 votes/retractions above changed no card's resolved status on their own, consistent with the vote-weight gate's own "no machine tipping" rule (§3 item 3 / theory.md §4/§7a). Arc closed. |
Wave-1 topline: of the 10,437-card cohort, 5 cards resolved a
genuine new printing match via wave-1's own Stage D pass (pass 4);
2,701 additional cards (full-pool scoped, not exclusive to wave-1)
received a lands-identity vote from the same re-scanned evidence
(pass 3); the closing recompute (pass 5) confirms none of this moved
any card's resolved/no_match status. The wave-1 slice of the Bug-A tail
is closed — only the remaining ~6,535-card tail (16,972 − 10,437) is
still open, and per the owner-ratified 2026-07-24 sequencing (see "What
remains open" item 1 above) it does not get a further batch pass: it is
routed to Stage E streaming's shakedown cohort
(docs/proposals/stage-e-streaming.md
§6/§7, issues #153/#418).
Understanding the system
- Overview
- Documentation-Process
- Theory
- Identification-Pipeline
- Pipeline-Fidelity-Gate
- Federation-v1
- Vote-System
- Readiness-Audit
- License-Provenance
- Upstreaming-Conventions
- Drift-Log
- Upstream-Wiki-Drift
- Printing-Tags
- Catalog-Completion-Plan
- Moderation
- Card-DOM-API
- PDF-Generator
- Print-Export-Page
- Google-Drive-Connect
- Grid-Selector
- Image-CDN
- Local-File-Source
Using it
Operating it
Folded into other pages