Replies: 1 comment
|
Branch: Closing as superseded by #3958 (Source-Check Quality Push), which explicitly pivots away from the claims-pipeline approach proposed here. The same verified-by-default vision is preserved in #3209 as the vision anchor. — Discussion review 2026-04-08. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Verified-by-Default Data Enrichment — Incremental Rollout Plan
TL;DR
94% of TableBase records have zero source-check coverage despite the verification infrastructure being fully built (source-check orchestrator, claims pipeline, inline verification, pre-submit gates, verification dots UI — all shipped, mostly unused). The source-check orchestrator handles all 13 record types today but has a routing bug that prevented type-specific runs. This plan fixes that bug, then systematically activates existing verification across record types, starting with personnel and expanding outward. Core code work is ~8-10 sessions; the remaining weeks are operational runs (backfills, triage, URL fixes). Phase 1 starts with a 1-line bug fix.
Problem
Data enters the system unverified. The verification infrastructure exists but agents bypass it — verification is optional and adds latency. Result: ~250 personnel records with 0 verdicts, ~3,000 grants mostly unchecked, and no enforcement preventing more unverified data from entering.
The fix isn't more infrastructure — it's activating what exists, backfilling coverage, and gradually tightening enforcement.
Current State
crux tb verify)tb verify personneldoesn't passtableparam → verifies ALL typesproposed_claims, claim-verification worker)pre-submit-verification.ts)write-inline-verdicts.ts)/api/resources/suggest)next_check_duefield)Key Decisions
Decision: Orchestrator is the primary verification path for both backfill and ongoing verification
Decision: Backfill first, enforce gradually
--forceoverride required)Decision: Focus on 5 high-value record types
Architecture
crux/commands/source-check-orchestrate.tstable: mappedto orchestrator optionscrux/commands/source-check-orchestrate.test.tstb verify personnelonly collects personnel recordscrux/lib/source-check/item-collectors.ts&&to||in stableId skip guard (line ~304)crux/lib/source-check/item-collectors.test.tscrux/lib/source-check/orchestrator.tscrux/lib/source-check/deterministic-matcher.tscrux/lib/source-check/record-fields.ts.claude/commands/review-tablebase-enrich.mdtb verifyafter submit" step to enrichment skillcrux/validate/validate-verification-coverage.ts--forceto bypassNo new PG tables. One small migration only if claims stretch goal is pursued (otherwise none).
Implementation Phases
Phase 1: Personnel Verification (Weeks 1-6) — 2-3 code sessions + ops
Goal: Personnel records show verification dots. Unverifiable rate drops below 15%.
[code]: Fix routing bug insource-check-orchestrate.ts— addtable: mappedto options (line ~115). Add test. RunWIKI_SERVER_ENV=prod pnpm crux tb verify personnel --dry-runto assess baseline.[ops]: Runpnpm crux tb verify personnel --batch(~$1.25). Report verdict distribution.[code]: Tighten name-resolution guard: change&&to||atitem-collectors.ts:304so records with unresolvable person names are skipped even if org name resolves. Add test.[ops]: Re-run on previously-skipped records.[ops]: Audit "unverifiable" results. Categorize: dead URL, wrong page, JS-rendered, paywall.[ops]: Fix source URLs usingsuggestResources(). Register top-50 org team pages inwebsite_source_pages. Re-run source-check on updated records.[ops]: Report: personnel verdict distribution before/after. Decide dot visibility threshold (Open Question Basic Search #2).Quality gates:
tb verify stats --table=personnelshows verdicts. Dots visible on 3+ org pages.Exit criteria: >60% of personnel have verdicts. Unverifiable <15%.
Phase 2: Grants + Workflow Integration (Weeks 7-14) — 3-4 code sessions + ops
Goal: Grants verified. Enrichment skill runs verification after each submission.
[ops]: Runpnpm crux tb verify grant --dry-run. Assess: how many have source URLs? What's the deterministic-match rate?[ops]: Run grants backfill in test batch (100 records). Verify deterministic matcher works on tabular grant sources.[code]: Update enrichment skill (.claude/commands/review-tablebase-enrich.md) to add post-submit verification step: aftercrux tb submit, runcrux tb verify --type=record --table=<type> --entity=<id>. This gives immediate feedback on what was just submitted.[code]: Add batch chunking to orchestrator for 3,000+ record sets (split into pages of 500, sequential API calls). Run full grants backfill (~$5-10).[ops]: Triage contradicted grants. Fix source URLs. Re-run.[code]: Extend deterministic matcher to investments. Add stableId skip guards for funding-round, benchmark-result inrecord-fields.ts.[ops]: Run investments backfill (~1,000 records). Run funding-rounds backfill (~500).Quality gates: Grants >80% coverage. Deterministic matcher handles investments. Enrichment skill includes post-submit verification.
Exit criteria: Financial record types show dots on directory pages. <5% contradicted.
Phase 3: Remaining Types + Advisory Enforcement (Weeks 15-22) — 2-3 code sessions + ops
Goal: All priority types verified. Advisory warnings when records lack verification.
[ops]: Publications backfill (DOI resolution for deterministic checks where possible). FactBase facts backfill (crux tb verify --type=fact).[code]: Add advisory enforcement — enrichment skill logs a warning when records are submitted without runningtb verifyafterward. The gate check (validate-verification-coverage.ts) starts printing warnings for manifests with zero verification.[ops]: Benchmark results, entity events, remaining types — run orchestrator ad-hoc.[ops]: Global audit: coverage across all types. Spot-check 10 "confirmed" verdicts per type for false-positive rate.Quality gates: All 5 priority types >50% coverage. False-positive rate <10%.
Exit criteria: E2200 dashboard shows coverage across all major types. Advisory warnings active.
Phase 4: Soft Enforcement (Weeks 23-30) — 1-2 code sessions + ops
Goal: Submitting without verification requires explicit opt-out.
[code]: Soft enforcement invalidate-verification-coverage.ts— fail gate check for manifests with zero verification unless--forceis used. Log all--forceusages.[ops]: Monitor: do any agent sessions hit the soft block? Fix workflows that don't include verification.[ops]: First quarterly recheck — run orchestrator on all types, targeting stale verdicts (>90 days old). Fix any verdict flips.[ops]: Personnel enrichment sessions for under-covered orgs (using updated skill with post-submit verification).Quality gates: Soft enforcement active. <5
--forceoverrides per week.Exit criteria: All priority types >70% coverage. Enrichment sessions routinely include verification.
Phase 5: Hard Enforcement + Steady State (Weeks 31-50) — background ops
Goal: Verification is structural norm. System self-maintains.
[code]: Hard enforcement for personnel (>70% coverage).requireVerification=trueon personnel sync endpoint for agent sessions.[ops]: Expand hard enforcement to each type as it reaches >70%. Continue enrichment.[ops]: Second quarterly recheck. Assess whether automated scheduling is justified by flip rate.[ops]: If flip rate >5%, implement simple recheck cron. Otherwise, continue quarterly manual runs.[ops]: Annual retrospective: total coverage, cost, false-positive rate, flip rate. Evaluate stretch goals.Quality gates: Hard enforcement for personnel. Quarterly rechecks running.
Exit criteria: 0 unverified personnel records submitted in last 30 days.
Stretch Goals (evaluated after Phase 3)
These are valuable but not essential for the core verification pipeline:
Claims pipeline integration: Wire
proposed_claimsinto enrichment as an audit trail for new writes. Requires solving the sequencing issue (claims need to reference records that don't exist yet — either use placeholder IDs or verify claims before record creation). Evaluate at Week 15 based on whether the audit trail gap is actually a problem.crux research <entity>command: Unified discover → resources → verify → write pipeline. Evaluate at Week 20 based on whether the enrichment skill update (Phase 2, Week 9) is sufficient.Website-as-data-feed extraction (LLM fact extraction from website snapshots into FactBase #3655): LLM extraction from org team page snapshots. Separate feature — evaluate at Week 50 retrospective.
Automated recheck scheduling (Add automated recheck scheduling via GitHub Actions #3661): Cron-based recheck using
next_check_duefield. Evaluate at Week 40 based on quarterly recheck flip rate.Scope Cuts
Quality & Verification Infrastructure
recordType === 'grant'guards and expand all.tb verify statsbefore/after.tb verify <type> --dry-run --limit=5spot-check.WIKI_SERVER_ENV=prodcommands.Risks & Mitigations
tb verify personnelrouting bug&&to||guard change + testOpen Questions
Dot visibility threshold: Should dots be hidden for a record type until >50% of that org's records are checked? Showing 5 green dots out of 27 may mislead users. Decide after Phase 1 data.
Unverifiable records under enforcement: When hard enforcement activates, should "unverifiable" records be allowed? Many legitimate records have PDF/LinkedIn sources that can't be machine-read. Proposed: allow if source URL is a valid Resource record.
Recheck frequency: Quarterly manual sufficient, or does flip rate justify automation? Decide at Week 40 with data.
Claims pipeline value: Is the orchestrator-only path sufficient, or does the audit trail from claims justify the integration complexity? Decide at Week 15 with operational experience.
Rejected Approaches
crux researchcommand first: 5-7 sessions before any visible value. Backfill delivers verdicts in Week 1.Related Issues & Discussions
Red Team Log
Round 1 — Technical Critic
Blocking (fixed):
tb verify personnelrouting bug — command verifies ALL types, not just personnel. Themappedvalue is computed but never passed astablein options. Fixed: moved to Week 1, 1-line fix.Significant (fixed): Claims pipeline requires running job worker — added deployment audit to stretch goals prerequisite.
Significant (fixed): Enforcement at Week 47 too late — moved to gradual ramp starting Week 18.
Significant (addressed): Two verification paths with no dedup — simplified to orchestrator-only as primary path. Claims moved to stretch goal. Dedup is moot if only one path is active.
Significant (noted): No orchestrator integration test — routing fix test added to Week 1.
Minor (fixed):
isResolvableNameguard description clarified —&&to||.Round 1 — UX Critic
High (addressed): 38% unverifiable orange dots make site look worse — added Open Question #1 (dot visibility threshold), URL remediation before dots become meaningful.
High (fixed): 37-week gap before single command — moved
crux researchto stretch goal. Enrichment skill update at Week 9 covers the workflow with existing commands.Medium (fixed): Enforcement cliff — 3-stage ramp (advisory → soft → hard).
Round 1 — Scope Critic
Applied: Compressed 10 phases → 5 phases. Cut website extraction, automated recheck,
crux researchcommand to stretch goals. Focus on 5 types not 13.Applied: Made claims pipeline a stretch goal, not core. Orchestrator alone covers the use case.
Debated: Scope critic suggested 14 weeks total. Plan retains 50-week timeline because: (a) user explicitly asked for 50-week trajectory, (b) core code is ~10 sessions but operational runs (backfills, triage, URL fixes) take real elapsed time, (c) Phase 5 is genuine steady-state maintenance.
Round 2 — Revision Agent
Addressed:
verification_sourcecolumn dropped — avoids unnecessary migration. Existingchecker_modeldistinguishes paths if needed.Addressed:
submit-verifiedsequencing issue (claims need target_id) — moved claims integration to stretch goal rather than trying to solve the design problem now.Addressed: Enforcement mechanism checks manifests vs
tb submit— clarified that enforcement applies to both gate check (manifests) and enrichment skill (workflow).Noted: Reframing agent argued plan could be 6 weeks if claims are dropped entirely. This is true for the code work, but operational execution (running backfills across types, fixing URLs, triaging results) takes real calendar time. The 50-week plan reflects a background project cadence (1-2 hours/week), not padding.
Round 2 — Implementability Agent
Confirmed: Phases 1-2 are concrete and implementable. Phase 1 (Weeks 1-2 code parts) could be one session.
Addressed: Weeks marked as
[code]vs[ops]for clarity.Noted: Phases 4-5 are vague because they're operational, not engineering. This is appropriate — they don't need code specs, they need execution.
Fixed:
crux/tablebase/directory doesn't exist — submit-verified moved to stretch goals, directory issue moot.All reactions