Replies: 2 comments
|
Status (April 2026): Not started. The verification infrastructure is mature (unified tables, LLM source-checking, CLI orchestration), but verification remains post-hoc. No write-time verification gates exist. This is an architectural direction discussion — keeping open for future consideration. |
0 replies
|
Superseded by Discussion #3993 (Source-Check → Sourcing Rename — Implementation Plan). The architecture and naming decisions in this discussion have been consolidated into the 7-phase sourcing rename plan. Key changes: routes become |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context & Motivation
Today we have a mature verification infrastructure — unified
source_check_evidence+source_check_verdictstables, LLM-powered source-checking viacrux source-check orchestrate, semantic diff for wiki pages, citation accuracy tracking — but it all runs post-hoc. Data goes into the system first, verification happens later (if at all). The result:crux tb tablebase submitand records appear in production immediatelyThe core problem: verification is an afterthought, not a prerequisite. We need data and verification to be submitted together, atomically, as a single operation.
What triggered this
During a recent Apart Research enrichment session, an agent:
This is a 10-step manual workflow where verification is step 2 of 10 and easily skipped. We need it to be step 0.
Problem Statement
Three disconnected loops:
Should be one loop:
Current State: What Exists
crux source-check orchestrate)crux/lib/semantic-diff/)improve --applynextCheckDuefield)Related discussions: #2950 (unified verification architecture), #2928 (websites as data feeds), #3087 (external research patterns), #3119 (source-check expansion).
The TableBase Change Tracking Problem
This is the most urgent gap. TableBase data is PG-primary, so changes bypass git entirely:
When an agent submits personnel records, they go straight to production PG. No PR, no diff, no review, no verification. The only audit trail is
createdAt/updatedAttimestamps.Proposed Architecture: Verified-by-Default Data Pipeline
Design Principle
No data enters the system without a verification verdict. The verdict may be "verified", "unverifiable", or "contradicted" — but it must exist. Unverified data is rejected at the API level.
Component 1: API-Level Verification Enforcement
Modify all wiki-server submit endpoints to require verification alongside data:
The API:
verification(unless--skip-verificationflag, which is logged as an override and visible on dashboards)source_check_verdictsin the same transactionThis applies to all PG-primary tables: personnel, grants, funding rounds, investments, divisions, benchmark results, funding programs, policy stakeholders.
Estimated effort: ~4-8 hours. Modify each submit endpoint's Zod schema + handler. The
source_check_verdictswrite infrastructure already exists.Component 2: PR Manifests for TableBase Changes
Since PG changes bypass git, agents write a manifest file to their PR branch documenting what they submitted:
Gate check: New gate rule
tablebase-manifest-coveragethat:skippedverificationWhat this gives you:
Component 3: PG Audit Log
Add a
tablebase_audit_logtable for complete change history:Every write to any TableBase table also writes to the audit log. This gives git-like history without changing the write path.
Dashboard integration: New "Recent TableBase Changes" section on the system health dashboard showing recent mutations, grouped by session and PR.
Component 4: FactBase Gate Rule (Source-Verification Coverage)
For YAML-based data (FactBase facts, entity descriptions), add a gate rule:
pnpm crux w validate gate --fix # Now includes: source-verification-coverageThis rule:
packages/factbase/data/things/*.yamlagainst the base branchsource:URL againstsource_check_verdictsImplementation: Parse the YAML diff, extract fact IDs, query wiki-server verdicts API. ~4-6 hours.
Component 5:
crux researchCommand (Verified Research)One command replaces the 3-agent manual workflow:
Steps:
Then
crux enrich applyreads the bundle and writes to all layers — but only verified claims. Contradicted claims are logged. Unverifiable claims are flagged for human review.Key insight: Research and verification happen in the same step, not as separate agent dispatches. The LLM sees the source page and checks the claim in one pass, which is both faster and more accurate than checking after the fact.
Component 6: Continuous Re-verification (Phase 4 Simplified)
The simplest version of "website-as-data-feed" doesn't need website snapshots. It just re-runs source-check on existing records periodically:
Daily cron:
pnpm crux source-check recheck --needs-refresh --budget=5nextCheckDuefield insource_check_verdicts(Auto-recheck scheduler for verification verdicts #3021)Verdict flip detection: If a verdict changes from
verifiedtocontradicted:verification:contradictedDashboard widget: "Recently Flipped Verdicts" on Source Checks (E2200) showing verdicts that changed in the last 7 days.
This captures ~90% of the value of full website monitoring with ~10% of the effort. Full website snapshots + structured diff detection (Discussion #2928) becomes a later phase.
Phasing
crux researchcommand (Component 5)Open Questions
Should
--skip-verificationexist? It's pragmatic for bootstrapping/migration but creates a loophole. Options: (a) allow but log prominently, (b) allow but require human approval, (c) don't allow at all.What verdict threshold is acceptable? Should we allow
unverifiableclaims to be written? Many source URLs are JS-rendered (Verification pipeline can't verify JS-rendered pages #3008) and return empty content. Blocking these would prevent ~20% of legitimate data additions.Should contradicted claims be auto-rejected or flagged? Auto-rejection is safer but may block legitimate updates where the source has changed. Flagging requires human review bandwidth.
How should this interact with auto-update? The auto-update pipeline already fetches news and modifies pages. Should it also be required to verify its changes? This would slow it down significantly.
Manifest file lifecycle: Should manifests be permanent (growing archive in
data/tablebase-manifests/) or ephemeral (deleted after merge, with PG audit log as the permanent record)?Retroactive verification: 94% of existing TableBase records have no verdicts. Should we batch-verify existing data before enforcing the new rules, or grandfather existing records and only enforce for new submissions?
Relationship to Existing Discussions
source_check_evidence+source_check_verdictstables decided therenextCheckDueAll reactions