Replies: 4 comments
|
Status (April 2026): The data sources system (tables, API, CLI) shipped in PRs #3571-#3572. The verification pipeline described here (structured data verification against source snapshots) extends the existing source-check system. See also #3589 for the UI integration plan. |
Decomposed into implementable issuesBased on the current infrastructure state (data_sources/source_snapshots tables shipped in #3571-3572, deterministic matcher working for grants, source parsers tested), here are 3 focused next-step issues: Phase 0: Quick fixes
Phase 3: Expand deterministic verification
What already shipped
Not yet filed (future phases)
|
|
Superseded by Discussion #3993 (Source-Check → Sourcing Rename — Implementation Plan). The architecture and naming decisions in this discussion have been consolidated into the 7-phase sourcing rename plan. Key changes: routes become |
|
Branch: Closing as superseded. The data sources / structured verification work is now consolidated under #3993 (Rename source-check → sourcing) as part of the broader cleanup. See also #4017 (Data integrity epic). — Discussion review 2026-04-08. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
916 of 2,486 source-checked records (37%) are "unverifiable." The source-check system verifies structured claims against source URLs using LLM reading comprehension, but this approach fails for tabular data sources (CSVs, APIs, HTML tables).
Breakdown by record type:
Core insight
For tabular data (grants from CSVs, personnel from databases), verification should be deterministic row-matching, not LLM reading comprehension. Parse the source CSV, find a matching row by key fields (grantee + amount + date). 100% accurate, zero cost.
Research
Investigated: DCAT (W3C), Frictionless Data/Table Schema, Schema.org Dataset, CKAN, Wikidata provenance model (stated in + record ID + retrieved date), git-scraping (Simon Willison), Datasette, RFC 7111 CSV fragment identifiers, Evidence.dev, ClaimDB benchmark (ICLR 2025), DVC, medallion architecture.
Key finding: No existing standard models provenance at the individual record level. Wikidata's three-part citation model is closest. Deterministic row-matching against tabular sources is more reliable than LLM verification (ClaimDB showed best models only reach 82.7% on large-table verification).
Architecture decisions
data_sourcestable (not a resource sub-table) — data sources identified by source ID, not URL; some have no stable URLsource_snapshotstable — stores raw content permanently in PG (~728MB/yr at worst); content-hash dedupImplementation phases
Phase 0: Quick fixes (3 independent PRs)
Phase 1: Schema foundation
data_sourcestable (id, format, access_method, record_type, fetch_url, column_mapping, verification_config, etc.)source_snapshotstable (data_source_id, snapshot_hash, raw_content, fetched_at; UNIQUE dedup)Phase 2: Grant pipeline integration
Phase 3: Deterministic verification
Phase 4: CLI + Dashboard
crux tb data-sources list|show|snapshot|verify|healthExpected impact
~400-460 of 916 unverifiable records fixed (~44-50%).
Full plan
Detailed implementation plan with file paths, migration SQL, and risk analysis:
~/.claude/plans/whimsical-noodling-hippo.mdAll reactions