Replies: 1 comment
Superseded WorkThe following issues have been closed as superseded by this plan:
Merged PRs that added Still-valid issues that will need terminology updates after the rename lands: #3444, #3900, #3548, #3609, #3546. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Overview
The current "source-check" system is the primary infrastructure for tracking where structured data comes from and whether it's accurate. But it has naming problems, architectural inconsistencies, and redundant mechanisms that make it confusing.
This discussion proposes:
sourcecolumn across ~20 tablessourceResourceIdFK columnsentity_resourcesas a separate system (entity-level, not record-level)Why "Sourcing" Instead of "Source-Check"
"Source-check" implies a post-hoc verification step. But this system also handles:
"Sourcing" covers all of these naturally — "we're sourcing this grant" works for both adding and verifying sources.
Precedent: Wikidata calls these "references" (attached to statements). W3C PROV uses "Entity/Activity/Agent." ClaimReview separates "Claim" from "ClaimReview." Our system is closest to ClaimReview's model — records are claims, sourcing evidence is the review.
Table Renames
source_check_evidencerecord_sourcessource_check_verdictssource_verdictsrecord.source(text)Why keep
sourceas the column name?After analysis, renaming
source→import_sourceon ~20 tables is high cost for marginal benefit. Thesourcecolumn is well-understood: "URL where this data came from." The real fix is ensuringrecord_sources(the many-to-many table) is the authoritative source list, with the inlinesourcecolumn being a convenience/seed value.Why remove
sourceResourceId?The
sourceResourceIdFK on 5 tables (personnel, grants, funding_rounds, investments, equity_positions) is redundant withrecord_sources.resourceId. The evidence table already links records to resources with richer metadata (verdict, confidence, extracted quotes). Maintaining both creates drift risk.Schema After Migration
Key design decisions
isPrimarystays as explicit boolean — deriving from highest confidence is fragile (shifts silently on re-check)discoveryMethodadded — single text column tracking how the source-record link was created ("import", "manual", "llm_discovery", "source_check")fieldNamekept — deterministic matcher uses it for amount/date verification. Default NULL = whole row.record_sourceskeyed on(recordType, recordId, COALESCE(resourceId, sourceUrl))— resource-based dedup when available, URL-based fallbackentity_resourcesstays separate — different concept (entity relationships vs record provenance)Migration Plan
~8-10 PRs over 6-10 weeks:
source-checkin code)ALTER TABLE RENAME)/api/sourcing+ server file renamesourceResourceIdfrom 5 tablesKey migration notes:
ALTER TABLE RENAMEis instant (metadata-only, no table lock)/source-checks→/sourcingfor SEOOpen Questions
sourcecolumn on records toimport_source? The plan above keeps it as-is. Renaming ~20 tables is high-cost. But the current name implies authoritativeness it doesn't have.discoveryMethodbe added torecord_sources? Low cost but increases schema surface./sourcing/grant/123vs/sources/grant/123vs/evidence/grant/123SourceCheckDot→SourcingDotorVerificationDot? The dot component is used in ~22 files.Tasks
Finish March 2026 rename: kill verification residue + remove sourceResourceId #4001 — Finish March 2026 rename: kill verification residue + remove sourceResourceId
Rename Drizzle symbols: sourceCheckEvidence → recordSources, sourceCheckVerdicts → sourceVerdicts #4002 — Rename Drizzle symbols: sourceCheckEvidence → recordSources, sourceCheckVerdicts → sourceVerdicts
All reactions