Replies: 4 comments
Updated coverage data (actual totals)The gap is much worse than initially estimated:
Grants are the biggest gap by far — 5,904 records with only 4% checked. Funding programs and funding rounds have zero coverage. The Coefficient Giving funding tab shows this clearly: every single grant row shows an unverified gray dot. This changes the cost/strategy discussion significantly. At ~5,900 unchecked grants, the batch API cost for a full backfill is non-trivial. |
Session results (2026-04-03)Coverage before/after
What was done
Remaining gapsAll remaining gaps are
The daily cron will process these via LLM fallback (~$0.002/record). Data issues found
|
Status Update (2026-04-04)PRs merged since last session updateSince the Apr 3 session results, 10+ related PRs have landed: Direct source-check improvements:
Display/rendering fixes:
Issue tracker
AssessmentThe big win from this strategy was the grant backfill (4% → 94% at zero LLM cost via deterministic row-matching). Verification dot UX has improved significantly: #3774 fixed the misleading display where "never checked" looked identical to "unverifiable", #3816 extended dots to FactBase facts, and #3788 laid groundwork for better source caching with a proper Answers emerging for original questions:
Next priorities:
|
Uh oh!
There was an error while loading. Please reload this page.
Context
Source-check verification dots now appear across all directory pages (personnel, grants, divisions, funding programs, funding rounds, investments). But the actual coverage is spotty:
Current state (2026-04-03)
Global source-check stats:
Per-org example (Anthropic): Only 5 of 27 personnel records have been checked (19%). All 5 came back "unverifiable."
Verdict distribution across all 3,044 checks:
Related open issues
Questions to decide
Priority ordering: Should we backfill unchecked records first (Run source-check backfill across all TableBase record types #3667) or improve unverifiable ones (Improve source-check coverage for 869 unverifiable records #3609)? Backfilling gives users coverage expectations; improving unverifiable gives better data quality.
Automation cadence: Should source-checks run on a schedule (daily/weekly via GitHub Actions, Add automated recheck scheduling via GitHub Actions #3661), on record creation, or manually? The batch API is cheap (50% discount) but still has cost.
Target coverage: What is the minimum acceptable coverage before we show verification dots on a page? Showing dots when only 5/27 records are checked is arguably worse than showing nothing — it implies the unchecked records are fine (especially given RecordVerificationDot: no visual distinction between 'never checked' and 'unverifiable' #3740, where "never checked" and "unverifiable" look identical).
Source URL quality: 30% of checked records come back "unverifiable." Is the problem bad source URLs, paywalled content, or records that genuinely cannot be verified? Should we invest in better source discovery (Auto-suggest better source URLs when source-check returns unverifiable #3548) before running more checks?
Cost budget: At current Haiku pricing, what is the estimated cost to reach 95%+ coverage across all record types? Is this a one-time backfill or an ongoing expense?
Display threshold: Should we hide verification dots for a record type on an org page until some minimum % of that org records have been checked? E.g., dont show personnel dots unless >=80% of this orgs personnel have been checked.
All reactions