Replies: 5 comments
|
Status (April 2026): Audit document with fix proposals. Source-check orchestration and recheck commands exist. The specific fixes proposed here (usability improvements, verdict accuracy, UI gaps) should be decomposed into individual issues for implementation. |
Decomposed into IssuesAudit reviewed and decomposed into 3 focused implementation issues covering the highest-impact fixes: New issues from this discussion
Already addressed or covered by existing issues
Implementation order
|
|
Starting work on Fixes 5-6: Unverifiable subcategories in dashboard UI + tab restructure. Branch: |
|
Completed Fixes 1-5 (partially). PR #3735 adds dead link count to dashboard (Fix 5, step 1). Fixes 1-4 were already shipped in PR #3681. Remaining for this discussion:
Status update: 5/6 fixes done, 1 remaining (Fix 6 — large UI work). |
|
Superseded by Discussion #3993 (Source-Check → Sourcing Rename — Implementation Plan). The architecture and naming decisions in this discussion have been consolidated into the 7-phase sourcing rename plan. Key changes: routes become |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem Statement
The Source Check system has significant usability and data quality issues that undermine its value as a verification tool. An audit of the production data and codebase reveals six interrelated problems.
Production Data Snapshot (March 2026)
FactBase facts checked (1,444 of 2,417):
Other record types: Personnel 81% checked, Divisions 81%, Publications 77%, Investments 78%, Grants 4.4%, Benchmarks 0%, Wiki pages 0%, Citations 0%.
Source URL coverage in YAML: 70.9% of facts have a source URL. 29.1% don't (mostly
descriptionandnotable-forproperties). 38% of source URLs come from Wikidata/Wikipedia.Problem 1: "Avg Confidence" is a meaningless headline stat
The dashboard shows "89% confidence" as a headline number. This is the LLM's self-reported certainty about its own verdict, not an accuracy rate. High confidence on an "unverifiable" verdict means "I'm very sure I can't verify this" — the opposite of useful. Averaging across verdict types makes it doubly meaningless.
What users actually want to know: "How much of our data is verified correct?" and "How much is wrong?"
Problem 2: "Unverifiable" is a bucket for 4+ different problems
32% of checked facts are "unverifiable," but this conflates distinct situations requiring different remediations:
sourcefield at all (609 facts, 29% of total)Each needs different action. Currently they're all lumped together with no way to filter or prioritize.
Problem 3: Source URLs don't link to Resources
The
source_check_evidencetable has aresourceIdFK column. The API accepts it. The frontend (FootnoteTooltip.tsx) already renders "View source details" links whenresourceIdis present. ButstoreSourceCheckEvidence()inverdict-handler.tsnever callslookupResourceByUrl(). The entire infrastructure is wired up except for one 5-line change.The
claim-verificationhandler even hasresourceIdin its job params but doesn't pass it through.Source check pages show raw URLs as external links instead of linking to the rich Resource pages that already exist for many of these sources.
Problem 4: No broken link detection
The source-fetcher reads from the
citation_contentcache. If a URL isn't cached, it enqueues aresource-ingestjob (fire-and-forget) and returnsnot_cached. There's no feedback loop — the checker never learns whether the URL was dead.The
citation_contenttable DOES trackfetchStatus(200, 404, etc.), but source checks don't query it. Dead URLs just silently become "unverifiable" verdicts with no differentiation from other causes.Problem 5: Claims/entities are confusing in the dashboard
The Verdicts tab shows:
A user looking at this table can't tell what claim was checked or why it matters. There's no link back to the entity page where the fact appears.
Problem 6: Dashboard stats hide the real picture
The headline stats ("Total Records", "Checked %", "Avg Confidence", "Contradicted", "Needs Recheck") don't answer the fundamental question: "Should I trust the data on this wiki?"
There's no breakdown of:
Proposed Fixes
Fix 1: Resource Linking (Small effort, no dependencies)
In
storeSourceCheckEvidence(), auto-lookupresourceIdvialookupResourceByUrl()whensourceUrlis provided:Then update the source-check detail page to show "View resource details" links (same pattern as FootnoteTooltip). Run a one-time backfill for existing evidence.
Fix 2: Headline Stat Redesign (Medium effort, no dependencies)
Replace "Avg Confidence" with stats that answer the actual question:
Row 1 — Data Quality:
Row 2 — Action Items:
Key insight: separate data quality (of what we've checked) from data coverage (how much we've checked). "Accuracy rate" should be Confirmed / (Confirmed + Contradicted + Outdated), excluding unverifiable — that's a coverage issue, not a quality issue.
Confidence should move to a "System Health" tab, shown per-verdict-type.
Fix 3: Claims Display Improvement (Medium effort, no dependencies)
Restructure the Verdicts table columns:
{property} = {expectedValue} (asOf)— the actual thing being checkedexpectedValuealready present in evidence recordsFix 4: Broken Link Detection (Medium effort, depends on source-fetcher)
Step 1: Before calling the LLM, query
citation_contentfor the URL'sfetchStatus. If 4xx/5xx, skip the LLM call and storeverdict: unverifiable+fetchStatus: dead_link. Saves money and provides better data.Step 2: Add dashboard filter for "Source Status": live (2xx), dead (4xx/5xx), paywall, not fetched.
Step 3: Add Action Queue entry type: "N source URLs are dead — find replacements," grouped by entity.
Fix 5: Unverifiable Subcategories (Medium-large effort, needs schema + prompt changes)
Part A — Source-fetcher error classification: Store the fetcher error type (
dead_link,paywall,not_cached,topic_mismatch) as afetchStatusfield alongside the evidence. Dashboard can then show "32% unverifiable" broken down into causes.Part B — LLM prompt refinement: When source IS fetched but doesn't confirm the claim, distinguish "source doesn't discuss this topic at all" from "source discusses it but is ambiguous." Both are currently "unverifiable."
Fix 6: Full Dashboard Restructure (Large effort, builds on 2-5)
Tab 1: Data Quality (new default)
Tab 2: Coverage (improved)
Tab 3: Action Queue (improved)
Tab 4: System Health (new)
Suggested Implementation Order
Fixes 1-3 could ship together as a first PR. Fix 4-5 as a second. Fix 6 as a polish pass.
All reactions