Replies: 10 comments
|
Clarification on wiki page verdicts: We already have infrastructure for showing verification results on wiki pages. See Kalshi (E537) which shows "Citation verification: 42 verified, 7 flagged, 4 unchecked of 87 total" via the The existing citation verification system (
So the answer to the open question "Should verdicts show on wiki pages?" is: they already do for citations — the real task is unifying the backend so all verification flows through one system, not building new UI. |
Source-Check System Audit — Comprehensive Problems & ImprovementsGenerated 2026-03-25 from 3 parallel review agents (Playwright UX audit, deep code review, data flow audit). Executive SummaryThe source-check system has solid fundamentals — unified schema, working pipeline, multiple display surfaces — but suffers from fragmentation, inconsistency, and missing context. The display layer shows raw internal IDs where names should be, duplicates evidence without deduplication, and lacks claim context. The pipeline sends opaque stableIds to the LLM checker, produces duplicate evidence rows on re-runs, and doesn't set entityId for most record types. 32 code issues found across 14 files. 3 GitHub issues filed during this session (#3136, #3137, #3138). Part 1: Display Layer Issues1.1 Missing Claim Context (CRITICAL UX)Problem: The source-check detail page (
Impact: Users see "confirmed 95%" but have to read the full reasoning text to understand what was confirmed. Fix: The detail page should show a structured "Claim" section at the top:
1.2 Duplicate Evidence Rows (HIGH)Problem: Re-running source checks against the same record produces new evidence rows without removing old ones. Example: Filed: #3136 Partial fix in this PR: Evidence is now grouped by source URL and deduplicated in the display layer (same verdict + similar notes collapsed with "X similar checks" badge). But the underlying data still has duplicates. 1.3 Raw Internal IDs Shown to Users (HIGH)Problem: Multiple surfaces show opaque IDs:
1.4 Three Different Display Surfaces, Three Different UX (MEDIUM)
Impact: Users see different information and behavior depending on which surface they use. 1.5 Truncation Length Inconsistencies (MEDIUM)
1.6 Checker Model Shown as Raw String (LOW)
Part 2: Pipeline Issues2.1 LLM Checker Receives Raw StableIds (HIGH)Problem: The source-check prompt builder sends raw internal stableIds (e.g., "ULjDXpSLCI") as part of the record data. The LLM correctly flags these as unverifiable. Filed: #3137 Fix: 2.2 entityId Not Set for Personnel/Division Verdicts (HIGH)Problem: When submitting verdicts, the source-check runner doesn't set Filed: #3138 Fix: For personnel, set entityId = orgEntityId or personEntityId. For divisions, set entityId = organizationId. 2.3 No Deduplication on Re-runs (MEDIUM)Problem: Running source checks again produces new evidence rows. The Filed: #3136 2.4 Stale Checker Models in Evidence (MEDIUM)Problem: Old evidence rows use 2.5 Only 3 of 15+ Record Types Checked (LOW — already tracked)Covered in Discussion #3119. Part 3: Architectural Issues3.1 Three Separate Name Resolution Paths (MEDIUM)
If resolution logic changes for a record type, it must be updated in 3 places. Fix: Consolidate to one path — all resolution through the wiki-server API, with the proxy handling entity resolution locally as a fast path. 3.2 Client-Side Search on Server-Fetched Pages (MEDIUM)The
Fix: Move search to server-side (add 3.3 Duplicate VerdictBadge Definitions (LOW — partially fixed)Was 3 copies, now 2 (E2200 viewer imports from shared). The 3.4 Verdict Priority Defined in 3 Places (LOW — partially fixed)
Part 4: Missing Features4.1 No Recheck SchedulingThe 4.2 No Confidence Threshold FilterUsers can filter by verdict type but not by confidence range (e.g., "show only <50% confidence"). 4.3 No Bulk OperationsNo way to "recheck all contradicted verdicts" or "clear all outdated evidence" from the UI. 4.4 No Revision HistoryNo way to see how a verdict changed over time (was it confirmed, then contradicted after a recheck?). 4.5 No ExportNo CSV/JSON export for external analysis. 4.6 No "Needs Recheck" FilterThe stat card shows the count but there's no filter button to show only items needing recheck. Part 5: Security & Validation5.1 No Rate Limiting on Proxy APIs (LOW)The 3 proxy routes don't implement rate limiting. A malicious client could flood the wiki-server. 5.2 Parameter Naming Inconsistency (LOW)
Issues Filed This Session
Discussion Filed
Changes Shipped in This PR (#3135)
|
Playwright UX Audit — 63 Issues FoundFull audit document: see Critical Issues (fix immediately)
High Priority
Data Quality Issues Found in Live Data
Issues by Severity
|
Connection to Discussion #2950 (Unified Verification Architecture)The current implementation is a direct descendant of the design in #2950:
What's missing from the #2950 vision
Recommended next steps (based on #2950)
|
Integration Audit: Source-Checks Are More Connected Than ExpectedInvestigation revealed the source-check system is already widely integrated: Already showing verification badges (10+ locations)
Gaps remaining
Key finding:
|
Data Quality Audit: Critical Pipeline FindingsAdversarial review of live source-check data revealed systemic quality issues: ~70% False Positive Rate in ContradictionsOf 40 contradicted verdicts, only ~10-12 are genuine errors. Three LLM failure modes:
Verdict Aggregation BugStale contradictions persist after re-checks confirm data. Example: Elon Musk net worth has a recent "confirmed" check but overall verdict still shows "contradicted" from older check. The aggregation uses "worst verdict wins" instead of "most recent verdict wins". Web Scraping QualityMany "unverifiable" verdicts stem from the scraper returning nav/menu HTML instead of page body content (CSIS Wadhwani Center, ForecastBench). This is a systematic fetching problem. Personnel Name Resolution Still Broken (~60%)Despite wiki-server improvements, ~60% of personnel records show raw hash IDs (e.g., "boXCNeZK3w @ Epoch AI" = Maria de la Lama). Entity column is "-" for ALL 185 personnel. Coverage: 94% UncheckedOnly 805 of 12,439 records (6%) have any source checks. Largest gaps: grants (5,831), citations (3,490), wiki pages (835). Issues Filed
|
|
Status (April 2026): Partially done. Source-check orchestration exists ( |
Decomposed into 4 GitHub IssuesBroke this discussion down into concrete, implementable issues based on what exists vs what's missing: Issues Created
Already Covered by Existing IssuesSeveral concerns from the discussion are tracked elsewhere:
Recommended Execution Order
|
|
Superseded by Discussion #3993 (Source-Check → Sourcing Rename — Implementation Plan). The architecture and naming decisions in this discussion have been consolidated into the 7-phase sourcing rename plan. Key changes: routes become |
Update 2026-04-10: Phase 6 addition — unattended per-org curation runnerContext from the Anthropic pilot (Linear milestone "Anthropic → 100% green"): the manual playbook of "for one org, fix unverifiable records by swapping in better source URLs, re-verify, iterate" took Anthropic personnel from 41% → 100% verified-or-better. But this required close human oversight and doesn't scale to the hundreds of orgs on the wiki. What Phases 1–5 cover vs. what they don'tThe existing 5-phase plan in this discussion covers:
The missing layer: none of these phases describe the proactive, per-org curation loop that actually drives "confirmed+partial → 100%" for an entity. Phases 1–5 are reactive infrastructure (tables, schemas, schedulers, dashboards). The thing that takes an org and fixes its unverifiable records by finding better sources is not scoped anywhere. What the Anthropic pilot showedThe manual loop per org is:
Per-org cost: ~$0.25 for a 25-record org. Fully automatable once QUA-210 (agentic source-discover v2) and QUA-218 (fix auto-enqueue blind spot — shipped as PR #4113) land. Proposed Phase 6: Fleet curation runnerA new
Target: fleet-wide confirmed+partial rate improves ≥3% per daily run, converging toward 95%+ for the top 100 orgs within a quarter. Tracked as: QUA-220 — Dependencies
Open question for Phase 6Should the fleet runner also touch wiki pages (Phase 3 territory), or stay strictly on TableBase records? Pilot suggests starting with TableBase (grants, personnel, funding-rounds, divisions, investments) because (a) the verification is deterministic for many fields, (b) per-entity filtering is already supported in Status update on Phases 1–5
|
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context
The Source Checks dashboard (E2200) currently shows 725 verdicts across only 3 record types: fact (477), personnel (185), and division (63). Meanwhile, the wiki has ~700 MDX pages, ~1,000 entities, ~3,000 grants, ~1,000 publications, and thousands of other structured records — all with source URLs that could be verified but aren't.
The verification infrastructure is already built (unified
source_check_verdicts+source_check_evidencetables, LLM batch checker, dashboard with filters), but it's only being used for a fraction of the data. This discussion proposes expanding source-checking into a comprehensive data integrity system covering all three data layers.Current State
What Works
source_check_verdictsuses a polymorphicrecordTypefield — any record type workscrux tb source-check-records --table=<name>runs checksWhat's Checked Today
crux fb source-check— separate system, results in unified tablecrux tb source-check-recordscrux tb source-check-recordsWhat's NOT Checked
<R>linkssourceUrlurl,doi,arxiv_idsourcesourcesourcesourceUrlsourcesourceThe Wiki Page Gap (Critical)
Wiki pages are the primary user-facing content and currently have the least systematic verification. There are several existing systems that partially cover this, but they're fragmented:
Existing Wiki Verification Systems (Fragmented)
Citation accuracy (
citation_quotestable,/internal/citation-accuracy) — Checks individual[^N]footnote quotes against their source URLs. HasquoteVerified,accuracyVerdict,accuracyScorefields. This works but only covers citations that have been explicitly footnoted.Hallucination risk (
hallucination_risk_snapshots,/internal/hallucination-risk) — Scores entire pages for unsupported claims. Produces a risk score but doesn't verify individual claims.Citation health (computed at build time) — Counts how many footnotes have been quote-verified. Shown as a banner on wiki pages.
Semantic diff (
crux/lib/semantic-diff/) — Extracts factual claims before/after page edits and checks for contradictions. Runs during the improve pipeline but doesn't persist verdicts.What's Missing for Wiki Pages
<EntityLink id="X">components reference entities that may have been renamed, merged, or deleted.Proposed: Wiki Page Verification Pipeline
Treat each wiki page as a "record" in the source-check system:
recordType = "wiki-page",recordId = page slugsource_check_verdictstableProposed Architecture
Phase 1: Expand TableBase Coverage (Low-Hanging Fruit)
The
crux tb source-check-recordspipeline already works. Expand it to all registered tables:Phase 2: Unify FactBase Source-Checking
Currently
crux fb source-checkis a separate system that happens to write to the same verdicts table. Unify it:crux tb source-check-records --table=factpathcrux fb source-checkcommand (or make it an alias)Phase 3: Wiki Page Verification (Highest Impact)
This is the most ambitious but highest-value phase:
Step 1: Claim extraction
Step 2: Cross-reference checking
Step 3: Source verification
Step 4: Staleness detection
lastEditedvs. entitylastUpdatedvs. auto-update news datesPhase 4: Recheck Scheduling
The
nextCheckDuefield exists insource_check_verdictsbut nothing reads it:nextCheckDue < now()Phase 5: Unified Dashboard
The E2200 dashboard already supports this — it groups by
recordTypedynamically. As new types flow in, they appear automatically. Improvements:Cost Estimates
Open Questions
fieldName-level verdicts but nobody uses them. Worth the granularity?crux w improveautomatically check verdicts before/after editing a page?Next Steps
Looking for feedback on:
All reactions