You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Deterministic Wikidata Source-Check Matcher — Implementation Plan
TL;DR
~425 FactBase facts sourced from Wikidata are verified by an LLM reading noisy HTML — expensive ($4.25/week) and less reliable than a direct API comparison. This plan adds a deterministic Wikidata matcher: extract Q-numbers from source URLs, batch-fetch structured claims via wbgetentities, and compare property values directly. Phase 1 covers 12 property types (~414 facts, 97%). Numeric properties (revenue, headcount — 11 facts) are deferred to a follow-up. Top open question: whether lazy label resolution is fast enough or if batch pre-fetch for labels is needed.
Problem
The weekly source-check spends ~$4.25 and ~425 LLM calls verifying Wikidata-sourced facts by converting HTML to text and asking Claude Haiku to judge. Wikidata has a structured JSON API that returns exact property values. The deterministic matcher for grants (PR #3571) proved the "deterministic first, LLM fallback" architecture works — this extends it to Wikidata facts.
Current State
Asset
Status
Gap
crux/lib/source-check/item-verifier.ts
Has tryDeterministicMatch() for grants (line 227)
No Wikidata matching — facts fall through to LLM
crux/lib/source-check/fuzzy-match.ts
nameMatches(), amountMatches(), dateMatches()
Missing urlMatches() for website/wikipedia-url
crux/commands/wikidata-enrich.ts
WIKIDATA_PROPERTY_MAP (10 PIDs, lines 60-154)
Enrichment direction only; verification needs reverse map
Duplicated in people/enrich.ts — shared API code extracted into matcher
packages/factbase/data/things/*.yaml
425 facts, 138 unique QIDs, 137 entities
QIDs only in source URLs; ~136 facts have trailing commas in URLs
crux/lib/source-check/source-check-recheck.ts
Recheck always uses LLM
Would cause verdict flip-flopping without Wikidata integration
Proposed Approach
Add tryWikidataMatch() in item-verifier.ts that intercepts fact items with Wikidata source URLs. Uses wbgetentities API (batched 50 QIDs/call) with lazy memo cache for label resolution. Falls through to LLM on any failure.
Per-fact verification (tryWikidataMatch):
1. Extract QID from source URL via regex /Q(\d+)/
2. Fetch entity claims from memo cache (or wbgetentities API)
3. Map FactBase property → Wikidata PID
4. Filter claims: preferred rank first, then normal. Never deprecated.
5. Extract + normalize Wikidata value
- Entity-refs (Q84): resolve label via memo cache (or wbgetentities)
- Sitelinks (wikipedia-url): construct URL from title
6. Compare against fact value using typed comparison
7. Return VerifyResult or null (→ LLM fallback)
Batch path optimization: In orchestrator.ts, scan items before verification loop → collect unique QIDs → batch pre-fetch entities (3 calls for 138 QIDs). Individual facts then hit the memo cache with zero API calls.
Realtime path (--entity=X): Lazy per-fact fetching with memo cache. First fact for a QID triggers one API call; subsequent facts for the same QID are free. Typically 1-2 API calls per entity.
Key Decisions
wbgetentities not SPARQL: Simpler, supports batching (50 QIDs/call = 3 calls for all entities), no SPARQL endpoint queue. SPARQL stays in enrichment commands.
Lazy memo cache for labels, not orchestrator pre-fetch: Lazy approach avoids orchestrator coupling. Each label resolution is one wbgetentities call, cached for reuse. Worst case for full corpus: ~80 individual label calls vs 6 batched. Acceptable for a weekly job.
New wikidata-matcher.ts, not extending deterministic-matcher.ts: Different data shape (structured claims vs tabular rows). Clean module boundary.
12 properties in v1, numeric deferred: revenue and headcount need Wikidata quantity parsing ("+500" with unit URIs, currency handling). 11 facts. Deferred to follow-up.
No circuit breaker: try/catch returning null is sufficient for a weekly batch job. LLM fallback is the circuit breaker.
No separate shared API module: Single consumer (the matcher). Extract when a second consumer appears.
Test each property type with mock API fixtures. Edge cases: multi-valued claims, missing properties, deprecated claims, trailing-comma URLs, sitelink title with special chars, non-English label fallback
crux/lib/source-check/item-verifier.ts
modify (~15 lines)
Add tryWikidataMatch() call for kind === 'fact' items with Wikidata URLs, before LLM path
crux/lib/source-check/orchestrator.ts
modify (~20 lines)
Batch path: scan items for QIDs, pre-fetch entities via batched wbgetentities before verification loop
Exit criteria:WIKI_SERVER_ENV=prod pnpm crux fb source-check --entity=80000-hours --verbose shows all facts with checkerModel: 'wikidata-api'.
Rollout sequence: Per-entity smoke tests → compare verdict distribution → full corpus run. If >20% contradicted rate (vs ~5% from LLM runs), investigate before proceeding.
Scope Cuts
Numeric properties (revenue, headcount): 11 facts. Need Wikidata quantity parsing ("+500" with unit URIs, currency comparison). Follow-up issue.
Enrichment consolidation: Triplicated API code in orgs.ts/people/enrich.ts/wikidata-enrich.ts works fine. Consolidate when a second consumer of shared API client emerges. Follow-up issue.
Stored wikidata-qid property: QIDs extracted from source URLs at verification time. Persistent QID field is a separate enhancement.
Auto-correction of contradicted facts: Matcher stores verdicts but doesn't auto-fix. Improve pipeline handles corrections.
Dashboard UI changes: Existing /source-checks page displays all checker models. Only change: add 'wikidata-api' case to formatCheckerModel() (~1 line).
Quality & Verification Infrastructure
Tests:wikidata-matcher.test.ts (~200 lines). Mock Wikidata API at HTTP level. Test each of 12 property types. Edge cases: multi-valued P112 (4 founders), missing P169, deprecated claim rank, trailing comma URLs, sitelink title with parentheses ("Asana (software)"), entity with no English label (fallback to null → LLM), multi-QID entity (facts from different QIDs on same FactBase entity).
Deploy steps: No new env vars, no migrations. Standard PR merge. Wikidata API is public (no auth).
Documentation: N/A — the matcher follows the existing tryDeterministicMatch pattern. Property map in wikidata-matcher.ts is self-documenting.
Monitoring: After first weekly run, compare verdict distribution (confirmed/contradicted/partial counts) against previous LLM-based run. If contradicted count increases by >2x, file an issue.
Risks & Mitigations
Risk
Severity
Mitigation
Entity-ref label resolution (8/12 properties)
Significant
Lazy memo cache. nameMatches() handles normalization. Fallback: no English label → return null → LLM.
Multi-valued properties (founder, education)
Significant
OR semantics: match any Wikidata value. partial if some found, contradicted only if none found.
Trailing commas in ~136 source URLs
Minor
Regex /Q(\d+)/ extraction (not URL parsing). Test with exact format.
Deprecated Wikidata claims (stale CEO, old HQ)
Significant
Filter: preferred rank first, then normal. Never deprecated. Test with fixture containing deprecated + preferred claims for same PID.
Wikipedia URL from sitelinks (title, not URL)
Significant
Construct URL: https://en.wikipedia.org/wiki/${title.replace(/ /g, '_')}. Test with parenthesized titles. Fetch with props=claims|sitelinks.
Verdict flip-flopping on rechecks
Significant
Integrate tryWikidataMatch() into recheck path. Same matcher runs on rechecks.
Systematic bug overwrites 414 aggregate verdicts
Significant
Rollout: per-entity first, compare distributions, then full run. Old LLM evidence preserved in evidence table (different checkerModel in dedup key).
API downtime during weekly run
Minor
try/catch returns null → LLM fallback. No worse than today.
Open Questions
Lazy vs batch label resolution: Lazy is simpler (no orchestrator coupling) but makes more individual API calls (~80 vs ~6 batched). For a weekly job, is the difference (~60 seconds) worth the complexity? Current plan: lazy.
Rate limit setting: Existing enrichment uses 1s/request. Wikidata allows ~200 req/min (300ms). Should the matcher use the faster rate? Current plan: match existing 1s setting for safety.
Rejected Approaches
A: Minimal (skip entity-ref properties): 76% coverage. The effort difference is small (~50 lines for label resolution).
C1: Enrichment-as-Verification: Couples verification to enrichment correctness. Uses SPARQL (heavier).
C2: Snapshot and Match: Adds infrastructure (cache table, refresh job) for a problem batched wbgetentities solves in 3 calls.
C3: Wikidata Watches Us: Needs nonexistent import timestamps. First run can't shortcircuit.
D: 4-phase incremental: More total effort for same outcome.
Red Team Log
Technical critic — 2 blocking, 10 significant:
[FIXED] Blocking: realtime execution path not addressed → added tryWikidataMatch in verifySingleItem
[FIXED] Blocking: recheck system causes verdict flip-flopping → integrated matcher into recheck path
[FIXED] Claim rank filtering (preferred > normal, never deprecated)
[FIXED] Wikipedia URL from sitelinks needs URL construction from title
[FIXED] Trailing commas affect ~136 facts, not ~25
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Deterministic Wikidata Source-Check Matcher — Implementation Plan
TL;DR
~425 FactBase facts sourced from Wikidata are verified by an LLM reading noisy HTML — expensive ($4.25/week) and less reliable than a direct API comparison. This plan adds a deterministic Wikidata matcher: extract Q-numbers from source URLs, batch-fetch structured claims via
wbgetentities, and compare property values directly. Phase 1 covers 12 property types (~414 facts, 97%). Numeric properties (revenue, headcount — 11 facts) are deferred to a follow-up. Top open question: whether lazy label resolution is fast enough or if batch pre-fetch for labels is needed.Problem
The weekly source-check spends ~$4.25 and ~425 LLM calls verifying Wikidata-sourced facts by converting HTML to text and asking Claude Haiku to judge. Wikidata has a structured JSON API that returns exact property values. The deterministic matcher for grants (PR #3571) proved the "deterministic first, LLM fallback" architecture works — this extends it to Wikidata facts.
Current State
crux/lib/source-check/item-verifier.tstryDeterministicMatch()for grants (line 227)crux/lib/source-check/fuzzy-match.tsnameMatches(),amountMatches(),dateMatches()urlMatches()for website/wikipedia-urlcrux/commands/wikidata-enrich.tsWIKIDATA_PROPERTY_MAP(10 PIDs, lines 60-154)crux/commands/orgs.tsgetWikidataClaims(),getWikidataLabel(),fetchWithRetry()people/enrich.ts— shared API code extracted into matcherpackages/factbase/data/things/*.yamlcrux/lib/source-check/source-check-recheck.tsProposed Approach
Add
tryWikidataMatch()initem-verifier.tsthat intercepts fact items with Wikidata source URLs. UseswbgetentitiesAPI (batched 50 QIDs/call) with lazy memo cache for label resolution. Falls through to LLM on any failure.Batch path optimization: In
orchestrator.ts, scan items before verification loop → collect unique QIDs → batch pre-fetch entities (3 calls for 138 QIDs). Individual facts then hit the memo cache with zero API calls.Realtime path (
--entity=X): Lazy per-fact fetching with memo cache. First fact for a QID triggers one API call; subsequent facts for the same QID are free. Typically 1-2 API calls per entity.Key Decisions
wbgetentitiesnot SPARQL: Simpler, supports batching (50 QIDs/call = 3 calls for all entities), no SPARQL endpoint queue. SPARQL stays in enrichment commands.wbgetentitiescall, cached for reuse. Worst case for full corpus: ~80 individual label calls vs 6 batched. Acceptable for a weekly job.wikidata-matcher.ts, not extendingdeterministic-matcher.ts: Different data shape (structured claims vs tabular rows). Clean module boundary.revenueandheadcountneed Wikidata quantity parsing ("+500" with unit URIs, currency handling). 11 facts. Deferred to follow-up.try/catchreturning null is sufficient for a weekly batch job. LLM fallback is the circuit breaker.Architecture
crux/lib/source-check/wikidata-matcher.tswbgetentitieswith memo cache, entity-ref label resolution, sitelink URL construction, typed comparison, claim rank filtering,tryWikidataMatch()entry pointcrux/lib/source-check/wikidata-matcher.test.tscrux/lib/source-check/item-verifier.tstryWikidataMatch()call forkind === 'fact'items with Wikidata URLs, before LLM pathcrux/lib/source-check/orchestrator.tswbgetentitiesbefore verification loopcrux/lib/source-check/fuzzy-match.tsurlMatches(): strip trailing /, normalize http↔https, strip wwwcrux/lib/source-check/fuzzy-match.test.tsurlMatches()crux/commands/source-check-recheck.tstryWikidataMatch()before LLM call to prevent verdict flip-floppingProperty-to-PID verification map (12 properties, ~414 facts):
websiteurlMatches()— normalize trailing /, http/https, wwwwikipedia-urlurlMatches()founded-datedateMatches()— hierarchical year/month/daycountrynameMatches()headquartersnameMatches()founded-by-namenameMatches(), multi-valued OReducationnameMatches(), multi-valued ORlegal-structurenameMatches()born-yeardateMatches()ceonameMatches(), prefer rankparent-organizationnameMatches()descriptionrevenueP21397headcountP11284Verdict mapping:
confirmed(confidence 0.95)contradicted(confidence 0.90)unverifiable(confidence 0.80)null(fall through to LLM)partial(confidence 0.70)Evidence format:
extractedValue:"Wikidata P159 (headquarters) = Q62 (San Francisco)"— raw QID + resolved labelreasoning:"[wikidata-api] FactBase value 'San Francisco' matches Wikidata P159 label 'San Francisco' (fuzzy name match). Claim rank: preferred."checkerModel:'wikidata-api'Implementation Phase: Wikidata Matcher — M (1 session)
Goal: All 414 Wikidata-sourced facts (12 properties) verified deterministically via API. Zero LLM cost for these facts.
wikidata-matcher.ts— QID extraction (regex/Q(\d+)/handles trailing commas), property map, memo cache for entities + labels, typed comparison per property, claim rank filtering (preferred > normal, never deprecated), sitelink URL construction, multi-valued OR for founder/educationwikidata-matcher.test.ts— all 12 property types with mock fixtures, edge casesurlMatches()tofuzzy-match.ts+ testsitem-verifier.ts—tryWikidataMatch()call for Wikidata-sourced factsorchestrator.ts— batch QID pre-fetchsource-check-recheck.ts— addtryWikidataMatch()before LLMWIKI_SERVER_ENV=prod pnpm crux fb source-check --entity=80000-hours --verboseWIKI_SERVER_ENV=prod pnpm crux fb source-check --entity=deepmind --verbosecheckerModel: 'wikidata-api'Quality gates:
pnpm testpasses (all existing + new tests)pnpm buildsucceedspnpm crux fb source-check --dry-runshows correct item countExit criteria:
WIKI_SERVER_ENV=prod pnpm crux fb source-check --entity=80000-hours --verboseshows all facts withcheckerModel: 'wikidata-api'.Rollout sequence: Per-entity smoke tests → compare verdict distribution → full corpus run. If >20% contradicted rate (vs ~5% from LLM runs), investigate before proceeding.
Scope Cuts
"+500"with unit URIs, currency comparison). Follow-up issue.orgs.ts/people/enrich.ts/wikidata-enrich.tsworks fine. Consolidate when a second consumer of shared API client emerges. Follow-up issue.wikidata-qidproperty: QIDs extracted from source URLs at verification time. Persistent QID field is a separate enhancement./source-checkspage displays all checker models. Only change: add'wikidata-api'case toformatCheckerModel()(~1 line).Quality & Verification Infrastructure
Tests:
wikidata-matcher.test.ts(~200 lines). Mock Wikidata API at HTTP level. Test each of 12 property types. Edge cases: multi-valued P112 (4 founders), missing P169, deprecated claim rank, trailing comma URLs, sitelink title with parentheses ("Asana (software)"), entity with no English label (fallback to null → LLM), multi-QID entity (facts from different QIDs on same FactBase entity).Integration smoke test:
UI verification: N/A — no UI changes beyond 1-line
formatCheckerModeladdition.Deploy steps: No new env vars, no migrations. Standard PR merge. Wikidata API is public (no auth).
Documentation: N/A — the matcher follows the existing
tryDeterministicMatchpattern. Property map inwikidata-matcher.tsis self-documenting.Monitoring: After first weekly run, compare verdict distribution (confirmed/contradicted/partial counts) against previous LLM-based run. If contradicted count increases by >2x, file an issue.
Risks & Mitigations
nameMatches()handles normalization. Fallback: no English label → return null → LLM.partialif some found,contradictedonly if none found./Q(\d+)/extraction (not URL parsing). Test with exact format.https://en.wikipedia.org/wiki/${title.replace(/ /g, '_')}. Test with parenthesized titles. Fetch withprops=claims|sitelinks.tryWikidataMatch()into recheck path. Same matcher runs on rechecks.checkerModelin dedup key).try/catchreturns null → LLM fallback. No worse than today.Open Questions
Rejected Approaches
wbgetentitiessolves in 3 calls.Red Team Log
Technical critic — 2 blocking, 10 significant:
Scope critic — key cuts adopted:
Integration critic — 1 blocking, 4 significant:
All reactions