Replies: 3 comments
Update: WS2 — PG-first for Website Source RegistryPer discussion, YAML is wrong for this. Auto-update has ~30 static RSS feeds — YAML works. Website sources will be hundreds of entries (80+ orgs × multiple pages, plus people/project/government sites), tightly coupled to PG entity records, and need operational state. PG table from the start. Revised schema: -- Websites we actively track as data sources
CREATE TABLE website_sources (
id VARCHAR(10) PRIMARY KEY,
domain TEXT NOT NULL,
entity_id TEXT REFERENCES entities(stable_id) ON DELETE SET NULL,
entity_display_name TEXT,
reliability TEXT DEFAULT 'medium', -- high | medium | low
refresh_interval_days INTEGER DEFAULT 30,
enabled BOOLEAN DEFAULT true,
notes TEXT,
last_run_at TIMESTAMPTZ,
last_error TEXT,
consecutive_failures INTEGER DEFAULT 0,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE UNIQUE INDEX idx_ws_domain ON website_sources(domain);
CREATE INDEX idx_ws_entity ON website_sources(entity_id);
-- Individual pages within a website that we track
CREATE TABLE website_source_pages (
id VARCHAR(10) PRIMARY KEY,
source_id VARCHAR(10) NOT NULL REFERENCES website_sources(id) ON DELETE CASCADE,
path TEXT NOT NULL, -- e.g., "/about", "/team"
page_role TEXT, -- about | team | research | pricing | careers | docs | other
extract_targets JSONB, -- ["headcount", "headquarters", "leadership", ...]
refresh_interval_days INTEGER, -- overrides source-level default if set
enabled BOOLEAN DEFAULT true,
last_fetched_at TIMESTAMPTZ,
last_snapshot_id VARCHAR(10) REFERENCES page_snapshots(id),
last_content_hash TEXT,
created_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE UNIQUE INDEX idx_wsp_source_path ON website_source_pages(source_id, path);
CREATE INDEX idx_wsp_role ON website_source_pages(page_role);Benefits over YAML:
Also removes the "migrate to PG later" step from the plan — just start there. |
|
Status (April 2026): Schema exists, pipeline not built. Done:
Not done:
|
Implementation Issues FiledDecomposed the remaining work into 3 sequential issues, each scoped to a single agent session:
Dependency chain: #3652 → #3655 → #3660 (each builds on the previous) What already exists: |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Vision
Treat organization websites as first-class data feeds — not just citation links, but active sources of structured facts that continuously flow into TableBase and FactBase. An org's about page is one of the most authoritative, richest sources of facts about that entity. The system should periodically fetch these pages, extract structured data via LLM, and pipe it into entity records automatically.
Current state: Website resources have
type: "web"— a catch-all that lumps org homepages, news articles, tool landing pages, and reference docs together. We store one URL, maybe a cached page fetch incitation_content, and a manually written summary. Facts from websites are embedded as prose in wiki pages ("Anthropic has 1,500 employees") with no structured representation or staleness detection.Target state: Org websites are registered as structured data sources with tracked sub-pages. A periodic pipeline fetches pages, stores dated snapshots, extracts facts via LLM, and upserts them into TableBase/FactBase with full provenance. Every fact traces back to a specific snapshot: "we fetched anthropic.com/about on 2025-11-15, it said X, we extracted Y." Staleness detection flags when source content changes.
Relationship to other work:
data/auto-update/sources.yaml) — similar pattern (fetch → extract → update), but for news. This extends it to structured data extraction from primary sources.Core Concept: Snapshots as Evidence
The central design constraint: every fact extracted from a website must trace back to a dated snapshot of a specific page. This connects two systems:
A snapshot is the unit of evidence. The chain is:
When the page changes (new fetch shows different content_hash), the system knows:
This is different from the current
citation_contenttable, which stores only the latest fetch with no history and no link to extracted facts.The Pipeline
Workstreams
WS1: Content Lifecycle on Resources (~2h)
Add a
content_lifecyclefield to theresourcestable.immutableversionedevergreenephemeralBackfill via heuristics + LLM batch classification. This field guides verification scheduling — only
evergreenandversionedresources need periodic re-checking.WS2: Website Source Registry (~4h)
YAML config defining which websites to track and what to extract from each page.
Start with YAML (version-controlled, consistent with
auto-update/sources.yaml). Migrate to PG if the list grows large or needs API management.WS3: Snapshot Storage (~4h)
Extend the data model so snapshots are preserved with history, not just latest-fetch. Two options:
Option A: Extend
citation_contentwith historyAdd a
citation_content_historytable that preserves previous fetches:citation_contentremains the "current" view. History preserves the snapshot facts were extracted from.Option B: Dedicated
page_snapshotstableRecommendation: Option B. Snapshots are a distinct concept from citation content caching. They're evidence records with extraction metadata. Keeping them separate avoids overloading
citation_contentwith temporal/extraction concerns.WS4: Extraction Pipeline (~12h)
The core capability: fetch → snapshot → extract → upsert.
Components:
source-fetcher.ts(unified in PR refactor: unify URL fetching — source-fetcher as single fetch layer #2895). Fetch page, create snapshot record.snapshot_idas provenance.Extraction prompt design:
Cost: ~$0.01-0.03/page. 80 orgs × 4 pages × monthly = ~$3-10/month.
WS5: Resource Type Split (~3h)
Split
webinto meaningful subtypes:website— an org/project's website (evergreen, multi-page source)article— standalone web page with article-like content (news, essays not on known forums)Add
resource_websitessub-table (like existingresource_papers,resource_forum_posts):Backfill by classifying existing
webresources using URL patterns + domain heuristics.WS6: Staleness Detection & Verification (~4h)
Connect snapshots to the verification loop:
content_hash→ previous snapshot's facts are "stale."sourcepointing to a snapshot. When the snapshot is stale, the fact is flagged.WS7: Auto-Update Integration (~3h)
Website extraction shares infrastructure with auto-update (scheduling, cost tracking, run history, dashboards). Consider making it a new source type:
Reuses auto-update's scheduling, run history (
auto_update_runstable), and dashboard infrastructure (E914, E915).Key Design Decisions
Snapshot storage: extend
citation_contentor new table? Recommend newpage_snapshotstable — snapshots are evidence records, not just a cache.YAML vs PG for source registry? Start with YAML (version-controlled, simple). Migrate to PG if list grows past ~100 sources.
Extraction granularity: Per-source config defines target fields. This is more reliable than "extract everything" and cheaper. Can add a "discover" mode later that identifies what's extractable on a page.
Auto-update vs human review for changes: Numeric fields (headcount, funding) → auto-update. Structural fields (leadership, mission) → flag for review. Configurable per property.
Scale priority: Start with top 20 orgs by page traffic, expand to all ~80 with directory pages.
Integration with auto-update: Share infrastructure (scheduling, dashboards) but keep extraction logic separate — auto-update produces prose updates, website extraction produces structured facts. Different output handlers, same runner framework.
Estimated Effort
WS1, WS2, WS3 can start in parallel. WS4 is the core work. WS5-7 follow.
Success Metrics
All reactions