Replies: 6 comments
Resource Pipeline v2: End-to-End Agent Workflow & Future ConcernsThis comment expands the discussion beyond the original 9 problems (mostly resolved) into a broader view of what the resource pipeline needs to become. The framing: an agent should be able to submit URLs, get structured results back, and use those results for source-checking — all without manual intervention. The Target Agent WorkflowToday's state: Step 1 exists ( Issue A: Agent Batch Submission & Polling APIProblem: An agent submitting 10 URLs has no way to track them as a group. It would need to call Proposed solution: Extend
The job infrastructure already supports batch creation ( Partial completion: When 7/10 URLs are done, the agent should be able to start using those 7 while waiting for the remaining 3. The status endpoint should return per-URL results as they complete. Issue B: Wayback Machine Fallback in resource-ingestProblem: Current state: The shared source-fetcher already has:
Fix: Make Issue C: Content Type HandlingProblem: Many resources are PDFs, videos, or structured data — not HTML pages.
The Issue D: Multi-Page Resources (Sites vs. Pages)Problem: A single "resource" like Options:
Option 3 is lowest effort and may be sufficient. The data model already captures individual pages as separate resources — the missing piece is browsing them grouped by site. Agent workflow implication: When an agent submits Issue E: Duplicate Resource DetectionProblem: The same paper might exist at arxiv.org, semanticscholar.org, the author's site, and the journal. We could have 4 resource records for one paper. Detection signals:
Impact: Wasted LLM costs (enriching the same paper 4 times), confusing citation counts, fragmented source-check results. Issue F: Resource Staleness & Periodic Re-ingestionProblem: Content changes over time. A policy document gets updated, a blog post gets corrected, an organization restructures their site. We cache content once and never refresh. The Proposed: A scheduled sweep (daily or weekly) that re-ingests resources based on lifecycle:
Issue G: Cost Tracking & Budget GuardsProblem: Each Proposed:
Issue H: Source-Check Cannot Self-Recover from Cache MissesProblem: When source-check hits a URL that's not in Current state: Three commands ( Proposed: When source-check encounters
This makes the pipeline self-healing: source-check never fetches directly, but it signals the resource pipeline to fill the gap. Issue I: No Query API for Verification StatusProblem: Evidence and verdicts are stored in the DB, but there are no GET endpoints to query them. Agents cannot discover:
Proposed endpoints: Issue J: Prioritized IngestionProblem: Not all resources are equally important. A paper cited by 40 wiki pages should be ingested and enriched before an obscure blog post cited once. The bulk enqueue command processes resources in arbitrary order. Proposed: Priority scoring for resource-ingest jobs:
Issue K: Rate Limiting & Domain PolitenessProblem: Ingesting 5,000 URLs hits thousands of domains. Some rate-limit, some might ban our IP. The current pipeline has per-request timeouts but no per-domain throttling. Proposed: Per-domain concurrency limits in the worker:
Issue L: Resource Discovery & SuggestionsProblem: Wiki pages cite URLs that aren't in the resources table. There's citation indexing that creates resources from footnotes, but what about inline links? Or URLs in FactBase facts? Or resources an LLM could suggest based on entity context? Proposed: When enriching an entity or page, the LLM could suggest additional resources:
These suggestions get added to a "suggested resources" queue for agent review. Issue M: Three Separate Verification PipelinesProblem: Three independent commands handle verification with different fetch strategies and output formats:
Impact: Agent confusion on which to use. Bugs need fixing in 3 places. Inconsistent verdicts. Proposed: Consolidate into one pipeline with scope filters: Summary: Priority OrderPhase 1 — Make the pipeline self-healing (highest impact):
Phase 2 — Agent batch workflow:
Phase 3 — Content quality & coverage:
Phase 4 — Scale & operations:
|
|
Status (April 2026): Audit findings documented. Resource enrichment pipeline ( |
Issues Filed from This AuditAfter reviewing the current codebase, most of the core pipeline problems (Problems 2-5, Issues A/B/H from the comment) have been addressed in the pipeline v2 work (March 2026). Specifically:
Three remaining high-impact gaps have been filed as issues:
Lower-priority items from the comment (Issues C/D/E/F/G/J/K/L/M) are not yet addressed but are less urgent — they represent scale/polish improvements rather than architectural gaps. |
Status Update — April 6, 2026This discussion is effectively resolved. All 7 original problems plus the 3 filed follow-up issues have been addressed across a series of PRs. Entity-Resources Cleanup (PRs #3936, #3943, #3951, #3965)Replaced the fragmented source/resource system with a proper
Table seeded: 4,182 rows in production (publisher + wiki citation passes). Architecture Now
Follow-up Issues (all resolved)
Remaining Minor Items
Recommend closing this discussion. |
Rename source-check → sourcing — Implementation PlanTL;DRThe "source-check" system has three coexisting naming conventions (source-check, verification, SourceCheck), a name that implies post-hoc verification when it also handles write-time sourcing, and redundant ProblemThe source-check system is the wiki's primary infrastructure for record-level provenance (linking structured data to their backing source URLs with verification verdicts). Three problems:
Current State
Proposed ApproachCode-first, tables-last (learned from red-team: Order:
Key Decisions
Implementation PhasesPhase 0: Remove sourceResourceId + finish March rename — SGoal: Clean foundation. One consistent name ("source-check") before introducing "sourcing".
Files: ~30. Risk: Low (mechanical, no DB table changes). Phase 1: Drizzle symbol rename (code points at old tables) — SGoal: All TypeScript uses new symbol names. PG tables unchanged.
Files: ~20. Risk: Low (TypeScript catches all broken refs). Phase 2: Server route + shared helpers rename — SGoal:
Files: ~15. Risk: Low (alias keeps old path working). Phase 3a: Frontend component rename — MGoal: Components say "Sourcing" instead of "SourceCheck".
Files: ~28. Risk: Medium (many consumers, but mechanical). Phase 3b: Frontend page + dashboard rename — SGoal: Users see
Files: ~20. Risk: Low (redirects catch broken links). Phase 3c: CLI + workflows rename — MGoal: CLI says
Files: ~45. Risk: Medium (largest single phase, conflict risk). Phase 4: DB table rename + cleanup — SGoal: PG tables match code. Clean break.
Files: ~15. Risk: Low (all code already uses new names, this is just the PG rename). Scope Cuts
Quality & Verification Infrastructure
Risks & Mitigations
Open Questions
Rejected Approaches
Red Team LogRound 1 findings (3 agents): Technical critic:
Scope critic:
UX critic: N/A (internal infrastructure, no user-facing UX changes beyond nav label) |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Overview
The resource/sources system has accumulated significant duplication, dead code, and missing automation. This discussion catalogs all the problems and proposes a unified cleanup plan organized around one architectural principle.
Core Principle: Resources First, Source Checks Second
The entire resource system should be understood as two layers with a strict dependency:
Layer 1 — Resource Pipeline (build the library): Fetch URLs, cache content, generate summaries. This runs proactively when resources are created or periodically refreshed. After it completes, every resource should have cached full text (or a documented reason why not), a summary, and enrichment metadata.
Layer 2 — Source Check (verify claims against the library): Compare wiki claims against resources that already exist in our database. Source-check should never fetch URLs itself — it reads from
citation_contentand nothing else. If a resource hasn't been cached yet, that's a Layer 1 problem, not something source-check should work around.Today these layers are tangled: source-check has its own fetcher, resource-verify discards content, enrichment is manual, and the content cache is empty. Every problem below traces back to this architectural confusion.
Problem 1: Three Overlapping Public Resource Pages
We have three separate URLs that all try to show the same thing — a browsable list of resources:
/sources/resources/sources, plus type-breakdown stats/wiki/E874<ResourcesIndex>— a stub component never ported from Astro/sourcesis a superset of/resources. The Resources tab on/sourcesrenders the exact sameResourcesTablecomponent with the sameResourceRowtype./resourcesjust adds some extra stat cards (type counts, "Enriched" count) that/sourcesdoesn't have./wiki/E874is a dead page —ResourcesIndexis listed in the stub components array inmdx-components.tsxand renders a placeholder warning. It has never worked in the Next.js version.Additionally, there are four wiki MDX pages in the sources area:
content/docs/sources/sources-overview.mdx<SourcesOverviewContent />/resourcesand/publicationscontent/docs/sources/factbase-resources.mdx<FBResourcesContent />content/docs/sources/factbase-publications.mdx<FBPublicationsContent />content/docs/tools/resources.mdx<ResourcesIndex />(stub)The "Resources" concept is spread across 7 different URLs (
/sources,/resources,/resources/[id],/publications,/wiki/E1049,/wiki/E1043,/wiki/E874) with overlapping content and inconsistent naming.Solution
content/docs/tools/resources.mdx) — broken stub. Redirect/wiki/E874to/sources./resourcesinto/sources— merge the extra stat cards (type breakdown, enrichment pipeline progress) into the/sourcesResources tab. Then redirect/resources→/sources./resources/[id]as the detail page for individual resources./sourcesApp Router pages. Consider redirecting to the App Router versions.ResourcesIndexfrom the stub list inmdx-components.tsxonce E874 is deleted./sources, but/resources,/wiki/E874,/wiki/E1043, and/wiki/E1049are all independently reachable. After consolidation, ensure nav only links to canonical URLs.Problem 2: The Content Cache Is Empty — resource-verify Discards Fetched Content
This is the root cause of most pipeline problems. The
resource-verifyjob handler fetches URLs to check liveness but throws away the HTML after computing a hash. Thecitation_contenttable — designed to cache full text — is essentially empty.Production data (from a recent investigation):
fetched_atIS NOT NULLcitation_contentEven the 563 resources marked as "fetched" have zero cached content.
fetched_atis set byresource-verify, but the content is discarded.Impact on Layer 2 (source-check):
citation_contentfirst → empty → falls back to cold HTTP fetchThis violates the core principle: source-check should read from a populated library, not do its own fetching.
Solution
Make
resource-verifycache content incitation_contentwhen it successfully fetches a URL. This is the single highest-leverage change — it populates the library that both enrichment and source-check depend on.The handler should also persist
fetch_status,last_fetched_at, andcontent_hashback to the resource record (currently these are returned inJobHandlerResult.databut never written to the DB).Related issues: #3449, #3452, #3457, #3472. PR: #3455.
Problem 3: Enrichment Pipeline Is Entirely Manual
The enrichment pipeline that generates summaries is a 3-stage manual CLI process:
Each stage must be run manually, in order, with separate polling for batch completion. When a new resource is added or verified, nothing triggers enrichment.
Current state: 5,326 resources, only 1,861 have summaries (35%).
Why batch-only is a problem
The Anthropic Batch API gives a 50% discount but:
Solution:
resource-enrichjob handler + automatic chainingFollowing the "resources first" principle, enrichment should be part of the resource pipeline — triggered automatically, not manually. Create a
resource-enrichjob type that:{ resourceId, url }as paramscitation_content(populated by the fixedresource-verify— see Problem 2)enrichment_statusto'enriched'Automatic chaining: When
resource-verifycompletes successfully withstatus === 'reachable', it enqueues aresource-enrichjob. The full automated pipeline becomes:Cost: ~1-2¢ per resource (standard API) vs ~0.5¢ (batch). The 2x markup is worth it for full automation.
The existing batch pipeline can remain as a bulk backfill tool for processing thousands at once, but the per-resource job handler is the primary path for new resources.
Merging classify into enrich
The classify stage (
classify.ts) generatesresource_subtype,resource_purpose, andcontext_noteusing Haiku. But the enrich stage also generatesresource_purposeandcontext_note, overriding classification results. The only unique classify output isresource_subtypeandsub_table.For the per-resource
resource-enrichhandler: combine classification and enrichment into a single Sonnet call. One LLM call per resource, not two. The batch pipeline can keep separate stages for cost optimization if desired.Problem 4: Three Separate Source-Fetching Implementations
Three independent systems fetch web content with no shared code:
crux/lib/search/source-fetcher.tsfetchSource()fetch-all.ts)crux/lib/source-check/source-fetcher.tsfetchSourceContent()crux/lib/job-handlers/resource-verify.tsfetch()All three detect paywalls and soft 404s, check HTTP status codes, and handle timeouts — but with different implementations. The source fetcher has domain-aware strategies (Firecrawl, Wayback, forum APIs); the other two don't.
Under the "resources first" principle, only the resource pipeline should fetch URLs. Source-check should read from
citation_contentonly.Solution
resource-verifyuse the shared source-fetcher (fetchSource()) instead of rolling its ownfetch()— this gives it domain-aware strategies, Firecrawl, Wayback Machine supportresource-verifycaches content incitation_contentas a side effect (see Problem 2) — this eliminates the need for the separatefetch-allCLI stagecitation_contentreader with no HTTP fallback. If the cache misses, report "resource not cached" — don't try to fetchfetch-allCLI command becomes a bulk backfill tool (re-cache everything), not the primary caching pathEnd state: one fetching implementation (source-fetcher), used by one system (resource-verify), writing to one cache (
citation_content). Source-check only reads.Problem 5: Enrichment Status Not Visible in Stats or Dashboards
The
/api/resources/statsendpoint returns:{ "totalResources": 5326, "withMetadata": 1861, "fetched": 563, "orphanedCount": "???", "byType": { "web": 3981, "paper": 465, ... } }Missing: No breakdown by
enrichment_status(pending,fetched,classified,enriched,reviewed). No way to see pipeline progress — how many resources are stuck at each stage. No visibility into how full the content cache is.Solution
Add
byEnrichmentStatusandcontentCacheStatsto the stats endpoint:{ "byEnrichmentStatus": { "null": 2800, "pending": 200, "fetched": 400, "classified": 65, "enriched": 1861, "reviewed": 0 }, "contentCacheStats": { "withFullText": 0, "withMetadataOnly": 14, "uncached": 5312 } }Surface this on the consolidated
/sourcespage as a pipeline progress indicator. This makes it immediately obvious whether the resource pipeline is healthy and how much of the library is populated for source-check.Problem 6:
fetch_statusvsenrichment_statusConfusionThe resource schema has two status-like fields:
fetch_status— the internal dashboard assumes values"full"|"metadata-only"|"unfetched", but nothing in the codebase consistently writes these valuesenrichment_status— the enrichment pipeline's state machine (pending→fetched→classified→enriched→reviewed)These overlap confusingly.
resource-verifydoesn't write to either (Problem 2).fetch-allwrites toenrichment_statusbut notfetch_status. TheUpdateResourceFetchStatusSchemaexists inapi-types.tsbut its usage is unclear.Solution
Audit whether
fetch_statusserves a distinct purpose fromenrichment_status. If it's about content availability (full text vs metadata-only vs nothing), consider making it a computed field derived fromcitation_contentstate. If it's dead, deprecate it.With the resource pipeline properly caching content (Problem 2), the content availability question can be answered by querying
citation_contentdirectly rather than maintaining a separate status field.Problem 7: Dead / Tactical Code
crux/resource-enrichment/fix-enrichment-status.tsenrichment_status— tactical fix for a batch size bugResourcesIndexstub inmdx-components.tsxcrux/resource-enrichment/fetch-all.tsresource-verify(Problem 2)crux/lib/source-check/source-fetcher.ts/source/[id]/page.tsx/sources/publications/[id]/page.tsxCurrent State of the Pipeline (What Happens Today)
Target State
Action Items (Priority Order)
These are ordered by dependency — each step enables the next:
Phase 1: Fix the resource pipeline (enables everything else)
Make
resource-verifycache content + persist results — the single highest-leverage change. Populatescitation_content, writesfetch_status/content_hashto resource records. (Problems 2, 4)Create
resource-enrichjob handler — combined classify+enrich in one Sonnet call, triggered automatically afterresource-verify. (Problems 3, 6)Add enrichment/cache stats to API —
byEnrichmentStatusandcontentCacheStatsin/api/resources/stats. (Problem 5)Phase 2: Simplify source-check (now that the library is populated)
citation_contentonly. If cache miss, report it instead of fetching. (Problem 4)Phase 3: Clean up the UI
Consolidate resource pages — merge
/resourcesinto/sources, delete E874, add pipeline progress indicator, fix nav. (Problems 1, 7, 9)Audit
fetch_statusfield — document or deprecate. (Problem 6)Delete dead code —
fix-enrichment-status.ts,ResourcesIndexstub. (Problem 7)Questions for Discussion
/sourcesbe the canonical URL, or rename to something clearer like/library?All reactions