Replies: 2 comments
Linked Issues
Both issues are part of this epic. Phase 2 sub-issues will be created as implementation begins. |
0 replies
Phase 0-1 Complete (2026-04-02)Phase 0: Free Enrichment
Phase 1: LLM Batch Enrichment
Impact
Both #3576 and #3577 closed. Ready for Phase 2 (bulk discovery). |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Goal
Massively expand the wiki's resource/publication coverage for EA-related organizations. Currently the wiki has 5,346 resources in PG but metadata coverage is poor (9% have authors, 9% have dates, 5% have abstracts), and there's no systematic publication discovery for the 592 tracked organizations.
Target: Go from ~5K sparse resources to ~30-60K well-enriched resources covering the major outputs of EA/AI safety organizations, forums, and academic researchers.
Current State (as of 2026-04-02)
Resources in PG
Literature.yaml
Organization Coverage
Phase 0: Enrich Existing Resources ($0 cost)
Issue: #3576
Run the existing free enrichment pipeline on all 5,346 current resources:
Expected outcome: >50% author coverage, >50% date coverage on existing resources.
Phase 1: LLM Batch Enrichment (~$115)
Issue: #3577
After Phase 0:
Expected outcome: >80% with summaries, >50% linked to org entities.
Phase 2: Bulk Resource Discovery (NEW)
This is the big push. Multiple independent discovery channels, each targeting different source types.
2A. LessWrong / EA Forum / Alignment Forum — Global Post Import
The single biggest win. ~19,000-37,000 high-quality posts available via free GraphQL API.
API Capabilities (verified 2026-04-02)
Both LW and EAF expose a public GraphQL API at:
https://www.lesswrong.com/graphqlhttps://forum.effectivealtruism.org/graphqlhttps://www.alignmentforum.org/graphqlKey query: Global post listing with filtering and sorting:
Full markdown content also available (no scraping needed):
API Constraints
totalCountreturns null on global queries — must paginate until emptykarmaThresholdworks as a server-side filter (efficient, reduces result set before pagination)Measured Post Counts (karma >= 10)
Combined total (karma >= 10): ~36,972 posts
Cost Estimates by Karma Threshold
Recommendation: Start with karma >= 30 (~19K posts, ~$44 LLM cost). This captures substantive posts while filtering spam and low-effort content. Can always lower the threshold later.
Full Cost Breakdown (karma >= 30)
Implementation Plan
What already exists:
crux/lib/forum-api.ts— GraphQL client for LW/EAF (needs AF added toALL_FORUMS)crux/scripts/fetch-forum-posts.ts— Single-author post discovery + resource creationcrux/resource-enrichment/enrich-forums.ts— Enriches existing forum resources with metadataPOST /api/resources/batch— Batch upsert (200/batch)POST /api/resources/suggest— Creates stubs + auto-chains enrichment jobsresource_forum_postssub-table with karma, commentCount, tags, curated, authorUsername fieldsWhat needs building:
ALL_FORUMSinforum-api.tsgetGlobalPosts(forum, { karmaThreshold, after, before, sortedBy })— paginated global querycrux resources discover-forumswith options:--karma=30— minimum karma threshold--forums=lw,eaf,af— which forums--after=2020-01-01— start date--limit=5000— max resources to create--dry-run— preview without creating--apply— create resources_idprefix or title+author match)enrich-forumsfor metadata + optionally chain LLM enrichmentEstimated engineering: ~4-6 hours
Cross-Post Handling
Posts are often cross-posted between LessWrong and Alignment Forum. The
resource_forum_poststable hascrossPostedFromandcanonicalForumfields (currently unused). The discovery command should:_id(AF posts often share LW_id)2B. Forum Author Sweep — Known Researchers
Complements 2A by ensuring we have complete publication lists for known EA researchers.
What exists
crux/lib/forum-api.ts:fetchAuthorPosts(slug)— fetches all posts by author across forumscrux/scripts/fetch-forum-posts.ts— creates resources from author's postsdata/entities/people.yaml— person entities with some having forum slugsImplementation
lesswrongSlugor similar fieldsfetchAuthorPosts()for eachauthorEntityIdsAdvantage over 2A: Directly links posts to person entities (2A creates resources but doesn't know which wiki person wrote them).
Estimated engineering: ~2-3 hours (mostly building the person→slug mapping)
2C. Forum Tag-Based Discovery
LW/EAF support querying posts by tag. Many orgs and topics have dedicated tags.
GraphQL Query (not yet implemented, needs verification)
Note: The exact
termssyntax for tag filtering needs verification against the LW/EAF API. ThefilterSettingsapproach is used by the website but may not be exposed via the public GraphQL API. An alternative is to fetch tag pages directly.Use Cases
Estimated engineering: ~3-4 hours (API exploration + command implementation)
2D. Semantic Scholar — Author-Chain Organization Discovery ($0)
For academic orgs, discover papers via known researchers' Semantic Scholar profiles.
What exists
crux/lib/search/providers/semantic-scholar.ts— S2 API providersearchAuthors(query)→ returns authorId, affiliationsgetAuthorPapers(authorId)→ returns all papers by authorcrux/resource-enrichment/enrich-papers.ts— enriches paper resources with S2 metadataThe Pattern
Constraints
Estimated Scale
Cost: $0 | Time: ~10 minutes | Engineering: ~4 hours
2E. OpenAlex API — Institution-Level Paper Discovery ($0)
OpenAlex is the most powerful option for "find all papers from org X" — it supports native institution-level search.
API
Advantages over Semantic Scholar
Estimated Scale
Could discover 5,000-15,000 papers across all tracked organizations — many of which won't be in the wiki yet.
Cost: $0 | Engineering: ~6-8 hours (new provider module needed)
2F. Org Research Page Scraping ($0 or Firecrawl credits)
Many orgs have
/researchor/publicationspages listing their work.Known Research Pages
What exists
crux/lib/search/source-fetcher.ts— fetches URLs, handles Firecrawl, forum APIs, paywallscrux/lib/search/fetch-strategies.ts— domain-aware routingImplementation
publicationsUrlfield to org entities (or hardcode top 20)Cost: $0 if using basic fetch, or Firecrawl credits for JS-heavy sites
Engineering: ~6 hours
2G. Auto-Update RSS Expansion
The auto-update system already monitors RSS feeds for some orgs. Expanding this catches ongoing publications.
Currently Configured
Gap
Opportunity
Existing Tooling Inventory
The codebase already has most of the infrastructure for this push:
crux/lib/forum-api.tscrux/resource-enrichment/enrich-forums.tscrux/lib/search/providers/semantic-scholar.tscrux/resource-enrichment/enrich-crossref.tsPOST /api/resources/batchPOST /api/resources/suggestcrux/resource-enrichment/classify.tscrux/resource-enrichment/enrich.tscrux/resource-enrichment/cross-reference.tscrux/lib/search/fetch-strategies.tscrux/resource-lookup.tscrux/scripts/fetch-forum-posts.tscrux/lib/search/source-fetcher.tscrux/commands/people/link-resources.tscrux/auto-update/feed-fetcher.tsforum-api.tsdiscover-forumscommandLinked Issues
Priority Order
Total Estimated Impact
All reactions