Skip to content

Repository files navigation

litgraph

Academic paper ingestion & search backed by ArcadeDB: keyword search, semantic (vector) search, and citation graph traversal. Currently ingests from arXiv, with more sources (e.g. PubMed) planned.

  • Storage: ArcadeDB (self-hosted, Apache-2.0) by default. A vector index handles semantic search, a full-text index handles keyword search, and the graph itself models the citation network.
  • Embeddings: SPECTER2 (allenai/specter2_base + allenai/specter2 proximity adapter, 768-dim), run locally via the adapters library — no external embedding API/cost. It is trained on citation networks, ideal for semantic search and relatedness of scientific and biomedical papers in a graph database.
  • Ingestion: historical backload from the Kaggle arXiv metadata snapshot, daily incremental fetch from the arXiv API, citation enrichment from Semantic Scholar.
  • Deferred: a FastAPI query layer and cron-based daily scheduling. For now, everything is a CLI command and a set of plain importable functions in litgraph.search.* that a future API layer can call directly.

Setup

uv sync --extra dev
cp .env.example .env   # fill in SEMANTIC_SCHOLAR_API_KEY
docker compose -f docker-compose.arcadedb.yml up -d
uv run litgraph init-db

Usage

# --- Step 1: Start up container
docker compose -f docker-compose.arcadedb.yml up -d

# --- Step 2: Choose any of the following:

# Backload a subset of the Kaggle arxiv-metadata-oai-snapshot.json(.gz)
# (download separately via `kaggle datasets download -d Cornell-University/arxiv`)
uv run litgraph backload --file /path/to/arxiv-metadata-oai-snapshot.json \
    --categories cs.AI,cs.CV --start-date 2023-01-01 --limit 5000

# Enrich ingested papers with Semantic Scholar citation data
uv run litgraph enrich --limit 500

# Pull new papers submitted since the last run (safe to run daily via cron later)
uv run litgraph fetch-daily --categories cs.CL,cs.LG

# Search
uv run litgraph search keyword "diffusion models"
uv run litgraph search semantic "generative models for images"

# Citation graph
uv run litgraph citations 1706.03762 --direction both --depth 2

Graph schema

Nodes

  • Paper {id, arxiv_id, s2_paper_id, title, abstract, categories, primary_category, published_date, updated_date, doi, journal_ref, comments, embedding, citation_count, reference_count, influential_citation_count, source, is_stub, fetched_at, enriched_at, embedded_at}id is arxiv_id when known, else s2:<s2_paper_id>. Citation targets outside the ingested set are written as lightweight stub nodes (is_stub: true) and get filled in automatically if that paper is later fully ingested.
  • Author {name}, Category {code}

Relationships

  • (:Author)-[:AUTHORED]->(:Paper)
  • (:Paper)-[:IN_CATEGORY]->(:Category)
  • (:Paper)-[:CITES]->(:Paper)

Known limitations

  • Author disambiguation: authors are merged by normalized name string, not a stable ID — two different people with the same name become one node.
  • Semantic Scholar's batch endpoint caps citations/references per paper rather than returning the full list; landmark papers with huge citation counts are undercounted in the graph even though citation_count/reference_count on the node reflect the true totals. Upgrading to the paginated /paper/{id}/citations and /paper/{id}/references endpoints is the natural next step if exhaustive edges are needed.
  • enrich only processes papers that have never been enriched (enriched_at IS NULL); there's no re-enrichment of stale citation counts yet.

Tests

uv run pytest

About

A search engine for research papers backed by Neo4j's citation/knowledge graph DB

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages