An AI-assisted toolkit for academic paper analysis, systematic reviews, corpus-scale evidence mapping, and review knowledge bases.
PaperScope is one of three linked but distinct tools:
| Tool | Does | Input → Output |
|---|---|---|
| PaperScope (this repo) | Analyzes the literature — bibliography / DOI / retraction QA, forensic metascience, systematic reviews, embeddings. | papers → checked analysis |
| LocalEvidence | Answers a clinical question from a library you own, grounded and cited. | a question + your library → a cited evidence pack |
| EvidenceViewer | Presents any source-backed artifact through one contract + viewer, every claim traceable to its source. | an EvidenceArtifact → a source-linked reading UI |
The pipeline: PaperScope analyzes → LocalEvidence answers (using PaperScope's fact-checking) → EvidenceViewer presents either one's output.
PaperScope is a Python toolkit for working with academic papers at both manuscript and corpus scale. It supports pre-submission checks on your own work, critical reads of someone else's manuscript, and AI-assisted scoping reviews where a review corpus becomes a queryable evidence base rather than a spreadsheet dump.
The core premise is that paper-level evaluation and corpus-level evaluation are inseparable. A paper is only meaningful relative to the literature it claims to extend, cite, contradict, ignore, or compress. PaperScope therefore treats "evaluate this paper" as a local view into "evaluate this corpus": citation checks, novelty, method resolution, overclaiming, forensic flags, and review synthesis all depend on knowing what the surrounding corpus looks like.
- Semantic analysis — embeds a manuscript and its literature into a shared vector space to catch citation misalignment, unsupported claims, abstract gaps, and missing related work
- Forensic statistics — 22 data-integrity checks (GRIM, GRIMMER, SPRITE, correlation bounds, p-value recalculation, Carlisle test, and more) based on Heathers (2025) An Introduction to Forensic Metascience
- Critical read — author profiling, method-resolution mismatch detection, overclaiming analysis
- Bibliography pipeline — citation extraction, DOI resolution, retraction detection, literature discovery
- Systematic literature reviews — JBI / PRISMA-ScR rails (harvest → screen → extract → validate → synthesise) for scoping reviews where your AI assistant is the screening/extraction engine. There is no bundled classifier:
screenandextractare SDK-agnostic seams (interface + abstaining stub) designed to be driven by the assistant you already use — Claude Code, Codex — through the CLI and JSONL contracts; the pipeline supplies the rails, the append-only audit trail, the human-adjudication queue, and a static-HTML review site. Reviews are protocol-as-data: one YAML defines PCC, search query blocks, screening rubric, charting schema, and aggregation rules. Thevalidatestep turns AI screening/extraction decisions into a human work queue (the model self-flags its low-confidence calls; the human adjudicates only those; flips reconcile back append-only) — seedocs/validate.md. See alsopaperscope/systematic_review/anddocs/systematic-review.md. - Review knowledge bases — the
systematic_review knowledge-baseandrater-comparesubcommands ship the first pieces: a corpus directory exports paper cards, cluster grouping, and quality flags as a self-contained bundle, and two rater passes yield field-level agreement plus Cohen's kappa for adjudication. Richer per-paper metadata, private source-object links, and a searchable collaborator portal remain on the roadmap. Seedocs/corpus-knowledge-base.md.
PaperScope is built for large review corpora: thousands of records, working evidence bases of well over a thousand papers, and AI-assisted charting. The knowledge-base layer's first pieces have shipped as systematic_review subcommands — knowledge-base (paper cards, cluster grouping, quality flags, exported as a self-contained bundle) and rater-compare (field-level agreement between two rater passes, with Cohen's kappa). Richer per-paper metadata, private source-object links, and a searchable collaborator portal remain on the roadmap. Discipline-specific rubrics, claims, and synthesis outputs stay in the caller project; the generic machinery is pulled back into PaperScope.
For large reviews the useful product is not just "screen and aggregate" but a corpus knowledge base that lets a reader ask questions such as:
- What is this cluster of papers about?
- Which papers support this claim?
- Which studies are externally validated?
- Which papers are likely off-scope, mismatched, or methodologically weak?
- Where are the source PDFs or text extracts, and what can be shared publicly?
git clone https://github.com/todd866/paperscope.git
cd paperscope
pip install -r requirements.txt
export PAPERSCOPE_EMAIL="you@university.edu" # for OpenAlex/CrossRef polite poolHeavier optional dependencies ship commented out in requirements.txt — e.g. playwright (only needed for systematic_review browser-harvest) and pyarrow (methodological-audit clustering). Uncomment what you use.
For better embeddings (optional but recommended):
pip install sentence-transformersWithout sentence-transformers, embedding-based tools fall back to TF-IDF (which requires scikit-learn, included in requirements).
PaperScope works as a CLI that any AI assistant can call. Two tested workflows:
Claude Code: Clone into your project or a known path. Claude reads the CLAUDE.md file and uses the CLI when working on papers. You can also use --plugin-dir for skill auto-invocation:
claude --plugin-dir /path/to/paperscopeCodex: Copy AGENTS.md into your paper project directory. Codex reads it and calls the CLI when relevant:
cp /path/to/paperscope/AGENTS.md /path/to/your/paper/# Full analysis: citation alignment, novelty, strength heatmap
python3 -m paperscope analyze paper.tex --literature text/
# Abstract coverage check
python3 -m paperscope abstract-check paper.tex
# Journal semantic fit ranking
python3 -m paperscope journal-fit paper.tex -j "BioSystems" "PLOS ONE"
# Semantic diff between revisions
python3 -m paperscope revision-diff old.tex new.tex
# Find missing related work (needs PAPERSCOPE_EMAIL)
python3 -m paperscope related paper.tex
# Cross-paper dependency / argument graph over a research program
python3 -m paperscope argument-graph /path/to/research/program/ -o graph.png
# Independent (non-self) citation uptake of a published DOI (needs PAPERSCOPE_EMAIL)
python3 -m paperscope citation-uptake 10.1016/j.biosystems.2025.105608# Full critical read of an external paper
python3 -m paperscope critical-read paper.pdf
# With explicit author names (skips auto-extraction)
python3 -m paperscope critical-read paper.pdf --authors "Alice Smith" "Bob Jones"
# Offline mode (skip OpenAlex author lookup)
python3 -m paperscope critical-read paper.pdf --skip-author-lookupRuns four analyses: author/COI profiling, method-resolution mismatch detection, missing complementary methods, and overclaiming detection.
# Build an annotated reading copy from a notes spec (JSON or YAML)
python3 -m paperscope annotate paper.pdf notes.json -o annotated.pdfTurns a PDF + a list of notes — each pinning an anchor phrase on a page to a colour-coded header + body (TEACH / DEF / STRENGTH / CRIT) — into a reading copy with highlighted, numbered passages, interleaved "annotator's notes" commentary pages, a colour-key front page, and an optional one-screen summary + figure appendix. Substrate-free: all paper-specific content lives in the spec, so the same tool builds a teaching copy, a referee's markup, or a collaborator's. Anchors that don't bind are reported (the note is still emitted, badge-only). Spec format and a programmatic API (build_annotated_pdf) are documented in paperscope/analysis/annotate.py; see examples/annotate/.
# Table mode: transcribe a paper's Table 1 into a JSON spec, get verdicts
python3 -m paperscope forensic table1.json
# Text mode: statcheck-style p recalculation straight from the paper
python3 -m paperscope forensic paper.pdf # also accepts .txt / .md
# Text mode + an annotated reading copy with FAIL/FLAG findings highlighted
python3 -m paperscope forensic paper.pdf --annotate annotated.pdfTable mode runs GRIM, GRIMMER, SD-range, variance-ratio, and Carlisle checks over transcribed summary statistics (schema documented in paperscope/analysis/forensic_report.py; the data entry is manual, the checks are automated). Text mode extracts reported t/F/chi2/r/z statistics from the paper text and recomputes each p-value — the approach pioneered by statcheck (Nuijten et al. 2016) — treating every printed number as a rounding interval so honest rounding never produces a false accusation. Every verdict is one of PASS / FLAG / FAIL / UNDETERMINED: FAIL means arithmetically impossible as printed, FLAG means suspicious but not proven, and a parsing problem yields UNDETERMINED, never FAIL. Reports are written as JSON alongside the console output; worked demo in examples/forensic/.
# Or import individual checks in Python
from paperscope.analysis.forensic_stats import grim, grim_percentage, correlation_bound
print(grim(mean="18.72", n=22)) # GRIM test (fails at 2dp)
print(grim_percentage(percentage=53.2, n=25, dp=1)) # GRIM applied to percentages
print(correlation_bound(0.10, 0.30, 0.05)) # impossible r (|r| > 1)# Run the calibration battery: sensitivity + specificity of the checks
python3 -m paperscope forensic-calibrate
# Add your own case directories (repeatable); write the full report JSON
python3 -m paperscope forensic-calibrate --cases my_cases/ -o calib.json
# Fail the process on any mismatch (for CI); otherwise it always exits 0
python3 -m paperscope forensic-calibrate --strictThe calibration harness measures both sides of the cardinal rule of a forensic tool: sensitivity (does it still catch real, arithmetically-verifiable errors?) and specificity (does it leave valid data alone — never a false accusation, the worst failure?). Each case is a known-answer scenario — a table spec and/or a chunk of prose whose errors are planted by construction — with an expected block asserting which checks must fire (must_detect) and which must stay silent (must_pass). Running the whole battery on every change to the checks or the cases is the regression gate: if a change stops catching a planted error or starts flagging valid data, a case flips to MISMATCH. The gate is wired into the test suite (tests/test_calibration.py), and the CLI is a report (always exit 0) unless --strict is passed.
A case file is JSON at calibration/cases/<slug>.json:
{
"meta": {"label": "...", "source": "synthetic",
"ground_truth": "why the expected findings are known",
"notes": "..."},
"table": { "<run_table_checks schema — see forensic_report.py>": "..." },
"text": "<prose with reported t/F/chi2/r/z statistics>",
"expected": {
"must_detect": [{"check": "grim", "target_contains": "treatment", "min_verdict": "FAIL"}],
"must_pass": [{"check": "grim", "target_contains": "control"}]
}
}Either table or text may be null. Verdict ordering for min_verdict is FAIL > FLAG > (PASS, UNDETERMINED): a must_detect of min_verdict FLAG is satisfied by FAIL or FLAG; a must_pass is satisfied only by a matching finding that is exactly PASS.
Public / private split (same pattern as the systematic-review corpus). The public repo ships synthetic, neutral cases only — surveys, generic measurements, abstract "effect of X on Y" studies whose ground truth is arithmetic, never medical. Point the PAPERSCOPE_CALIBRATION_DIR env var (colon-separated for several) at a private directory to add domain-specific cases; those run alongside the built-in battery without ever entering the public corpus:
PAPERSCOPE_CALIBRATION_DIR=/path/to/private/cases python3 -m paperscope forensic-calibrate# Extract citations from LaTeX
python3 -m paperscope extract /path/to/paper/
# Resolve missing DOIs via CrossRef
python3 -m paperscope resolve bibliography.json
# Verify DOIs and detect retractions
python3 -m paperscope verify bibliography.json
# Discover new papers matching your research profile
python3 -m paperscope harvest --config config.yaml
# Download open-access PDFs and extract text
python3 -m paperscope ingest /path/to/literature/
# Harvest depth-2 references from CrossRef (references-of-references)
python3 -m paperscope depth2 /path/to/literature/
# Generate an HTML page for manually sourcing the missing (paywalled) PDFs
python3 -m paperscope sourcing-page /path/to/literature/ paper.tex
# Pre-submission citation check
python3 -m paperscope pre-submit paper.tex --bib bibliography.jsonpython3 -m paperscope paper-site ./site \
--title "Bayesian Descriptions Are Not Mechanisms"Scaffolds a paper-library-backed Next.js paper reader with the shared MD3 visual
contract used across the author's paper sites: native web manuscript first,
downloadable PDFs second, inline citation/detail controls, side-panel reference
context, and paper-library source status. Reference records can carry
role/what/why/caution/contexts fields so the sidebar teaches the
reader what each source is doing instead of dumping raw citation snippets. When
those fields are absent, the scaffold falls back to conservative academic and
clinical source-type explanations. Sidebar prose resolves citation markers, cite
keys, and author-year labels back into the same reference panel; optional
native_href and source_href fields add panel actions without making citation
clicks launch a new tab directly.
LocalEvidence is intended to call the same generator in medical mode rather than
maintain a separate paper-reader fork.
ingest writes into a transient per-project literature/ folder and re-fetches
every project. If you use paperscope across many papers and reviews, stand up a
permanent, machine-wide paper library instead: one deduped catalog (by
DOI/MD5/PMID) with standing semantic search and a snapshot/restore safety net,
sitting on top of paperscope's acquisition and embeddings. A paper enters once and
is never re-fetched. See docs/permanent-library.md
for the pattern and examples/permanent-library/ for
a copy-and-adapt reference skeleton.
cp -r examples/permanent-library ~/paper-library && cd ~/paper-library
export PAPERSCOPE_HOME=/path/to/paperscope
python3 library.py pull 10.1016/j.biosystems.2025.105608 --title "..."
python3 library.py search "active inference free energy" -k 10# Show the composed Boolean query / per-block counts (sanity-check the strategy)
python3 -m paperscope.systematic_review search myreview.yaml --show-query
python3 -m paperscope.systematic_review search myreview.yaml --block-counts
# Harvest MEDLINE into records.jsonl
python3 -m paperscope.systematic_review search myreview.yaml
# Aggregate charted JSONL into synthesis tables
python3 -m paperscope.systematic_review aggregate myreview.yaml
# PRISMA-ScR flow from records + screening JSONL
python3 -m paperscope.systematic_review prisma --config myreview.yaml
# Acquire PDFs for the included set (OA via Unpaywall + EZProxy queue for the tail)
python3 -m paperscope.systematic_review acquire myreview.yaml
# Static HTML review site (Covidence-style record pages, no JS)
python3 -m paperscope.systematic_review build-site --config myreview.yaml --out ./review-site
# Corpus knowledge-base bundle (paper-cards.jsonl + clusters.json + manifest.json)
python3 -m paperscope.systematic_review knowledge-base --config myreview.yaml --out ./kb
# Field-level disagreement between two raters' screening/extraction JSONL
python3 -m paperscope.systematic_review rater-compare \
--rater-a rater_a.jsonl --rater-b rater_b.jsonl --kappa-field decision --table
# Optional institutional-access browser harvest for the paywalled tail
python3 -m paperscope.systematic_review browser-harvest \
--config myreview.yaml \
--user-data-dir "$HOME/Library/Application Support/Google/Chrome" \
--profile-directory Default \
--group-by-publisher \
--inter-paper-delay 5Module README: paperscope/systematic_review/README.md. Design + roadmap: docs/systematic-review.md.
Corpus knowledge-base roadmap: docs/corpus-knowledge-base.md.
22 functions based on techniques from Heathers (2025) An Introduction to Forensic Metascience (DOI: 10.5281/zenodo.14871843).
| Check | Detects |
|---|---|
grim() |
Impossible means for integer-valued instruments (Likert scales, bounded questionnaires, counts) |
grim_column() |
GRIM across a column of means, inferring decimal places column-wide |
grim_row() |
Cross-cell GRIM: baseline/end/change in one row constrain each other's precision |
grimmer() |
Impossible SDs for integer data (extends GRIM to standard deviations) |
grim_percentage() |
Impossible percentages from discrete counts (GRIM applied to percentages; debit() remains as a deprecated alias) |
sprite() |
Whether any valid dataset can produce the reported mean + SD |
correlation_bound() |
Impossible pre/post/change SD combinations (implied |r| > 1) |
check_ttest_paired() |
Recalculates paired t-test p-values from reported statistics |
check_ttest_independent() |
Recalculates independent t-test p-values |
check_anova_oneway() |
Recalculates one-way ANOVA F and p from group statistics |
check_chi_squared() |
Recalculates chi-squared from contingency tables |
sample_size_from_t() |
Back-calculates n from reported t and p |
effect_size_consistency() |
Cross-checks Cohen's d, p-values, and confidence intervals |
carlisle_stouffer_fisher() |
Tests whether Table 1 baseline p-values are too well-balanced |
check_sd_se_confusion() |
Flags likely SD/SE mix-ups given data range |
quick_sd_check() |
Checks SD plausibility against data range bounds |
check_contingency_table() |
Verifies row/column marginal totals are consistent |
benfords_law() |
Tests first-digit distribution against Benford's law |
variance_ratio_test() |
Flags suspiciously similar or divergent group variances |
check_frozen_sds() |
Flags suspiciously constant SDs across timepoints |
check_change_arithmetic() |
Verifies End - Baseline = reported Change |
check_sd_positive() |
Flags negative standard deviations |
Semantic analysis (embedding-based tools):
- LaTeX is cleaned to plain text, split into ~200-word overlapping chunks
- Chunks encoded using sentence-transformers (all-MiniLM-L6-v2, 384-dim), or TF-IDF as fallback
- Cosine similarity matrices between paper chunks and literature chunks power the analysis modules
Critical read (external papers):
- PDF text extracted via PyMuPDF
- Sections auto-detected (methods, results, discussion, conclusions)
- Four independent analyses: author COI, resolution mismatch, missing methods, overclaiming
Forensic statistics (data integrity):
- Reviewer transcribes summary statistics from paper tables
- 22 automated checks test internal consistency
- Results classified as pass, flag (suspicious), or fail (impossible)
The paper (PDF) (March 2026) describes the embedding-analysis core; the forensic-statistics and systematic-review modules postdate it. For those, Heathers (2025) is the forensic reference and docs/systematic-review.md is the design document.
See examples/annotate/ for a worked annotation spec, examples/forensic/ for a synthetic mini-paper with planted errors driving all three forensic CLI modes (table, text, --annotate), and examples/permanent-library/ for a reference paper-library integration.
PaperScope is developed using a multi-model feedback loop:
- Claude Code (Opus) writes features and runs analyses
- Codex reviews the code and output, producing detailed findings
- The human routes Codex's feedback back to Claude for fixes
- Repeat until Codex stops finding issues
This loop catches bugs that either model would miss alone -- Claude builds fast but can be overconfident about its own output; Codex is a thorough critic but doesn't write the fixes. The human's role is routing and judgment: deciding which findings matter and when the code is done.
The forensic statistics module went through three review passes this way, fixing 11 issues including empty-corpus crashes, false-negative verdicts, broken retraction detection, and over-aggressive author name filtering.
- Python 3.9+
numpy,scipy,scikit-learn,requests(all inrequirements.txt)PyMuPDFfor PDF text extractionsentence-transformers(optional, recommended -- better embeddings than TF-IDF fallback)matplotlib,networkx(optional -- for visualization)
MIT -- see LICENSE.