Skip to content

Releases: amos689/paper-preflight

v0.8.0

Choose a tag to compare

@amos689 amos689 released this 08 Oct 13:33
fc38637

Books, software, models, datasets, RFCs and web links are checked against their own registries; invented venues and near-miss titles are caught; and the project gains settings, a Python API, a JSON Schema and a page for every rule. On a held-out week of real papers, run once: 0.9 false positives per 100 references (1,127 references), at 0.43 s per reference.

Fewer references left undecided

  • Books without a DOI are confirmed by Open Library: Deep Learning (Goodfellow, Bengio and Courville), Pearl 1988, Cormen 2022, ... A record confirms a book only when its title, an author and an edition's year all fit; it never accuses.
  • Software linked from an entry is confirmed by GitHub, PyPI or CRAN; models and datasets by Hugging Face or OpenML. A repository, package or OpenML dataset that does not exist is worth a look (REF019, an info).
  • RFCs are verified by their number (type={RFC}, number={8446}, "RFC 791", rfc-editor.org links).
  • Web pages are never judged true or false, but a linked page that is gone and that the Wayback Machine never archived is reported (REF021, an info).
  • References too new to be indexed are remembered and searched for again once a day; check --recheck searches at once.

On the 280 development papers, these verify 48 books, 47 pieces of software and 29 models and datasets that were undecided, with no warning or error gained.

More fabrications caught

  • Invented venues (REF020): a work found only as a preprint, cited with a journal or conference no catalogue has (OpenAlex, dblp, Crossref).
  • Reworded titles of arXiv papers, when arXiv's record shows every title the paper has had.
  • Authors dropped from the middle of a list are named in the info about omitted authors.

HALLMARK test_public, any issue: recall 90.3% → 93.5% at a 1.9% false-positive rate; invented venues 78% → 97%, near-miss titles 70% → 85%.

For projects, scripts and agents

  • Project settings in paper-preflight.toml or [tool.paper-preflight]: rules and entries to ignore, severities, optional sources to turn off, fail-on.
  • A Python API: paper_preflight.check_paper(path) returns frozen dataclasses.
  • A JSON Schema for the JSON report (docs/schema/check-report.schema.json).
  • bib fix --level unsafe cites a preprint as its published version (entry type, venue, year, volume, pages, DOI; the eprint is kept).
  • A page for every rule (docs/rules/, English and Chinese), also printed by explain and shown with SARIF alerts in GitHub code scanning; a user guide (docs/README.md); a compatibility policy for 1.0 (docs/stability.md).
  • NCBI_API_KEY is now used (PubMed at ten requests a second), and the GitHub Action's cache is saved on every run.

Fixed

  • All seven false positives of the 0.7.0 held-out week: articles a registry dates only by their online appearance, books digitised later, registries' notation in titles (<tex-math>, "M sub sun"), a paper's code offered as its published version.
  • Two entries sharing a DOI are reported again (CIT004) when they are papers in different proceedings (a regression in 0.7.0).
  • SARIF reports linked to rule pages that did not exist.

Measured

0.7.0 0.8.0
Held-out week of real papers, false positives per 100 references 0.7 (1,014 refs) 0.9 (1,127 refs)*
HALLMARK test_public, any issue: recall / false-positive rate 90.3% / 1.9% 93.5% / 1.9%
HALLMARK test_public, fabrication: recall / false-positive rate 50.2% / 0.6% 50.2% / 0.6%
GPTZero's 151 hallucinated references, flagged / verified 139 / 0 139 / 0
Cold cache, seconds per reference 0.43 0.43

* The week measuring 0.8.0 (papers first submitted 2026-06-24..30) is one sampled from before the first evaluation week and never used until then, so that the release need not wait for next week's papers. It was checked months after its papers were written, when the works they cite are better indexed, and may read a little better than a fresh week; the next fresh week will be run as a control. Every flag was reviewed by hand: 132 real problems, 10 false positives, 1 unclear.

Full list of changes: CHANGELOG.

v0.7.0

Choose a tag to compare

@amos689 amos689 released this 08 Oct 08:25
cb93bb9

Manuscripts beyond LaTeX: Word, Markdown, Quarto, R Markdown and Typst, and bibliographies in CSL-JSON, RIS and YAML. On a new held-out week of real papers, run once: 0.7 false positives per 100 references (1,014 references), at 0.43 s per reference.

New inputs

uvx paper-preflight check paper.docx     # Word
uvx paper-preflight check paper.qmd      # Markdown, Quarto, R Markdown: [@key]
uvx paper-preflight check paper.typ      # Typst: @key and #bibliography(...)
uvx paper-preflight check refs.json      # CSL-JSON; also .ris, .yml/.yaml
  • Word (.docx).
    • Citations inserted by Zotero, Mendeley or EndNote carry the manager's own record of each work in the document's field codes, and these are read first.
    • Otherwise Word's source manager is read; failing that, the reference list as typed, after its "References" heading.
    • Positions in the report are paragraph numbers. The document is never edited.
  • Markdown, Quarto, R Markdown, Pandoc. The keys cited are checked against the bibliography the front matter or _quarto.yml names, as for a LaTeX project. A Quarto book is read across all its chapters and the files they include; a bookdown book across its rmd_files.
  • Typst. @key and #cite(<key>) are checked against #bibliography(...), with labels, comments and packages left out.
  • Bibliography formats. CSL-JSON, RIS, Typst's Hayagriva and Pandoc's CSL YAML. The same works in any of them are read exactly as from BibTeX.

How it was checked

  • Word. Each development paper's 924 references were written into Word documents and checked against the paper's own .bib. Same verdict:

    Document Same verdict
    Reference-manager field codes 99.6%
    Typed list, APA 94.0%
    Typed list, IEEE 96.9%
    Typed list, NLM 96.0%
  • Real documents.

    • Word: 12 manuscripts from Zenodo.
    • Markdown, Quarto and R Markdown: three JOSS papers, a Quarto book (mlr3book) and a bookdown book (Tidy Modeling with R).
    • Typst: a paper with its Hayagriva bibliography.
    • Real CSL-JSON and RIS exports.
    • The cited keys match a plain search in every project.

Fixed

Found on those documents; these apply to every input:

  • Publisher URLs: a publisher's page URL that carries the DOI and more path is no longer read as a second, non-existent DOI.
  • Garbled names: names a registry stored with the wrong encoding ("DuÅ¡ica") are read as meant.
  • Double surnames: an author cited with both Spanish surnames is matched to a record with the first only.
  • Chapters of one book: chapters that each give the book's DOI are no longer reported as duplicates.
  • Government gazettes: Federal Register notices are not called "not found".
  • Plain-text reader:
    • NLM titles in numbered parts ("Dust. IV. The Silicate ...");
    • years next to arXiv numbers;
    • APA editions, translations and "et al." after initials;
    • GOST titles of two sentences;
    • initials beyond A–Z.

Measured

0.6.0 0.7.0
New held-out week of real papers, false positives per 100 references 1.5, then 0.8 (868, 786 refs) 0.7 (1,014 refs)
Development batches replayed (13 weeks, 260 papers) no verdict or finding changed
GPTZero's 151 hallucinated references, flagged / verified 139 / 0 139 / 0
HALLMARK test_public, fabrication recall / false-positive rate 50.2% / 1.9% 50.2% / 1.9%
PDF reference lists, same verdict as the .bib 749 / 814 (92%) 756 / 820 (92%)

Install

uvx paper-preflight check path/to/paper

v0.6.0

Choose a tag to compare

@amos689 amos689 released this 08 Oct 06:43
96967e4

Faster, with fewer false alarms, and it now reads references written whole in a note. Two new held-out weeks of real papers were each run once: the first gave 1.5 false positives per 100 references (868 references); the second, run after the fixes for the first, gave 0.8 (786 references).

Faster

  • About twice as fast on a cold cache. A live run now takes about half a second per reference: 0.47–0.48 s on the two held-out weeks, against 1.22 s the week before.
  • How:
    • Crossref's polite pool is used at the limit its own answers announce, whenever a contact address (PAPER_PREFLIGHT_EMAIL) is set. Every source slows down at once if an answer announces a stricter limit.
    • Semantic Scholar is asked about an entry as soon as dblp and Crossref have both answered for it.

Fixed

Seventeen false positives from the last two held-out weeks:

  • ACM Digital Library links (dl.acm.org/doi/10.5555/...) are no longer read as DOIs.
  • Venue names: "Knowledge Discovery & Data Mining" is recognised as KDD.
  • Preprint archives: IACR ePrint and ECCC are never treated as a preprint's published version.
  • Registries' own errors:
    • garbled UTF-8, HTML entities, LaTeX font switches and AAS markup are cleaned from titles and names;
    • "Prof." given as a first name is ignored;
    • Cambridge Books Online's digitisation dates are not compared;
    • a dblp year contradicted by its own key and the DOI is corrected.
  • Version differences:
    • an author's name is reported only if no record of the work has it;
    • a preprint's title is only a hint when the entry cites the journal version.
  • Two clearer messages: a name with "Jr." where BibTeX reads the given name, and a DataCite record that only has an arXiv paper's latest version.

Added

  • References written whole in an @misc note are now read as plain text. One paper wrote all 46 of its references this way; 37 are now verified, where none were before.
  • The plain-text reader also takes links in angle brackets, name particles ("de Sá"), two-sentence titles, and a year in brackets.

Measured

0.5.3 0.6.0
New held-out week of real papers, false positives per 100 references 1.3 (1,067 refs) 1.5, then 0.8 (868, 786 refs)
Seconds per reference, live run 1.22 0.47
GPTZero's 151 hallucinated references, flagged / verified 139 / 0 139 / 0
HALLMARK test_public, fabrication recall / false-positive rate 50.2% / 1.9% 50.2% / 1.9%

Install

uvx paper-preflight check path/to/paper

v0.5.3

Choose a tag to compare

@amos689 amos689 released this 08 Oct 05:07
60b8a8e

Fewer false alarms on real papers. On a new held-out week (1,067 references), run once with these changes: 1.3 false positives per 100 references, alongside 83 real problems.

Fixed

Eleven of the previous held-out week's thirteen false positives:

  • Double surnames cited by their first part. dblp's "Sabela Ramos Garea" is the entry's "Ramos, Sabela". Both surnames must have four letters or more.
  • Group authors. Names ending in "Laboratory", "Laboratories" or "Institute" (The International Brain Laboratory) are groups, not missing people.
  • "et al." inside a name. "Paul Ralph et al." written as one name now cuts the author list short.
  • Subtitles Crossref keeps apart. Crossref search results are compared with their subtitle ("How Many Interviews Are Enough?"). A record found only that way does not stop the Semantic Scholar rescue, so a reissue cannot hide the original (Marr's Vision, 1982). On the development batches ten more books are verified this way.
  • Years:
    • a workshop paper matches its arXiv version a year later;
    • a book printed in September or later also counts the next year, its copyright year (Cover & Thomas; Rasmussen & Williams).
  • Venues:
    • dblp's challenge volumes ("HECKTOR@MICCAI");
    • law reviews and "Open Review" are treated as sources no index covers;
    • a URL in the journal field counts as a link.

Measured

0.5.2 0.5.3
New held-out week of real papers, false positives per 100 references 1.3 (975 refs) 1.3 (1,067 refs)
GPTZero's 151 hallucinated references, flagged / verified 139 / 0 139 / 0
HALLMARK test_public, fabrication recall / false-positive rate 50.2% / 1.9% 50.2% / 1.9%

Install

uvx paper-preflight check path/to/paper

paper-preflight 0.5.2

Choose a tag to compare

@amos689 amos689 released this 07 Oct 17:02
dadc708

Fewer false alarms, in English papers and for Chinese-language works cited in English. On a new held-out week of real papers (975 references), run once with these changes: 1.3 false positives per 100 references (0.5.1's week: 2.4).

Fixed

  • Chinese-language works cited in English ("... (in Chinese)", the journal's English name, pinyin authors) are no longer called "not found". The open indexes hold these works under their Chinese titles, if at all. Their titles and authors are now compared as the registries store them.
    • GitHub sample. On a sample of such references from GitHub .bib files, the share with a warning or error fell from 27% to 8%.
    • PMC sample. On known-real references in this form from PubMed Central, "not found" fell from 60% to 0.
  • "The DOI does not exist" (REF002) is confirmed with the Handle API. doi.org's agency lookup also gives that answer when a handle server is down.
  • Eleven of the previous held-out week's nineteen false positives:
    • subtitles left out;
    • titles cut short;
    • dataset names after titles;
    • middle-name nicknames;
    • joint meetings and ACM SIG newsletters;
    • Black Hat talks;
    • an MNRAS volume year;
    • a footnote run into a registry's title.

Measured

0.5.1 0.5.2
New held-out week of real papers, false positives per 100 references 2.4 (780 refs) 1.3 (975 refs)
GPTZero's 151 hallucinated references, flagged / verified 139 / 0 139 / 0
HALLMARK test_public, fabrication recall / false-positive rate 50.2% / 1.9% 50.2% / 1.9%

Install

uvx paper-preflight check path/to/paper

paper-preflight 0.5.1

Choose a tag to compare

@amos689 amos689 released this 07 Oct 12:12
fda7944

Two fixes, one of them for compliance.

Fixed

  • A CNKI DOI is no longer resolved. doi.org answers metadata requests for CNKI's DOIs with a redirect to chndoi.org, whose robots.txt disallows every agent, and 0.4.x and 0.5.0 followed that redirect once per cited CNKI DOI. Now only doi.org's agency lookup is used: the DOI exists, and the reference is otherwise left undetermined.
  • Eight of the ten false positives from 0.5.0's held-out week of real papers:
    • author lists that a registry cut short (ESO's DataCite records, KISTI's) when the arXiv record has everyone;
    • collaboration names ("MAGPI Team");
    • "RDKit Contributors" style credits;
    • software cited through its Zenodo concept DOI;
    • a title cut short, now reported as a title missing words (REF012), not as wrong authors.

Measured

  • Development batches. On the eight batches of real papers used for development, the changes remove 8 false positives and find 2 titles cut short; no real problem is lost.

  • HALLMARK and GPTZero. Both are unchanged.

  • A new held-out week of real papers (780 references), run once with these changes: 2.4 false positives per 100 references. That is over our 1.5 target and above 0.5.0's week (1.0). None of the 19 comes from this release's changes. They are:

    • other forms of one person's name;
    • pages of a reference site cited with its handbook's arXiv ID;
    • real works no source indexes or whose title was cut short;
    • dataset names appended to titles;
    • venue name forms.

    They are the next fixes, to be measured on yet another held-out week.

Install

uvx paper-preflight check path/to/paper

paper-preflight 0.5.0

Choose a tag to compare

@amos689 amos689 released this 07 Oct 10:53
755991d

Catches more, abstains less, misfires less. On a new held-out week of real papers, collected before any change in this release: 1.0 false positives per 100 references (0.4.0: 1.2).

Added

  • Journal articles cited without a title ("MNRAS 249, 523", as astronomy and physics cite) are looked up on Crossref by journal, volume and first page, with the first author's surname. A record counts only at exactly those coordinates and with the same first author; nothing found is never "not found". On an astronomy paper that cites 86 works, most of them this way, undetermined references fall from 88% to 9%.

Changed

  • Short titles at a venue dblp indexes in full (three or four words, a past year) can be "not found". A found title that begins with the cited one, such as a real paper cited by its first words, still means "cannot determine".
  • One work under two records whose titles differ only in hyphens ("Trade-off", "Tradeoff") is no longer ambiguous. Invented authors on such a real title are reported.
  • Plain-text and PDF reference lists: a name with a lower-case part ("Yun chen Chen") no longer stops the authors and title from being read.

Fixed

  • bib fix --apply wrote advice into the file. An invalid DOI or arXiv ID was replaced by "(remove or correct the field)" or "(correct the arXiv ID)", even at the safe level. Such fields are now left for you, and Zotero's "2311.07911 [cs]" gets the bare ID.
  • Seven of the previous held-out batch's nine false positives:
    • names written "Rouse D. M.";
    • affiliation marks inside registry names;
    • databases cited by their access year;
    • IEEE early-access years;
    • workshops in a joint dblp volume;
    • Substack posts;
    • technical reports written as conference papers.

Measured

0.4.0 / 0.4.1 0.5.0
Real papers, held-out week, false positives per 100 references 1.2 (753 refs) 1.0 (962 refs)
GPTZero's 151 hallucinated references, flagged / verified 135 / 0 139 / 0
HALLMARK test_public, fabrication recall 49.0% 50.2%
HALLMARK test_public, any-issue recall 89.1% 90.3%
HALLMARK test_public, false-positive rate 2.2% 1.9%

Install

uvx paper-preflight check path/to/paper

paper-preflight 0.4.1

Choose a tag to compare

@amos689 amos689 released this 07 Oct 04:58
fd4e9a6

Smoother first runs. Findings are unchanged: a replay of the sixth held-out batch of real papers gives identical results.

Changed

  • An overloaded arXiv no longer leaves the run "incomplete". DataCite registers every arXiv paper, so when the arXiv API does not answer, those references are still verified. Previously the run was marked incomplete anyway (RUN001, exit code 2). Now it completes, and a new info finding, RUN002, says which references were checked through DataCite and that withdrawals and earlier version titles were not. If DataCite does not answer either, the run is incomplete as before.
  • Repeated suggestions are one line. On a sound paper the report used to list "this preprint has since been published" a dozen times. In the text report, REF015 (published preprints) and REF016 (available DOIs) are now summarised in one line once a rule has three or more findings; --details lists each. JSON, SARIF and all counts are unchanged.
  • MCP tools describe their parameters. Every parameter now has a description in the tool schema (none had one), and each tool says when to use it instead of the others.
  • README: a logo, badges and a centred header, in English and Chinese, which also show on PyPI and Glama. The roadmap no longer lists an online demo as done.

Install

uvx paper-preflight check path/to/paper

v0.4.0

Choose a tag to compare

@amos689 amos689 released this 05 Oct 07:48
d44231c

Caught in the wild, tried in the browser. paper-preflight 0.4.0 measures itself on hallucinations found in real NeurIPS and ICLR papers, gets an online demo, installs into more coding agents, and lets your agent judge whether cited works support your sentences.

pip install --upgrade paper-preflight

Measured on hallucinations that got past peer review

GPTZero published 151 hallucinated references it found in NeurIPS 2025 papers and ICLR 2026 submissions, each confirmed by its staff. We pasted them as plain text, as the papers printed them:

References Flagged Cannot determine Verified
151 135 (89%) 16 0
  • Not one is verified.
  • The 16 left undecided are:
    • web pages and blog posts;
    • titles too short to search with confidence;
    • real titles given with invented authors where several works share the title;
    • two references too garbled to take apart.
  • A caveat: GPTZero's own search tool found these, so this is recall on hallucinations a search can find. Details: evals/results/gptzero.md.

New

  • Online demo (Hugging Face Space, in space/). You can:

    • give an arXiv ID;
    • upload a .bib, .bbl, .tex, .txt or .pdf, or a .zip of an Overleaf project;
    • paste references.

    You get the report, the suggested .bib fixes, and JSON and SARIF files to download. Files are deleted after the check; no API keys are held.

  • More agents. npx skills add amos689/paper-preflight (or gh skill install) installs the skill into Codex, Gemini CLI, GitHub Copilot, Cursor and other agents.

  • Your agent judges citation support. The MCP tool preflight_cited_passages returns, for each sentence citing a work, the claim and the cited work's best passages. Your agent confirms a claim only with a quote copied word for word, and never calls a citation wrong.

    • On 100 citations from our AI-labelled gold set, a Claude agent confirmed 39% of the real ones, against 9% for the local model.
    • Every confirmation was at least partially supported.
    • The labels come from the same model family, so take this as indicative.
  • support checks names. A citation that only names what it cites ("Adam \cite{kingma}") is confirmed when the cited work's title carries the name.

Fixed

  • The plain-text reader (pasted lists, .txt, PDFs) reads more references in full:
    • accented first names;
    • titles that end in a question mark;
    • venues that start with an edition or a year;
    • arXiv links broken by a space.
  • False alarms from the sixth batch of real papers. 17 of its 20 false positives are gone:
    • registry records in odd forms: DataCite creators stored as one name, Zenodo release titles, entities escaped twice, journal codes, lost letters;
    • Semantic Scholar's venues;
    • four more ways of writing one name;
    • a work cited without its subtitle.

How it measures

On a seventh batch of real papers, collected after every fix in this release (20 arXiv papers first submitted 2026-08-12..18, every flag reviewed by hand): 753 references, 86 flags, 77 real problems and 9 false positives: 1.2 false positives per 100 references (0.3.0 had 1.9 on its held-out batch).

Full list: CHANGELOG.

v0.3.0

Choose a tag to compare

@amos689 amos689 released this 04 Oct 15:44
d1d68ec

Does the cited work say it? paper-preflight 0.3.0 adds an experimental support command. For each citation in a LaTeX paper, it looks in the cited work for a passage that says what the citing sentence claims. The release also fixes the false alarms that a fifth week of real papers turned up.

pip install "paper-preflight[support]==0.3.0"
paper-preflight support path/to/paper --download-model --all

support (experimental)

  • Where it reads. It reads each cited work's text: the arXiv source, an open-access full text or PDF, or the abstract when nothing more can be had.
  • How it scores. The passages are ranked for the claim. A small local model, HHEM-2.1-open (0.4 GB, Apache-2.0), scores the best four.
  • What it says. "Confirmed", with the passage quoted word for word. Otherwise "could not confirm", with the reason: no text, only the abstract, or no passage close enough.
  • What it never says. It never says a citation is wrong.
  • What stays local. Claims are scored on your machine. Only the cited works' identifiers go out, to fetch their text, which is cached.
  • Other options. arxiv:<id> works as a target. --model minicheck and --model factcg choose other verifiers. --format json gives everything.

How it measures

The gold set has 298 pairs of a citing sentence and a cited work, from 55 arXiv papers under CC licences. Of these, 250 are real citations, and 48 are mis-citations made by swapping the cited work.

The labels were made by AI models, not by experts:

  1. two annotators (Claude Sonnet and Claude Opus);
  2. a third (Claude Fable) for the pairs they disagreed on;
  3. then an adjudication and a 20% spot check.
"Confirmed" was right 97% [95% CI 85%, 99%] (34 of 35)
Real citations confirmed 12% (31 of 250)
Mis-citations confirmed 0
  • It confirms little. Most real citations paraphrase loosely, point at a dataset or method, or rest on text that is not openly available.
  • A low score is no evidence of a mis-citation. A low score meant a mis-citation 23–45% of the time, so support does not accuse. In real papers about 3% of citations are not supported by the cited work.

Every number is in evals/results/support.md.

What's fixed in check

  • Names written another way are one person:
    • initials without dots ("Brown, JR", as Google Scholar exports them);
    • a generational suffix ("Smith IV");
    • Danny for Daniel.
  • An organisation is not an author. One named "… Research" first on arXiv is not taken for the first author.
  • Titles keep words set in \texttt, \textsf, \textup, \textmd or \mbox. Before, "\texttt{torch.compile}: …" was read as ": …".
  • These no longer count as a reworded title:
    • a chapter's title field that also names its book ("…, in New Phenomena in Subnuclear Physics");
    • TeX math left in dblp's titles.
  • Books and years:
    • a cited book is no longer bound to a later chapter that reprints it;
    • a volume of a multi-volume book matches a record that leaves the volume out of its title;
    • a JMLR paper may carry the year after dblp's volume year.

On the fifth batch of real papers, whose false positives these fixes were made for, 6 of its 19 false positives remain (0.6 per 100 references), and all 66 real problems are still flagged. These are development numbers. A sixth batch, collected after this release, will measure the fixes held out.

Full list: CHANGELOG.