Skip to content

v0.3.0 — Text extraction: a PDF's own words, page by page

Latest

Choose a tag to compare

@benbergner benbergner released this 13 Sep 19:06
· 31 commits to main since this release

Added

  • pdf_extract_text — the text a PDF already holds:
    pdf_extract_text(ref, pages=None, output="text", layout=False).
    • This is the cheap, exact path. The characters come from the file, so a page is
      milliseconds and nothing is guessed — reach for it first, and for pdf_ocr only
      when a page carries no text at all. pdf_check_text tells you which you have.
    • Text arrives per page, with the page number beside it, so an answer can cite
      page 8 instead of quoting a document-sized blob. Each page also carries its
      characters and words.
    • pages takes the same ranges as the other tools: "1-20", "3", "1,5,9-12"
      or "all".

Reading something book-length

One call returns about 50,000 characters — roughly 12,500 tokens, which is what a
client can hold alongside a conversation. Past that it stops, sets truncated: True,
and names what it didn't reach in pages_remaining, with stopped_for saying which
budget it met and the summary spelling out the next call.

output="txt" is the other way through: the whole extraction goes into one .txt
artifact and the result keeps the counts alone, so any length is a single call. Pages
in the file are separated by a form feed, the plain-text page break, so
text.split("\f") gives them back. output="both" adds the text inline up to the
same budget and names the pages that live in the file only, in text_omitted_pages.
Pass the artifact to export to keep it.

Text worth a second look

Most PDFs say what their glyphs mean, in a /ToUnicode map. A page that has no map
and comes back holding characters that decoded to nothing is flagged
text_suspect, with unmappable_characters counting them and suspect_pages
collecting the page numbers. Both signals have to appear together, because plenty of
perfectly readable files are missing the map — which keeps the flag rare enough to be
worth acting on. When it shows up, pdf_render_pages shows you what the page really
says and pdf_ocr with force=True re-reads it.

empty_pages names the pages that hold no text at all, and a document where nothing
does points you at pdf_check_text and pdf_ocr.

Forms and tables

layout=True keeps the page's own spacing instead of reading it as prose, which is
what a form or a table is: values stay under their headings and columns stay apart.
Prose reads better without it, so it's per call rather than the default. A page whose
layout can't be worked out comes back as prose and is named in pages_read_as_plain,
so asking for spacing never costs you the text.

No new dependencies

Unlike pdf_ocr, this needs nothing beyond the package — no tesseract, no system
install. It reads the text that's already in the file.

If you already have it

uvx benspdf-mcp resolves the latest version, so there's nothing to do.