Repository navigation
Added
pdf_extract_text— the text a PDF already holds:
pdf_extract_text(ref, pages=None, output="text", layout=False).- This is the cheap, exact path. The characters come from the file, so a page is
milliseconds and nothing is guessed — reach for it first, and forpdf_ocronly
when a page carries no text at all.pdf_check_texttells you which you have. - Text arrives per page, with the page number beside it, so an answer can cite
page 8 instead of quoting a document-sized blob. Each page also carries its
charactersandwords. pagestakes the same ranges as the other tools:"1-20","3","1,5,9-12"
or"all".
- This is the cheap, exact path. The characters come from the file, so a page is
Reading something book-length
One call returns about 50,000 characters — roughly 12,500 tokens, which is what a
client can hold alongside a conversation. Past that it stops, sets truncated: True,
and names what it didn't reach in pages_remaining, with stopped_for saying which
budget it met and the summary spelling out the next call.
output="txt" is the other way through: the whole extraction goes into one .txt
artifact and the result keeps the counts alone, so any length is a single call. Pages
in the file are separated by a form feed, the plain-text page break, so
text.split("\f") gives them back. output="both" adds the text inline up to the
same budget and names the pages that live in the file only, in text_omitted_pages.
Pass the artifact to export to keep it.
Text worth a second look
Most PDFs say what their glyphs mean, in a /ToUnicode map. A page that has no map
and comes back holding characters that decoded to nothing is flagged
text_suspect, with unmappable_characters counting them and suspect_pages
collecting the page numbers. Both signals have to appear together, because plenty of
perfectly readable files are missing the map — which keeps the flag rare enough to be
worth acting on. When it shows up, pdf_render_pages shows you what the page really
says and pdf_ocr with force=True re-reads it.
empty_pages names the pages that hold no text at all, and a document where nothing
does points you at pdf_check_text and pdf_ocr.
Forms and tables
layout=True keeps the page's own spacing instead of reading it as prose, which is
what a form or a table is: values stay under their headings and columns stay apart.
Prose reads better without it, so it's per call rather than the default. A page whose
layout can't be worked out comes back as prose and is named in pages_read_as_plain,
so asking for spacing never costs you the text.
No new dependencies
Unlike pdf_ocr, this needs nothing beyond the package — no tesseract, no system
install. It reads the text that's already in the file.
If you already have it
uvx benspdf-mcp resolves the latest version, so there's nothing to do.