Repository navigation
Releases: benbergner/BensPDF
Release list
v0.3.0 — Text extraction: a PDF's own words, page by page
Added
pdf_extract_text— the text a PDF already holds:
pdf_extract_text(ref, pages=None, output="text", layout=False).- This is the cheap, exact path. The characters come from the file, so a page is
milliseconds and nothing is guessed — reach for it first, and forpdf_ocronly
when a page carries no text at all.pdf_check_texttells you which you have. - Text arrives per page, with the page number beside it, so an answer can cite
page 8 instead of quoting a document-sized blob. Each page also carries its
charactersandwords. pagestakes the same ranges as the other tools:"1-20","3","1,5,9-12"
or"all".
- This is the cheap, exact path. The characters come from the file, so a page is
Reading something book-length
One call returns about 50,000 characters — roughly 12,500 tokens, which is what a
client can hold alongside a conversation. Past that it stops, sets truncated: True,
and names what it didn't reach in pages_remaining, with stopped_for saying which
budget it met and the summary spelling out the next call.
output="txt" is the other way through: the whole extraction goes into one .txt
artifact and the result keeps the counts alone, so any length is a single call. Pages
in the file are separated by a form feed, the plain-text page break, so
text.split("\f") gives them back. output="both" adds the text inline up to the
same budget and names the pages that live in the file only, in text_omitted_pages.
Pass the artifact to export to keep it.
Text worth a second look
Most PDFs say what their glyphs mean, in a /ToUnicode map. A page that has no map
and comes back holding characters that decoded to nothing is flagged
text_suspect, with unmappable_characters counting them and suspect_pages
collecting the page numbers. Both signals have to appear together, because plenty of
perfectly readable files are missing the map — which keeps the flag rare enough to be
worth acting on. When it shows up, pdf_render_pages shows you what the page really
says and pdf_ocr with force=True re-reads it.
empty_pages names the pages that hold no text at all, and a document where nothing
does points you at pdf_check_text and pdf_ocr.
Forms and tables
layout=True keeps the page's own spacing instead of reading it as prose, which is
what a form or a table is: values stay under their headings and columns stay apart.
Prose reads better without it, so it's per call rather than the default. A page whose
layout can't be worked out comes back as prose and is named in pages_read_as_plain,
so asking for spacing never costs you the text.
No new dependencies
Unlike pdf_ocr, this needs nothing beyond the package — no tesseract, no system
install. It reads the text that's already in the file.
If you already have it
uvx benspdf-mcp resolves the latest version, so there's nothing to do.
v0.2.0 — OCR: read a scan, and hand back a searchable copy
0.1.0 shipped without OCR and said so. This adds it.
Added
pdf_ocr— reads pages that carry no text, and can return the document with
a real text layer added:pdf_ocr(ref, pages=None, lang="eng", dpi=200, output="text", force=False).output="text"for the text alone,"pdf"for the searchable copy as an
artifact,"both"for both. The copy is your original with an invisible text
layer over each page — the scan is untouched, so pages look exactly as they
did and the file barely grows, but the text now selects, searches and copies
like any other PDF's.- Per page you get the text,
chars,words,mean_confidenceout of 100 and
low_confidence_words. Confidence is the number that tells you when to look at
the page yourself, since unreliable OCR still reads fluently. - A page that already has text is listed in
skippedrather than read again;
force=Trueoverrides that when the existing text came out of some other OCR
as gibberish. - Most documents are one call: up to 50 pages, up to 40 seconds of recognizing,
whichever runs out first. A longer one comes back withtruncated: True,
pages_remainingand a summary spelling out the next call. Pass each result
into the next call and the layers accumulate in one document; overlapping
ranges are safe, because a page that already has its layer is skipped.
pdf_check_textreports OCR'd pages. They come backocred: true, counted
byocred_pages, and the summary says the text is recognition output rather
than the document's own.
Requires tesseract, only for OCR
pdf_ocr uses tesseract, which is a
program rather than a Python package: brew install tesseract,
sudo apt install tesseract-ocr, or winget install UB-Mannheim.TesseractOCR.
It's looked for when you call pdf_ocr, not at startup, so every other tool works
without it and installing it later needs no reinstall or restart. Until it's there
the error names the command for your platform. Language packs install separately;
lang="deu" on an English-only machine tells you which languages you do have
rather than reading German as English.
Fixed
pdf_check_texton a freshly OCR'd document saidscanned,needs_ocr: True.
The text layer goes on top of the scan rather than replacing it, so the page still
reads as a full-page image, and a sparse page's character count stayed under the
threshold — both signals still said "scan" about a document that had just been
read. It now recognizes the pagespdf_ocrhas read.
If you already have it
uvx benspdf-mcp resolves the latest version, so there's nothing to do beyond
installing tesseract if you want OCR.
v0.1.1 — one-click install, and a registry listing
No changes to the tools. Every verb behaves exactly as in 0.1.0 — this release
is about being easier to install and easier to find.
Added
- One-click install for VS Code and Cursor, from buttons at the top of the README.
- Claude Code setup, which was missing entirely:
claude mcp add benspdf -- uvx benspdf-mcp. - Explicit setup sections for Cursor, Windsurf and Continue, each naming its own config file rather than pointing at another client's.
- Listed on the official MCP registry as
io.github.benbergner/benspdf.
Fixed
- The Continue instructions were wrong. Continue uses YAML with
mcpServersas a
list, not the JSON object every other client here takes, so the old "same shape
as above" produced a config it couldn't read. - The Cursor config and install link were missing
"type": "stdio", which Cursor
documents as required for stdio servers.
If you already have it
uvx benspdf-mcp resolves the latest version, so there is nothing to do. If you
pinned a version, there is no functional reason to move.
v0.1.0 — first release
First published release. BensPDF gives an AI assistant a set of PDF tools that
run on your own machine, over the Model Context Protocol.
Tools
pdf_page_count— how many pages, reading only the page tree so it stays cheap on long documentspdf_metadata— title, author, dates, producer, keeping the Info dictionary and the XMP packet separate so you can see where they disagreepdf_check_text— whether a file is readable text or a scan that needs OCRpdf_page_layout— page sizes, orientation, rotation and print boxes, grouped by shape rather than listed per pagepdf_check_access— encryption, and what the file asks viewers to permitpdf_render_pages— pages as images, for when appearance is the content: handwriting, signatures, charts, checkbox state
Results that produce files land in a scratch workspace and come back as ids, so
steps chain. export is the only tool that writes into your folders. list_artifacts
and discard manage what's in the workspace, and create_test_pdf_file gives you
something to try the tools on.
Install
Install uv, then point
your MCP client at uvx benspdf-mcp. Config blocks for Claude Desktop, VS Code,
Kiro, Cursor and the Codex clients are in the README.
Not in this release
Extracting and editing. There's no text extraction, no OCR, and no split, merge or
rotate. pdf_check_text will tell you a document's text is extractable; getting it
out is the next thing to land.
Privacy
Your PDFs are read locally and never uploaded. With a hosted model your questions
still reach that model — pair the tools with a local Ollama model through the
bundled benspdf-cli and nothing leaves the machine.
Alpha. Tool names and result shapes may still change.