Skip to content

Releases: benbergner/BensPDF

v0.3.0 — Text extraction: a PDF's own words, page by page

Choose a tag to compare

@benbergner benbergner released this 13 Sep 19:06

Added

  • pdf_extract_text — the text a PDF already holds:
    pdf_extract_text(ref, pages=None, output="text", layout=False).
    • This is the cheap, exact path. The characters come from the file, so a page is
      milliseconds and nothing is guessed — reach for it first, and for pdf_ocr only
      when a page carries no text at all. pdf_check_text tells you which you have.
    • Text arrives per page, with the page number beside it, so an answer can cite
      page 8 instead of quoting a document-sized blob. Each page also carries its
      characters and words.
    • pages takes the same ranges as the other tools: "1-20", "3", "1,5,9-12"
      or "all".

Reading something book-length

One call returns about 50,000 characters — roughly 12,500 tokens, which is what a
client can hold alongside a conversation. Past that it stops, sets truncated: True,
and names what it didn't reach in pages_remaining, with stopped_for saying which
budget it met and the summary spelling out the next call.

output="txt" is the other way through: the whole extraction goes into one .txt
artifact and the result keeps the counts alone, so any length is a single call. Pages
in the file are separated by a form feed, the plain-text page break, so
text.split("\f") gives them back. output="both" adds the text inline up to the
same budget and names the pages that live in the file only, in text_omitted_pages.
Pass the artifact to export to keep it.

Text worth a second look

Most PDFs say what their glyphs mean, in a /ToUnicode map. A page that has no map
and comes back holding characters that decoded to nothing is flagged
text_suspect, with unmappable_characters counting them and suspect_pages
collecting the page numbers. Both signals have to appear together, because plenty of
perfectly readable files are missing the map — which keeps the flag rare enough to be
worth acting on. When it shows up, pdf_render_pages shows you what the page really
says and pdf_ocr with force=True re-reads it.

empty_pages names the pages that hold no text at all, and a document where nothing
does points you at pdf_check_text and pdf_ocr.

Forms and tables

layout=True keeps the page's own spacing instead of reading it as prose, which is
what a form or a table is: values stay under their headings and columns stay apart.
Prose reads better without it, so it's per call rather than the default. A page whose
layout can't be worked out comes back as prose and is named in pages_read_as_plain,
so asking for spacing never costs you the text.

No new dependencies

Unlike pdf_ocr, this needs nothing beyond the package — no tesseract, no system
install. It reads the text that's already in the file.

If you already have it

uvx benspdf-mcp resolves the latest version, so there's nothing to do.

v0.2.0 — OCR: read a scan, and hand back a searchable copy

Choose a tag to compare

@benbergner benbergner released this 13 Sep 17:28

0.1.0 shipped without OCR and said so. This adds it.

Added

  • pdf_ocr — reads pages that carry no text, and can return the document with
    a real text layer added: pdf_ocr(ref, pages=None, lang="eng", dpi=200, output="text", force=False).
    • output="text" for the text alone, "pdf" for the searchable copy as an
      artifact, "both" for both. The copy is your original with an invisible text
      layer over each page — the scan is untouched, so pages look exactly as they
      did and the file barely grows, but the text now selects, searches and copies
      like any other PDF's.
    • Per page you get the text, chars, words, mean_confidence out of 100 and
      low_confidence_words. Confidence is the number that tells you when to look at
      the page yourself, since unreliable OCR still reads fluently.
    • A page that already has text is listed in skipped rather than read again;
      force=True overrides that when the existing text came out of some other OCR
      as gibberish.
    • Most documents are one call: up to 50 pages, up to 40 seconds of recognizing,
      whichever runs out first. A longer one comes back with truncated: True,
      pages_remaining and a summary spelling out the next call. Pass each result
      into the next call and the layers accumulate in one document; overlapping
      ranges are safe, because a page that already has its layer is skipped.
  • pdf_check_text reports OCR'd pages. They come back ocred: true, counted
    by ocred_pages, and the summary says the text is recognition output rather
    than the document's own.

Requires tesseract, only for OCR

pdf_ocr uses tesseract, which is a
program rather than a Python package: brew install tesseract,
sudo apt install tesseract-ocr, or winget install UB-Mannheim.TesseractOCR.

It's looked for when you call pdf_ocr, not at startup, so every other tool works
without it and installing it later needs no reinstall or restart. Until it's there
the error names the command for your platform. Language packs install separately;
lang="deu" on an English-only machine tells you which languages you do have
rather than reading German as English.

Fixed

  • pdf_check_text on a freshly OCR'd document said scanned, needs_ocr: True.
    The text layer goes on top of the scan rather than replacing it, so the page still
    reads as a full-page image, and a sparse page's character count stayed under the
    threshold — both signals still said "scan" about a document that had just been
    read. It now recognizes the pages pdf_ocr has read.

If you already have it

uvx benspdf-mcp resolves the latest version, so there's nothing to do beyond
installing tesseract if you want OCR.

v0.1.1 — one-click install, and a registry listing

Choose a tag to compare

@benbergner benbergner released this 12 Sep 20:51

No changes to the tools. Every verb behaves exactly as in 0.1.0 — this release
is about being easier to install and easier to find.

Added

  • One-click install for VS Code and Cursor, from buttons at the top of the README.
  • Claude Code setup, which was missing entirely: claude mcp add benspdf -- uvx benspdf-mcp.
  • Explicit setup sections for Cursor, Windsurf and Continue, each naming its own config file rather than pointing at another client's.
  • Listed on the official MCP registry as io.github.benbergner/benspdf.

Fixed

  • The Continue instructions were wrong. Continue uses YAML with mcpServers as a
    list, not the JSON object every other client here takes, so the old "same shape
    as above" produced a config it couldn't read.
  • The Cursor config and install link were missing "type": "stdio", which Cursor
    documents as required for stdio servers.

If you already have it

uvx benspdf-mcp resolves the latest version, so there is nothing to do. If you
pinned a version, there is no functional reason to move.

v0.1.0 — first release

Choose a tag to compare

@benbergner benbergner released this 12 Sep 18:04

First published release. BensPDF gives an AI assistant a set of PDF tools that
run on your own machine, over the Model Context Protocol.

Tools

  • pdf_page_count — how many pages, reading only the page tree so it stays cheap on long documents
  • pdf_metadata — title, author, dates, producer, keeping the Info dictionary and the XMP packet separate so you can see where they disagree
  • pdf_check_text — whether a file is readable text or a scan that needs OCR
  • pdf_page_layout — page sizes, orientation, rotation and print boxes, grouped by shape rather than listed per page
  • pdf_check_access — encryption, and what the file asks viewers to permit
  • pdf_render_pages — pages as images, for when appearance is the content: handwriting, signatures, charts, checkbox state

Results that produce files land in a scratch workspace and come back as ids, so
steps chain. export is the only tool that writes into your folders. list_artifacts
and discard manage what's in the workspace, and create_test_pdf_file gives you
something to try the tools on.

Install

Install uv, then point
your MCP client at uvx benspdf-mcp. Config blocks for Claude Desktop, VS Code,
Kiro, Cursor and the Codex clients are in the README.

Not in this release

Extracting and editing. There's no text extraction, no OCR, and no split, merge or
rotate. pdf_check_text will tell you a document's text is extractable; getting it
out is the next thing to land.

Privacy

Your PDFs are read locally and never uploaded. With a hosted model your questions
still reach that model — pair the tools with a local Ollama model through the
bundled benspdf-cli and nothing leaves the machine.

Alpha. Tool names and result shapes may still change.