Repository navigation
v0.2.0 — OCR: read a scan, and hand back a searchable copy
0.1.0 shipped without OCR and said so. This adds it.
Added
pdf_ocr— reads pages that carry no text, and can return the document with
a real text layer added:pdf_ocr(ref, pages=None, lang="eng", dpi=200, output="text", force=False).output="text"for the text alone,"pdf"for the searchable copy as an
artifact,"both"for both. The copy is your original with an invisible text
layer over each page — the scan is untouched, so pages look exactly as they
did and the file barely grows, but the text now selects, searches and copies
like any other PDF's.- Per page you get the text,
chars,words,mean_confidenceout of 100 and
low_confidence_words. Confidence is the number that tells you when to look at
the page yourself, since unreliable OCR still reads fluently. - A page that already has text is listed in
skippedrather than read again;
force=Trueoverrides that when the existing text came out of some other OCR
as gibberish. - Most documents are one call: up to 50 pages, up to 40 seconds of recognizing,
whichever runs out first. A longer one comes back withtruncated: True,
pages_remainingand a summary spelling out the next call. Pass each result
into the next call and the layers accumulate in one document; overlapping
ranges are safe, because a page that already has its layer is skipped.
pdf_check_textreports OCR'd pages. They come backocred: true, counted
byocred_pages, and the summary says the text is recognition output rather
than the document's own.
Requires tesseract, only for OCR
pdf_ocr uses tesseract, which is a
program rather than a Python package: brew install tesseract,
sudo apt install tesseract-ocr, or winget install UB-Mannheim.TesseractOCR.
It's looked for when you call pdf_ocr, not at startup, so every other tool works
without it and installing it later needs no reinstall or restart. Until it's there
the error names the command for your platform. Language packs install separately;
lang="deu" on an English-only machine tells you which languages you do have
rather than reading German as English.
Fixed
pdf_check_texton a freshly OCR'd document saidscanned,needs_ocr: True.
The text layer goes on top of the scan rather than replacing it, so the page still
reads as a full-page image, and a sparse page's character count stayed under the
threshold — both signals still said "scan" about a document that had just been
read. It now recognizes the pagespdf_ocrhas read.
If you already have it
uvx benspdf-mcp resolves the latest version, so there's nothing to do beyond
installing tesseract if you want OCR.