Skip to content

v0.2.0 — OCR: read a scan, and hand back a searchable copy

Choose a tag to compare

@benbergner benbergner released this 13 Sep 17:28
· 31 commits to main since this release

0.1.0 shipped without OCR and said so. This adds it.

Added

  • pdf_ocr — reads pages that carry no text, and can return the document with
    a real text layer added: pdf_ocr(ref, pages=None, lang="eng", dpi=200, output="text", force=False).
    • output="text" for the text alone, "pdf" for the searchable copy as an
      artifact, "both" for both. The copy is your original with an invisible text
      layer over each page — the scan is untouched, so pages look exactly as they
      did and the file barely grows, but the text now selects, searches and copies
      like any other PDF's.
    • Per page you get the text, chars, words, mean_confidence out of 100 and
      low_confidence_words. Confidence is the number that tells you when to look at
      the page yourself, since unreliable OCR still reads fluently.
    • A page that already has text is listed in skipped rather than read again;
      force=True overrides that when the existing text came out of some other OCR
      as gibberish.
    • Most documents are one call: up to 50 pages, up to 40 seconds of recognizing,
      whichever runs out first. A longer one comes back with truncated: True,
      pages_remaining and a summary spelling out the next call. Pass each result
      into the next call and the layers accumulate in one document; overlapping
      ranges are safe, because a page that already has its layer is skipped.
  • pdf_check_text reports OCR'd pages. They come back ocred: true, counted
    by ocred_pages, and the summary says the text is recognition output rather
    than the document's own.

Requires tesseract, only for OCR

pdf_ocr uses tesseract, which is a
program rather than a Python package: brew install tesseract,
sudo apt install tesseract-ocr, or winget install UB-Mannheim.TesseractOCR.

It's looked for when you call pdf_ocr, not at startup, so every other tool works
without it and installing it later needs no reinstall or restart. Until it's there
the error names the command for your platform. Language packs install separately;
lang="deu" on an English-only machine tells you which languages you do have
rather than reading German as English.

Fixed

  • pdf_check_text on a freshly OCR'd document said scanned, needs_ocr: True.
    The text layer goes on top of the scan rather than replacing it, so the page still
    reads as a full-page image, and a sparse page's character count stayed under the
    threshold — both signals still said "scan" about a document that had just been
    read. It now recognizes the pages pdf_ocr has read.

If you already have it

uvx benspdf-mcp resolves the latest version, so there's nothing to do beyond
installing tesseract if you want OCR.