Skip to content

v17.11.0

Latest

Choose a tag to compare

@github-actions github-actions released this 28 Aug 22:15
e63fab4

Enhancements

  • New --max-ocr-image-mpixels downsamples the image sent to OCR when a page
    exceeds the given size, which bounds the largest consumer of memory. The
    visible page is never downsampled, so output appearance is unaffected in every
    mode; what it trades is OCR accuracy on very high resolution scans. A 34
    megapixel page that peaks at 492 MB peaks at 325 MB under
    --max-ocr-image-mpixels 8, and recognizes the same text. See the "Memory"
    section of the performance documentation for how to size a memory limit.
  • watcher.py (the watcher extra) gained a configurable output layout and
    conflict policy. These are watcher-only changes; they do not affect the
    ocrmypdf library API.
    • New OCR_OUTPUT_STRUCTURE setting (--output-structure): FLAT
      (default, all outputs directly in the destination directory),
      YEAR_MONTH ({destination}/{year}/{month}/{filename}, same layout as
      the old OCR_OUTPUT_DIRECTORY_YEAR_MONTH=1), or HIERARCHY, which
      mirrors the input directory tree under the destination, e.g.
      input/a/b/c.pdfoutput/a/b/c.pdf.
    • New OCR_ON_CONFLICT setting (--on-conflict) controls what happens
      when the intended destination file already exists: SUFFIX (default)
      writes name (1).pdf, name (2).pdf, ... in the OS style; SKIP logs
      and leaves the file unprocessed; OVERWRITE is the old behavior.
      Behavior change: the default is now SUFFIX, so existing output
      files are no longer silently overwritten.
    • Both settings now apply equally to OCR_OUTPUT_DIRECTORY and to the
      archive directory used by OCR_ON_SUCCESS_ARCHIVE; previously the
      archive directory was always flat and silently overwrote on a name
      collision.
    • Output and archive filenames and directory components are sanitized
      for filesystems more restrictive than the input side (e.g. an SMB
      share): characters illegal on Windows/SMB (<>:"/\|?* and control
      characters) are replaced with _, trailing dots/spaces are stripped,
      and reserved DOS device names (CON, PRN, AUX, NUL, COM1-9,
      LPT1-9) are prefixed with _.
    • OCR_OUTPUT_DIRECTORY_YEAR_MONTH is now deprecated in favor of
      OCR_OUTPUT_STRUCTURE=YEAR_MONTH. It is still honored and logs a
      deprecation warning; if both are set, OCR_OUTPUT_STRUCTURE wins.
    • See the "Watched folders with watcher.py" section of the batch
      processing documentation for the full description, including a note
      for SMB users about client-side directory/file-info caching delays.

Performance

  • The image sent to OCR is no longer decoded and re-encoded when nothing needs
    to change it. On a page with no pre-existing text to mask and no filtering
    plugin -- the ordinary case for a scanned document -- the rasterized page
    already is the OCR image, so it is linked rather than rewritten. Producing
    it was costing 3.3 seconds and a full-size buffer on a 34 megapixel page, out
    of about 15 seconds for the whole file. Building the mask is now deferred
    until a text area actually needs blanking, since that step is what forced the
    decode. The minimum Pillow version is raised to 12, because deciding whether
    anything decoded the image reads an attribute whose shape settled in Pillow 11
    (pi-heif already required Pillow 11.1, so the effective floor barely moves).
  • Reduced peak memory on files with large images by about a third at default
    settings. On a 34 megapixel page the process tree peaked at 752 MB and now
    peaks at 492 MB. Two changes account for it: rasterizing a page no longer
    allocates a third full-page buffer to correct PDFium's rounding of the
    rendered size by a pixel or two, and freed heap memory is now returned to the
    operating system before the OCR engine runs, instead of counting against our
    resident set for as long as the engine is working. Rasterization is also
    faster, since correcting the size no longer resamples the whole page.
  • Several OCR jobs may now run concurrently in a single Python process. The API
    previously held a lock for the whole duration of ocrmypdf.ocr(), so a second
    call in another thread had to wait for the first to finish. Plugin state is
    now guarded by a readers-writer lock: installing a plugin set takes it
    exclusively, and a job holds it shared for its run. An in-flight job therefore
    cannot have the plugin infrastructure it depends on replaced underneath it,
    while jobs that are past installation proceed concurrently.
  • A plugin set is now installed once per interpreter and reused, rather than
    being reinstalled on every call. Previously each ocrmypdf.ocr() call
    re-executed plugin modules given as file paths and rebound them in
    sys.modules. Plugins must not rely on being re-executed for each job, and
    must not store per-job state on the plugin manager, which concurrent jobs
    requesting the same plugin set now share.
  • Jobs requesting different plugin sets serialize against each other, since
    installing the second set must wait for the first set's jobs to finish. Use
    the same plugin set across concurrent jobs, or separate processes.
  • Image optimization is substantially faster on documents with large images.
    An image stored as /FlateDecode with a PNG predictor already holds exactly
    what a PNG IDAT chunk holds, so it is now repackaged as a PNG directly
    instead of being decoded to a bitmap and re-encoded. On a 6-page document
    containing one 9000x9000 image, the optimization step went from 3.4s to 0.6s,
    and to 0.04s together with the JPEG change below; total runtime went from
    10.0s to 7.0s. Output is unchanged: the compressed
    data is reused verbatim. Images that are not in a directly repackageable form
    still take the previous path.
  • During image optimization, we decoded all JPEGs, even if the code
    path was an optimization setting with the decoded JPEG would be never be
    re-encoded (below --optimize 2). We now decode only on code paths that
    use the decoded JPEG. Output is unchanged.
  • An uncompressed image (one with no /Filter) no longer produces a spurious
    "could not be processed by the optimizer" warning. Such an image raised
    IndexError internally, which the optimizer's best-effort handler caught and
    reported as a warning; it is now recognized and skipped quietly.
  • Removed an unreachable branch in the image optimizer that claimed to handle
    1 bit per component images in an ICC-based colorspace. An earlier check sends
    every 1 bpc image to the JBIG2 pass, which handles ICC-based images by
    neutralizing the profile before extracting, so the branch could never run.
  • Removed the process-wide lock that serialized worker pools across all
    Executor instances. Executor.pool_lock is retained but no longer acquired,
    and is deprecated; it will be removed in a future major release. The invariant
    it protected - that only one progress bar renders on the shared console - is
    now enforced by the progress bar, which disables itself if another bar already
    owns the console.
  • Note that N concurrent jobs each configured with jobs=M may now spawn up to
    N*M workers, where previously they were serialized to M. Size jobs
    accordingly.
  • Known limitation of concurrent in-process jobs: they must use the same
    max_image_mpixels. Pillow's decompression-bomb limit is interpreter-global
    and the last job to set it wins. Use separate processes to run jobs with
    differing configurations.

Fixes

  • Fixed a latent use-after-free in the pypdfium2 rasterizer. to_pil() lets
    Pillow alias PDFium's bitmap buffer for some formats -- grayscale renders
    among them, which is every mono and grayscale page -- and the buffer was freed
    immediately afterwards, leaving Pillow reading memory PDFium had released.
  • Windows: OCRmyPDF no longer prints [WinError 2] The system cannot find the file specified warnings while it searches for Ghostscript and Tesseract
    (#1671). These
    messages came from probing registry keys that simply don't exist when the
    programs were installed by a package manager such as Scoop, or not installed
    at all. Since the search then continues elsewhere and usually succeeds, these
    failures are normal, and are now logged at debug level, naming the location
    that was searched. If a program genuinely cannot be found, OCRmyPDF still
    reports that as an error.
  • Windows: fixed a crash when the PROGRAMFILES environment variable pointed to
    a folder that does not exist.
  • --mode strip failed to remove OCR text layers that OCRmyPDF itself
    produced ({issue}1730). OCRmyPDF grafts its text layer as a Form XObject
    and stripping only examined the page content stream, so the invisible text
    was never found. The same flaw made --mode redo stack a second text layer
    on top of the old one instead of replacing it. Stripping now descends into
    Form XObjects. Thanks @Anai-Guo ({issue}1732).
  • Stripping now also resolves page /Resources inherited from an ancestor
    /Pages node, rather than only looking at the page's own resources.
  • A text layer that becomes empty after stripping is now removed from the page
    instead of being left behind as a vestigial Form XObject husk.
  • Setting clean_final on an existing options object in the Python API no
    longer leaves clean unset. --clean-final implies --clean, but the rule
    lived in a field validator that only ran while the options object was being
    constructed, so assigning to the attribute afterwards silently skipped it.
  • Validation errors for out-of-range or misspelled options now name the option
    they are about, and their wording comes from Pydantic rather than being
    hand-written, so it has changed slightly. For example, --jobs 999 now
    reports --jobs: Input should be less than or equal to 256. The set of
    accepted values is unchanged.