Skip to content
Adds text to PDF files using the cuneiform OCR software
Branch: master
Clone or download
Latest commit 4cbab8e Feb 14, 2019
Permalink
Type Name Latest commit message Commit time
Failed to load latest commit information.
COPYRIGHT
pdfocr.1 added support for unpaper Dec 27, 2015
pdfocr.rb fixes #37 Feb 14, 2019
readme.md

readme.md

pdfocr

pdfocr adds an OCR text layer to scanned PDF files, allowing them to be searched. It currently depends on Ruby 1.8.7 or above, and uses ocropus, cuneiform, or tesseract for performing OCR.

Using

To use, run:

pdfocr -i input.pdf -o output.pdf

For more details, see the manpage.

Dependencies

pdfocr requires tesseract and hocr2pdf. These can be provided by installing the packages tesseract-ocr, tesseract-ocr-eng (or other languages you need), and exactimage from your distribution.

Credits

pdfocr was written by Geza Kovacs

pdfocr is hosted at http://github.com/gkovacs/pdfocr

Christian Pietsch added tesseract support.

You can’t perform that action at this time.