Skip to content
Adds text to PDF files using the cuneiform OCR software
Branch: master
Clone or download
Latest commit 4cbab8e Feb 14, 2019
Type Name Latest commit message Commit time
Failed to load latest commit information.
pdfocr.1 added support for unpaper Dec 27, 2015
pdfocr.rb fixes #37 Feb 14, 2019


pdfocr adds an OCR text layer to scanned PDF files, allowing them to be searched. It currently depends on Ruby 1.8.7 or above, and uses ocropus, cuneiform, or tesseract for performing OCR.


To use, run:

pdfocr -i input.pdf -o output.pdf

For more details, see the manpage.


pdfocr requires tesseract and hocr2pdf. These can be provided by installing the packages tesseract-ocr, tesseract-ocr-eng (or other languages you need), and exactimage from your distribution.


pdfocr was written by Geza Kovacs

pdfocr is hosted at

Christian Pietsch added tesseract support.

You can’t perform that action at this time.