Adds text to PDF files using the cuneiform OCR software
Ruby Roff
Switch branches/tags
Nothing to show
Clone or download
Latest commit 7a7a567 May 15, 2017
Failed to load latest commit information.
COPYRIGHT added manpage, README, and COPYRIGHT Apr 17, 2010
pdfocr.1 added support for unpaper Dec 27, 2015
pdfocr.rb added support for unpaper Dec 27, 2015 dependencies May 15, 2017


pdfocr adds an OCR text layer to scanned PDF files, allowing them to be searched. It currently depends on Ruby 1.8.7 or above, and uses ocropus, cuneiform, or tesseract for performing OCR.


To use, run:

pdfocr -i input.pdf -o output.pdf

For more details, see the manpage.


pdfocr requires tesseract and hocr2pdf. These can be provided by installing the packages tesseract-ocr, tesseract-ocr-eng (or other languages you need), and exactimage from your distribution.


pdfocr was written by Geza Kovacs

pdfocr is hosted at

Christian Pietsch added tesseract support.