Skip to content
Adds text to PDF files using the cuneiform OCR software
Ruby Roff
Branch: master
Clone or download
Fetching latest commit…
Cannot retrieve the latest commit at this time.
Type Name Latest commit message Commit time
Failed to load latest commit information.
COPYRIGHT added manpage, README, and COPYRIGHT Apr 17, 2010
pdfocr.1 added support for unpaper Dec 27, 2015
pdfocr.rb fixes #37 Feb 14, 2019 dependencies May 15, 2017


pdfocr adds an OCR text layer to scanned PDF files, allowing them to be searched. It currently depends on Ruby 1.8.7 or above, and uses ocropus, cuneiform, or tesseract for performing OCR.


To use, run:

pdfocr -i input.pdf -o output.pdf

For more details, see the manpage.


pdfocr requires tesseract and hocr2pdf. These can be provided by installing the packages tesseract-ocr, tesseract-ocr-eng (or other languages you need), and exactimage from your distribution.


pdfocr was written by Geza Kovacs

pdfocr is hosted at

Christian Pietsch added tesseract support.

You can’t perform that action at this time.