PDFBox And OCR: This repo is for various tests with PDFBOX and OCR (Tess4J) Currently Working Features From Sample PDF: Creating pdf(s) Appending pages to existing pdf Splitting pdf into single pdfs Splitting pdf into single pdfs by range Splitting pdf into png image files Extracting all text from pdf Goals For This Repo I want to create an invoice scanner that extracts relevant information Priority Info To Pull From Invoices: Customer Invoice Number Order Number Date Tracking Number Serial Number Etc