Repository navigation
Proposal: adaptive OCR fast path for born-digital PDFs #4238
Replies: 1 comment
|
Thanks for the numbers. My company has been looking at the same cost from the other side, and I think most of this pipeline is already in docling. The 8x on your file probably comes from a region-selection rule rather than a missing fast path. What the default already does. Since 2.116 (#3710), the default OCR mode is Why it doesn't behave that way on a form. Since 2.121 (#3981), a cluster also goes to OCR if any bitmap or vector shape touches it, even when it has a good text layer ( One caution before treating the OCR pass as unnecessary: 176 → 40 texts is a big drop. The layout postprocessor drops clusters that contain no text cells. So most of those ~136 elements are probably regions where docling found no programmatic text and got the text only from OCR. On an insurance form, the usual suspects are:
If that's what they are, The same point applies to tables. Table structure is predicted from the page image, so an unchanged shape with OCR on or off is expected. The cell text is the part worth comparing. Where the text layer is good, the OCR run can actually be the worse one (#4139). Where your proposal goes beyond what docling does now is the quality check. The PDF-aware selector trusts any cluster that has text cells. A text layer that exists but decodes badly (fonts with no ToUnicode map, private-use codepoints, mis-encoded glyphs) is never OCR'd. A cheap per-cluster check along those lines, sending the cluster to OCR when it fails, would fit neatly into the existing selector without a new pipeline. That seems like the piece worth designing here, once the shape trigger is settled. |
Uh oh!
There was an error while loading. Please reload this page.
Proposal: adaptive OCR fast path for born-digital PDFs
Hi Docling team,
we ran a small benchmark on a 6-page born-digital insurance-style PDF and found a potentially useful performance optimization opportunity.
All document-specific and customer-identifying data have been removed from the examples below.
Test document
Results
Default pipeline with OCR enabled
Docling reported:
The structural quality was very good, especially for table reconstruction and document hierarchy.
Same document with
--no-ocrDocling reported:
This is roughly a 7.9x reduction in Docling processing time.
Interesting result
The table reconstruction was unchanged with OCR disabled.
For example, the main table retained the same logical shape:
All 6 detected tables had the same logical shape and content structure with OCR enabled and disabled.
OCR did add substantially more text elements and logical groups, but it did not improve the table reconstruction in this born-digital PDF.
Suggestion
Would it make sense to add an adaptive fast path for born-digital PDFs?
Possible pipeline:
In other words, OCR could be selectively activated per page or even per region instead of being part of the full default path when native text is already reliable.
This could potentially preserve most of Docling's structural quality while significantly reducing latency on born-digital documents.
Why this may be useful
In our test, the main performance cost appeared unnecessary for table reconstruction:
The trade-off was mainly in richer text extraction and grouping rather than table accuracy.
A hybrid mode such as:
might provide a useful middle ground between full OCR quality and fast native-PDF processing.
We would be happy to provide additional anonymized benchmark details if useful.
All reactions