Make the Azure Document Intelligence model configurable (prebuilt-layout in addition to prebuilt-read) #14299
Unanswered
isi4d
asked this question in
Feature Requests
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Description
Description
The remote OCR parser currently calls the Azure Document Intelligence
prebuilt-readmodel. I would like to request an optional setting to select the model, e.g.PAPERLESS_REMOTE_OCR_MODEL, defaulting toprebuilt-readso nothing changes for existing users.Why
1. Table structure is lost. With
prebuilt-read, tabular documents come back as a flat list of cell values, one per line, with no relationship between them. A line-item table from a repair quote arrives as: description, part number, quantity, unit price — each on its own line, in reading order. The text is searchable, but the structure needed for any later extraction (totals, amounts, quantities) is gone.prebuilt-layoutreturns tables as tables.2. Layout appears to recognise characters better, not just structure. Microsoft's own changelog states that the Layout model received "improvements to the OCR model for scanned text targeting improvements for single characters, boxed text, and dense text documents". That is relevant for forms and handwritten entries, which is where
prebuilt-readcurrently struggles most in my archive. On a printed quote,prebuilt-readmisread a licence plate and several phone numbers, while getting VIN, customer numbers and all monetary amounts correct.Compatibility
Per Microsoft's model overview, searchable PDF output is supported for both
prebuilt-readandprebuilt-layout, so the existing archive PDF generation path should be unaffected.Cost
prebuilt-layoutis priced higher per page thanprebuilt-read. Keepingprebuilt-readas the default means no existing user sees a cost change; those who need structure can opt in knowingly.Alternatives considered
Routing the call through a proxy that rewrites the model, and building a patched image. Both work around a one-line difference and create maintenance burden for something the parser could expose directly.
Other
No response
All reactions