Replies: 1 comment
|
@heejun32 yes, and the behaviour you describe is real. I checked content-core's routing: with the engine on The fix belongs in content-core, the extraction library Open Notebook uses. Open Notebook only passes the engine choice through. I opened lfnovo/content-core#98 with the scope; since you offered to implement it, it's yours there. Two design points to settle in the PR:
Once it's released in content-core, Open Notebook picks it up with a dependency bump. Status: accepted → upstream (lfnovo/content-core#98). |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What are you trying to do?
I'm ingesting PDF sources (lecture slides, papers, scanned handouts) into Open Notebook. Some are text-based exports (e.g. from PowerPoint/Keynote) and some are actual scans, so the ingestion pipeline needs to decide, per document, whether OCR is required before extracting text.
What feels difficult or missing today?
Today, when Docling is enabled (OPEN_NOTEBOOK_ENABLE_DOCLING=true), scanned PDFs and images go through OCR, but there doesn't seem to be a lightweight pre-check that determines whether a PDF actually needs OCR before the heavier Docling/OCR pipeline runs. For text-based PDFs (a large share of real documents, including most lecture slide exports), this means the OCR/Docling path may still be invoked even though the text layer is already extractable, adding unnecessary latency and resource cost.
What outcome would help?
Text-based PDFs should skip the OCR/Docling path entirely and go straight to fast, local text extraction, while only the pages that genuinely lack a text layer get routed to OCR. Ideally this decision happens automatically and cheaply (on the order of milliseconds per document), without the user needing to know in advance whether a given PDF requires OCR.
How do you handle this today?
No response
Additional context, examples, or possible directions
No response
How would you like to participate?
Before posting
All reactions