[4.0.7] pdf识别,文字不全,但转为png图片识别,文字全了 #5580
Unanswered
doublex640
asked this question in
Q&A
Replies: 1 comment
|
@doublex640 这页 PDF 没有文字层,pypdfium2 和 pdfminer 都读到 0 字符,所以两条路都走 OCR。我把同一页按 100/150/200/432 dpi 各渲染一遍再识别,六列字每次都认全了,432 dpi 反而比 100 dpi 少一个字,所以渲染分辨率在这页上不是分水岭。你导出的那张 2151x3488 PNG,也只是页内那张 300 dpi 扫描图的重采样。 你提到的 左边那栏是用哪个后端跑的,pipeline 还是 vlm,完整命令是什么?能把 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
我这有一本 pdf,其中有一页文字始终识别不全(丢失部分文字),于是我单独抽取这一页内容,抽取成 pdf 和 png 图片这两种格式分别 ocr。也供你们测试。


经济解释 卷1 科学说需求_12.pdf
另外,在 ocr 这一页 pdf 时,日志中有大量的 Ignoring wrong pointing object。
以下是识别结果,左侧是 pdf 识别结果,右侧是 png 识别结果。
All reactions