Replies: 5 comments
|
I first sold OCR 4 Decades Ago for recognising single drawn or written symbols and OCR is roughly as good now as it was then. That is it needs intelligence to spot the failures. So today we have infant like AI that can improve over the next 40 years much much faster. Ghostscript and MuPDF which is the base of SumatraPDF can do fairly well so compare the old image test I re ran this week via SumatraPDF and copy the characters to see how they compare. But the contrast between foreground and background will always need to be cleaned up whatever the thresholding system used. Note your browser (Edge) is often the best as using online AI so is able to compensate by looking across the web for fuzzy character patterns |
|
I tried out Ghostscript and the result is not very satisfying, see attached pdf. That's the type of "dirty scanned file" I was talking about. If it were a brand new digital scan, I imagine it would give you a very nice document at the end, but that's not the case here. |





Uh oh!
There was an error while loading. Please reload this page.
Hey there,
Been wanted to ask it ages ago, but never got the chance: which OCR software would you wholeheartedly recommend that can reliable work on not so clean, scanned documents, ebooks, pdfs? I'm planning to delve into some Letter Pages on ancient (50-60 years old) scanned comic books, however I'm not getting a whole lot of success with various software: switching characters, seeing dirt on the page as a symbol, can't differentiate between
8(Scharacters etc.Are there good free alternatives or at that point a purchasable, high end option is the only good choice?
What do you guys think?
Thank you in advance!
All reactions