Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I suppose the best approach is to combine OCR techniques while taking hints from the PDF structure.


An order of magnitude increase of time is very significant. If you're just processing a few documents with a lot of human oversight you may be right, but it's definitely not a generalised best approach, at least going by the article.


Or maybe use a text corpus to infer likely next words between 2 lines of characters so we can infer the reading order and columns using the text.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: