Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
amelius
on March 3, 2020
|
parent
|
context
|
favorite
| on:
What's so hard about PDF text extraction?
I suppose the best approach is to combine OCR techniques while taking hints from the PDF structure.
wyattpeak
on March 3, 2020
|
next
[–]
An order of magnitude increase of time is very significant. If you're just processing a few documents with a lot of human oversight you may be right, but it's definitely not a generalised best approach, at least going by the article.
ldenoue
on March 3, 2020
|
prev
[–]
Or maybe use a text corpus to infer likely next words between 2 lines of characters so we can infer the reading order and columns using the text.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: