> By looking at the content, understanding what it is talking about and knowing that vegetables are washed before chopping, we can determine that A C B D is the correct order. Determining this algorithmically is a difficult problem.
Sorry, this is a bit off-topic regarding PDF extraction, but it distracted me greatly while reading...
I'm pretty sure the intention was A B C D (cut then wash). Not sure why the author would not use alphabet order for the recipe...
[edit] Sorry, I made it read to a colleague and he mentioned the A B C D annotations were probably not in the original document. This was not clear at all for me while reading, and if they are not included it's indeed hard to find the correct paragraph order.
Even if the ABCD was in the original document, how would the computer figure out it's supposed to indicate the order?
And of course, even if the letters were there in the original document, it would be clear to a human that they're incorrect because it doesn't make sense to wash vegetables after cutting.
Sorry, this is a bit off-topic regarding PDF extraction, but it distracted me greatly while reading...
I'm pretty sure the intention was A B C D (cut then wash). Not sure why the author would not use alphabet order for the recipe...
[edit] Sorry, I made it read to a colleague and he mentioned the A B C D annotations were probably not in the original document. This was not clear at all for me while reading, and if they are not included it's indeed hard to find the correct paragraph order.