[
https://issues.apache.org/jira/browse/TIKA-4898?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18116414#comment-18116414
]
Tim Allison commented on TIKA-4898:
-----------------------------------
It uses the heuristic of dropping any startxref with a value of 0. So, y, it
does not report incremental updates on that file, and that heuristic works on a
couple of thousand files I checked on recently. It is not 100% accurate.
This is just byte sniffing, and it would trigger on comments... for example.
If you have recommendations for making this more accurate, let me know. Is this
something that we could track during the parse inside PDFBox?
This particular issue focuses on keeping the same, ahem, quality, but getting
there faster. :D
> Improve xref scanning efficiency
> --------------------------------
>
> Key: TIKA-4898
> URL: https://issues.apache.org/jira/browse/TIKA-4898
> Project: Tika
> Issue Type: Task
> Reporter: Tim Allison
> Priority: Minor
> Attachments: PDFJS-6078-screen-annotations.pdf
>
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)