[
https://issues.apache.org/jira/browse/TIKA-4883?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18113418#comment-18113418
]
ASF GitHub Bot commented on TIKA-4883:
--------------------------------------
tballison merged PR #3145:
URL: https://github.com/apache/tika/pull/3145
> Don't dump raw text when OCR is triggered in AUTO mode in PDFs
> --------------------------------------------------------------
>
> Key: TIKA-4883
> URL: https://issues.apache.org/jira/browse/TIKA-4883
> Project: Tika
> Issue Type: Task
> Reporter: Tim Allison
> Priority: Minor
>
> In AUTO OCR mode in PDFs, if there's not enough content or if there are a lot
> of characters missing unicode mappings, we trigger OCR on that page.
> In the current implementation, we write that content to the page, and then
> make the determination. If OCR is triggered, we append the OCR content to the
> end of the page.
> It would be cleaner in AUTO mode, if we cached the page content, ran the
> score and then decided whether or not to run OCR WITHOUT including the
> original page content before making that determination.
> Or, in short, don't dump trash then concatenate OCR... go only with OCR in
> AUTO mode, when triggered.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)