[
https://issues.apache.org/jira/browse/TIKA-4946?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Tim Allison updated TIKA-4946:
------------------------------
Description:
Some vlm-based OCR tools will hallucinate an entire journal article page's
worth of fluent content if we send a blank image.
Ideally, the ocr tools would prevent this, but we should try to do something on
our side to prevent this.
Not sure where to put this: in the PDFParser or in the OCR engines or both?
was:
Some vlm-based OCR tools will hallucinate an entire journal article page's
worth of fluent content if we send a blank image.
Ideally, the ocr tools would prevent this, but we should try to do something on
our side to prevent this.
Not sure where to put this: in the PDFParser or in the OCR engines?
> Don't send blank pages for OCR
> ------------------------------
>
> Key: TIKA-4946
> URL: https://issues.apache.org/jira/browse/TIKA-4946
> Project: Tika
> Issue Type: Task
> Reporter: Tim Allison
> Priority: Major
>
> Some vlm-based OCR tools will hallucinate an entire journal article page's
> worth of fluent content if we send a blank image.
> Ideally, the ocr tools would prevent this, but we should try to do something
> on our side to prevent this.
> Not sure where to put this: in the PDFParser or in the OCR engines or both?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)