[
https://issues.apache.org/jira/browse/TIKA-4946?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121960#comment-18121960
]
Tim Allison edited comment on TIKA-4946 at 10/2/26 12:32 PM:
-------------------------------------------------------------
Claude recommends the following. If we have anyone with actual image processing
experience who wants to chime in, please do.
{noformat}
1. Contrast check (standard deviation)
A blank page is nearly uniform, so its grayscale standard deviation is tiny.
bashmagick page.png -colorspace Gray -format "%[fx:standard_deviation]\n" info:
Blank scans usually come out well under ~0.02, and text pages are much higher.
Shave the borders first (-shave 3%x3%) so scanner edges, shadows, and punch
holes don't inflate the number.
2. Ink ratio (fraction of dark pixels)
bashmagick page.png -colorspace Gray -shave 3%x3% -median 3 \
-threshold 60% -negate -format "%[fx:mean]\n" info:
-median 3 (or -despeck) removes speckle noise before thresholding. Use a fixed
threshold here, not -auto-threshold otsu. On a uniform page, Otsu still splits
the noise into two classes, and a blank sheet can come out as ~50% "ink." This
one bites people.
3. Trim test
Trim away everything close to the background color and see what's left.
bashmagick page.png -shave 3%x3% -fuzz 15% -trim -format "%w %h\n" info:
A blank page collapses to something tiny, like 1 1, while a page with content
keeps a real bounding box. This is quick and surprisingly effective.
4. Connected components (text-likeness)
Text produces many small blobs; noise produces a few specks.
bashmagick page.png -colorspace Gray -shave 3%x3% -threshold 60% -negate \
-define connected-components:verbose=true \
-define connected-components:area-threshold=15 \
-connected-components 8 null: | tail -n +2 | wc -l
A real text page typically has hundreds or thousands of components, and a blank
one has a handful. This is the best ImageMagick-only discriminator, because it
also catches pages that have marks but no text (a stray line, a stamp edge).
{noformat}
was (Author: [email protected]):
Claude recommends the following. If we have anyone with actual image processing
experience who wants to chime in, please do.
{noformat}
1. Contrast check (standard deviation)
A blank page is nearly uniform, so its grayscale standard deviation is tiny.
bashmagick page.png -colorspace Gray -format "%[fx:standard_deviation]\n" info:
Blank scans usually come out well under ~0.02, and text pages are much higher.
Shave the borders first (-shave 3%x3%) so scanner edges, shadows, and punch
holes don't inflate the number.
2. Ink ratio (fraction of dark pixels)
bashmagick page.png -colorspace Gray -shave 3%x3% -median 3 \
-threshold 60% -negate -format "%[fx:mean]\n" info:
-median 3 (or -despeck) removes speckle noise before thresholding. Use a fixed
threshold here, not -auto-threshold otsu. On a uniform page, Otsu still splits
the noise into two classes, and a blank sheet can come out as ~50% "ink." This
one bites people.
3. Trim test
Trim away everything close to the background color and see what's left.
bashmagick page.png -shave 3%x3% -fuzz 15% -trim -format "%w %h\n" info:
A blank page collapses to something tiny, like 1 1, while a page with content
keeps a real bounding box. This is quick and surprisingly effective.
4. Connected components (text-likeness)
Text produces many small blobs; noise produces a few specks.
bashmagick page.png -colorspace Gray -shave 3%x3% -threshold 60% -negate \
-define connected-components:verbose=true \
-define connected-components:area-threshold=15 \
-connected-components 8 null: | tail -n +2 | wc -l
A real text page typically has hundreds or thousands of components, and a blank
one has a handful. This is the best ImageMagick-only discriminator, because it
also catches pages that have marks but no text (a stray line, a stamp edge).
{noformat}
> Don't send blank pages for OCR
> ------------------------------
>
> Key: TIKA-4946
> URL: https://issues.apache.org/jira/browse/TIKA-4946
> Project: Tika
> Issue Type: Task
> Reporter: Tim Allison
> Priority: Major
>
> Some vlm-based OCR tools will hallucinate an entire journal article page's
> worth of content if we send a blank image.
> Ideally, the ocr tools would prevent this, but we should try to do something
> on our side to prevent this.
> Not sure where to put this: in the PDFParser or in the OCR engines?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)