[ 
https://issues.apache.org/jira/browse/TIKA-4901?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18116611#comment-18116611
 ] 

Hudson commented on TIKA-4901:
------------------------------

SUCCESS: Integrated in Jenkins build Tika » tika-main-jdk17 #1651 (See 
[https://ci-builds.apache.org/job/Tika/job/tika-main-jdk17/1651/])
TIKA-4901: don't inflate embedded doc twice for digesting (#3191) (github: 
[https://github.com/apache/tika/commit/c309b2a3dfa1323188eddd53adb4632791537883])
* (edit) CHANGES.txt
* (edit) tika-core/src/main/java/org/apache/tika/io/TikaInputStream.java
* (edit) tika-core/src/main/java/org/apache/tika/io/ReopenableSource.java
* (edit) tika-core/src/main/java/org/apache/tika/digest/InputStreamDigester.java
* (edit) tika-core/src/main/java/org/apache/tika/io/TikaInputSource.java
* (edit) tika-core/src/test/java/org/apache/tika/io/ReopenableSourceTest.java
* (edit) 
tika-core/src/test/java/org/apache/tika/digest/InputStreamDigesterTest.java


> Don't inflate an embedded document twice to digest it
> -----------------------------------------------------
>
>                 Key: TIKA-4901
>                 URL: https://issues.apache.org/jira/browse/TIKA-4901
>             Project: Tika
>          Issue Type: Improvement
>            Reporter: Tim Allison
>            Priority: Minor
>             Fix For: 4.1.0
>
>
> :robot: description
> {noformat}
>    InputStreamDigester reads a stream to hash it, then rewinds so the parse 
> can read it again. For a file that costs nothing — the page cache serves 
>   the second read at ~6 GB/s. For an embedded document it is not a file: 
> every ReopenableSource is a re-openable supplier, and re-opening a zip entry
>    means inflating it again, measured at 480–670 MB/s. The page cache holds 
> the compressed container, so it does not help.
>   
>    TikaInputSource gains an advisory tryRetainInMemory(), default false. 
> ReopenableSource implements it by reusing its existing in-memory drain and 
>   budget reservation, refusing when the length is unknown, when the content 
> exceeds the 1 MB floor with no budget, or when the budget cannot cover it
>    — so it never spills and never starts a drain that cannot finish. 
> ensureOpen() and seekTo() now serve from the retained content when there is 
>    some. InputStreamDigester asks for retention after enableRewind; 
> FileSource and CachingSource take the default and are unaffected.
>   
>    Full recursive parse of 120 zips × 16 × 128 KB entries with embedded 
> digesting on, only tika-core swapped, 3 reps × 2 rounds: wall 3.875–4.081 s → 
>    3.047–3.141 s, CPU 4.59–5.64 s → 3.78–4.83 s. Output identical (digests 
> and extracted lengths checksum the same in every run). The durable number 
>    is the absolute one — about 3.4 ms per MB of embedded content, one inflate 
> — since the ratio depends on how expensive the embedded parser is; text 
>    entries make the inflate share look large.
>   
>   Limits: entries above the 1 MB floor need budget headroom or retention is 
> refused, and zip is the best case — an OLE2 document stream or a PST item
>   re-opens more cheaply.
>   Also fixes a latent bug in the same method: ensureOpen() at a non-zero 
> position re-read from byte 0 while continuing to report positions as if it 
>   had not.
> {noformat}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to