This is an automated email from the ASF dual-hosted git repository.
tballison pushed a commit to branch docs/4.0.x
in repository https://gitbox.apache.org/repos/asf/tika.git
The following commit(s) were added to refs/heads/docs/4.0.x by this push:
new aed69bb971 TIKA-4835 -- update performance doc (#3069)
aed69bb971 is described below
commit aed69bb9719c617712fc1d7f953575e916028f8a
Author: Tim Allison <[email protected]>
AuthorDate: Tue Aug 25 13:46:17 2026 -0400
TIKA-4835 -- update performance doc (#3069)
---
docs/modules/ROOT/pages/pipes/performance.adoc | 165 +++++++++++++++++++++----
1 file changed, 140 insertions(+), 25 deletions(-)
diff --git a/docs/modules/ROOT/pages/pipes/performance.adoc
b/docs/modules/ROOT/pages/pipes/performance.adoc
index 612d685c12..f35296d633 100644
--- a/docs/modules/ROOT/pages/pipes/performance.adoc
+++ b/docs/modules/ROOT/pages/pipes/performance.adoc
@@ -26,13 +26,14 @@ for it.
[NOTE]
====
-Closing the throughput gap of the default isolated mode — *without* giving up
the
-crash/OOM isolation it provides — is an area of active work. Treat the figures
on
-this page as a snapshot of the 4.0.0 release, not a fixed ceiling: expect the
-isolated-mode gap to narrow in future releases. Note also that this gap is
-specific to the *upload* endpoints on small documents — for file-system inputs
-and outputs the fetch/emit endpoints already match or beat 3.x while staying
-fully isolated (see <<endpoint-choice>>).
+The tables in the first half of this page were measured on the **4.0.0
+release** and are kept as that snapshot. After 4.0.0 shipped we traced the
+batch slowdown against 3.x to a specific cause — temp-file volume, not the
+pipes architecture — and fixed it for **4.1.0**. The findings, the fix, and a
+worked configuration from the box where it was diagnosed are in
+<<found-in-400>> and <<corpora-case-study>>. The upload-endpoint gap on tiny
+documents (<<endpoint-choice>>) is a separate, smaller effect and still
+applies.
====
== Deployment shapes
@@ -68,7 +69,7 @@ one parsing JVM serving all concurrency, with the
front-end/watchdog restarting
it on failure. The practical 4.x default choice is between **per-client**
(strongest isolation) and **shared-server** (highest throughput).
-== Throughput
+== Throughput (4.0.0 measurements)
The pipes per-request overhead — temp-spool of large payloads, socket IPC, and
result serialization — is roughly *fixed per request*. It therefore dominates
@@ -99,7 +100,7 @@ Two things to note:
smaller, independently-warming worker heaps.
[#endpoint-choice]
-== Endpoint choice: uploading bytes vs fetch-and-emit
+== Endpoint choice: uploading bytes vs fetch-and-emit (4.0.0 measurements)
How a document reaches the parser matters as much as the parsing mode. The
classic endpoints (`/tika`, `/rmeta`, `/meta`) receive the document *in the
HTTP
@@ -222,19 +223,121 @@ while a worker restarts.
xref:pipes/cpu-sizing.adoc[Forked-JVM CPU and Heap Sizing]; configure
per-parse
limits with xref:pipes/timeouts.adoc[Timeouts].
-== Benchmarking your own workload
+[#found-in-400]
+== What we found in 4.0.0, and what 4.1.0 changes
+
+Our own regression testing runs `tika-app` in batch mode over a
1.2-million-file
+corpus (file system in, file system out) on a box with spinning disks. That run
+took about 4 hours on Tika 3.x and about 7.5 hours on 4.0.0. The investigation
+that followed is worth summarizing, because the cause was not where the
+architecture suggested it would be.
+
+=== The cause: temp-file volume
+
+4.0.0 wrote **8–30 times more temp bytes** than 3.x for the same documents:
+
+* Digesting an embedded document (MD5/SHA-256 per embedded object) buffered a
+ rewindable copy of it that spilled to a temp file past 1 MB — one file per
+ embedded object, hundreds of thousands of them over a large corpus.
+* Several parsers and detectors asked for a `java.io.File` even when the
+ document was already in memory: the JPEG/TIFF/WebP metadata extractors, the
+ OLE2 container detector, the OpenDocument parser's inline pictures, the
+ digest of translated embedded streams, and the PDF incremental-update scan
+ each wrote the bytes out just to read them back.
+
+On a spinning-disk host where the temp directory, the corpus, and the outputs
+share spindles, every temp byte is a seek taken away from a corpus read or an
+extract write. Wall clock tracked temp volume almost linearly.
+
+=== What it was *not*
+
+Each of these was measured and ruled out, so they need not be re-chased:
+
+* Pipes IPC and result passback — about 1% of worker time.
+* The driver's emit path — with the default `DYNAMIC` strategy the workers
+ already write nearly all extract bytes themselves; more emitter threads made
+ no difference.
+* Reading each container twice for the digest pre-pass — the second read is
+ served from the page cache; disk reads were equal to or lower than 3.x's.
+* The parsers — on identical embedded objects most 4.x parsers are as fast or
+ faster; the JPEG parser is 4x faster in isolation.
+* 4.x extracting more embedded objects (it does, about 3% more) — negligible
+ cost.
+
+=== The fix (4.1.0, unreleased at the time of writing)
+
+* TIKA-4828/TIKA-4829: embedded zip entries are re-read from the archive on
+ rewind instead of being copied, and a process-wide `CacheMemoryBudget`
+ (seeded by the forked worker, tunable via
+ `-Dtika.pipes.cacheMemoryBudgetBytes` in `forkedJvmArgs`, `<=0` disables)
+ governs how much rewindable content stays in memory.
+* TIKA-4835: the parsers and detectors above no longer spool in-memory input
+ to disk; they rewind or read through a seekable channel, within the same
+ budget, and fall back to a file only past it.
+
+Measured on the diagnosis box (20,000 randomly sampled files of the corpus,
+page cache evicted before each run, extracts written to the corpus disk):
+
+[cols="3,1,1"]
+|===
+|Build and configuration |Temp written |Wall
-The numbers above are illustrative. Throughput depends on your document mix,
-document sizes, requested output format, concurrency, host CPU/heap, and disk
-speed (large payloads spool to a temp directory). Measure with *your* corpus:
+|Tika 3.x, 10 consumer threads (three runs) |0.33 GB |245 / 287 / 328 s
+|4.0.0, per-client, 8 workers (the 7.5 h shape) |10.5 GB |500 s
+|4.1.0 with TIKA-4828/29 only, corrected config |2.8 GB |341 s
+|4.1.0, shared server, 10 threads |0.26 GB |277 s
+|4.1.0, per-client, 7 workers |1.1 GB |287 s
+|===
-* Fix concurrency equal to the worker count so the comparison is apples to
- apples.
-* Exclude a warm-up phase — forked workers pay a one-time fork + JIT cost on
- their first requests.
-* Hold the output format constant across the versions or modes you compare.
-* Watch peak RSS across the whole process tree (front-end plus workers), not
- just one process.
+4.1.0 writes less temp than 3.x did, and matches or beats 3.x throughput while
+keeping process isolation. These are subset measurements on one host; we have
+not re-timed the full run, and remote emitters (Solr, OpenSearch, S3) were not
+measured.
+
+[#corpora-case-study]
+== A worked configuration: the regression-test box
+
+The diagnosis box is an 8-core/16-thread Ryzen with 62 GB of RAM and two
+spinning disks in RAID1, holding the temp directory, the 4 TB corpus, and the
+outputs on the same pair of spindles. The corpus is far larger than RAM, so
+every run is effectively cold-cache. The configuration we settled on:
+
+* **Per-client mode, `numClients = 7`.** Tika auto-injects
+ `-XX:ActiveProcessorCount` per fork as `(cores − 2) / numClients` but only
+ when that slice is at least 2. On 16 logical cores, 8 workers give 1.75 and
+ the cap is *skipped* — eight JVMs each sizing GC and JIT for 16 cores, which
+ is what the 7.5 h run did. Seven workers get 2 cores each and the cap
+ applies. Per-client cost about 4% versus shared-server here (287 s vs 277 s)
+ and dropped no files, where the shared worker loses the in-flight documents
+ of every other client when one document crashes it.
+* **`-Xmx4g` per fork.** The cache budget clamps to a quarter of the fork heap,
+ so this gives each worker 1 GB of in-memory rewind space; seven of them
+ leave about 30 GB for the page cache, which matters more than heap on a
+ cold-cache corpus. Archive-heavy corpora do better with `-Xmx6g` (1.5 GB
+ budget) — the tar/gz subset only reached parity with the budget raised.
+* Digest MD5 (SHA-256 measured within noise), default emit strategy, temp
+ directory left on disk.
+
+=== Reading your own deployment
+
+Three questions decided the result above, and they are cheap to answer for any
+box:
+
+* **Do temp, corpus, and outputs share spindles?** `/proc/mdstat`,
+ `lsblk`, or your cloud volume layout will say. If they do, temp volume is
+ wall clock; if temp is on separate fast storage, the 4.0.0 regression may
+ never have shown.
+* **Is the corpus larger than RAM?** If so, benchmark cold — evict the page
+ cache (or use a subset you have not touched) and write outputs to the real
+ destination. Warm-cache runs with outputs on tmpfs hid this entire problem
+ from us for weeks.
+* **Does your `numClients` trip the cap skip?** Check the startup log for the
+ `ActiveProcessorCount` decision; if it reports `skipped`, lower
+ `numClients` or set the cap yourself in `forkedJvmArgs`.
+
+Beyond those: fix concurrency equal to the worker count when comparing, exclude
+a warm-up phase, hold the output format constant, and watch peak RSS across the
+whole process tree rather than one JVM.
== Appendix: approaches considered and set aside
@@ -256,8 +359,20 @@ close it, recorded here so they need not be re-litigated:
* **Swapping the garbage collector.** ParallelGC helped tiny documents
marginally
and hurt larger ones — no reliable win over the default across a mixed
corpus.
-What is left is structural: the fixed per-request IPC + temp-spool + result
-serialization cost, and running several CPU-partitioned JVMs instead of one.
The
-productive directions are shrinking that per-request cost (RAM-disk temp
directory,
-keeping more payloads inline, leaner serialization) and fork-pool sizing — not
a
-single JVM flag.
+* **Driver-side emitter parallelism (`numEmitters`) and worker-direct emit
+ (`EMIT_ALL`) for file-system output.** Neither moved the batch numbers: under
+ the default `DYNAMIC` strategy the workers already write nearly all extract
+ bytes directly.
+* **A RAM-disk temp directory.** It does recover the 4.0.0 batch loss, and it
is
+ a useful *diagnostic* (if moving temp to tmpfs makes a run fast, temp volume
+ is your problem) — but do not run with it. Temp on tmpfs is bounded only by
+ RAM: a single large archive expanding into it can exhaust memory for every
+ process on the host, starve the page cache a cold-corpus run depends on, or
+ count against a container's memory limit and get the pod evicted. A slow run
+ is recoverable; that is not. 4.1.0 removes the temp writes instead of hiding
+ them.
+
+What remains structural on the *upload* endpoints is the fixed per-request IPC
+
+result-serialization cost and running several CPU-partitioned JVMs instead of
one;
+the productive directions there are keeping more payloads inline, leaner
+serialization, and fork-pool sizing — not a single JVM flag.