DominikSuess opened a new pull request, #3143:
URL: https://github.com/apache/jackrabbit-oak/pull/3143
## Summary
Two isolated fixes in `oak-search-mongot`, found while investigating a ~18s
query-latency
anomaly during a Mongot-vs-Elasticsearch benchmark evaluation. Both were
empirically
verified against a live MongoDB Atlas cluster (not just unit tests).
### 1. `MongotIndex.aggregateCursor()` — cursor batch size (primary bug)
`aggregateCursor()` hardcoded the MongoDB driver's cursor `batchSize` to
`fetchSizes[0]` (10) from `MongotIndexDefinition.getQueryFetchSizes()`
(default
`{10, 100, 1000}`). Each batch is a full network round trip, so a 2000+ hit
query required ~200 round trips.
`fetchSizes[1]`/`fetchSizes[2]` were declared but **never read anywhere** —
dead config, despite the array's shape suggesting adaptive/tiered batching
was
intended but never implemented.
**Fix:** use `Arrays.stream(fetchSizes).max().getAsInt()` (the largest
configured tier) instead of the smallest.
**Empirical validation** (direct `mongosh` run against the live Atlas
cluster, same pipeline, varying only `batchSize`):
| batchSize | latency |
|---|---|
| 10 | 24,943 ms |
| 100 | 2,812 ms |
| 1000 | 759 ms |
Rationale for safety: the MongoDB driver only holds one batch in memory at a
time regardless of the configured `batchSize` (not the whole result set), and
the server enforces a hard ~16MB cap per batch response, so a larger
`batchSize` can't cause unbounded per-batch memory/time.
**Documented follow-up (out of scope here):** there's no `$limit` pushdown in
`MongotIndex.pipeline()`, and the driver's cursor `batchSize` is fixed for
the
cursor's lifetime once set. A truly adaptive/limit-aware batching scheme
would
need a row-limit hint threaded from the backend-neutral `IndexPlan`/`Filter`
down through the `FulltextIndex` SPI (`oak-search` module) — flagged in code
comments as a candidate for a shared, implementation-neutral improvement
rather than a mongot-only special case.
### 2. `MongotIndex.streamingCursor()` — `Cursor.getSize()` drains the cursor
`streamingCursor()` wired `rows::getSize` (`MongotResultIterator.getSize()`)
as the `Cursor`'s `SizeEstimator`. That fully drains the entire result cursor
— through the same batched round trips fixed above — the moment anything
calls `Cursor.getSize()` / JCR's `RowIterator.getSize()` (a common "N results
found" UI pattern), even when the caller never intends to iterate all rows.
Meanwhile `MongotIndex.getSizeEstimator(IndexPlan)` was already implemented
(cheap, index-level `numDocs()`) but never actually invoked anywhere in this
class — dead code — unlike `ElasticIndex`, which wires its equivalent,
query-filtered `getDocCountFor()` into its own non-facet query path.
`numDocs()` itself isn't a like-for-like replacement: it's a whole-index
count, not filtered by the query's predicate, so reusing it as-is would
silently return the wrong (larger) count.
**Fix:** added a query-specific `countMatches()` helper that appends `$count`
to the already-built pipeline — a cheap server-side count with no document
bodies returned — and wired that into `streamingCursor()` in place of the
drain-based estimator. The pre-existing `getSizeEstimator(IndexPlan)`
override
is now documented as intentionally unused (kept only to satisfy the
`FulltextIndex` SPI contract).
## Validation
- Full reactor build (`mvn install -DskipTests`, JDK 17) succeeds.
- `oak-search-mongot` unit tests: 91 run, 87 pass. The remaining 4 failures
are pre-existing, unrelated to this change — they require Docker
(`testcontainers`) for the Mongot test server, which isn't available in the
sandbox this was verified in.
- Fix 1 root cause was independently confirmed by directly querying the
live Atlas cluster used for the benchmark evaluation (see table above),
not just by code inspection.
## Scope
Changes are isolated entirely to
`oak-search-mongot/src/main/java/org/apache/jackrabbit/oak/plugins/index/mongot/query/MongotIndex.java`,
in two commits (primary bug first).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]