zhuxiangyi opened a new pull request, #9955: URL: https://github.com/apache/paimon/pull/9955
### Purpose `DataEvolutionFullTextRead.eval()` asks the native full-text reader for `candidateLimit(rowRangeStart, rowRangeEnd)` — every row of the split range — instead of the user's `limit`. The reader scores and returns every matching document of the shard, Java materializes them all into a `HashMap<Long, Float>`, and only then `topK(limit)` keeps the requested rows. A search for 10 rows over a 200,000-row shard materializes ~200,000 candidates; the cost and memory of every full-text query grow with the number of matching documents instead of with `limit`. This dates from #8308, when compound queries (`multi_match`, cross-column `boolean`, `boost`) were composed in the Java read layer from per-column leaf reads, so a leaf had to return every candidate. Since #8467 the whole JSON DSL is executed by the native engine, which scores compound queries per document inside its own top-k collector, so the per-split top-k merged across splits already equals the global top-k. The primary-key path (`PrimaryKeyFullTextBucketSearch`) already passes the user `limit`. ### Change Pass `limit` to `FullTextSearch` and drop `candidateLimit`. Everything else (per-split include bitmaps, merge, raw fallback, final `topK(limit)`) is unchanged. ### Performance `NativeFullTextRowFilterTest.benchmarkRowFilterStrategies` (native engine, 200,000 rows, 100 categories, `limit = 10`, best of 10 after warm-up, same machine, the only change being this one file): | scenario | before | after | |---|---|---| | no filter | 307.6 ms | **2.4 ms** | | row filter via btree, 1% selective | 6.9 ms | 3.7 ms | | row filter via btree, ~99% of rows | 326.4 ms | **6.0 ms** | | over-fetch `limit × 100` + client-side filter | 320.5 ms | 5.9 ms | | unindexed column, `scalar-index.search-mode=full` | 42.8 ms | 9.9 ms | The whole benchmark drops from 16.2 s to 2.9 s. The remaining cost of the dense-filter case is the bitmap handed to the engine; the remaining cost of the last case is reading the filter column. ### Tests No new tests: the result of every search is unchanged, which the existing suites pin. `FullTextSearchBuilderTest` (44, including `testCompoundFullTextSearchUsesFullLeafCandidatesBeforeFinalTopK`, whose boost query is scored per document by both the test indexer and the native engine), `paimon-core` `table.source.*` + `globalindex.**` (361), the `paimon-full-text` module against the native engine (including the filter-then-rank equality checks), and Spark `FullTextSearchTest` + `HybridSearchTest` (20) pass. ### API and Format No API, option, or format changes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
