zhuxiangyi opened a new pull request, #9955:
URL: https://github.com/apache/paimon/pull/9955

   ### Purpose
   
   `DataEvolutionFullTextRead.eval()` asks the native full-text reader for 
`candidateLimit(rowRangeStart, rowRangeEnd)` — every row of the split range — 
instead of the user's `limit`. The reader scores and returns every matching 
document of the shard, Java materializes them all into a `HashMap<Long, 
Float>`, and only then `topK(limit)` keeps the requested rows. A search for 10 
rows over a 200,000-row shard materializes ~200,000 candidates; the cost and 
memory of every full-text query grow with the number of matching documents 
instead of with `limit`.
   
   This dates from #8308, when compound queries (`multi_match`, cross-column 
`boolean`, `boost`) were composed in the Java read layer from per-column leaf 
reads, so a leaf had to return every candidate. Since #8467 the whole JSON DSL 
is executed by the native engine, which scores compound queries per document 
inside its own top-k collector, so the per-split top-k merged across splits 
already equals the global top-k. The primary-key path 
(`PrimaryKeyFullTextBucketSearch`) already passes the user `limit`.
   
   ### Change
   
   Pass `limit` to `FullTextSearch` and drop `candidateLimit`. Everything else 
(per-split include bitmaps, merge, raw fallback, final `topK(limit)`) is 
unchanged.
   
   ### Performance
   
   `NativeFullTextRowFilterTest.benchmarkRowFilterStrategies` (native engine, 
200,000 rows, 100 categories, `limit = 10`, best of 10 after warm-up, same 
machine, the only change being this one file):
   
   | scenario | before | after |
   |---|---|---|
   | no filter | 307.6 ms | **2.4 ms** |
   | row filter via btree, 1% selective | 6.9 ms | 3.7 ms |
   | row filter via btree, ~99% of rows | 326.4 ms | **6.0 ms** |
   | over-fetch `limit × 100` + client-side filter | 320.5 ms | 5.9 ms |
   | unindexed column, `scalar-index.search-mode=full` | 42.8 ms | 9.9 ms |
   
   The whole benchmark drops from 16.2 s to 2.9 s. The remaining cost of the 
dense-filter case is the bitmap handed to the engine; the remaining cost of the 
last case is reading the filter column.
   
   ### Tests
   
   No new tests: the result of every search is unchanged, which the existing 
suites pin. `FullTextSearchBuilderTest` (44, including 
`testCompoundFullTextSearchUsesFullLeafCandidatesBeforeFinalTopK`, whose boost 
query is scored per document by both the test indexer and the native engine), 
`paimon-core` `table.source.*` + `globalindex.**` (361), the `paimon-full-text` 
module against the native engine (including the filter-then-rank equality 
checks), and Spark `FullTextSearchTest` + `HybridSearchTest` (20) pass.
   
   ### API and Format
   
   No API, option, or format changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to