JingsongLi opened a new pull request, #10050: URL: https://github.com/apache/paimon/pull/10050
## What changed - Allow the BTree TopN file selector to stop after several selected files jointly provide enough guaranteed rows with disjoint row-id ranges. - Add a bounded BTree equality lookup for ordinary `LIMIT` scans. It reads at most the requested number of row IDs and skips the BTree posting when metadata proves that every row in an index file has the same value. - Use the bounded path only for a single equality predicate with complete index coverage and unchanged data files. Queries with residual filters, deletion vectors, partial coverage, or additional data pruning keep the existing path. ## Performance Local planning benchmark on 100,000 rows, `f1 = 'dog' LIMIT 10`, with 5 warmups and 20 measured iterations. The comparison used the same scan with a no-op bucket filter to force the existing full-index path; values below are median `plan()` times on the same JVM. Data reads were not timed. | Value distribution | Bounded lookup | Existing path | Planned candidate rows | | --- | ---: | ---: | ---: | | Alternating dog/cat (50,000 matches) | 2.34 ms | 3.40 ms | 10 vs 50,000 | | All dog (100,000 matches) | 1.43 ms | 1.54 ms | 10 vs 100,000 | ## Verification - `mvn -o -q -pl paimon-core -am -Pfast-build -DfailIfNoTests=false -DwildcardSuites=none -Dtest=BtreeGlobalIndexTableTest,BTreeTopNIndexFileSelectorTest,BTreeIndexReaderTest test` (148 tests passed) - `mvn -o -q -pl paimon-core -am -DskipTests compile` - `mvn -o -q -pl paimon-common,paimon-core spotless:check` - `git diff --check` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
