JingsongLi opened a new pull request, #888:
URL: https://github.com/apache/paimon-rust/pull/888

   ### Purpose
   
   Improve the Data Evolution scalar-index query-to-read path, especially when 
predicates produce large posting lists or many disjoint row ranges, and add 
Java-compatible BTree V2 postings without changing the default write format.
   
   Related Rust tracking issue: #227 (extends the BTree capability; does not 
close the broader multimodal tracking issue).
   
   Related Java work: apache/paimon#10007, apache/paimon#9957, 
apache/paimon#9844 and apache/paimon#9834. The changes adapt their applicable 
optimizations to Rust rather than introducing a new index-versus-scan cost 
model.
   
   ### Brief change log
   
   | Java optimization | Rust implementation or existing equivalent |
   |---|---|
   | apache/paimon#10007: large-result transport and selected-row traversal | 
Use native bitmap range iteration and shift range endpoints instead of 
reinserting every RowID into another bitmap. Seek directly to overlapping 
selection ranges for `.row` blocks, avoiding repeated prefix traversal. Parquet 
already uses native RowSelection. |
   | apache/paimon#9957: avoid redundant range sorting | Rust already skips 
sorting canonical ranges; use the trusted sorted path for intersections to 
avoid redundant validation too. Still normalize unordered shard results. |
   | apache/paimon#9844: bounded same-column queries | Tighten all 
lower/upper/BETWEEN bounds into one interval, including strictness ties. 
Contradictory bounds return no indexed candidates without opening index files. 
Preserve indexed coverage for unindexed-tail handling. |
   | apache/paimon#9834: metadata all-match | Prove complete, null-free source 
domains from all same-domain key shards and their row counts. Apply the proof 
per conjunct, so a wide all-matching bound does not trigger a posting scan 
alongside a selective equality. Proven conjuncts do not consume the existing 
fallback index-scan byte budget. |
   
   Also add BTree V2 SINGLE, DELTA and portable ROARING posting read/write 
support, adaptive encoding, and malformed-input validation. Include fixtures 
generated by the Java production writer (uncompressed and actual LZ4-compressed 
blocks), plus byte-for-byte uncompressed writer interoperability checks.
   
   No new RowID-count budget, row-range memory option or blanket scan fallback 
is introduced. Residual predicates and FAST/FULL/DETAIL coverage behavior 
remain intact; a declined sibling predicate must not turn retained candidates 
into an empty result. Bitmap floating range candidates stay conservative, and 
bound proofs follow residual filtering's bit-preserving floating-point order.
   
   #### End-to-end local performance
   
   Compared with unmodified Rust `f33dee15`, with only the same benchmark 
fixture added to the baseline:
   
   - 1,200,000 synthetic rows, default Parquet settings, Data Evolution, row 
tracking and FAST search, memory FileIO.
   - Physical column groups: `[id, name]` and `[payload]` in separate Parquet 
files covering the same RowID domain. The fixture uses the production 
partial-column writer/commit path and checks metadata and `payload = id * 17` 
after reading.
   - Three warmups, seven measured rounds, median milliseconds. Index-on/off 
alternates against the same snapshot within each run; baseline/optimized 
binaries run separately on identical synthetic contents.
   - Includes planning, complete Arrow reads, collection, ID sorting and 
payload-alignment checks. Excludes table/index creation; exact expected-ID 
comparison is outside the timer.
   
   | Case | Upstream V1 | Optimized V1 | Optimized V2 | No index (optimized V1 
run) |
   |---|---:|---:|---:|---:|
   | 90% match | 59.432 | 44.949 | 37.324 | 28.800 |
   | 10%, 120,000 discrete ranges | 29.531 | 27.614 | 27.077 | 20.473 |
   | Point, one row | 1.581 | 1.541 | 1.534 | 17.807 |
   | Point AND all-matching same-field bound | 17.998 | 1.543 | 1.512 | 18.134 |
   | Multiple bounds, final 99 rows | 32.409 | 1.647 | 1.637 | 18.838 |
   | Point AND 90% other-field condition | 23.807 | 8.559 | 2.459 | 22.270 |
   | All match | 47.830 | 24.554 | 24.562 | 24.530 |
   | Single bounded interval, 99 rows | 1.655 | 1.661 | 1.643 | 18.200 |
   
   The corresponding V1 reductions are approximately 24% for 90% matches, 64% 
for point-AND-wide, 49% for all-match, 91% for the redundant same-field bound 
and 95% for multiple bounds. Point and single bounded lookups are effectively 
unchanged. V2 helps dense postings more than sparse ones.
   
   **These local Table API results are not production/Spark SQL speedups. Broad 
indexed results still lose to no-index scanning in this workload.** Remote 
storage latency, cold caches and peak RSS were not measured.
   
   #### Focused microbenchmarks
   
   - `.row` selection only, 1,172 blocks and 120,000 disjoint ranges: 90% 
selection **74.920 -> 1.515 ms**, 10% selection **73.435 -> 0.609 ms**; one 
continuous range **0.711 -> 0.687 ms**. This excludes file decoding and is not 
a Parquet/SQL speedup.
   - RowID transport: 1.08 million matches in 120,000 ranges **13.745 -> 0.652 
ms**; 120,000 singletons **2.160 -> 0.663 ms**. Twelve million continuous 
RowIDs no longer require per-row traversal (**163.944 -> 0.002 ms**, below 
useful query-level timing precision afterward).
   - Dense 1.08 million-row posting query: V1 **6.266 ms**, V2 **0.028 ms**. 
Sparse 120,000-row posting: **1.324 / 1.268 ms**. Uncompressed fixture sizes: 
**3,583,659 / 275,940 bytes**.
   
   ### Tests
   
   Verified locally on macOS ARM64 with Rust 1.95, release builds and the 
existing unchanged Cargo.lock:
   
   | Command | Result |
   |---|---|
   | `cargo test --offline --release -p paimon --lib` | 2947 passed, 6 ignored |
   | `cargo test --offline --release -p paimon --lib --features fulltext` | 
3023 passed, 6 ignored |
   | `cargo test --offline --release -p paimon-datafusion --test procedures 
btree` | 3 passed; V1/V2 SQL cases include redundant bounds, multiple bounds, 
contradictory ranges and OR siblings |
   | `cargo test --offline --release -p paimon --test scan_planning_parity_test 
--test pk_vector_batch_test` | 15 + 10 passed |
   | `cargo clippy --offline --release -p paimon --lib --tests -- -D warnings` 
| Passed |
   | `cargo fmt --all -- --check` and `git diff --check` | Passed |
   
   Regression coverage includes exact query counts, empty-range zero I/O, 12 
key types across V1/V2, null/incomplete/key-sharded metadata, indexed coverage 
and fallback, negative/positive NaN bounds, high-32-bit range boundaries, 
unordered/overlapping `.row` selections, and column-group key/payload 
alignment. The redundant-bound and empty-range tests were confirmed failing 
before the fix. There are no timing thresholds in correctness tests.
   
   Run the four ignored manual benchmarks with:
   
   ```sh
   cargo test --offline --release -p paimon --lib benchmark_ -- --ignored 
--nocapture --test-threads=1
   ```
   
   The full workspace/all-feature CI matrix, Vortex, MSRV 1.94 and production 
engine/storage workloads were not run locally.
   
   ### API and Format
   
   - V1 remains the default BTree write format; existing V1 reads remain 
supported.
   - `btree-index.file-version=2` explicitly enables V2 writing through 
table/build options; the writer also exposes a version selector.
   - The reader supports mixed V1/V2 files. Upgrade every reader before 
enabling V2; existing files do not need rebuilding merely to read mixed 
versions.
   - No public row-range representation replacement or search-mode semantic 
change.
   
   ### Documentation
   
   Update SQL documentation for the version option, compatibility requirements 
and per-conjunct all-match/bounded-query behavior. Document fixture contents 
and regeneration in `crates/paimon/testdata/btree/README.md`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to