sundapeng opened a new pull request, #1058:
URL: https://github.com/apache/paimon-rust/pull/1058

   ### Purpose
   
   The integer `In` / `NotIn` fast path in the residual filter builds a SipHash 
`HashSet` of the literals on every evaluation and collects the mask row by row 
through `Option<bool>`. Predicates such as `(k = 1 AND step IN (0, 100, ..., 
182)) OR ...` over 16 keys evaluate many such lists against every decoded 
batch, and the hashing and per-row collection become a large part of the read.
   
   ### Brief change log
   
   - Sort and deduplicate the converted literals once per call; test membership 
with a bitmap when the literal span is below 65,536, otherwise with binary 
search.
   - Build the mask with `BooleanBuffer::collect_bool` instead of collecting 
`Option<bool>`.
   - Semantics are unchanged: null rows are `false` for both `In` and `NotIn`, 
and a literal that does not fit the column type still falls back to the general 
path.
   
   ### Tests
   
   - New unit test covering the bitmap and binary-search paths with nulls, 
negative values and duplicate literals, for both `In` and `NotIn`.
   - `cargo test -p paimon --lib -- residual`, `cargo fmt --all -- --check`, 
`cargo clippy --locked --all-targets --workspace --features fulltext,vortex -- 
-D warnings`.
   - Benchmark: in a DataLoader workload whose lookup is `(episode_index = E 
AND step_index IN (84 values)) OR ...` over 16 episodes, the lookup's read CPU 
per batch fell from 0.46 to 0.35 core-seconds (-24%) with identical output.
   
   ### API and Format
   
   No.
   
   ### Documentation
   
   No.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to