sundapeng opened a new pull request, #1058: URL: https://github.com/apache/paimon-rust/pull/1058
### Purpose The integer `In` / `NotIn` fast path in the residual filter builds a SipHash `HashSet` of the literals on every evaluation and collects the mask row by row through `Option<bool>`. Predicates such as `(k = 1 AND step IN (0, 100, ..., 182)) OR ...` over 16 keys evaluate many such lists against every decoded batch, and the hashing and per-row collection become a large part of the read. ### Brief change log - Sort and deduplicate the converted literals once per call; test membership with a bitmap when the literal span is below 65,536, otherwise with binary search. - Build the mask with `BooleanBuffer::collect_bool` instead of collecting `Option<bool>`. - Semantics are unchanged: null rows are `false` for both `In` and `NotIn`, and a literal that does not fit the column type still falls back to the general path. ### Tests - New unit test covering the bitmap and binary-search paths with nulls, negative values and duplicate literals, for both `In` and `NotIn`. - `cargo test -p paimon --lib -- residual`, `cargo fmt --all -- --check`, `cargo clippy --locked --all-targets --workspace --features fulltext,vortex -- -D warnings`. - Benchmark: in a DataLoader workload whose lookup is `(episode_index = E AND step_index IN (84 values)) OR ...` over 16 episodes, the lookup's read CPU per batch fell from 0.46 to 0.35 core-seconds (-24%) with identical output. ### API and Format No. ### Documentation No. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
