lucasfang opened a new issue, #307: URL: https://github.com/apache/paimon-cpp/issues/307
Labels: enhancement ## Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ## Motivation Evaluating a comparison predicate (`=`, `<>`, `<`, `<=`, `>`, `>=`) on a batch goes through `NullFalseLeafBinaryFunction::Test`, which materializes the whole column into `Literal` objects — one heap allocation and one value hash per row — and then compares each row against the literal one at a time. That is `O(rows)` allocations plus `O(rows)` scalar comparisons per batch, even though a comparison only needs one vectorized pass over the column. On a wide scan with a selective filter this per-row `Literal` churn is a measurable part of predicate evaluation. This is the sibling of the `IN` / `NOT IN` optimization that already probes a batch with `arrow::compute::is_in`; the comparison functions are the other `LeafFunction` family still on the row-by-row path. ## Solution Compare the whole batch against the literal with one `arrow::compute` comparison kernel (`equal`, `not_equal`, `less`, `less_equal`, `greater`, `greater_equal`) whenever the literal and the column allow it, which turns evaluation into `O(1)` setup — build one scalar per batch — plus an `O(rows)` vectorized compare with no `Literal` per row. Every `LeafFunction` is a shared stateless singleton, so the scalar cannot be cached on the function and is built per batch. - Move `NullFalseLeafBinaryFunction::Test(array, literals, pool)` out of the header into a new `null_false_leaf_binary_function.cpp`. - Build the comparison scalar with the existing `LiteralConverter::ConvertLiteralsToArray`, the same call `IN` writes its value set with, and probe the column with the kernel that matches the function type. - Take the kernel path only where a kernel agrees with `Literal::CompareTo`: `BOOLEAN`, `TINYINT`, `SMALLINT`, `INT`, `BIGINT`, `DATE`, `STRING`, `BINARY`, and `DECIMAL` / `TIMESTAMP` whose column carries the scale / the unit of the literal and no time zone. `FLOAT` and `DOUBLE` keep the row-by-row path because `FieldsComparator::CompareFloatingPoint` orders `-0.0 < +0.0` and makes every NaN equal to every NaN, where an IEEE-754 kernel says `-0.0 == +0.0` and that no NaN compares to anything. - Gate the kernel path on the layouts the row-by-row path already accepts, so which columns evaluate and which report an error stays exactly the same: the string pattern functions (`STARTS_WITH`, `ENDS_WITH`, `CONTAINS`, `LIKE`), a decimal column of another scale, a timestamp column of another unit or with a time zone, a literal whose field type disagrees with the column, and the layouts `ConvertLiteralsFromArray` rejects all keep the row-by-row path and its error reporting. - As a side effect the kernel decodes a dictionary column, so a row that points at a null dictionary value is now false for every comparison instead of reading the empty value the slot holds, which an empty literal used to match. ## Anything else? No public API, storage format, or protocol change; the work is internal to `src/paimon/common/predicate/`. The statistics and min/max pruning path keeps the existing per-`Literal` comparison, since it compares a single min/max pair rather than a batch. ## Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
