lucasfang opened a new issue, #307:
URL: https://github.com/apache/paimon-cpp/issues/307

   
   Labels: enhancement
   
   ## Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ## Motivation
   
   Evaluating a comparison predicate (`=`, `<>`, `<`, `<=`, `>`, `>=`) on a 
batch goes through `NullFalseLeafBinaryFunction::Test`, which materializes the 
whole column into `Literal` objects — one heap allocation and one value hash 
per row — and then compares each row against the literal one at a time. That is 
`O(rows)` allocations plus `O(rows)` scalar comparisons per batch, even though 
a comparison only needs one vectorized pass over the column. On a wide scan 
with a selective filter this per-row `Literal` churn is a measurable part of 
predicate evaluation. This is the sibling of the `IN` / `NOT IN` optimization 
that already probes a batch with `arrow::compute::is_in`; the comparison 
functions are the other `LeafFunction` family still on the row-by-row path.
   
   ## Solution
   
   Compare the whole batch against the literal with one `arrow::compute` 
comparison kernel (`equal`, `not_equal`, `less`, `less_equal`, `greater`, 
`greater_equal`) whenever the literal and the column allow it, which turns 
evaluation into `O(1)` setup — build one scalar per batch — plus an `O(rows)` 
vectorized compare with no `Literal` per row. Every `LeafFunction` is a shared 
stateless singleton, so the scalar cannot be cached on the function and is 
built per batch.
   
   - Move `NullFalseLeafBinaryFunction::Test(array, literals, pool)` out of the 
header into a new `null_false_leaf_binary_function.cpp`.
   - Build the comparison scalar with the existing 
`LiteralConverter::ConvertLiteralsToArray`, the same call `IN` writes its value 
set with, and probe the column with the kernel that matches the function type.
   - Take the kernel path only where a kernel agrees with `Literal::CompareTo`: 
`BOOLEAN`, `TINYINT`, `SMALLINT`, `INT`, `BIGINT`, `DATE`, `STRING`, `BINARY`, 
and `DECIMAL` / `TIMESTAMP` whose column carries the scale / the unit of the 
literal and no time zone. `FLOAT` and `DOUBLE` keep the row-by-row path because 
`FieldsComparator::CompareFloatingPoint` orders `-0.0 < +0.0` and makes every 
NaN equal to every NaN, where an IEEE-754 kernel says `-0.0 == +0.0` and that 
no NaN compares to anything.
   - Gate the kernel path on the layouts the row-by-row path already accepts, 
so which columns evaluate and which report an error stays exactly the same: the 
string pattern functions (`STARTS_WITH`, `ENDS_WITH`, `CONTAINS`, `LIKE`), a 
decimal column of another scale, a timestamp column of another unit or with a 
time zone, a literal whose field type disagrees with the column, and the 
layouts `ConvertLiteralsFromArray` rejects all keep the row-by-row path and its 
error reporting.
   - As a side effect the kernel decodes a dictionary column, so a row that 
points at a null dictionary value is now false for every comparison instead of 
reading the empty value the slot holds, which an empty literal used to match.
   
   ## Anything else?
   
   No public API, storage format, or protocol change; the work is internal to 
`src/paimon/common/predicate/`. The statistics and min/max pruning path keeps 
the existing per-`Literal` comparison, since it compares a single min/max pair 
rather than a batch.
   
   ## Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to