FrankChen021 opened a new pull request, #20390:
URL: https://github.com/apache/druid/pull/20390
### Description
This PR improves native expression vector processing by reducing null-vector
propagation and allocation overhead in common numeric expression chains.
The changes:
- return `null` null-vectors from scalar and SIMD unary/binary processor
bases when the output batch contains no nulls;
- record scalar null presence only when the existing null branch is taken,
avoiding a loop-carried accumulation on every row;
- reuse per-processor buffers for numeric long/double casts instead of
allocating a converted array for every evaluation;
- replace stream-based fallback numeric conversions with simple loops;
- use dedicated bound-constant processors with direct
add/subtract/multiply/divide loops for same-type long and double arithmetic
when the JDK Vector API path is not selected;
- add correctness tests and a JMH benchmark covering long chains, mixed
long/double chains, constants, nullability, vector sizes 128–2048, and Vector
API on/off.
SIMD selection continues to take precedence when the JDK Vector API is
enabled. There are no configuration or expression-semantics changes.
### Benchmark
The benchmark compares this branch against exact base commit
`b2b9baedca090ac70e7293a239904a6d99fa6254` using the same JMH harness and JVM.
Environment: arm64, Temurin OpenJDK 25. Results below use vector size 512, five
warmup iterations, eight measurement iterations, three forks, and 300 ms per
iteration.
Lower `ns/op` is better. Speedup is calculated as `baseline / optimized -
1`. The null-vector shapes are: **Absent**, where bindings return `null`; **All
false**, where bindings return a non-null `boolean[]` with no null rows; and
**Sparse**, where every sixteenth row is null. Each pair of rows is the same
expression and null-vector shape with the Vector API disabled and enabled.
| Workload | Expression | Null vector | Vector API | Master ns/op |
Optimized ns/op | Speedup |
|---|---|---:|---:|---:|---:|---:|
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Absent | Off | 8,873 |
7,052 | **25.8%** |
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Absent | On | 1,813 |
1,737 | **4.4%** |
| | | | | | | |
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | All false | Off | 9,797 |
9,274 | **5.6%** |
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | All false | On | 2,435 |
2,213 | **10.0%** |
| | | | | | | |
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Sparse | Off | 10,308 |
10,414 | **-1.0%** |
| Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Sparse | On | 2,418 |
2,366 | **2.2%** |
| | | | | | | |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Absent | Off | 1,336 | 948 |
**41.0%** |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Absent | On | 1,033 | 935 |
**10.5%** |
| | | | | | | |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | All false | Off | 6,495 |
1,404 | **362.6%** |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | All false | On | 1,420 |
1,336 | **6.3%** |
| | | | | | | |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Sparse | Off | 6,883 | 2,722
| **152.9%** |
| Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Sparse | On | 1,419 | 1,393 |
**1.9%** |
| | | | | | | |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | Absent | Off | 5,171 | 496 |
**943.6%** |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | Absent | On | 745 | 710 |
**4.8%** |
| | | | | | | |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | All false | Off | 5,938 | 621 |
**856.3%** |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | All false | On | 776 | 791 |
**-2.0%** |
| | | | | | | |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | Sparse | Off | 5,711 | 873 |
**553.9%** |
| Constant-heavy | `(((l1 + 7) * 3) - 11)` | Sparse | On | 768 | 765 |
**0.4%** |
#### Constant-bound arithmetic
`LongBivariateLongsConstantProcessor` and
`DoubleBivariateDoublesConstantProcessor` specialize same-type arithmetic with
one bound constant. The general bivariate processor materializes the constant
as a second input vector and invokes an operation function for every row. The
specialized processors retain the constant as a scalar, choose the operation
and operand order once per batch, and execute direct primitive arithmetic
loops. The double specialization exists for the same reason as the long
version: it removes the redundant constant vector and keeps the hot
double-arithmetic loop monomorphic and readily inlineable by HotSpot.
For the constant-heavy expression with absent null vectors, the optimized
non-Vector-API path is faster than the Vector API path (496 versus 710 ns/op).
The scalar specialization reads only the varying input array and applies its
scalar constants directly. SIMD selection still takes precedence when the
Vector API is enabled, so that path uses the general bivariate processors,
materializes each constant operand as a vector, and loads both operand arrays
for every operation. At a 512-row batch and for this short arithmetic chain,
that extra materialization and load/store overhead outweighs SIMD throughput.
This is specific to the constant-heavy case: the longer all-variable chain
remains substantially faster with the Vector API enabled.
Two of the 18 baseline-to-optimized comparisons are small regressions: the
sparse scalar long chain at -1.0%, and the all-false SIMD constant chain at
-2.0%. The direct constant processors are not present in the first case, while
the second case's 99.9% confidence intervals overlap the baseline. These small
results are sensitive to run-to-run JVM/code-layout variation; neither offsets
the gains in the corresponding expression family or null-vector shape.
Geometric-mean speedups at vector size 512:
- long chains: **7.5%**;
- mixed numeric chains: **64.4%**;
- constant-heavy chains: **196.0%**;
- all 18 configurations: **73.6%**.
### Testing
```text
mvn -ntp test -pl processing \
-Dtest='org.apache.druid.math.expr.VectorExprResultConsistencyTest,org.apache.druid.math.expr.VectorExprResultConsistencyVectorApiTest'
\
-Dsurefire.failIfNoSpecifiedTests=false \
-Pskip-static-checks -Dweb.console.skip=true -T1C
```
Result: 69 tests run, 0 failures, 0 errors, 0 skipped.
The benchmark module also packages successfully with the new JMH benchmark.
#### Release note
Native vectorized numeric expressions reduce intermediate null-vector and
numeric-conversion allocation overhead, improving representative
expression-chain performance without changing query behavior or configuration.
<hr>
##### Key changed/added classes in this PR
* `CastToDoubleVectorProcessor`
* `CastToLongVectorProcessor`
* `LongBivariateLongsConstantProcessor`
* `DoubleBivariateDoublesConstantProcessor`
* `SimdProcessorUtils`
* `ExpressionVectorProcessorBenchmark`
<hr>
This PR has:
- [x] been self-reviewed.
- [x] a release note entry in the PR description.
- [x] added Javadocs for the new processor classes.
- [x] added unit tests covering the new code paths.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]