neilconway opened a new pull request, #11227:
URL: https://github.com/apache/arrow-rs/pull/11227
# Which issue does this PR close?
- N/A
# Rationale for this change
`unary_opt` is much slower than `unary`, for two reasons:
1. It guarantees that the closure will only be invoked on valid (non-null)
array elements
2. It records whether each closure succeeded in a bitmap which would need to
be shared across SIMD lanes
In Property 2, this is slow because the read-modify-write of a shared value
prevents vectorization. This is relatively easy to avoid: instead of recording
closure success in a single bitmap, instead use a temporary array of 64 bytes
to record the success of 64 closure invocations. We can then combine the values
in that array into a `u64` whose bits represent the success of the individual
closures in this chunk, and then the result validity bitmap is easy to assemble
from those `u64` chunks.
We don't use this strategy when the input array has nulls (since checking
for nulls introduce additional data dependencies that prevent vectorization
anyway) or when the array element size is larger than 8 bytes.
Benchmarks:
- null_if_overflow_precision_128: 9.471 → 9.489 µs, +0.2%
- null_if_overflow_precision_256: 12.111 → 12.105 µs, −0.0%
- primitive_array_unary/unary_opt/no_input_nulls/1024: 0.446 → 0.182 µs,
−59.3%
- primitive_array_unary/unary_opt/no_input_nulls/8192: 3.143 → 1.308 µs,
−58.4%
- primitive_array_unary/unary_opt/no_input_nulls/65536: 24.909 → 9.982 µs,
−59.9%
- primitive_array_unary/unary_opt/20pct_input_nulls/1024: 0.639 → 0.636
µs, −0.5%
- primitive_array_unary/unary_opt/20pct_input_nulls/8192: 4.698 → 4.695
µs, −0.1%
- primitive_array_unary/unary_opt/20pct_input_nulls/65536: 37.480 → 37.327
µs, −0.4%
- primitive_array_unary/unary_opt_narrow/no_input_nulls/1024: 0.570 →
0.230 µs, −59.7%
- primitive_array_unary/unary_opt_narrow/no_input_nulls/8192: 4.174 →
1.506 µs, −63.9%
- primitive_array_unary/unary_opt_narrow/no_input_nulls/65536: 36.901 →
11.945 µs, −67.6%
- primitive_array_unary/unary_opt_narrow/20pct_input_nulls/1024: 0.684 →
0.681 µs, −0.4%
- primitive_array_unary/unary_opt_narrow/20pct_input_nulls/8192: 4.995 →
4.998 µs, +0.1%
- primitive_array_unary/unary_opt_narrow/20pct_input_nulls/65536: 42.103 →
42.267 µs, +0.4%
To validate that this change improves the end-to-end performance of
call-sites of `unary_opt`, I added benchmarks for the `cast` kernel on
null-free arrays:
- cast int64 to int32 no_nulls 512: 277.900 → 160.810 ns, −42.1%
- cast float64 to int32 no_nulls 512: 282.070 → 208.000 ns, −26.3%
- cast date64 to date32 no_nulls 512: 285.100 → 224.840 ns, −21.1%
# What changes are included in this PR?
* Optimize `unary_opt` as described above
* Add null-free benchmark cases for cast kernels
* Extend `unary_opt` benchmark to range over batch sizes (1k and 8k rows,
not just 64k)
# Are these changes tested?
Yes; existing tests pass, new test added.
# Are there any user-facing changes?
No.
# AI usage
Developed with Claude Code (Fable 5.1), reviewed with Codex (Astra 6). I
reviewed, revised, and understand the resulting code.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]