yongster opened a new pull request, #11245:
URL: https://github.com/apache/arrow-rs/pull/11245
# Which issue does this PR close?
- Closes #11243.
# Rationale for this change
`Float16 -> Float32` and `Float32 -> Float16` use the generic numeric cast.
The safe path calls `unary_opt(num_cast)`, so every element builds an `Option`
and the validity bitmap is rebuilt. These two conversions cannot fail: every
`f16` widens exactly, and every `f32` rounds to an `f16`.
That path is what `FixedSizeList<Float16, 768/1536>` hits when embedding
values are widened or narrowed. `half::slice::HalfFloatSliceExt` already
converts a whole slice, and on targets with `fp16` or `f16c` it uses the
hardware conversion. This is separate from #10955, which does not cover
`half::f16`.
The new embedding benchmarks are in a separate PR so they can land on `main`
first. This PR does not include them.
# What changes are included in this PR?
- `cast_float16_to_float32` and `cast_float32_to_float16` fill the output
values with `convert_to_f32_slice` / `convert_from_f32_slice` and clone the
input null buffer.
- The two dispatch arms point at those helpers. Every other numeric cast is
unchanged.
- The output buffer is a normal zeroed `Vec`. No `unsafe`.
- Tests for all 65,536 `f16` bit patterns, 1,000,000 sampled `f32` patterns,
nulls, slices, empty arrays, `safe: false`, and a nullable `FixedSizeList`
child.
Null slots are converted along with the valid values, then the original
validity is reused. Arrow equality ignores bytes in null slots. `CastOptions {
safe: false }` still succeeds and matches the safe result, because the
conversion never fails.
# Are these changes tested?
Yes. `cargo test -p arrow-cast --test cast float16` (5 tests) and the
existing `test_cast_from_f32` / `test_cast_from_f64` pass. `cargo clippy -p
arrow-cast --all-targets --no-deps -- -D warnings` passes.
Local remeasurement of `cast()` itself, Apple Silicon
(`target_feature=fp16`), rustc 1.97, release + thin LTO, median of 51 × 20
calls, `FixedSizeList` of 1024 rows. Baseline is `fa337f843` before this patch.
After the patch, `cast()` matches a standalone slice prototype (about 1.00×).
| Shape | Direction | Before | After | Speedup |
|---|---|---:|---:|---:|
| 1024 × 768 | f16 → f32 | 218 µs | 65 µs | 3.33× |
| 1024 × 768 | f32 → f16 | 206 µs | 57 µs | 3.59× |
| 1024 × 1536 | f16 → f32 | 421 µs | 126 µs | 3.34× |
| 1024 × 1536 | f32 → f16 | 410 µs | 113 µs | 3.63× |
| 1024 × 768, 10% null | f16 → f32 | 366 µs | 66 µs | 5.59× |
| 1024 × 768, 10% null | f32 → f16 | 354 µs | 56 µs | 6.31× |
Bit identity against scalar `f16::to_f32` / `f16::from_f32` held for every
`f16` pattern and for the 1,000,000 `f32` sample, which is the same conversion
the previous `num_cast` path used.
# Are there any user-facing changes?
No API change. Logical values and validity are unchanged. Physical bytes in
null slots may differ, because those slots are now converted instead of left as
zero. They are not part of the array's logical value.
AI assistance: the helpers and tests were drafted with AI, then reviewed and
remeasured locally against scalar `half` conversion and the previous `cast`
results.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]