geoffreyclaude opened a new pull request, #25912: URL: https://github.com/apache/datafusion/pull/25912
Stacked on #25186. Review the [optimization-only diff](https://github.com/apache/datafusion/compare/ba584cd25592c7868402ef48e0b3646bb143e6ea..e1bd187acd4b57cb291a4f66f2767daf045d80af). This is a draft until the parent merges; the GitHub diff against `main` currently also includes the parent's functional changes. ## Which issue does this PR close? None. ## Rationale for this change Dynamic `IN` evaluation over dictionary-encoded floats can spend unnecessary time revalidating unchanged dictionary keys during signed-zero normalization. The existing generic `has_float_leaf` path already normalizes dictionary arrays correctly; this PR specializes that path to reduce its overhead while preserving results. ## What changes are included in this PR? - Reuses dictionary keys when normalized values change, avoiding dictionary `ArrayData` reconstruction and key validation. - Drops redundant all-valid bitmaps from rewritten values so comparisons can reuse key validity without scanning every key for value nulls. - Adds a focused four-case benchmark in a separate commit before the optimization. ## What is the testing strategy for this PR? The existing dictionary regression checks normalized zeros, preserved keys and NaN payloads, scalar consistency, and idempotence. It now also checks that rewriting all-valid values removes redundant validity and lets logical nulls reuse the keys' bitmap. The benchmark verifies expected results outside the timed evaluation. Benchmark commit f48613118b uses the generic normalization path. The immediately following commit, e1bd187acd, introduces the complete dictionary optimization, including redundant-validity removal. The four Float32 cases cross negative-zero/no-op inputs with all-valid/absent values bitmaps (8,192 rows, 16 dictionary values). The [adriangbot benchmark results](https://github.com/apache/datafusion/pull/25186#issuecomment-5915237922) compare those commits on GKE (`c4a-standard-32`, Arm Neoverse-V2). Physical-expression evaluation times, in µs, as reported by the bot: | Contains `-0.0` | Values bitmap | Generic | Optimized | | --- | --- | ---: | ---: | | No | Absent | 23.0 ± 0.02 | 22.0 ± 0.01 | | No | All-valid | 45.2 ± 0.04 | 44.3 ± 0.03 | | Yes | Absent | 28.9 ± 0.02 | 22.2 ± 0.01 | | Yes | All-valid | 29.0 ± 0.02 | 22.4 ± 0.02 | Rewriting cases took about 23% less time (approximately 1.30× speedup); no-rewrite cases took about 2–4% less time. These measurements compare the complete specialization against the generic path, directly isolating the performance change. To compare the two commits with Criterion: ```bash git switch --detach f48613118b cargo bench --profile release-nonlto -p datafusion-physical-expr --bench dictionary_float_zero -- --save-baseline before git switch --detach e1bd187acd cargo bench --profile release-nonlto -p datafusion-physical-expr --bench dictionary_float_zero -- --baseline before ``` ## Are there any user-facing changes? Dictionary float normalization and dependent expression evaluation become faster. SQL results, null semantics, and public APIs are unchanged. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
