neilconway opened a new pull request, #26114: URL: https://github.com/apache/datafusion/pull/26114
## Which issue does this PR close? - N/A ## Rationale for this change Various cleanups and improvements for these string benchmarks: * The benchmarks all used fixed-length strings, which is atypical for real-world workloads. Fixed-length inputs make life much easier for the branch predictor, and that can have a very significant impact when microbenchmarking and tuning these functions. * Add/remove benchmark cases to improve coverage * The return type of several UDF invocations was incorrect (e.g., Utf8 vs Utf8View). * Make benchmark coverage consistent between Utf8 and Utf8View * Use 8k batch size consistently * Refactoring and code cleanup This PR also adjusts the `substr_index` benchmarks to use an 8k batch size for consistency, but it didn't suffer from most of the other issues described above. ## What changes are included in this PR? See above. ## What is the testing strategy for this PR? Benchmark changes only. ## Are there any user-facing changes? No. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
