HippoBaro commented on PR #11267: URL: https://github.com/apache/arrow-rs/pull/11267#issuecomment-5938203615
For context, the 1%-cardinality cases have been on this branch for quite a while; the fixed-cardinality 20/100/400 cases were added later. > is the new string_dictionary_1pct markedly different from the other low cardinality tests? They all seem to have similar execution times. The main difference is that the 1%-cardinality cases include 25% nulls, whereas the fixed-cardinality cases contain none. My intention was consistency with the existing benchmarks: general-purpose cases such as primitive, bool, string, and list_primitive use 25% null density, with `_non_null` variants explicitly covering non-null data. The 1%-cardinality cases also form a consistent set of low-cardinality counterparts to the high-cardinality dictionary benchmarks across `Int32`, `Int64`, `Float64`, strings, and `Decimal128`, rather than being string-specific cases. Perhaps we should append `_non_null` to the fixed-cardinality string benchmark names to make that distinction clearer and align them with that naming convention? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
