devanbenz commented on code in PR #10136: URL: https://github.com/apache/arrow-rs/pull/10136#discussion_r4085673619
########## arrow-buffer/src/util/bit_util.rs: ########## @@ -19,6 +19,45 @@ use crate::bit_chunk_iterator::BitChunks; +/// Parallel bit extract: for each set bit in `mask`, extract the +/// corresponding bit from `value` and pack them contiguously into the low +/// bits of the return value. +/// +/// Equivalent to the x86 BMI2 `PEXT` instruction. When compiled with the +/// `bmi2` target feature enabled (for example `-C target-cpu=x86-64-v3`) +/// this lowers to the hardware `pext` instruction; otherwise it falls back +/// to a portable scalar loop. +// +// Replace with `value.compress(mask)` when `uint_gather_scatter_bits` is +// stabilised: <https://github.com/rust-lang/rust/issues/149069> +#[inline] +pub fn compress(value: u64, mask: u64) -> u64 { + #[cfg(all(target_arch = "x86_64", target_feature = "bmi2"))] Review Comment: > Have you measured this on AMD Zen 1 or Zen 2? No, we will likely need a machine that has Zen 1 or 2. I don't have access to one without deploying some sort of cloud virtual machine. > I think you have a Zen 1 or Zen 2 machine. Could you run cargo bench -p arrow-select --bench filter_bits at this PR's head with and without RUSTFLAGS="-C target-cpu=native" and post the numbers here? It would be amazing if Andy has the capacity and time to run this on his given he has one :saluting_face: Thanks for checking with him. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
