devanbenz commented on PR #10136:
URL: https://github.com/apache/arrow-rs/pull/10136#issuecomment-5822684602

   > > What do you think about having the fallback walk whichever set of bits 
is smaller? The current loop runs once per set bit in mask, so at 9/10 density 
it takes about 58 iterations per word. Each iteration is 9 instructions on 
aarch64:
   > 
   > In my opinion, we should merge this PR and then improve the fallback code 
as a follow on issue / PR
   > 
   > It seems like this PR is already faster even with the somewhat basic 
scalar fallback. I am sure we can all then geek out trying to improve the 
performance of the fallback with more crazy bithacks (this is a good thing)
   > 
   > I looked over comments, and it looks to me like these are the only 
remaining outstanding comments about adding some additional asserts
   > 
   >     * [feat(perf): Improve filter performance with per word bit filtering 
(and `BMI` when supported) #10136 
(comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4087289204)
 from me and @mbutrovich
   > 
   >     * [feat(perf): Improve filter performance with per word bit filtering 
(and `BMI` when supported) #10136 
(comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4094224975)
 from @mbutrovich
   > 
   > 
   > For fun, I will re-run my benchmark run with the latest fixes
   
   I've gone ahead and added the suggested `debug_asserts`. I've also added a 
new benchmark for FSB which includes nulls to stress the code path in 
`filter_bits`. 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to