mightsleep commented on code in PR #11271:
URL: https://github.com/apache/arrow-rs/pull/11271#discussion_r4137291274
##########
arrow-select/benches/filter_bits.rs:
##########
@@ -86,5 +111,44 @@ fn add_benchmark(c: &mut Criterion) {
}
}
-criterion_group!(benches, add_benchmark);
+/// `filter_bits` on masks too large for the branch predictor to learn: the
+/// cases above repeat one 1024-word mask, whose per-word branches a recent
+/// core learns, which makes random masks look faster than a real filter.
+/// Lazy strategies only, which compress word by word
+fn add_large_benchmark(c: &mut Criterion) {
+ const SIZE: usize = 1 << 22;
Review Comment:
Quick update: 100 batches of 8K rows keep the branch predictor out on
Neoverse-N2, but Zen 5 learns masks that repeat every iteration up to at least
16K words (100 batches are 12.8K). The bot's Neoverse-V2 may well learn them
too. So I'd keep 8K-row batches, as in coalesce_kernels, but make 512 of them
(64K words), past what Zen 5 learns:
| Kept 1/64, 8K-row batches, main, Zen 5 | branch misses per word | ns per
word |
|---|---|---|
| 16K words of masks per iteration | 0.005 | 1.9 |
| 64K words of masks per iteration | 0.74 | 6.0 |
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]