HippoBaro commented on code in PR #11267:
URL: https://github.com/apache/arrow-rs/pull/11267#discussion_r4170070339
##########
parquet/Cargo.toml:
##########
@@ -94,6 +94,9 @@ object_store = { workspace = true, features = ["azure", "fs"]
}
opendal = { version = "0.59.1", default-features = false, features =
["services-memory"] }
sysinfo = { version = "0.39.6", default-features = false, features =
["system"] }
+[target.'cfg(target_os = "linux")'.dev-dependencies]
Review Comment:
Sorry, I thought this change was documented in the commit message.
The reason for it is that Rust uses the system allocator by default, which
is generally tuned for throughput. It achieves that in part by aggressively
caching freed allocations by size classes to avoid repeated syscalls.
On my benchmark host, I can reliably observe up to ~2x performance drift
based solely on benchmark execution order. In practice, that means `cargo bench
arrow_writer` and `cargo bench arrow_writer -- some_filter` can produce
materially different results for the same matching benchmarks, depending on
whether allocator caches have already been warmed up by earlier benchmarks in
the invocation that filters.
I pulled a fair bit of hair out tracking this down. Jemalloc is a good fit
here because it gives us enough control over allocator behavior to minimize
this kind of cross-benchmark pollution. With the current configuration, I can
no longer reproduce the large execution-order-dependent variance I was seeing
before.
The allocator configuration is here:
https://github.com/apache/arrow-rs/pull/11267/changes#diff-4cc46e66635cf845f7b12621356c189294fcf21a6dda0f1e4786cad34445dce2R21-R31
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]