HippoBaro commented on code in PR #11267:
URL: https://github.com/apache/arrow-rs/pull/11267#discussion_r4170070339


##########
parquet/Cargo.toml:
##########
@@ -94,6 +94,9 @@ object_store = { workspace = true, features = ["azure", "fs"] 
}
 opendal = { version = "0.59.1", default-features = false, features = 
["services-memory"] }
 sysinfo = { version = "0.39.6", default-features = false, features = 
["system"] }
 
+[target.'cfg(target_os = "linux")'.dev-dependencies]

Review Comment:
   Sorry, I thought this change was documented in the commit message.
   
   The reason for it is that Rust uses the system allocator by default, which 
is generally tuned for throughput. It achieves that in part by aggressively 
caching freed allocations by size classes to avoid repeated syscalls.
   
   On my benchmark host, I can reliably observe up to ~2x performance drift 
based solely on benchmark execution order. In practice, that means `cargo bench 
arrow_writer` and `cargo bench arrow_writer -- some_filter` can produce 
materially different results for the same matching benchmarks, depending on 
whether allocator caches have already been warmed up by earlier benchmarks in 
the invocation that filters.
   
   I pulled a fair bit of hair out tracking this down. Jemalloc is a good fit 
here because it gives us enough control over allocator behavior to minimize 
this kind of cross-benchmark pollution. With the current configuration, I can 
no longer reproduce the large execution-order-dependent variance I was seeing 
before.
   
   The allocator configuration is here:
   
https://github.com/apache/arrow-rs/pull/11267/changes#diff-4cc46e66635cf845f7b12621356c189294fcf21a6dda0f1e4786cad34445dce2R21-R31



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to