HippoBaro commented on PR #11267:
URL: https://github.com/apache/arrow-rs/pull/11267#issuecomment-5876792873
For some of my recent work, I’ve set up a benchmarking environment designed
to make results as deterministic as possible. The benchmarks are now much more
reliable than the CI bot here, where I’ve seen performance drift by as much as
2x between runs.
That said, even with a dedicated bare-metal host and a tuned TuneD profile,
I still see up to ~10% variance between runs, especially on heavily sparse
inputs.
I’m sharing the TuneD profile here in case anyone finds it useful:
```ini
[main]
summary=Deterministic single-threaded benchmarking on Graviton bare metal
include=cpu-partitioning
[variables]
# Cores listed here are isolated: removed from the scheduler's
load-balancing
# domain, get nohz_full + rcu_nocbs, and are excluded from
systemd/workqueue
# default affinity. All other cores remain fully schedulable.
# no_balance_cores == isolated_cores forces isolcpus in the kernel cmdline
# so the scheduler can never place other tasks on these cores.
isolated_cores=42
no_balance_cores=42
[cpu]
governor=performance
# No-op on hardware with no cpufreq/P-state interface (e.g. AWS Graviton):
# there is no DVFS to fight in the first place (fixed frequency).
[vm]
transparent_hugepages=never
[sysctl]
kernel.numa_balancing=0
vm.swappiness=0
kernel.watchdog=0
kernel.nmi_watchdog=0
```
Once configured like so, no process, including kernel tasks, may run on
core 42, unless specifically invoked using `taskset -c 42 ...` @adriangb I
believe you contributed to, or are responsible for, our benchmark bot. The
above might be of interest to you!
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]