Krzysztof WilczyƄski reviewed this off-list while looking at carrying it
in a distribution kernel, and raised two points. I'd like to bring them
here, with numbers, before sending a v2. Cc'ing him.

1. ftrace_cmp_addr() duplicates ftrace_cmp_ips().

Agreed. For v2 I have dropped ftrace_cmp_addr() and moved
ftrace_cmp_ips() up, above the FTRACE_MCOUNT_MAX_OFFSET block, so that
ftrace_process_locs() still sees it on architectures that don't define
FTRACE_MCOUNT_MAX_OFFSET. That builds on x86_64 and arm64, and on x86_64
with CONFIG_MODULES=n.

2. Should the sort use sort_nonatomic()?

I timed the collection and the sort with ktime_get_ns() on a Ryzen 3
3200U (v7.3-rc5 plus this patch, debug printk only):

  module     addresses   collect     sort()
  amdgpu        64983    0.41 ms    36.4 ms
  mac80211       8199    0.11 ms    12.0 ms
  nouveau       19354    0.06 ms     9.6 ms
  kvm            7993    0.07 ms     7.9 ms
  radeon         9882    0.03 ms     4.6 ms
  all 141 modules loaded at boot:
                         1.09 ms    94.3 ms

So the sort dominates the new code, and amdgpu spends 36 ms in it.
It runs in ftrace_module_enable() before ftrace_lock is taken, so it is
sleepable and is preempted normally under full or lazy preemption.
What it does not have is a resched point, which only matters for
PREEMPT_NONE and PREEMPT_VOLUNTARY. sort_nonatomic() would add one, but
commit 340e3c5165d4 ("iommu/arm-smmu-v3: Replace sort_nonatomic() with
sort()") removed its last caller with the intent of dropping it, so I'd
rather not add a user. For comparison, the lookups this replaces took
about 4 s for amdgpu on this machine under ftrace_lock, and needed the
cond_resched() from commit 4099b98203d6 ("ftrace: Fix softlockup in
ftrace_module_enable").

Is 36 ms without a resched point acceptable here, or would you prefer
something else?

For reference, the cold-boot A/B with v1 and the v2 change, on the same
machine with a distribution kernel (linux-omarchy 7.2.5, amdgpu from the
initramfs, three boots each):

                     stock     v1        v2
  amdgpu probed      6.17 s    2.03 s    2.03 s
  kernel (systemd)   6.64 s    2.48 s    2.50 s
  modprobe radeon    155 ms     83 ms     83 ms
  modprobe nouveau   468 ms    141 ms    138 ms

(medians; radeon and nouveau have no device on this machine). The list
of functions in available_filter_functions is identical across the three
(84301 entries), and none of the boots logged a warning or soft lockup.

The same three kernels on a faster machine, a Ryzen AI MAX+ 395
(Strix Halo), three boots each:

                     stock     v1        v2
  amdgpu probed      4.84 s    3.80 s    3.81 s
  modprobe radeon     47 ms     32 ms     32 ms
  modprobe nouveau   130 ms     60 ms     59 ms

The saving is smaller there, about 1 s for amdgpu rather than 4 s, as
expected with a cheaper per-symbol lookup. available_filter_functions is
again identical across the three (84676 entries), and no boot logged a
warning or soft lockup. The systemd kernel time is left out because on
this machine it includes a LUKS passphrase prompt.

Unless there are other comments, I'll send v2 with the change in (1)
and these numbers in the changelog in a few days.

Thanks,
Lawrence

Reply via email to