From: Lawrence Lin <deduce@gmail.com>
To: Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: "Krzysztof Wilczyński" <kwilczynski@kernel.org>,
"Petr Pavlu" <petr.pavlu@suse.com>,
"Stanislaw Gruszka" <stf_xl@wp.pl>,
linux-modules@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] ftrace: Avoid quadratic symbol lookups in ftrace_module_enable()
Date: Sat, 3 Oct 2026 22:03:58 -0500 [thread overview]
Message-ID: <20261004030438.434327-1-deduce@gmail.com> (raw)
In-Reply-To: <20261003-ftrace-mod-bsearch-v1-1-92e2fd2d80ff@gmail.com>
Krzysztof Wilczyński reviewed this off-list while looking at carrying it
in a distribution kernel, and raised two points. I'd like to bring them
here, with numbers, before sending a v2. Cc'ing him.
1. ftrace_cmp_addr() duplicates ftrace_cmp_ips().
Agreed. For v2 I have dropped ftrace_cmp_addr() and moved
ftrace_cmp_ips() up, above the FTRACE_MCOUNT_MAX_OFFSET block, so that
ftrace_process_locs() still sees it on architectures that don't define
FTRACE_MCOUNT_MAX_OFFSET. That builds on x86_64 and arm64, and on x86_64
with CONFIG_MODULES=n.
2. Should the sort use sort_nonatomic()?
I timed the collection and the sort with ktime_get_ns() on a Ryzen 3
3200U (v7.3-rc5 plus this patch, debug printk only):
module addresses collect sort()
amdgpu 64983 0.41 ms 36.4 ms
mac80211 8199 0.11 ms 12.0 ms
nouveau 19354 0.06 ms 9.6 ms
kvm 7993 0.07 ms 7.9 ms
radeon 9882 0.03 ms 4.6 ms
all 141 modules loaded at boot:
1.09 ms 94.3 ms
So the sort dominates the new code, and amdgpu spends 36 ms in it.
It runs in ftrace_module_enable() before ftrace_lock is taken, so it is
sleepable and is preempted normally under full or lazy preemption.
What it does not have is a resched point, which only matters for
PREEMPT_NONE and PREEMPT_VOLUNTARY. sort_nonatomic() would add one, but
commit 340e3c5165d4 ("iommu/arm-smmu-v3: Replace sort_nonatomic() with
sort()") removed its last caller with the intent of dropping it, so I'd
rather not add a user. For comparison, the lookups this replaces took
about 4 s for amdgpu on this machine under ftrace_lock, and needed the
cond_resched() from commit 4099b98203d6 ("ftrace: Fix softlockup in
ftrace_module_enable").
Is 36 ms without a resched point acceptable here, or would you prefer
something else?
For reference, the cold-boot A/B with v1 and the v2 change, on the same
machine with a distribution kernel (linux-omarchy 7.2.5, amdgpu from the
initramfs, three boots each):
stock v1 v2
amdgpu probed 6.17 s 2.03 s 2.03 s
kernel (systemd) 6.64 s 2.48 s 2.50 s
modprobe radeon 155 ms 83 ms 83 ms
modprobe nouveau 468 ms 141 ms 138 ms
(medians; radeon and nouveau have no device on this machine). The list
of functions in available_filter_functions is identical across the three
(84301 entries), and none of the boots logged a warning or soft lockup.
The same three kernels on a faster machine, a Ryzen AI MAX+ 395
(Strix Halo), three boots each:
stock v1 v2
amdgpu probed 4.84 s 3.80 s 3.81 s
modprobe radeon 47 ms 32 ms 32 ms
modprobe nouveau 130 ms 60 ms 59 ms
The saving is smaller there, about 1 s for amdgpu rather than 4 s, as
expected with a cheaper per-symbol lookup. available_filter_functions is
again identical across the three (84676 entries), and no boot logged a
warning or soft lockup. The systemd kernel time is left out because on
this machine it includes a LUKS passphrase prompt.
Unless there are other comments, I'll send v2 with the change in (1)
and these numbers in the changelog in a few days.
Thanks,
Lawrence
next prev parent reply other threads:[~2026-10-04 3:07 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 16:27 Lawrence Lin via B4 Relay
2026-10-04 3:03 ` Lawrence Lin [this message]
2026-10-04 9:00 ` David Laight
2026-10-04 17:05 ` Lawrence Lin
2026-10-04 9:09 ` Steven Rostedt
2026-10-04 17:05 ` Lawrence Lin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261004030438.434327-1-deduce@gmail.com \
--to=deduce@gmail.com \
--cc=kwilczynski@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-modules@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=petr.pavlu@suse.com \
--cc=rostedt@goodmis.org \
--cc=stf_xl@wp.pl \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®