mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Lawrence Lin <deduce@gmail.com>
To: Steven Rostedt <rostedt@goodmis.org>,
	Masami Hiramatsu <mhiramat@kernel.org>,
	Mark Rutland <mark.rutland@arm.com>,
	Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: "Krzysztof Wilczyński" <kwilczynski@kernel.org>,
	"Petr Pavlu" <petr.pavlu@suse.com>,
	"Stanislaw Gruszka" <stf_xl@wp.pl>,
	linux-modules@vger.kernel.org,
	linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] ftrace: Avoid quadratic symbol lookups in ftrace_module_enable()
Date: Sat,  3 Oct 2026 22:03:58 -0500	[thread overview]
Message-ID: <20261004030438.434327-1-deduce@gmail.com> (raw)
In-Reply-To: <20261003-ftrace-mod-bsearch-v1-1-92e2fd2d80ff@gmail.com>

Krzysztof Wilczyński reviewed this off-list while looking at carrying it
in a distribution kernel, and raised two points. I'd like to bring them
here, with numbers, before sending a v2. Cc'ing him.

1. ftrace_cmp_addr() duplicates ftrace_cmp_ips().

Agreed. For v2 I have dropped ftrace_cmp_addr() and moved
ftrace_cmp_ips() up, above the FTRACE_MCOUNT_MAX_OFFSET block, so that
ftrace_process_locs() still sees it on architectures that don't define
FTRACE_MCOUNT_MAX_OFFSET. That builds on x86_64 and arm64, and on x86_64
with CONFIG_MODULES=n.

2. Should the sort use sort_nonatomic()?

I timed the collection and the sort with ktime_get_ns() on a Ryzen 3
3200U (v7.3-rc5 plus this patch, debug printk only):

  module     addresses   collect     sort()
  amdgpu        64983    0.41 ms    36.4 ms
  mac80211       8199    0.11 ms    12.0 ms
  nouveau       19354    0.06 ms     9.6 ms
  kvm            7993    0.07 ms     7.9 ms
  radeon         9882    0.03 ms     4.6 ms
  all 141 modules loaded at boot:
                         1.09 ms    94.3 ms

So the sort dominates the new code, and amdgpu spends 36 ms in it.
It runs in ftrace_module_enable() before ftrace_lock is taken, so it is
sleepable and is preempted normally under full or lazy preemption.
What it does not have is a resched point, which only matters for
PREEMPT_NONE and PREEMPT_VOLUNTARY. sort_nonatomic() would add one, but
commit 340e3c5165d4 ("iommu/arm-smmu-v3: Replace sort_nonatomic() with
sort()") removed its last caller with the intent of dropping it, so I'd
rather not add a user. For comparison, the lookups this replaces took
about 4 s for amdgpu on this machine under ftrace_lock, and needed the
cond_resched() from commit 4099b98203d6 ("ftrace: Fix softlockup in
ftrace_module_enable").

Is 36 ms without a resched point acceptable here, or would you prefer
something else?

For reference, the cold-boot A/B with v1 and the v2 change, on the same
machine with a distribution kernel (linux-omarchy 7.2.5, amdgpu from the
initramfs, three boots each):

                     stock     v1        v2
  amdgpu probed      6.17 s    2.03 s    2.03 s
  kernel (systemd)   6.64 s    2.48 s    2.50 s
  modprobe radeon    155 ms     83 ms     83 ms
  modprobe nouveau   468 ms    141 ms    138 ms

(medians; radeon and nouveau have no device on this machine). The list
of functions in available_filter_functions is identical across the three
(84301 entries), and none of the boots logged a warning or soft lockup.

The same three kernels on a faster machine, a Ryzen AI MAX+ 395
(Strix Halo), three boots each:

                     stock     v1        v2
  amdgpu probed      4.84 s    3.80 s    3.81 s
  modprobe radeon     47 ms     32 ms     32 ms
  modprobe nouveau   130 ms     60 ms     59 ms

The saving is smaller there, about 1 s for amdgpu rather than 4 s, as
expected with a cheaper per-symbol lookup. available_filter_functions is
again identical across the three (84676 entries), and no boot logged a
warning or soft lockup. The systemd kernel time is left out because on
this machine it includes a LUKS passphrase prompt.

Unless there are other comments, I'll send v2 with the change in (1)
and these numbers in the changelog in a few days.

Thanks,
Lawrence

  reply	other threads:[~2026-10-04  3:07 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03 16:27 Lawrence Lin via B4 Relay
2026-10-04  3:03 ` Lawrence Lin [this message]
2026-10-04  9:00 ` David Laight
2026-10-04 17:05   ` Lawrence Lin
2026-10-04  9:09 ` Steven Rostedt
2026-10-04 17:05   ` Lawrence Lin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261004030438.434327-1-deduce@gmail.com \
    --to=deduce@gmail.com \
    --cc=kwilczynski@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=rostedt@goodmis.org \
    --cc=stf_xl@wp.pl \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®