mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jay Wang <wanjay@amazon.com>
To: <alexei.starovoitov@gmail.com>
Cc: <abuehaze@amazon.com>, <alan.maguire@oracle.com>,
	<andrii@kernel.org>, <arnd@arndb.de>, <bpf@vger.kernel.org>,
	<daniel@iogearbox.net>, <doebel@amazon.de>, <eddyz87@gmail.com>,
	<jay.wang.upstream@gmail.com>, <jolsa@kernel.org>,
	<linux-kbuild@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<linux-modules@vger.kernel.org>, <martin.lau@linux.dev>,
	<mcgrof@kernel.org>, <memxor@gmail.com>, <mpohlack@amazon.de>,
	<nathan@kernel.org>, <nsc@kernel.org>, <ojeda@kernel.org>,
	<petr.pavlu@suse.com>, <rust-for-linux@vger.kernel.org>,
	<samitolvanen@google.com>, <yonghong.song@linux.dev>,
	<rostedt@goodmis.org>, <mhiramat@kernel.org>,
	<mathieu.desnoyers@efficios.com>,
	<linux-trace-kernel@vger.kernel.org>
Subject: Re: [PATCH bpf-next v3 4/9] bpf: take the vmlinux BTF from the btf_vmlinux module
Date: Thu, 1 Oct 2026 23:51:42 +0000	[thread overview]
Message-ID: <20261001235142.4626-1-wanjay@amazon.com> (raw)
In-Reply-To: <DLP3USQDIMJG.1FUZXEXOX38TR@gmail.com>

On Sat, Sep 26, 2026 at 08:29:12AM +0000, Alexei Starovoitov wrote:
> On Fri, Sep 25, 2026 at 10:42 PM Jay Wang <wanjay@amazon.com> wrote:
> > +	if (!data && load) {
> > +		/*
> > +		 * The module notifier installs the BTF before init_module()
> > +		 * returns, so it is either there after this or the module is
> > +		 * not available (yet).  Not cached: a later call retries,
> > +		 * e.g. once the module becomes reachable on the root fs.
> > +		 */
> > +		request_module("btf_vmlinux");
>
> This will deadlock.
> event_btf_ids_read() calls btf_get_module_btf(NULL) with event_mutex
> held. With =m and BTF not loaded yet
>   cat /sys/kernel/tracing/events/sched/sched_switch/btf_ids
> gets here and waits for modprobe. modprobe gets to
> trace_module_notify() which takes event_mutex.
> lockdep doesn't see it.
>
> print_function_args() gets here for every line of the trace,
> from ftrace_dump() with irqs off too. When the module is not
> installed that is one modprobe per line.
>
> Every caller of bpf_get_btf_vmlinux() and bpf_find_btf_id() was
> written for a function that doesn't wait for user space.

Right.  Given that fixing those callers one by one does not work, as
there are too many and every new one would have to know, I went with a
different mechanism in v4:
https://lore.kernel.org/bpf/20261001225214.12351-1-wanjay@amazon.com/

In short, the lookups no longer load the BTF; only requests from user
space do, at their start.

The lookups cannot wait for the module, because they run in places that
must not sleep, or that hold locks the module load needs too: under
event_mutex in your btf_ids case, with interrupts off in ftrace_dump(),
inside a running BPF program.  Loading the module means waiting for
modprobe, so waiting in any of them can deadlock or sleep where it must
not.

Loading at the start of a request from user space is enough, because
that is where every need for the BTF begins: a program or map that uses
kernel types, a read of the BTF file, a probe event with BTF arguments.
At that point the request holds nothing yet, so it can safely wait for
modprobe.  And once it has loaded the BTF, the lookups it makes later
find it there, so they have no reason to wait.  Only the kernel's own
early users, the kfunc and struct_ops registrations at boot and the BTF
of modules loaded before it, come before any such request, and they are
queued until the BTF arrives.

With =m, bpf_get_btf_vmlinux() and bpf_find_btf_id() never load the
module and never sleep.  Until the BTF is loaded they fail as on a
kernel without BTF (NULL, -EINVAL), so every existing caller keeps the
assumptions it was written with.  The only function that loads the BTF
is a new bpf_load_btf_vmlinux(), called at the start of such a request,
with nothing held that loading a module needs:

 - on entry to the bpf() syscall: a BPF_PROG_LOAD, BPF_MAP_CREATE or
   BPF_BTF_LOAD that failed while the BTF was missing is run once more
   after loading it, the way tc and nf_tables retry after loading a
   module.  The verifier itself never waits, and bpf_sys_bpf() does not
   go through that entry;
 - BPF_BTF_GET_NEXT_ID with CAP_SYS_ADMIN, and loading a syscall
   program, since a light skeleton loader loads its programs while it
   runs;
 - read() of /sys/kernel/btf/vmlinux.  Not mmap(): it runs with the
   caller's mmap_lock held, and a uprobe registration holds event_mutex
   while it takes the mmap_lock of every mm mapping the probed file,
   which would close a loop through the module load.  mmap() fails
   until the BTF is loaded, and libbpf then falls back to read();
 - read() of the sysfs file of a module built against a distilled base
   (.BTF.base), through a work item: the reader holds the file's kernfs
   active reference, which MODULE_STATE_GOING drains with the module
   notifier chain held;
 - reading btf_ids (before event_mutex), creating probe events with BTF
   arguments (the parser holds only dyn_event_ops_mutex, which no module
   load takes), and mounting bpffs with delegate options that name
   commands or types.

print_function_args() only uses the BTF if it is already loaded and
never waits for it, so no modprobe per line.  Both of your cases now
work as the first BTF user after boot: the cat of btf_ids loads the BTF
before taking event_mutex and prints the ids, and func-args with sysrq-z
prints the functions without arguments, as on a kernel without BTF; the
arguments show up once a request from user space has loaded the BTF.

Jay

  reply	other threads:[~2026-10-01 23:51 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 22:42 [PATCH bpf-next v3 0/9] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 1/9] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 2/9] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 3/9] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 4/9] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-09-25 23:23   ` bot+bpf-ci
2026-10-01 23:51     ` Jay Wang
2026-09-26  8:29   ` Alexei Starovoitov
2026-10-01 23:51     ` Jay Wang [this message]
2026-09-25 22:42 ` [PATCH bpf-next v3 5/9] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-09-25 23:34   ` bot+bpf-ci
2026-10-01 23:52     ` Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 6/9] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 7/9] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-09-25 23:23   ` bot+bpf-ci
2026-10-01 23:52     ` Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 8/9] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-09-25 23:23   ` bot+bpf-ci
2026-10-01 23:53     ` Jay Wang
2026-09-25 22:42 ` [PATCH bpf-next v3 9/9] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-09-25 23:34   ` bot+bpf-ci
2026-10-01 23:53     ` Jay Wang
2026-09-28 10:00   ` Alan Maguire
     [not found]     ` <DM6PR18MB2666F6B235D5B356AA157A7DA98D2@DM6PR18MB2666.namprd18.prod.outlook.com>
2026-09-29 18:25       ` Alan Maguire
2026-10-01 23:53         ` Jay Wang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001235142.4626-1-wanjay@amazon.com \
    --to=wanjay@amazon.com \
    --cc=abuehaze@amazon.com \
    --cc=alan.maguire@oracle.com \
    --cc=alexei.starovoitov@gmail.com \
    --cc=andrii@kernel.org \
    --cc=arnd@arndb.de \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=doebel@amazon.de \
    --cc=eddyz87@gmail.com \
    --cc=jay.wang.upstream@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=linux-kbuild@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mcgrof@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mhiramat@kernel.org \
    --cc=mpohlack@amazon.de \
    --cc=nathan@kernel.org \
    --cc=nsc@kernel.org \
    --cc=ojeda@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=rostedt@goodmis.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=samitolvanen@google.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®