mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jay Wang <wanjay@amazon.com>
To: <ihor.solodrai@linux.dev>, <bpf@vger.kernel.org>,
	<ast@kernel.org>, <daniel@iogearbox.net>, <andrii@kernel.org>,
	<eddyz87@gmail.com>, <memxor@gmail.com>
Cc: <alan.maguire@oracle.com>, <martin.lau@linux.dev>,
	<yonghong.song@linux.dev>, <jolsa@kernel.org>, <qmo@kernel.org>,
	<nathan@kernel.org>, <nsc@kernel.org>,
	<linux-kbuild@vger.kernel.org>, <linux@weissschuh.net>,
	<christian@heusel.eu>, <mcgrof@kernel.org>, <petr.pavlu@suse.com>,
	<samitolvanen@google.com>, <linux-modules@vger.kernel.org>,
	<rostedt@goodmis.org>, <mhiramat@kernel.org>,
	<mathieu.desnoyers@efficios.com>,
	<linux-trace-kernel@vger.kernel.org>, <acme@kernel.org>,
	<namhyung@kernel.org>, <irogers@google.com>,
	<linux-perf-users@vger.kernel.org>, <jikos@kernel.org>,
	<bentiss@kernel.org>, <linux-input@vger.kernel.org>,
	<tj@kernel.org>, <void@manifault.com>, <arighi@nvidia.com>,
	<changwoo@igalia.com>, <sched-ext@lists.linux.dev>,
	<shuah@kernel.org>, <linux-kselftest@vger.kernel.org>,
	<ojeda@kernel.org>, <rust-for-linux@vger.kernel.org>,
	<arnd@arndb.de>, <linux-doc@vger.kernel.org>,
	<linux-kernel@vger.kernel.org>, <abuehaze@amazon.com>,
	<doebel@amazon.de>, <mpohlack@amazon.de>,
	<jay.wang.upstream@gmail.com>
Subject: Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
Date: Sun, 4 Oct 2026 22:21:27 +0000	[thread overview]
Message-ID: <20261004222127.31128-1-wanjay@amazon.com> (raw)
In-Reply-To: <44315425-c505-46fe-8379-f19d78ff3abb@linux.dev>

>    - If I know I don't need BPF on the workload, and these 5.4 MB
>      bite me, I can just turn off BPF

They cannot, even when they know they do not need it.  As the cover
letter says, the cloud provider usually ships one kernel build to all of
its customers.  Customers install their own packages and workloads around
that kernel, but they do not build or customize the kernel.  Turning
BPF off is not a switch they have.

> What I find strange is for a user of this scale *to not know* whether
> BPF will be used in the workload or not, which is what lazy load may
> solve.

They do know.  They run predictable workloads on a large number of
small instances, and BPF is not part of them.  They spread the load by
resource usage, so at that scale every byte matters: more free memory
on each instance means fewer instances in total.

>    * Compress it.
[...]
>      Keep zstd-compressed blob in the kernel image and decompress and
>      parse it synchronously on first use.

Thanks for putting this forward.  This compression approach is simpler 
than loading a module.  I expect the overall code complexity to stay similar
to v3, since both need to find the first user and defer the module BTF parsing
until then, but it should be easier to merge into mainline because it avoids
the request_module() deadlocks.

Would you like to prepare it for a formal submission, or would you
like me to do so?  Either works for me.  Ideally it lands in time for
the next LTS kernel, so that we can adopt it there.

From what I can see, most of the work left is in that deferral.
Whoever needs the BTF first also parses the waiting module BTFs and
replays the queued registrations, struct_ops ->init() included, under
whatever locks it holds at that point, e.g. event_mutex or
bpf_verifier_lock, or inside the module notifier.  So the
cand_cache_mutex ABBA you found may not be the only deadlock.  Besides
that, only x86_64 was touched so far, so it needs extending to the
other architectures as well.

      parent reply	other threads:[~2026-10-04 22:21 UTC|newest]

Thread overview: 26+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 22:52 Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-02 11:48   ` Alexei Starovoitov
2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02  9:47   ` Alan Maguire
2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02  7:34   ` Jay Wang
2026-10-02 10:05     ` Alan Maguire
2026-10-02 20:58     ` Ihor Solodrai
2026-10-03  6:38       ` Alexei Starovoitov
2026-10-03 11:45         ` Alan Maguire
2026-10-03 12:19           ` Alexei Starovoitov
2026-10-04 22:21       ` Jay Wang [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261004222127.31128-1-wanjay@amazon.com \
    --to=wanjay@amazon.com \
    --cc=abuehaze@amazon.com \
    --cc=acme@kernel.org \
    --cc=alan.maguire@oracle.com \
    --cc=andrii@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=arnd@arndb.de \
    --cc=ast@kernel.org \
    --cc=bentiss@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=changwoo@igalia.com \
    --cc=christian@heusel.eu \
    --cc=daniel@iogearbox.net \
    --cc=doebel@amazon.de \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=irogers@google.com \
    --cc=jay.wang.upstream@gmail.com \
    --cc=jikos@kernel.org \
    --cc=jolsa@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-input@vger.kernel.org \
    --cc=linux-kbuild@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=linux@weissschuh.net \
    --cc=martin.lau@linux.dev \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mcgrof@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mhiramat@kernel.org \
    --cc=mpohlack@amazon.de \
    --cc=namhyung@kernel.org \
    --cc=nathan@kernel.org \
    --cc=nsc@kernel.org \
    --cc=ojeda@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=qmo@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=samitolvanen@google.com \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®