From: Ihor Solodrai <ihor.solodrai@linux.dev>
To: Jay Wang <wanjay@amazon.com>,
bpf@vger.kernel.org, ast@kernel.org, daniel@iogearbox.net,
andrii@kernel.org, eddyz87@gmail.com, memxor@gmail.com
Cc: alan.maguire@oracle.com, martin.lau@linux.dev,
yonghong.song@linux.dev, jolsa@kernel.org, qmo@kernel.org,
nathan@kernel.org, nsc@kernel.org, linux-kbuild@vger.kernel.org,
linux@weissschuh.net, christian@heusel.eu, mcgrof@kernel.org,
petr.pavlu@suse.com, samitolvanen@google.com,
linux-modules@vger.kernel.org, rostedt@goodmis.org,
mhiramat@kernel.org, mathieu.desnoyers@efficios.com,
linux-trace-kernel@vger.kernel.org, acme@kernel.org,
namhyung@kernel.org, irogers@google.com,
linux-perf-users@vger.kernel.org, jikos@kernel.org,
bentiss@kernel.org, linux-input@vger.kernel.org, tj@kernel.org,
void@manifault.com, arighi@nvidia.com, changwoo@igalia.com,
sched-ext@lists.linux.dev, shuah@kernel.org,
linux-kselftest@vger.kernel.org, ojeda@kernel.org,
rust-for-linux@vger.kernel.org, arnd@arndb.de,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
abuehaze@amazon.com, doebel@amazon.de, mpohlack@amazon.de,
jay.wang.upstream@gmail.com
Subject: Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
Date: Fri, 2 Oct 2026 13:58:06 -0700 [thread overview]
Message-ID: <44315425-c505-46fe-8379-f19d78ff3abb@linux.dev> (raw)
In-Reply-To: <20261002073422.14253-1-wanjay@amazon.com>
Hi Jay, thank you for taking the time to reply.
On 2026-10-02 12:34 a.m., Jay Wang wrote:
>> A question I had: is it really worth 2k of complicated kernel code to
>> *maybe sometimes* save <10Mb of memory?
>
> There are many reasons this is worth it. The main ones:
>
> 1. This did not start from a number we picked, but from a real use
> case: users who run large fleets of small instances, 1 GiB of
> memory or less, with their workloads sized to fit. We are not able
> to disclose more detail, but for them 5.4 MB on every instance is
> real money, and it can be exactly what pushes a workload over its
> memory budget and onto the next instance size. That is why we took
> this on, even though we knew it would not be a small change.
Ok, good to know there is a real customer for this.
Obviously I don't know any details of the use-case apart from what
you've shared. But bear with me trying the customer's hat on:
- I am running a workload on a legion of 1GiB instances
- I care about memory footprint a lot, because it directly
translates into the amount of compute I pay for
- If I know I don't need BPF on the workload, and these 5.4 MB
bite me, I can just turn off BPF
- If I know I need BPF, then those 5.4 MB must be loaded anyways, so
I have to look for savings elsewhere
What I find strange is for a user of this scale *to not know* whether
BPF will be used in the workload or not, which is what lazy load may
solve.
One thing I can imagine is an opaque workload: untrusted (AI agents
etc), end-user-defined (including BPF usage), or confidential. But in
these cases, what are the chances that BPF is used there? They are
high, BPF is very widespread.
I can also imagine normal workload not needing BPF, but when something
goes wrong an observability or security thing turning on that loads
BPF programs. But this contradicts the "5.4 MB may push workload over
1GiB" premise: you still need to reserve/swap memory for just-in-time
BPF-based tools, otherwise you OOM.
I of course may be missing something, happy to be corrected.
>
> 2. Loading code and data only when a system needs them is what kernel
> modules exist for. And most of them take well under ~1 MB once
> loaded (nf_conntrack, overlay, vfat); even big ones like ext4, kvm
> and btrfs stay around 1-2.5 MB. The vmlinux BTF is 5.4 MB.
True. At the same time BPF without vmlinux BTF is barely useful. And
it is trusted by the verifier. Both are addressed by BTF being
built-in in the kernel image.
>
> 3. We are not alone in trying to keep BTF out of memory until it is
> needed. The inline BTF work [1] plans to deliver its data, which
> is even larger, through a module too. The .BTF.link record and the
> resolve_btfids option that serve both are already part of this
> series (patch 10), so this is not machinery for the vmlinux BTF
> alone.
The inline data is different. It's bigger by it's nature, and it has
more specific users: tracing tools. The userspace tools are able to
themselves decide whether to use that data. The kernel only needs to
make it available.
In comparison, as you yourself noted in the cover:
CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext and bpf-lsm
all depend on vmlinux BTF
>
>> Do those tiny VMs that you target run systemd with BPF LSM?
>
> Not necessarily, and that is the image's choice. Even with BPF LSM
> built in and bpf in CONFIG_LSM, memory-conscious users can turn it off
> at boot with an lsm= list that leaves it out, and then systemd does not
> load restrict_fs at all.
>
> And even where user space is in the way, which as above it does not
> have to be, the answer is to fix user space so it gets this win too,
> not to give up on it.
>
>> If the answer is "yes" for the majority of them,
>
> "The majority" is also hard to pin down here: what runs at boot
> depends on the user space packages built into each image, and on the
> workload.
Yes, it is hard to pin down. Which is why we first need to figure out
whether the alleged memory savings actually help anyone.
Adding code to the kernel is not free. More complexity means bigger
bug surface: more opportunities for concurrency bugs, security bugs
etc. Especially runtime loads with retries and stuff. I'm sure you
understand the future cost of all that.
So the bias shouldn't be "let's implement a big thing that might
hypothetically help someone". It's backwards IMO.
>
> So the question should be what it takes to make this work, not whether
> to drop it because it is complicated. If there are ways to cut it
> down, or issues found in the implementation, we would be happy to take
> them and go through every one.
The question is in the trade-off.
Would you run a million line python program to search for a string in
a text file? No, that's absurd. But you could.
I'm not saying 2k line patch series is unacceptable in principle. It
may be justified. But I am not convinced it is justified in this case.
Let's say we (as in kernel devs) decided that we want to reduce
vmlinux BTF size, and only load it when necessary. Here are a couple
of alternatives, simpler in comparison to CONFIG_DEBUG_INFO_BTF=m,
although still a bit complex:
* Compress it.
$ ls -lah /sys/kernel/btf/vmlinux
-r--r--r-- 1 root root 6.7M Jul 18 01:54 /sys/kernel/btf/vmlinux
$ tar --zstd -cf /tmp/vmlinux.btf.zstd -C /sys/kernel/btf vmlinux
$ ls -lah /tmp/vmlinux.btf.zstd
-rw-r--r-- 1 isolodrai users 2.2M Oct 2 10:11 /tmp/vmlinux.btf.zstd
3x win right there.
Keep zstd-compressed blob in the kernel image and decompress and
parse it synchronously on first use.
Here is a prototype (vibe-coded, obviously):
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.zstd-20261002
* Use an external blob.
Ship /lib/modules/$r/vmlinux.btf, checked against a SHA-256
recorded in the image. The catch is a synchronous file I/O under
caller locks, so this is less straightforward.
But this should cover embedded case that Alan has mentioned.
Here is a prototype:
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.external-20261002
Both alternatives avoid kernel-module loading machine. If we think
harder, we may come up with something even more clever.
I don't particularly prefer one approach over the other, including the
module one. Maintainers probably have a better intuition on that.
The point I'm trying to make is that if you could achieve similar
memory savings with say 500-line well encapsulated diff (or even
better: by refactoring, simplifying, deleting code) it would have
already landed.
And if it was a 50-line patch, the discussion on whether anyone cares
wouldn't be relevant, because it's so cheap.
As it stands, it's not clear (to me at least) in what circumstances
we'll get the promised memory savings at all.
>
>
> [1] https://lore.kernel.org/bpf/20260916074118.1007116-1-alan.maguire@oracle.com/
prev parent reply other threads:[~2026-10-02 20:58 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 22:52 Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45 ` bot+bpf-ci
2026-10-02 11:48 ` Alexei Starovoitov
2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45 ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29 ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02 9:47 ` Alan Maguire
2026-10-02 4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02 7:34 ` Jay Wang
2026-10-02 10:05 ` Alan Maguire
2026-10-02 20:58 ` Ihor Solodrai [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=44315425-c505-46fe-8379-f19d78ff3abb@linux.dev \
--to=ihor.solodrai@linux.dev \
--cc=abuehaze@amazon.com \
--cc=acme@kernel.org \
--cc=alan.maguire@oracle.com \
--cc=andrii@kernel.org \
--cc=arighi@nvidia.com \
--cc=arnd@arndb.de \
--cc=ast@kernel.org \
--cc=bentiss@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=changwoo@igalia.com \
--cc=christian@heusel.eu \
--cc=daniel@iogearbox.net \
--cc=doebel@amazon.de \
--cc=eddyz87@gmail.com \
--cc=irogers@google.com \
--cc=jay.wang.upstream@gmail.com \
--cc=jikos@kernel.org \
--cc=jolsa@kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-input@vger.kernel.org \
--cc=linux-kbuild@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-modules@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=linux@weissschuh.net \
--cc=martin.lau@linux.dev \
--cc=mathieu.desnoyers@efficios.com \
--cc=mcgrof@kernel.org \
--cc=memxor@gmail.com \
--cc=mhiramat@kernel.org \
--cc=mpohlack@amazon.de \
--cc=namhyung@kernel.org \
--cc=nathan@kernel.org \
--cc=nsc@kernel.org \
--cc=ojeda@kernel.org \
--cc=petr.pavlu@suse.com \
--cc=qmo@kernel.org \
--cc=rostedt@goodmis.org \
--cc=rust-for-linux@vger.kernel.org \
--cc=samitolvanen@google.com \
--cc=sched-ext@lists.linux.dev \
--cc=shuah@kernel.org \
--cc=tj@kernel.org \
--cc=void@manifault.com \
--cc=wanjay@amazon.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®