mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
@ 2026-10-01 22:52 Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
                   ` (12 more replies)
  0 siblings, 13 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Based on and tested against bpf-next commit b5a4aa31abd6 ("bpf, cgroup:
Fix cgroup struct_ops query for a second attach type").

This series makes CONFIG_DEBUG_INFO_BTF a tristate, so that it can be
set to =m.  With =m the vmlinux BTF is carried by a module, btf_vmlinux,
that the kernel loads the first time user space asks for something that
needs the BTF.  On a system where nothing does, that saves ~5.4 MB of
RAM with a distribution config; on a system that uses BTF, it behaves as
with =y.  =y itself is untouched.

Problem
-------

The vmlinux BTF that CONFIG_DEBUG_INFO_BTF=y builds into the kernel
image takes ~5.4 MB of memory, resident from boot whether anything uses
it or not.  On small instances that is not negligible.

A distribution cannot simply turn it off for the users who do not need
it: it ships one kernel build for all its users, and BTF is not debug
info anymore.  CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext and
bpf-lsm all depend on it, so =n takes those away from everyone who does
use them.

Hence this series adds CONFIG_DEBUG_INFO_BTF=m: the BTF becomes an
on-demand module.  Users who never use BTF get the memory back; for
users who do, the first request that needs it loads it, and everything
works as with =y.

Approach
--------

Do not remove anything, defer it.  The BTF is generated exactly as
before, but with =m it is not part of the kernel image: it is packed
into a module, btf_vmlinux.ko, which the kernel loads itself when the
BTF is first needed.  Once loaded, the BTF stays.

Making that work ran into six problems.  The first is where the module
can be loaded at all; the next three are existing components that rely
on the vmlinux BTF being present from boot, which a loadable module
cannot provide; the last two follow from the BTF no longer being part of
the image.

1. Loading.  Problem: loading btf_vmlinux runs modprobe and waits for
   it, and the module load takes locks of its own (event_mutex, the
   module notifier chain, ...).  bpf_get_btf_vmlinux() and
   bpf_find_btf_id() are called with such locks held, from running BPF
   programs, and with interrupts disabled, so they must not load it.
   Solution: they never do.  With =m they return nothing until the BTF
   is loaded, as on a kernel without BTF, and do not even sleep.  A new
   bpf_load_btf_vmlinux() loads it, and is only called at the start of
   a request from user space, holding no lock that loading a module
   needs: on entry to the bpf() system call, which runs a
   BPF_PROG_LOAD, BPF_MAP_CREATE or BPF_BTF_LOAD that failed for want
   of the BTF once more after loading it, the way tc and nf_tables
   retry after loading a module; for BPF_BTF_GET_NEXT_ID and for
   loading a light skeleton loader; for read() of
   /sys/kernel/btf/vmlinux (mmap(), which runs under the caller's
   mmap_lock, only maps it once it is loaded); and for the tracefs and
   bpffs requests that need the BTF (btf_ids files, probe events with
   BTF arguments, bpffs delegate options that name commands or types).
   It returns once the BTF of the modules loaded before is registered
   too, also to concurrent callers.  Everything else, such as the
   ftrace argument printer, only uses the BTF if it is already there.

2. Verifier.  Problem: bpf_check() fetched the vmlinux BTF for every
   program, so the first socket filter at boot would have needed it on
   every system.  Solution: fetch the BTF only where kernel types enter
   a program (attach_btf, kfunc calls, ksyms, map pointer access,
   helpers that take or return kernel pointers, the program context
   type table).  A program using none of these never touches it.

3. Initcall registrations.  Problem: kfunc, dtor kfunc and struct_ops
   registrations run from initcalls and need the parsed BTF, which
   would again pull it in at boot.  Solution: queue them and apply the
   queue when the BTF is parsed, before it is published, so no program
   can ever see a vmlinux BTF that lacks its kfuncs or struct_ops.

4. Module BTF.  Problem: module BTF is split BTF against the vmlinux
   BTF and was parsed at module load.  A module loaded before the
   vmlinux BTF cannot be parsed yet, and the module notifier cannot
   load btf_vmlinux (that would nest a module load inside a module
   load).  Solution: keep the module's BTF aside (the same copy
   btf_parse_module() makes with =y, so module BTF costs the same in
   both; the saving is the vmlinux BTF only), expose it in
   /sys/kernel/btf right away, and parse and register it, together with
   the module's own kfuncs and struct_ops, once the vmlinux BTF arrives;
   for a module that is still initializing then, before it counts as
   live.  Until that is done, searches of the module BTFs fail and are
   retried rather than miss a module.  Module BTF thus works regardless
   of load order, including out-of-tree modules with a .BTF.base, whose
   sysfs reader waits for the relocation.

5. Trust.  Problem: the verifier treats the BTF as the description of
   the kernel's types, so a carrier from a different build must be
   refused even if vermagic lets it load.  Solution: record the name of
   the carrier and the size and SHA-256 of the BTF in the image, in a
   .BTF.link section that resolve_btfids fills in after the final link,
   and check the module against it when it loads.  The recorded size
   also lets /sys/kernel/btf/vmlinux report its final size before the
   load, which tooling expects.

6. Tooling.  Problem: pahole --btf_base for module BTF and external
   module builds read .BTF from the vmlinux ELF, while no boot image may
   carry it.  Solution: generate module BTF against vmlinux.unstripped,
   which keeps .BTF, and strip it from vmlinux, so that no image made
   from vmlinux carries it; the in-tree tools and samples that generate
   vmlinux.h look for vmlinux.unstripped first, and the pacman and rpm
   packages keep the BTF where it now is.

CONFIG_BPF_PRELOAD is made unavailable with =m: its preloaded programs
attach through the vmlinux BTF, so every bpffs mount, which systemd does
at boot, would load it and defeat the point.

Relation to the inline BTF series
---------------------------------

Alan's inline BTF series [2], now in bpf-next, adds inline function
information to BTF, which is even larger than the BTF itself, and its
cover letter leaves delivering that information on demand, through a
module and sysfs, to a follow-up.  When Alexei suggested taking the
same route for the vmlinux BTF [1], Alan pointed out where the
difficulty would lie [3].

However, that approach cannot be used directly for the vmlinux BTF,
because of what depends on the data:

1. Boot-time consumers.  Nothing needs inline information at boot, so
   loading it late only concerns the sysfs file.  The core vmlinux BTF
   is needed at boot by the verifier, by kfunc and struct_ops
   registrations, and by every module's BTF, and it is looked up from
   contexts that cannot wait for a module to load.  Therefore we need
   to make each of those work without it: the verifier fetches it only
   when a program brings kernel types in, the registrations are queued
   and replayed once it is parsed, and the BTF is loaded only where a
   request from user space starts.  That is most of this series.

2. Module BTF.  Module BTF is split against the vmlinux BTF, so a module
   loaded before it has nothing to be parsed against, and the module
   notifier cannot load btf_vmlinux itself.  Therefore we need to keep
   the module's BTF aside, create its sysfs file right away, and parse
   and register it once the vmlinux BTF arrives; the file of a module
   built against a distilled base (.BTF.base) serves it once relocated,
   the others serve the raw bytes as they are.  Alan named this as the
   hard part; it works here regardless of whether the module loads
   before or after the vmlinux BTF.

The inline information would get a carrier module of its own when that
follow-up comes, since a system may want the one but not the other, and
share the rest: Alan proposed the .BTF.link record and the resolve_btfids
option that fills it in [7] with both in mind, and this series uses them
as proposed (patch 10 is his).

Patches Structure
-----------------

Patch 1 (refactor, no functional change): btf_parse_module() takes the
vmlinux BTF as an argument and can adopt an existing copy of the module
BTF data instead of duplicating it.  Needed so that module BTF kept
aside at load time can be parsed later without moving the buffer the
sysfs file points at.

Patch 2 (refactor, no functional change): splits the kfunc, dtor kfunc
and struct_ops registration functions into "find the BTF for the owner"
and "add the registration to this BTF", so the second half can be
replayed on a queued registration.

Patch 3 (verifier): stops fetching the vmlinux BTF up front in
bpf_check() and fetches it where kernel types enter a program instead.
Adds bpf_peek_btf_vmlinux() for the helpers that run in program
context.  With =y the BTF is parsed at boot anyway, so this is
invisible there.

Patch 4 (carrier and loading): the runtime side of taking the vmlinux
BTF from the btf_vmlinux module: copy it out of the module in the BTF
module notifier after checking it against .BTF.link (struct btf_link:
the carrier's name, and the size and SHA-256 of the BTF);
bpf_get_btf_vmlinux() never loads it, bpf_load_btf_vmlinux() does, from
the bpf() system call entry (with the retry), BTF_GET_NEXT_ID, light
skeleton loaders and read() of /sys/kernel/btf/vmlinux, which has its
size known from boot; mmap() only maps it once it is loaded.  All under
IS_MODULE(CONFIG_DEBUG_INFO_BTF), so unreachable until patch 12.

Patch 5 (tracing, bpffs): loads the BTF at the start of the tracefs and
bpffs requests that need it (btf_ids, probe events with BTF arguments,
bpffs delegate options that name commands or types), before event_mutex
where that matters; the ftrace argument printer and bpffs show_options
only use it if it is there.

Patch 6 (vmlinux registrations): queues kfunc, dtor kfunc and struct_ops
registrations for vmlinux made from initcalls and applies them when the
BTF is parsed, before it is published, by pointer or by id.

Patch 7 (module BTF): keeps the BTF of modules loaded before the vmlinux
BTF, with their own queued registrations, and parses, registers and
publishes it when the vmlinux BTF arrives, before a module that is still
initializing counts as live; bpf_load_btf_vmlinux() returns once that is
done, and searches of module BTFs meanwhile fail and are retried.

Patch 8 (.BTF.base sysfs): gives such a module with a .BTF.base its
/sys/kernel/btf file from load, with a reader that waits for the
relocation.  Patches 4-8 are unreachable until patch 12.

Patch 9 (preparation, no functional change): the #ifdef, Makefile and
Kconfig checks of CONFIG_DEBUG_INFO_BTF that must hold for both =y and
=m use IS_ENABLED(), $(subst m,y,...) and DEBUG_INFO_BTF=n, in bpf,
tracing, netfilter, xfrm, Rust and modules.

Patch 10 (resolve_btfids, Alan's): a --btf_link
<section>:<module>:<raw BTF file> option for the final --patch_btfids
pass, which fills in the <section>.link record: the module name, and the
SHA-256 and size of the BTF, in the byte order of the ELF file.

Patch 11 (tools, samples): the in-tree tools and samples that generate
vmlinux.h from a build tree (bpftool, the bpf, hid and sched_ext
selftests, sched_ext, samples/bpf and hid, perf) look for
vmlinux.unstripped before vmlinux, which has no .BTF with =m; with =y
both have the same BTF.

Patch 12 (kbuild and Kconfig): makes CONFIG_DEBUG_INFO_BTF a tristate;
with =m links .BTF into vmlinux as a non-loadable section, has
resolve_btfids fill in .BTF.link after the final link, builds
btf_vmlinux.ko with the vmlinux .BTF as its payload, strips .BTF from
vmlinux (module BTF is generated against vmlinux.unstripped), keeps the
BTF in the pacman and rpm packages, excludes CONFIG_BPF_PRELOAD, and
documents the option.

Patches 1-3 and 9-11 are independently useful or neutral; 4-8 are dead
code until 12 flips the switch, which keeps each bisect step building
and behaving as before.

Testing
-------

Tested with 1 GiB of memory, same tree, =y against =m, both with
CONFIG_DEBUG_INFO_BTF_MODULES=y.  The on-demand behaviour is easy to
see by hand on an =m kernel:

  # lsmod | grep btf_vmlinux
      -> nothing: the BTF is not loaded at boot.
  # ls -la /sys/kernel/btf/vmlinux
      -> the file exists with its final size (from .BTF.link), although
         the BTF behind it is not loaded yet.
  # modprobe ext4 nf_conntrack
  # ls /sys/kernel/btf/
      -> ext4, nf_conntrack, ... appear immediately, although their
         BTF is only kept aside, not parsed: there is no vmlinux BTF to
         parse it against yet.
  # lsmod | grep btf_vmlinux
      -> still nothing: loading modules does not load the vmlinux BTF.
  # cat /sys/kernel/btf/ext4 > /dev/null
  # lsmod | grep btf_vmlinux
      -> still nothing: a module's BTF file is served from the raw copy,
         reading it does not need the vmlinux BTF.
  # grep VmallocUsed /proc/meminfo
      -> baseline.
  # cat /sys/kernel/btf/vmlinux > /dev/null
  # lsmod | grep btf_vmlinux
      -> btf_vmlinux ... [permanent]: the first use loaded it, and it
         cannot be unloaded.
  # grep VmallocUsed /proc/meminfo
      -> up by ~5.5 MB: the BTF copy, allocated only now.
  # bpftool btf list
      -> vmlinux and every loaded module now have BTF ids; the modules
         loaded before were parsed and registered on the way.

The same with an out-of-tree module (built with M=, so its BTF is split
against a distilled base, .BTF.base), on a fresh boot:

  # insmod btf_extmod.ko
  # ls -la /sys/kernel/btf/btf_extmod
      -> the file exists with its final size, although the BTF behind
         it is only valid once relocated against the vmlinux BTF.
  # lsmod | grep btf_vmlinux
      -> nothing: loading the module does not load the vmlinux BTF.
  # cat /sys/kernel/btf/btf_extmod > /dev/null
  # lsmod | grep btf_vmlinux
      -> btf_vmlinux ... [permanent]: reading this file loaded the
         vmlinux BTF, relocated the module's BTF and then returned it.
  # bpftool btf dump file /sys/kernel/btf/btf_extmod
      -> the module's own types, resolved against the vmlinux BTF.
  # rmmod btf_extmod
      -> unloads normally; its file goes away with it.

Results:

 - MemTotal is ~5.4 MB higher with =m while the BTF is unused, which is
   the size of the .BTF section.  Once the BTF is in use, MemFree is the
   same within run-to-run noise.
 - Each of these, as the first user of the BTF on a fresh boot, loads it
   and works: a kprobe program calling bpf_get_current_task_btf(), a
   syscall program calling kfuncs, a struct_ops map for
   tcp_congestion_ops, BTF and an array map with a kptr to task_struct,
   a light skeleton loader, read() of /sys/kernel/btf/vmlinux, libbpf's
   access to it (its mmap() fails until the BTF is loaded and it reads
   instead), BPF_BTF_GET_NEXT_ID, kprobe events with BTF arguments
   (argument names, $retval, $current), a tracepoint's btf_ids file, a
   bpffs mount with delegate options that name commands, and four such
   programs loaded at once.  With the in-tree bpftool and
   clang-built programs: bpftool btf dump and btf list, a socket filter
   with a global subprogram taking struct __sk_buff *, a raw_tp program
   calling bpf_snprintf_btf(), a CO-RE field read, and a program calling
   a kfunc of an out-of-tree module loaded before the vmlinux BTF.
 - A socket filter, and one that fails verification, do not load it;
   neither do bpffs mounts with delegate_*=any or hex masks,
   /proc/self/mountinfo with bpffs delegate options, the ftrace
   argument printer (also from sysrq-z), and BPF_BTF_GET_NEXT_ID
   without CAP_SYS_ADMIN.
 - Six BPF_BTF_GET_NEXT_ID users and a module tracepoint's btf_ids read,
   started at once right after modules were loaded, all see every
   module BTF.  A module whose kfunc registration was still queued,
   because the BTF arrived during its init, has a working kfunc once it
   is live.
 - Without btf_vmlinux.ko installed, the program, sysfs, BTF id, kptr,
   btf_ids and probe event cases fail or degrade as on a kernel without
   BTF, without delay (a btf_ids file tries once, not once per read);
   once it is installed, the next request loads it.
 - Modules loaded before the trigger (ext4, nf_conntrack, which
   registers kfuncs from its init, xfrm_interface) get BTF ids once the
   BTF is loaded; nf_nat loaded afterwards takes the usual path.
 - stat() of /sys/kernel/btf/vmlinux reports the final size before the
   load; fstat/read/mmap agree afterwards.
 - A carrier with one byte of .BTF changed is refused with -EINVAL.
   The .BTF.link of the kernel names btf_vmlinux and matches the BTF in
   btf_vmlinux.ko, size and SHA-256, on x86-64, on an i386 build and on
   an LLVM=1 build (clang and ld.lld 19).
   resolve_btfids --btf_link fills in the record of 32- and 64-bit,
   little- and big-endian objects (arm, m68k, parisc, arm64, riscv64,
   x86-64), also with two links in one run, and fails with a message
   on a missing or wrongly sized section, a missing or empty BTF file,
   or a module name that does not fit.
 - lockdep and kmemleak kernels are clean in all of the above.
 - =y and =n build and behave as before; =m without module BTF works;
   every patch builds on its own.  make localmodconfig keeps =m when
   btf_vmlinux is loaded; bpftool built from an =m tree takes its
   vmlinux.h from vmlinux.unstripped; the rpm spec refuses a debuginfo
   build whose find-debuginfo cannot keep .BTF.
 - The boot image on disk shrinks by the compressed BTF with =m
   (15.0 MB to 13.2 MB here).

Changes since v3 [6]:

 - Reworked where the BTF gets loaded (Alexei).  v3 loaded it from
   bpf_get_btf_vmlinux() and bpf_find_btf_id(), whose callers were not
   written for a function that waits for user space: the btf_ids file
   under event_mutex, the ftrace argument printer with interrupts off.
   Now neither of them loads; bpf_load_btf_vmlinux() does, only at the
   start of a request from user space, and the bpf() system call runs
   PROG_LOAD, MAP_CREATE or BTF_LOAD once more if it failed for want of
   the BTF (patch 4).  The tracefs and bpffs entry points moved to a
   patch of their own (patch 5).  The CO-RE candidate lookup is back to
   the upstream code: it fetched the BTF before cand_cache_mutex only
   because fetching could load it.
 - A light skeleton loader loads programs from within the running
   program, so it must not rely on them loading the BTF; bpf() loads it
   when user space loads the loader (patch 4).  This also covers
   bpf_btf_find_by_name_kind().
 - Patch 8: the sysfs reader of a kept .BTF.base module no longer loads
   the vmlinux BTF itself but has a work item do it: it holds the file's
   kernfs active reference, which MODULE_STATE_GOING drains with the
   module notifier chain held, and loading btf_vmlinux needs that chain
   (bpf-ci).
 - Patch 8: a kept .BTF.base module whose BTF then fails to parse no
   longer holds on to its raw BTF until it is unloaded; its reader never
   serves that data (Sashiko).
 - Patch 5: the tracefs btf_ids file loads the vmlinux BTF before taking
   event_mutex, which btf_vmlinux's trace module notifier takes; the
   ftrace function argument printer (func-args, funcgraph-args) never
   loads it, it runs from ftrace_dump() with interrupts disabled
   (bpf-ci).
 - Patch 4: with =m only an allocation failure of the vmlinux BTF parse
   is retried, a broken BTF is remembered as with =y; and
   /sys/kernel/btf/vmlinux serves the raw BTF even if it does not parse,
   as with =y (bpf-ci).
 - Patch 4: BPF_BTF_GET_NEXT_ID only loads the BTF for callers that may
   enumerate BTF ids (CAP_SYS_ADMIN).
 - Patch 4: mmap() of /sys/kernel/btf/vmlinux does not load the BTF: it
   runs under the caller's mmap_lock, a uprobe registration holds
   event_mutex while it takes the mmap_lock of every mm that maps the
   probed file, and the module load takes event_mutex.  It fails until
   the BTF is loaded; libbpf then falls back to read().
 - Patch 7: bpf_load_btf_vmlinux() returns only once the BTF of the
   modules loaded before it is registered, also to concurrent callers;
   until then, searches of the module BTFs by name fail and are retried,
   so a kptr to a module type is not taken for a local one and
   BPF_BTF_GET_NEXT_ID does not miss a module.
 - Patches 6-7: the registrations a module made while initializing are
   applied before it counts as live, in order, and applying loops, as
   for vmlinux; the vmlinux BTF and the kept module BTF have their ids
   reserved before their registrations are applied, and installed after.
   A module whose BTF cannot be kept for lack of memory fails to load,
   as with =y.
 - Patch 5: bpffs delegate_*=any and numeric masks do not load the BTF;
   a btf_ids file loads it on the first read only.
 - Patch 5: bpffs shows its delegate_* options (/proc/*/mountinfo, under
   namespace_sem) without loading the BTF.
 - Patches 6-7: the comment and changelog on the vmlinux registration
   queue's lock state the current reason (bpf-ci); when the vmlinux BTF
   fails to parse for good, the queued vmlinux registrations are freed
   and the queue closed, and the kept module BTF entries become dead,
   instead of waiting for ever.
 - Patches 4, 10, 12: the record of the BTF in the image is a section
   named .BTF.link, after .gnu_debuglink, laid out as Alan proposed for
   this series and the inline BTF follow-up alike [7] (struct btf_link:
   the carrier's module name, the SHA-256 and the size of the BTF).
   resolve_btfids fills it in when it patches .BTF_ids after the final
   link, with Alan's --btf_link option (patch 10), instead of gen-btf.sh;
   the kernel takes the name of the module to load from it.  The zeroed
   record is defined in C next to struct btf_link, so the first link
   needs no placeholder object, and resolve_btfids checks that the
   section has the size of the record (Alan).
 - Patch 12: changelog and btf.rst name what does not work with =m as
   with =y (BPF_PRELOAD, users before btf_vmlinux.ko can be loaded, the
   vmlinux BTF id) and list the requests that load the BTF (bpf-ci);
   they also name module loading policy, the check of the BTF of
   modules loaded before it, and tools that look for the BTF in memory.
 - Patch 11 (new): the in-tree tools and samples take the vmlinux BTF
   from vmlinux.unstripped first.  Patch 12: the pacman debug package
   ships vmlinux.unstripped with =m, the rpm spec refuses a debuginfo
   build that would strip the .BTF of btf_vmlinux.ko,
   make localmodconfig maps btf_vmlinux.ko to its option, and the build
   checks that the carrier got a BTF.
 - Rebased onto current bpf-next.

Changes since v2 [5]:

 - Rebased onto current bpf-next; v2 no longer applied there.
 - Patch 8: the RUST and GENDWARFKSYMS pahole restrictions, written as
   "depends on !DEBUG_INFO_BTF", now say DEBUG_INFO_BTF=n, so they
   still hold with =m (found while checking the Sashiko question on
   bool options depending on DEBUG_INFO_BTF, which Kconfig handles:
   a bool whose dependency is m can still be y).

Changes since v1 [4]:

 - Split for review: v1 patch 5 is now patches 5-7 (vmlinux
   registrations, module BTF, .BTF.base sysfs), v1 patch 6 is now
   patches 8-9 (preparation of the existing checks, the switch).
 - Fetch sites added for the program context type table
   (bpf_ctx_convert: global subprograms taking the context, ctx access
   of tracing/EXT programs), for bpf_snprintf_btf()/bpf_seq_printf_btf()
   and for CO-RE candidate lookup, which now fetches before taking
   cand_cache_mutex (Sashiko, bpf-ci, Jiri).  Without
   CONFIG_DEBUG_INFO_BTF the new helper check is skipped, so nothing
   changes there.
 - Module BTF is published only after its deferred registrations are
   applied; only modules past MODULE_STATE_LIVE are replayed, with the
   module pinned; a module still in init has its queue applied at LIVE.
   Fixes the concurrent registration and COMING-module lifetime issues
   (Sashiko, bpf-ci).
 - The sysfs reader of a module with .BTF.base waits for the module's
   BTF to be relocated instead of serving the raw data (bpf-ci); sysfs
   files are no longer removed from the deferred parse path, and
   MODULE_STATE_GOING removes them outside btf_module_mutex.
 - A module whose deferred BTF fails to parse or get an id keeps the
   buffer its sysfs file serves; nothing is freed under a reader
   (Sashiko).
 - The vmlinux registration queue has its own mutex; no lock is taken
   under btf_vmlinux_lock that leads back to it (bpf-ci).
 - A failed parse is not cached with =m (Sashiko).
 - Only the carrier depends on vmlinux in Makefile.modfinal; POSIX dd
   instead of head -c (Sashiko).
 - .BTF stripped from vmlinux with =m, so no boot image carries it,
   also where the image is an ELF copy of vmlinux; module BTF is
   generated against vmlinux.unstripped (Alan).
 - Kconfig help: initramfs note, module BTF accounting (Alan).
 - btf_struct_ops_add() renamed btf_struct_ops_register() (bpf-ci);
   its stub only defined where used.
 - Tests with bpftool/libbpf userspace and an out-of-tree .BTF.base
   module added (Alan).

[1] https://lore.kernel.org/all/20260917000201.25581-2-wanjay@amazon.com/
[2] https://lore.kernel.org/bpf/20260916074118.1007116-1-alan.maguire@oracle.com/
[3] https://lore.kernel.org/all/33592fca-88a8-44aa-8d94-40e1e604554e@oracle.com/
[4] https://lore.kernel.org/bpf/20260923053948.30617-1-wanjay@amazon.com/
[5] https://lore.kernel.org/bpf/20260925211314.5118-1-wanjay@amazon.com/
[6] https://lore.kernel.org/bpf/20260925224229.1850-1-wanjay@amazon.com/
[7] https://lore.kernel.org/bpf/dbbb6cde-9889-4445-abf6-0486c380925d@oracle.com/

Alan Maguire (1):
  resolve_btfids: add --btf_link to fill in .BTF.link records

Jay Wang (11):
  bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the
    data
  bpf: split the kfunc, dtor kfunc and struct_ops registration bodies
  bpf: fetch the vmlinux BTF where kernel types enter a program
  bpf: take the vmlinux BTF from the btf_vmlinux module
  bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests
    start
  bpf: defer vmlinux kfunc and struct_ops registrations
  bpf: keep module BTF until the vmlinux BTF is available
  bpf: expose deferred .BTF.base module BTF in sysfs from module load
  bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate
  tools, samples: take the vmlinux BTF from vmlinux.unstripped first
  kbuild, bpf: allow building the vmlinux BTF as a module

 Documentation/bpf/btf.rst                  |   68 ++
 Makefile                                   |    8 +-
 include/asm-generic/vmlinux.lds.h          |   32 +-
 include/linux/bpf.h                        |   15 +
 include/linux/btf.h                        |   11 +
 include/linux/btf_ids.h                    |    2 +-
 include/linux/compiler_types.h             |    2 +-
 include/linux/module.h                     |    2 +-
 include/trace/trace_events.h               |    2 +-
 init/Kconfig                               |    2 +-
 kernel/bpf/Makefile                        |    6 +-
 kernel/bpf/bpf_struct_ops.c                |    3 +-
 kernel/bpf/btf.c                           | 1085 ++++++++++++++++++--
 kernel/bpf/btf_vmlinux.c                   |   23 +
 kernel/bpf/inode.c                         |   44 +-
 kernel/bpf/preload/Kconfig                 |    4 +
 kernel/bpf/syscall.c                       |   51 +-
 kernel/bpf/sysfs_btf.c                     |   97 +-
 kernel/bpf/verifier.c                      |  172 +++-
 kernel/module/Kconfig                      |    2 +-
 kernel/module/main.c                       |    4 +-
 kernel/trace/bpf_trace.c                   |    3 +-
 kernel/trace/trace_events.c                |   10 +
 kernel/trace/trace_output.c                |    7 +
 kernel/trace/trace_probe.c                 |   16 +
 kernel/trace/trace_syscalls.c              |    6 +-
 lib/Kconfig.debug                          |   30 +-
 net/netfilter/Makefile                     |    6 +-
 net/xfrm/Makefile                          |    4 +-
 samples/bpf/Makefile                       |    6 +-
 samples/hid/Makefile                       |    6 +-
 scripts/Makefile.modfinal                  |   28 +-
 scripts/Makefile.vmlinux                   |    5 +
 scripts/gen-btf.sh                         |   53 +-
 scripts/link-vmlinux.sh                    |   25 +-
 scripts/package/PKGBUILD                   |    7 +
 scripts/package/kernel.spec                |    4 +
 scripts/package/mkspec                     |    7 +
 tools/bpf/bpftool/Makefile                 |    6 +-
 tools/bpf/resolve_btfids/main.c            |  217 +++-
 tools/perf/bpf_skel.mak                    |    8 +-
 tools/sched_ext/Makefile                   |    6 +-
 tools/testing/selftests/bpf/Makefile       |    6 +-
 tools/testing/selftests/hid/Makefile       |    6 +-
 tools/testing/selftests/sched_ext/Makefile |    6 +-
 45 files changed, 1936 insertions(+), 177 deletions(-)
 create mode 100644 kernel/bpf/btf_vmlinux.c


base-commit: b5a4aa31abd6fe90009b63e35dc18c67d041ec0c
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
                   ` (11 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Make btf_parse_module() take the vmlinux BTF as an argument instead of
fetching it with bpf_get_btf_vmlinux(), and add a data_owned flag: when
set, the passed .BTF data is an already kvmalloc()ed copy that the new
btf takes ownership of on success (on failure the caller keeps it).

Factor the sysfs file creation out of the module notifier into
btf_module_sysfs_add() and the teardown into btf_module_free(), and set
btf_mod->module right after the allocation rather than under the mutex.

No functional change.  This prepares for CONFIG_DEBUG_INFO_BTF=m, where a
module can be loaded before the vmlinux BTF is available: its .BTF is
then copied and exposed in sysfs first and parsed later, at which point
the parser must take the copy as is so that the sysfs file keeps
pointing at valid data.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/btf.c | 105 +++++++++++++++++++++++++++++------------------
 1 file changed, 64 insertions(+), 41 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 8cc17a1cd25c..c27b929f84f5 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -6911,16 +6911,20 @@ __u32 btf_relocate_id(const struct btf *btf, __u32 id)
 
 #ifdef CONFIG_DEBUG_INFO_BTF_MODULES
 
-static struct btf *btf_parse_module(const char *module_name, const void *data,
-				    unsigned int data_size, void *base_data,
-				    unsigned int base_data_size)
+/*
+ * Parse split module BTF against @vmlinux_btf.  @data is the module's .BTF
+ * section; if @data_owned, it is an already kvmalloc()ed copy that the new
+ * btf takes ownership of on success (on failure the caller keeps it).
+ */
+static struct btf *btf_parse_module(const char *module_name, struct btf *vmlinux_btf,
+				    void *data, unsigned int data_size, bool data_owned,
+				    void *base_data, unsigned int base_data_size)
 {
-	struct btf *btf = NULL, *vmlinux_btf, *base_btf = NULL;
+	struct btf *btf = NULL, *base_btf = NULL;
 	struct btf_verifier_env *env = NULL;
 	struct bpf_verifier_log *log;
 	int err = 0;
 
-	vmlinux_btf = bpf_get_btf_vmlinux();
 	if (IS_ERR(vmlinux_btf))
 		return vmlinux_btf;
 	if (!vmlinux_btf)
@@ -6957,7 +6961,10 @@ static struct btf *btf_parse_module(const char *module_name, const void *data,
 	btf->named_start_id = 0;
 	strscpy(btf->name, module_name);
 
-	btf->data = kvmemdup(data, data_size, GFP_KERNEL | __GFP_NOWARN);
+	if (data_owned)
+		btf->data = data;
+	else
+		btf->data = kvmemdup(data, data_size, GFP_KERNEL | __GFP_NOWARN);
 	if (!btf->data) {
 		err = -ENOMEM;
 		goto errout;
@@ -7000,7 +7007,8 @@ static struct btf *btf_parse_module(const char *module_name, const void *data,
 	if (!IS_ERR(base_btf) && base_btf != vmlinux_btf)
 		btf_free(base_btf);
 	if (btf) {
-		kvfree(btf->data);
+		if (!data_owned)
+			kvfree(btf->data);
 		kvfree(btf->types);
 		kfree(btf);
 	}
@@ -9005,6 +9013,48 @@ static DEFINE_MUTEX(btf_module_mutex);
 
 static void purge_cand_cache(struct btf *btf);
 
+static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
+				void *data, size_t data_size)
+{
+	struct bin_attribute *attr;
+	int err;
+
+	if (!IS_ENABLED(CONFIG_SYSFS))
+		return 0;
+
+	attr = kzalloc_obj(*attr);
+	if (!attr)
+		return -ENOMEM;
+
+	sysfs_bin_attr_init(attr);
+	attr->attr.name = name;
+	attr->attr.mode = 0444;
+	attr->size = data_size;
+	attr->private = data;
+	attr->read = sysfs_bin_attr_simple_read;
+
+	err = sysfs_create_bin_file(btf_kobj, attr);
+	if (err) {
+		pr_warn("failed to register module [%s] BTF in sysfs: %d\n",
+			name, err);
+		kfree(attr);
+		return err;
+	}
+
+	btf_mod->sysfs_attr = attr;
+	return 0;
+}
+
+static void btf_module_free(struct btf_module *btf_mod)
+{
+	if (btf_mod->sysfs_attr)
+		sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
+	purge_cand_cache(btf_mod->btf);
+	btf_put(btf_mod->btf);
+	kfree(btf_mod->sysfs_attr);
+	kfree(btf_mod);
+}
+
 static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			     void *module)
 {
@@ -9025,7 +9075,10 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			err = -ENOMEM;
 			goto out;
 		}
-		btf = btf_parse_module(mod->name, mod->btf_data, mod->btf_data_size,
+		btf_mod->module = module;
+
+		btf = btf_parse_module(mod->name, bpf_get_btf_vmlinux(),
+				       mod->btf_data, mod->btf_data_size, false,
 				       mod->btf_base_data, mod->btf_base_data_size);
 		if (IS_ERR(btf)) {
 			kfree(btf_mod);
@@ -9047,37 +9100,12 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 
 		purge_cand_cache(NULL);
 		mutex_lock(&btf_module_mutex);
-		btf_mod->module = module;
 		btf_mod->btf = btf;
 		list_add(&btf_mod->list, &btf_modules);
 		mutex_unlock(&btf_module_mutex);
 
-		if (IS_ENABLED(CONFIG_SYSFS)) {
-			struct bin_attribute *attr;
-
-			attr = kzalloc_obj(*attr);
-			if (!attr)
-				goto out;
-
-			sysfs_bin_attr_init(attr);
-			attr->attr.name = btf->name;
-			attr->attr.mode = 0444;
-			attr->size = btf->data_size;
-			attr->private = btf->data;
-			attr->read = sysfs_bin_attr_simple_read;
-
-			err = sysfs_create_bin_file(btf_kobj, attr);
-			if (err) {
-				pr_warn("failed to register module [%s] BTF in sysfs: %d\n",
-					mod->name, err);
-				kfree(attr);
-				err = 0;
-				goto out;
-			}
-
-			btf_mod->sysfs_attr = attr;
-		}
-
+		/* not fatal, the module BTF is usable without the sysfs file */
+		btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
 		break;
 	case MODULE_STATE_LIVE:
 		mutex_lock(&btf_module_mutex);
@@ -9104,12 +9132,7 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			 */
 			btf_free_id(btf_mod->btf);
 			list_del(&btf_mod->list);
-			if (btf_mod->sysfs_attr)
-				sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
-			purge_cand_cache(btf_mod->btf);
-			btf_put(btf_mod->btf);
-			kfree(btf_mod->sysfs_attr);
-			kfree(btf_mod);
+			btf_module_free(btf_mod);
 			break;
 		}
 		mutex_unlock(&btf_module_mutex);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
                   ` (10 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Split __register_btf_kfunc_id_set(), register_btf_id_dtor_kfuncs() and
__register_bpf_struct_ops() into the part that looks up the BTF for the
owner and the part that adds the registration to a given BTF:
btf_kfunc_id_set_add(), btf_dtor_kfuncs_add() and
btf_struct_ops_register().

In is_valid_value_type(), look up bpf_struct_ops_common_value in the btf
the function was given rather than in the btf_vmlinux global.  The id is
a vmlinux id and a module BTF resolves it through its base, so the result
is the same; the function already uses the passed btf for every other
lookup.

No functional change.  With CONFIG_DEBUG_INFO_BTF=m, registrations made
from initcalls before the vmlinux BTF is available are queued and applied
later by the BTF parsing code, which needs the add-to-this-btf half on
its own; the struct_ops ones are applied before the parsed vmlinux BTF is
published, i.e. while btf_vmlinux is still NULL.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/bpf_struct_ops.c |  3 +-
 kernel/bpf/btf.c            | 91 ++++++++++++++++++++++---------------
 2 files changed, 57 insertions(+), 37 deletions(-)

diff --git a/kernel/bpf/bpf_struct_ops.c b/kernel/bpf/bpf_struct_ops.c
index 1178acd72296..bf3004908d15 100644
--- a/kernel/bpf/bpf_struct_ops.c
+++ b/kernel/bpf/bpf_struct_ops.c
@@ -103,7 +103,8 @@ static bool is_valid_value_type(struct btf *btf, s32 value_id,
 	}
 	member = btf_type_member(vt);
 	mt = btf_type_by_id(btf, member->type);
-	common_value_type = btf_type_by_id(btf_vmlinux,
+	/* a vmlinux id resolves through the base BTF of a module BTF too */
+	common_value_type = btf_type_by_id(btf,
 					   st_ops_ids[IDX_ST_OPS_COMMON_VALUE_ID]);
 	if (mt != common_value_type) {
 		pr_warn("The first member of %s should be bpf_struct_ops_common_value\n",
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index c27b929f84f5..dfd1af8c2ac5 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -9708,11 +9708,26 @@ u32 *btf_kfunc_is_modify_return(const struct btf *btf, u32 kfunc_btf_id,
 	return btf_kfunc_id_set_contains(btf, BTF_KFUNC_HOOK_FMODRET, kfunc_btf_id);
 }
 
+static int btf_kfunc_id_set_add(struct btf *btf, enum btf_kfunc_hook hook,
+				const struct btf_kfunc_id_set *kset)
+{
+	int ret, i;
+
+	for (i = 0; i < kset->set->cnt; i++) {
+		ret = btf_check_kfunc_protos(btf, btf_relocate_id(btf, kset->set->pairs[i].id),
+					     kset->set->pairs[i].flags);
+		if (ret)
+			return ret;
+	}
+
+	return btf_populate_kfunc_set(btf, hook, kset);
+}
+
 static int __register_btf_kfunc_id_set(enum btf_kfunc_hook hook,
 				       const struct btf_kfunc_id_set *kset)
 {
 	struct btf *btf;
-	int ret, i;
+	int ret;
 
 	btf = btf_get_module_btf(kset->owner);
 	if (!btf)
@@ -9720,16 +9735,7 @@ static int __register_btf_kfunc_id_set(enum btf_kfunc_hook hook,
 	if (IS_ERR(btf))
 		return PTR_ERR(btf);
 
-	for (i = 0; i < kset->set->cnt; i++) {
-		ret = btf_check_kfunc_protos(btf, btf_relocate_id(btf, kset->set->pairs[i].id),
-					     kset->set->pairs[i].flags);
-		if (ret)
-			goto err_out;
-	}
-
-	ret = btf_populate_kfunc_set(btf, hook, kset);
-
-err_out:
+	ret = btf_kfunc_id_set_add(btf, hook, kset);
 	btf_put(btf);
 	return ret;
 }
@@ -9821,21 +9827,13 @@ static int btf_check_dtor_kfuncs(struct btf *btf, const struct btf_id_dtor_kfunc
 	return 0;
 }
 
-/* This function must be invoked only from initcalls/module init functions */
-int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32 add_cnt,
-				struct module *owner)
+static int btf_dtor_kfuncs_add(struct btf *btf, const struct btf_id_dtor_kfunc *dtors,
+			       u32 add_cnt)
 {
 	struct btf_id_dtor_kfunc_tab *tab;
-	struct btf *btf;
 	u32 tab_cnt, i;
 	int ret;
 
-	btf = btf_get_module_btf(owner);
-	if (!btf)
-		return check_btf_kconfigs(owner, "dtor kfuncs");
-	if (IS_ERR(btf))
-		return PTR_ERR(btf);
-
 	if (add_cnt >= BTF_DTOR_KFUNC_MAX_CNT) {
 		pr_err("cannot register more than %d kfunc destructors\n", BTF_DTOR_KFUNC_MAX_CNT);
 		ret = -E2BIG;
@@ -9892,6 +9890,23 @@ int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32 add_c
 end:
 	if (ret)
 		btf_free_dtor_kfunc_tab(btf);
+	return ret;
+}
+
+/* This function must be invoked only from initcalls/module init functions */
+int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32 add_cnt,
+				struct module *owner)
+{
+	struct btf *btf;
+	int ret;
+
+	btf = btf_get_module_btf(owner);
+	if (!btf)
+		return check_btf_kconfigs(owner, "dtor kfuncs");
+	if (IS_ERR(btf))
+		return PTR_ERR(btf);
+
+	ret = btf_dtor_kfuncs_add(btf, dtors, add_cnt);
 	btf_put(btf);
 	return ret;
 }
@@ -10530,32 +10545,36 @@ bpf_struct_ops_find(struct btf *btf, u32 type_id)
 	return NULL;
 }
 
-int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops)
+static int btf_struct_ops_register(struct btf *btf, struct bpf_struct_ops *st_ops)
 {
 	struct bpf_verifier_log *log;
-	struct btf *btf;
-	int err = 0;
-
-	btf = btf_get_module_btf(st_ops->owner);
-	if (!btf)
-		return check_btf_kconfigs(st_ops->owner, "struct_ops");
-	if (IS_ERR(btf))
-		return PTR_ERR(btf);
+	int err;
 
 	log = kzalloc_obj(*log, GFP_KERNEL | __GFP_NOWARN);
-	if (!log) {
-		err = -ENOMEM;
-		goto errout;
-	}
+	if (!log)
+		return -ENOMEM;
 
 	log->level = BPF_LOG_KERNEL;
 
 	err = btf_add_struct_ops(btf, st_ops, log);
 
-errout:
 	kfree(log);
-	btf_put(btf);
+	return err;
+}
 
+int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops)
+{
+	struct btf *btf;
+	int err;
+
+	btf = btf_get_module_btf(st_ops->owner);
+	if (!btf)
+		return check_btf_kconfigs(st_ops->owner, "struct_ops");
+	if (IS_ERR(btf))
+		return PTR_ERR(btf);
+
+	err = btf_struct_ops_register(btf, st_ops);
+	btf_put(btf);
 	return err;
 }
 EXPORT_SYMBOL_GPL(__register_bpf_struct_ops);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
                   ` (9 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

bpf_check() fetches the vmlinux BTF up front for every program, whether
the program uses kernel types or not.  With the upcoming
CONFIG_DEBUG_INFO_BTF=m the BTF is a module that is loaded on demand,
and since systemd loads socket filters at boot, a program that needs the
BTF only because bpf_check() asked for it would pull it in on every
system, whether anything uses BTF or not.

Stop fetching up front and fetch at the points where kernel types enter
the verifier state instead:

 - bpf_add_kfunc_call(), for the first kfunc call of a program;
 - check_pseudo_btf_id(), for ldimm64 of a kernel variable;
 - check_ptr_to_map_access(), for accessing a map pointer's fields;
 - check_helper_call(), when the helper's prototype takes or returns a
   PTR_TO_BTF_ID, or is bpf_snprintf_btf()/bpf_seq_printf_btf(), which
   take the kernel type id inside a struct btf_ptr instead
   (helper_uses_vmlinux_btf());
 - the program context type table, bpf_ctx_convert, which
   btf_parse_vmlinux() fills in: global subprograms taking the context
   (btf_prepare_func_args()) and context access of tracing and EXT
   programs (btf_translate_to_vmlinux()) go through it, so its readers
   fetch the BTF (bpf_ctx_convert_type()).

Together with the existing fetches in bpf_prog_load() for attach_btf, in
struct_ops map creation and in CO-RE candidate lookup, every way kernel
types enter a program goes through one of these sites.  A program that
uses none of them, such as a socket filter, no longer touches the vmlinux
BTF.

The two callers that used the result of bpf_get_btf_vmlinux() without
checking it, btf_prepare_func_args() and btf_check_kfunc_name(), now do.

bpf_snprintf_btf() and bpf_seq_printf_btf() run in program context, where
the verifier has made sure the BTF is there.  Add bpf_peek_btf_vmlinux(),
which returns the parsed vmlinux BTF or NULL without parsing anything,
and use it there and for the "malformed BTF" check in bpf_check(); the
helpers fail with -EINVAL if the BTF is not parsed, as they do on a
kernel without BTF.

With CONFIG_DEBUG_INFO_BTF=y the vmlinux BTF is parsed at boot by the
first kfunc registration, and without BTF the new helper check is
skipped, so nothing changes for either.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 include/linux/bpf.h      |  1 +
 kernel/bpf/btf.c         | 34 ++++++++++++++++++---
 kernel/bpf/verifier.c    | 66 +++++++++++++++++++++++++++++++++++-----
 kernel/trace/bpf_trace.c |  3 +-
 4 files changed, 91 insertions(+), 13 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..e46a14809dd4 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -3183,6 +3183,7 @@ static inline s32 bpf_call_args_imm(s16 idx)
 #endif
 
 struct btf *bpf_get_btf_vmlinux(void);
+struct btf *bpf_peek_btf_vmlinux(void);
 
 /* Map specifics */
 struct xdp_frame;
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index dfd1af8c2ac5..9896a30eeac4 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -6521,12 +6521,25 @@ static u8 bpf_ctx_convert_map[] = {
 #undef BPF_MAP_TYPE
 #undef BPF_LINK_TYPE
 
+/*
+ * bpf_ctx_convert.t is filled in by btf_parse_vmlinux(), which bpf_check()
+ * does not trigger up front any more.  The program context types are kernel
+ * types too: fetch the vmlinux BTF here, as at the other points where kernel
+ * types enter a program.
+ */
+static const struct btf_type *bpf_ctx_convert_type(void)
+{
+	if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux()))
+		return NULL;
+	return bpf_ctx_convert.t;
+}
+
 static const struct btf_type *find_canonical_prog_ctx_type(enum bpf_prog_type prog_type)
 {
 	const struct btf_type *conv_struct;
 	const struct btf_member *ctx_type;
 
-	conv_struct = bpf_ctx_convert.t;
+	conv_struct = bpf_ctx_convert_type();
 	if (!conv_struct)
 		return NULL;
 	/* prog_type is valid bpf program type. No need for bounds check. */
@@ -6542,7 +6555,7 @@ static int find_kern_ctx_type_id(enum bpf_prog_type prog_type)
 	const struct btf_type *conv_struct;
 	const struct btf_member *ctx_type;
 
-	conv_struct = bpf_ctx_convert.t;
+	conv_struct = bpf_ctx_convert_type();
 	if (!conv_struct)
 		return -EFAULT;
 	/* prog_type is valid bpf program type. No need for bounds check. */
@@ -6801,7 +6814,11 @@ int get_kern_ctx_btf_id(struct bpf_verifier_log *log, enum bpf_prog_type prog_ty
 	const struct btf_type *kctx_type;
 	u32 kctx_type_id;
 
-	conv_struct = bpf_ctx_convert.t;
+	conv_struct = bpf_ctx_convert_type();
+	if (!conv_struct) {
+		bpf_log(log, "btf_vmlinux is malformed\n");
+		return -EINVAL;
+	}
 	/* get member for kernel ctx type */
 	kctx_member = btf_type_member(conv_struct) + bpf_ctx_convert_map[prog_type] * 2 + 1;
 	kctx_type_id = kctx_member->type;
@@ -8630,7 +8647,10 @@ int btf_prepare_func_args(struct bpf_verifier_env *env, int subprog)
 			if (kern_type_id < 0)
 				return kern_type_id;
 
+			/* present: btf_get_ptr_to_btf_id() found the candidate in it */
 			vmlinux_btf = bpf_get_btf_vmlinux();
+			if (IS_ERR_OR_NULL(vmlinux_btf))
+				return -EINVAL;
 			ref_t = btf_type_by_id(vmlinux_btf, kern_type_id);
 			if (!btf_type_is_struct(ref_t)) {
 				tname = __btf_name_by_offset(vmlinux_btf, t->name_off);
@@ -9368,12 +9388,18 @@ static int btf_check_kfunc_name(struct btf *btf, const char *func_name, u32 kind
 #ifdef CONFIG_DEBUG_INFO_BTF_MODULES
 	struct btf_module *btf_mod, *tmp;
 #endif
+	struct btf *vmlinux_btf;
 	s32 id;
 
 	if (!btf_is_module(btf))
 		return 0;
 
-	id = btf_find_by_name_kind(bpf_get_btf_vmlinux(), func_name, kind);
+	/* a module BTF only exists once the vmlinux BTF is parsed */
+	vmlinux_btf = bpf_get_btf_vmlinux();
+	if (IS_ERR_OR_NULL(vmlinux_btf))
+		return -EINVAL;
+
+	id = btf_find_by_name_kind(vmlinux_btf, func_name, kind);
 	if (id >= 0) {
 		pr_err("kfunc %s (id: %d) is already present in vmlinux.\n",
 		       func_name, id);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index b840b3eb9b22..f375c5dad4b5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -2882,7 +2882,8 @@ int bpf_add_kfunc_call(struct bpf_verifier_env *env, u32 func_id, u16 offset)
 	tab = prog_aux->kfunc_tab;
 	btf_tab = prog_aux->kfunc_btf_tab;
 	if (!tab) {
-		if (!btf_vmlinux) {
+		/* kernel types enter the program here, see bpf_check() */
+		if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) {
 			verbose(env, "calling kernel function is not supported without CONFIG_DEBUG_INFO_BTF\n");
 			return -ENOTSUPP;
 		}
@@ -6530,7 +6531,8 @@ static int check_ptr_to_map_access(struct bpf_verifier_env *env,
 	u32 btf_id;
 	int ret;
 
-	if (!btf_vmlinux) {
+	/* kernel types enter the program here, see bpf_check() */
+	if (IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) {
 		verbose(env, "map_ptr access not supported without CONFIG_DEBUG_INFO_BTF\n");
 		return -ENOTSUPP;
 	}
@@ -12060,6 +12062,24 @@ static int release_reg(struct bpf_verifier_env *env, struct bpf_reg_state *reg,
 	return err;
 }
 
+/* Does calling helper @func_id bring kernel BTF types into the program? */
+static bool helper_uses_vmlinux_btf(enum bpf_func_id func_id,
+				    const struct bpf_func_proto *fn)
+{
+	int i;
+
+	/* these take the kernel type id in a struct btf_ptr, not in a register */
+	if (func_id == BPF_FUNC_snprintf_btf || func_id == BPF_FUNC_seq_printf_btf)
+		return true;
+	if (base_type(fn->ret_type) == RET_PTR_TO_BTF_ID)
+		return true;
+	for (i = 0; i < MAX_BPF_FUNC_ARGS; i++) {
+		if (base_type(fn->arg_type[i]) == ARG_PTR_TO_BTF_ID)
+			return true;
+	}
+	return false;
+}
+
 static int check_helper_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 			     int *insn_idx_p)
 {
@@ -12129,6 +12149,18 @@ static int check_helper_call(struct bpf_verifier_env *env, struct bpf_insn *insn
 		return err;
 	}
 
+	/*
+	 * Helpers that take or return kernel BTF pointers bring kernel types
+	 * into the program, see bpf_check().  Without CONFIG_DEBUG_INFO_BTF
+	 * they keep failing as they always did.
+	 */
+	if (IS_ENABLED(CONFIG_DEBUG_INFO_BTF) && helper_uses_vmlinux_btf(func_id, fn) &&
+	    IS_ERR_OR_NULL(bpf_get_btf_vmlinux())) {
+		verbose(env, "helper %s#%d is not supported without vmlinux BTF\n",
+			func_id_name(func_id), func_id);
+		return -ENOTSUPP;
+	}
+
 	if (fn->might_sleep && !in_sleepable_context(env)) {
 		const char *suggestion;
 
@@ -19812,12 +19844,13 @@ static int check_pseudo_btf_id(struct bpf_verifier_env *env,
 			return -EINVAL;
 		}
 	} else {
-		if (!btf_vmlinux) {
+		/* kernel types enter the program here, see bpf_check() */
+		btf = bpf_get_btf_vmlinux();
+		if (IS_ERR_OR_NULL(btf)) {
 			verbose(env, "kernel is missing BTF, make sure CONFIG_DEBUG_INFO_BTF=y is specified in Kconfig.\n");
 			return -EINVAL;
 		}
-		btf_get(btf_vmlinux);
-		btf = btf_vmlinux;
+		btf_get(btf);
 	}
 
 	err = __check_pseudo_btf_id(env, insn, aux, btf);
@@ -21871,6 +21904,17 @@ struct btf *bpf_get_btf_vmlinux(void)
 	return btf;
 }
 
+/*
+ * The vmlinux BTF if it has been parsed already, else NULL.  Unlike
+ * bpf_get_btf_vmlinux() this never parses anything: for a running BPF
+ * program, and for code that only uses the BTF if it happens to be there.
+ */
+struct btf *bpf_peek_btf_vmlinux(void)
+{
+	/* Pairs with the smp_store_release() in bpf_get_btf_vmlinux() */
+	return smp_load_acquire(&btf_vmlinux);
+}
+
 /*
  * The add_fd_from_fd_array() is executed only if fd_array_cnt is non-zero. In
  * this case expect that every file descriptor in the array is either a map or
@@ -22470,7 +22514,13 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret)
 		goto err_prep;
 
-	bpf_get_btf_vmlinux();
+	/*
+	 * The vmlinux BTF is not fetched up front, only at the points where
+	 * kernel types enter the program (attach_btf, kfuncs, ksyms, map
+	 * pointers, BTF-typed helpers, the context type table): with
+	 * CONFIG_DEBUG_INFO_BTF=m it is a module, which a program that uses
+	 * no kernel types must not need.
+	 */
 
 	/* Serialize verification of unprivileged programs. */
 	if (!is_priv)
@@ -22491,10 +22541,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 
 	mark_verifier_state_clean(env);
 
-	if (IS_ERR(btf_vmlinux)) {
+	if (IS_ERR(bpf_peek_btf_vmlinux())) {
 		/* Either gcc or pahole or kernel are broken. */
 		verbose(env, "in-kernel BTF is malformed\n");
-		ret = PTR_ERR(btf_vmlinux);
+		ret = PTR_ERR(bpf_peek_btf_vmlinux());
 		goto skip_full_check;
 	}
 
diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c
index 195f78db9bda..c022b2877f0b 100644
--- a/kernel/trace/bpf_trace.c
+++ b/kernel/trace/bpf_trace.c
@@ -1015,7 +1015,8 @@ static int bpf_btf_printf_prepare(struct btf_ptr *ptr, u32 btf_ptr_size,
 	if (btf_ptr_size != sizeof(struct btf_ptr))
 		return -EINVAL;
 
-	*btf = bpf_get_btf_vmlinux();
+	/* Called from a running program: only use the BTF if it is parsed. */
+	*btf = bpf_peek_btf_vmlinux();
 
 	if (IS_ERR_OR_NULL(*btf))
 		return IS_ERR(*btf) ? PTR_ERR(*btf) : -EINVAL;
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (2 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 23:45   ` bot+bpf-ci
  2026-10-02 11:48   ` Alexei Starovoitov
  2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
                   ` (8 subsequent siblings)
  12 siblings, 2 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Add the runtime side of delivering the vmlinux BTF as a module: with
CONFIG_DEBUG_INFO_BTF=m the BTF is carried by a module named btf_vmlinux
and installed by the BTF module notifier when it loads.  Nothing in this
patch is reachable yet: CONFIG_DEBUG_INFO_BTF is still a bool and every
new path is under IS_MODULE(CONFIG_DEBUG_INFO_BTF); the kbuild side and
the Kconfig change follow.

With CONFIG_DEBUG_INFO_BTF=y the vmlinux BTF, 5.4 MiB on x86-64 with a
distribution config, is part of the kernel image and resident from boot
whether anything uses it or not.  Most systems never do.  Carrying it in
a module that is loaded on first use makes the memory a cost of using
BTF rather than of having a kernel that supports it.

btf_vmlinux_data() hands out the raw vmlinux BTF: from __start_BTF with
=y, or with =m from a vmalloc_user() copy that the notifier makes when
the btf_vmlinux module loads.  The kernel is linked with a record of the
BTF the module carries, .BTF.link (struct btf_link, zeroed here and
filled in by resolve_btfids at the end of the link in a later patch),
much as .gnu_debuglink describes a separate debug info file: the name of
the module, and the size and SHA-256 of the BTF.  The verifier trusts
the BTF as the description of this kernel's types, so the notifier only
accepts a payload of that size and SHA-256 from the module of that name.
A carrier from another build is refused with -EINVAL even if vermagic
lets it load, and a second carrier is ignored.  The copy is never freed:
as with =y, the BTF stays for the lifetime of the kernel, and the
carrier has no exit.

Loading the module waits for user space (modprobe), and the module's
notifiers take locks of their own, event_mutex among them.  The callers
of bpf_get_btf_vmlinux() and bpf_find_btf_id() were not written for
that: some hold locks the module load needs, some run with interrupts
disabled.  So neither of them loads anything.  With =m,
bpf_get_btf_vmlinux() returns the BTF if it is loaded and NULL
otherwise, like a kernel without BTF, and counts the miss; it no longer
sleeps at all.  A new function, bpf_load_btf_vmlinux(), is the only one
that loads: it calls request_module() for the carrier outside
btf_vmlinux_lock, so that the notifier never waits for its caller, then
parses the BTF.  The notifier installs the copy before init_module()
returns, so the data is either there afterwards or the module is not
available (yet); that is not cached, the next call tries again.  A parse
that fails for lack of memory is not remembered either; any other
failure is a broken BTF, the same with every attempt, and is remembered
as with =y.

bpf_load_btf_vmlinux() is called only at the start of a request from
user space, in process context, holding no lock that loading a module
needs:

 - on entry to the bpf() system call: a BPF_PROG_LOAD, BPF_MAP_CREATE
   or BPF_BTF_LOAD that fails while the vmlinux BTF was found missing is
   run once more after loading it, the way tc and nf_tables retry a
   request after loading a module.  A failure of these commands leaves
   nothing behind, and bpf_sys_bpf() and kern_sys_bpf() do not go
   through the system call entry, so nothing loads from within a
   running BPF program.  The miss count is global, so a command that
   fails for another reason while a concurrent one misses is retried
   too, and fails again the same way;
 - for BPF_BTF_GET_NEXT_ID, for callers that may enumerate BTF ids
   (CAP_SYS_ADMIN, as bpf_obj_get_next_id() checks): kernel BTFs get
   their ids when the vmlinux BTF is parsed, and whoever enumerates BTF
   ids wants them;
 - when user space loads a syscall program: a light skeleton loader
   loads programs while it runs (bpf_sys_bpf()), and those may need the
   BTF, which cannot be loaded from there;
 - for read() of /sys/kernel/btf/vmlinux, see below.

The next patch adds the tracefs and bpffs entry points.

The module notifier, btf_parse_module() and the btf_data fields in struct
module are compiled for CONFIG_DEBUG_INFO_BTF_MODULES or =m; with =m and
no module BTF, the notifier only recognizes the carrier.

/sys/kernel/btf/vmlinux exists from boot with its final size, which is
known from .BTF.link before the BTF is loaded, so stat() works before the
load, which is what the btf_sysfs selftest does.  The first read() loads
it, and the raw BTF is served even if it does not parse, as with =y.
mmap() does not load it: it runs with the caller's mmap_lock held, and
the module load takes event_mutex (trace_module_notify()), which a task
registering a uprobe holds while it takes the mmap_lock of every mm that
maps the probed file.  Before the BTF is loaded mmap() fails, and libbpf,
which tries mmap() first, falls back to read(); afterwards it maps the
vmalloc_user() copy with remap_vmalloc_range().

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 include/linux/bpf.h    |   9 +++
 include/linux/btf.h    |   1 +
 include/linux/module.h |   2 +-
 kernel/bpf/btf.c       | 177 +++++++++++++++++++++++++++++++++++++++--
 kernel/bpf/syscall.c   |  51 +++++++++++-
 kernel/bpf/sysfs_btf.c |  97 +++++++++++++++++++++-
 kernel/bpf/verifier.c  |  91 +++++++++++++++++++--
 kernel/module/main.c   |   4 +-
 8 files changed, 412 insertions(+), 20 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index e46a14809dd4..d812dbc683ae 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -3184,6 +3184,15 @@ static inline s32 bpf_call_args_imm(s16 idx)
 
 struct btf *bpf_get_btf_vmlinux(void);
 struct btf *bpf_peek_btf_vmlinux(void);
+struct btf *bpf_load_btf_vmlinux(void);
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+unsigned int bpf_btf_vmlinux_misses(void);
+#else
+static inline unsigned int bpf_btf_vmlinux_misses(void)
+{
+	return 0;
+}
+#endif
 
 /* Map specifics */
 struct xdp_frame;
diff --git a/include/linux/btf.h b/include/linux/btf.h
index 4b63bb91550a..81e6c65fe5f6 100644
--- a/include/linux/btf.h
+++ b/include/linux/btf.h
@@ -602,6 +602,7 @@ __u32 *btf_field_iter_next(struct btf_field_iter *it);
 const char *btf_name_by_offset(const struct btf *btf, u32 offset);
 const char *btf_str_by_offset(const struct btf *btf, u32 offset);
 struct btf *btf_parse_vmlinux(void);
+void *btf_vmlinux_data(u32 *size, bool load);
 struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog);
 u32 *btf_kfunc_flags(const struct btf *btf, u32 kfunc_btf_id, const struct bpf_prog *prog);
 int btf_kfunc_check_flag(const struct btf *btf, u32 kfunc_btf_id, u32 flag);
diff --git a/include/linux/module.h b/include/linux/module.h
index 96cc98568eea..82734996a862 100644
--- a/include/linux/module.h
+++ b/include/linux/module.h
@@ -497,7 +497,7 @@ struct module {
 	unsigned int num_bpf_raw_events;
 	struct bpf_raw_event_map *bpf_raw_events;
 #endif
-#ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_INFO_BTF)
 	unsigned int btf_data_size;
 	unsigned int btf_base_data_size;
 	void *btf_data;
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 9896a30eeac4..2aa9d4b3f438 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -29,6 +29,7 @@
 #include <linux/string.h>
 #include <linux/sysfs.h>
 #include <linux/overflow.h>
+#include <crypto/sha2.h>
 #include <linux/bitops.h>
 
 #include <net/netfilter/nf_bpf_link.h>
@@ -6487,10 +6488,83 @@ static struct btf *btf_parse(const union bpf_attr *attr, bpfptr_t uattr,
 	return ERR_PTR(err);
 }
 
+#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF)
 extern char __start_BTF[];
 extern char __stop_BTF[];
+#endif
 extern struct btf *btf_vmlinux;
 
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+/*
+ * With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is not part of the kernel
+ * image.  The btf_vmlinux module carries it in its .BTF section; when the
+ * module loads, btf_module_notify() copies the section here.  The copy is
+ * made with vmalloc_user() so that /sys/kernel/btf/vmlinux can be mmap()ed
+ * as with the built-in BTF.  Set once, never cleared: like the built-in
+ * BTF, once present it stays for the lifetime of the kernel.
+ *
+ * The kernel is linked with .BTF.link, which describes the BTF the module
+ * carries, much as .gnu_debuglink describes a separate debug info file: the
+ * name of the module, and the size and SHA-256 of the BTF.  It is zeroed
+ * here and filled in by resolve_btfids at the end of the link (see
+ * scripts/link-vmlinux.sh), which also checks that the section is as large
+ * as this struct, whose name field is sized like the one in struct module.
+ * The size makes /sys/kernel/btf/vmlinux report its size before the BTF is
+ * loaded, the hash makes sure only the BTF this kernel was built with is
+ * accepted.
+ */
+struct btf_link {
+	char module_name[MODULE_NAME_LEN];
+	u8 sha256[SHA256_DIGEST_SIZE];
+	u32 btf_size;
+} __packed;
+
+static const struct btf_link __btf_vmlinux_link __section(".BTF.link") __used;
+
+/* Read through the linker symbol, the compiler would fold the zeroes above */
+extern const struct btf_link __start_BTF_link[];
+#define btf_vmlinux_link (__start_BTF_link[0])
+
+static void *btf_vmlinux_raw;
+#endif
+
+/**
+ * btf_vmlinux_data - get the raw vmlinux BTF
+ * @size: where to store the size of the BTF, also when it is not loaded yet
+ * @load: with CONFIG_DEBUG_INFO_BTF=m, load the btf_vmlinux module if the
+ *	  BTF is not present yet; waits for user space, see
+ *	  bpf_load_btf_vmlinux()
+ *
+ * Return: the raw BTF, or NULL if it is not available.
+ */
+void *btf_vmlinux_data(u32 *size, bool load)
+{
+#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF)
+	*size = __stop_BTF - __start_BTF;
+	return __start_BTF;
+#elif IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+	/* Pairs with the smp_store_release() in btf_vmlinux_module_coming() */
+	void *data = smp_load_acquire(&btf_vmlinux_raw);
+
+	if (!data && load) {
+		/*
+		 * The module notifier installs the BTF before init_module()
+		 * returns, so it is either there after this or the module is
+		 * not available (yet).  Not cached: a later call retries,
+		 * e.g. once the module becomes reachable on the root fs.
+		 */
+		request_module("%s", btf_vmlinux_link.module_name);
+		/* Same pairing as above */
+		data = smp_load_acquire(&btf_vmlinux_raw);
+	}
+	*size = btf_vmlinux_link.btf_size;
+	return data;
+#else
+	*size = 0;
+	return NULL;
+#endif
+}
+
 #define BPF_MAP_TYPE(_id, _ops)
 #define BPF_LINK_TYPE(_id, _name)
 static union {
@@ -6891,15 +6965,22 @@ struct btf *btf_parse_vmlinux(void)
 	struct btf_verifier_env *env = NULL;
 	struct bpf_verifier_log *log;
 	struct btf *btf;
+	void *data;
+	u32 size;
 	int err;
 
+	/* The caller made sure the BTF is present, see bpf_load_btf_vmlinux() */
+	data = btf_vmlinux_data(&size, false);
+	if (!data)
+		return ERR_PTR(-ENOENT);
+
 	env = kzalloc_obj(*env, GFP_KERNEL | __GFP_NOWARN);
 	if (!env)
 		return ERR_PTR(-ENOMEM);
 
 	log = &env->log;
 	log->level = BPF_LOG_KERNEL;
-	btf = btf_parse_base(env, "vmlinux", __start_BTF, __stop_BTF - __start_BTF);
+	btf = btf_parse_base(env, "vmlinux", data, size);
 	if (IS_ERR(btf))
 		goto err_out;
 
@@ -6907,6 +6988,7 @@ struct btf *btf_parse_vmlinux(void)
 	bpf_ctx_convert.t = btf_type_by_id(btf, bpf_ctx_convert_btf_id[0]);
 	err = btf_alloc_id(btf);
 	if (err) {
+		bpf_ctx_convert.t = NULL;
 		btf_free(btf);
 		btf = ERR_PTR(err);
 	}
@@ -6926,7 +7008,7 @@ __u32 btf_relocate_id(const struct btf *btf, __u32 id)
 	return btf->base_id_map[id];
 }
 
-#ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_INFO_BTF)
 
 /*
  * Parse split module BTF against @vmlinux_btf.  @data is the module's .BTF
@@ -7032,7 +7114,7 @@ static struct btf *btf_parse_module(const char *module_name, struct btf *vmlinux
 	return ERR_PTR(err);
 }
 
-#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */
+#endif /* CONFIG_DEBUG_INFO_BTF_MODULES || CONFIG_DEBUG_INFO_BTF=m */
 
 struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog)
 {
@@ -9019,7 +9101,16 @@ enum {
 	BTF_MODULE_F_LIVE = (1 << 0),
 };
 
-#ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+/*
+ * The module notifier registers module BTF (CONFIG_DEBUG_INFO_BTF_MODULES)
+ * and picks up the vmlinux BTF from the btf_vmlinux module
+ * (CONFIG_DEBUG_INFO_BTF=m).
+ */
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+#define BTF_MODULE_NOTIFIER 1
+#endif
+
+#ifdef BTF_MODULE_NOTIFIER
 struct btf_module {
 	struct list_head list;
 	struct module *module;
@@ -9075,6 +9166,63 @@ static void btf_module_free(struct btf_module *btf_mod)
 	kfree(btf_mod);
 }
 
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+/*
+ * The btf_vmlinux module carries the vmlinux BTF in its .BTF section
+ * (scripts/gen-btf.sh).  Keep a copy; the module is only the carrier and
+ * has no BTF of its own.
+ */
+static int btf_vmlinux_module_coming(struct module *mod)
+{
+	u8 sha256sum[SHA256_DIGEST_SIZE];
+	void *data;
+
+	if (btf_vmlinux_raw)
+		return 0;
+
+	/*
+	 * The verifier trusts the BTF as the description of this kernel's
+	 * types, so a BTF from a different build must not get in even if
+	 * the module otherwise loads (same release string, same vermagic).
+	 */
+	if (mod->btf_data_size != btf_vmlinux_link.btf_size) {
+		pr_err("module [%s]: BTF size %u does not match this kernel (%u)\n",
+		       mod->name, mod->btf_data_size, btf_vmlinux_link.btf_size);
+		return -EINVAL;
+	}
+	sha256(mod->btf_data, mod->btf_data_size, sha256sum);
+	if (memcmp(sha256sum, btf_vmlinux_link.sha256, sizeof(sha256sum))) {
+		pr_err("module [%s]: BTF does not match this kernel\n", mod->name);
+		return -EINVAL;
+	}
+
+	data = vmalloc_user(mod->btf_data_size);
+	if (!data)
+		return -ENOMEM;
+	memcpy(data, mod->btf_data, mod->btf_data_size);
+
+	/* Pairs with the smp_load_acquire() in btf_vmlinux_data() */
+	smp_store_release(&btf_vmlinux_raw, data);
+	return 0;
+}
+
+/* The module named in .BTF.link */
+static bool btf_is_vmlinux_carrier(const struct module *mod)
+{
+	return !strcmp(mod->name, btf_vmlinux_link.module_name);
+}
+#else
+static int btf_vmlinux_module_coming(struct module *mod)
+{
+	return 0;
+}
+
+static bool btf_is_vmlinux_carrier(const struct module *mod)
+{
+	return false;
+}
+#endif
+
 static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			     void *module)
 {
@@ -9083,9 +9231,17 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 	struct btf *btf;
 	int err = 0;
 
-	if (mod->btf_data_size == 0 ||
-	    (op != MODULE_STATE_COMING && op != MODULE_STATE_LIVE &&
-	     op != MODULE_STATE_GOING))
+	if (op != MODULE_STATE_COMING && op != MODULE_STATE_LIVE &&
+	    op != MODULE_STATE_GOING)
+		goto out;
+
+	if (btf_is_vmlinux_carrier(mod)) {
+		if (op == MODULE_STATE_COMING)
+			err = btf_vmlinux_module_coming(mod);
+		goto out;
+	}
+
+	if (!IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || mod->btf_data_size == 0)
 		goto out;
 
 	switch (op) {
@@ -9173,7 +9329,7 @@ static int __init btf_module_init(void)
 }
 
 fs_initcall(btf_module_init);
-#endif /* CONFIG_DEBUG_INFO_BTF_MODULES */
+#endif /* BTF_MODULE_NOTIFIER */
 
 struct module *btf_try_get_module(const struct btf *btf)
 {
@@ -10136,6 +10292,11 @@ static void purge_cand_cache(struct btf *btf)
 	__purge_cand_cache(btf, module_cand_cache, MODULE_CAND_CACHE_SIZE);
 	mutex_unlock(&cand_cache_mutex);
 }
+#elif defined(BTF_MODULE_NOTIFIER)
+/* CONFIG_DEBUG_INFO_BTF=m without module BTF: nothing is ever cached */
+static void purge_cand_cache(struct btf *btf)
+{
+}
 #endif
 
 static struct bpf_cand_cache *
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index ac52f4ae414c..e5c4c0aa776e 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -3002,6 +3002,17 @@ static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_at
 	if (is_perfmon_prog_type(type) && !bpf_token_capable(token, CAP_PERFMON))
 		goto put_token;
 
+	/*
+	 * CONFIG_DEBUG_INFO_BTF=m: a light skeleton loader is a syscall
+	 * program that loads programs while it runs (bpf_sys_bpf()), and
+	 * those may need the vmlinux BTF, which cannot be loaded from within
+	 * a running program.  Load it now, while user space loads the loader;
+	 * if that fails, the loader works as on a kernel without BTF.
+	 */
+	if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && type == BPF_PROG_TYPE_SYSCALL &&
+	    !uattr.is_kernel)
+		bpf_load_btf_vmlinux();
+
 	multi_func = is_tracing_multi(attr->expected_attach_type);
 
 	/* attach_prog_fd/attach_btf_obj_fd can specify fd of either bpf_prog
@@ -6435,6 +6446,14 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
 					  &map_idr, &map_idr_lock);
 		break;
 	case BPF_BTF_GET_NEXT_ID:
+		/*
+		 * With CONFIG_DEBUG_INFO_BTF=m the kernel BTFs get ids when the
+		 * vmlinux BTF is loaded; whoever enumerates them wants them.
+		 * Only for callers bpf_obj_get_next_id() lets through.
+		 */
+		if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) &&
+		    ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN))
+			bpf_load_btf_vmlinux();
 		err = bpf_obj_get_next_id(&attr, uattr.user,
 					  &btf_idr, &btf_idr_lock);
 		break;
@@ -6525,10 +6544,40 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
 	return err;
 }
 
+/*
+ * With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is loaded on demand, but never
+ * from within a command: loading waits for user space, and a command may hold
+ * locks or run from a BPF program (bpf_sys_bpf()).  A command that needs the
+ * BTF while it is not loaded fails as it would without BTF.  If the command
+ * is one whose failure leaves nothing behind, load the BTF here, on entry
+ * from user space with nothing held, and run the command once more.
+ */
+static bool bpf_btf_vmlinux_retry(int cmd, unsigned int misses)
+{
+	switch (cmd & ~BPF_COMMON_ATTRS) {
+	case BPF_PROG_LOAD:
+	case BPF_MAP_CREATE:
+	case BPF_BTF_LOAD:
+		break;
+	default:
+		return false;
+	}
+	if (bpf_btf_vmlinux_misses() == misses)
+		return false;
+	return !IS_ERR_OR_NULL(bpf_load_btf_vmlinux());
+}
+
 SYSCALL_DEFINE5(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size,
 		struct bpf_common_attr __user *, uattr_common, unsigned int, size_common)
 {
-	return __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
+	unsigned int misses = bpf_btf_vmlinux_misses();
+	int err;
+
+	err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
+	if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && err < 0 && bpf_btf_vmlinux_retry(cmd, misses))
+		err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common),
+				size_common);
+	return err;
 }
 
 static bool syscall_prog_is_valid_access(int off, int size,
diff --git a/kernel/bpf/sysfs_btf.c b/kernel/bpf/sysfs_btf.c
index 9cbe15ce3540..296eaf642786 100644
--- a/kernel/bpf/sysfs_btf.c
+++ b/kernel/bpf/sysfs_btf.c
@@ -9,8 +9,13 @@
 #include <linux/sysfs.h>
 #include <linux/mm.h>
 #include <linux/io.h>
+#include <linux/bpf.h>
 #include <linux/btf.h>
+#include <linux/vmalloc.h>
 
+struct kobject *btf_kobj;
+
+#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF)
 /* See scripts/link-vmlinux.sh, gen_btf() func for details */
 extern char __start_BTF[];
 extern char __stop_BTF[];
@@ -49,12 +54,98 @@ static struct bin_attribute bin_attr_btf_vmlinux __ro_after_init = {
 	.mmap = btf_sysfs_vmlinux_mmap,
 };
 
-struct kobject *btf_kobj;
-
-static int __init btf_vmlinux_init(void)
+static void __init btf_sysfs_vmlinux_init(void)
 {
 	bin_attr_btf_vmlinux.private = __start_BTF;
 	bin_attr_btf_vmlinux.size = __stop_BTF - __start_BTF;
+}
+
+#else /* CONFIG_DEBUG_INFO_BTF=m */
+
+/*
+ * The BTF is carried by the btf_vmlinux module and only loaded when
+ * something needs it.  Its size is known from the start, so the file has
+ * its final size from boot; the first read() loads the BTF, mmap() maps it
+ * once it is loaded.
+ */
+static void *btf_sysfs_vmlinux_load(u32 *size)
+{
+	/*
+	 * Loads the module, parses the BTF and registers module BTFs.  The
+	 * raw BTF is served even if it does not parse, as with =y.
+	 */
+	bpf_load_btf_vmlinux();
+	return btf_vmlinux_data(size, false);
+}
+
+static ssize_t btf_sysfs_vmlinux_read(struct file *filp, struct kobject *kobj,
+				      const struct bin_attribute *attr,
+				      char *buf, loff_t off, size_t count)
+{
+	u32 size;
+	void *data = btf_sysfs_vmlinux_load(&size);
+
+	if (!data)
+		return -ENODEV;
+
+	/* sysfs clamps @off and @count to attr->size, which is @size */
+	memcpy(buf, data + off, count);
+	return count;
+}
+
+static int btf_sysfs_vmlinux_mmap(struct file *filp, struct kobject *kobj,
+				  const struct bin_attribute *attr,
+				  struct vm_area_struct *vma)
+{
+	size_t vm_size = vma->vm_end - vma->vm_start;
+	void *data;
+	u32 size;
+
+	if (vma->vm_pgoff)
+		return -EINVAL;
+
+	if (vma->vm_flags & (VM_WRITE | VM_EXEC | VM_MAYSHARE))
+		return -EACCES;
+
+	if (vm_size > PAGE_ALIGN(attr->size))
+		return -EINVAL;
+
+	/*
+	 * Not loaded from here: this runs with the caller's mmap_lock held,
+	 * and loading the module waits for modprobe, whose module load takes
+	 * event_mutex in the trace module notifier, which a task registering
+	 * a uprobe holds while it takes the mmap_lock of each mm that maps
+	 * the probed file.  Until the BTF is loaded, e.g. by a read(), mmap()
+	 * fails; libbpf then falls back to read().
+	 */
+	data = btf_vmlinux_data(&size, false);
+	if (!data)
+		return -ENODEV;
+
+	vm_flags_mod(vma, VM_DONTDUMP, VM_MAYEXEC | VM_MAYWRITE);
+	/* the copy was made with vmalloc_user() for this purpose */
+	return remap_vmalloc_range(vma, data, 0);
+}
+
+static struct bin_attribute bin_attr_btf_vmlinux __ro_after_init = {
+	.attr = { .name = "vmlinux", .mode = 0444, },
+	.read = btf_sysfs_vmlinux_read,
+	.mmap = btf_sysfs_vmlinux_mmap,
+};
+
+static void __init btf_sysfs_vmlinux_init(void)
+{
+	u32 size;
+
+	/* known before the BTF is loaded, see .BTF.link */
+	btf_vmlinux_data(&size, false);
+	bin_attr_btf_vmlinux.size = size;
+}
+#endif
+
+static int __init btf_vmlinux_init(void)
+{
+	btf_sysfs_vmlinux_init();
 
 	if (bin_attr_btf_vmlinux.size == 0)
 		return 0;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f375c5dad4b5..895e7feb6429 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19847,7 +19847,7 @@ static int check_pseudo_btf_id(struct bpf_verifier_env *env,
 		/* kernel types enter the program here, see bpf_check() */
 		btf = bpf_get_btf_vmlinux();
 		if (IS_ERR_OR_NULL(btf)) {
-			verbose(env, "kernel is missing BTF, make sure CONFIG_DEBUG_INFO_BTF=y is specified in Kconfig.\n");
+			verbose(env, "kernel is missing BTF, make sure CONFIG_DEBUG_INFO_BTF is specified in Kconfig (with =m, that btf_vmlinux can be loaded).\n");
 			return -EINVAL;
 		}
 		btf_get(btf);
@@ -21881,11 +21881,26 @@ int bpf_check_attach_btf_id_multi(struct btf *btf, struct bpf_prog *prog, u32 bt
 	return 0;
 }
 
+/* CONFIG_DEBUG_INFO_BTF=m: lookups that found the vmlinux BTF not loaded */
+static atomic_t btf_vmlinux_misses = ATOMIC_INIT(0);
+
+/*
+ * Returns the parsed vmlinux BTF, NULL if the kernel has none, or an ERR_PTR
+ * if it is malformed.  Never waits for user space.
+ *
+ * With CONFIG_DEBUG_INFO_BTF=m the BTF is in the btf_vmlinux module, and
+ * this does not load it: until bpf_load_btf_vmlinux() has, it returns NULL,
+ * as without BTF, and counts the miss for bpf_btf_vmlinux_misses().
+ */
 struct btf *bpf_get_btf_vmlinux(void)
 {
-	/* Pairs with the smp_store_release() on the parse path below. */
+	/* Pairs with the smp_store_release() on the parse paths. */
 	struct btf *btf = smp_load_acquire(&btf_vmlinux);
 
+	if (!btf && IS_MODULE(CONFIG_DEBUG_INFO_BTF)) {
+		atomic_inc(&btf_vmlinux_misses);
+		return NULL;
+	}
 	if (!btf && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)) {
 		mutex_lock(&btf_vmlinux_lock);
 		btf = btf_vmlinux;
@@ -21906,15 +21921,81 @@ struct btf *bpf_get_btf_vmlinux(void)
 
 /*
  * The vmlinux BTF if it has been parsed already, else NULL.  Unlike
- * bpf_get_btf_vmlinux() this never parses anything: for a running BPF
- * program, and for code that only uses the BTF if it happens to be there.
+ * bpf_get_btf_vmlinux() this never parses anything and, with
+ * CONFIG_DEBUG_INFO_BTF=m, does not count a miss: for a running BPF program,
+ * and for code that only uses the BTF if it happens to be there.
  */
 struct btf *bpf_peek_btf_vmlinux(void)
 {
-	/* Pairs with the smp_store_release() in bpf_get_btf_vmlinux() */
+	/* Pairs with the smp_store_release() on the parse paths */
 	return smp_load_acquire(&btf_vmlinux);
 }
 
+/**
+ * bpf_load_btf_vmlinux - get the vmlinux BTF, loading it if necessary
+ *
+ * Like bpf_get_btf_vmlinux(), but with CONFIG_DEBUG_INFO_BTF=m it loads the
+ * btf_vmlinux module if the BTF is not there yet and parses it.  Loading the
+ * module waits for user space (modprobe), and the notifiers of the module
+ * load take locks of their own, event_mutex among them.  So this is only
+ * called at the start of a request from user space, in process context,
+ * holding no lock that loading a module may need; everything else uses
+ * bpf_get_btf_vmlinux() or bpf_peek_btf_vmlinux().
+ *
+ * If the module cannot be loaded, returns NULL like a kernel without BTF;
+ * the next call tries again.
+ */
+struct btf *bpf_load_btf_vmlinux(void)
+{
+	struct btf *btf;
+	u32 size;
+
+	might_sleep();
+	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF))
+		return bpf_get_btf_vmlinux();
+
+	/* Pairs with the smp_store_release() below */
+	btf = smp_load_acquire(&btf_vmlinux);
+	if (btf)
+		return btf;
+
+	/* Outside btf_vmlinux_lock, the module's notifier must not wait for us */
+	if (!btf_vmlinux_data(&size, true))
+		return NULL;
+
+	mutex_lock(&btf_vmlinux_lock);
+	btf = btf_vmlinux;
+	if (!btf) {
+		btf = btf_parse_vmlinux();
+		/*
+		 * An allocation failure is not remembered, the next caller
+		 * retries.  Anything else is a broken BTF, the same one with
+		 * every attempt, and is remembered as with =y.
+		 */
+		if (IS_ERR(btf) && PTR_ERR(btf) == -ENOMEM) {
+			mutex_unlock(&btf_vmlinux_lock);
+			return btf;
+		}
+		/* As in bpf_get_btf_vmlinux(): publish after the parse */
+		smp_store_release(&btf_vmlinux, btf);
+	}
+	mutex_unlock(&btf_vmlinux_lock);
+	return btf;
+}
+
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+/*
+ * bpf(2) samples this before a command and, if the command failed and the
+ * count moved, loads the vmlinux BTF and runs the command once more.  The
+ * count is global: a command that failed for another reason while a
+ * concurrent one missed the BTF is run again as well, and fails the same way.
+ */
+unsigned int bpf_btf_vmlinux_misses(void)
+{
+	return atomic_read(&btf_vmlinux_misses);
+}
+#endif
+
 /*
  * The add_fd_from_fd_array() is executed only if fd_array_cnt is non-zero. In
  * this case expect that every file descriptor in the array is either a map or
diff --git a/kernel/module/main.c b/kernel/module/main.c
index d0e1e0bd2ad0..694c4bc7e679 100644
--- a/kernel/module/main.c
+++ b/kernel/module/main.c
@@ -2718,7 +2718,7 @@ static int find_module_sections(struct module *mod, struct load_info *info)
 					   sizeof(*mod->bpf_raw_events),
 					   &mod->num_bpf_raw_events);
 #endif
-#ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_INFO_BTF)
 	mod->btf_data = any_section_objs(info, ".BTF", 1, &mod->btf_data_size);
 	mod->btf_base_data = any_section_objs(info, ".BTF.base", 1,
 					      &mod->btf_base_data_size);
@@ -3172,7 +3172,7 @@ static noinline int do_init_module(struct module *mod)
 		mod->mem[type].size = 0;
 	}
 
-#ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF_MODULES) || IS_MODULE(CONFIG_DEBUG_INFO_BTF)
 	/* .BTF is not SHF_ALLOC and will get removed, so sanitize pointers */
 	mod->btf_data = NULL;
 	mod->btf_base_data = NULL;
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (3 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
                   ` (7 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

With CONFIG_DEBUG_INFO_BTF=m, bpf_get_btf_vmlinux() and
bpf_find_btf_id() do not load the vmlinux BTF: loading waits for user
space, and their callers were not written for that.  Besides the bpf()
system call and /sys/kernel/btf/vmlinux, some tracefs and bpffs requests
need the BTF.  Load it at the start of those, with
bpf_load_btf_vmlinux(), and have the code that only uses the BTF if it
happens to be there peek:

 - Reading a tracepoint's btf_ids file (events/*/btf_ids): built-in
   events use the vmlinux BTF, and module BTF is only registered once
   that is loaded.  event_btf_ids_read() looks the BTF up under
   event_mutex, which the trace notifier of the btf_vmlinux module
   takes, so load it before taking the mutex, on the first read() of
   the file (not again for the one that returns EOF).
 - Probe events with BTF arguments ($argN, argument names, $retval,
   $current, typecasts): the parser loads the BTF before looking a
   function or struct up.  It holds dyn_event_ops_mutex, which loading
   a module never takes: besides event creation, only
   dyn_event_register() takes it, from built-in init code.  Loading
   here also keeps a $retval from silently losing its type.
 - The ftrace function argument printer (func-args, funcgraph-args)
   runs in the trace output path, which includes ftrace_dump() with
   interrupts disabled.  It only prints the arguments if the BTF is
   already loaded, and never loads it: that also avoids one modprobe
   per trace line when the module is not installed.
 - bpffs: parsing delegate_* mount options (fs_context) loads the BTF
   when a value names commands or types, which are looked up in it;
   "any" and numeric masks need no BTF and do not load it.  Showing the
   options in /proc/*/mountinfo runs under namespace_sem; it only uses
   the names if the BTF is already there and falls back to hex, as it
   already does without BTF.

With CONFIG_DEBUG_INFO_BTF=y bpf_load_btf_vmlinux() is
bpf_get_btf_vmlinux() and the BTF is parsed at boot, so nothing changes.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/inode.c          | 44 ++++++++++++++++++++++++-------------
 kernel/trace/trace_events.c | 10 +++++++++
 kernel/trace/trace_output.c |  7 ++++++
 kernel/trace/trace_probe.c  | 16 ++++++++++++++
 4 files changed, 62 insertions(+), 15 deletions(-)

diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842..d05bbb61a593 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -658,7 +658,11 @@ struct bpffs_btf_enums {
 	const struct btf_type *attach_t;
 };
 
-static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
+/*
+ * @load: load the vmlinux BTF if necessary (CONFIG_DEBUG_INFO_BTF=m), see
+ * bpf_load_btf_vmlinux(); otherwise only use it if it is already parsed.
+ */
+static int find_bpffs_btf_enums(struct bpffs_btf_enums *info, bool load)
 {
 	struct {
 		const struct btf_type **type;
@@ -674,7 +678,7 @@ static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
 
 	memset(info, 0, sizeof(*info));
 
-	btf = bpf_get_btf_vmlinux();
+	btf = load ? bpf_load_btf_vmlinux() : bpf_peek_btf_vmlinux();
 	if (IS_ERR(btf))
 		return PTR_ERR(btf);
 	if (!btf)
@@ -795,8 +799,11 @@ static int bpf_show_options(struct seq_file *m, struct dentry *root)
 	    opts->delegate_progs || opts->delegate_attachs) {
 		struct bpffs_btf_enums info;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
+		/*
+		 * ignore errors, fallback to hex; this runs under
+		 * namespace_sem, so do not load the BTF from here
+		 */
+		(void)find_bpffs_btf_enums(&info, false);
 
 		mask = (1ULL << __MAX_BPF_CMD) - 1;
 		seq_print_delegate_opts(m, "delegate_cmds",
@@ -1052,35 +1059,33 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 	case OPT_DELEGATE_MAPS:
 	case OPT_DELEGATE_PROGS:
 	case OPT_DELEGATE_ATTACHS: {
-		struct bpffs_btf_enums info;
-		const struct btf_type *enum_t;
+		struct bpffs_btf_enums info = {};
+		const struct btf_type **enum_t;
+		bool enums_tried = false;
 		const char *enum_pfx;
-		u64 *delegate_msk, msk = 0;
+		u64 *delegate_msk, msk = 0, num;
 		char *p, *str;
 		int val;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
-
 		switch (opt) {
 		case OPT_DELEGATE_CMDS:
 			delegate_msk = &opts->delegate_cmds;
-			enum_t = info.cmd_t;
+			enum_t = &info.cmd_t;
 			enum_pfx = "BPF_";
 			break;
 		case OPT_DELEGATE_MAPS:
 			delegate_msk = &opts->delegate_maps;
-			enum_t = info.map_t;
+			enum_t = &info.map_t;
 			enum_pfx = "BPF_MAP_TYPE_";
 			break;
 		case OPT_DELEGATE_PROGS:
 			delegate_msk = &opts->delegate_progs;
-			enum_t = info.prog_t;
+			enum_t = &info.prog_t;
 			enum_pfx = "BPF_PROG_TYPE_";
 			break;
 		case OPT_DELEGATE_ATTACHS:
 			delegate_msk = &opts->delegate_attachs;
-			enum_t = info.attach_t;
+			enum_t = &info.attach_t;
 			enum_pfx = "BPF_";
 			break;
 		default:
@@ -1089,9 +1094,18 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 
 		str = param->string;
 		while ((p = strsep(&str, ":"))) {
+			/*
+			 * Only names need the vmlinux BTF: "any" and numbers do
+			 * not load it.  Ignore errors, fallback to hex.
+			 */
+			if (strcmp(p, "any") && kstrtou64(p, 0, &num) && !enums_tried) {
+				(void)find_bpffs_btf_enums(&info, true);
+				enums_tried = true;
+			}
+
 			if (strcmp(p, "any") == 0) {
 				msk |= ~0ULL;
-			} else if (find_btf_enum_const(info.btf, enum_t, enum_pfx, p, &val)) {
+			} else if (find_btf_enum_const(info.btf, *enum_t, enum_pfx, p, &val)) {
 				msk |= 1ULL << val;
 			} else {
 				err = kstrtou64(p, 0, &msk);
diff --git a/kernel/trace/trace_events.c b/kernel/trace/trace_events.c
index 30c0ddf90887..c887ac6a4857 100644
--- a/kernel/trace/trace_events.c
+++ b/kernel/trace/trace_events.c
@@ -23,6 +23,7 @@
 #include <linux/sort.h>
 #include <linux/slab.h>
 #include <linux/delay.h>
+#include <linux/bpf.h>
 #include <linux/btf.h>
 
 #include <trace/events/sched.h>
@@ -2245,6 +2246,15 @@ event_btf_ids_read(struct file *filp, char __user *ubuf, size_t cnt, loff_t *ppo
 	char buf[128];
 	int len;
 
+	/*
+	 * Built-in events use the vmlinux BTF, and with CONFIG_DEBUG_INFO_BTF=m
+	 * module BTF is only registered once that is loaded.  Loading it loads
+	 * a module, whose trace notifier takes event_mutex: load it before
+	 * taking that, and only for the first read, not again for the EOF one.
+	 */
+	if (!*ppos)
+		bpf_load_btf_vmlinux();
+
 	/* Module unload could free call->class and ids[] mid-read. */
 	scoped_guard(mutex, &event_mutex) {
 		file = event_file_file(filp);
diff --git a/kernel/trace/trace_output.c b/kernel/trace/trace_output.c
index a5ad76175d10..1f346e524ee6 100644
--- a/kernel/trace/trace_output.c
+++ b/kernel/trace/trace_output.c
@@ -739,6 +739,13 @@ void print_function_args(struct trace_seq *s, unsigned long *args,
 	if (lookup_symbol_name(func, name))
 		goto out;
 
+	/*
+	 * This can run with interrupts disabled (ftrace_dump()): only use
+	 * the vmlinux BTF if it is parsed, never load it from here.
+	 */
+	if (IS_ERR_OR_NULL(bpf_peek_btf_vmlinux()))
+		goto out;
+
 	/* TODO: Pass module name here too */
 	t = btf_find_func_proto(name, &btf);
 	if (IS_ERR_OR_NULL(t))
diff --git a/kernel/trace/trace_probe.c b/kernel/trace/trace_probe.c
index 804442b2f7d2..d53ee1ef820c 100644
--- a/kernel/trace/trace_probe.c
+++ b/kernel/trace/trace_probe.c
@@ -531,6 +531,19 @@ static const char *fetch_type_from_btf_type(struct btf *btf,
 	return NULL;
 }
 
+/*
+ * Arguments described by BTF need the vmlinux BTF.  With
+ * CONFIG_DEBUG_INFO_BTF=m it may not be loaded yet, so load it before looking
+ * anything up.  Parsing holds dyn_event_ops_mutex, which loading a module
+ * never takes: besides event creation, only dyn_event_register() takes it,
+ * from built-in init code.
+ */
+static void trace_probe_load_btf(void)
+{
+	lockdep_assert_held(&dyn_event_ops_mutex);
+	bpf_load_btf_vmlinux();
+}
+
 static int query_btf_context(struct traceprobe_parse_context *ctx)
 {
 	const struct btf_param *param;
@@ -544,6 +557,7 @@ static int query_btf_context(struct traceprobe_parse_context *ctx)
 	if (!ctx->funcname)
 		return -EINVAL;
 
+	trace_probe_load_btf();
 	type = btf_find_func_proto(ctx->funcname, &btf);
 	if (!type)
 		return -ENOENT;
@@ -762,6 +776,7 @@ static int parse_btf_arg(char *varname,
 	if (!strcmp(varname, "$current")) {
 		code->op = FETCH_OP_CURRENT;
 		/* If no typecast is specified for $current, use task_struct by default */
+		trace_probe_load_btf();
 		ret = bpf_find_btf_id("task_struct", BTF_KIND_STRUCT, &ctx->struct_btf);
 		if (ret < 0) {
 			trace_probe_log_err(ctx->offset, NO_BTF_ENTRY);
@@ -890,6 +905,7 @@ static int query_btf_struct(const char *sname, struct traceprobe_parse_context *
 		ctx->struct_btf = NULL;
 	}
 
+	trace_probe_load_btf();
 	id = bpf_find_btf_id(sname, BTF_KIND_STRUCT, &btf);
 	if (id < 0)
 		return id;
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (4 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
                   ` (6 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is loaded on first use.  For
that to save anything, nothing may pull it in at boot.  The verifier no
longer does since the previous patches, but register_btf_kfunc_id_set(),
register_btf_id_dtor_kfuncs() and register_bpf_struct_ops() for vmlinux
run from initcalls and need the parsed BTF.

Queue them instead (btf_defer_reg()) and apply them in
btf_parse_vmlinux(), before the BTF is published, so that no program can
see a vmlinux BTF without its kfuncs and struct_ops, not even by id: the
id is reserved before the queue is applied, so that nothing can fail once
it is, and installed afterwards (btf_reserve_id(), btf_install_id()).
Applying a struct_ops runs its ->init(), which registers the kfunc sets
of its hook; those land back on the queue, so it is drained in a loop
until a pass adds nothing, and only then are new registrations applied
directly.  The dtor arrays are copied: every caller in the tree builds
them on the stack of its initcall.

The queue has its own lock, btf_vmlinux_regs_mutex: it is drained under
btf_vmlinux_lock, and with its own lock btf_module_mutex, which the
module notifier takes, is never taken under btf_vmlinux_lock, so the two
stay unordered.

A parse failure other than -ENOMEM is remembered by
bpf_load_btf_vmlinux(), so nothing would ever apply the queue: it is then
freed and closed, and later registrations fail as they do with =y.

Module registrations are not queued yet; the next patch does that
together with deferring the module BTF itself.  With =y the BTF is
present from boot and nothing is queued.  Nothing here is reachable
until the Kconfig symbol becomes a tristate.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/btf.c | 271 ++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 268 insertions(+), 3 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 2aa9d4b3f438..96241dc62dc3 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -1786,13 +1786,13 @@ static int btf_add_type(struct btf_verifier_env *env, struct btf_type *t)
 	return 0;
 }
 
-static int btf_alloc_id(struct btf *btf)
+static int __btf_alloc_id(struct btf *btf, struct btf *entry)
 {
 	int id;
 
 	idr_preload(GFP_KERNEL);
 	spin_lock_bh(&btf_idr_lock);
-	id = idr_alloc_cyclic(&btf_idr, btf, 1, INT_MAX, GFP_ATOMIC);
+	id = idr_alloc_cyclic(&btf_idr, entry, 1, INT_MAX, GFP_ATOMIC);
 	if (id > 0)
 		btf->id = id;
 	spin_unlock_bh(&btf_idr_lock);
@@ -1804,6 +1804,29 @@ static int btf_alloc_id(struct btf *btf)
 	return id > 0 ? 0 : id;
 }
 
+static int btf_alloc_id(struct btf *btf)
+{
+	return __btf_alloc_id(btf, btf);
+}
+
+/*
+ * A BTF that gets registrations applied after it is parsed, with
+ * CONFIG_DEBUG_INFO_BTF=m, has its id reserved before, so that nothing can
+ * fail once they are applied, and installed after, so that nobody finds it
+ * by id in the meantime (btf_get_fd_by_id() and idr_get_next() skip NULL).
+ */
+static int btf_reserve_id(struct btf *btf)
+{
+	return __btf_alloc_id(btf, NULL);
+}
+
+static void btf_install_id(struct btf *btf)
+{
+	spin_lock_bh(&btf_idr_lock);
+	idr_replace(&btf_idr, btf, btf->id);
+	spin_unlock_bh(&btf_idr_lock);
+}
+
 static void btf_free_id(struct btf *btf)
 {
 	unsigned long flags;
@@ -6960,6 +6983,9 @@ static struct btf *btf_parse_base(struct btf_verifier_env *env, const char *name
 	return ERR_PTR(err);
 }
 
+static void btf_apply_deferred_vmlinux_regs(struct btf *btf);
+static void btf_drop_deferred_vmlinux_regs(void);
+
 struct btf *btf_parse_vmlinux(void)
 {
 	struct btf_verifier_env *env = NULL;
@@ -6986,13 +7012,29 @@ struct btf *btf_parse_vmlinux(void)
 
 	/* btf_parse_vmlinux() runs under btf_vmlinux_lock */
 	bpf_ctx_convert.t = btf_type_by_id(btf, bpf_ctx_convert_btf_id[0]);
-	err = btf_alloc_id(btf);
+	err = btf_reserve_id(btf);
 	if (err) {
 		bpf_ctx_convert.t = NULL;
 		btf_free(btf);
 		btf = ERR_PTR(err);
+		goto err_out;
 	}
+
+	/*
+	 * With CONFIG_DEBUG_INFO_BTF=m, kfunc, dtor kfunc and struct_ops
+	 * registrations for vmlinux made before the BTF was available were
+	 * queued; apply them now, before the BTF becomes visible to anyone,
+	 * by id or through bpf_get_btf_vmlinux().
+	 */
+	btf_apply_deferred_vmlinux_regs(btf);
+	btf_install_id(btf);
 err_out:
+	/*
+	 * Any failure but -ENOMEM is remembered by bpf_load_btf_vmlinux(), so
+	 * the queued registrations would never be applied.
+	 */
+	if (IS_ERR(btf) && PTR_ERR(btf) != -ENOMEM)
+		btf_drop_deferred_vmlinux_regs();
 	btf_verifier_env_free(env);
 	return btf;
 }
@@ -9110,6 +9152,33 @@ enum {
 #define BTF_MODULE_NOTIFIER 1
 #endif
 
+/*
+ * CONFIG_DEBUG_INFO_BTF=m: a kfunc, dtor kfunc or struct_ops registration
+ * made while the BTF it applies to is not available yet.  Kept until the BTF
+ * arrives, see btf_defer_reg().
+ */
+enum btf_deferred_reg_kind {
+	BTF_DEFERRED_KFUNC_SET,
+	BTF_DEFERRED_DTOR_KFUNCS,
+	BTF_DEFERRED_STRUCT_OPS,
+};
+
+struct btf_deferred_reg {
+	struct list_head list;
+	enum btf_deferred_reg_kind kind;
+	union {
+		struct {
+			enum btf_kfunc_hook hook;
+			const struct btf_kfunc_id_set *kset;
+		} kfunc;
+		struct {
+			const struct btf_id_dtor_kfunc *dtors;
+			u32 cnt;
+		} dtor;
+		struct bpf_struct_ops *st_ops;
+	};
+};
+
 #ifdef BTF_MODULE_NOTIFIER
 struct btf_module {
 	struct list_head list;
@@ -9905,12 +9974,22 @@ static int btf_kfunc_id_set_add(struct btf *btf, enum btf_kfunc_hook hook,
 	return btf_populate_kfunc_set(btf, hook, kset);
 }
 
+static int btf_defer_reg(struct module *owner, const struct btf_deferred_reg *tmpl);
+
 static int __register_btf_kfunc_id_set(enum btf_kfunc_hook hook,
 				       const struct btf_kfunc_id_set *kset)
 {
+	struct btf_deferred_reg tmpl = {
+		.kind = BTF_DEFERRED_KFUNC_SET,
+		.kfunc = { .hook = hook, .kset = kset },
+	};
 	struct btf *btf;
 	int ret;
 
+	ret = btf_defer_reg(kset->owner, &tmpl);
+	if (ret)
+		return ret > 0 ? 0 : ret;
+
 	btf = btf_get_module_btf(kset->owner);
 	if (!btf)
 		return check_btf_kconfigs(kset->owner, "kfunc");
@@ -10079,9 +10158,17 @@ static int btf_dtor_kfuncs_add(struct btf *btf, const struct btf_id_dtor_kfunc *
 int register_btf_id_dtor_kfuncs(const struct btf_id_dtor_kfunc *dtors, u32 add_cnt,
 				struct module *owner)
 {
+	struct btf_deferred_reg tmpl = {
+		.kind = BTF_DEFERRED_DTOR_KFUNCS,
+		.dtor = { .dtors = dtors, .cnt = add_cnt },
+	};
 	struct btf *btf;
 	int ret;
 
+	ret = btf_defer_reg(owner, &tmpl);
+	if (ret)
+		return ret > 0 ? 0 : ret;
+
 	btf = btf_get_module_btf(owner);
 	if (!btf)
 		return check_btf_kconfigs(owner, "dtor kfuncs");
@@ -10751,9 +10838,17 @@ static int btf_struct_ops_register(struct btf *btf, struct bpf_struct_ops *st_op
 
 int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops)
 {
+	struct btf_deferred_reg tmpl = {
+		.kind = BTF_DEFERRED_STRUCT_OPS,
+		.st_ops = st_ops,
+	};
 	struct btf *btf;
 	int err;
 
+	err = btf_defer_reg(st_ops->owner, &tmpl);
+	if (err)
+		return err > 0 ? 0 : err;
+
 	btf = btf_get_module_btf(st_ops->owner);
 	if (!btf)
 		return check_btf_kconfigs(st_ops->owner, "struct_ops");
@@ -10765,8 +10860,178 @@ int __register_bpf_struct_ops(struct bpf_struct_ops *st_ops)
 	return err;
 }
 EXPORT_SYMBOL_GPL(__register_bpf_struct_ops);
+#elif defined(BTF_MODULE_NOTIFIER)
+static int btf_struct_ops_register(struct btf *btf, struct bpf_struct_ops *st_ops)
+{
+	return -EOPNOTSUPP;
+}
 #endif
 
+/*
+ * CONFIG_DEBUG_INFO_BTF=m: registrations for vmlinux made before its BTF is
+ * available wait in btf_vmlinux_deferred_regs until btf_parse_vmlinux()
+ * applies them.
+ */
+#ifdef BTF_MODULE_NOTIFIER
+/*
+ * The queue has its own lock: it is drained under btf_vmlinux_lock, and
+ * with its own lock btf_module_mutex, which the module notifier takes, is
+ * never taken under btf_vmlinux_lock, so the two stay unordered.
+ */
+static DEFINE_MUTEX(btf_vmlinux_regs_mutex);
+static LIST_HEAD(btf_vmlinux_deferred_regs);
+/* Set when the vmlinux BTF is parsed; new registrations apply directly */
+static bool btf_vmlinux_regs_closed;
+
+/*
+ * Queue @tmpl if the BTF for @owner is not available yet.  Returns 1 if the
+ * registration was queued and is to be considered done, 0 if the caller has
+ * to apply it, or -ENOMEM.  Only vmlinux registrations are queued so far.
+ */
+static int btf_defer_reg(struct module *owner, const struct btf_deferred_reg *tmpl)
+{
+	struct list_head *head = NULL;
+	struct btf_deferred_reg *reg;
+
+	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF) || owner)
+		return 0;
+
+	guard(mutex)(&btf_vmlinux_regs_mutex);
+	if (!btf_vmlinux_regs_closed)
+		head = &btf_vmlinux_deferred_regs;
+	if (!head)
+		return 0;
+
+	reg = kmemdup(tmpl, sizeof(*reg), GFP_KERNEL);
+	if (!reg)
+		return -ENOMEM;
+	/*
+	 * kfunc id sets and struct_ops are static data of their owner, but
+	 * the dtor arrays are commonly built on the stack of the initcall.
+	 */
+	if (reg->kind == BTF_DEFERRED_DTOR_KFUNCS) {
+		reg->dtor.dtors = kmemdup_array(tmpl->dtor.dtors, tmpl->dtor.cnt,
+						sizeof(*tmpl->dtor.dtors), GFP_KERNEL);
+		if (!reg->dtor.dtors) {
+			kfree(reg);
+			return -ENOMEM;
+		}
+	}
+	list_add_tail(&reg->list, head);
+	return 1;
+}
+
+static void btf_free_deferred_reg(struct btf_deferred_reg *reg)
+{
+	if (reg->kind == BTF_DEFERRED_DTOR_KFUNCS)
+		kfree(reg->dtor.dtors);
+	kfree(reg);
+}
+
+static const char *btf_deferred_reg_name(const struct btf_deferred_reg *reg)
+{
+	switch (reg->kind) {
+	case BTF_DEFERRED_KFUNC_SET:	return "kfunc set";
+	case BTF_DEFERRED_DTOR_KFUNCS:	return "dtor kfuncs";
+	case BTF_DEFERRED_STRUCT_OPS:	return "struct_ops";
+	}
+	return "?";
+}
+
+static int btf_apply_deferred_reg(struct btf *btf, const struct btf_deferred_reg *reg)
+{
+	switch (reg->kind) {
+	case BTF_DEFERRED_KFUNC_SET:
+		return btf_kfunc_id_set_add(btf, reg->kfunc.hook, reg->kfunc.kset);
+	case BTF_DEFERRED_DTOR_KFUNCS:
+		return btf_dtor_kfuncs_add(btf, reg->dtor.dtors, reg->dtor.cnt);
+	case BTF_DEFERRED_STRUCT_OPS:
+		return btf_struct_ops_register(btf, reg->st_ops);
+	}
+	return -EINVAL;
+}
+
+/* Apply and free the registrations in @regs to @btf. */
+static void btf_apply_deferred_regs(struct btf *btf, struct list_head *regs)
+{
+	struct btf_deferred_reg *reg, *tmp;
+	int err;
+
+	list_for_each_entry_safe(reg, tmp, regs, list) {
+		err = btf_apply_deferred_reg(btf, reg);
+		if (err)
+			pr_warn("failed to register deferred %s for [%s] BTF: %d\n",
+				btf_deferred_reg_name(reg), btf->name, err);
+		list_del(&reg->list);
+		btf_free_deferred_reg(reg);
+	}
+}
+
+/*
+ * The vmlinux BTF has just been parsed; apply the registrations that waited
+ * for it.  Runs under btf_vmlinux_lock, before @btf is published, so nothing
+ * can observe a vmlinux BTF without its kfuncs and struct_ops.
+ *
+ * Applying a registration can queue further ones: a struct_ops ->init()
+ * registers the kfuncs of its hook.  Those cannot be applied directly, the
+ * BTF is only published once we are done, so the queue stays open until a
+ * pass applies nothing new, and only then are registrations applied directly.
+ */
+static void btf_apply_deferred_vmlinux_regs(struct btf *btf)
+{
+	LIST_HEAD(regs);
+
+	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF))
+		return;
+
+	mutex_lock(&btf_vmlinux_regs_mutex);
+	while (!list_empty(&btf_vmlinux_deferred_regs)) {
+		list_splice_init(&btf_vmlinux_deferred_regs, &regs);
+		mutex_unlock(&btf_vmlinux_regs_mutex);
+		btf_apply_deferred_regs(btf, &regs);
+		mutex_lock(&btf_vmlinux_regs_mutex);
+	}
+	btf_vmlinux_regs_closed = true;
+	mutex_unlock(&btf_vmlinux_regs_mutex);
+}
+
+/*
+ * The vmlinux BTF will not become available: free the queued registrations
+ * and close the queue, so that later ones fail as they do with =y.
+ */
+static void btf_drop_deferred_vmlinux_regs(void)
+{
+	struct btf_deferred_reg *reg, *tmp;
+	LIST_HEAD(regs);
+
+	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF))
+		return;
+
+	mutex_lock(&btf_vmlinux_regs_mutex);
+	list_splice_init(&btf_vmlinux_deferred_regs, &regs);
+	btf_vmlinux_regs_closed = true;
+	mutex_unlock(&btf_vmlinux_regs_mutex);
+
+	list_for_each_entry_safe(reg, tmp, &regs, list) {
+		list_del(&reg->list);
+		btf_free_deferred_reg(reg);
+	}
+}
+#else
+static int btf_defer_reg(struct module *owner, const struct btf_deferred_reg *tmpl)
+{
+	return 0;
+}
+
+static void btf_apply_deferred_vmlinux_regs(struct btf *btf)
+{
+}
+
+static void btf_drop_deferred_vmlinux_regs(void)
+{
+}
+#endif /* BTF_MODULE_NOTIFIER */
+
 bool btf_param_match_suffix(const struct btf *btf,
 			    const struct btf_param *arg,
 			    const char *suffix)
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (5 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 23:45   ` bot+bpf-ci
  2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
                   ` (5 subsequent siblings)
  12 siblings, 1 reply; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Module BTF is split BTF against the vmlinux BTF and is parsed in the
module notifier.  With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF may not be
loaded yet when a module loads, and the notifier cannot load btf_vmlinux
(that would nest a module load in a module load).

So a module loaded before the vmlinux BTF keeps a copy of its .BTF and
.BTF.base and gets a list entry with btf == NULL; its kfunc, dtor kfunc
and struct_ops registrations wait on that entry.  Without a .BTF.base
the data is final and is exposed in /sys/kernel/btf right away (the raw
bytes need no parsing); with one, parsing relocates the data in place,
so its file is created once parsed, as with =y.  The next patch creates
that file earlier.

When the vmlinux BTF arrives, btf_parse_deferred_modules() parses the
kept copies (the copy is the one btf_parse_module() makes anyway, so an
existing sysfs file keeps pointing at valid data), applies the waiting
registrations and only then publishes the BTF, so nobody sees a module
BTF without its kfuncs; its id is reserved before the registrations are
applied and installed after.  Applying walks the module list
(btf_check_kfunc_name()) and so happens with btf_module_mutex dropped
and the module pinned, and it loops, as for vmlinux, because a struct_ops
->init() registers kfuncs in turn.  A module that is still initializing
when the vmlinux BTF arrives is only published; its init is still
queueing registrations, later ones queue behind them so that the order
is kept, and MODULE_STATE_LIVE applies them once init is done, before
the module counts as live, as with =y.  This also keeps a module whose
init fails from being touched after it is freed.

bpf_load_btf_vmlinux() returns only once the kept module BTF is
registered, also to callers that come while another one is doing it.
Until then a search of the module BTFs by name (bpf_find_btf_id(), and
in-kernel CO-RE candidate search) may miss a module that is loaded;
rather than taking a missing type for a local one, or caching the
candidates, it counts a vmlinux BTF miss and fails, and bpf() runs the
command again after waiting.

A module whose BTF cannot be kept for lack of memory fails to load, as
with =y when its BTF fails to parse, unless
CONFIG_MODULE_ALLOW_BTF_MISMATCH.

A module whose BTF turns out to mismatch at that point is already
running and keeps running without BTF, with a warning; its entry stays,
dead, until the module goes, and keeps the raw data a sysfs file may
serve.  With =y such a module would have been refused at load time
unless CONFIG_MODULE_ALLOW_BTF_MISMATCH; that check only applies to
modules loaded after the vmlinux BTF.  If the vmlinux BTF itself fails
to parse for good, every kept module becomes such a dead entry.

Walkers of the module BTF list skip entries whose BTF is not parsed yet.
With =y the BTF is present from boot and the notifier takes the existing
path.  Still nothing is reachable until the Kconfig symbol becomes a
tristate.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 include/linux/bpf.h   |   5 +
 include/linux/btf.h   |  10 ++
 kernel/bpf/btf.c      | 333 +++++++++++++++++++++++++++++++++++++++---
 kernel/bpf/verifier.c |  39 +++--
 4 files changed, 359 insertions(+), 28 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index d812dbc683ae..f6634600467e 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -3187,11 +3187,16 @@ struct btf *bpf_peek_btf_vmlinux(void);
 struct btf *bpf_load_btf_vmlinux(void);
 #if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
 unsigned int bpf_btf_vmlinux_misses(void);
+void bpf_btf_vmlinux_miss(void);
 #else
 static inline unsigned int bpf_btf_vmlinux_misses(void)
 {
 	return 0;
 }
+
+static inline void bpf_btf_vmlinux_miss(void)
+{
+}
 #endif
 
 /* Map specifics */
diff --git a/include/linux/btf.h b/include/linux/btf.h
index 81e6c65fe5f6..b546531dd6f4 100644
--- a/include/linux/btf.h
+++ b/include/linux/btf.h
@@ -603,6 +603,16 @@ const char *btf_name_by_offset(const struct btf *btf, u32 offset);
 const char *btf_str_by_offset(const struct btf *btf, u32 offset);
 struct btf *btf_parse_vmlinux(void);
 void *btf_vmlinux_data(u32 *size, bool load);
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+void btf_parse_deferred_modules(void);
+bool btf_deferred_modules_pending(void);
+#else
+static inline void btf_parse_deferred_modules(void) {}
+static inline bool btf_deferred_modules_pending(void)
+{
+	return false;
+}
+#endif
 struct btf *bpf_prog_get_target_btf(const struct bpf_prog *prog);
 u32 *btf_kfunc_flags(const struct btf *btf, u32 kfunc_btf_id, const struct bpf_prog *prog);
 int btf_kfunc_check_flag(const struct btf *btf, u32 kfunc_btf_id, u32 flag);
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 96241dc62dc3..4d51fb212218 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -719,6 +719,12 @@ s32 bpf_find_btf_id(const char *name, u32 kind, struct btf **btf_p)
 		return ret;
 	}
 
+	/* the modules loaded before the vmlinux BTF may not be registered yet */
+	if (btf_deferred_modules_pending()) {
+		bpf_btf_vmlinux_miss();
+		return -EINVAL;
+	}
+
 	/* If name is not found in vmlinux's BTF then search in module's BTFs */
 	spin_lock_bh(&btf_idr_lock);
 	idr_for_each_entry(&btf_idr, btf, id) {
@@ -9155,7 +9161,7 @@ enum {
 /*
  * CONFIG_DEBUG_INFO_BTF=m: a kfunc, dtor kfunc or struct_ops registration
  * made while the BTF it applies to is not available yet.  Kept until the BTF
- * arrives, see btf_defer_reg().
+ * arrives, see btf_defer_reg() and btf_apply_deferred_regs().
  */
 enum btf_deferred_reg_kind {
 	BTF_DEFERRED_KFUNC_SET,
@@ -9180,12 +9186,30 @@ struct btf_deferred_reg {
 };
 
 #ifdef BTF_MODULE_NOTIFIER
+static void btf_free_deferred_regs(struct list_head *regs);
+static void btf_apply_deferred_regs(struct btf *btf, struct list_head *regs);
+
 struct btf_module {
 	struct list_head list;
 	struct module *module;
 	struct btf *btf;
 	struct bin_attribute *sysfs_attr;
 	int flags;
+	/*
+	 * CONFIG_DEBUG_INFO_BTF=m: a module loaded before the vmlinux BTF is
+	 * available cannot have its BTF parsed yet.  Its .BTF and .BTF.base
+	 * sections are copied here and parsed once the vmlinux BTF arrives
+	 * (btf_parse_deferred_modules()); @btf is NULL until then.
+	 * Registrations of the module's kfuncs, dtor kfuncs and struct_ops
+	 * wait in @deferred_regs.
+	 */
+	void *data;
+	void *base_data;
+	u32 data_size;
+	u32 base_data_size;
+	struct list_head deferred_regs;
+	/* the kept BTF turned out unusable; the entry stays until the module goes */
+	bool gone;
 };
 
 static LIST_HEAD(btf_modules);
@@ -9229,8 +9253,14 @@ static void btf_module_free(struct btf_module *btf_mod)
 {
 	if (btf_mod->sysfs_attr)
 		sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
-	purge_cand_cache(btf_mod->btf);
-	btf_put(btf_mod->btf);
+	if (btf_mod->btf) {
+		purge_cand_cache(btf_mod->btf);
+		btf_put(btf_mod->btf);
+	} else {
+		kvfree(btf_mod->data);
+		kvfree(btf_mod->base_data);
+	}
+	btf_free_deferred_regs(&btf_mod->deferred_regs);
 	kfree(btf_mod->sysfs_attr);
 	kfree(btf_mod);
 }
@@ -9280,6 +9310,66 @@ static bool btf_is_vmlinux_carrier(const struct module *mod)
 {
 	return !strcmp(mod->name, btf_vmlinux_link.module_name);
 }
+
+/*
+ * The vmlinux BTF is not available yet and must not be loaded from the
+ * module notifier (that would nest a module load into a module load).  Keep
+ * the module's BTF for btf_parse_deferred_modules().
+ *
+ * Without a .BTF.base section the .BTF data is final and can be exposed in
+ * sysfs right away, it needs no parsing.  With one, parsing relocates the
+ * data in place against the vmlinux BTF, so the file is created afterwards,
+ * as with =y where it also only appears once the BTF is parsed.
+ */
+static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
+{
+	btf_mod->data = kvmemdup(mod->btf_data, mod->btf_data_size,
+				 GFP_KERNEL | __GFP_NOWARN);
+	if (!btf_mod->data)
+		return -ENOMEM;
+	btf_mod->data_size = mod->btf_data_size;
+
+	if (mod->btf_base_data) {
+		btf_mod->base_data = kvmemdup(mod->btf_base_data,
+					      mod->btf_base_data_size,
+					      GFP_KERNEL | __GFP_NOWARN);
+		if (!btf_mod->base_data) {
+			kvfree(btf_mod->data);
+			return -ENOMEM;
+		}
+		btf_mod->base_data_size = mod->btf_base_data_size;
+	} else {
+		/* not fatal, the module BTF is usable without the sysfs file */
+		btf_module_sysfs_add(btf_mod, mod->name, btf_mod->data,
+				     btf_mod->data_size);
+	}
+
+	list_add(&btf_mod->list, &btf_modules);
+	return 0;
+}
+
+/*
+ * Apply the registrations queued for @btf_mod to @btf, and those that
+ * applying them queues in turn: a struct_ops ->init() registers the kfuncs
+ * of its hook.  Called and returns with btf_module_mutex held, which is
+ * dropped while applying, as that walks btf_modules
+ * (btf_check_kfunc_name()); the module is pinned meanwhile, so the entry
+ * stays.  A module that is going has its queue freed with the entry.
+ */
+static void btf_module_apply_regs(struct btf_module *btf_mod, struct btf *btf)
+{
+	LIST_HEAD(regs);
+
+	if (list_empty(&btf_mod->deferred_regs) || !try_module_get(btf_mod->module))
+		return;
+	while (!list_empty(&btf_mod->deferred_regs)) {
+		list_splice_init(&btf_mod->deferred_regs, &regs);
+		mutex_unlock(&btf_module_mutex);
+		btf_apply_deferred_regs(btf, &regs);
+		mutex_lock(&btf_module_mutex);
+	}
+	module_put(btf_mod->module);
+}
 #else
 static int btf_vmlinux_module_coming(struct module *mod)
 {
@@ -9290,6 +9380,15 @@ static bool btf_is_vmlinux_carrier(const struct module *mod)
 {
 	return false;
 }
+
+static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
+{
+	return 0;
+}
+
+static void btf_module_apply_regs(struct btf_module *btf_mod, struct btf *btf)
+{
+}
 #endif
 
 static int btf_module_notify(struct notifier_block *nb, unsigned long op,
@@ -9321,6 +9420,26 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			goto out;
 		}
 		btf_mod->module = module;
+		INIT_LIST_HEAD(&btf_mod->deferred_regs);
+
+		if (IS_MODULE(CONFIG_DEBUG_INFO_BTF)) {
+			mutex_lock(&btf_module_mutex);
+			/* Pairs with the publication in bpf_load_btf_vmlinux() */
+			if (!smp_load_acquire(&btf_vmlinux)) {
+				err = btf_module_defer(btf_mod, mod);
+				mutex_unlock(&btf_module_mutex);
+				if (err) {
+					pr_warn("failed to keep module [%s] BTF: %d\n",
+						mod->name, err);
+					kfree(btf_mod);
+					/* as a module BTF that fails to parse with =y */
+					if (IS_ENABLED(CONFIG_MODULE_ALLOW_BTF_MISMATCH))
+						err = 0;
+				}
+				goto out;
+			}
+			mutex_unlock(&btf_module_mutex);
+		}
 
 		btf = btf_parse_module(mod->name, bpf_get_btf_vmlinux(),
 				       mod->btf_data, mod->btf_data_size, false,
@@ -9358,6 +9477,16 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			if (btf_mod->module != module)
 				continue;
 
+			/*
+			 * The vmlinux BTF arrived while this module was
+			 * initializing: btf_parse_deferred_modules() parsed its
+			 * BTF but left the registrations its init queued to us.
+			 * Apply them before the module counts as live: with =y
+			 * all of them are done by then, and nothing may use the
+			 * module's kfuncs or struct_ops while they are added.
+			 */
+			if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && btf_mod->btf)
+				btf_module_apply_regs(btf_mod, btf_mod->btf);
 			btf_mod->flags |= BTF_MODULE_F_LIVE;
 			break;
 		}
@@ -9375,7 +9504,8 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			 * btf_try_get_module() on such BTFs will fail. This may
 			 * be called again on btf_put(), but it's ok to do so.
 			 */
-			btf_free_id(btf_mod->btf);
+			if (btf_mod->btf)
+				btf_free_id(btf_mod->btf);
 			list_del(&btf_mod->list);
 			btf_module_free(btf_mod);
 			break;
@@ -9398,6 +9528,127 @@ static int __init btf_module_init(void)
 }
 
 fs_initcall(btf_module_init);
+
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+/* The modules kept aside have been registered, see btf_parse_deferred_modules() */
+static bool btf_deferred_modules_done;
+
+/*
+ * A kept module whose BTF cannot be used after all.  The module is loaded
+ * and stays, so there is no way to reject it: the entry stays on the list,
+ * dead, until the module goes.  A sysfs file it has keeps serving the raw
+ * data, which is kept for that.
+ */
+static void btf_module_dead(struct btf_module *btf_mod, const char *what, int err)
+{
+	pr_warn("failed to %s module [%s] BTF: %d\n", what, btf_mod->module->name, err);
+	kvfree(btf_mod->base_data);
+	btf_mod->base_data = NULL;
+	btf_free_deferred_regs(&btf_mod->deferred_regs);
+	btf_mod->gone = true;
+}
+
+/*
+ * CONFIG_DEBUG_INFO_BTF=m: the vmlinux BTF has just become available.  Parse
+ * the BTF of the modules that were loaded before it, and apply the
+ * registrations that waited for them.  Called from bpf_load_btf_vmlinux()
+ * once btf_vmlinux is published, serialized by it, with no locks held.
+ *
+ * A module's BTF is published (btf_mod->btf set, id installed) only after
+ * its queued registrations are applied, so nobody sees a module BTF without
+ * its kfuncs and struct_ops, as with the vmlinux BTF; its id is reserved
+ * before, so that nothing can fail once they are.  Applying walks
+ * btf_modules (btf_check_kfunc_name()) and so needs the mutex dropped; the
+ * module is pinned for that, and the scan restarts afterwards.  A module
+ * that is still initializing is only published: its init is still queueing
+ * registrations, and MODULE_STATE_LIVE applies them once it is done.
+ *
+ * Until this has run, a search of the module BTFs may miss one of these
+ * modules: see btf_deferred_modules_pending().
+ */
+void btf_parse_deferred_modules(void)
+{
+	/* Pairs with the publication in bpf_load_btf_vmlinux() */
+	struct btf *vmlinux_btf = smp_load_acquire(&btf_vmlinux);
+	struct btf_module *btf_mod;
+	bool parsed = false;
+	struct btf *btf;
+	int err;
+
+	if (!vmlinux_btf || !btf_deferred_modules_pending())
+		return;
+
+	mutex_lock(&btf_module_mutex);
+	if (IS_ERR(vmlinux_btf)) {
+		/* remembered failure: the kept modules can never be parsed */
+		list_for_each_entry(btf_mod, &btf_modules, list) {
+			if (!btf_mod->btf && !btf_mod->gone)
+				btf_module_dead(btf_mod, "parse", PTR_ERR(vmlinux_btf));
+		}
+		mutex_unlock(&btf_module_mutex);
+		goto done;
+	}
+restart:
+	list_for_each_entry(btf_mod, &btf_modules, list) {
+		if (btf_mod->btf || btf_mod->gone)
+			continue;
+
+		btf = btf_parse_module(btf_mod->module->name, vmlinux_btf,
+				       btf_mod->data, btf_mod->data_size, true,
+				       btf_mod->base_data, btf_mod->base_data_size);
+		if (IS_ERR(btf)) {
+			/* on failure the caller keeps the data */
+			btf_module_dead(btf_mod, "validate", PTR_ERR(btf));
+			continue;
+		}
+		/* btf->data is btf_mod->data now, the sysfs file keeps pointing at valid data */
+		kvfree(btf_mod->base_data);
+		btf_mod->base_data = NULL;
+
+		err = btf_reserve_id(btf);
+		if (err) {
+			/* give the data back to the entry, the sysfs file may serve it */
+			btf->data = NULL;
+			btf_free(btf);
+			btf_module_dead(btf_mod, "register", err);
+			continue;
+		}
+
+		if (btf_mod->flags & BTF_MODULE_F_LIVE)
+			btf_module_apply_regs(btf_mod, btf);
+
+		/* modules with .BTF.base get their sysfs file now, the data is relocated */
+		if (!btf_mod->sysfs_attr)
+			btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
+		btf_mod->data = NULL;
+		btf_mod->btf = btf;
+		btf_install_id(btf);
+		parsed = true;
+		/* the list may have changed while the mutex was dropped */
+		goto restart;
+	}
+	mutex_unlock(&btf_module_mutex);
+
+	if (parsed)
+		purge_cand_cache(NULL);
+done:
+	/* Pairs with the smp_load_acquire() in btf_deferred_modules_pending() */
+	smp_store_release(&btf_deferred_modules_done, true);
+}
+
+/*
+ * The vmlinux BTF is published before the BTF of the modules loaded ahead of
+ * it is registered by btf_parse_deferred_modules().  Until then, a search of
+ * the module BTFs may miss a module: such searches count a miss and fail, so
+ * that bpf() runs the command again after bpf_load_btf_vmlinux(), which
+ * waits for the modules.
+ */
+bool btf_deferred_modules_pending(void)
+{
+	/* Pairs with the smp_store_release() in btf_parse_deferred_modules() */
+	return !smp_load_acquire(&btf_deferred_modules_done);
+}
+#endif /* IS_MODULE(CONFIG_DEBUG_INFO_BTF) */
 #endif /* BTF_MODULE_NOTIFIER */
 
 struct module *btf_try_get_module(const struct btf *btf)
@@ -9450,8 +9701,11 @@ struct btf *btf_get_module_btf(const struct module *module)
 		if (btf_mod->module != module)
 			continue;
 
-		btf_get(btf_mod->btf);
-		btf = btf_mod->btf;
+		/* NULL while waiting for the vmlinux BTF (CONFIG_DEBUG_INFO_BTF=m) */
+		if (btf_mod->btf) {
+			btf_get(btf_mod->btf);
+			btf = btf_mod->btf;
+		}
 		break;
 	}
 	mutex_unlock(&btf_module_mutex);
@@ -9634,7 +9888,8 @@ static int btf_check_kfunc_name(struct btf *btf, const char *func_name, u32 kind
 #ifdef CONFIG_DEBUG_INFO_BTF_MODULES
 	guard(mutex)(&btf_module_mutex);
 	list_for_each_entry_safe(btf_mod, tmp, &btf_modules, list) {
-		if (btf_mod->btf == btf)
+		/* skip ourselves and, with CONFIG_DEBUG_INFO_BTF=m, unparsed BTF */
+		if (btf_mod->btf == btf || !btf_mod->btf)
 			continue;
 		id = btf_find_by_name_kind(btf_mod->btf, func_name, kind);
 		if (id >= 0) {
@@ -10492,6 +10747,12 @@ bpf_core_find_cands(struct bpf_core_ctx *ctx, u32 local_type_id)
 		return cc;
 
 check_modules:
+	/* the modules loaded before the vmlinux BTF may not be registered yet */
+	if (btf_deferred_modules_pending()) {
+		bpf_btf_vmlinux_miss();
+		return ERR_PTR(-EINVAL);
+	}
+
 	/* cands is a pointer to stack here and cands->cnt == 0 */
 	cc = check_cand_cache(cands, module_cand_cache, MODULE_CAND_CACHE_SIZE);
 	if (cc)
@@ -10868,15 +11129,17 @@ static int btf_struct_ops_register(struct btf *btf, struct bpf_struct_ops *st_op
 #endif
 
 /*
- * CONFIG_DEBUG_INFO_BTF=m: registrations for vmlinux made before its BTF is
- * available wait in btf_vmlinux_deferred_regs until btf_parse_vmlinux()
- * applies them.
+ * CONFIG_DEBUG_INFO_BTF=m: registrations made before the BTF they apply to
+ * is available.  Registrations for vmlinux wait in btf_vmlinux_deferred_regs
+ * until btf_parse_vmlinux() applies them; registrations for a module wait in
+ * its struct btf_module, under btf_module_mutex, until
+ * btf_parse_deferred_modules() does.
  */
 #ifdef BTF_MODULE_NOTIFIER
 /*
- * The queue has its own lock: it is drained under btf_vmlinux_lock, and
- * with its own lock btf_module_mutex, which the module notifier takes, is
- * never taken under btf_vmlinux_lock, so the two stay unordered.
+ * The vmlinux queue has its own lock: it is drained under btf_vmlinux_lock,
+ * and with its own lock btf_module_mutex, which the module notifier takes,
+ * is never taken under btf_vmlinux_lock, so the two stay unordered.
  */
 static DEFINE_MUTEX(btf_vmlinux_regs_mutex);
 static LIST_HEAD(btf_vmlinux_deferred_regs);
@@ -10886,19 +11149,38 @@ static bool btf_vmlinux_regs_closed;
 /*
  * Queue @tmpl if the BTF for @owner is not available yet.  Returns 1 if the
  * registration was queued and is to be considered done, 0 if the caller has
- * to apply it, or -ENOMEM.  Only vmlinux registrations are queued so far.
+ * to apply it, or -ENOMEM.
  */
 static int btf_defer_reg(struct module *owner, const struct btf_deferred_reg *tmpl)
 {
 	struct list_head *head = NULL;
 	struct btf_deferred_reg *reg;
+	struct btf_module *btf_mod;
 
-	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF) || owner)
+	if (!IS_MODULE(CONFIG_DEBUG_INFO_BTF))
 		return 0;
 
-	guard(mutex)(&btf_vmlinux_regs_mutex);
-	if (!btf_vmlinux_regs_closed)
-		head = &btf_vmlinux_deferred_regs;
+	guard(mutex)(owner ? &btf_module_mutex : &btf_vmlinux_regs_mutex);
+	if (!owner) {
+		if (!btf_vmlinux_regs_closed)
+			head = &btf_vmlinux_deferred_regs;
+	} else {
+		list_for_each_entry(btf_mod, &btf_modules, list) {
+			if (btf_mod->module != owner)
+				continue;
+			/*
+			 * Wait until the entry's BTF is published, and after that
+			 * behind registrations that are still queued (a module
+			 * that was initializing when its BTF was published), so
+			 * that they are applied in order.  A dead entry has no BTF
+			 * to register with, as with =y.
+			 */
+			if (!btf_mod->gone &&
+			    (!btf_mod->btf || !list_empty(&btf_mod->deferred_regs)))
+				head = &btf_mod->deferred_regs;
+			break;
+		}
+	}
 	if (!head)
 		return 0;
 
@@ -10951,7 +11233,10 @@ static int btf_apply_deferred_reg(struct btf *btf, const struct btf_deferred_reg
 	return -EINVAL;
 }
 
-/* Apply and free the registrations in @regs to @btf. */
+/*
+ * Apply and free the registrations in @regs to @btf.  For a module BTF the
+ * caller holds a reference on @btf and makes sure the owning module stays.
+ */
 static void btf_apply_deferred_regs(struct btf *btf, struct list_head *regs)
 {
 	struct btf_deferred_reg *reg, *tmp;
@@ -10967,6 +11252,16 @@ static void btf_apply_deferred_regs(struct btf *btf, struct list_head *regs)
 	}
 }
 
+static void btf_free_deferred_regs(struct list_head *regs)
+{
+	struct btf_deferred_reg *reg, *tmp;
+
+	list_for_each_entry_safe(reg, tmp, regs, list) {
+		list_del(&reg->list);
+		btf_free_deferred_reg(reg);
+	}
+}
+
 /*
  * The vmlinux BTF has just been parsed; apply the registrations that waited
  * for it.  Runs under btf_vmlinux_lock, before @btf is published, so nothing
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 895e7feb6429..f6205894fe11 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -21935,18 +21935,22 @@ struct btf *bpf_peek_btf_vmlinux(void)
  * bpf_load_btf_vmlinux - get the vmlinux BTF, loading it if necessary
  *
  * Like bpf_get_btf_vmlinux(), but with CONFIG_DEBUG_INFO_BTF=m it loads the
- * btf_vmlinux module if the BTF is not there yet and parses it.  Loading the
- * module waits for user space (modprobe), and the notifiers of the module
- * load take locks of their own, event_mutex among them.  So this is only
- * called at the start of a request from user space, in process context,
- * holding no lock that loading a module may need; everything else uses
- * bpf_get_btf_vmlinux() or bpf_peek_btf_vmlinux().
+ * btf_vmlinux module if the BTF is not there yet, parses it, and registers
+ * the BTF of the modules that were loaded before it; it returns once all of
+ * that is done, also to a caller that comes while another one is doing it.
+ * Loading the module waits for user space (modprobe), and the notifiers of
+ * the module load take locks of their own, event_mutex among them.  So this
+ * is only called at the start of a request from user space, in process
+ * context, holding no lock that loading a module may need; everything else
+ * uses bpf_get_btf_vmlinux() or bpf_peek_btf_vmlinux().
  *
  * If the module cannot be loaded, returns NULL like a kernel without BTF;
  * the next call tries again.
  */
 struct btf *bpf_load_btf_vmlinux(void)
 {
+	/* Held from loading until the module BTF kept aside is registered */
+	static DEFINE_MUTEX(load_mutex);
 	struct btf *btf;
 	u32 size;
 
@@ -21956,13 +21960,14 @@ struct btf *bpf_load_btf_vmlinux(void)
 
 	/* Pairs with the smp_store_release() below */
 	btf = smp_load_acquire(&btf_vmlinux);
-	if (btf)
+	if (btf && !btf_deferred_modules_pending())
 		return btf;
 
-	/* Outside btf_vmlinux_lock, the module's notifier must not wait for us */
-	if (!btf_vmlinux_data(&size, true))
+	/* Outside the locks, the module's notifier must not wait for us */
+	if (!btf && !btf_vmlinux_data(&size, true))
 		return NULL;
 
+	mutex_lock(&load_mutex);
 	mutex_lock(&btf_vmlinux_lock);
 	btf = btf_vmlinux;
 	if (!btf) {
@@ -21974,12 +21979,22 @@ struct btf *bpf_load_btf_vmlinux(void)
 		 */
 		if (IS_ERR(btf) && PTR_ERR(btf) == -ENOMEM) {
 			mutex_unlock(&btf_vmlinux_lock);
+			mutex_unlock(&load_mutex);
 			return btf;
 		}
 		/* As in bpf_get_btf_vmlinux(): publish after the parse */
 		smp_store_release(&btf_vmlinux, btf);
 	}
 	mutex_unlock(&btf_vmlinux_lock);
+
+	/*
+	 * Until this is done, the module BTFs may lack a module, which
+	 * btf_deferred_modules_pending() tells the searches of module BTFs,
+	 * and which is why concurrent callers wait for it on the mutex.
+	 */
+	btf_parse_deferred_modules();
+	mutex_unlock(&load_mutex);
+
 	return btf;
 }
 
@@ -21994,6 +22009,12 @@ unsigned int bpf_btf_vmlinux_misses(void)
 {
 	return atomic_read(&btf_vmlinux_misses);
 }
+
+/* A lookup that needs the BTF of a module that is not registered yet */
+void bpf_btf_vmlinux_miss(void)
+{
+	atomic_inc(&btf_vmlinux_misses);
+}
 #endif
 
 /*
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (6 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
                   ` (4 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

A module with a .BTF.base section (built out of tree) that is loaded
before the vmlinux BTF only gets its /sys/kernel/btf file once the
vmlinux BTF has been loaded and its BTF relocated, because the raw .BTF
is only valid against the distilled base and relocation rewrites it in
place.  Until then the module is missing from /sys/kernel/btf, unlike
with =y, and reading its file cannot trigger the load.

Create the file at module load instead, with its final size: relocation
only rewrites type ids and string offsets, never the length.  Its reader,
btf_module_sysfs_read_deferred(), has the vmlinux BTF loaded, which parses
and relocates the kept modules, then waits until this module's BTF is
published (btf_mod->ready) before serving it, so no unrelocated or
half-relocated data is ever visible.  If the module goes away or its BTF
turns out unusable first (btf_mod->gone), or the vmlinux BTF cannot be
loaded, the read fails with -ENODEV.

The reader does not load the vmlinux BTF itself but queues a work item
that does: it holds the file's kernfs active reference, which
MODULE_STATE_GOING waits for when it removes the file, with the module
notifier chain held; loading btf_vmlinux takes that chain, so with a
writer queued on it the two would wait for each other.  The reader's own
wait ends when the module goes.

Because that reader may be waiting for btf_parse_deferred_modules(), and
removing a sysfs file waits for its readers, nothing removes a sysfs file
from that path (a failed entry
stays dead until the module goes, which the previous patch already
arranged), and MODULE_STATE_GOING removes the file after dropping
btf_module_mutex, which the reader may need to get there.  A dead entry
keeps its raw data only for a file that serves it as is, the deferred
reader never does.

Modules without .BTF.base are unchanged: their data is final and is
served as is.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/btf.c | 144 ++++++++++++++++++++++++++++++++++++++++-------
 1 file changed, 124 insertions(+), 20 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 4d51fb212218..3c3aba0fc4cb 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -9208,17 +9208,27 @@ struct btf_module {
 	u32 data_size;
 	u32 base_data_size;
 	struct list_head deferred_regs;
-	/* the kept BTF turned out unusable; the entry stays until the module goes */
+	/*
+	 * For the sysfs reader of a module whose data is only final once its
+	 * BTF is relocated: @ready once @btf is published, @gone once the
+	 * entry is dead (parse failed or module going).  Both only ever go
+	 * from false to true; waiters sleep on btf_module_wq.
+	 */
+	bool ready;
 	bool gone;
 };
 
 static LIST_HEAD(btf_modules);
 static DEFINE_MUTEX(btf_module_mutex);
+static DECLARE_WAIT_QUEUE_HEAD(btf_module_wq);
 
 static void purge_cand_cache(struct btf *btf);
 
 static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
-				void *data, size_t data_size)
+				void *private, size_t size,
+				ssize_t (*read)(struct file *, struct kobject *,
+						const struct bin_attribute *,
+						char *, loff_t, size_t))
 {
 	struct bin_attribute *attr;
 	int err;
@@ -9233,9 +9243,9 @@ static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
 	sysfs_bin_attr_init(attr);
 	attr->attr.name = name;
 	attr->attr.mode = 0444;
-	attr->size = data_size;
-	attr->private = data;
-	attr->read = sysfs_bin_attr_simple_read;
+	attr->size = size;
+	attr->private = private;
+	attr->read = read;
 
 	err = sysfs_create_bin_file(btf_kobj, attr);
 	if (err) {
@@ -9249,8 +9259,14 @@ static int btf_module_sysfs_add(struct btf_module *btf_mod, const char *name,
 	return 0;
 }
 
+/*
+ * Called with btf_module_mutex NOT held: removing the sysfs file waits for
+ * readers to leave, and a deferred reader may need the mutex to get there.
+ */
 static void btf_module_free(struct btf_module *btf_mod)
 {
+	WRITE_ONCE(btf_mod->gone, true);
+	wake_up_all(&btf_module_wq);
 	if (btf_mod->sysfs_attr)
 		sysfs_remove_bin_file(btf_kobj, btf_mod->sysfs_attr);
 	if (btf_mod->btf) {
@@ -9311,15 +9327,87 @@ static bool btf_is_vmlinux_carrier(const struct module *mod)
 	return !strcmp(mod->name, btf_vmlinux_link.module_name);
 }
 
+/*
+ * sysfs reader for a module kept aside with a .BTF.base section: its .BTF is
+ * split against the distilled base and only becomes valid split BTF against
+ * the vmlinux BTF once relocated, which rewrites the buffer in place.  So
+ * have the vmlinux BTF loaded (which parses and relocates the kept modules),
+ * then wait until this module's BTF is published.  The size does not change:
+ * relocation only rewrites ids and string offsets.
+ *
+ * The reader does not load the vmlinux BTF itself: it holds the file's
+ * kernfs active reference, which MODULE_STATE_GOING waits for when it
+ * removes the file with the module notifier chain held, and loading
+ * btf_vmlinux needs that chain.  A work item loads it, and the reader waits
+ * in a way that the module going away (@gone) ends.
+ */
+static bool btf_module_published(struct btf_module *btf_mod)
+{
+	/* Pairs with the smp_store_release() of @ready after btf_mod->btf is set */
+	return smp_load_acquire(&btf_mod->ready);
+}
+
+/* Bumped after each load attempt by btf_vmlinux_load_work */
+static atomic_t btf_vmlinux_load_seq = ATOMIC_INIT(0);
+
+static void btf_vmlinux_load_workfn(struct work_struct *work)
+{
+	bpf_load_btf_vmlinux();
+	atomic_inc(&btf_vmlinux_load_seq);
+	wake_up_all(&btf_module_wq);
+}
+
+static DECLARE_WORK(btf_vmlinux_load_work, btf_vmlinux_load_workfn);
+
+/* The vmlinux BTF could not be had: a load attempt ended without it. */
+static bool btf_vmlinux_load_failed(int seq)
+{
+	return atomic_read(&btf_vmlinux_load_seq) != seq &&
+	       IS_ERR_OR_NULL(bpf_peek_btf_vmlinux());
+}
+
+static ssize_t btf_module_sysfs_read_deferred(struct file *filp, struct kobject *kobj,
+					      const struct bin_attribute *attr,
+					      char *buf, loff_t off, size_t count)
+{
+	struct btf_module *btf_mod = attr->private;
+	int seq = atomic_read(&btf_vmlinux_load_seq);
+	struct btf *vmlinux_btf = bpf_peek_btf_vmlinux();
+	int err;
+
+	if (IS_ERR(vmlinux_btf))
+		return -ENODEV;
+	if (!vmlinux_btf)
+		queue_work(system_unbound_wq, &btf_vmlinux_load_work);
+
+	/*
+	 * Another thread may still be relocating and publishing it; if the
+	 * module goes away or its BTF turns out unusable, btf_module_free()
+	 * or btf_parse_deferred_modules() set @gone and wake us.
+	 */
+	err = wait_event_interruptible(btf_module_wq,
+				       btf_module_published(btf_mod) ||
+				       READ_ONCE(btf_mod->gone) ||
+				       btf_vmlinux_load_failed(seq));
+	if (err)
+		return err;
+	if (!btf_module_published(btf_mod))
+		return -ENODEV;
+
+	/* sysfs clamps @off and @count to attr->size == btf->data_size */
+	memcpy(buf, btf_mod->btf->data + off, count);
+	return count;
+}
+
 /*
  * The vmlinux BTF is not available yet and must not be loaded from the
  * module notifier (that would nest a module load into a module load).  Keep
  * the module's BTF for btf_parse_deferred_modules().
  *
- * Without a .BTF.base section the .BTF data is final and can be exposed in
- * sysfs right away, it needs no parsing.  With one, parsing relocates the
- * data in place against the vmlinux BTF, so the file is created afterwards,
- * as with =y where it also only appears once the BTF is parsed.
+ * The sysfs file is created right away with its final size, as with =y.
+ * Without a .BTF.base section the .BTF data is final and is served as is;
+ * with one, it is only valid once relocated, so its reader waits for that
+ * (btf_module_sysfs_read_deferred()).
  */
 static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
 {
@@ -9338,10 +9426,12 @@ static int btf_module_defer(struct btf_module *btf_mod, struct module *mod)
 			return -ENOMEM;
 		}
 		btf_mod->base_data_size = mod->btf_base_data_size;
-	} else {
 		/* not fatal, the module BTF is usable without the sysfs file */
+		btf_module_sysfs_add(btf_mod, mod->name, btf_mod, btf_mod->data_size,
+				     btf_module_sysfs_read_deferred);
+	} else {
 		btf_module_sysfs_add(btf_mod, mod->name, btf_mod->data,
-				     btf_mod->data_size);
+				     btf_mod->data_size, sysfs_bin_attr_simple_read);
 	}
 
 	list_add(&btf_mod->list, &btf_modules);
@@ -9469,7 +9559,8 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 		mutex_unlock(&btf_module_mutex);
 
 		/* not fatal, the module BTF is usable without the sysfs file */
-		btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
+		btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size,
+				     sysfs_bin_attr_simple_read);
 		break;
 	case MODULE_STATE_LIVE:
 		mutex_lock(&btf_module_mutex);
@@ -9507,8 +9598,10 @@ static int btf_module_notify(struct notifier_block *nb, unsigned long op,
 			if (btf_mod->btf)
 				btf_free_id(btf_mod->btf);
 			list_del(&btf_mod->list);
+			mutex_unlock(&btf_module_mutex);
+			/* off the list, nobody else can find it now */
 			btf_module_free(btf_mod);
-			break;
+			goto out;
 		}
 		mutex_unlock(&btf_module_mutex);
 		break;
@@ -9536,23 +9629,34 @@ static bool btf_deferred_modules_done;
 /*
  * A kept module whose BTF cannot be used after all.  The module is loaded
  * and stays, so there is no way to reject it: the entry stays on the list,
- * dead, until the module goes.  A sysfs file it has keeps serving the raw
- * data, which is kept for that.
+ * dead, until the module goes.  Its sysfs file stays too, its reader, or
+ * the caller of this function, may be inside it right now: a .BTF.base
+ * reader wakes up and fails without touching the data, which can go; a
+ * plain one keeps serving the raw data, which is kept for that.
  */
 static void btf_module_dead(struct btf_module *btf_mod, const char *what, int err)
 {
 	pr_warn("failed to %s module [%s] BTF: %d\n", what, btf_mod->module->name, err);
+	if (!btf_mod->sysfs_attr ||
+	    btf_mod->sysfs_attr->read != sysfs_bin_attr_simple_read) {
+		kvfree(btf_mod->data);
+		btf_mod->data = NULL;
+	}
 	kvfree(btf_mod->base_data);
 	btf_mod->base_data = NULL;
 	btf_free_deferred_regs(&btf_mod->deferred_regs);
-	btf_mod->gone = true;
+	WRITE_ONCE(btf_mod->gone, true);
+	wake_up_all(&btf_module_wq);
 }
 
 /*
  * CONFIG_DEBUG_INFO_BTF=m: the vmlinux BTF has just become available.  Parse
  * the BTF of the modules that were loaded before it, and apply the
  * registrations that waited for them.  Called from bpf_load_btf_vmlinux()
- * once btf_vmlinux is published, serialized by it, with no locks held.
+ * once btf_vmlinux is published, serialized by it, with no locks held.  The
+ * sysfs reader of one of these modules may be waiting for it (see
+ * btf_module_sysfs_read_deferred()), which is why no sysfs file is removed
+ * here.
  *
  * A module's BTF is published (btf_mod->btf set, id installed) only after
  * its queued registrations are applied, so nobody sees a module BTF without
@@ -9617,12 +9721,12 @@ void btf_parse_deferred_modules(void)
 		if (btf_mod->flags & BTF_MODULE_F_LIVE)
 			btf_module_apply_regs(btf_mod, btf);
 
-		/* modules with .BTF.base get their sysfs file now, the data is relocated */
-		if (!btf_mod->sysfs_attr)
-			btf_module_sysfs_add(btf_mod, btf->name, btf->data, btf->data_size);
 		btf_mod->data = NULL;
 		btf_mod->btf = btf;
 		btf_install_id(btf);
+		/* Pairs with the smp_load_acquire() in btf_module_sysfs_read_deferred() */
+		smp_store_release(&btf_mod->ready, true);
+		wake_up_all(&btf_module_wq);
 		parsed = true;
 		/* the list may have changed while the mutex was dropped */
 		goto restart;
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (7 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
                   ` (3 subsequent siblings)
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

The next patch makes CONFIG_DEBUG_INFO_BTF a tristate.  With =m, Kconfig
defines CONFIG_DEBUG_INFO_BTF_MODULE instead of CONFIG_DEBUG_INFO_BTF,
so every check that must hold for both =y and =m has to be written for
it:

 - #ifdef CONFIG_DEBUG_INFO_BTF becomes #if IS_ENABLED(...) where the
   generated BTF and its id tables must be the same for =y and =m: the
   .BTF_ids tables (btf_ids.h), the BTF type tags (compiler_types.h), and
   the tracepoint and syscall BTF ids (trace_events.h, trace_syscalls.c).
   Leaving them would silently produce empty id sets with =m.

 - obj-$(CONFIG_DEBUG_INFO_BTF) and include-$(CONFIG_DEBUG_INFO_BTF)
   become $(subst m,y,...) where the object is built into the kernel
   regardless: sysfs_btf.o, the netfilter and xfrm kfunc objects, and
   scripts/Makefile.btf.  Otherwise =m would try to build them as
   modules (xfrm_state_bpf.o fails modpost for lack of MODULE_LICENSE)
   or skip the BTF generation flags.

 - "depends on !DEBUG_INFO_BTF" becomes "depends on DEBUG_INFO_BTF=n"
   for RUST and GENDWARFKSYMS: with =m the BTF is generated as with =y,
   so the pahole restrictions they express still apply, but !m is m,
   which a bool option takes as y.

No functional change: CONFIG_DEBUG_INFO_BTF is still a bool, for which
IS_ENABLED() and #ifdef agree, $(subst m,y,y) is y and "=n" is "!".

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 Makefile                       | 3 ++-
 include/linux/btf_ids.h        | 2 +-
 include/linux/compiler_types.h | 2 +-
 include/trace/trace_events.h   | 2 +-
 init/Kconfig                   | 2 +-
 kernel/bpf/Makefile            | 2 +-
 kernel/module/Kconfig          | 2 +-
 kernel/trace/trace_syscalls.c  | 6 +++---
 net/netfilter/Makefile         | 6 +++---
 net/xfrm/Makefile              | 4 ++--
 10 files changed, 16 insertions(+), 15 deletions(-)

diff --git a/Makefile b/Makefile
index 751a08643bf8..f561516e1735 100644
--- a/Makefile
+++ b/Makefile
@@ -1208,7 +1208,8 @@ endif
 # include additional Makefiles when needed
 include-y			:= scripts/Makefile.warn
 include-$(CONFIG_DEBUG_INFO)	+= scripts/Makefile.debug
-include-$(CONFIG_DEBUG_INFO_BTF)+= scripts/Makefile.btf
+# CONFIG_DEBUG_INFO_BTF is a tristate; BTF is generated for both y and m
+include-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += scripts/Makefile.btf
 include-$(CONFIG_KASAN)		+= scripts/Makefile.kasan
 include-$(CONFIG_KCSAN)		+= scripts/Makefile.kcsan
 include-$(CONFIG_KMSAN)		+= scripts/Makefile.kmsan
diff --git a/include/linux/btf_ids.h b/include/linux/btf_ids.h
index 8b5a9ee92513..c665afff100e 100644
--- a/include/linux/btf_ids.h
+++ b/include/linux/btf_ids.h
@@ -22,7 +22,7 @@ struct btf_id_set8 {
 	} pairs[];
 };
 
-#ifdef CONFIG_DEBUG_INFO_BTF
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF)
 
 #include <linux/compiler.h> /* for __PASTE */
 #include <linux/compiler_attributes.h> /* for __maybe_unused */
diff --git a/include/linux/compiler_types.h b/include/linux/compiler_types.h
index c5921f139007..a90a99849cee 100644
--- a/include/linux/compiler_types.h
+++ b/include/linux/compiler_types.h
@@ -34,7 +34,7 @@
  * Skipped when running bindgen due to a libclang issue;
  * see https://github.com/rust-lang/rust-bindgen/issues/2244.
  */
-#if defined(CONFIG_DEBUG_INFO_BTF) && defined(CONFIG_PAHOLE_HAS_BTF_TAG) && \
+#if IS_ENABLED(CONFIG_DEBUG_INFO_BTF) && defined(CONFIG_PAHOLE_HAS_BTF_TAG) && \
 	__has_attribute(btf_type_tag) && !defined(__BINDGEN__)
 # define BTF_TYPE_TAG(value) __attribute__((btf_type_tag(#value)))
 #else
diff --git a/include/trace/trace_events.h b/include/trace/trace_events.h
index 93011f800d0f..2a0098929771 100644
--- a/include/trace/trace_events.h
+++ b/include/trace/trace_events.h
@@ -398,7 +398,7 @@ static inline notrace int trace_event_get_offsets_##call(		\
 #define _TRACE_PERF_INIT(call)
 #endif /* CONFIG_PERF_EVENTS */
 
-#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF)
+#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)
 /*
  * Per-template BTF id list, populated at link time by resolve_btfids:
  *   [0] FUNC   __bpf_trace_<call>     (the BPF dispatcher)
diff --git a/init/Kconfig b/init/Kconfig
index 8583d9f06c52..9f9562f813c5 100644
--- a/init/Kconfig
+++ b/init/Kconfig
@@ -2257,7 +2257,7 @@ config RUST
 	depends on !MODVERSIONS || GENDWARFKSYMS
 	depends on !GCC_PLUGIN_RANDSTRUCT
 	depends on !RANDSTRUCT
-	depends on !DEBUG_INFO_BTF || (PAHOLE_HAS_LANG_EXCLUDE && !LTO)
+	depends on DEBUG_INFO_BTF=n || (PAHOLE_HAS_LANG_EXCLUDE && !LTO)
 	depends on !CFI || HAVE_CFI_ICALL_NORMALIZE_INTEGERS_RUSTC
 	select CFI_ICALL_NORMALIZE_INTEGERS if CFI
 	depends on !KASAN || CC_IS_CLANG
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..0b7db88f1bed 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -41,7 +41,7 @@ ifeq ($(CONFIG_INET),y)
 obj-$(CONFIG_BPF_SYSCALL) += reuseport_array.o
 endif
 ifeq ($(CONFIG_SYSFS),y)
-obj-$(CONFIG_DEBUG_INFO_BTF) += sysfs_btf.o
+obj-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += sysfs_btf.o
 endif
 ifeq ($(CONFIG_BPF_JIT),y)
 obj-$(CONFIG_BPF_SYSCALL) += bpf_struct_ops.o
diff --git a/kernel/module/Kconfig b/kernel/module/Kconfig
index 43b1bb01fd27..da49cb984b0d 100644
--- a/kernel/module/Kconfig
+++ b/kernel/module/Kconfig
@@ -197,7 +197,7 @@ config GENDWARFKSYMS
 	# X86, requires pahole before commit 47dcb534e253 ("btf_encoder: Stop
 	# indexing symbols for VARs") or after commit 9810758003ce ("btf_encoder:
 	# Verify 0 address DWARF variables are in ELF section").
-	depends on !X86 || !DEBUG_INFO_BTF || PAHOLE_VERSION < 128 || PAHOLE_VERSION > 129
+	depends on !X86 || DEBUG_INFO_BTF=n || PAHOLE_VERSION < 128 || PAHOLE_VERSION > 129
 	help
 	  Calculate symbol versions from DWARF debugging information using
 	  gendwarfksyms. Requires DEBUG_INFO to be enabled.
diff --git a/kernel/trace/trace_syscalls.c b/kernel/trace/trace_syscalls.c
index e35744049e3f..7a0d59c308c2 100644
--- a/kernel/trace/trace_syscalls.c
+++ b/kernel/trace/trace_syscalls.c
@@ -1304,7 +1304,7 @@ struct trace_event_functions exit_syscall_print_funcs = {
 	.trace		= print_syscall_exit,
 };
 
-#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF)
+#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)
 /* BTF id lists for the shared sys_enter/sys_exit dispatcher tracepoints. */
 BTF_ID_LIST(syscall_enter_btf_ids)
 BTF_ID(func,   __bpf_trace_sys_enter)
@@ -1321,7 +1321,7 @@ struct trace_event_class __refdata event_class_syscall_enter = {
 	.fields_array	= syscall_enter_fields_array,
 	.get_fields	= syscall_get_enter_fields,
 	.raw_init	= init_syscall_trace,
-#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF)
+#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)
 	.btf_ids	= syscall_enter_btf_ids,
 #endif
 };
@@ -1336,7 +1336,7 @@ struct trace_event_class __refdata event_class_syscall_exit = {
 	},
 	.fields		= LIST_HEAD_INIT(event_class_syscall_exit.fields),
 	.raw_init	= init_syscall_trace,
-#if defined(CONFIG_BPF_EVENTS) && defined(CONFIG_DEBUG_INFO_BTF)
+#if defined(CONFIG_BPF_EVENTS) && IS_ENABLED(CONFIG_DEBUG_INFO_BTF)
 	.btf_ids	= syscall_exit_btf_ids,
 #endif
 };
diff --git a/net/netfilter/Makefile b/net/netfilter/Makefile
index 6bf74d488a29..a2c7f00794d2 100644
--- a/net/netfilter/Makefile
+++ b/net/netfilter/Makefile
@@ -18,7 +18,7 @@ nf_conntrack-$(CONFIG_NF_CT_PROTO_GRE) += nf_conntrack_proto_gre.o
 ifeq ($(CONFIG_NF_CONNTRACK),m)
 nf_conntrack-$(CONFIG_DEBUG_INFO_BTF_MODULES) += nf_conntrack_bpf.o
 else ifeq ($(CONFIG_NF_CONNTRACK),y)
-nf_conntrack-$(CONFIG_DEBUG_INFO_BTF) += nf_conntrack_bpf.o
+nf_conntrack-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += nf_conntrack_bpf.o
 endif
 
 obj-$(CONFIG_NETFILTER) = netfilter.o
@@ -65,7 +65,7 @@ nf_nat-$(CONFIG_NF_NAT_OVS) += nf_nat_ovs.o
 ifeq ($(CONFIG_NF_NAT),m)
 nf_nat-$(CONFIG_DEBUG_INFO_BTF_MODULES) += nf_nat_bpf.o
 else ifeq ($(CONFIG_NF_NAT),y)
-nf_nat-$(CONFIG_DEBUG_INFO_BTF) += nf_nat_bpf.o
+nf_nat-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += nf_nat_bpf.o
 endif
 
 # NAT helpers
@@ -147,7 +147,7 @@ nf_flow_table-$(CONFIG_NF_FLOW_TABLE_PROCFS) += nf_flow_table_procfs.o
 ifeq ($(CONFIG_NF_FLOW_TABLE),m)
 nf_flow_table-$(CONFIG_DEBUG_INFO_BTF_MODULES) += nf_flow_table_bpf.o
 else ifeq ($(CONFIG_NF_FLOW_TABLE),y)
-nf_flow_table-$(CONFIG_DEBUG_INFO_BTF) += nf_flow_table_bpf.o
+nf_flow_table-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += nf_flow_table_bpf.o
 endif
 
 obj-$(CONFIG_NF_FLOW_TABLE_INET) += nf_flow_table_inet.o
diff --git a/net/xfrm/Makefile b/net/xfrm/Makefile
index 5a1787587cb3..b7f6e5046a0e 100644
--- a/net/xfrm/Makefile
+++ b/net/xfrm/Makefile
@@ -8,7 +8,7 @@ xfrm_interface-$(CONFIG_XFRM_INTERFACE) += xfrm_interface_core.o
 ifeq ($(CONFIG_XFRM_INTERFACE),m)
 xfrm_interface-$(CONFIG_DEBUG_INFO_BTF_MODULES) += xfrm_interface_bpf.o
 else ifeq ($(CONFIG_XFRM_INTERFACE),y)
-xfrm_interface-$(CONFIG_DEBUG_INFO_BTF) += xfrm_interface_bpf.o
+xfrm_interface-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += xfrm_interface_bpf.o
 endif
 
 obj-$(CONFIG_XFRM) := xfrm_policy.o xfrm_state.o xfrm_hash.o \
@@ -23,4 +23,4 @@ obj-$(CONFIG_XFRM_IPCOMP) += xfrm_ipcomp.o
 obj-$(CONFIG_XFRM_INTERFACE) += xfrm_interface.o
 obj-$(CONFIG_XFRM_IPTFS) += xfrm_iptfs.o
 obj-$(CONFIG_XFRM_ESPINTCP) += espintcp.o
-obj-$(CONFIG_DEBUG_INFO_BTF) += xfrm_state_bpf.o
+obj-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += xfrm_state_bpf.o
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (8 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 23:29   ` bot+bpf-ci
  2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
                   ` (2 subsequent siblings)
  12 siblings, 1 reply; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

From: Alan Maguire <alan.maguire@oracle.com>

With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is not part of the kernel
image but carried by a module, and the kernel has to know, before that
module is loaded, which module that is, how large the BTF is and which
BTF it was built with.  It keeps that in a .BTF.link record (struct
btf_link): the name of the module, NUL-padded to the target's
__MODULE_NAME_LEN, the SHA-256 of the BTF, and its size in the byte
order of the target.  Inline BTF delivered as a module, a planned
follow-up, needs the same for its own section, in .BTF.inline.link.

Fill these records in where .BTF_ids is patched, after the final link,
when the BTF is final: the --patch_btfids pass takes one or more

    --btf_link <section>:<module>:<raw BTF file>

options, and writes into the existing <section>.link section of the ELF
file the module name and the size and SHA-256 of the raw BTF file,
computed with the SHA-256 implementation of the libbpf the tool links.
The section must have the size of the record for the target's ELF
class, 92 bytes for 64-bit and 96 for 32-bit, which also catches the
kernel and the tool disagreeing on its layout.  Without --btf_link
nothing changes.

Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Assisted-by: OpenAI Codex (GPT 5.6)
[adapted subject and commit message; rebased onto bpf-next without the
 inline BTF changes; take the ELF class and byte order from the ELF
 header in patch_btf_link() rather than from elf_collect(), set up libelf
 there too, report each failure, refuse an overlong section name and
 --btf_link without --patch_btfids, document the option]
Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 tools/bpf/resolve_btfids/main.c | 217 +++++++++++++++++++++++++++++++-
 1 file changed, 213 insertions(+), 4 deletions(-)

diff --git a/tools/bpf/resolve_btfids/main.c b/tools/bpf/resolve_btfids/main.c
index 37d7e7224207..787ed5818c24 100644
--- a/tools/bpf/resolve_btfids/main.c
+++ b/tools/bpf/resolve_btfids/main.c
@@ -69,6 +69,11 @@
  *   - rewrites the prototype of KF_IMPLICIT_ARGS kfuncs.
  *
  * These kfunc annotations were historically produced by pahole.
+ *
+ * With --patch_btfids, --btf_link <section>:<module>:<raw BTF file> also fills
+ * in the <section>.link section of the ELF file, for BTF that a module carries
+ * (struct btf_link in kernel/bpf/btf.c): the module name, and the size and
+ * SHA-256 of the raw BTF.
  */
 
 #define  _GNU_SOURCE
@@ -92,8 +97,14 @@
 #include <subcmd/parse-options.h>
 
 #define BTF_IDS_SECTION	".BTF_ids"
+#define BTF_LINK_MODULE_NAME_MAX	64
+#define LIBBPF_SHA256_DIGEST_LENGTH	32
 #define BTF_ID_PREFIX	"__BTF_ID__"
 
+/* from libbpf, which this tool is statically linked with */
+void libbpf_sha256(const void *data, size_t len,
+		   __u8 out[LIBBPF_SHA256_DIGEST_LENGTH]);
+
 #define BTF_STRUCT	"struct"
 #define BTF_UNION	"union"
 #define BTF_TYPEDEF	"typedef"
@@ -172,6 +183,19 @@ struct object {
 	u32 addr_syms_cap;
 };
 
+struct btf_link {
+	char *value;
+	char *section;
+	char *module;
+	char *btf_path;
+};
+
+struct btf_links {
+	struct btf_link *links;
+	u32 cnt;
+	u32 cap;
+};
+
 #define DECL_TAG_FASTCALL "bpf_fastcall"
 #define DECL_TAG_KFUNC "bpf_kfunc"
 
@@ -259,6 +283,48 @@ static int __ensure_mem(void **data, u32 *cap, u32 cnt, size_t elem_sz)
 #define ensure_mem(arr_ptr, cap_ptr, cnt) \
 	__ensure_mem((void **)(arr_ptr), (cap_ptr), (cnt), sizeof(**(arr_ptr)))
 
+static int parse_btf_link(const struct option *opt, const char *arg, int unset)
+{
+	struct btf_links *links = opt->value;
+	struct btf_link *link;
+	char *separator;
+
+	if (unset)
+		return -EINVAL;
+	if (ensure_mem(&links->links, &links->cap, links->cnt + 1))
+		return -ENOMEM;
+	link = &links->links[links->cnt];
+	memset(link, 0, sizeof(*link));
+	link->value = strdup(arg);
+	if (!link->value)
+		return -ENOMEM;
+	link->section = link->value;
+	separator = strchr(link->section, ':');
+	if (!separator || separator == link->section)
+		goto err_value;
+	*separator++ = '\0';
+	link->module = separator;
+	separator = strchr(link->module, ':');
+	if (!separator || separator == link->module || !separator[1])
+		goto err_value;
+	*separator++ = '\0';
+	link->btf_path = separator;
+	links->cnt++;
+	return 0;
+err_value:
+	free(link->value);
+	return -EINVAL;
+}
+
+static void free_btf_links(struct btf_links *links)
+{
+	u32 i;
+
+	for (i = 0; i < links->cnt; i++)
+		free(links->links[i].value);
+	free(links->links);
+}
+
 static bool is_btf_id(const char *name)
 {
 	return name && !strncmp(name, BTF_ID_PREFIX, sizeof(BTF_ID_PREFIX) - 1);
@@ -1738,9 +1804,143 @@ static int patch_btfids(const char *btfids_path, const char *elf_path)
 	return err;
 }
 
+/*
+ * Fill in the <section>.link record of an ELF file: the name of the module
+ * that carries <section>, NUL-padded to __MODULE_NAME_LEN of the target, the
+ * SHA-256 of the raw BTF in link->btf_path, and its size in the byte order of
+ * the target.  The record must already be there, with exactly that size.
+ */
+static int patch_btf_link(const char *elf_path, const struct btf_link *link)
+{
+	size_t shdrstrndx, module_name_len, link_size;
+	void *raw_btf_data;
+	Elf_Scn *scn = NULL;
+	char section[128];
+	int fd, err = -1;
+	FILE *btf_file;
+	u32 raw_btf_size, btf_size;
+	Elf_Data *data;
+	GElf_Ehdr ehdr;
+	struct stat st;
+	GElf_Shdr sh;
+	char *name;
+	Elf *elf;
+
+	if (stat(link->btf_path, &st) < 0) {
+		pr_err("FAILED to stat %s: %s\n", link->btf_path, strerror(errno));
+		return -1;
+	}
+	if (!st.st_size || st.st_size > UINT_MAX) {
+		pr_err("FAILED: %s has an unexpected size\n", link->btf_path);
+		return -1;
+	}
+	raw_btf_size = st.st_size;
+	raw_btf_data = malloc(raw_btf_size);
+	if (!raw_btf_data) {
+		pr_err("FAILED to allocate %u bytes for %s\n", raw_btf_size, link->btf_path);
+		return -ENOMEM;
+	}
+	btf_file = fopen(link->btf_path, "rb");
+	if (!btf_file) {
+		pr_err("FAILED to open %s: %s\n", link->btf_path, strerror(errno));
+		goto out_data;
+	}
+	if (fread(raw_btf_data, raw_btf_size, 1, btf_file) != 1) {
+		pr_err("FAILED to read %s\n", link->btf_path);
+		fclose(btf_file);
+		goto out_data;
+	}
+	fclose(btf_file);
+
+	if (snprintf(section, sizeof(section), "%s.link", link->section) >= (int)sizeof(section)) {
+		pr_err("FAILED: section name %s.link is too long\n", link->section);
+		goto out_data;
+	}
+
+	elf_version(EV_CURRENT);
+	fd = open(elf_path, O_RDWR);
+	if (fd < 0) {
+		pr_err("FAILED to open %s: %s\n", elf_path, strerror(errno));
+		goto out_data;
+	}
+	elf = elf_begin(fd, ELF_C_RDWR_MMAP, NULL);
+	if (!elf) {
+		pr_err("FAILED cannot create ELF descriptor: %s\n", elf_errmsg(-1));
+		goto out_close;
+	}
+	elf_flagelf(elf, ELF_C_SET, ELF_F_LAYOUT);
+	if (!gelf_getehdr(elf, &ehdr)) {
+		pr_err("FAILED cannot get ELF header: %s\n", elf_errmsg(-1));
+		goto out_elf;
+	}
+	/* __MODULE_NAME_LEN is 64 - sizeof(unsigned long) */
+	if (ehdr.e_ident[EI_CLASS] == ELFCLASS32) {
+		module_name_len = BTF_LINK_MODULE_NAME_MAX - 4;
+	} else if (ehdr.e_ident[EI_CLASS] == ELFCLASS64) {
+		module_name_len = BTF_LINK_MODULE_NAME_MAX - 8;
+	} else {
+		pr_err("FAILED: unknown ELF class of %s\n", elf_path);
+		goto out_elf;
+	}
+	if (strlen(link->module) >= module_name_len) {
+		pr_err("FAILED: module name %s is too long\n", link->module);
+		goto out_elf;
+	}
+
+	if (elf_getshdrstrndx(elf, &shdrstrndx)) {
+		pr_err("FAILED cannot get shdr str ndx\n");
+		goto out_elf;
+	}
+	while ((scn = elf_nextscn(elf, scn))) {
+		if (gelf_getshdr(scn, &sh) != &sh) {
+			pr_err("FAILED to get section header\n");
+			goto out_elf;
+		}
+		name = elf_strptr(elf, shdrstrndx, sh.sh_name);
+		if (name && !strcmp(name, section))
+			break;
+	}
+	if (!scn) {
+		pr_err("FAILED: section %s not found in %s\n", section, elf_path);
+		goto out_elf;
+	}
+	link_size = module_name_len + LIBBPF_SHA256_DIGEST_LENGTH + sizeof(u32);
+	data = elf_getdata(scn, NULL);
+	if (!data || !data->d_buf || data->d_size != link_size) {
+		pr_err("FAILED: section %s in %s is not %zu bytes of data\n",
+		       section, elf_path, link_size);
+		goto out_elf;
+	}
+
+	memset(data->d_buf, 0, data->d_size);
+	memcpy(data->d_buf, link->module, strlen(link->module) + 1);
+	libbpf_sha256(raw_btf_data, raw_btf_size,
+		      (u8 *)data->d_buf + module_name_len);
+	btf_size = ehdr.e_ident[EI_DATA] == ELFDATANATIVE ? raw_btf_size :
+							     bswap_32(raw_btf_size);
+	memcpy((u8 *)data->d_buf + module_name_len + LIBBPF_SHA256_DIGEST_LENGTH,
+	       &btf_size, sizeof(btf_size));
+
+	pr_debug("Filled in %s of %s for module %s\n", section, elf_path, link->module);
+
+	elf_flagdata(data, ELF_C_SET, ELF_F_DIRTY);
+	if (elf_update(elf, ELF_C_WRITE) < 0) {
+		pr_err("FAILED to update ELF file %s: %s\n", elf_path, elf_errmsg(-1));
+		goto out_elf;
+	}
+	err = 0;
+out_elf:
+	elf_end(elf);
+out_close:
+	close(fd);
+out_data:
+	free(raw_btf_data);
+	return err;
+}
+
 static const char * const resolve_btfids_usage[] = {
 	"resolve_btfids [<options>] <ELF object>",
-	"resolve_btfids --patch_btfids <.BTF_ids file> <ELF object>",
+	"resolve_btfids --patch_btfids <.BTF_ids file> [--btf_link <section>:<module>:<BTF file>]... <ELF object>",
 	NULL
 };
 
@@ -1758,6 +1958,7 @@ int main(int argc, const char **argv)
 		.sets     = RB_ROOT,
 	};
 	const char *btfids_path = NULL;
+	struct btf_links btf_links = {};
 	bool fatal_warnings = false;
 	bool resolve_btfids = true;
 	char out_path[PATH_MAX];
@@ -1773,21 +1974,28 @@ int main(int argc, const char **argv)
 			    "turn warnings into errors"),
 		OPT_BOOLEAN(0, "distill_base", &obj.distill_base,
 			    "distill --btf_base and emit .BTF.base section data"),
+		OPT_CALLBACK(0, "btf_link", &btf_links, "section:module:btf-file",
+			     "patch a BTF link (with --patch_btfids)", parse_btf_link),
 		OPT_STRING(0, "patch_btfids", &btfids_path, "file",
 			   "path to .BTF_ids section data blob to patch into ELF file"),
 		OPT_END()
 	};
 	int err = -1;
+	u32 i;
 
 	argc = parse_options(argc, argv, btfid_options, resolve_btfids_usage,
 			     PARSE_OPT_STOP_AT_NON_OPTION);
-	if (argc != 1)
+	if (argc != 1 || (btf_links.cnt && !btfids_path))
 		usage_with_options(resolve_btfids_usage, btfid_options);
 
 	obj.path = argv[0];
 
-	if (btfids_path)
-		return patch_btfids(btfids_path, obj.path);
+	if (btfids_path) {
+		err = patch_btfids(btfids_path, obj.path);
+		for (i = 0; !err && i < btf_links.cnt; i++)
+			err = patch_btf_link(obj.path, &btf_links.links[i]);
+		goto out;
+	}
 
 	if (elf_collect(&obj))
 		goto out;
@@ -1845,6 +2053,7 @@ int main(int argc, const char **argv)
 	if (!(fatal_warnings && warnings))
 		err = 0;
 out:
+	free_btf_links(&btf_links);
 	btf__free(obj.base_btf);
 	btf__free(obj.btf);
 	btf_id__free_all(&obj.structs);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (9 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
  2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
  12 siblings, 0 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

With CONFIG_DEBUG_INFO_BTF=m, which the next patch makes possible, the
vmlinux BTF is not part of vmlinux: scripts/Makefile.vmlinux strips .BTF
from it, and only vmlinux.unstripped, from which vmlinux is made, keeps
it.  The tools and samples that generate vmlinux.h from a kernel build
tree take the first vmlinux they find, so after an =m build they would
pick one without BTF and fail, e.g. the bpftool skeletons when bpftool is
built from the kernel tree right after the kernel.

Look for vmlinux.unstripped before vmlinux in each of these lists.  Every
kernel build produces it, and with =y it carries the same BTF as vmlinux,
so nothing changes there.  perf filters the candidates for a .BTF
section, but its unanchored grep also matches .BTF_ids and .BTF.link,
which vmlinux keeps with =m; match the section name exactly.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 samples/bpf/Makefile                       | 6 +++++-
 samples/hid/Makefile                       | 6 +++++-
 tools/bpf/bpftool/Makefile                 | 6 +++++-
 tools/perf/bpf_skel.mak                    | 8 ++++++--
 tools/sched_ext/Makefile                   | 6 +++++-
 tools/testing/selftests/bpf/Makefile       | 6 +++++-
 tools/testing/selftests/hid/Makefile       | 6 +++++-
 tools/testing/selftests/sched_ext/Makefile | 6 +++++-
 8 files changed, 41 insertions(+), 9 deletions(-)

diff --git a/samples/bpf/Makefile b/samples/bpf/Makefile
index c28e65046986..fb5cd2d5b6f8 100644
--- a/samples/bpf/Makefile
+++ b/samples/bpf/Makefile
@@ -305,8 +305,12 @@ $(obj)/$(TRACE_HELPERS): TPROGS_CFLAGS := $(TPROGS_CFLAGS) -D__must_check=
 
 -include $(BPF_SAMPLES_PATH)/Makefile.target
 
-VMLINUX_BTF_PATHS ?= $(abspath $(if $(O),$(O)/vmlinux))				\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(abspath $(if $(O),$(O)/vmlinux.unstripped))		\
+		     $(abspath $(if $(O),$(O)/vmlinux))				\
+		     $(abspath $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped))	\
 		     $(abspath $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux))	\
+		     $(abspath $(objtree)/vmlinux.unstripped)			\
 		     $(abspath $(objtree)/vmlinux)
 VMLINUX_BTF ?= $(abspath $(firstword $(wildcard $(VMLINUX_BTF_PATHS))))
 
diff --git a/samples/hid/Makefile b/samples/hid/Makefile
index db5a077c77fc..84566fbe4a5b 100644
--- a/samples/hid/Makefile
+++ b/samples/hid/Makefile
@@ -162,8 +162,12 @@ $(obj)/hid_surface_dial.o: $(obj)/hid_surface_dial.skel.h
 
 -include $(HID_SAMPLES_PATH)/Makefile.target
 
-VMLINUX_BTF_PATHS ?= $(abspath $(if $(O),$(O)/vmlinux))				\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(abspath $(if $(O),$(O)/vmlinux.unstripped))		\
+		     $(abspath $(if $(O),$(O)/vmlinux))				\
+		     $(abspath $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped))	\
 		     $(abspath $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux))	\
+		     $(abspath $(objtree)/vmlinux.unstripped)			\
 		     $(abspath $(objtree)/vmlinux)
 VMLINUX_BTF ?= $(abspath $(firstword $(wildcard $(VMLINUX_BTF_PATHS))))
 
diff --git a/tools/bpf/bpftool/Makefile b/tools/bpf/bpftool/Makefile
index b0f7168e7943..0e25bffd0d72 100644
--- a/tools/bpf/bpftool/Makefile
+++ b/tools/bpf/bpftool/Makefile
@@ -236,8 +236,12 @@ $(BOOTSTRAP_OBJS): $(LIBBPF_BOOTSTRAP)
 OBJS = $(patsubst %.c,$(OUTPUT)%.o,$(SRCS)) $(OUTPUT)disasm.o
 $(OBJS): $(LIBBPF) $(LIBBPF_INTERNAL_HDRS)
 
-VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux)				\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)			\
+		     $(if $(O),$(O)/vmlinux)				\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)	\
+		     ../../../vmlinux.unstripped			\
 		     ../../../vmlinux					\
 		     /sys/kernel/btf/vmlinux				\
 		     /boot/vmlinux-$(shell uname -r)
diff --git a/tools/perf/bpf_skel.mak b/tools/perf/bpf_skel.mak
index f2559de39f96..01b07dd77692 100644
--- a/tools/perf/bpf_skel.mak
+++ b/tools/perf/bpf_skel.mak
@@ -44,8 +44,12 @@ $(BPFTOOL):
 	$(Q)CFLAGS= $(MAKE) -C ../bpf/bpftool OUTPUT=$(SKEL_TOOL_TMP_OUT)/ bootstrap
 
 # Paths to search for a kernel to generate vmlinux.h from.
-VMLINUX_BTF_ELF_PATHS ?= $(if $(O),$(O)/vmlinux)			\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_ELF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)		\
+		     $(if $(O),$(O)/vmlinux)				\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)	\
+		     ../../vmlinux.unstripped				\
 		     ../../vmlinux					\
 		     /boot/vmlinux-$(shell uname -r)
 
@@ -56,7 +60,7 @@ VMLINUX_BTF_BTF_PATHS ?= /sys/kernel/btf/vmlinux
 VMLINUX_BTF_ELF_ABSPATHS ?= $(abspath $(wildcard $(VMLINUX_BTF_ELF_PATHS)))
 VMLINUX_BTF_PATHS ?= $(shell for file in $(VMLINUX_BTF_ELF_ABSPATHS); \
 			do \
-				if [ -f $$file ] && ($(READELF) -S "$$file" | grep -q .BTF); \
+				if [ -f $$file ] && ($(READELF) -SW "$$file" | grep -q "[[:space:]]\.BTF[[:space:]]"); \
 				then \
 					echo "$$file"; \
 				fi; \
diff --git a/tools/sched_ext/Makefile b/tools/sched_ext/Makefile
index 21554f089692..cf6842904934 100644
--- a/tools/sched_ext/Makefile
+++ b/tools/sched_ext/Makefile
@@ -73,8 +73,12 @@ HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
 RESOLVE_BTFIDS := $(HOST_BUILD_DIR)/resolve_btfids/resolve_btfids
 DEFAULT_BPFTOOL := $(HOST_OUTPUT_DIR)/sbin/bpftool
 
-VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux)					\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)				\
+		     $(if $(O),$(O)/vmlinux)					\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)		\
+		     ../../vmlinux.unstripped					\
 		     ../../vmlinux						\
 		     /sys/kernel/btf/vmlinux					\
 		     /boot/vmlinux-$(shell uname -r)
diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index afa589a27b15..e6a841e03219 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -140,8 +140,12 @@ endif
 # Some utility functions use LLVM libraries
 $(OUTPUT)/jit_disasm_helpers.o: CFLAGS += $(LLVM_CFLAGS)
 
-VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux)				\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)			\
+		     $(if $(O),$(O)/vmlinux)				\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)	\
+		     ../../../../vmlinux.unstripped			\
 		     ../../../../vmlinux				\
 		     /sys/kernel/btf/vmlinux				\
 		     /boot/vmlinux-$(shell uname -r)
diff --git a/tools/testing/selftests/hid/Makefile b/tools/testing/selftests/hid/Makefile
index 2f423de83147..8b85382be254 100644
--- a/tools/testing/selftests/hid/Makefile
+++ b/tools/testing/selftests/hid/Makefile
@@ -81,8 +81,12 @@ endif
 HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
 RESOLVE_BTFIDS := $(HOST_BUILD_DIR)/resolve_btfids/resolve_btfids
 
-VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux)				\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)			\
+		     $(if $(O),$(O)/vmlinux)				\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)	\
+		     ../../../../vmlinux.unstripped			\
 		     ../../../../vmlinux				\
 		     /sys/kernel/btf/vmlinux				\
 		     /boot/vmlinux-$(shell uname -r)
diff --git a/tools/testing/selftests/sched_ext/Makefile b/tools/testing/selftests/sched_ext/Makefile
index 3cfe90e0f34f..f9a911c5ce33 100644
--- a/tools/testing/selftests/sched_ext/Makefile
+++ b/tools/testing/selftests/sched_ext/Makefile
@@ -37,8 +37,12 @@ HOST_LIBBPF_OUTPUT := $(OBJ_DIR)/host/libbpf/
 HOST_LIBBPF_DESTDIR := $(OUTPUT_DIR)/host/
 HOST_DESTDIR := $(OUTPUT_DIR)/host/
 
-VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux)					\
+# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF
+VMLINUX_BTF_PATHS ?= $(if $(O),$(O)/vmlinux.unstripped)				\
+		     $(if $(O),$(O)/vmlinux)					\
+		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux.unstripped)	\
 		     $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux)		\
+		     ../../../../vmlinux.unstripped				\
 		     ../../../../vmlinux					\
 		     /sys/kernel/btf/vmlinux					\
 		     /boot/vmlinux-$(shell uname -r)
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (10 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
@ 2026-10-01 22:52 ` Jay Wang
  2026-10-02  9:47   ` Alan Maguire
  2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
  12 siblings, 1 reply; 22+ messages in thread
From: Jay Wang @ 2026-10-01 22:52 UTC (permalink / raw)
  To: bpf, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

Make CONFIG_DEBUG_INFO_BTF a tristate.  With =m the vmlinux BTF is not
part of the kernel image: it is carried by a new module, btf_vmlinux, and
loaded the first time user space asks for something that needs it.
Otherwise it behaves as with =y, with the exceptions btf.rst lists:
BPF_PRELOAD; users of kernel types that run before the module can be
loaded (from the kernel command line or the boot configuration, or
before the root file system is mounted if btf_vmlinux.ko is not in the
initramfs), or on a system that does not let them load modules; the BTF
check of modules loaded before the vmlinux BTF; code that assumes the
vmlinux BTF has id 1; and tools that look for the BTF in vmlinux or in
kernel memory.  The 5.4 MiB of read-only data (distribution config) is
simply not there on systems where nothing uses it.

The only way to save that memory today is CONFIG_DEBUG_INFO_BTF=n, which
a distribution cannot ship: one binary goes to every user, and off takes
BTF away from the users of CO-RE, fentry/fexit, kfuncs, struct_ops,
sched_ext or bpf-lsm.  Whether BTF is used is a property of the
workload, not of the build, so let the first user decide.

The BTF is generated as before, but with =m the .BTF section is linked
as a non-loadable section (like .comment), so the kernel image does not
load it, and the final step that makes vmlinux from vmlinux.unstripped
strips it: every boot image made from vmlinux, whether a raw binary or
an ELF copy, is without it.  Module BTF is generated against
vmlinux.unstripped, which keeps it.  .BTF_ids stays loadable, the
verifier needs it once the BTF is loaded.  The zeroed .BTF.link record
of the kernel (struct btf_link, checked by the module notifier) is filled
in where .BTF_ids is patched, after the final link: resolve_btfids
--btf_link .BTF:btf_vmlinux:.tmp_vmlinux1.BTF writes the name of the
carrier and the size and SHA-256 of the BTF into it, so with =m
gen-btf.sh keeps .tmp_vmlinux1.BTF for that step.
kernel/bpf/btf_vmlinux.c is an empty carrier module, built through
obj-$(CONFIG_DEBUG_INFO_BTF) so that make localmodconfig maps it to its
option; scripts/gen-btf.sh gives it the vmlinux .BTF as its own .BTF
section instead of generating split BTF for it, and fails if it found
none, so that one module depends on vmlinux with =m; with
CONFIG_DEBUG_INFO_BTF_MODULES all of them do, as before.

The packages keep the BTF where it now is: make pacman-pkg puts
vmlinux.unstripped into the debug package next to vmlinux with =m, for
the BTF of external modules, and make rpm-pkg refuses to build a
debuginfo package with a find-debuginfo that cannot keep .BTF (no
--keep-section), which would strip the payload of btf_vmlinux.ko.

CONFIG_BPF_PRELOAD is not selectable with =m: its iterator programs
attach through the vmlinux BTF, so every bpffs mount (systemd does one
at boot) would load it and defeat the point.  Module BTF is still kept
when a module loads, as with =y; the saving is the vmlinux BTF only.
Programs that need kernel types before the root file system is mounted
need btf_vmlinux.ko in the initramfs; the Kconfig help says so.

The module has no exit: once loaded the BTF stays, as with =y.  The
runtime side -- loading the module at the start of the requests that need
it, checking it against .BTF.link, deferring kfunc and struct_ops
registrations and module BTF until it arrives -- and the
IS_ENABLED()/$(subst m,y,...) preparation of the existing checks are in
the preceding patches; this one makes it selectable.

Tested with 1 GiB of memory, same tree, =y vs =m, both with
CONFIG_DEBUG_INFO_BTF_MODULES=y:

 - MemTotal is ~5.4 MB higher with =m while the BTF is unused: the size
   of the .BTF section.
 - stat() of /sys/kernel/btf/vmlinux reports the BTF size before it is
   loaded, as the btf_sysfs selftest expects.
 - With BTF in use, MemFree is the same within run-to-run noise.
 - Modules loaded before the trigger (ext4, nf_conntrack and its kfuncs,
   xfrm_interface) appear in /sys/kernel/btf immediately and get BTF ids
   once the BTF is loaded.  A socket filter, and one that fails
   verification, load without loading the module.  Each of these, as the
   first user of the BTF, loads it and works: a kprobe program calling
   bpf_get_current_task_btf(), a syscall program calling kfuncs, a
   struct_ops map, BTF and a map with a kptr to task_struct, a light
   skeleton loader, read() of /sys/kernel/btf/vmlinux, mmap() of it as
   libbpf does it (it fails, then libbpf reads), BPF_BTF_GET_NEXT_ID,
   kprobe events with BTF arguments, a tracepoint's btf_ids file, a
   bpffs mount with delegate options that name commands, and four
   programs loaded at once.  /proc/self/mountinfo, bpffs mounts with
   delegate_*=any or hex masks, the ftrace argument printer (also from
   sysrq-z) and BPF_BTF_GET_NEXT_ID without CAP_SYS_ADMIN do not load it.
 - Six BPF_BTF_GET_NEXT_ID users and a module tracepoint's btf_ids read
   at once, right after modules were loaded, all see every module BTF.
 - Without btf_vmlinux.ko installed, all of these fail or degrade as on a
   kernel without BTF, promptly; once it is installed, the next request
   loads it.
 - A carrier module with one byte of its .BTF changed is refused with
   "BTF does not match this kernel" and leaves no state behind.
 - After the load: fstat/read/mmap of /sys/kernel/btf/vmlinux, a
   struct_ops map for tcp_congestion_ops, a syscall program calling the
   bpf_task_from_pid()/bpf_task_release() kfuncs, and modules loaded
   afterwards (nf_nat) all work as with =y.
 - =m without DEBUG_INFO_BTF_MODULES, and =y, build and pass the same
   tests.
 - An i386 kernel builds with =m; its .BTF.link, 96 bytes there against
   92 on x86-64, also matches its carrier.  So does that of an LLVM=1
   build (clang and ld.lld 19).
 - make localmodconfig keeps =m with btf_vmlinux loaded; the rpm spec
   parses, and refuses a find-debuginfo without --keep-section.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 Documentation/bpf/btf.rst         | 68 +++++++++++++++++++++++++++++++
 Makefile                          |  5 ++-
 include/asm-generic/vmlinux.lds.h | 32 ++++++++++++++-
 kernel/bpf/Makefile               |  4 ++
 kernel/bpf/btf_vmlinux.c          | 23 +++++++++++
 kernel/bpf/preload/Kconfig        |  4 ++
 lib/Kconfig.debug                 | 30 +++++++++++++-
 scripts/Makefile.modfinal         | 28 +++++++++----
 scripts/Makefile.vmlinux          |  5 +++
 scripts/gen-btf.sh                | 53 ++++++++++++++++++++++--
 scripts/link-vmlinux.sh           | 25 +++++++++---
 scripts/package/PKGBUILD          |  7 ++++
 scripts/package/kernel.spec       |  4 ++
 scripts/package/mkspec            |  7 ++++
 14 files changed, 276 insertions(+), 19 deletions(-)
 create mode 100644 kernel/bpf/btf_vmlinux.c

diff --git a/Documentation/bpf/btf.rst b/Documentation/bpf/btf.rst
index 29de1222c3e7..7d44374b67ba 100644
--- a/Documentation/bpf/btf.rst
+++ b/Documentation/bpf/btf.rst
@@ -1276,6 +1276,74 @@ format.::
             .long   58
             .long   8206                    # Line 8 Col 14
 
+6.1 Kernel BTF
+--------------
+
+With CONFIG_DEBUG_INFO_BTF=y the BTF of the kernel is generated at link time
+from its DWARF and placed in the .BTF section of vmlinux, which is read-only
+data of the kernel image. It is available as /sys/kernel/btf/vmlinux and, if
+CONFIG_DEBUG_INFO_BTF_MODULES is set, module BTF is generated as split BTF
+against it and available as /sys/kernel/btf/<module>.
+
+With CONFIG_DEBUG_INFO_BTF=m the same BTF is generated, but it is not part of
+the kernel image or of the vmlinux ELF file (vmlinux.unstripped in the build
+tree keeps it, for module BTF generation). It is delivered by the
+btf_vmlinux module, which the kernel loads the first time user space asks for
+something that needs the BTF: reading /sys/kernel/btf/vmlinux, enumerating
+kernel BTF objects (BPF_BTF_GET_NEXT_ID, with CAP_SYS_ADMIN), a BPF program,
+map or BTF object that uses kernel types (an attach_btf_id, a kfunc call, a
+ksym, a map pointer, a helper that takes or returns a kernel BTF pointer, a
+struct_ops map, a kptr to a kernel type), loading a light skeleton loader (a
+syscall program), a kprobe or fprobe event with BTF arguments, a tracepoint's
+btf_ids file, or mounting bpffs with delegate_* options that name commands or
+types. mmap() of /sys/kernel/btf/vmlinux does not load it and fails until it
+is loaded; libbpf then reads the file instead.
+
+Loading the module waits for user space (modprobe), so it only happens at the
+start of such a request, holding no lock that loading a module needs. The code
+that uses the BTF never loads it: bpf_get_btf_vmlinux() and bpf_find_btf_id()
+return nothing while it is not loaded, as on a kernel without BTF, and a bpf()
+command that fails because of that is run once more after the system call has
+loaded the BTF. Until then no memory is used for it; afterwards it behaves as
+with =y, except as described below. In particular:
+
+  * /sys/kernel/btf/vmlinux exists from boot with its final size.
+  * Modules loaded before the vmlinux BTF are exposed in /sys/kernel/btf right
+    away, their BTF is parsed and gets a BTF id once the vmlinux BTF is
+    loaded, together with their kfunc and struct_ops registrations. A request
+    that loads the vmlinux BTF returns once that is done.
+  * kfunc, dtor kfunc and struct_ops registrations of the kernel itself are
+    applied before the BTF becomes visible.
+  * The kernel only accepts the BTF it was built with: the name of the module,
+    and the size and SHA-256 of the BTF, are recorded in the kernel when it is
+    linked (.BTF.link), and the module is checked against them.
+  * Once loaded the BTF stays; the module cannot be unloaded.
+
+If the module is not available (not installed, or the root file system is not
+mounted yet), the kernel behaves as one built without BTF and tries again next
+time; probe events defined on the kernel command line or in the boot
+configuration cannot use BTF arguments for that reason. The module is also not
+loaded where the system does not let the task that needs the BTF load modules:
+with kernel.modules_disabled set, when the security policy does not allow the
+module request, or when modprobe is configured not to load it. Such systems can
+load btf_vmlinux at boot instead, e.g. through modules-load.d.
+
+CONFIG_BPF_PRELOAD is not available with =m: its iterators attach through the
+vmlinux BTF, so mounting bpffs would load it. Some users only use the BTF if it
+is already loaded: bpf_snprintf_btf() and bpf_seq_printf_btf(), which run in
+program context, the ftrace function argument printer (func-args,
+funcgraph-args), which can run with interrupts disabled, and the names of bpffs
+delegate_* options in ``/proc/*/mountinfo``. The vmlinux BTF gets its BTF id
+when it is loaded, so it is not necessarily id 1. Tools that look for the BTF
+in kernel memory, or in a crash dump, through the __start_BTF and __stop_BTF
+symbols do not find it.
+
+A module loaded before the vmlinux BTF is loaded cannot have its BTF checked
+against it yet. Its BTF is checked when the vmlinux BTF arrives, and if it
+does not match, the module keeps running without BTF, with a warning: without
+CONFIG_MODULE_ALLOW_BTF_MISMATCH such a module is only refused if it loads
+after the vmlinux BTF.
+
 7. Testing
 ==========
 
diff --git a/Makefile b/Makefile
index f561516e1735..7ce5d478abd3 100644
--- a/Makefile
+++ b/Makefile
@@ -1745,8 +1745,9 @@ endif
 #
 
 # *.ko are usually independent of vmlinux, but CONFIG_DEBUG_INFO_BTF_MODULES
-# is an exception.
-ifdef CONFIG_DEBUG_INFO_BTF_MODULES
+# is an exception, and so is the btf_vmlinux module with CONFIG_DEBUG_INFO_BTF=m,
+# which carries the vmlinux BTF.
+ifneq ($(CONFIG_DEBUG_INFO_BTF_MODULES)$(filter m,$(CONFIG_DEBUG_INFO_BTF)),)
 KBUILD_BUILTIN := y
 modules: vmlinux
 endif
diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index b2988aa12f66..cf5a3b35b33b 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -674,8 +674,19 @@
 
 /*
  * .BTF
+ *
+ * With CONFIG_DEBUG_INFO_BTF=y the vmlinux BTF is loaded as read-only data and
+ * bounded by __start_BTF/__stop_BTF.  With CONFIG_DEBUG_INFO_BTF=m it is
+ * linked as a non-loadable section (see BTF_NONALLOC in ELF_DETAILS), so that
+ * module BTF generation can read it from vmlinux.unstripped; it is stripped
+ * from vmlinux (scripts/Makefile.vmlinux), and the btf_vmlinux module carries
+ * a copy and provides it on demand at runtime.
+ * What is loaded instead is .BTF.link (struct btf_link in kernel/bpf/btf.c):
+ * the name of that module and the size and SHA-256 of the BTF, which
+ * resolve_btfids fills in after the final link (scripts/link-vmlinux.sh).
+ * .BTF_ids is needed by the kernel in both cases.
  */
-#ifdef CONFIG_DEBUG_INFO_BTF
+#if IS_BUILTIN(CONFIG_DEBUG_INFO_BTF)
 #define BTF								\
 	. = ALIGN(PAGE_SIZE);						\
 	.BTF : AT(ADDR(.BTF) - LOAD_OFFSET) {				\
@@ -685,10 +696,28 @@
 	.BTF_ids : AT(ADDR(.BTF_ids) - LOAD_OFFSET) {			\
 		*(.BTF_ids)						\
 	}
+#elif IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+#define BTF								\
+	. = ALIGN(8);							\
+	.BTF.link : AT(ADDR(.BTF.link) - LOAD_OFFSET) {			\
+		BOUNDED_SECTION_BY(.BTF.link, _BTF_link)		\
+	}								\
+	. = ALIGN(PAGE_SIZE);						\
+	.BTF_ids : AT(ADDR(.BTF_ids) - LOAD_OFFSET) {			\
+		*(.BTF_ids)						\
+	}
 #else
 #define BTF
 #endif
 
+#if IS_MODULE(CONFIG_DEBUG_INFO_BTF)
+/* quoted: BTF is a macro, an unquoted .BTF here would expand it */
+#define BTF_NONALLOC							\
+		".BTF" 0 : { *(".BTF") }
+#else
+#define BTF_NONALLOC
+#endif
+
 /*
  * Init task
  */
@@ -849,6 +878,7 @@
 /* Required sections not related to debugging. */
 #define ELF_DETAILS							\
 		.comment 0 : { *(.comment) }				\
+		BTF_NONALLOC						\
 		.symtab 0 : { *(.symtab) }				\
 		.strtab 0 : { *(.strtab) }				\
 		.shstrtab 0 : { *(.shstrtab) }				\
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 0b7db88f1bed..bde2ae68908f 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -43,6 +43,10 @@ endif
 ifeq ($(CONFIG_SYSFS),y)
 obj-$(subst m,y,$(CONFIG_DEBUG_INFO_BTF)) += sysfs_btf.o
 endif
+# With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is carried by this module
+ifeq ($(CONFIG_DEBUG_INFO_BTF),m)
+obj-$(CONFIG_DEBUG_INFO_BTF) += btf_vmlinux.o
+endif
 ifeq ($(CONFIG_BPF_JIT),y)
 obj-$(CONFIG_BPF_SYSCALL) += bpf_struct_ops.o
 obj-$(CONFIG_BPF_SYSCALL) += cpumask.o
diff --git a/kernel/bpf/btf_vmlinux.c b/kernel/bpf/btf_vmlinux.c
new file mode 100644
index 000000000000..8d89b4bb3c43
--- /dev/null
+++ b/kernel/bpf/btf_vmlinux.c
@@ -0,0 +1,23 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Carrier module for the vmlinux BTF when CONFIG_DEBUG_INFO_BTF=m.
+ *
+ * This module has no code of its own.  Its .BTF section is a copy of the
+ * vmlinux BTF (see scripts/gen-btf.sh), which the BTF module notifier in
+ * kernel/bpf/btf.c recognizes by module name and installs as the vmlinux BTF.
+ * The kernel loads it on demand, the first time the vmlinux BTF is needed.
+ *
+ * There is deliberately no module_exit(): once the BTF is in use it cannot
+ * be taken away again, exactly as with CONFIG_DEBUG_INFO_BTF=y.
+ */
+#include <linux/init.h>
+#include <linux/module.h>
+
+static int __init btf_vmlinux_init(void)
+{
+	return 0;
+}
+module_init(btf_vmlinux_init);
+
+MODULE_DESCRIPTION("BTF type information for vmlinux");
+MODULE_LICENSE("GPL");
diff --git a/kernel/bpf/preload/Kconfig b/kernel/bpf/preload/Kconfig
index aef7b0bc96d6..b1600bdce7a0 100644
--- a/kernel/bpf/preload/Kconfig
+++ b/kernel/bpf/preload/Kconfig
@@ -6,6 +6,10 @@ menuconfig BPF_PRELOAD
 	# The dependency on !COMPILE_TEST prevents it from being enabled
 	# in allmodconfig or allyesconfig configurations
 	depends on !COMPILE_TEST
+	# The preloaded iterators attach through the vmlinux BTF, so with
+	# CONFIG_DEBUG_INFO_BTF=m every bpffs mount would load the BTF, which
+	# defeats the point of =m on any system that mounts bpffs at boot.
+	depends on DEBUG_INFO_BTF!=m
 	help
 	  This builds kernel module with several embedded BPF programs that are
 	  pinned into BPF FS mount point as human readable files that are
diff --git a/lib/Kconfig.debug b/lib/Kconfig.debug
index 134b15a44625..8c671621006f 100644
--- a/lib/Kconfig.debug
+++ b/lib/Kconfig.debug
@@ -396,7 +396,7 @@ config DEBUG_INFO_SPLIT
 	  Incompatible with older versions of ccache.
 
 config DEBUG_INFO_BTF
-	bool "Generate BTF type information"
+	tristate "Generate BTF type information"
 	depends on !DEBUG_INFO_SPLIT && !DEBUG_INFO_REDUCED
 	depends on !GCC_PLUGIN_RANDSTRUCT || COMPILE_TEST
 	depends on BPF_SYSCALL
@@ -408,6 +408,30 @@ config DEBUG_INFO_BTF
 	  Turning this on requires pahole v1.22 or later, which will convert
 	  DWARF type info into equivalent deduplicated BTF type info.
 
+	  If built as a module (=m), the vmlinux BTF is not part of the
+	  kernel image.  It is carried by the btf_vmlinux module, which is
+	  loaded on demand the first time the BTF is needed: when a BPF
+	  program requires kernel type information, or when
+	  /sys/kernel/btf/vmlinux is read.  Until then, no memory is
+	  spent on it.  The vmlinux ELF file does not carry the BTF
+	  either; module BTF is generated against vmlinux.unstripped, and
+	  tools that read the BTF from a file can use that or
+	  /sys/kernel/btf/vmlinux.
+
+	  Module BTF (DEBUG_INFO_BTF_MODULES) is kept when a module loads,
+	  as with =y, and registered once the vmlinux BTF is available; the
+	  saving is the vmlinux BTF only.
+
+	  If BPF programs that use kernel types run before the root file
+	  system is mounted, put btf_vmlinux.ko into the initramfs: until
+	  the module can be loaded, such programs fail as on a kernel
+	  without BTF.  The kernel requests the module itself, so the
+	  processes that use BTF must be allowed to cause a module load
+	  (kernel.modules_disabled, the security policy's module_request);
+	  otherwise load btf_vmlinux at boot, e.g. through modules-load.d.
+	  Not compatible with BPF_PRELOAD, whose iterators would load the
+	  BTF at every bpffs mount.
+
 config PAHOLE_HAS_BTF_TAG
 	def_bool PAHOLE_VERSION >= 123
 	depends on CC_IS_CLANG
@@ -442,6 +466,10 @@ config MODULE_ALLOW_BTF_MISMATCH
 	  this option will still load module BTF where possible but ignore
 	  it when a mismatch is found.
 
+	  With DEBUG_INFO_BTF=m, a module loaded before the vmlinux BTF can
+	  only be checked once that is loaded; it is then kept without BTF
+	  on a mismatch, as with this option.
+
 config GDB_SCRIPTS
 	bool "Provide GDB scripts for kernel debugging"
 	help
diff --git a/scripts/Makefile.modfinal b/scripts/Makefile.modfinal
index 01a37ec872b9..201444f51b7d 100644
--- a/scripts/Makefile.modfinal
+++ b/scripts/Makefile.modfinal
@@ -38,20 +38,34 @@ quiet_cmd_ld_ko_o = LD [M]  $@
 		$(KBUILD_LDFLAGS_MODULE) $(LDFLAGS_MODULE)		\
 		-T $(objtree)/scripts/module.lds -o $@ $(filter %.o, $^)
 
+# The ELF file with the vmlinux BTF: with CONFIG_DEBUG_INFO_BTF=m the BTF is
+# stripped from vmlinux (scripts/Makefile.vmlinux), vmlinux.unstripped keeps it.
+btf-vmlinux := $(objtree)/vmlinux$(if $(filter m,$(CONFIG_DEBUG_INFO_BTF)),.unstripped)
+
 quiet_cmd_btf_ko = BTF [M] $@
       cmd_btf_ko = 							\
-	if [ ! -f $(objtree)/vmlinux ]; then				\
-		printf "Skipping BTF generation for %s due to unavailability of vmlinux\n" $@ 1>&2; \
+	if [ ! -f $(btf-vmlinux) ]; then				\
+		printf "Skipping BTF generation for %s due to unavailability of %s\n" $@ $(notdir $(btf-vmlinux)) 1>&2; \
 	else	\
-		$(CONFIG_SHELL) $(srctree)/scripts/gen-btf.sh --btf_base $(objtree)/vmlinux $@; \
+		$(CONFIG_SHELL) $(srctree)/scripts/gen-btf.sh --btf_base $(btf-vmlinux) $@; \
 	fi;
 
-# Re-generate module BTFs if either module's .ko or vmlinux changed
-%.ko: %.o %.mod.o .module-common.o $(objtree)/scripts/module.lds $(and $(CONFIG_DEBUG_INFO_BTF_MODULES),$(KBUILD_BUILTIN),$(objtree)/vmlinux) FORCE
-	+$(call if_changed,ld_ko_o)
+# Modules that get a .BTF section: all of them with CONFIG_DEBUG_INFO_BTF_MODULES,
+# otherwise only the vmlinux BTF carrier module with CONFIG_DEBUG_INFO_BTF=m.
 ifdef CONFIG_DEBUG_INFO_BTF_MODULES
-	+$(if $(newer-prereqs),$(call cmd,btf_ko))
+btf-modules := $(modules:%.o=%.ko)
+else ifeq ($(CONFIG_DEBUG_INFO_BTF),m)
+btf-modules := $(filter %/btf_vmlinux.ko,$(modules:%.o=%.ko))
+# Only the carrier depends on vmlinux, not every module
+ifdef KBUILD_BUILTIN
+$(btf-modules): $(btf-vmlinux)
+endif
 endif
+
+# Re-generate module BTFs if either module's .ko or vmlinux changed
+%.ko: %.o %.mod.o .module-common.o $(objtree)/scripts/module.lds $(and $(CONFIG_DEBUG_INFO_BTF_MODULES),$(KBUILD_BUILTIN),$(btf-vmlinux)) FORCE
+	+$(call if_changed,ld_ko_o)
+	+$(if $(and $(filter $@,$(btf-modules)),$(newer-prereqs)),$(call cmd,btf_ko))
 	+$(call cmd,check_tracepoint)
 
 targets += $(modules:%.o=%.ko) $(modules:%.o=%.mod.o) .module-common.o
diff --git a/scripts/Makefile.vmlinux b/scripts/Makefile.vmlinux
index fcae1e432d9a..557db1ee1f3b 100644
--- a/scripts/Makefile.vmlinux
+++ b/scripts/Makefile.vmlinux
@@ -86,6 +86,11 @@ remove-section-$(CONFIG_ARCH_VMLINUX_NEEDS_RELOCS) += '.rel*' '!.rel*.dyn'
 # for compatibility with binutils < 2.32
 # https://sourceware.org/git/?p=binutils-gdb.git;a=commit;h=c12d9fa2afe7abcbe407a00e15719e1a1350c2a7
 remove-section-$(CONFIG_ARCH_VMLINUX_NEEDS_RELOCS) += '.rel.*'
+# With CONFIG_DEBUG_INFO_BTF=m the btf_vmlinux module carries the vmlinux BTF;
+# only vmlinux.unstripped keeps it, for module BTF generation.
+ifeq ($(CONFIG_DEBUG_INFO_BTF),m)
+remove-section-y += .BTF
+endif
 
 remove-symbols := -w --strip-unneeded-symbol='__mod_device_table__*'
 
diff --git a/scripts/gen-btf.sh b/scripts/gen-btf.sh
index 8ca96eb10a69..780f7b4bb214 100755
--- a/scripts/gen-btf.sh
+++ b/scripts/gen-btf.sh
@@ -22,6 +22,14 @@
 #   - ${1}.btf.o ready for linking into vmlinux
 #   - ${1}.BTF_ids with .BTF_ids data blob
 # This output is consumed by scripts/link-vmlinux.sh
+#
+# With CONFIG_DEBUG_INFO_BTF=m the .BTF section in ${1}.btf.o is not
+# allocatable, so the kernel image does not carry the BTF; vmlinux.unstripped
+# does, for module BTF generation, and scripts/Makefile.vmlinux strips it from
+# vmlinux.  ${1}.BTF is kept too: scripts/link-vmlinux.sh has resolve_btfids
+# record its size and SHA-256 in .BTF.link (struct btf_link).  The
+# btf_vmlinux module gets no BTF of its own; its .BTF section is a copy of the
+# vmlinux BTF, extracted from --btf_base.
 
 set -e
 
@@ -60,6 +68,10 @@ is_enabled() {
 	grep -q "^$1=y" ${objtree}/include/config/auto.conf
 }
 
+is_module() {
+	grep -q "^$1=m" ${objtree}/include/config/auto.conf
+}
+
 case "${KBUILD_VERBOSE}" in
 *1*)
 	set -x
@@ -83,13 +95,21 @@ gen_btf_o()
 {
 	btf_data=${ELF_FILE}.btf.o
 
+	# CONFIG_DEBUG_INFO_BTF=m: .BTF stays non-allocatable, kept in
+	# vmlinux.unstripped for module BTF but not loaded; the btf_vmlinux
+	# module provides it at runtime.
+	btf_flags=alloc,readonly
+	if is_module CONFIG_DEBUG_INFO_BTF; then
+		btf_flags=readonly
+	fi
+
 	# Create ${btf_data} which contains just .BTF section but no symbols. Add
-	# SHF_ALLOC because .BTF will be part of the vmlinux image. --strip-all
+	# SHF_ALLOC (=y) because .BTF will be part of the vmlinux image. --strip-all
 	# deletes all symbols including __start_BTF and __stop_BTF, which will
 	# be redefined in the linker script.
 	echo "" | ${CC} ${CLANG_FLAGS} ${KBUILD_CPPFLAGS} ${KBUILD_CFLAGS} -fno-lto -c -x c -o ${btf_data} -
 	${OBJCOPY} --add-section .BTF=${ELF_FILE}.BTF \
-		--set-section-flags .BTF=alloc,readonly ${btf_data}
+		--set-section-flags .BTF=${btf_flags} ${btf_data}
 	${OBJCOPY} --only-section=.BTF --strip-all ${btf_data}
 
 	# Change e_type to ET_REL so that it can be used to link final vmlinux.
@@ -120,7 +140,10 @@ embed_btf_data()
 cleanup()
 {
 	rm -f "${ELF_FILE}.BTF.1"
-	rm -f "${ELF_FILE}.BTF"
+	# CONFIG_DEBUG_INFO_BTF=m: vmlinux's .BTF is needed for .BTF.link
+	if [ "${BTFGEN_MODE}" = "module" ] || ! is_module CONFIG_DEBUG_INFO_BTF; then
+		rm -f "${ELF_FILE}.BTF"
+	fi
 	if [ "${BTFGEN_MODE}" = "module" ]; then
 		rm -f "${ELF_FILE}.BTF.base"
 		rm -f "${ELF_FILE}.BTF_ids"
@@ -133,6 +156,30 @@ if [ -n "${BTF_BASE}" ]; then
 	BTFGEN_MODE="module"
 fi
 
+# CONFIG_DEBUG_INFO_BTF=m: the btf_vmlinux module carries the vmlinux BTF
+# itself.  Its own types are of no interest, so instead of generating split
+# BTF for it, copy the (non-loadable) .BTF section of --btf_base
+# (vmlinux.unstripped) into the module.
+# The kernel recognizes the module by name and treats its .BTF as base BTF.
+case "${BTFGEN_MODE}:${ELF_FILE}" in
+module:*/btf_vmlinux.ko)
+	if is_module CONFIG_DEBUG_INFO_BTF; then
+		# -O binary only emits allocatable sections; make .BTF one for
+		# the extraction.  ${BTF_BASE} itself is not modified.
+		${OBJCOPY} -O binary --only-section=.BTF			\
+			--set-section-flags .BTF=alloc,load,readonly	\
+			"${BTF_BASE}" "${ELF_FILE}.BTF"
+		# objcopy succeeds with an empty file if there is no .BTF
+		if [ ! -s "${ELF_FILE}.BTF" ]; then
+			echo >&2 "error: no .BTF section in ${BTF_BASE}"
+			exit 1
+		fi
+		${OBJCOPY} --add-section .BTF="${ELF_FILE}.BTF" "${ELF_FILE}"
+		exit 0
+	fi
+	;;
+esac
+
 gen_btf_data
 
 case "${BTFGEN_MODE}" in
diff --git a/scripts/link-vmlinux.sh b/scripts/link-vmlinux.sh
index ab0b8125c8cb..68b99234eab8 100755
--- a/scripts/link-vmlinux.sh
+++ b/scripts/link-vmlinux.sh
@@ -37,6 +37,15 @@ is_enabled() {
 	grep -q "^$1=y" include/config/auto.conf
 }
 
+is_module() {
+	grep -q "^$1=m" include/config/auto.conf
+}
+
+# =y or =m
+is_set() {
+	grep -q "^$1=[ym]" include/config/auto.conf
+}
+
 # Nice output in kbuild format
 # Will be suppressed by "make -s"
 info()
@@ -195,6 +204,7 @@ fi
 
 btf_vmlinux_bin_o=
 btfids_vmlinux=
+btf_link=
 kallsymso=
 strip_debug=
 generate_map=
@@ -211,17 +221,17 @@ if is_enabled CONFIG_KALLSYMS; then
 	kallsyms .tmp_vmlinux0.syms .tmp_vmlinux0.kallsyms
 fi
 
-if is_enabled CONFIG_KALLSYMS || is_enabled CONFIG_DEBUG_INFO_BTF; then
+if is_enabled CONFIG_KALLSYMS || is_set CONFIG_DEBUG_INFO_BTF; then
 
 	# The kallsyms linking does not need debug symbols, but the BTF does.
-	if ! is_enabled CONFIG_DEBUG_INFO_BTF; then
+	if ! is_set CONFIG_DEBUG_INFO_BTF; then
 		strip_debug=1
 	fi
 
 	vmlinux_link .tmp_vmlinux1
 fi
 
-if is_enabled CONFIG_DEBUG_INFO_BTF; then
+if is_set CONFIG_DEBUG_INFO_BTF; then
 	info BTF .tmp_vmlinux1
 	if ! ${CONFIG_SHELL} ${srctree}/scripts/gen-btf.sh .tmp_vmlinux1; then
 		echo >&2 "Failed to generate BTF for vmlinux"
@@ -230,6 +240,11 @@ if is_enabled CONFIG_DEBUG_INFO_BTF; then
 	fi
 	btf_vmlinux_bin_o=.tmp_vmlinux1.btf.o
 	btfids_vmlinux=.tmp_vmlinux1.BTF_ids
+	if is_module CONFIG_DEBUG_INFO_BTF; then
+		# The btf_vmlinux module carries the BTF; .BTF.link names it and
+		# holds the size and SHA-256 of the BTF (struct btf_link).
+		btf_link="--btf_link .BTF:btf_vmlinux:.tmp_vmlinux1.BTF"
+	fi
 fi
 
 if is_enabled CONFIG_KALLSYMS; then
@@ -287,9 +302,9 @@ fi
 
 vmlinux_link "${VMLINUX}"
 
-if is_enabled CONFIG_DEBUG_INFO_BTF; then
+if is_set CONFIG_DEBUG_INFO_BTF; then
 	info BTFIDS ${VMLINUX}
-	${RESOLVE_BTFIDS} --patch_btfids ${btfids_vmlinux} ${VMLINUX}
+	${RESOLVE_BTFIDS} --patch_btfids ${btfids_vmlinux} ${btf_link} ${VMLINUX}
 fi
 
 mksysmap "${VMLINUX}" System.map
diff --git a/scripts/package/PKGBUILD b/scripts/package/PKGBUILD
index 66e4b6a37783..b66b5e9f1ef1 100644
--- a/scripts/package/PKGBUILD
+++ b/scripts/package/PKGBUILD
@@ -122,6 +122,13 @@ _package-debug(){
 	mkdir -p "${builddir}"
 	ln -sr "${debugdir}/vmlinux" "${builddir}/vmlinux"
 
+	# With CONFIG_DEBUG_INFO_BTF=m only vmlinux.unstripped has the BTF, which
+	# external modules need for theirs (scripts/Makefile.modfinal)
+	if grep -q CONFIG_DEBUG_INFO_BTF=m include/config/auto.conf; then
+		install -Dt "${debugdir}" -m644 vmlinux.unstripped
+		ln -sr "${debugdir}/vmlinux.unstripped" "${builddir}/vmlinux.unstripped"
+	fi
+
 	echo "Installing unstripped vDSO(s)..."
 	${MAKE} INSTALL_MOD_PATH="${pkgdir}/usr" vdso_install
 }
diff --git a/scripts/package/kernel.spec b/scripts/package/kernel.spec
index 46e80970f723..c884308d1e85 100644
--- a/scripts/package/kernel.spec
+++ b/scripts/package/kernel.spec
@@ -76,6 +76,10 @@ This package provides debug information for the kernel image and modules from th
 %if %{with_keep_section}
 %global _find_debuginfo_opts -r --keep-section .BTF --keep-section .BTF.base
 %else
+%if %{with_btf_vmlinux_module}
+# With CONFIG_DEBUG_INFO_BTF=m, btf_vmlinux.ko carries the vmlinux BTF in .BTF
+%{error:find-debuginfo cannot keep .BTF, which btf_vmlinux.ko needs; build without debuginfo (--without debuginfo)}
+%endif
 %global _find_debuginfo_opts -r
 %endif
 
diff --git a/scripts/package/mkspec b/scripts/package/mkspec
index c604f8c174e2..d28d77ef8b73 100755
--- a/scripts/package/mkspec
+++ b/scripts/package/mkspec
@@ -65,6 +65,13 @@ fi
 echo "%define with_debuginfo_manual $with_debuginfo_manual"
 echo "%define with_debuginfo_rpm $with_debuginfo_rpm"
 
+# the btf_vmlinux module must keep its .BTF section (see kernel.spec)
+if grep -q CONFIG_DEBUG_INFO_BTF=m include/config/auto.conf; then
+echo '%define with_btf_vmlinux_module 1'
+else
+echo '%define with_btf_vmlinux_module 0'
+fi
+
 cat<<EOF
 %define ARCH ${ARCH}
 %define KERNELRELEASE ${KERNELRELEASE}
-- 
2.47.3


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records
  2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
@ 2026-10-01 23:29   ` bot+bpf-ci
  0 siblings, 0 replies; 22+ messages in thread
From: bot+bpf-ci @ 2026-10-01 23:29 UTC (permalink / raw)
  To: wanjay, bpf, ast, daniel, andrii, eddyz87, memxor
  Cc: alan.maguire, martin.lau, yonghong.song, jolsa, ihor.solodrai,
	qmo, nathan, nsc, linux-kbuild, linux, christian, mcgrof,
	petr.pavlu, samitolvanen, linux-modules, rostedt, mhiramat,
	mathieu.desnoyers, linux-trace-kernel, acme, namhyung, irogers,
	linux-perf-users, jikos, bentiss, linux-input, tj, void, arighi,
	changwoo, sched-ext, shuah, linux-kselftest, ojeda,
	rust-for-linux, arnd, linux-doc, linux-kernel, abuehaze, doebel,
	mpohlack, jay.wang.upstream, ast, andrii, daniel, martin.lau

[-- Attachment #1: Type: text/plain, Size: 2905 bytes --]

> resolve_btfids: add --btf_link to fill in .BTF.link records

TL;DR: Not a bug: patch_btf_link() duplicates most of patch_btfids()'s ELF
open/find-section/update code, so each --btf_link reopens and rewrites the
file; suggests a shared helper or a single ELF handle.

> diff --git a/tools/bpf/resolve_btfids/main.c b/tools/bpf/resolve_btfids/main.c
> --- a/tools/bpf/resolve_btfids/main.c
> +++ b/tools/bpf/resolve_btfids/main.c

[ ... ]

> +static int patch_btf_link(const char *elf_path, const struct btf_link *link)
> +{

[ ... ]

> +	elf_version(EV_CURRENT);
> +	fd = open(elf_path, O_RDWR);
> +	if (fd < 0) {
> +		pr_err("FAILED to open %s: %s\n", elf_path, strerror(errno));
> +		goto out_data;
> +	}
> +	elf = elf_begin(fd, ELF_C_RDWR_MMAP, NULL);
> +	if (!elf) {
> +		pr_err("FAILED cannot create ELF descriptor: %s\n", elf_errmsg(-1));
> +		goto out_close;
> +	}
> +	elf_flagelf(elf, ELF_C_SET, ELF_F_LAYOUT);

[ ... ]

> +	if (elf_getshdrstrndx(elf, &shdrstrndx)) {
> +		pr_err("FAILED cannot get shdr str ndx\n");
> +		goto out_elf;
> +	}
> +	while ((scn = elf_nextscn(elf, scn))) {
> +		if (gelf_getshdr(scn, &sh) != &sh) {
> +			pr_err("FAILED to get section header\n");
> +			goto out_elf;
> +		}
> +		name = elf_strptr(elf, shdrstrndx, sh.sh_name);
> +		if (name && !strcmp(name, section))
> +			break;
> +	}

[ ... ]

> +	elf_flagdata(data, ELF_C_SET, ELF_F_DIRTY);
> +	if (elf_update(elf, ELF_C_WRITE) < 0) {
> +		pr_err("FAILED to update ELF file %s: %s\n", elf_path, elf_errmsg(-1));
> +		goto out_elf;
> +	}

[ ... ]

> @@ -1773,21 +1974,28 @@ int main(int argc, const char **argv)
[ ... ]
> -	if (btfids_path)
> -		return patch_btfids(btfids_path, obj.path);
> +	if (btfids_path) {
> +		err = patch_btfids(btfids_path, obj.path);
> +		for (i = 0; !err && i < btf_links.cnt; i++)
> +			err = patch_btf_link(obj.path, &btf_links.links[i]);
> +		goto out;
> +	}

This isn't a bug, but patch_btf_link() repeats most of what patch_btfids()
does just above it: elf_version(), open(O_RDWR), elf_begin(ELF_C_RDWR_MMAP),
elf_flagelf(ELF_F_LAYOUT), the elf_getshdrstrndx() plus elf_nextscn() /
gelf_getshdr() / elf_strptr() / strcmp() loop that finds a section by name,
and the elf_flagdata(ELF_F_DIRTY) / elf_update() / elf_end() / close()
teardown.  Only the section name and what gets written into data->d_buf
differ.

With the loop in main(), a --patch_btfids run with N --btf_link options
opens, maps and rewrites the same ELF file N + 1 times.

Could the two share a small helper that opens the ELF and returns the
Elf_Data of a named section, or could the .BTF.link records be written while
.BTF_ids is being patched, using the same ELF handle?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36938681172

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module
  2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
@ 2026-10-01 23:45   ` bot+bpf-ci
  2026-10-02 11:48   ` Alexei Starovoitov
  1 sibling, 0 replies; 22+ messages in thread
From: bot+bpf-ci @ 2026-10-01 23:45 UTC (permalink / raw)
  To: wanjay, bpf, ast, daniel, andrii, eddyz87, memxor
  Cc: alan.maguire, martin.lau, yonghong.song, jolsa, ihor.solodrai,
	qmo, nathan, nsc, linux-kbuild, linux, christian, mcgrof,
	petr.pavlu, samitolvanen, linux-modules, rostedt, mhiramat,
	mathieu.desnoyers, linux-trace-kernel, acme, namhyung, irogers,
	linux-perf-users, jikos, bentiss, linux-input, tj, void, arighi,
	changwoo, sched-ext, shuah, linux-kselftest, ojeda,
	rust-for-linux, arnd, linux-doc, linux-kernel, abuehaze, doebel,
	mpohlack, jay.wang.upstream, ast, andrii, daniel, martin.lau

[-- Attachment #1: Type: text/plain, Size: 4325 bytes --]

> bpf: take the vmlinux BTF from the btf_vmlinux module

TL;DR: Retrying __sys_bpf() in bpf() may break BPF_PROG_LOAD: the failed first
run can write the kernel's record size into uattr, so the retry may fail with
-EINVAL instead of -E2BIG and libbpf won't recover.

> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c

[ ... ]

> @@ -6525,10 +6544,40 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
>  	return err;
>  }
>
> +/*
> + * With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is loaded on demand, but never
> + * from within a command: loading waits for user space, and a command may hold
> + * locks or run from a BPF program (bpf_sys_bpf()).  A command that needs the
> + * BTF while it is not loaded fails as it would without BTF.  If the command
> + * is one whose failure leaves nothing behind, load the BTF here, on entry
> + * from user space with nothing held, and run the command once more.
> + */

[ ... ]

>  SYSCALL_DEFINE5(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size,
>  		struct bpf_common_attr __user *, uattr_common, unsigned int, size_common)
>  {
> -	return __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
> +	unsigned int misses = bpf_btf_vmlinux_misses();
> +	int err;
> +
> +	err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
> +	if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && err < 0 && bpf_btf_vmlinux_retry(cmd, misses))
> +		err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common),
> +				size_common);
> +	return err;
>  }

Is it safe to run __sys_bpf() a second time on the same uattr?  The comment
and the commit message say a failed BPF_PROG_LOAD "leaves nothing behind" and
"fails again the same way", but the first run can write into the caller's
attr.

When the func_info, line_info or core_relo records are larger than the
kernel's struct and have a non-zero tail, the failing run writes the
kernel's record size back into the attr, which is an input field:

kernel/bpf/check_btf.c:
	if (ret == -E2BIG) {
		verbose(env, "nonzero tailing record in func info");
		/* set the size kernel expects so loader can zero
		 * out the rest of the record.
		 */
		if (copy_to_bpfptr_offset(uattr,
				  offsetof(union bpf_attr, func_info_rec_size),
				  &min_size, sizeof(min_size)))

The same happens for line_info_rec_size in check_btf_line() and for
core_relo_rec_size in bpf_check_core_relo().

With CONFIG_DEBUG_INFO_BTF=m, suppose the global miss counter moves while
this load runs.  That happens if any bpf_get_btf_vmlinux() call misses, for
example another process loading BPF, or btf_module_notify() for a module that
is loading.  Suppose the carrier module can also be loaded.  Then
bpf_btf_vmlinux_retry() returns true and the second __sys_bpf() copies the
attr in again.

func_info_rec_size (or line_info_rec_size) now holds the kernel's size, but
the user's func_info/line_info array is still laid out with the original,
larger stride.  prepare_btf_func() and check_btf_line() then read record i
from the wrong offset, landing in the previous record's non-zero tail, and
fail with -EINVAL ("same or smaller insn offset", "Invalid
line_info[%u].insn_off") instead of -E2BIG.

libbpf's bpf_prog_load() recovers from E2BIG only here:

	while (errno == E2BIG && (!finfo || !linfo))

which rebuilds the records with the size the kernel wrote back.  With EINVAL
that path is skipped, so a load that works without this patch (E2BIG, libbpf
trims the records, then success) now fails, and the verifier log shows a
misleading message.

A miss in the same run does not trigger this, because it fails earlier
(attach_btf, CO-RE candidates, kfunc).  Only a concurrent miss before the
vmlinux BTF is first loaded does, for example during boot.

Should the retry be skipped when err == -E2BIG, or more generally when the
first run may have written into uattr?  kernel/bpf/syscall.c is not touched
by later commits in the series, so nothing there changes this.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36938681172

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available
  2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
@ 2026-10-01 23:45   ` bot+bpf-ci
  0 siblings, 0 replies; 22+ messages in thread
From: bot+bpf-ci @ 2026-10-01 23:45 UTC (permalink / raw)
  To: wanjay, bpf, ast, daniel, andrii, eddyz87, memxor
  Cc: alan.maguire, martin.lau, yonghong.song, jolsa, ihor.solodrai,
	qmo, nathan, nsc, linux-kbuild, linux, christian, mcgrof,
	petr.pavlu, samitolvanen, linux-modules, rostedt, mhiramat,
	mathieu.desnoyers, linux-trace-kernel, acme, namhyung, irogers,
	linux-perf-users, jikos, bentiss, linux-input, tj, void, arighi,
	changwoo, sched-ext, shuah, linux-kselftest, ojeda,
	rust-for-linux, arnd, linux-doc, linux-kernel, abuehaze, doebel,
	mpohlack, jay.wang.upstream, ast, andrii, daniel, martin.lau

[-- Attachment #1: Type: text/plain, Size: 1538 bytes --]

> bpf: keep module BTF until the vmlinux BTF is available

TL;DR: Nits only: in kernel/bpf/btf.c, btf_drop_deferred_vmlinux_regs() could
reuse the new btf_free_deferred_regs() helper, and the module and vmlinux
apply loops look duplicated. No bug.

> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> index 96241dc62dc3..4d51fb212218 100644
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c

[ ... ]

> @@ -10967,6 +11252,16 @@ static void btf_apply_deferred_regs(struct btf *btf, struct list_head *regs)
>  	}
>  }
>
> +static void btf_free_deferred_regs(struct list_head *regs)
> +{
> +	struct btf_deferred_reg *reg, *tmp;
> +
> +	list_for_each_entry_safe(reg, tmp, regs, list) {
> +		list_del(&reg->list);
> +		btf_free_deferred_reg(reg);
> +	}
> +}
> +

This isn't a bug, but could btf_drop_deferred_vmlinux_regs() call the new
btf_free_deferred_regs(&regs) instead of keeping its own copy of the loop?

The tail of btf_drop_deferred_vmlinux_regs() still open-codes the same
thing on its local list:

    list_for_each_entry_safe(reg, tmp, &regs, list) {
        list_del(&reg->list);
        btf_free_deferred_reg(reg);
    }

The splice / unlock / apply / relock loop in btf_module_apply_regs() is
also the same as the one in btf_apply_deferred_vmlinux_regs(), just with a
different mutex.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36938681172

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
  2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
                   ` (11 preceding siblings ...)
  2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
@ 2026-10-02  4:36 ` Ihor Solodrai
  2026-10-02  7:34   ` Jay Wang
  12 siblings, 1 reply; 22+ messages in thread
From: Ihor Solodrai @ 2026-10-02  4:36 UTC (permalink / raw)
  To: Jay Wang, bpf, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Quentin Monnet, Nathan Chancellor, Nicolas Schier, linux-kbuild,
	Thomas Weißschuh, Christian Heusel, Luis Chamberlain,
	Petr Pavlu, Sami Tolvanen, linux-modules, Steven Rostedt,
	Masami Hiramatsu, Mathieu Desnoyers, linux-trace-kernel,
	Arnaldo Carvalho de Melo, Namhyung Kim, Ian Rogers,
	linux-perf-users, Jiri Kosina, Benjamin Tissoires, linux-input,
	Tejun Heo, David Vernet, Andrea Righi, Changwoo Min, sched-ext,
	Shuah Khan, linux-kselftest, Miguel Ojeda, rust-for-linux,
	Arnd Bergmann, linux-doc, linux-kernel, Hazem Mohamed Abuelfotoh,
	Bjoern Doebel, Martin Pohlack, jay.wang.upstream

On 2026-10-01 3:52 p.m., Jay Wang wrote:
> Based on and tested against bpf-next commit b5a4aa31abd6 ("bpf, cgroup:
> Fix cgroup struct_ops query for a second attach type").
> 
> This series makes CONFIG_DEBUG_INFO_BTF a tristate, so that it can be
> set to =m.  With =m the vmlinux BTF is carried by a module, btf_vmlinux,
> that the kernel loads the first time user space asks for something that
> needs the BTF.  On a system where nothing does, that saves ~5.4 MB of
> RAM with a distribution config; on a system that uses BTF, it behaves as
> with =y.  =y itself is untouched.
> 
> Problem
> -------
> 
> The vmlinux BTF that CONFIG_DEBUG_INFO_BTF=y builds into the kernel
> image takes ~5.4 MB of memory, resident from boot whether anything uses
> it or not.  On small instances that is not negligible.
> 
> A distribution cannot simply turn it off for the users who do not need
> it: it ships one kernel build for all its users, and BTF is not debug
> info anymore.  CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext and
> bpf-lsm all depend on it, so =n takes those away from everyone who does
> use them.
> 
> Hence this series adds CONFIG_DEBUG_INFO_BTF=m: the BTF becomes an
> on-demand module.  Users who never use BTF get the memory back; for
> users who do, the first request that needs it loads it, and everything
> works as with =y.
> 
> [...]
> 
>   Documentation/bpf/btf.rst                  |   68 ++
>   Makefile                                   |    8 +-
>   include/asm-generic/vmlinux.lds.h          |   32 +-
>   include/linux/bpf.h                        |   15 +
>   include/linux/btf.h                        |   11 +
>   include/linux/btf_ids.h                    |    2 +-
>   include/linux/compiler_types.h             |    2 +-
>   include/linux/module.h                     |    2 +-
>   include/trace/trace_events.h               |    2 +-
>   init/Kconfig                               |    2 +-
>   kernel/bpf/Makefile                        |    6 +-
>   kernel/bpf/bpf_struct_ops.c                |    3 +-
>   kernel/bpf/btf.c                           | 1085 ++++++++++++++++++--
>   kernel/bpf/btf_vmlinux.c                   |   23 +
>   kernel/bpf/inode.c                         |   44 +-
>   kernel/bpf/preload/Kconfig                 |    4 +
>   kernel/bpf/syscall.c                       |   51 +-
>   kernel/bpf/sysfs_btf.c                     |   97 +-
>   kernel/bpf/verifier.c                      |  172 +++-
>   kernel/module/Kconfig                      |    2 +-
>   kernel/module/main.c                       |    4 +-
>   kernel/trace/bpf_trace.c                   |    3 +-
>   kernel/trace/trace_events.c                |   10 +
>   kernel/trace/trace_output.c                |    7 +
>   kernel/trace/trace_probe.c                 |   16 +
>   kernel/trace/trace_syscalls.c              |    6 +-
>   lib/Kconfig.debug                          |   30 +-
>   net/netfilter/Makefile                     |    6 +-
>   net/xfrm/Makefile                          |    4 +-
>   samples/bpf/Makefile                       |    6 +-
>   samples/hid/Makefile                       |    6 +-
>   scripts/Makefile.modfinal                  |   28 +-
>   scripts/Makefile.vmlinux                   |    5 +
>   scripts/gen-btf.sh                         |   53 +-
>   scripts/link-vmlinux.sh                    |   25 +-
>   scripts/package/PKGBUILD                   |    7 +
>   scripts/package/kernel.spec                |    4 +
>   scripts/package/mkspec                     |    7 +
>   tools/bpf/bpftool/Makefile                 |    6 +-
>   tools/bpf/resolve_btfids/main.c            |  217 +++-
>   tools/perf/bpf_skel.mak                    |    8 +-
>   tools/sched_ext/Makefile                   |    6 +-
>   tools/testing/selftests/bpf/Makefile       |    6 +-
>   tools/testing/selftests/hid/Makefile       |    6 +-
>   tools/testing/selftests/sched_ext/Makefile |    6 +-
>   45 files changed, 1936 insertions(+), 177 deletions(-)

This series caught my attention because of the sheer amount of the code
and the iteration speed. I've tried to read the cover (it was
tiresome, please ask your "agent" to be concise), and fed the series
to my bot.

A question I had: is it really worth 2k of complicated kernel code to
*maybe sometimes* save <10Mb of memory?

And then the bot found this:

     What systemd does at every boot

     PID1 loads and attaches a BTF-dependent LSM program at every boot

       src/core/manager.c:1026-1059, in manager_new():

       if (FLAGS_SET(test_run_flags, MANAGER_TEST_RUN_MINIMAL)) {
               ...
       } else {
               ...
               (void) bpf_restrict_fs_setup(m);
       }

https://github.com/systemd/systemd/blob/main/src/core/manager.c#L1026-L1059
https://github.com/systemd/systemd/blob/main/src/core/bpf-restrict-fs.c#L27-L85

Do those tiny VMs that you target run systemd with BPF LSM?
I assume you target concrete VMs, not "small instances" in a vacuum.

You can check this by booting them with =m and then:
   $ lsmod | grep btf_vmlinux
   $ systemctl --version  # look for +BPF_FRAMEWORK
   $ cat /sys/kernel/security/lsm

If the answer is "yes" for the majority of them, the whole project is
moot, because the module will be loaded immediately on boot.

My bot found more stuff implementation-wise...

I suggest to slow down, stop burning tokens for a bit, and first try
to figure out the important details:
   - who will actually benefit from these memory savings and do they care?
   - is it worth it in terms of implementation complexity?

Even for LLM it's easier to deal with a 100 lines of code than with 2k.

pw-bot: cr

>   create mode 100644 kernel/bpf/btf_vmlinux.c
> 
> 
> base-commit: b5a4aa31abd6fe90009b63e35dc18c67d041ec0c


^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
  2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
@ 2026-10-02  7:34   ` Jay Wang
  2026-10-02 10:05     ` Alan Maguire
  2026-10-02 20:58     ` Ihor Solodrai
  0 siblings, 2 replies; 22+ messages in thread
From: Jay Wang @ 2026-10-02  7:34 UTC (permalink / raw)
  To: ihor.solodrai, bpf, ast, daniel, andrii, eddyz87, memxor
  Cc: alan.maguire, martin.lau, yonghong.song, jolsa, qmo, nathan, nsc,
	linux-kbuild, linux, christian, mcgrof, petr.pavlu, samitolvanen,
	linux-modules, rostedt, mhiramat, mathieu.desnoyers,
	linux-trace-kernel, acme, namhyung, irogers, linux-perf-users,
	jikos, bentiss, linux-input, tj, void, arighi, changwoo,
	sched-ext, shuah, linux-kselftest, ojeda, rust-for-linux, arnd,
	linux-doc, linux-kernel, abuehaze, doebel, mpohlack,
	jay.wang.upstream

> A question I had: is it really worth 2k of complicated kernel code to
> *maybe sometimes* save <10Mb of memory?

There are many reasons this is worth it.  The main ones:

1. This did not start from a number we picked, but from a real use
   case: users who run large fleets of small instances, 1 GiB of
   memory or less, with their workloads sized to fit.  We are not able
   to disclose more detail, but for them 5.4 MB on every instance is
   real money, and it can be exactly what pushes a workload over its
   memory budget and onto the next instance size.  That is why we took
   this on, even though we knew it would not be a small change.

2. Loading code and data only when a system needs them is what kernel
   modules exist for.  And most of them take well under ~1 MB once
   loaded (nf_conntrack, overlay, vfat); even big ones like ext4, kvm
   and btrfs stay around 1-2.5 MB.  The vmlinux BTF is 5.4 MB.

3. We are not alone in trying to keep BTF out of memory until it is
   needed.  The inline BTF work [1] plans to deliver its data, which
   is even larger, through a module too.  The .BTF.link record and the
   resolve_btfids option that serve both are already part of this
   series (patch 10), so this is not machinery for the vmlinux BTF
   alone.

> Do those tiny VMs that you target run systemd with BPF LSM?

Not necessarily, and that is the image's choice.  Even with BPF LSM
built in and bpf in CONFIG_LSM, memory-conscious users can turn it off
at boot with an lsm= list that leaves it out, and then systemd does not
load restrict_fs at all.

And even where user space is in the way, which as above it does not
have to be, the answer is to fix user space so it gets this win too,
not to give up on it.

> If the answer is "yes" for the majority of them,

"The majority" is also hard to pin down here: what runs at boot
depends on the user space packages built into each image, and on the
workload.

So the question should be what it takes to make this work, not whether
to drop it because it is complicated.  If there are ways to cut it
down, or issues found in the implementation, we would be happy to take
them and go through every one.


[1] https://lore.kernel.org/bpf/20260916074118.1007116-1-alan.maguire@oracle.com/

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module
  2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
@ 2026-10-02  9:47   ` Alan Maguire
  0 siblings, 0 replies; 22+ messages in thread
From: Alan Maguire @ 2026-10-02  9:47 UTC (permalink / raw)
  To: wanjay
  Cc: abuehaze, acme, alan.maguire, andrii, arighi, arnd, ast, bentiss,
	bpf, changwoo, christian, daniel, doebel, eddyz87, ihor.solodrai,
	irogers, jay.wang.upstream, jikos, jolsa, linux-doc, linux-input,
	linux-kbuild, linux-kernel, linux-kselftest, linux-modules,
	linux-perf-users, linux-trace-kernel, linux, martin.lau,
	mathieu.desnoyers, mcgrof, memxor, mhiramat, mpohlack, namhyung,
	nathan, nsc, ojeda, petr.pavlu, qmo, rostedt, rust-for-linux,
	samitolvanen, sched-ext, shuah, tj, void, yonghong.song


[snip]

> diff --git a/scripts/package/kernel.spec b/scripts/package/kernel.spec
> index 46e80970f723..c884308d1e85 100644
> --- a/scripts/package/kernel.spec
> +++ b/scripts/package/kernel.spec
> @@ -76,6 +76,10 @@ This package provides debug information for the kernel image and modules from th
>  %if %{with_keep_section}
>  %global _find_debuginfo_opts -r --keep-section .BTF --keep-section .BTF.base
>  %else
> +%if %{with_btf_vmlinux_module}
> +# With CONFIG_DEBUG_INFO_BTF=m, btf_vmlinux.ko carries the vmlinux BTF in .BTF
> +%{error:find-debuginfo cannot keep .BTF, which btf_vmlinux.ko needs; build without debuginfo (--without debuginfo)}
> +%endif
>  %global _find_debuginfo_opts -r
>  %endif

I think we need a --keep-section for .BTF.link above too?

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
  2026-10-02  7:34   ` Jay Wang
@ 2026-10-02 10:05     ` Alan Maguire
  2026-10-02 20:58     ` Ihor Solodrai
  1 sibling, 0 replies; 22+ messages in thread
From: Alan Maguire @ 2026-10-02 10:05 UTC (permalink / raw)
  To: wanjay
  Cc: abuehaze, acme, alan.maguire, andrii, arighi, arnd, ast, bentiss,
	bpf, changwoo, christian, daniel, doebel, eddyz87, ihor.solodrai,
	irogers, jay.wang.upstream, jikos, jolsa, linux-doc, linux-input,
	linux-kbuild, linux-kernel, linux-kselftest, linux-modules,
	linux-perf-users, linux-trace-kernel, linux, martin.lau,
	mathieu.desnoyers, mcgrof, memxor, mhiramat, mpohlack, namhyung,
	nathan, nsc, ojeda, petr.pavlu, qmo, rostedt, rust-for-linux,
	samitolvanen, sched-ext, shuah, tj, void, yonghong.song

>> A question I had: is it really worth 2k of complicated kernel code to
>> *maybe sometimes* save <10Mb of memory?
> 
> There are many reasons this is worth it.  The main ones:
> 
> 1. This did not start from a number we picked, but from a real use
>    case: users who run large fleets of small instances, 1 GiB of
>    memory or less, with their workloads sized to fit.  We are not able
>    to disclose more detail, but for them 5.4 MB on every instance is
>    real money, and it can be exactly what pushes a workload over its
>    memory budget and onto the next instance size.  That is why we took
>    this on, even though we knew it would not be a small change.
> 
> 2. Loading code and data only when a system needs them is what kernel
>    modules exist for.  And most of them take well under ~1 MB once
>    loaded (nf_conntrack, overlay, vfat); even big ones like ext4, kvm
>    and btrfs stay around 1-2.5 MB.  The vmlinux BTF is 5.4 MB.
> 
> 3. We are not alone in trying to keep BTF out of memory until it is
>    needed.  The inline BTF work [1] plans to deliver its data, which
>    is even larger, through a module too.  The .BTF.link record and the
>    resolve_btfids option that serve both are already part of this
>    series (patch 10), so this is not machinery for the vmlinux BTF
>    alone.

The other consideration that embedded folks have raised is disk footprint;
apparently by moving vmlinux .BTF to a module that is a win for them due
to the way embedded systems are partitioned (vmlinux image on a different
partition from modules). And for them a lower memory footprint is a win too.
So from my perspective it is worth doing, especially if the only other option is 
to switch BTF off and lose most BPF functionality. Between small VMs
and embedded systems that could amount to a lot of systems.

However it may make sense to approach via a multi-phase delivery. The inline 
stuff is pretty close, so if I can land that with the resolve_btfids support
for --btf_link that reduces the footprint of this work somewhat. Then perhaps
it might make sense to split into a prerequisite series that does the prep work
for module-delivered BTF (roughly patches 1-5), avoiding issues around module
load by reorganizing when we access vmlinux BTF. Then the remainder - stuff
which is tied more closely to module vmlinux BTF delivery - would be a smaller 
series (I'm planning on sending out the inline kbuild stuff later today all
going well). Others may have different suggestions, but that may make the work
easier to land; the multi-series approach definitely helped with the inline work
FWIW.

Alan

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module
  2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
  2026-10-01 23:45   ` bot+bpf-ci
@ 2026-10-02 11:48   ` Alexei Starovoitov
  1 sibling, 0 replies; 22+ messages in thread
From: Alexei Starovoitov @ 2026-10-02 11:48 UTC (permalink / raw)
  To: Jay Wang, bpf, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Alan Maguire, Martin KaFai Lau, Yonghong Song, Jiri Olsa,
	Ihor Solodrai, Quentin Monnet, Nathan Chancellor, Nicolas Schier,
	linux-kbuild, Thomas Weißschuh, Christian Heusel,
	Luis Chamberlain, Petr Pavlu, Sami Tolvanen, linux-modules,
	Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
	linux-trace-kernel, Arnaldo Carvalho de Melo, Namhyung Kim,
	Ian Rogers, linux-perf-users, Jiri Kosina, Benjamin Tissoires,
	linux-input, Tejun Heo, David Vernet, Andrea Righi, Changwoo Min,
	sched-ext, Shuah Khan, linux-kselftest, Miguel Ojeda,
	rust-for-linux, Arnd Bergmann, linux-doc, linux-kernel,
	Hazem Mohamed Abuelfotoh, Bjoern Doebel, Martin Pohlack,
	jay.wang.upstream

On Thu, Oct 01, 2026 at 10:52 PM Jay Wang <wanjay@amazon.com> wrote:
> +	err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
> +	if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && err < 0 && bpf_btf_vmlinux_retry(cmd, misses))
> +		err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common),
> +				size_common);

Running the command twice when it failed and a global counter moved
is not going to fly.
Missing BTF is not always an error.
With =m and BTF not loaded yet take
  int prog(struct xdp_md *ctx)
with func_info. That is everything libbpf loads.
do_check_common() calls btf_prepare_func_args(env, 0).
find_canonical_prog_ctx_type() returns NULL, btf_is_prog_ctx_type()
returns false, the arg becomes ARG_PTR_TO_MEM and
func_info_aux[0].unreliable is set.
The prog loads. err == 0. No retry.
From then on freplace of that prog fails with
"Cannot replace static functions" and fentry/fexit see its args
as scalars, even after BTF is loaded.
With =y both work.

That is the only reason why "a socket filter does not load it".
In v3 such prog loaded BTF.
systemd built with BPF_FRAMEWORK loads
  int sd_bind4(struct bpf_sock_addr *ctx)
from manager_setup_cgroup() at every boot.

The result of the verification cannot depend on who touched BTF first.

bpf-ci is right too. The first run writes into uattr.

It's starting to feel that this is dead end.
So much complexity to save few Mbyte.

pw-bot: cr

^ permalink raw reply	[flat|nested] 22+ messages in thread

* Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
  2026-10-02  7:34   ` Jay Wang
  2026-10-02 10:05     ` Alan Maguire
@ 2026-10-02 20:58     ` Ihor Solodrai
  1 sibling, 0 replies; 22+ messages in thread
From: Ihor Solodrai @ 2026-10-02 20:58 UTC (permalink / raw)
  To: Jay Wang, bpf, ast, daniel, andrii, eddyz87, memxor
  Cc: alan.maguire, martin.lau, yonghong.song, jolsa, qmo, nathan, nsc,
	linux-kbuild, linux, christian, mcgrof, petr.pavlu, samitolvanen,
	linux-modules, rostedt, mhiramat, mathieu.desnoyers,
	linux-trace-kernel, acme, namhyung, irogers, linux-perf-users,
	jikos, bentiss, linux-input, tj, void, arighi, changwoo,
	sched-ext, shuah, linux-kselftest, ojeda, rust-for-linux, arnd,
	linux-doc, linux-kernel, abuehaze, doebel, mpohlack,
	jay.wang.upstream

Hi Jay, thank you for taking the time to reply.

On 2026-10-02 12:34 a.m., Jay Wang wrote:
>> A question I had: is it really worth 2k of complicated kernel code to
>> *maybe sometimes* save <10Mb of memory?
> 
> There are many reasons this is worth it.  The main ones:
> 
> 1. This did not start from a number we picked, but from a real use
>     case: users who run large fleets of small instances, 1 GiB of
>     memory or less, with their workloads sized to fit.  We are not able
>     to disclose more detail, but for them 5.4 MB on every instance is
>     real money, and it can be exactly what pushes a workload over its
>     memory budget and onto the next instance size.  That is why we took
>     this on, even though we knew it would not be a small change.

Ok, good to know there is a real customer for this.

Obviously I don't know any details of the use-case apart from what
you've shared. But bear with me trying the customer's hat on:

   - I am running a workload on a legion of 1GiB instances
   - I care about memory footprint a lot, because it directly
     translates into the amount of compute I pay for
   - If I know I don't need BPF on the workload, and these 5.4 MB
     bite me, I can just turn off BPF
   - If I know I need BPF, then those 5.4 MB must be loaded anyways, so
     I have to look for savings elsewhere

What I find strange is for a user of this scale *to not know* whether
BPF will be used in the workload or not, which is what lazy load may
solve.

One thing I can imagine is an opaque workload: untrusted (AI agents
etc), end-user-defined (including BPF usage), or confidential.  But in
these cases, what are the chances that BPF is used there? They are
high, BPF is very widespread.

I can also imagine normal workload not needing BPF, but when something
goes wrong an observability or security thing turning on that loads
BPF programs. But this contradicts the "5.4 MB may push workload over
1GiB" premise: you still need to reserve/swap memory for just-in-time
BPF-based tools, otherwise you OOM.

I of course may be missing something, happy to be corrected.

> 
> 2. Loading code and data only when a system needs them is what kernel
>     modules exist for.  And most of them take well under ~1 MB once
>     loaded (nf_conntrack, overlay, vfat); even big ones like ext4, kvm
>     and btrfs stay around 1-2.5 MB.  The vmlinux BTF is 5.4 MB.

True. At the same time BPF without vmlinux BTF is barely useful. And
it is trusted by the verifier. Both are addressed by BTF being
built-in in the kernel image.

> 
> 3. We are not alone in trying to keep BTF out of memory until it is
>     needed.  The inline BTF work [1] plans to deliver its data, which
>     is even larger, through a module too.  The .BTF.link record and the
>     resolve_btfids option that serve both are already part of this
>     series (patch 10), so this is not machinery for the vmlinux BTF
>     alone.

The inline data is different. It's bigger by it's nature, and it has
more specific users: tracing tools. The userspace tools are able to
themselves decide whether to use that data. The kernel only needs to
make it available.

In comparison, as you yourself noted in the cover:

     CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext and bpf-lsm
     all depend on vmlinux BTF

> 
>> Do those tiny VMs that you target run systemd with BPF LSM?
> 
> Not necessarily, and that is the image's choice.  Even with BPF LSM
> built in and bpf in CONFIG_LSM, memory-conscious users can turn it off
> at boot with an lsm= list that leaves it out, and then systemd does not
> load restrict_fs at all.
> 
> And even where user space is in the way, which as above it does not
> have to be, the answer is to fix user space so it gets this win too,
> not to give up on it.
> 
>> If the answer is "yes" for the majority of them,
> 
> "The majority" is also hard to pin down here: what runs at boot
> depends on the user space packages built into each image, and on the
> workload.

Yes, it is hard to pin down. Which is why we first need to figure out
whether the alleged memory savings actually help anyone.

Adding code to the kernel is not free. More complexity means bigger
bug surface: more opportunities for concurrency bugs, security bugs
etc. Especially runtime loads with retries and stuff. I'm sure you
understand the future cost of all that.

So the bias shouldn't be "let's implement a big thing that might
hypothetically help someone". It's backwards IMO.

> 
> So the question should be what it takes to make this work, not whether
> to drop it because it is complicated.  If there are ways to cut it
> down, or issues found in the implementation, we would be happy to take
> them and go through every one.

The question is in the trade-off.

Would you run a million line python program to search for a string in
a text file? No, that's absurd. But you could.

I'm not saying 2k line patch series is unacceptable in principle. It
may be justified. But I am not convinced it is justified in this case.

Let's say we (as in kernel devs) decided that we want to reduce
vmlinux BTF size, and only load it when necessary. Here are a couple
of alternatives, simpler in comparison to CONFIG_DEBUG_INFO_BTF=m,
although still a bit complex:

   * Compress it.

     $ ls -lah /sys/kernel/btf/vmlinux
     -r--r--r-- 1 root root 6.7M Jul 18 01:54 /sys/kernel/btf/vmlinux
     $ tar --zstd -cf /tmp/vmlinux.btf.zstd -C /sys/kernel/btf vmlinux
     $ ls -lah /tmp/vmlinux.btf.zstd
     -rw-r--r-- 1 isolodrai users 2.2M Oct  2 10:11 /tmp/vmlinux.btf.zstd

     3x win right there.

     Keep zstd-compressed blob in the kernel image and decompress and
     parse it synchronously on first use.

     Here is a prototype (vibe-coded, obviously):
  
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.zstd-20261002

   * Use an external blob.

     Ship /lib/modules/$r/vmlinux.btf, checked against a SHA-256
     recorded in the image. The catch is a synchronous file I/O under
     caller locks, so this is less straightforward.
     But this should cover embedded case that Alan has mentioned.

     Here is a prototype:
  
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.external-20261002

Both alternatives avoid kernel-module loading machine. If we think
harder, we may come up with something even more clever.

I don't particularly prefer one approach over the other, including the
module one. Maintainers probably have a better intuition on that.

The point I'm trying to make is that if you could achieve similar
memory savings with say 500-line well encapsulated diff (or even
better: by refactoring, simplifying, deleting code) it would have
already landed.

And if it was a 50-line patch, the discussion on whether anyone cares
wouldn't be relevant, because it's so cheap.

As it stands, it's not clear (to me at least) in what circumstances
we'll get the promised memory savings at all.

> 
> 
> [1] https://lore.kernel.org/bpf/20260916074118.1007116-1-alan.maguire@oracle.com/


^ permalink raw reply	[flat|nested] 22+ messages in thread

end of thread, other threads:[~2026-10-02 20:58 UTC | newest]

Thread overview: 22+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-02 11:48   ` Alexei Starovoitov
2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02  9:47   ` Alan Maguire
2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02  7:34   ` Jay Wang
2026-10-02 10:05     ` Alan Maguire
2026-10-02 20:58     ` Ihor Solodrai

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®