mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: bot+bpf-ci@kernel.org
To: wanjay@amazon.com,bpf@vger.kernel.org,ast@kernel.org,daniel@iogearbox.net,andrii@kernel.org,eddyz87@gmail.com,memxor@gmail.com
Cc: alan.maguire@oracle.com,martin.lau@linux.dev,yonghong.song@linux.dev,jolsa@kernel.org,ihor.solodrai@linux.dev,qmo@kernel.org,nathan@kernel.org,nsc@kernel.org,linux-kbuild@vger.kernel.org,linux@weissschuh.net,christian@heusel.eu,mcgrof@kernel.org,petr.pavlu@suse.com,samitolvanen@google.com,linux-modules@vger.kernel.org,rostedt@goodmis.org,mhiramat@kernel.org,mathieu.desnoyers@efficios.com,linux-trace-kernel@vger.kernel.org,acme@kernel.org,namhyung@kernel.org,irogers@google.com,linux-perf-users@vger.kernel.org,jikos@kernel.org,bentiss@kernel.org,linux-input@vger.kernel.org,tj@kernel.org,void@manifault.com,arighi@nvidia.com,changwoo@igalia.com,sched-ext@lists.linux.dev,shuah@kernel.org,linux-kselftest@vger.kernel.org,ojeda@kernel.org,rust-for-linux@vger.kernel.org,arnd@arndb.de,linux-doc@vger.kernel.org,linux-kernel@vger.kernel.org,abuehaze@amazon.com,doebel@amazon.de,mpohlack@amazon.de,jay.wang.upstream@gmail.com,ast@kernel.org,andrii@kernel.org,daniel@iogearbox.net,martin.lau@kern
	el.org,eddyz87@gmail.com,yonghong.song@linux.dev,mason@kernel.org,ihor.solodrai@linux.dev
Subject: Re: [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module
Date: Thu,  1 Oct 2026 23:45:55 +0000 (UTC)	[thread overview]
Message-ID: <02c101b324038913fe17ddd47d66f51774a541d7ed4d04796e643cee7ee65da7@mail.kernel.org> (raw)
In-Reply-To: <20261001225214.12351-5-wanjay@amazon.com>

[-- Attachment #1: Type: text/plain, Size: 4325 bytes --]

> bpf: take the vmlinux BTF from the btf_vmlinux module

TL;DR: Retrying __sys_bpf() in bpf() may break BPF_PROG_LOAD: the failed first
run can write the kernel's record size into uattr, so the retry may fail with
-EINVAL instead of -E2BIG and libbpf won't recover.

> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c

[ ... ]

> @@ -6525,10 +6544,40 @@ static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size,
>  	return err;
>  }
>
> +/*
> + * With CONFIG_DEBUG_INFO_BTF=m the vmlinux BTF is loaded on demand, but never
> + * from within a command: loading waits for user space, and a command may hold
> + * locks or run from a BPF program (bpf_sys_bpf()).  A command that needs the
> + * BTF while it is not loaded fails as it would without BTF.  If the command
> + * is one whose failure leaves nothing behind, load the BTF here, on entry
> + * from user space with nothing held, and run the command once more.
> + */

[ ... ]

>  SYSCALL_DEFINE5(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size,
>  		struct bpf_common_attr __user *, uattr_common, unsigned int, size_common)
>  {
> -	return __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
> +	unsigned int misses = bpf_btf_vmlinux_misses();
> +	int err;
> +
> +	err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common), size_common);
> +	if (IS_MODULE(CONFIG_DEBUG_INFO_BTF) && err < 0 && bpf_btf_vmlinux_retry(cmd, misses))
> +		err = __sys_bpf(cmd, USER_BPFPTR(uattr), size, USER_BPFPTR(uattr_common),
> +				size_common);
> +	return err;
>  }

Is it safe to run __sys_bpf() a second time on the same uattr?  The comment
and the commit message say a failed BPF_PROG_LOAD "leaves nothing behind" and
"fails again the same way", but the first run can write into the caller's
attr.

When the func_info, line_info or core_relo records are larger than the
kernel's struct and have a non-zero tail, the failing run writes the
kernel's record size back into the attr, which is an input field:

kernel/bpf/check_btf.c:
	if (ret == -E2BIG) {
		verbose(env, "nonzero tailing record in func info");
		/* set the size kernel expects so loader can zero
		 * out the rest of the record.
		 */
		if (copy_to_bpfptr_offset(uattr,
				  offsetof(union bpf_attr, func_info_rec_size),
				  &min_size, sizeof(min_size)))

The same happens for line_info_rec_size in check_btf_line() and for
core_relo_rec_size in bpf_check_core_relo().

With CONFIG_DEBUG_INFO_BTF=m, suppose the global miss counter moves while
this load runs.  That happens if any bpf_get_btf_vmlinux() call misses, for
example another process loading BPF, or btf_module_notify() for a module that
is loading.  Suppose the carrier module can also be loaded.  Then
bpf_btf_vmlinux_retry() returns true and the second __sys_bpf() copies the
attr in again.

func_info_rec_size (or line_info_rec_size) now holds the kernel's size, but
the user's func_info/line_info array is still laid out with the original,
larger stride.  prepare_btf_func() and check_btf_line() then read record i
from the wrong offset, landing in the previous record's non-zero tail, and
fail with -EINVAL ("same or smaller insn offset", "Invalid
line_info[%u].insn_off") instead of -E2BIG.

libbpf's bpf_prog_load() recovers from E2BIG only here:

	while (errno == E2BIG && (!finfo || !linfo))

which rebuilds the records with the size the kernel wrote back.  With EINVAL
that path is skipped, so a load that works without this patch (E2BIG, libbpf
trims the records, then success) now fails, and the verifier log shows a
misleading message.

A miss in the same run does not trigger this, because it fails earlier
(attach_btf, CO-RE candidates, kfunc).  Only a concurrent miss before the
vmlinux BTF is first loaded does, for example during boot.

Should the retry be skipped when err == -E2BIG, or more generally when the
first run may have written into uattr?  kernel/bpf/syscall.c is not touched
by later commits in the series, so nothing there changes this.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36938681172

  reply	other threads:[~2026-10-01 23:45 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45   ` bot+bpf-ci [this message]
2026-10-02 11:48   ` Alexei Starovoitov
2026-10-01 22:52 ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02  9:47   ` Alan Maguire
2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02  7:34   ` Jay Wang
2026-10-02 10:05     ` Alan Maguire
2026-10-02 20:58     ` Ihor Solodrai
2026-10-03  6:38       ` Alexei Starovoitov
2026-10-03 11:45         ` Alan Maguire
2026-10-03 12:19           ` Alexei Starovoitov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=02c101b324038913fe17ddd47d66f51774a541d7ed4d04796e643cee7ee65da7@mail.kernel.org \
    --to=bot+bpf-ci@kernel.org \
    --cc=abuehaze@amazon.com \
    --cc=acme@kernel.org \
    --cc=alan.maguire@oracle.com \
    --cc=andrii@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=arnd@arndb.de \
    --cc=ast@kernel.org \
    --cc=bentiss@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=changwoo@igalia.com \
    --cc=christian@heusel.eu \
    --cc=daniel@iogearbox.net \
    --cc=doebel@amazon.de \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=irogers@google.com \
    --cc=jay.wang.upstream@gmail.com \
    --cc=jikos@kernel.org \
    --cc=jolsa@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-input@vger.kernel.org \
    --cc=linux-kbuild@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=linux@weissschuh.net \
    --cc=martin.lau@kern \
    --cc=martin.lau@linux.dev \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mcgrof@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mhiramat@kernel.org \
    --cc=mpohlack@amazon.de \
    --cc=namhyung@kernel.org \
    --cc=nathan@kernel.org \
    --cc=nsc@kernel.org \
    --cc=ojeda@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=qmo@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=samitolvanen@google.com \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    --cc=wanjay@amazon.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®