mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jay Wang <wanjay@amazon.com>
To: <bpf@vger.kernel.org>, Alexei Starovoitov <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	Andrii Nakryiko <andrii@kernel.org>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	Kumar Kartikeya Dwivedi <memxor@gmail.com>
Cc: "Alan Maguire" <alan.maguire@oracle.com>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"Jiri Olsa" <jolsa@kernel.org>,
	"Ihor Solodrai" <ihor.solodrai@linux.dev>,
	"Quentin Monnet" <qmo@kernel.org>,
	"Nathan Chancellor" <nathan@kernel.org>,
	"Nicolas Schier" <nsc@kernel.org>,
	linux-kbuild@vger.kernel.org,
	"Thomas Weißschuh" <linux@weissschuh.net>,
	"Christian Heusel" <christian@heusel.eu>,
	"Luis Chamberlain" <mcgrof@kernel.org>,
	"Petr Pavlu" <petr.pavlu@suse.com>,
	"Sami Tolvanen" <samitolvanen@google.com>,
	linux-modules@vger.kernel.org,
	"Steven Rostedt" <rostedt@goodmis.org>,
	"Masami Hiramatsu" <mhiramat@kernel.org>,
	"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
	linux-trace-kernel@vger.kernel.org,
	"Arnaldo Carvalho de Melo" <acme@kernel.org>,
	"Namhyung Kim" <namhyung@kernel.org>,
	"Ian Rogers" <irogers@google.com>,
	linux-perf-users@vger.kernel.org,
	"Jiri Kosina" <jikos@kernel.org>,
	"Benjamin Tissoires" <bentiss@kernel.org>,
	linux-input@vger.kernel.org, "Tejun Heo" <tj@kernel.org>,
	"David Vernet" <void@manifault.com>,
	"Andrea Righi" <arighi@nvidia.com>,
	"Changwoo Min" <changwoo@igalia.com>,
	sched-ext@lists.linux.dev, "Shuah Khan" <shuah@kernel.org>,
	linux-kselftest@vger.kernel.org,
	"Miguel Ojeda" <ojeda@kernel.org>,
	rust-for-linux@vger.kernel.org, "Arnd Bergmann" <arnd@arndb.de>,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
	"Hazem Mohamed Abuelfotoh" <abuehaze@amazon.com>,
	"Bjoern Doebel" <doebel@amazon.de>,
	"Martin Pohlack" <mpohlack@amazon.de>,
	jay.wang.upstream@gmail.com
Subject: [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start
Date: Thu, 1 Oct 2026 22:52:07 +0000	[thread overview]
Message-ID: <20261001225214.12351-6-wanjay@amazon.com> (raw)
In-Reply-To: <20261001225214.12351-1-wanjay@amazon.com>

With CONFIG_DEBUG_INFO_BTF=m, bpf_get_btf_vmlinux() and
bpf_find_btf_id() do not load the vmlinux BTF: loading waits for user
space, and their callers were not written for that.  Besides the bpf()
system call and /sys/kernel/btf/vmlinux, some tracefs and bpffs requests
need the BTF.  Load it at the start of those, with
bpf_load_btf_vmlinux(), and have the code that only uses the BTF if it
happens to be there peek:

 - Reading a tracepoint's btf_ids file (events/*/btf_ids): built-in
   events use the vmlinux BTF, and module BTF is only registered once
   that is loaded.  event_btf_ids_read() looks the BTF up under
   event_mutex, which the trace notifier of the btf_vmlinux module
   takes, so load it before taking the mutex, on the first read() of
   the file (not again for the one that returns EOF).
 - Probe events with BTF arguments ($argN, argument names, $retval,
   $current, typecasts): the parser loads the BTF before looking a
   function or struct up.  It holds dyn_event_ops_mutex, which loading
   a module never takes: besides event creation, only
   dyn_event_register() takes it, from built-in init code.  Loading
   here also keeps a $retval from silently losing its type.
 - The ftrace function argument printer (func-args, funcgraph-args)
   runs in the trace output path, which includes ftrace_dump() with
   interrupts disabled.  It only prints the arguments if the BTF is
   already loaded, and never loads it: that also avoids one modprobe
   per trace line when the module is not installed.
 - bpffs: parsing delegate_* mount options (fs_context) loads the BTF
   when a value names commands or types, which are looked up in it;
   "any" and numeric masks need no BTF and do not load it.  Showing the
   options in /proc/*/mountinfo runs under namespace_sem; it only uses
   the names if the BTF is already there and falls back to hex, as it
   already does without BTF.

With CONFIG_DEBUG_INFO_BTF=y bpf_load_btf_vmlinux() is
bpf_get_btf_vmlinux() and the BTF is parsed at boot, so nothing changes.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/inode.c          | 44 ++++++++++++++++++++++++-------------
 kernel/trace/trace_events.c | 10 +++++++++
 kernel/trace/trace_output.c |  7 ++++++
 kernel/trace/trace_probe.c  | 16 ++++++++++++++
 4 files changed, 62 insertions(+), 15 deletions(-)

diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842..d05bbb61a593 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -658,7 +658,11 @@ struct bpffs_btf_enums {
 	const struct btf_type *attach_t;
 };
 
-static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
+/*
+ * @load: load the vmlinux BTF if necessary (CONFIG_DEBUG_INFO_BTF=m), see
+ * bpf_load_btf_vmlinux(); otherwise only use it if it is already parsed.
+ */
+static int find_bpffs_btf_enums(struct bpffs_btf_enums *info, bool load)
 {
 	struct {
 		const struct btf_type **type;
@@ -674,7 +678,7 @@ static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
 
 	memset(info, 0, sizeof(*info));
 
-	btf = bpf_get_btf_vmlinux();
+	btf = load ? bpf_load_btf_vmlinux() : bpf_peek_btf_vmlinux();
 	if (IS_ERR(btf))
 		return PTR_ERR(btf);
 	if (!btf)
@@ -795,8 +799,11 @@ static int bpf_show_options(struct seq_file *m, struct dentry *root)
 	    opts->delegate_progs || opts->delegate_attachs) {
 		struct bpffs_btf_enums info;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
+		/*
+		 * ignore errors, fallback to hex; this runs under
+		 * namespace_sem, so do not load the BTF from here
+		 */
+		(void)find_bpffs_btf_enums(&info, false);
 
 		mask = (1ULL << __MAX_BPF_CMD) - 1;
 		seq_print_delegate_opts(m, "delegate_cmds",
@@ -1052,35 +1059,33 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 	case OPT_DELEGATE_MAPS:
 	case OPT_DELEGATE_PROGS:
 	case OPT_DELEGATE_ATTACHS: {
-		struct bpffs_btf_enums info;
-		const struct btf_type *enum_t;
+		struct bpffs_btf_enums info = {};
+		const struct btf_type **enum_t;
+		bool enums_tried = false;
 		const char *enum_pfx;
-		u64 *delegate_msk, msk = 0;
+		u64 *delegate_msk, msk = 0, num;
 		char *p, *str;
 		int val;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
-
 		switch (opt) {
 		case OPT_DELEGATE_CMDS:
 			delegate_msk = &opts->delegate_cmds;
-			enum_t = info.cmd_t;
+			enum_t = &info.cmd_t;
 			enum_pfx = "BPF_";
 			break;
 		case OPT_DELEGATE_MAPS:
 			delegate_msk = &opts->delegate_maps;
-			enum_t = info.map_t;
+			enum_t = &info.map_t;
 			enum_pfx = "BPF_MAP_TYPE_";
 			break;
 		case OPT_DELEGATE_PROGS:
 			delegate_msk = &opts->delegate_progs;
-			enum_t = info.prog_t;
+			enum_t = &info.prog_t;
 			enum_pfx = "BPF_PROG_TYPE_";
 			break;
 		case OPT_DELEGATE_ATTACHS:
 			delegate_msk = &opts->delegate_attachs;
-			enum_t = info.attach_t;
+			enum_t = &info.attach_t;
 			enum_pfx = "BPF_";
 			break;
 		default:
@@ -1089,9 +1094,18 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 
 		str = param->string;
 		while ((p = strsep(&str, ":"))) {
+			/*
+			 * Only names need the vmlinux BTF: "any" and numbers do
+			 * not load it.  Ignore errors, fallback to hex.
+			 */
+			if (strcmp(p, "any") && kstrtou64(p, 0, &num) && !enums_tried) {
+				(void)find_bpffs_btf_enums(&info, true);
+				enums_tried = true;
+			}
+
 			if (strcmp(p, "any") == 0) {
 				msk |= ~0ULL;
-			} else if (find_btf_enum_const(info.btf, enum_t, enum_pfx, p, &val)) {
+			} else if (find_btf_enum_const(info.btf, *enum_t, enum_pfx, p, &val)) {
 				msk |= 1ULL << val;
 			} else {
 				err = kstrtou64(p, 0, &msk);
diff --git a/kernel/trace/trace_events.c b/kernel/trace/trace_events.c
index 30c0ddf90887..c887ac6a4857 100644
--- a/kernel/trace/trace_events.c
+++ b/kernel/trace/trace_events.c
@@ -23,6 +23,7 @@
 #include <linux/sort.h>
 #include <linux/slab.h>
 #include <linux/delay.h>
+#include <linux/bpf.h>
 #include <linux/btf.h>
 
 #include <trace/events/sched.h>
@@ -2245,6 +2246,15 @@ event_btf_ids_read(struct file *filp, char __user *ubuf, size_t cnt, loff_t *ppo
 	char buf[128];
 	int len;
 
+	/*
+	 * Built-in events use the vmlinux BTF, and with CONFIG_DEBUG_INFO_BTF=m
+	 * module BTF is only registered once that is loaded.  Loading it loads
+	 * a module, whose trace notifier takes event_mutex: load it before
+	 * taking that, and only for the first read, not again for the EOF one.
+	 */
+	if (!*ppos)
+		bpf_load_btf_vmlinux();
+
 	/* Module unload could free call->class and ids[] mid-read. */
 	scoped_guard(mutex, &event_mutex) {
 		file = event_file_file(filp);
diff --git a/kernel/trace/trace_output.c b/kernel/trace/trace_output.c
index a5ad76175d10..1f346e524ee6 100644
--- a/kernel/trace/trace_output.c
+++ b/kernel/trace/trace_output.c
@@ -739,6 +739,13 @@ void print_function_args(struct trace_seq *s, unsigned long *args,
 	if (lookup_symbol_name(func, name))
 		goto out;
 
+	/*
+	 * This can run with interrupts disabled (ftrace_dump()): only use
+	 * the vmlinux BTF if it is parsed, never load it from here.
+	 */
+	if (IS_ERR_OR_NULL(bpf_peek_btf_vmlinux()))
+		goto out;
+
 	/* TODO: Pass module name here too */
 	t = btf_find_func_proto(name, &btf);
 	if (IS_ERR_OR_NULL(t))
diff --git a/kernel/trace/trace_probe.c b/kernel/trace/trace_probe.c
index 804442b2f7d2..d53ee1ef820c 100644
--- a/kernel/trace/trace_probe.c
+++ b/kernel/trace/trace_probe.c
@@ -531,6 +531,19 @@ static const char *fetch_type_from_btf_type(struct btf *btf,
 	return NULL;
 }
 
+/*
+ * Arguments described by BTF need the vmlinux BTF.  With
+ * CONFIG_DEBUG_INFO_BTF=m it may not be loaded yet, so load it before looking
+ * anything up.  Parsing holds dyn_event_ops_mutex, which loading a module
+ * never takes: besides event creation, only dyn_event_register() takes it,
+ * from built-in init code.
+ */
+static void trace_probe_load_btf(void)
+{
+	lockdep_assert_held(&dyn_event_ops_mutex);
+	bpf_load_btf_vmlinux();
+}
+
 static int query_btf_context(struct traceprobe_parse_context *ctx)
 {
 	const struct btf_param *param;
@@ -544,6 +557,7 @@ static int query_btf_context(struct traceprobe_parse_context *ctx)
 	if (!ctx->funcname)
 		return -EINVAL;
 
+	trace_probe_load_btf();
 	type = btf_find_func_proto(ctx->funcname, &btf);
 	if (!type)
 		return -ENOENT;
@@ -762,6 +776,7 @@ static int parse_btf_arg(char *varname,
 	if (!strcmp(varname, "$current")) {
 		code->op = FETCH_OP_CURRENT;
 		/* If no typecast is specified for $current, use task_struct by default */
+		trace_probe_load_btf();
 		ret = bpf_find_btf_id("task_struct", BTF_KIND_STRUCT, &ctx->struct_btf);
 		if (ret < 0) {
 			trace_probe_log_err(ctx->offset, NO_BTF_ENTRY);
@@ -890,6 +905,7 @@ static int query_btf_struct(const char *sname, struct traceprobe_parse_context *
 		ctx->struct_btf = NULL;
 	}
 
+	trace_probe_load_btf();
 	id = bpf_find_btf_id(sname, BTF_KIND_STRUCT, &btf);
 	if (id < 0)
 		return id;
-- 
2.47.3


  parent reply	other threads:[~2026-10-01 22:53 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-02 11:48   ` Alexei Starovoitov
2026-10-01 22:52 ` Jay Wang [this message]
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29   ` bot+bpf-ci
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02  9:47   ` Alan Maguire
2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02  7:34   ` Jay Wang
2026-10-02 10:05     ` Alan Maguire

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001225214.12351-6-wanjay@amazon.com \
    --to=wanjay@amazon.com \
    --cc=abuehaze@amazon.com \
    --cc=acme@kernel.org \
    --cc=alan.maguire@oracle.com \
    --cc=andrii@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=arnd@arndb.de \
    --cc=ast@kernel.org \
    --cc=bentiss@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=changwoo@igalia.com \
    --cc=christian@heusel.eu \
    --cc=daniel@iogearbox.net \
    --cc=doebel@amazon.de \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=irogers@google.com \
    --cc=jay.wang.upstream@gmail.com \
    --cc=jikos@kernel.org \
    --cc=jolsa@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-input@vger.kernel.org \
    --cc=linux-kbuild@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=linux@weissschuh.net \
    --cc=martin.lau@linux.dev \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mcgrof@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mhiramat@kernel.org \
    --cc=mpohlack@amazon.de \
    --cc=namhyung@kernel.org \
    --cc=nathan@kernel.org \
    --cc=nsc@kernel.org \
    --cc=ojeda@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=qmo@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=samitolvanen@google.com \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®