mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] binfmt: fixes for kres reports
@ 2026-09-18  9:19 Christian Brauner
  2026-09-18  9:19 ` [PATCH 1/2] binfmt_misc: fix OOB read in bpf_binprm_select_interp() Christian Brauner
  2026-09-18  9:20 ` [PATCH 2/2] binfmt_misc: fix racy checks in bpf set_interp kfuncs Christian Brauner
  0 siblings, 2 replies; 3+ messages in thread
From: Christian Brauner @ 2026-09-18  9:19 UTC (permalink / raw)
  To: Chris Mason, linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-mm, linux-kernel, bpf,
	Christian Brauner (Amutable)

I asked Chris to run his kres tooling on the binfmt with bpf changes
merged for this cycle. It found two issues that are fixed in this
series. I reproduced both of them.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Chris Mason (2):
      binfmt_misc: fix OOB read in bpf_binprm_select_interp()
      binfmt_misc: fix racy checks in bpf set_interp kfuncs

 fs/binfmt_misc_bpf.c | 35 ++++++++++++++++++++++++++++++++---
 1 file changed, 32 insertions(+), 3 deletions(-)
---
base-commit: 5179241521401ef364294128bb43cdbce7252457
change-id: 20260918-work-binfmt_misc-fixes-98c4fe6db27d


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/2] binfmt_misc: fix OOB read in bpf_binprm_select_interp()
  2026-09-18  9:19 [PATCH 0/2] binfmt: fixes for kres reports Christian Brauner
@ 2026-09-18  9:19 ` Christian Brauner
  2026-09-18  9:20 ` [PATCH 2/2] binfmt_misc: fix racy checks in bpf set_interp kfuncs Christian Brauner
  1 sibling, 0 replies; 3+ messages in thread
From: Christian Brauner @ 2026-09-18  9:19 UTC (permalink / raw)
  To: Chris Mason, linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-mm, linux-kernel, bpf,
	Christian Brauner (Amutable)

From: Chris Mason <mason@kernel.org>

bpf_binprm_select_interp() checks the name its load program passes with
strnlen(name, name__sz) and then hands the same buffer to
binfmt_misc_find_interp(), which compares it with an unbounded strcmp().
The buffer can be a BPF map value that another CPU rewrites between the
two reads. If the terminating NUL is overwritten in that window, strcmp()
reads past the name__sz bytes the verifier checked. That is an
out-of-bounds read of up to 31 bytes of whatever follows the checked
name__sz bytes.

The verifier checks the name and name__sz pair with BPF_READ | BPF_WRITE,
so a writable array map value is an accepted argument.
bpf(BPF_MAP_UPDATE_ELEM) on an array map copies the new value over the
old one in place and takes no lock. The NUL that strnlen() finds can be
overwritten before strcmp() reads the buffer again:

    CPU0                                   CPU1
    bpf_binprm_select_interp()
      strnlen(name, name__sz)
        finds the NUL inside name__sz
                                           bpf(BPF_MAP_UPDATE_ELEM)
                                             array_map_update_elem()
                                               copy_map_value()
                                                 overwrites the NUL
      binfmt_misc_find_interp()
        strcmp(interp->name, name)
          reads past name__sz

strnlen() proves that a NUL lies inside name__sz only at the moment it
runs. The map update on CPU1 takes no lock, so it can store over the NUL
right after. The lookup on CPU0 then walks the live buffer again, once
per bound interpreter:

    fs/binfmt_misc.c:binfmt_misc_find_interp

        list_for_each_entry(interp, interps, list)
                if (!strcmp(interp->name, name))
                        return interp;

strcmp() stops at the first mismatch or at the end of interp->name.
bm_entry_add_interp() caps a bound name at BINFMT_MISC_INTERP_NAME_MAX
(32) bytes, so strcmp() reads at most 33 bytes of name. The smallest
name__sz the kfunc accepts is 2, which leaves up to 31 bytes read beyond
the checked extent. The handler's own load program has to pass a
writable map value, and something has to store into it while the kfunc
runs. The window between strnlen() and strcmp() is short, but with a
BPF_F_MMAPABLE array the store is a plain user space write into the
mapped value, so a loop can hit it without a single bpf() call.

Copy the name into a stack buffer of BINFMT_MISC_INTERP_NAME_MAX + 1
bytes, terminate it, and look up the copy. The memcpy() length is below
name__sz, so the copy stays inside the extent the verifier checked, and
the BPF buffer is not read again afterwards.

Return -ENOENT first for a name longer than BINFMT_MISC_INTERP_NAME_MAX.
bm_entry_add_interp() rejects a longer name, and the only other binding
site attaches the empty name. No entry can bind such a name, so that
lookup already ended in -ENOENT and no result changes.

Check the first byte of the copy and return -EINVAL if it is NUL, as the
existing "!len" test does for an empty name. Only an 'F' entry binds the
empty name and a 'B' entry cannot carry 'F', so without that check a
racing store of NUL to byte 0 would look up a name no entry binds and
end in -ENOENT rather than -EINVAL. A NUL stored further into the name
only shortens it to another name the program could have passed anyway.

binfmt_misc_find_interp() itself is left alone: entry_attach_interpreter()
calls it with a kernel string, and this kfunc now calls it with a private
copy.

Fixes: 6ec7c96bee30 ("binfmt_misc: let a 'B' entry bind its interpreters")
Signed-off-by: Chris Mason <mason@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/binfmt_misc_bpf.c | 15 ++++++++++++++-
 1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/fs/binfmt_misc_bpf.c b/fs/binfmt_misc_bpf.c
index 91576ff05911..ce1bc78e8511 100644
--- a/fs/binfmt_misc_bpf.c
+++ b/fs/binfmt_misc_bpf.c
@@ -176,6 +176,7 @@ __bpf_kfunc int bpf_binprm_select_interp(struct linux_binprm *bprm,
 					 const char *name, size_t name__sz)
 {
 	const struct binfmt_misc_interp *interp;
+	char buf[BINFMT_MISC_INTERP_NAME_MAX + 1];
 	size_t len;
 	char *path;
 
@@ -184,8 +185,20 @@ __bpf_kfunc int bpf_binprm_select_interp(struct linux_binprm *bprm,
 	len = strnlen(name, name__sz);
 	if (len == name__sz || !len)
 		return -EINVAL;
+	/* No entry binds a longer name, so it cannot be found. */
+	if (len > BINFMT_MISC_INTERP_NAME_MAX)
+		return -ENOENT;
+
+	/*
+	 * The program may pass memory that is written to while this runs,
+	 * so look the name up in a private copy and check that instead.
+	 */
+	memcpy(buf, name, len);
+	buf[len] = '\0';
+	if (!buf[0])
+		return -EINVAL;
 
-	interp = binfmt_misc_find_interp(bprm->bpf_interps, name);
+	interp = binfmt_misc_find_interp(bprm->bpf_interps, buf);
 	if (!interp)
 		return -ENOENT;
 

-- 
2.53.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 2/2] binfmt_misc: fix racy checks in bpf set_interp kfuncs
  2026-09-18  9:19 [PATCH 0/2] binfmt: fixes for kres reports Christian Brauner
  2026-09-18  9:19 ` [PATCH 1/2] binfmt_misc: fix OOB read in bpf_binprm_select_interp() Christian Brauner
@ 2026-09-18  9:20 ` Christian Brauner
  1 sibling, 0 replies; 3+ messages in thread
From: Christian Brauner @ 2026-09-18  9:20 UTC (permalink / raw)
  To: Chris Mason, linux-fsdevel
  Cc: Alexander Viro, Jan Kara, linux-mm, linux-kernel, bpf,
	Christian Brauner (Amutable)

From: Chris Mason <mason@kernel.org>

bpf_binprm_set_interp() tests path[0] != '/' on the buffer its load
program passes and then reads the same buffer again to copy it with
kmemdup_nul(). The buffer can be a BPF map value that another CPU
rewrites between the two reads. If byte 0 is overwritten in that
window, the kfunc stages a relative or empty interpreter path. The
staged path is not checked again, so open_exec() resolves a relative
path against the working directory of the task doing the exec.
bpf_binprm_set_interp_arg() has the same pattern for its "!len" test
and can stage an empty argument, which the interpreter then receives
as an empty argv entry.

The verifier checks the path and path__sz pair with BPF_READ |
BPF_WRITE, so a writable array map value is an accepted argument.
bpf(BPF_MAP_UPDATE_ELEM) on an array map copies the new value over the
old one in place and takes no lock. Both kfuncs are KF_SLEEPABLE and
allocate with GFP_KERNEL between the test and the copy, so the task
can sleep inside the window:

    load program                          bpf(BPF_MAP_UPDATE_ELEM)
    bpf_binprm_set_interp()
      strnlen(path, path__sz)
      path[0] != '/' is false
      kmemdup_nul(path, len, GFP_KERNEL)
        allocation may sleep
                                          array_map_update_elem()
                                            copy_map_value()
                                              rewrites byte 0
        copy reads path again
      bm_bpf_stage_selection()

The test in the load program's column proves what byte 0 held only at
the moment the test ran. The map update takes no lock, so it can store
to byte 0 right after. kmemdup_nul() then copies the rewritten bytes,
and bm_bpf_stage_selection() publishes them as bprm->bpf_interp.

The staged path is not checked again on its way to open_exec():

    load_misc_binary()
      entry_select_interpreter()     returns bprm->bpf_interp unchanged
      build_interp_argv()
        copy_string_kernel()         copies it as argv[0]
      bprm_change_interp()
        kstrdup()
      entry_open_interpreter()
        open_exec()                  unless a bound file is staged or
                                     the entry is an 'F' entry

None of these functions tests the first byte, and load_misc_binary()
hands the pointer to nothing else.

In bpf_binprm_set_interp_arg(), strnlen() finds a non-zero len, a NUL
is then stored to byte 0, and build_interp_argv() later copies the
empty bprm->bpf_interp_arg with copy_string_kernel().

The handler's own load program has to pass a writable map value, and
something has to store into it while the kfunc runs. The allocation can
sleep inside the window, and with a BPF_F_MMAPABLE array the store is a
plain user space write into the mapped value, so a loop can hit it
without a single bpf() call.

Check the private copy in both kfuncs, so that the string that gets
staged is the string that was checked. bpf_binprm_select_interp()
already looks its name up in a private copy for the same reason. The
remaining tests work on path__sz, arg__sz or the local len, and the
copy length is len, so the copy stays inside the extent the verifier
checked.

Results of bpf_binprm_set_interp() with the check on the copy:

- A NUL stored to byte 0 fails interp[0] != '/' and gets -EINVAL.
- For len == 0, kmemdup_nul() returns an empty string, so an empty path
  still gets -EINVAL.
- A NUL stored further into the string only shortens it to another
  absolute path, or another non-empty argument, that the program could
  have passed anyway.
- A path that both lacks the leading '/' and is PATH_MAX or longer now
  gets -ENAMETOOLONG instead of -EINVAL.
- A path that is empty or lacks the leading '/' is now rejected after
  the copy rather than before it, so such a call makes an allocation
  and returns -ENOMEM instead of -EINVAL if that allocation fails.

bpf_binprm_set_interp_arg() still rejects an empty argument before
allocating, so its results are unchanged apart from the raced case
fixed here.

Both new checks run before the previously staged string is freed or
replaced. A failing call frees only its own allocation and leaves the
earlier selection in place, as the -ENOMEM path already does.

Fixes: b4bfe2f6b011 ("binfmt_misc: add binfmt_misc_ops bpf struct_ops")
Signed-off-by: Chris Mason <mason@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/binfmt_misc_bpf.c | 20 ++++++++++++++++++--
 1 file changed, 18 insertions(+), 2 deletions(-)

diff --git a/fs/binfmt_misc_bpf.c b/fs/binfmt_misc_bpf.c
index ce1bc78e8511..a3e26e8a4027 100644
--- a/fs/binfmt_misc_bpf.c
+++ b/fs/binfmt_misc_bpf.c
@@ -141,8 +141,6 @@ __bpf_kfunc int bpf_binprm_set_interp(struct linux_binprm *bprm,
 	len = strnlen(path, path__sz);
 	if (len == path__sz)
 		return -EINVAL;
-	if (path[0] != '/')
-		return -EINVAL;
 	if (len >= PATH_MAX)
 		return -ENAMETOOLONG;
 
@@ -150,6 +148,15 @@ __bpf_kfunc int bpf_binprm_set_interp(struct linux_binprm *bprm,
 	if (!interp)
 		return -ENOMEM;
 
+	/*
+	 * The program may pass memory that is written to while this runs,
+	 * so check the private copy and not the buffer it was made from.
+	 */
+	if (interp[0] != '/') {
+		kfree(interp);
+		return -EINVAL;
+	}
+
 	bm_bpf_stage_selection(bprm, interp, NULL);
 	return 0;
 }
@@ -241,6 +248,15 @@ __bpf_kfunc int bpf_binprm_set_interp_arg(struct linux_binprm *bprm,
 	if (!val)
 		return -ENOMEM;
 
+	/*
+	 * The program may pass memory that is written to while this runs,
+	 * so check the private copy and not the buffer it was made from.
+	 */
+	if (!val[0]) {
+		kfree(val);
+		return -EINVAL;
+	}
+
 	kfree(bprm->bpf_interp_arg);
 	bprm->bpf_interp_arg = val;
 	return 0;

-- 
2.53.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-18  9:20 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-18  9:19 [PATCH 0/2] binfmt: fixes for kres reports Christian Brauner
2026-09-18  9:19 ` [PATCH 1/2] binfmt_misc: fix OOB read in bpf_binprm_select_interp() Christian Brauner
2026-09-18  9:20 ` [PATCH 2/2] binfmt_misc: fix racy checks in bpf set_interp kfuncs Christian Brauner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®