* [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring)
@ 2026-09-24 19:26 Jann Horn
2026-09-28 7:09 ` Peter Zijlstra
0 siblings, 1 reply; 2+ messages in thread
From: Jann Horn @ 2026-09-24 19:26 UTC (permalink / raw)
To: Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Darren Hart,
Davidlohr Bueso, André Almeida
Cc: kernel list, Jens Axboe, io-uring
== Summary ==
FUT_OFF_MMSHARED futex waiters can outlive the mm_struct that the
futex_key refers to because of futex_requeue() [tested] and probably
also via io_uring [untested].
This doesn't have much practical impact, I think at most it could:
1. cause missed wakeups for non-private futexes in anonymous memory
2. allow unrelated processes to guess addresses of FUT_OFF_MMSHARED futexes.
I propose we handle FUT_OFF_MMSHARED like FUT_OFF_INODE and put a
unique 64-bit ID into struct futex_mm_data, which can be used in futex
keys. That should address both futex_requeue() and io_uring. If that
sounds good, I'd be happy to write a patch.
== Detailed description ==
Commit 222993395ed3 ("futex: Remove pointless mmgrap() + mmdrop()") claims:
"We always set 'key->private.mm' to 'current->mm', getting an extra
reference on 'current->mm' is quite pointless, because as long as the
task is blocked it isn't going to go away."
This assumes that a FUT_OFF_MMSHARED futex waiter is always waiting on
its own mm_struct. That assumption does not hold in the following
scenario:
1. process A starts waiting on a FUT_OFF_INODE key with
futex_wait(shm_addr, ...)
2. process B uses futex_requeue({{.uaddr=shm_addr},
{.uaddr=anon_addr}}, ...) to move the waiter to a FUT_OFF_MMSHARED
futex associated with the mm_struct of process B
3. process B exits and its mm_struct goes away
4. process C is created, with its mm_struct allocated at the same
kernel address where process B's used to be
5. process C calls futex_wake(anon_addr, ..., nr=1, ...), which
wrongly wakes process A (and won't wake other legitimate wakers
because of nr=1)
== io_uring interaction ==
io_uring permits asynchronous waiting on futexes; io_futex_wait_prep()
can use io_req_track_inflight() to ensure that the futex waiter will
be destroyed before the calling task goes through exit or exec, but
only does so for private futexes, not for FUT_OFF_MMSHARED waiters. I
think io_uring FUT_OFF_MMSHARED waiters could probably survive process
exit and outlive the mm_struct; but I haven't tested this.
== Fixing it ==
I see two approaches to fixing this:
1. Put a unique 64-bit ID into struct futex_mm_data, which can be used
in futex keys, similar to how FUT_OFF_INODE is already handled with
get_inode_sequence_number(). That should address both futex_requeue()
and io_uring.
2. Ensure that FUT_OFF_MMSHARED waiters can't wait on another
mm_struct by blocking attempts to requeue FUT_OFF_INODE waiters onto a
FUT_OFF_MMSHARED key. For io_uring, make io_req_track_inflight()
unconditional.
I like option 1 because it also avoids feeding a kernel pointer into a
hash function together with user input, which probably creates a
timing side channel leak of mm_struct pointers.
== Reproducer ==
Here is a reproducer that demonstrates receiving a futex wakeup from
futex_wake() on anonymous memory in an unrelated process:
#include <pthread.h>
#include <err.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <sys/wait.h>
#include <linux/futex.h>
#define FLAGS_SIZE_32 0x0002
#define SYSCHK(x) ({ \
typeof(x) __res = (x); \
if (__res == (typeof(x))-1) \
err(1, "SYSCHK(" #x ")"); \
__res; \
})
int futex_wait_until_nonzero(void *uaddr) {
return syscall(__NR_futex_wait, uaddr, /*val=*/0, /*mask=*/0x1,
/*flags=*/FLAGS_SIZE_32, NULL, 0);
}
int futex_wake(void *uaddr) {
return syscall(__NR_futex_wake, uaddr, /*mask=*/0x1, /*nr=*/100,
/*flags=*/FLAGS_SIZE_32);
}
int futex_requeue(void *src, void *dst) {
struct futex_waitv waiters[2] = {
{.val = 0, .uaddr = (unsigned long)src, .flags = FLAGS_SIZE_32},
{.val = 0, .uaddr = (unsigned long)dst, .flags = FLAGS_SIZE_32},
};
return syscall(__NR_futex_requeue, waiters, 0, /*nr_wake=*/0,
/*nr_requeue=*/100);
}
static char *shm;
static void *waiter_thread_fn(void *dummy) {
printf("[waiter_thread] futex_wait_until_nonzero(shm) = %d\n",
futex_wait_until_nonzero(shm));
return NULL;
}
int main(void) {
setbuf(stdout, NULL);
shm = SYSCHK(mmap(NULL, 0x1000, PROT_READ|PROT_WRITE,
MAP_SHARED|MAP_ANONYMOUS, -1, 0));
*shm = 0;
pthread_t waiter_thread;
pthread_create(&waiter_thread, NULL, waiter_thread_fn, NULL);
pid_t child1 = SYSCHK(fork());
if (child1 == 0) {
usleep(10000);
char *priv = SYSCHK(mmap((void*)0x10000, 0x1000,
PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_FIXED_NOREPLACE,
-1, 0));
printf("[child1] futex_requeue(shm, priv) = %d\n",
futex_requeue(shm, priv));
exit(0);
}
int wstatus;
SYSCHK(waitpid(child1, &wstatus, 0));
for (int i=0; i<10000; i++) {
pid_t child2 = SYSCHK(fork());
if (child2 == 0) {
char *priv = SYSCHK(mmap((void*)0x10000, 0x1000,
PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_FIXED_NOREPLACE,
-1, 0));
int wake_res = futex_wake(priv);
if (wake_res)
printf("[child2, iteration %d] futex_wake(priv) = %d\n", i, wake_res);
exit(0);
}
SYSCHK(waitpid(child2, &wstatus, 0));
}
exit(0);
}
Testing on current mainline (fe2ec83746e5), built without KASAN, I get
output like:
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 20] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 261] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 6] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 31] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3#
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring)
2026-09-24 19:26 [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring) Jann Horn
@ 2026-09-28 7:09 ` Peter Zijlstra
0 siblings, 0 replies; 2+ messages in thread
From: Peter Zijlstra @ 2026-09-28 7:09 UTC (permalink / raw)
To: Jann Horn
Cc: Thomas Gleixner, Ingo Molnar, Darren Hart, Davidlohr Bueso,
André Almeida, kernel list, Jens Axboe, io-uring
On Thu, Sep 24, 2026 at 09:26:45PM +0200, Jann Horn wrote:
> == Summary ==
> FUT_OFF_MMSHARED futex waiters can outlive the mm_struct that the
> futex_key refers to because of futex_requeue() [tested] and probably
> also via io_uring [untested].
>
> This doesn't have much practical impact, I think at most it could:
> 1. cause missed wakeups for non-private futexes in anonymous memory
> 2. allow unrelated processes to guess addresses of FUT_OFF_MMSHARED futexes.
>
> I propose we handle FUT_OFF_MMSHARED like FUT_OFF_INODE and put a
> unique 64-bit ID into struct futex_mm_data, which can be used in futex
> keys. That should address both futex_requeue() and io_uring. If that
> sounds good, I'd be happy to write a patch.
Whee, you do find the best problems :-)
> == Fixing it ==
> I see two approaches to fixing this:
>
> 1. Put a unique 64-bit ID into struct futex_mm_data, which can be used
> in futex keys, similar to how FUT_OFF_INODE is already handled with
> get_inode_sequence_number(). That should address both futex_requeue()
> and io_uring.
> 2. Ensure that FUT_OFF_MMSHARED waiters can't wait on another
> mm_struct by blocking attempts to requeue FUT_OFF_INODE waiters onto a
> FUT_OFF_MMSHARED key. For io_uring, make io_req_track_inflight()
> unconditional.
>
> I like option 1 because it also avoids feeding a kernel pointer into a
> hash function together with user input, which probably creates a
> timing side channel leak of mm_struct pointers.
I agree that option 1 is the sanest. Please write that patch.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-28 7:09 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-24 19:26 [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring) Jann Horn
2026-09-28 7:09 ` Peter Zijlstra
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®