mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring)
@ 2026-09-24 19:26 Jann Horn
  2026-09-28  7:09 ` Peter Zijlstra
  0 siblings, 1 reply; 2+ messages in thread
From: Jann Horn @ 2026-09-24 19:26 UTC (permalink / raw)
  To: Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Darren Hart,
	Davidlohr Bueso, André Almeida
  Cc: kernel list, Jens Axboe, io-uring

== Summary ==
FUT_OFF_MMSHARED futex waiters can outlive the mm_struct that the
futex_key refers to because of futex_requeue() [tested] and probably
also via io_uring [untested].

This doesn't have much practical impact, I think at most it could:
1. cause missed wakeups for non-private futexes in anonymous memory
2. allow unrelated processes to guess addresses of FUT_OFF_MMSHARED futexes.

I propose we handle FUT_OFF_MMSHARED like FUT_OFF_INODE and put a
unique 64-bit ID into struct futex_mm_data, which can be used in futex
keys. That should address both futex_requeue() and io_uring. If that
sounds good, I'd be happy to write a patch.

== Detailed description ==
Commit 222993395ed3 ("futex: Remove pointless mmgrap() + mmdrop()") claims:

"We always set 'key->private.mm' to 'current->mm', getting an extra
reference on 'current->mm' is quite pointless, because as long as the
task is blocked it isn't going to go away."

This assumes that a FUT_OFF_MMSHARED futex waiter is always waiting on
its own mm_struct. That assumption does not hold in the following
scenario:

1. process A starts waiting on a FUT_OFF_INODE key with
futex_wait(shm_addr, ...)
2. process B uses futex_requeue({{.uaddr=shm_addr},
{.uaddr=anon_addr}}, ...) to move the waiter to a FUT_OFF_MMSHARED
futex associated with the mm_struct of process B
3. process B exits and its mm_struct goes away
4. process C is created, with its mm_struct allocated at the same
kernel address where process B's used to be
5. process C calls futex_wake(anon_addr, ..., nr=1, ...), which
wrongly wakes process A (and won't wake other legitimate wakers
because of nr=1)

== io_uring interaction ==
io_uring permits asynchronous waiting on futexes; io_futex_wait_prep()
can use io_req_track_inflight() to ensure that the futex waiter will
be destroyed before the calling task goes through exit or exec, but
only does so for private futexes, not for FUT_OFF_MMSHARED waiters. I
think io_uring FUT_OFF_MMSHARED waiters could probably survive process
exit and outlive the mm_struct; but I haven't tested this.

== Fixing it ==
I see two approaches to fixing this:

1. Put a unique 64-bit ID into struct futex_mm_data, which can be used
in futex keys, similar to how FUT_OFF_INODE is already handled with
get_inode_sequence_number(). That should address both futex_requeue()
and io_uring.
2. Ensure that FUT_OFF_MMSHARED waiters can't wait on another
mm_struct by blocking attempts to requeue FUT_OFF_INODE waiters onto a
FUT_OFF_MMSHARED key. For io_uring, make io_req_track_inflight()
unconditional.

I like option 1 because it also avoids feeding a kernel pointer into a
hash function together with user input, which probably creates a
timing side channel leak of mm_struct pointers.

== Reproducer ==
Here is a reproducer that demonstrates receiving a futex wakeup from
futex_wake() on anonymous memory in an unrelated process:

#include <pthread.h>
#include <err.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <sys/wait.h>
#include <linux/futex.h>

#define FLAGS_SIZE_32   0x0002

#define SYSCHK(x) ({          \
  typeof(x) __res = (x);      \
  if (__res == (typeof(x))-1) \
    err(1, "SYSCHK(" #x ")"); \
  __res;                      \
})

int futex_wait_until_nonzero(void *uaddr) {
  return syscall(__NR_futex_wait, uaddr, /*val=*/0, /*mask=*/0x1,
/*flags=*/FLAGS_SIZE_32, NULL, 0);
}

int futex_wake(void *uaddr) {
  return syscall(__NR_futex_wake, uaddr, /*mask=*/0x1, /*nr=*/100,
/*flags=*/FLAGS_SIZE_32);
}

int futex_requeue(void *src, void *dst) {
  struct futex_waitv waiters[2] = {
    {.val = 0, .uaddr = (unsigned long)src, .flags = FLAGS_SIZE_32},
    {.val = 0, .uaddr = (unsigned long)dst, .flags = FLAGS_SIZE_32},
  };
  return syscall(__NR_futex_requeue, waiters, 0, /*nr_wake=*/0,
/*nr_requeue=*/100);
}

static char *shm;

static void *waiter_thread_fn(void *dummy) {
  printf("[waiter_thread] futex_wait_until_nonzero(shm) = %d\n",
futex_wait_until_nonzero(shm));
  return NULL;
}

int main(void) {
  setbuf(stdout, NULL);
  shm = SYSCHK(mmap(NULL, 0x1000, PROT_READ|PROT_WRITE,
MAP_SHARED|MAP_ANONYMOUS, -1, 0));
  *shm = 0;

  pthread_t waiter_thread;
  pthread_create(&waiter_thread, NULL, waiter_thread_fn, NULL);

  pid_t child1 = SYSCHK(fork());
  if (child1 == 0) {
    usleep(10000);
    char *priv = SYSCHK(mmap((void*)0x10000, 0x1000,
PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_FIXED_NOREPLACE,
-1, 0));
    printf("[child1] futex_requeue(shm, priv) = %d\n",
futex_requeue(shm, priv));
    exit(0);
  }
  int wstatus;
  SYSCHK(waitpid(child1, &wstatus, 0));

  for (int i=0; i<10000; i++) {
    pid_t child2 = SYSCHK(fork());
    if (child2 == 0) {
      char *priv = SYSCHK(mmap((void*)0x10000, 0x1000,
PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_FIXED_NOREPLACE,
-1, 0));
      int wake_res = futex_wake(priv);
      if (wake_res)
        printf("[child2, iteration %d] futex_wake(priv) = %d\n", i, wake_res);
      exit(0);
    }
    SYSCHK(waitpid(child2, &wstatus, 0));
  }
  exit(0);
}

Testing on current mainline (fe2ec83746e5), built without KASAN, I get
output like:

sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 20] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 261] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 6] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3# /host/h/test/futex-requeue-inode-to-mm/futex-requeue-inode-to-mm
[child1] futex_requeue(shm, priv) = 1
[child2, iteration 31] futex_wake(priv) = 1
[waiter_thread] futex_wait_until_nonzero(shm) = 0
sh-5.3#

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring)
  2026-09-24 19:26 [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring) Jann Horn
@ 2026-09-28  7:09 ` Peter Zijlstra
  0 siblings, 0 replies; 2+ messages in thread
From: Peter Zijlstra @ 2026-09-28  7:09 UTC (permalink / raw)
  To: Jann Horn
  Cc: Thomas Gleixner, Ingo Molnar, Darren Hart, Davidlohr Bueso,
	André Almeida, kernel list, Jens Axboe, io-uring

On Thu, Sep 24, 2026 at 09:26:45PM +0200, Jann Horn wrote:
> == Summary ==
> FUT_OFF_MMSHARED futex waiters can outlive the mm_struct that the
> futex_key refers to because of futex_requeue() [tested] and probably
> also via io_uring [untested].
> 
> This doesn't have much practical impact, I think at most it could:
> 1. cause missed wakeups for non-private futexes in anonymous memory
> 2. allow unrelated processes to guess addresses of FUT_OFF_MMSHARED futexes.
> 
> I propose we handle FUT_OFF_MMSHARED like FUT_OFF_INODE and put a
> unique 64-bit ID into struct futex_mm_data, which can be used in futex
> keys. That should address both futex_requeue() and io_uring. If that
> sounds good, I'd be happy to write a patch.

Whee, you do find the best problems :-)

> == Fixing it ==
> I see two approaches to fixing this:
> 
> 1. Put a unique 64-bit ID into struct futex_mm_data, which can be used
> in futex keys, similar to how FUT_OFF_INODE is already handled with
> get_inode_sequence_number(). That should address both futex_requeue()
> and io_uring.
> 2. Ensure that FUT_OFF_MMSHARED waiters can't wait on another
> mm_struct by blocking attempts to requeue FUT_OFF_INODE waiters onto a
> FUT_OFF_MMSHARED key. For io_uring, make io_req_track_inflight()
> unconditional.
> 
> I like option 1 because it also avoids feeding a kernel pointer into a
> hash function together with user input, which probably creates a
> timing side channel leak of mm_struct pointers.

I agree that option 1 is the sanest. Please write that patch.

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-28  7:09 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-24 19:26 [BUG] FUT_OFF_MMSHARED futex waiters can outlive mm_struct (because of futex_requeue() and maybe io_uring) Jann Horn
2026-09-28  7:09 ` Peter Zijlstra

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®