From: Peter Zijlstra <peterz@infradead.org>
To: Jann Horn <jannh@google.com>
Cc: "Hyunwoo Kim" <imv4bel@gmail.com>,
"Thomas Gleixner" <tglx@kernel.org>,
"Ingo Molnar" <mingo@redhat.com>,
"Darren Hart" <dvhart@infradead.org>,
"Davidlohr Bueso" <dave@stgolabs.net>,
"André Almeida" <andrealmeid@igalia.com>,
"kernel list" <linux-kernel@vger.kernel.org>
Subject: Re: [BUG] futex: scheduling-while-atomic because nested vfork can break guard(private_hash)
Date: Fri, 11 Sep 2026 10:36:39 +0200 [thread overview]
Message-ID: <20260911083639.GX776954@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <CAG48ez26PrUOb8er9TP1gWLsqVLU8TEDr+1Zs25fgH0++=zhUQ@mail.gmail.com>
On Thu, Sep 10, 2026 at 05:27:54PM +0200, Jann Horn wrote:
> Introduction
> ============
> The futex private hash implementation was written with the assumption
> that any MM used by more than one running task either has a private
> hash or has irrevocably opted out of having a private hash; but this
> isn't actually the case.
>
> Commit ee9dce44362b ("futex: Drop CLONE_THREAD requirement for private
> default hash alloc") tried to fix cases where processes can share an
> MM without triggering private hash allocation, but overlooked that
> nested vfork can also result in concurrently running processes sharing
> an MM. After that was discovered, commit bde023808364 ("futex: Fix
> race on the initial mm->futex.phash.ref allocation") instead changed
> futex_hash_allocate() to not rely on this assumption anymore; but
> other places in the futex code still make this assumption.
>
> Issue description
> ============
> exit_pi_state_list() contains this code to handle waiters on PI
> futexes held by the current task, which is exiting or going through
> execve:
> ```
What's with the markdown tags? The email is text/plain, so plain it
should be.
> /*
> * Ensure the hash remains stable (no resize) during the while loop
> * below. The hb pointer is acquired under the pi_lock so we can't block
> * on the mutex.
> */
> [...]
> guard(private_hash)(current->mm);
> [...]
> raw_spin_lock_irq(&curr->pi_lock);
> while (!list_empty(head)) {
> [...]
> if (1) {
> CLASS(hbr, hbr)(&key);
> ```
>
> CLASS(hbr, hbr)(&key) calls futex_hash(), which (if the MM is in the
> middle of switching between two futex_private_hash instances) can call
> futex_pivot_hash(), which can block on scoped_guard(mutex,
> &mm->futex.phash.lock). guard(private_hash) is supposed to have taken
> a reference on the futex_private_hash, which would prevent switching
> to another futex_private_hash instance; but guard(private_hash) will
> have been a no-op if the MM didn't have a futex_private_hash yet at
> that time.
*groan*.
> So it is possible to race as follows:
>
> First, use nested vfork to create three processes P1, P2, and P3 that
> share one MM and can run concurrently. That can be done like this:
>
> - P1 vforks to create P2'
> - P2' vforks to create P2
> - P2 sends SIGKILL to P2' to unblock P1
> - P1 vforks to create P3'
> - P3' vforks to create P3
> - P3 sends SIGKILL to P3' to unblock P1
>
> Then race like this (view this in monospace):
What a crazy ass world we live in that this has to be stated :-( Of
course you should view text/plain in monospace, anything else would be
absolutely insane. (And yes, I know outlook exists, but people using
that deserve all the pain that piece of shit gets them.)
> Impact, related issues
> ==========
> futex_wait_multiple_setup() follows the same pattern of using
> guard(private_hash)(current->mm) to ensure that a later CLASS(hbr,
> hbr)(&q->key) won't sleep and is probably affected by the same issue.
>
> __futex_unlock_pi() and requeue_pi_wake_futex() also use
> futex_private_hash(), but there the race seems benign.
>
> This issue also means that different futex users could disagree about
> the hash bucket that a private futex belongs in (older futex_hash()
> calls returning a pointer into the global hash, newer futex_hash()
> calls returning a pointer into the private hash).
Right, that transition is not supposed to be possible. Notably a single
thread cannot have (private) futex waiters, and we allocate the private
hash on cloning the second thread.
> Sidenote
> ========
> Maybe we should block CLONE_VFORK when current->vfork_done!=NULL ? It
> is clearly bad to allow nested vfork() such that a SIGKILL can cause
> two threads to run concurrently on the same userspace stack.
>
> But I'm not sure if that's a good fix for these futex issues - a fix
> that doesn't rely on such distant assumptions might be nicer...
So need_futex_hash_allocate_default() is explicitly excluding vfork from
causing a private hash to be allocated. Perhaps we should fix that.
I need more thinking (and wake-up juice) to see if that might perhaps
bring other problems with it.
> Reproducer
> ==========
> Tested at mainline commit 50d05c7c76c96b90462f24debacca971d2e86713,
Thanks, I'll go poke at this a bit.
next prev parent reply other threads:[~2026-09-11 8:36 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 15:27 Jann Horn
2026-09-11 8:36 ` Peter Zijlstra [this message]
2026-09-11 9:04 ` Peter Zijlstra
2026-09-11 11:51 ` Thomas Gleixner
2026-09-11 15:28 ` Jann Horn
2026-09-11 17:51 ` Davidlohr Bueso
2026-09-15 11:06 ` Sebastian Andrzej Siewior
2026-09-15 11:37 ` Peter Zijlstra
2026-09-11 15:31 ` Jann Horn
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260911083639.GX776954@noisy.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=andrealmeid@igalia.com \
--cc=dave@stgolabs.net \
--cc=dvhart@infradead.org \
--cc=imv4bel@gmail.com \
--cc=jannh@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®