From: "Chris Mason" <mason@kernel.org>
To: "Peter Zijlstra" <peterz@infradead.org>,
"Paul E. McKenney" <paulmck@kernel.org>
Cc: tglx@kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] futex: sample poll cookie after publishing new hash
Date: Wed, 30 Sep 2026 09:23:52 -0400 [thread overview]
Message-ID: <d6e1b8cf-ee16-4a8f-b2e5-abc52908e6e3@app.fastmail.com> (raw)
In-Reply-To: <20260930112156.GL88198@noisy.programming.kicks-ass.net>
On Wed, Sep 30, 2026, at 7:21 AM, Peter Zijlstra wrote:
> On Tue, Sep 29, 2026 at 02:12:53PM -0700, Paul E. McKenney wrote:
>> On Fri, Sep 25, 2026 at 05:58:13PM +0000, Chris Mason wrote:
>> > On Thu Aug 20, 2026 at 10:40 AM UTC, Peter Zijlstra wrote:
>> > > On Mon, Aug 17, 2026 at 06:01:42PM -0700, Chris Mason wrote:
>> > >> futex_ref_drop() may only skip its grace period when one has already
>> > >> elapsed since the current private hash was published:
>> > >>
>> > >> kernel/futex/core.c:futex_ref_drop
>> > >> if (poll_state_synchronize_rcu(mm->futex.phash.batches)) {
>> > >> /*
>> > >> * There was a grace-period, we can begin now.
>> > >> */
>> > >> __futex_ref_atomic_begin(fph);
>> > >> return;
>> > >> }
>> > >>
>> > >> The cookie it polls is sampled one statement before that publication:
>> > >>
>> > >> kernel/futex/core.c:__futex_pivot_hash
>> > >> new->state = FR_PERCPU;
>> > >> scoped_guard(rcu) {
>> > >> mmph->batches = get_state_synchronize_rcu();
>> > >> rcu_assign_pointer(mmph->hash, new);
>> > >> }
>> > >> kvfree_rcu(fph, rcu);
>> > >>
>> > >> get_state_synchronize_rcu() anchors its guarantee at the snapshot, so
>> > >> the cookie is cleared by the first grace period that starts from there
>> > >> on, including one starting between the two stores which never waited
>> > >> for a reader that loaded the old hash after it began.
>> > >>
>> > >> CPU 0 (resize) CPU 1 (futex_hash)
>> > >> ============== ==================
>> > >> __futex_pivot_hash()
>> > >> batches = get_state_...()
>> > >> grace period starts
>> > >> guard(rcu)
>> > >> fph = old hash
>> > >> rcu_assign_pointer(hash, new)
>> > >> kvfree_rcu(old hash)
>> > >> futex_hash_allocate()
>> > >> futex_ref_drop(new hash)
>> > >> poll_state_...() -> true
>> > >> __futex_ref_atomic_begin()
>> > >> atomic = LONG_MAX
>> > >> futex_ref_get(old hash) -> true
>> > >> spin_lock(&fph->queues[i].lock)
>> > >>
>> > >> The reference count lives in the mm and has just been biased for the
>> > >> new generation, so the stalled reader pins and then locks the retired
>> > >> hash that is already queued for free, and its later put is charged
>> > >> against the live generation.
>> > >>
>> > >> Fix by sampling the cookie after rcu_assign_pointer() publishes the
>> > >> new hash. Drop the surrounding scoped_guard(rcu) while at it: it only
>> > >> delayed completion of the prematurely anchored grace-period and serves
>> > >> no purpose once the cookie is sampled after publication. The writer
>> > >> side is serialized by mm->futex.phash.lock and neither
>> > >> rcu_assign_pointer() nor kvfree_rcu() requires a read-side section.
>> > >
>> > > God, how I hate reading AI output :-(
>> > >
>> > > Anyway, the thinking was that by holding rcu_read_lock(), the current
>> > > RCU-GP cannot change and the cookie and assignment are effectively
>> > > 'atomic'.
>> >
>> > >From what I can tell there are a few ways for new grace periods to start
>> > while we're holding rcu_read_lock(), synchronize_rcu_expedited() if no
>> > GP is currently in flight being the easiest?
>> >
>> > Anyway, this BUG_ON() fires for me:
>> >
>> > diff --git a/kernel/futex/core.c b/kernel/futex/core.c
>> > index a061f54b6..18d51c145 100644
>> > --- a/kernel/futex/core.c
>> > +++ b/kernel/futex/core.c
>> > @@ -215,6 +215,14 @@ static bool __futex_pivot_hash(struct mm_struct *mm, struct futex_private_hash *
>> > new->state = FR_PERCPU;
>> > scoped_guard(rcu) {
>> > mmph->batches = get_state_synchronize_rcu();
>> > + /*
>> > + * Fires only if a grace period started after the cookie
>> > + * above was taken, although rcu_read_lock() is held. That
>> > + * grace period then satisfies the stored cookie, yet it began
>> > + * before the new hash is published below.
>> > + */
>> > + BUG_ON(!same_state_synchronize_rcu(mmph->batches,
>> > + get_state_synchronize_rcu()));
>> > rcu_assign_pointer(mmph->hash, new);
>> > }
>> > kvfree_rcu(fph, rcu);
>> >
>> > -chris
>>
>> This matches my understanding of RCU. Although rcu_read_lock() will
>> prevent a *new* RCU grace period from ending, it will not prevent an *old*
>> one from ending or a new one from starting.
>
> So, let us consider the dual counter RCU.
>
> The RCU state is:
>
> counter[2];
> index;
>
> rcu_read_lock() would increment counter[index & 1], rcu_read_unlock() would
> decrement whatever counter it incremented (say rcu_read_lock() returns
> the index, like SRCU does).
>
> GP progression can either be sample or rcu_read_unlock() driven. Either
> way, when 'counter[!(index & 1)] == 0' is observed, there are no more
> observers of the old state (and callbacks can be ran).
>
> At this point we can flip the counter and start anew: index++;
>
> So index changes while counter[index & 1] != 0, which is somewhat
> fundamental to how the whole thing works.
>
> Now, suppose get_state_synchronize_rcu() is simply returning a snapshot
> of 'index', then same_state_synchronize_rcu() can indeed trigger. The
> index can get advanced.
>
> But that does not mean poll_state_synchronize_rcu() would return true;
> in the above scheme poll_state_synchronize_rcu() would be something
> like:
>
> if (index - snapshot < 2)
> return counter[snapshot & 1] == 0;
>
> return true;
>
> That is, not only does index need to advance, but also the counter needs
> to drain in order for it to report the GP is complete.
>
> And this is where I went wrong, when we hold rcu_read_lock() we pin the
> counter and thus the GP cannot complete. But as I now see, that is not
> sufficient, because there can be observers of the old state in the next
> GP too.
>
> So yes, the code is wrong and the patch is correct.
Ok, I'll resend against recent git with a better commit message, unless you've already queued up something better suited?
-chris
prev parent reply other threads:[~2026-09-30 13:24 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 1:01 [PATCH RFC] __futex_pivot_hash() race Chris Mason
2026-08-18 1:01 ` [PATCH] futex: sample poll cookie after publishing new hash Chris Mason
2026-08-20 10:40 ` Peter Zijlstra
2026-09-25 17:58 ` Chris Mason
2026-09-29 21:12 ` Paul E. McKenney
2026-09-30 11:21 ` Peter Zijlstra
2026-09-30 13:23 ` Chris Mason [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d6e1b8cf-ee16-4a8f-b2e5-abc52908e6e3@app.fastmail.com \
--to=mason@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=paulmck@kernel.org \
--cc=peterz@infradead.org \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®