From: "Chris Mason" <mason@kernel.org>
To: "Peter Zijlstra" <peterz@infradead.org>,
"Chris Mason" <mason@kernel.org>
Cc: <tglx@kernel.org>, <linux-kernel@vger.kernel.org>,
"Paul McKenney" <paulmck@kernel.org>
Subject: Re: [PATCH] futex: sample poll cookie after publishing new hash
Date: Fri, 25 Sep 2026 17:58:13 +0000 [thread overview]
Message-ID: <DLOLBXAIDH5J.2H6LPY4E4Q2RA@kernel.org> (raw)
In-Reply-To: <20260820104015.GE687043@noisy.programming.kicks-ass.net>
On Thu Aug 20, 2026 at 10:40 AM UTC, Peter Zijlstra wrote:
> On Mon, Aug 17, 2026 at 06:01:42PM -0700, Chris Mason wrote:
>> futex_ref_drop() may only skip its grace period when one has already
>> elapsed since the current private hash was published:
>>
>> kernel/futex/core.c:futex_ref_drop
>> if (poll_state_synchronize_rcu(mm->futex.phash.batches)) {
>> /*
>> * There was a grace-period, we can begin now.
>> */
>> __futex_ref_atomic_begin(fph);
>> return;
>> }
>>
>> The cookie it polls is sampled one statement before that publication:
>>
>> kernel/futex/core.c:__futex_pivot_hash
>> new->state = FR_PERCPU;
>> scoped_guard(rcu) {
>> mmph->batches = get_state_synchronize_rcu();
>> rcu_assign_pointer(mmph->hash, new);
>> }
>> kvfree_rcu(fph, rcu);
>>
>> get_state_synchronize_rcu() anchors its guarantee at the snapshot, so
>> the cookie is cleared by the first grace period that starts from there
>> on, including one starting between the two stores which never waited
>> for a reader that loaded the old hash after it began.
>>
>> CPU 0 (resize) CPU 1 (futex_hash)
>> ============== ==================
>> __futex_pivot_hash()
>> batches = get_state_...()
>> grace period starts
>> guard(rcu)
>> fph = old hash
>> rcu_assign_pointer(hash, new)
>> kvfree_rcu(old hash)
>> futex_hash_allocate()
>> futex_ref_drop(new hash)
>> poll_state_...() -> true
>> __futex_ref_atomic_begin()
>> atomic = LONG_MAX
>> futex_ref_get(old hash) -> true
>> spin_lock(&fph->queues[i].lock)
>>
>> The reference count lives in the mm and has just been biased for the
>> new generation, so the stalled reader pins and then locks the retired
>> hash that is already queued for free, and its later put is charged
>> against the live generation.
>>
>> Fix by sampling the cookie after rcu_assign_pointer() publishes the
>> new hash. Drop the surrounding scoped_guard(rcu) while at it: it only
>> delayed completion of the prematurely anchored grace-period and serves
>> no purpose once the cookie is sampled after publication. The writer
>> side is serialized by mm->futex.phash.lock and neither
>> rcu_assign_pointer() nor kvfree_rcu() requires a read-side section.
>
> God, how I hate reading AI output :-(
>
> Anyway, the thinking was that by holding rcu_read_lock(), the current
> RCU-GP cannot change and the cookie and assignment are effectively
> 'atomic'.
From what I can tell there are a few ways for new grace periods to start
while we're holding rcu_read_lock(), synchronize_rcu_expedited() if no
GP is currently in flight being the easiest?
Anyway, this BUG_ON() fires for me:
diff --git a/kernel/futex/core.c b/kernel/futex/core.c
index a061f54b6..18d51c145 100644
--- a/kernel/futex/core.c
+++ b/kernel/futex/core.c
@@ -215,6 +215,14 @@ static bool __futex_pivot_hash(struct mm_struct *mm, struct futex_private_hash *
new->state = FR_PERCPU;
scoped_guard(rcu) {
mmph->batches = get_state_synchronize_rcu();
+ /*
+ * Fires only if a grace period started after the cookie
+ * above was taken, although rcu_read_lock() is held. That
+ * grace period then satisfies the stored cookie, yet it began
+ * before the new hash is published below.
+ */
+ BUG_ON(!same_state_synchronize_rcu(mmph->batches,
+ get_state_synchronize_rcu()));
rcu_assign_pointer(mmph->hash, new);
}
kvfree_rcu(fph, rcu);
-chris
prev parent reply other threads:[~2026-09-25 17:58 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 1:01 [PATCH RFC] __futex_pivot_hash() race Chris Mason
2026-08-18 1:01 ` [PATCH] futex: sample poll cookie after publishing new hash Chris Mason
2026-08-20 10:40 ` Peter Zijlstra
2026-09-25 17:58 ` Chris Mason [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLOLBXAIDH5J.2H6LPY4E4Q2RA@kernel.org \
--to=mason@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=paulmck@kernel.org \
--cc=peterz@infradead.org \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®