mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Chris Mason" <mason@kernel.org>
To: "Peter Zijlstra" <peterz@infradead.org>,
	"Chris Mason" <mason@kernel.org>
Cc: <tglx@kernel.org>, <linux-kernel@vger.kernel.org>,
	"Paul McKenney" <paulmck@kernel.org>
Subject: Re: [PATCH] futex: sample poll cookie after publishing new hash
Date: Fri, 25 Sep 2026 17:58:13 +0000	[thread overview]
Message-ID: <DLOLBXAIDH5J.2H6LPY4E4Q2RA@kernel.org> (raw)
In-Reply-To: <20260820104015.GE687043@noisy.programming.kicks-ass.net>

On Thu Aug 20, 2026 at 10:40 AM UTC, Peter Zijlstra wrote:
> On Mon, Aug 17, 2026 at 06:01:42PM -0700, Chris Mason wrote:
>> futex_ref_drop() may only skip its grace period when one has already
>> elapsed since the current private hash was published:
>>
>>     kernel/futex/core.c:futex_ref_drop
>>         if (poll_state_synchronize_rcu(mm->futex.phash.batches)) {
>>                 /*
>>                  * There was a grace-period, we can begin now.
>>                  */
>>                 __futex_ref_atomic_begin(fph);
>>                 return;
>>         }
>>
>> The cookie it polls is sampled one statement before that publication:
>>
>>     kernel/futex/core.c:__futex_pivot_hash
>>         new->state = FR_PERCPU;
>>         scoped_guard(rcu) {
>>                 mmph->batches = get_state_synchronize_rcu();
>>                 rcu_assign_pointer(mmph->hash, new);
>>         }
>>         kvfree_rcu(fph, rcu);
>>
>> get_state_synchronize_rcu() anchors its guarantee at the snapshot, so
>> the cookie is cleared by the first grace period that starts from there
>> on, including one starting between the two stores which never waited
>> for a reader that loaded the old hash after it began.
>>
>>     CPU 0 (resize)                    CPU 1 (futex_hash)
>>     ==============                    ==================
>>     __futex_pivot_hash()
>>       batches = get_state_...()
>>                                       grace period starts
>>                                       guard(rcu)
>>                                       fph = old hash
>>       rcu_assign_pointer(hash, new)
>>       kvfree_rcu(old hash)
>>     futex_hash_allocate()
>>       futex_ref_drop(new hash)
>>         poll_state_...() -> true
>>         __futex_ref_atomic_begin()
>>           atomic = LONG_MAX
>>                                       futex_ref_get(old hash) -> true
>>                                       spin_lock(&fph->queues[i].lock)
>>
>> The reference count lives in the mm and has just been biased for the
>> new generation, so the stalled reader pins and then locks the retired
>> hash that is already queued for free, and its later put is charged
>> against the live generation.
>>
>> Fix by sampling the cookie after rcu_assign_pointer() publishes the
>> new hash. Drop the surrounding scoped_guard(rcu) while at it: it only
>> delayed completion of the prematurely anchored grace-period and serves
>> no purpose once the cookie is sampled after publication. The writer
>> side is serialized by mm->futex.phash.lock and neither
>> rcu_assign_pointer() nor kvfree_rcu() requires a read-side section.
>
> God, how I hate reading AI output :-(
>
> Anyway, the thinking was that by holding rcu_read_lock(), the current
> RCU-GP cannot change and the cookie and assignment are effectively
> 'atomic'.

From what I can tell there are a few ways for new grace periods to start
while we're holding rcu_read_lock(), synchronize_rcu_expedited() if no
GP is currently in flight being the easiest?

Anyway, this BUG_ON() fires for me:

diff --git a/kernel/futex/core.c b/kernel/futex/core.c
index a061f54b6..18d51c145 100644
--- a/kernel/futex/core.c
+++ b/kernel/futex/core.c
@@ -215,6 +215,14 @@ static bool __futex_pivot_hash(struct mm_struct *mm, struct futex_private_hash *
 	new->state = FR_PERCPU;
 	scoped_guard(rcu) {
 		mmph->batches = get_state_synchronize_rcu();
+		/*
+		 * Fires only if a grace period started after the cookie
+		 * above was taken, although rcu_read_lock() is held. That
+		 * grace period then satisfies the stored cookie, yet it began
+		 * before the new hash is published below.
+		 */
+		BUG_ON(!same_state_synchronize_rcu(mmph->batches,
+						   get_state_synchronize_rcu()));
 		rcu_assign_pointer(mmph->hash, new);
 	}
 	kvfree_rcu(fph, rcu);

-chris

      reply	other threads:[~2026-09-25 17:58 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-18  1:01 [PATCH RFC] __futex_pivot_hash() race Chris Mason
2026-08-18  1:01 ` [PATCH] futex: sample poll cookie after publishing new hash Chris Mason
2026-08-20 10:40   ` Peter Zijlstra
2026-09-25 17:58     ` Chris Mason [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLOLBXAIDH5J.2H6LPY4E4Q2RA@kernel.org \
    --to=mason@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=paulmck@kernel.org \
    --cc=peterz@infradead.org \
    --cc=tglx@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®