From: Shakeel Butt <shakeel.butt@linux.dev>
To: Peter Zijlstra <peterz@infradead.org>,
Suren Baghdasaryan <surenb@google.com>,
Johannes Weiner <hannes@cmpxchg.org>
Cc: Tejun Heo <tj@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Meta kernel team <kernel-team@meta.com>,
cgroups@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [PATCH] sched/psi: Remove the RCU wait from trigger destruction
Date: Sat, 10 Oct 2026 12:15:00 -0700 [thread overview]
Message-ID: <20261010191500.3895736-1-shakeel.butt@linux.dev> (raw)
In the Meta fleet, we found PSI trigger destruction waiting for an RCU
grace period while holding kernfs node_mutex, blocking other users [1].
File close and node drain hold this mutex while calling the release
callback. Cgroup removal and disabling cgroup.pressure also hold
cgroup_mutex during the drain.
The wait was added by commit 0e94682b73bf ("psi: introduce psi monitor")
to protect worker and trigger lookups. Commit 461daba06bdc
("psi: eliminate kthread_worker from psi trigger scheduling mechanism")
replaced the worker queue with a group timer. The scheduler now only
checks whether rtpoll_task is NULL; it uses neither the task nor the
trigger. Commit a06247c6804f ("psi: Fix uaf issue when psi trigger is
destroyed while being polled") removed trigger replacement and polling's
RCU lookup.
Trigger lists are protected by mutexes. Proc poll/select entries are
removed before their file references are dropped. eventpoll_release()
removes epoll entries before the file release callback.
Commit aff037078eca ("sched/psi: use kernfs polling functions for PSI
trigger polling") moved cgroup polling to a kernfs waitqueue whose
lifetime follows the file.
A late timer firing only wakes the group waitqueue. The system group is
static, and commit 5457025fa8ca ("sched/psi: Shut down rtpoll_timer in
psi_cgroup_free()") shuts down the cgroup timer before freeing the group.
Cgroup reclamation already waits for scheduler readers.
Remove synchronize_rcu() from psi_trigger_destroy() and use
rcu_access_pointer() for the worker NULL check. Keep kthread_stop()
outside the trigger mutex.
Tested trigger churn, notifications and cgroup removal in an 8-CPU VM
with KASAN, lockdep, and full and lazy preemption. No warnings were found.
Link: https://github.com/bpftrace/user-tools/tree/master/runnablelockmonitor [1]
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
---
kernel/sched/psi.c | 25 +++++--------------------
1 file changed, 5 insertions(+), 20 deletions(-)
diff --git a/kernel/sched/psi.c b/kernel/sched/psi.c
index 4e152410653d..7c5381423334 100644
--- a/kernel/sched/psi.c
+++ b/kernel/sched/psi.c
@@ -626,8 +626,6 @@ static void init_rtpoll_triggers(struct psi_group *group, u64 now)
static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long delay,
bool force)
{
- struct task_struct *task;
-
/*
* atomic_xchg should be called even when !force to provide a
* full memory barrier (see the comment inside psi_rtpoll_work).
@@ -635,19 +633,16 @@ static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long dela
if (atomic_xchg(&group->rtpoll_scheduled, 1) && !force)
return;
- rcu_read_lock();
-
- task = rcu_dereference(group->rtpoll_task);
/*
- * kworker might be NULL in case psi_trigger_destroy races with
- * psi_task_change (hotpath) which can't use locks
+ * Only test whether a worker is installed; do not dereference the task.
+ * A racing trigger destruction may leave the timer armed.
+ * psi_cgroup_free() shuts it down before freeing the group.
+ * The system PSI group is static.
*/
- if (likely(task))
+ if (likely(rcu_access_pointer(group->rtpoll_task)))
mod_timer(&group->rtpoll_timer, jiffies + delay);
else
atomic_set(&group->rtpoll_scheduled, 0);
-
- rcu_read_unlock();
}
static void psi_rtpoll_work(struct psi_group *group)
@@ -1488,22 +1483,12 @@ void psi_trigger_destroy(struct psi_trigger *t)
mutex_unlock(&group->rtpoll_trigger_lock);
}
- /*
- * Wait for psi_schedule_rtpoll_work RCU to complete its read-side
- * critical section before destroying the trigger and optionally the
- * rtpoll_task.
- */
- synchronize_rcu();
/*
* Stop kthread 'psimon' after releasing rtpoll_trigger_lock to prevent
* a deadlock while waiting for psi_rtpoll_work to acquire
* rtpoll_trigger_lock
*/
if (task_to_destroy) {
- /*
- * After the RCU grace period has expired, the worker
- * can no longer be found through group->rtpoll_task.
- */
kthread_stop(task_to_destroy);
atomic_set(&group->rtpoll_scheduled, 0);
}
--
2.53.0-Meta
reply other threads:[~2026-10-10 19:15 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261010191500.3895736-1-shakeel.butt@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=bsegall@google.com \
--cc=cgroups@vger.kernel.org \
--cc=dietmar.eggemann@arm.com \
--cc=hannes@cmpxchg.org \
--cc=juri.lelli@redhat.com \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=surenb@google.com \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®