mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] sched/psi: Remove the RCU wait from trigger destruction
@ 2026-10-10 19:15 Shakeel Butt
  0 siblings, 0 replies; only message in thread
From: Shakeel Butt @ 2026-10-10 19:15 UTC (permalink / raw)
  To: Peter Zijlstra, Suren Baghdasaryan, Johannes Weiner
  Cc: Tejun Heo, Ingo Molnar, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Meta kernel team, cgroups,
	linux-kernel

In the Meta fleet, we found PSI trigger destruction waiting for an RCU
grace period while holding kernfs node_mutex, blocking other users [1].
File close and node drain hold this mutex while calling the release
callback. Cgroup removal and disabling cgroup.pressure also hold
cgroup_mutex during the drain.

The wait was added by commit 0e94682b73bf ("psi: introduce psi monitor")
to protect worker and trigger lookups. Commit 461daba06bdc
("psi: eliminate kthread_worker from psi trigger scheduling mechanism")
replaced the worker queue with a group timer. The scheduler now only
checks whether rtpoll_task is NULL; it uses neither the task nor the
trigger. Commit a06247c6804f ("psi: Fix uaf issue when psi trigger is
destroyed while being polled") removed trigger replacement and polling's
RCU lookup.

Trigger lists are protected by mutexes. Proc poll/select entries are
removed before their file references are dropped. eventpoll_release()
removes epoll entries before the file release callback.
Commit aff037078eca ("sched/psi: use kernfs polling functions for PSI
trigger polling") moved cgroup polling to a kernfs waitqueue whose
lifetime follows the file.

A late timer firing only wakes the group waitqueue. The system group is
static, and commit 5457025fa8ca ("sched/psi: Shut down rtpoll_timer in
psi_cgroup_free()") shuts down the cgroup timer before freeing the group.
Cgroup reclamation already waits for scheduler readers.

Remove synchronize_rcu() from psi_trigger_destroy() and use
rcu_access_pointer() for the worker NULL check. Keep kthread_stop()
outside the trigger mutex.

Tested trigger churn, notifications and cgroup removal in an 8-CPU VM
with KASAN, lockdep, and full and lazy preemption. No warnings were found.

Link: https://github.com/bpftrace/user-tools/tree/master/runnablelockmonitor [1]
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
---
 kernel/sched/psi.c | 25 +++++--------------------
 1 file changed, 5 insertions(+), 20 deletions(-)

diff --git a/kernel/sched/psi.c b/kernel/sched/psi.c
index 4e152410653d..7c5381423334 100644
--- a/kernel/sched/psi.c
+++ b/kernel/sched/psi.c
@@ -626,8 +626,6 @@ static void init_rtpoll_triggers(struct psi_group *group, u64 now)
 static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long delay,
 				   bool force)
 {
-	struct task_struct *task;
-
 	/*
 	 * atomic_xchg should be called even when !force to provide a
 	 * full memory barrier (see the comment inside psi_rtpoll_work).
@@ -635,19 +633,16 @@ static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long dela
 	if (atomic_xchg(&group->rtpoll_scheduled, 1) && !force)
 		return;
 
-	rcu_read_lock();
-
-	task = rcu_dereference(group->rtpoll_task);
 	/*
-	 * kworker might be NULL in case psi_trigger_destroy races with
-	 * psi_task_change (hotpath) which can't use locks
+	 * Only test whether a worker is installed; do not dereference the task.
+	 * A racing trigger destruction may leave the timer armed.
+	 * psi_cgroup_free() shuts it down before freeing the group.
+	 * The system PSI group is static.
 	 */
-	if (likely(task))
+	if (likely(rcu_access_pointer(group->rtpoll_task)))
 		mod_timer(&group->rtpoll_timer, jiffies + delay);
 	else
 		atomic_set(&group->rtpoll_scheduled, 0);
-
-	rcu_read_unlock();
 }
 
 static void psi_rtpoll_work(struct psi_group *group)
@@ -1488,22 +1483,12 @@ void psi_trigger_destroy(struct psi_trigger *t)
 		mutex_unlock(&group->rtpoll_trigger_lock);
 	}
 
-	/*
-	 * Wait for psi_schedule_rtpoll_work RCU to complete its read-side
-	 * critical section before destroying the trigger and optionally the
-	 * rtpoll_task.
-	 */
-	synchronize_rcu();
 	/*
 	 * Stop kthread 'psimon' after releasing rtpoll_trigger_lock to prevent
 	 * a deadlock while waiting for psi_rtpoll_work to acquire
 	 * rtpoll_trigger_lock
 	 */
 	if (task_to_destroy) {
-		/*
-		 * After the RCU grace period has expired, the worker
-		 * can no longer be found through group->rtpoll_task.
-		 */
 		kthread_stop(task_to_destroy);
 		atomic_set(&group->rtpoll_scheduled, 0);
 	}
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2026-10-10 19:15 UTC | newest]

Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 19:15 [PATCH] sched/psi: Remove the RCU wait from trigger destruction Shakeel Butt

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®