From: wen.yang@linux.dev
To: Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>
Cc: Wen Yang <wen.yang@linux.dev>,
Vincent Guittot <vincent.guittot@linaro.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
linux-kernel@vger.kernel.org
Subject: [PATCH 2/2] sched/rt: add RT throttle statistics
Date: Tue, 2 Dec 2025 13:51:19 +0800 [thread overview]
Message-ID: <a732cb65347843e8b5fdbe182363c5438ac0916f.1764648076.git.wen.yang@linux.dev> (raw)
In-Reply-To: <cover.1764648076.git.wen.yang@linux.dev>
From: Wen Yang <wen.yang@linux.dev>
A priority inversion scenario can occur when a CFS task is starved
due to RT throttling. The scenario is as follows:
0. An rtmutex (e.g., softirq_ctrl.lock) is contended by both CFS
tasks (e.g., ksoftirqd) and RT tasks (e.g., ktimer).
1. An RT task 'A' (e.g., ktimer) acquired the rtmutex.
2. A CFS task 'B' (e.g., ksoftirqd) attempts to acquire the same
rtmutex and blocks.
3. A higher-priority RT task 'C' (e.g., stress-ng) runs for an
extended period, preempting task 'A' and causing the RT runqueue
to be throttled.
4. Once rt throttled, CFS task 'B' should run, but it remains blocked
because the lock is still held by the non-running RT task 'A'. This
can even lead to the CPU going idle.
5. When the rt throttle period ends, the high-priority RT task 'C'
resumes execution, and the cycle repeats, leading to indefinite
starvation of CFS task 'B'.
A typical stack trace for the blocked ksoftirqd shows it in a 'D'
(TASK_RTLOCK_WAIT) state, waiting on the lock:
ksoftirqd/5-61 [005] d...211 58212.064160: sched_switch: prev_comm=ksoftirqd/5 prev_pid=61 prev_prio=120 prev_state=D ==> next_comm=swapper/5 next_pid=0 next_prio=120
ksoftirqd/5-61 [005] d...211 58212.064161: <stack trace>
=> __schedule
=> schedule_rtlock
=> rtlock_slowlock_locked
=> rt_spin_lock
=> __local_bh_disable_ip
=> run_ksoftirqd
=> smpboot_thread_fn
=> kthread
=> ret_from_fork
This patch adds throttle_count to rt_rq, incremented on each throttling event
and displayed in print_rt_rq for /proc/sched_debug.
Thus user-space tools (e.g. stalld) can monitor throttle_comunt to detect
the huge CPU consumption by RT processes and find tasks in the
'TASK_RTLOCK_WAIT' state to handle priority inversion.
Signed-off-by: Wen Yang <wen.yang@linux.dev>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Juri Lelli <juri.lelli@redhat.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Ben Segall <bsegall@google.com>
Cc: Mel Gorman <mgorman@suse.de>
Cc: Valentin Schneider <vschneid@redhat.com>
Cc: linux-kernel@vger.kernel.org
---
kernel/sched/debug.c | 1 +
kernel/sched/rt.c | 1 +
kernel/sched/sched.h | 1 +
3 files changed, 3 insertions(+)
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index 41caa22e0680..8ed33c74e5a5 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -894,6 +894,7 @@ void print_rt_rq(struct seq_file *m, int cpu, struct rt_rq *rt_rq)
P(rt_throttled);
PN(rt_time);
PN(rt_runtime);
+ PU(throttle_count);
#endif
#undef PN
diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
index f1867fe8e5c5..88c659285c70 100644
--- a/kernel/sched/rt.c
+++ b/kernel/sched/rt.c
@@ -884,6 +884,7 @@ static int sched_rt_runtime_exceeded(struct rt_rq *rt_rq)
*/
if (likely(rt_b->rt_runtime)) {
rt_rq->rt_throttled = 1;
+ rt_rq->throttle_count++;
printk_deferred_once("sched: RT throttling activated\n");
} else {
/*
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index bbf513b3e76c..88119540e4d4 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -840,6 +840,7 @@ struct rt_rq {
int rt_throttled;
u64 rt_time; /* consumed RT time, goes up in update_curr_rt */
u64 rt_runtime; /* allotted RT time, "slice" from rt_bandwidth, RT sharing/balancing */
+ u64 throttle_count;
/* Nests inside the rq lock: */
raw_spinlock_t rt_runtime_lock;
--
2.25.1
prev parent reply other threads:[~2025-12-02 5:52 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-12-02 5:51 [PATCH 0/2] sched: expose RT throttling info to facilitate priority reversal processing wen.yang
2025-12-02 5:51 ` [PATCH 1/2] sched/debug: add explicit TASK_RTLOCK_WAIT printing wen.yang
2025-12-02 5:51 ` wen.yang [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a732cb65347843e8b5fdbe182363c5438ac0916f.1764648076.git.wen.yang@linux.dev \
--to=wen.yang@linux.dev \
--cc=bsegall@google.com \
--cc=dietmar.eggemann@arm.com \
--cc=juri.lelli@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®