mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Hui Su <sh_def@163.com>
To: peterz@infradead.org, mingo@redhat.com,
	tim.c.chen@linux.intel.com, yu.c.chen@intel.com,
	kprateek.nayak@amd.com
Cc: juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	connoro@google.com, jstultz@google.com, arighi@nvidia.com,
	tj@kernel.org, void@manifault.com, changwoo@igalia.com,
	linux-kernel@vger.kernel.org, sched-ext@lists.linux.dev
Subject: [PATCH v4 4/5] sched/rt: Fix RT watchdog accounting for proxy execution
Date: Wed,  9 Sep 2026 18:29:00 +0900	[thread overview]
Message-ID: <20260909092901.2989564-5-sh_def@163.com> (raw)
In-Reply-To: <20260909092901.2989564-1-sh_def@163.com>

Proxy execution separates the scheduling context in rq->donor from the
execution context in rq->curr.

When rq->donor belongs to the RT class, task_tick_rt() keeps RT
scheduling-class state, load tracking, and RR time-slice management
associated with rq->donor.

The RT watchdog is different. It looks up RLIMIT_RTTIME through its task
argument and updates that task's rt.timeout and posix_cputimers state.
update_curr_rt(), however, accounts elapsed task runtime to rq->curr, and
run_posix_cpu_timers() checks the execution task after the scheduler tick.

Keep the RT scheduling work on rq->donor, but run watchdog() for rq->curr.
The watchdog state then follows the task whose execution runtime advances.
The callback can also be dispatched for an RT execution context whose donor
belongs to another scheduling class. In that case, run only the watchdog
for rq->curr, without applying RT donor accounting or RR time-slice
management.

Reset rt.timeout when a task blocks under proxy execution. A mutex-blocked
task can remain on the runqueue and bypass ENQUEUE_WAKEUP, which normally
resets the RT watchdog interval. Also reset it when a task without an RT
policy is subsequently selected with neither the execution nor scheduling
context in the RT class. Preserve the timeout across ordinary scheduler
preemption and when an RT-policy task is temporarily PI-boosted into the
deadline class.

The donor/curr attribution was checked with both the in-kernel mutex
reproducer and a userspace owner in a two-node QEMU guest. FIFO and RR
donors produced donor-targeted watchdog traces on the baseline kernel and
execution-owner-targeted traces with this change; both proxy runs
completed successfully. A DL donor with an RT execution owner was also
exercised; the execution owner's timeout advanced and SIGXCPU was delivered
to it, while an RR execution owner's rt.time_slice remained unchanged in
execution-only callbacks.

A targeted rtmutex test set rt.timeout to 123 on a SCHED_FIFO owner. A DL
waiter then PI-boosted it into the deadline class. Without the policy
guard, scheduling the boosted owner reset the timeout to zero. With the
guard, it remained 123 during the boost and after deboosting. A FAIR owner
leaving an RT proxy interval still reset its timeout from 123 to zero.

Fixes: 7de9d4f94638 ("sched: Start blocked_on chain processing in find_proxy_task()")
Signed-off-by: Hui Su <sh_def@163.com>
---
 kernel/sched/core.c | 18 ++++++++++++++++++
 kernel/sched/rt.c   | 12 ++++++++++--
 2 files changed, 28 insertions(+), 2 deletions(-)

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 05e599665fdd..d8a785bec639 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -6768,6 +6768,14 @@ static bool try_to_block_task(struct rq *rq, struct task_struct *p,
 		return false;
 	}
 
+	/*
+	 * Proxy execution can keep a mutex-blocked task on the runqueue, so it
+	 * may not pass through ENQUEUE_WAKEUP, which normally resets the RT
+	 * watchdog interval.
+	 */
+	if (sched_proxy_exec() && p->rt.timeout)
+		p->rt.timeout = 0;
+
 	p->is_blocked = 1;
 
 	/*
@@ -7243,6 +7251,16 @@ static void __sched notrace __schedule(int sched_mode)
 		rq_set_donor(rq, next);
 	}
 
+	/*
+	 * End a previous RT proxy watchdog interval once neither context is
+	 * in the RT class. Preserve the interval for an RT-policy task that is
+	 * temporarily PI-boosted into the DL class.
+	 */
+	if (sched_proxy_exec() && !task_has_rt_policy(next) &&
+	    !rt_prio(next->prio) &&
+	    !rt_prio(rq->donor->prio) && next->rt.timeout)
+		next->rt.timeout = 0;
+
 picked:
 	clear_tsk_need_resched(prev);
 	clear_preempt_need_resched();
diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
index dd058a6ca06b..0fdb8edb7528 100644
--- a/kernel/sched/rt.c
+++ b/kernel/sched/rt.c
@@ -2542,15 +2542,23 @@ static void task_tick_rt(struct rq *rq, int queued)
 	struct task_struct *p = rq->donor;
 	struct sched_rt_entity *rt_se;
 
-	if (p->sched_class != &rt_sched_class)
+	if (p->sched_class != &rt_sched_class) {
+		/*
+		 * The RT callback can also be dispatched for an RT execution
+		 * context whose scheduling context belongs to another class.
+		 * Keep the watchdog tied to the task whose runtime is advancing.
+		 */
+		if (rq->curr->sched_class == &rt_sched_class)
+			watchdog(rq, rq->curr);
 		return;
+	}
 
 	rt_se = &p->rt;
 
 	update_curr_rt(rq);
 	update_rt_rq_load_avg(rq_clock_pelt(rq), rq, 1);
 
-	watchdog(rq, p);
+	watchdog(rq, rq->curr);
 
 	/*
 	 * RR tasks need a special form of time-slice management.
-- 
2.55.0


  parent reply	other threads:[~2026-09-09  9:32 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-09  9:28 [PATCH v4 0/5] sched: Handle split scheduling and execution contexts in task ticks Hui Su
2026-09-09  9:28 ` [PATCH v4 1/5] sched: Dispatch task ticks for donor and execution classes Hui Su
2026-09-09 10:42   ` Hui Su
2026-09-09 17:43     ` Andrea Righi
2026-09-10 10:49       ` Hui Su
2026-09-09  9:28 ` [PATCH v4 2/5] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-09  9:28 ` [PATCH v4 3/5] sched/cache: Drive cache " Hui Su
2026-09-09 11:03   ` Peter Zijlstra
2026-09-10 10:53     ` Hui Su
2026-09-09  9:29 ` Hui Su [this message]
2026-09-09 11:04   ` [PATCH v4 4/5] sched/rt: Fix RT watchdog accounting for proxy execution Peter Zijlstra
2026-09-10 10:54     ` Hui Su
2026-09-12 17:30     ` Hui Su
2026-09-09  9:29 ` [PATCH v4 5/5] sched/core: Fix donor slice accounting under " Hui Su
2026-09-10 10:55   ` Hui Su

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260909092901.2989564-5-sh_def@163.com \
    --to=sh_def@163.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=connoro@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tim.c.chen@linux.intel.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®