mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Hui Su <sh_def@163.com>
To: peterz@infradead.org, mingo@redhat.com,
	tim.c.chen@linux.intel.com, yu.c.chen@intel.com,
	kprateek.nayak@amd.com
Cc: juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	connoro@google.com, jstultz@google.com, arighi@nvidia.com,
	tj@kernel.org, void@manifault.com, changwoo@igalia.com,
	linux-kernel@vger.kernel.org, sched-ext@lists.linux.dev
Subject: [PATCH v4 5/5] sched/core: Fix donor slice accounting under proxy execution
Date: Wed,  9 Sep 2026 18:29:01 +0900	[thread overview]
Message-ID: <20260909092901.2989564-6-sh_def@163.com> (raw)
In-Reply-To: <20260909092901.2989564-1-sh_def@163.com>

Core scheduling uses __entity_slice_used() to decide whether the current
scheduling context has consumed enough of its slice to let a force-idled
SMT sibling run.

The check is correctly made against rq->donor, since the slice belongs
to the scheduling context. With proxy execution, however, task runtime
is accounted to rq->curr. The donor's sum_exec_runtime therefore does
not advance while another task executes on its behalf, causing

        se->sum_exec_runtime - se->prev_sum_exec_runtime

to remain near zero and preventing the force-idle reschedule from
triggering.

Using rq->curr is not correct either, as that would compare the execution
task's runtime against its own slice rather than the donor's slice.

Track the donor's task-clock timestamp when it is selected and measure
the elapsed service using se->exec_start. update_se() advances the donor's
exec_start from rq_clock_task() even under proxy execution, while the
accumulated task runtime itself is charged to rq->curr.

Keeping the comparison in the task-clock domain also avoids depending on
the entity's weight. A vruntime delta accumulated across different
weights cannot reliably be compared against a slice converted using only
the current weight.

Only snapshot task entities, as task_tick_core() performs the consumed
slice check on the donor task.

In non-proxy testing, the task-clock predicate matched the existing
sum_exec_runtime predicate across HZ=100/250/1000 and nice -10/0/+10.
Under proxy execution, the donor's sum_exec_runtime delta remained zero
while the task-clock delta advanced and triggered the force-idle
reschedule.

Fixes: aa4f74dfd42b ("sched: Fix runtime accounting w/ split exec & sched contexts")
Suggested-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Hui Su <sh_def@163.com>
---
 include/linux/sched.h |  3 +++
 kernel/sched/core.c   |  3 +++
 kernel/sched/fair.c   | 15 +++++++++------
 3 files changed, 15 insertions(+), 6 deletions(-)

diff --git a/include/linux/sched.h b/include/linux/sched.h
index 8b3d47a325cc..c32d9931129f 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -590,6 +590,9 @@ struct sched_entity {
 	u64				sum_exec_runtime;
 	u64				prev_sum_exec_runtime;
 	u64				vruntime;
+#ifdef CONFIG_SCHED_CORE
+	u64				core_sched_start;
+#endif
 	/* Approximated virtual lag: */
 	s64				vlag;
 	/* 'Protected' deadline, to give out minimum quantums: */
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index d8a785bec639..87178ec08392 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -4602,6 +4602,9 @@ static void __sched_fork(u64 clone_flags, struct task_struct *p)
 	p->se.prev_sum_exec_runtime	= 0;
 	p->se.nr_migrations		= 0;
 	p->se.vruntime			= 0;
+#ifdef CONFIG_SCHED_CORE
+	p->se.core_sched_start		= 0;
+#endif
 	p->se.vlag			= 0;
 	p->se.rel_deadline		= 0;
 	INIT_LIST_HEAD(&p->se.group_node);
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 43d558289856..0528dc4846f3 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -6502,6 +6502,10 @@ set_next_entity(struct cfs_rq *cfs_rq, struct sched_entity *se)
 	}
 
 	se->prev_sum_exec_runtime = se->sum_exec_runtime;
+#ifdef CONFIG_SCHED_CORE
+	if (entity_is_task(se))
+		se->core_sched_start = se->exec_start;
+#endif
 }
 
 static bool __dequeue_task(struct rq *rq, struct task_struct *p, int flags);
@@ -14788,14 +14792,13 @@ static void rq_offline_fair(struct rq *rq)
 static inline bool
 __entity_slice_used(struct sched_entity *se, int min_nr_tasks)
 {
-	u64 rtime = se->sum_exec_runtime - se->prev_sum_exec_runtime;
-	u64 slice = se->slice;
+	u64 rtime = se->exec_start - se->core_sched_start;
 
-	return (rtime * min_nr_tasks > slice);
+	return (rtime * min_nr_tasks > se->slice);
 }
 
 #define MIN_NR_TASKS_DURING_FORCEIDLE	2
-static inline void task_tick_core(struct rq *rq, struct task_struct *curr)
+static inline void task_tick_core(struct rq *rq, struct task_struct *donor)
 {
 	if (!sched_core_enabled(rq))
 		return;
@@ -14815,7 +14818,7 @@ static inline void task_tick_core(struct rq *rq, struct task_struct *curr)
 	 * if we need to give up the CPU.
 	 */
 	if (rq->core->core_forceidle_count && rq->cfs.h_nr_queued == 1 &&
-	    __entity_slice_used(&curr->se, MIN_NR_TASKS_DURING_FORCEIDLE))
+	    __entity_slice_used(&donor->se, MIN_NR_TASKS_DURING_FORCEIDLE))
 		resched_curr(rq);
 }
 
@@ -15049,7 +15052,7 @@ static int task_is_throttled_fair(struct task_struct *p, int cpu)
 	return throttled_hierarchy(cfs_rq);
 }
 #else /* !CONFIG_SCHED_CORE: */
-static inline void task_tick_core(struct rq *rq, struct task_struct *curr) {}
+static inline void task_tick_core(struct rq *rq, struct task_struct *donor) {}
 #endif /* !CONFIG_SCHED_CORE */
 
 /*
-- 
2.55.0


  parent reply	other threads:[~2026-09-09  9:31 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-09  9:28 [PATCH v4 0/5] sched: Handle split scheduling and execution contexts in task ticks Hui Su
2026-09-09  9:28 ` [PATCH v4 1/5] sched: Dispatch task ticks for donor and execution classes Hui Su
2026-09-09 10:42   ` Hui Su
2026-09-09 17:43     ` Andrea Righi
2026-09-10 10:49       ` Hui Su
2026-09-09  9:28 ` [PATCH v4 2/5] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-09  9:28 ` [PATCH v4 3/5] sched/cache: Drive cache " Hui Su
2026-09-09 11:03   ` Peter Zijlstra
2026-09-10 10:53     ` Hui Su
2026-09-09  9:29 ` [PATCH v4 4/5] sched/rt: Fix RT watchdog accounting for proxy execution Hui Su
2026-09-09 11:04   ` Peter Zijlstra
2026-09-10 10:54     ` Hui Su
2026-09-12 17:30     ` Hui Su
2026-09-09  9:29 ` Hui Su [this message]
2026-09-10 10:55   ` [PATCH v4 5/5] sched/core: Fix donor slice accounting under " Hui Su

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260909092901.2989564-6-sh_def@163.com \
    --to=sh_def@163.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=connoro@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tim.c.chen@linux.intel.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®