mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Hui Su <sh_def@163.com>
To: peterz@infradead.org, soolaugust@gmail.com, arighi@nvidia.com
Cc: mingo@redhat.com, kprateek.nayak@amd.com, juri.lelli@redhat.com,
	vincent.guittot@linaro.org, dietmar.eggemann@arm.com,
	rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de,
	vschneid@redhat.com, connoro@google.com, jstultz@google.com,
	linux-kernel@vger.kernel.org, sched-ext@lists.linux.dev
Subject: [PATCH v5 4/4] sched/core: Fix donor slice accounting under proxy execution
Date: Sun, 13 Sep 2026 15:47:22 +0900	[thread overview]
Message-ID: <20260913064722.1534766-5-sh_def@163.com> (raw)
In-Reply-To: <20260913064722.1534766-1-sh_def@163.com>

Core scheduling uses __entity_slice_used() to decide whether the current
scheduling context has consumed enough of its slice to let a force-idled
SMT sibling run.

The check is correctly made against rq->donor, since the slice belongs to
the scheduling context. With proxy execution, however, task runtime is
accounted to rq->curr. The donor's sum_exec_runtime therefore does not
advance while another task executes on its behalf, causing

	se->sum_exec_runtime - se->prev_sum_exec_runtime

to remain near zero and preventing the force-idle reschedule from
triggering.

Using rq->curr is not correct either, as that would compare the execution
task's runtime against its own slice rather than the donor's slice.

Track the donor's task-clock timestamp when it is selected and measure the
elapsed service using exec_start. update_se() advances the donor's
exec_start from rq_clock_task() even under proxy execution, while the
accumulated task runtime itself is charged to rq->curr.

Keep the selection baseline in struct rq: there is only one active donor
per runqueue. Storing it in every sched_entity needlessly increases the
size of the hot sched_entity structure, especially on 32-bit builds.

Proxy execution can switch rq->curr from one execution owner to another
without changing rq->donor. The scheduler performs a synthetic
put_prev_task()/set_next_task() pair for the unchanged donor in that case.
Preserve the baseline for that reselect, or service accumulated before the
owner handoff would be lost. Refresh it when selecting a different donor
or through an explicit set-next reactivation.

Keeping the comparison in the task-clock domain also avoids depending on
the entity's weight. A vruntime delta accumulated across different weights
cannot reliably be compared against a slice converted using only the
current weight.

Only snapshot task entities, as task_tick_core() performs the consumed
slice check on the donor task.

In non-proxy testing, the task-clock predicate matched the existing
sum_exec_runtime predicate across HZ=100/250/1000 and nice -10/0/+10.
Under proxy execution, the donor's sum_exec_runtime delta remained zero
while the task-clock delta advanced and triggered the force-idle
reschedule. A focused same-donor owner-handoff test split the donor's
service across two execution owners. Each owner's individual service stayed
below half of the donor slice, while their accumulated service exceeded the
threshold and triggered the force-idle reschedule. Both phases kept the
same core_sched_start baseline.

Fixes: aa4f74dfd42b ("sched: Fix runtime accounting w/ split exec & sched contexts")
Suggested-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Hui Su <sh_def@163.com>
---
 kernel/sched/fair.c  | 27 ++++++++++++++++++---------
 kernel/sched/sched.h |  2 ++
 2 files changed, 20 insertions(+), 9 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index ed08b287016a..9a7be87f1991 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -6469,7 +6469,7 @@ dequeue_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags)
 }
 
 static void
-set_next_entity(struct cfs_rq *cfs_rq, struct sched_entity *se)
+set_next_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, bool reset_core_slice)
 {
 	/* 'current' is not kept within the tree. */
 	if (se->on_rq) {
@@ -6502,6 +6502,10 @@ set_next_entity(struct cfs_rq *cfs_rq, struct sched_entity *se)
 	}
 
 	se->prev_sum_exec_runtime = se->sum_exec_runtime;
+#ifdef CONFIG_SCHED_CORE
+	if (reset_core_slice && entity_is_task(se))
+		rq_of(cfs_rq)->core_sched_start = se->exec_start;
+#endif
 }
 
 static bool __dequeue_task(struct rq *rq, struct task_struct *p, int flags);
@@ -14786,16 +14790,16 @@ static void rq_offline_fair(struct rq *rq)
 
 #ifdef CONFIG_SCHED_CORE
 static inline bool
-__entity_slice_used(struct sched_entity *se, int min_nr_tasks)
+__entity_slice_used(struct rq *rq, struct sched_entity *se, int min_nr_tasks)
 {
-	u64 rtime = se->sum_exec_runtime - se->prev_sum_exec_runtime;
-	u64 slice = se->slice;
+	/* exec_start advances with donor service under proxy execution. */
+	u64 rtime = se->exec_start - rq->core_sched_start;
 
-	return (rtime * min_nr_tasks > slice);
+	return (rtime * min_nr_tasks > se->slice);
 }
 
 #define MIN_NR_TASKS_DURING_FORCEIDLE	2
-static inline void task_tick_core(struct rq *rq, struct task_struct *curr)
+static inline void task_tick_core(struct rq *rq, struct task_struct *donor)
 {
 	if (!sched_core_enabled(rq))
 		return;
@@ -14815,7 +14819,7 @@ static inline void task_tick_core(struct rq *rq, struct task_struct *curr)
 	 * if we need to give up the CPU.
 	 */
 	if (rq->core->core_forceidle_count && rq->cfs.h_nr_queued == 1 &&
-	    __entity_slice_used(&curr->se, MIN_NR_TASKS_DURING_FORCEIDLE))
+	    __entity_slice_used(rq, &donor->se, MIN_NR_TASKS_DURING_FORCEIDLE))
 		resched_curr(rq);
 }
 
@@ -15049,7 +15053,7 @@ static int task_is_throttled_fair(struct task_struct *p, int cpu)
 	return throttled_hierarchy(cfs_rq);
 }
 #else /* !CONFIG_SCHED_CORE: */
-static inline void task_tick_core(struct rq *rq, struct task_struct *curr) {}
+static inline void task_tick_core(struct rq *rq, struct task_struct *donor) {}
 #endif /* !CONFIG_SCHED_CORE */
 
 /*
@@ -15257,6 +15261,11 @@ static void set_next_task_fair(struct rq *rq, struct task_struct *p, bool first)
 {
 	struct sched_entity *se = &p->se;
 	bool throttled = false;
+	/*
+	 * Reset for a new donor or reactivation, but preserve a same-donor
+	 * proxy reselect so service before an owner handoff is retained.
+	 */
+	bool reset_core_slice = !first || rq->donor != p;
 	struct cfs_rq *cfs_rq = &rq->cfs;
 	unsigned long weight = NICE_0_LOAD;
 	bool on_rq = se->on_rq;
@@ -15271,7 +15280,7 @@ static void set_next_task_fair(struct rq *rq, struct task_struct *p, bool first)
 
 		if (!IS_ENABLED(CONFIG_FAIR_GROUP_SCHED) ||
 		    !first || !cfs_rq->h_curr)
-			set_next_entity(cfs_rq, se);
+			set_next_entity(cfs_rq, se, reset_core_slice);
 
 		/* ensure bandwidth has been allocated on our new cfs_rq */
 		throttled |= account_cfs_rq_runtime(cfs_rq, 0);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 6a8deddc725b..f2f1e3642831 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1382,6 +1382,8 @@ struct rq {
 	unsigned int		core_forceidle_seq;
 	unsigned int		core_forceidle_occupation;
 	u64			core_forceidle_start;
+	/* Task-clock baseline for the current core-scheduling donor slice. */
+	u64			core_sched_start;
 	unsigned int		core_pick_in_flight;
 #endif /* CONFIG_SCHED_CORE */
 
-- 
2.55.0


      parent reply	other threads:[~2026-09-13  6:48 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-13  6:47 [PATCH v5 0/4] sched: Handle split scheduling and execution contexts in task ticks Hui Su
2026-09-13  6:47 ` [PATCH v5 1/4] sched: Dispatch task ticks for donor and execution classes Hui Su
2026-09-13 12:08   ` Kayra Cizmeci
2026-09-13 13:07     ` Hui Su
2026-09-13  6:47 ` [PATCH v5 2/4] sched/numa: Drive NUMA task tick from execution context Hui Su
2026-09-13  6:47 ` [PATCH v5 3/4] sched/cache: Drive cache " Hui Su
2026-09-13  6:47 ` Hui Su [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260913064722.1534766-5-sh_def@163.com \
    --to=sh_def@163.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=connoro@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=soolaugust@gmail.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®