* [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU
@ 2025-11-27 9:14 Huang Shijie
2025-11-27 9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie
2025-11-27 9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
0 siblings, 2 replies; 5+ messages in thread
From: Huang Shijie @ 2025-11-27 9:14 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak,
Huang Shijie
In the newidle balance, the rq->idle_stamp may set to a non-zero value
if it cannot pull any task.
In the wakeup, it will detect the rq->idle_stamp, and updates
the rq->avg_idle, then ends the CPU idle status by setting rq->idle_stamp
to zero.
Besides the wakeup, current code does not end the CPU idle status
when a task is moved to the idle CPU, such as fork/clone, execve,
or other cases.
This patch set tries to resolve it.
v2--> v3:
-- merge patch 3 into patch 2:
move update_rq_avg_idle() to enqueue_task().
v2: https://lkml.org/lkml/2025/11/27/214
v1--> v2:
-- Put update_rq_avg_idle() to activate_task()
-- Add Delay-dequeue task check.
v1: https://lkml.org/lkml/2025/11/24/97
Huang Shijie (2):
sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle
sched: update the rq->avg_idle when a task is moved to an idle CPU
kernel/sched/core.c | 36 ++++++++++++++++++++++++------------
kernel/sched/fair.c | 12 ++++++++----
2 files changed, 32 insertions(+), 16 deletions(-)
--
2.40.1
^ permalink raw reply [flat|nested] 5+ messages in thread* [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle 2025-11-27 9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie @ 2025-11-27 9:14 ` Huang Shijie 2025-11-27 9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie 1 sibling, 0 replies; 5+ messages in thread From: Huang Shijie @ 2025-11-27 9:14 UTC (permalink / raw) To: mingo, peterz, juri.lelli, vincent.guittot Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall, mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak, Huang Shijie Save the idle_stamp at the beginning of sched_balance_newidle(), if it cannot pull any task, set it for rq->idle_stamp. This patch does not change the logic of rq->idle_stamp. Signed-off-by: Huang Shijie <shijie@os.amperecomputing.com> --- kernel/sched/fair.c | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 769d7b7990df..c1a8fa043156 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12862,6 +12862,7 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf) u64 t0, t1, curr_cost = 0; struct sched_domain *sd; int pulled_task = 0; + u64 idle_stamp; update_misfit_status(NULL, this_rq); @@ -12877,7 +12878,9 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf) * for CPU_NEWLY_IDLE, such that we measure the this duration * as idle time. */ - this_rq->idle_stamp = rq_clock(this_rq); + idle_stamp = rq_clock(this_rq); + + this_rq->idle_stamp = 0; /* * Do not pull tasks towards !active CPUs... @@ -12989,10 +12992,11 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf) if (time_after(this_rq->next_balance, next_balance)) this_rq->next_balance = next_balance; - if (pulled_task) - this_rq->idle_stamp = 0; - else + if (!pulled_task) { + /* Set it here on purpose. */ + this_rq->idle_stamp = idle_stamp; nohz_newidle_balance(this_rq); + } rq_repin_lock(this_rq, rf); -- 2.40.1 ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU 2025-11-27 9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie 2025-11-27 9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie @ 2025-11-27 9:14 ` Huang Shijie 2025-11-27 10:12 ` K Prateek Nayak 1 sibling, 1 reply; 5+ messages in thread From: Huang Shijie @ 2025-11-27 9:14 UTC (permalink / raw) To: mingo, peterz, juri.lelli, vincent.guittot Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall, mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak, Huang Shijie In the newidle balance, the rq->idle_stamp may set to a non-zero value if it cannot pull any task. In the wakeup, it will detect the rq->idle_stamp, and updates the rq->avg_idle, then ends the CPU idle status by setting rq->idle_stamp to zero. Besides the wakeup, current code does not end the CPU idle status when a task is moved to the idle CPU, such as fork/clone, execve, or other cases. This patch introduces a helper: update_rq_avg_idle(). And uses it in enqueue_task(), so it will update the rq->avg_idle when a task is moved to an idle CPU at: -- wakeup -- fork/clone -- execve -- idle balance -- delayed dequeue task -- other cases Signed-off-by: Huang Shijie <shijie@os.amperecomputing.com> --- kernel/sched/core.c | 36 ++++++++++++++++++++++++------------ 1 file changed, 24 insertions(+), 12 deletions(-) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 0c4ff93eeb78..8531ef68ce76 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -2078,8 +2078,25 @@ unsigned long get_wchan(struct task_struct *p) return ip; } +static void update_rq_avg_idle(struct rq *rq) +{ + if (rq->idle_stamp) { + u64 delta = rq_clock(rq) - rq->idle_stamp; + u64 max = 2*rq->max_idle_balance_cost; + + update_avg(&rq->avg_idle, delta); + + if (rq->avg_idle > max) + rq->avg_idle = max; + + rq->idle_stamp = 0; + } +} + void enqueue_task(struct rq *rq, struct task_struct *p, int flags) { + int delayed = p->se.sched_delayed; + if (!(flags & ENQUEUE_NOCLOCK)) update_rq_clock(rq); @@ -2100,6 +2117,13 @@ void enqueue_task(struct rq *rq, struct task_struct *p, int flags) if (sched_core_enabled(rq)) sched_core_enqueue(rq, p); + + if (delayed) { + if (entity_eligible(cfs_rq_of(&p->se), &p->se)) + update_rq_avg_idle(rq); + } else { + update_rq_avg_idle(rq); + } } /* @@ -3645,18 +3669,6 @@ ttwu_do_activate(struct rq *rq, struct task_struct *p, int wake_flags, p->sched_class->task_woken(rq, p); rq_repin_lock(rq, rf); } - - if (rq->idle_stamp) { - u64 delta = rq_clock(rq) - rq->idle_stamp; - u64 max = 2*rq->max_idle_balance_cost; - - update_avg(&rq->avg_idle, delta); - - if (rq->avg_idle > max) - rq->avg_idle = max; - - rq->idle_stamp = 0; - } } /* -- 2.40.1 ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU 2025-11-27 9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie @ 2025-11-27 10:12 ` K Prateek Nayak 2025-11-28 6:30 ` Shijie Huang 0 siblings, 1 reply; 5+ messages in thread From: K Prateek Nayak @ 2025-11-27 10:12 UTC (permalink / raw) To: Huang Shijie, mingo, peterz, juri.lelli, vincent.guittot Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall, mgorman, linux-kernel, vschneid, vineethr Hello Huang Shijie, On 11/27/2025 2:44 PM, Huang Shijie wrote: > void enqueue_task(struct rq *rq, struct task_struct *p, int flags) > { > + int delayed = p->se.sched_delayed; > + > if (!(flags & ENQUEUE_NOCLOCK)) > update_rq_clock(rq); > > @@ -2100,6 +2117,13 @@ void enqueue_task(struct rq *rq, struct task_struct *p, int flags) > > if (sched_core_enabled(rq)) > sched_core_enqueue(rq, p); > + > + if (delayed) { > + if (entity_eligible(cfs_rq_of(&p->se), &p->se)) > + update_rq_avg_idle(rq); Question: Why do we want to treat the delayed case like this? If entity is not eligible, we want to consider that it hasn't even gone through a wakeup? Wouldn't this lead to the next wakeup seeing rq->idle_stamp to be non-zero and inaccurately account more idle time? Also if we've done newidle balance and the rq->idle_stamp is set, we cannot have delayed tasks since pick_next_task() would have dequeued all delayed tasks before reaching newidle balance. Just doing a update_rq_avg_idle() unconditionally should be fine. > + } else { > + update_rq_avg_idle(rq); > + } > } > > /* -- Thanks and Regards, Prateek ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU 2025-11-27 10:12 ` K Prateek Nayak @ 2025-11-28 6:30 ` Shijie Huang 0 siblings, 0 replies; 5+ messages in thread From: Shijie Huang @ 2025-11-28 6:30 UTC (permalink / raw) To: K Prateek Nayak, Huang Shijie, mingo, peterz, juri.lelli, vincent.guittot Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall, mgorman, linux-kernel, vschneid, vineethr On 27/11/2025 18:12, K Prateek Nayak wrote: > Also if we've done newidle balance and the rq->idle_stamp is > set, we cannot have delayed tasks since pick_next_task() would > have dequeued all delayed tasks before reaching newidle > balance. Yes, you are right. > Just doing a update_rq_avg_idle() unconditionally should be > fine. okay. Thanks Huang Shijie ^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2025-11-28 6:30 UTC | newest] Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2025-11-27 9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie 2025-11-27 9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie 2025-11-27 9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie 2025-11-27 10:12 ` K Prateek Nayak 2025-11-28 6:30 ` Shijie Huang
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®