mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU
@ 2025-11-27  9:14 Huang Shijie
  2025-11-27  9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie
  2025-11-27  9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
  0 siblings, 2 replies; 5+ messages in thread
From: Huang Shijie @ 2025-11-27  9:14 UTC (permalink / raw)
  To: mingo, peterz, juri.lelli, vincent.guittot
  Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
	mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak,
	Huang Shijie

In the newidle balance, the rq->idle_stamp may set to a non-zero value
if it cannot pull any task.

In the wakeup, it will detect the rq->idle_stamp, and updates
the rq->avg_idle, then ends the CPU idle status by setting rq->idle_stamp
to zero.

Besides the wakeup, current code does not end the CPU idle status
when a task is moved to the idle CPU, such as fork/clone, execve,
or other cases.

This patch set tries to resolve it.

v2--> v3:
  -- merge patch 3 into patch 2:
      move update_rq_avg_idle() to enqueue_task().

   v2: https://lkml.org/lkml/2025/11/27/214   

v1--> v2:
  -- Put update_rq_avg_idle() to activate_task()
  -- Add Delay-dequeue task check.	

   v1: https://lkml.org/lkml/2025/11/24/97


Huang Shijie (2):
  sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle
  sched: update the rq->avg_idle when a task is moved to an idle CPU

 kernel/sched/core.c | 36 ++++++++++++++++++++++++------------
 kernel/sched/fair.c | 12 ++++++++----
 2 files changed, 32 insertions(+), 16 deletions(-)

-- 
2.40.1


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle
  2025-11-27  9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
@ 2025-11-27  9:14 ` Huang Shijie
  2025-11-27  9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
  1 sibling, 0 replies; 5+ messages in thread
From: Huang Shijie @ 2025-11-27  9:14 UTC (permalink / raw)
  To: mingo, peterz, juri.lelli, vincent.guittot
  Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
	mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak,
	Huang Shijie

Save the idle_stamp at the beginning of sched_balance_newidle(),
if it cannot pull any task, set it for rq->idle_stamp.

This patch does not change the logic of rq->idle_stamp.

Signed-off-by: Huang Shijie <shijie@os.amperecomputing.com>
---
 kernel/sched/fair.c | 12 ++++++++----
 1 file changed, 8 insertions(+), 4 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 769d7b7990df..c1a8fa043156 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -12862,6 +12862,7 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf)
 	u64 t0, t1, curr_cost = 0;
 	struct sched_domain *sd;
 	int pulled_task = 0;
+	u64 idle_stamp;
 
 	update_misfit_status(NULL, this_rq);
 
@@ -12877,7 +12878,9 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf)
 	 * for CPU_NEWLY_IDLE, such that we measure the this duration
 	 * as idle time.
 	 */
-	this_rq->idle_stamp = rq_clock(this_rq);
+	idle_stamp = rq_clock(this_rq);
+
+	this_rq->idle_stamp = 0;
 
 	/*
 	 * Do not pull tasks towards !active CPUs...
@@ -12989,10 +12992,11 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf)
 	if (time_after(this_rq->next_balance, next_balance))
 		this_rq->next_balance = next_balance;
 
-	if (pulled_task)
-		this_rq->idle_stamp = 0;
-	else
+	if (!pulled_task) {
+		/* Set it here on purpose. */
+		this_rq->idle_stamp = idle_stamp;
 		nohz_newidle_balance(this_rq);
+	}
 
 	rq_repin_lock(this_rq, rf);
 
-- 
2.40.1


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU
  2025-11-27  9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
  2025-11-27  9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie
@ 2025-11-27  9:14 ` Huang Shijie
  2025-11-27 10:12   ` K Prateek Nayak
  1 sibling, 1 reply; 5+ messages in thread
From: Huang Shijie @ 2025-11-27  9:14 UTC (permalink / raw)
  To: mingo, peterz, juri.lelli, vincent.guittot
  Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
	mgorman, linux-kernel, vschneid, vineethr, kprateek.nayak,
	Huang Shijie

In the newidle balance, the rq->idle_stamp may set to a non-zero value
if it cannot pull any task.

In the wakeup, it will detect the rq->idle_stamp, and updates
the rq->avg_idle, then ends the CPU idle status by setting rq->idle_stamp
to zero.

Besides the wakeup, current code does not end the CPU idle status
when a task is moved to the idle CPU, such as fork/clone, execve,
or other cases.

This patch introduces a helper: update_rq_avg_idle().
And uses it in enqueue_task(), so it will update the rq->avg_idle
when a task is moved to an idle CPU at:
   -- wakeup
   -- fork/clone
   -- execve
   -- idle balance
   -- delayed dequeue task
   -- other cases

Signed-off-by: Huang Shijie <shijie@os.amperecomputing.com>
---
 kernel/sched/core.c | 36 ++++++++++++++++++++++++------------
 1 file changed, 24 insertions(+), 12 deletions(-)

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0c4ff93eeb78..8531ef68ce76 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -2078,8 +2078,25 @@ unsigned long get_wchan(struct task_struct *p)
 	return ip;
 }
 
+static void update_rq_avg_idle(struct rq *rq)
+{
+	if (rq->idle_stamp) {
+		u64 delta = rq_clock(rq) - rq->idle_stamp;
+		u64 max = 2*rq->max_idle_balance_cost;
+
+		update_avg(&rq->avg_idle, delta);
+
+		if (rq->avg_idle > max)
+			rq->avg_idle = max;
+
+		rq->idle_stamp = 0;
+	}
+}
+
 void enqueue_task(struct rq *rq, struct task_struct *p, int flags)
 {
+	int delayed = p->se.sched_delayed;
+
 	if (!(flags & ENQUEUE_NOCLOCK))
 		update_rq_clock(rq);
 
@@ -2100,6 +2117,13 @@ void enqueue_task(struct rq *rq, struct task_struct *p, int flags)
 
 	if (sched_core_enabled(rq))
 		sched_core_enqueue(rq, p);
+
+	if (delayed) {
+		if (entity_eligible(cfs_rq_of(&p->se), &p->se))
+			update_rq_avg_idle(rq);
+	} else {
+		update_rq_avg_idle(rq);
+	}
 }
 
 /*
@@ -3645,18 +3669,6 @@ ttwu_do_activate(struct rq *rq, struct task_struct *p, int wake_flags,
 		p->sched_class->task_woken(rq, p);
 		rq_repin_lock(rq, rf);
 	}
-
-	if (rq->idle_stamp) {
-		u64 delta = rq_clock(rq) - rq->idle_stamp;
-		u64 max = 2*rq->max_idle_balance_cost;
-
-		update_avg(&rq->avg_idle, delta);
-
-		if (rq->avg_idle > max)
-			rq->avg_idle = max;
-
-		rq->idle_stamp = 0;
-	}
 }
 
 /*
-- 
2.40.1


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU
  2025-11-27  9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
@ 2025-11-27 10:12   ` K Prateek Nayak
  2025-11-28  6:30     ` Shijie Huang
  0 siblings, 1 reply; 5+ messages in thread
From: K Prateek Nayak @ 2025-11-27 10:12 UTC (permalink / raw)
  To: Huang Shijie, mingo, peterz, juri.lelli, vincent.guittot
  Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
	mgorman, linux-kernel, vschneid, vineethr

Hello Huang Shijie,

On 11/27/2025 2:44 PM, Huang Shijie wrote:
>  void enqueue_task(struct rq *rq, struct task_struct *p, int flags)
>  {
> +	int delayed = p->se.sched_delayed;
> +
>  	if (!(flags & ENQUEUE_NOCLOCK))
>  		update_rq_clock(rq);
>  
> @@ -2100,6 +2117,13 @@ void enqueue_task(struct rq *rq, struct task_struct *p, int flags)
>  
>  	if (sched_core_enabled(rq))
>  		sched_core_enqueue(rq, p);
> +
> +	if (delayed) {
> +		if (entity_eligible(cfs_rq_of(&p->se), &p->se))
> +			update_rq_avg_idle(rq);

Question: Why do we want to treat the delayed case like this?

If entity is not eligible, we want to consider that it hasn't
even gone through a wakeup? Wouldn't this lead to the next
wakeup seeing rq->idle_stamp to be non-zero and inaccurately
account more idle time?

Also if we've done newidle balance and the rq->idle_stamp is
set, we cannot have delayed tasks since pick_next_task() would
have dequeued all delayed tasks before reaching newidle
balance.

Just doing a update_rq_avg_idle() unconditionally should be
fine.

> +	} else {
> +		update_rq_avg_idle(rq);
> +	}
>  }
>  
>  /*
-- 
Thanks and Regards,
Prateek


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU
  2025-11-27 10:12   ` K Prateek Nayak
@ 2025-11-28  6:30     ` Shijie Huang
  0 siblings, 0 replies; 5+ messages in thread
From: Shijie Huang @ 2025-11-28  6:30 UTC (permalink / raw)
  To: K Prateek Nayak, Huang Shijie, mingo, peterz, juri.lelli,
	vincent.guittot
  Cc: patches, cl, Shubhang, dietmar.eggemann, rostedt, bsegall,
	mgorman, linux-kernel, vschneid, vineethr


On 27/11/2025 18:12, K Prateek Nayak wrote:
> Also if we've done newidle balance and the rq->idle_stamp is
> set, we cannot have delayed tasks since pick_next_task() would
> have dequeued all delayed tasks before reaching newidle
> balance.
Yes, you are right.
> Just doing a update_rq_avg_idle() unconditionally should be
> fine.

okay.


Thanks

Huang Shijie


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2025-11-28  6:30 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2025-11-27  9:14 [PATCH v3 0/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
2025-11-27  9:14 ` [PATCH v3 1/2] sched/fair: set rq->idle_stamp at the end of the sched_balance_newidle Huang Shijie
2025-11-27  9:14 ` [PATCH v3 2/2] sched: update the rq->avg_idle when a task is moved to an idle CPU Huang Shijie
2025-11-27 10:12   ` K Prateek Nayak
2025-11-28  6:30     ` Shijie Huang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®