mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Chen Yu <yu.c.chen@intel.com>
To: Tim Chen <tim.c.chen@linux.intel.com>
Cc: Chen Yu <chen.yu@linux.dev>,
	Zhan Xusheng <zhanxusheng1024@gmail.com>, <peterz@infradead.org>,
	<mingo@redhat.com>, <juri.lelli@redhat.com>,
	<vincent.guittot@linaro.org>, <dietmar.eggemann@arm.com>,
	<rostedt@goodmis.org>, <bsegall@google.com>, <mgorman@suse.de>,
	<vschneid@redhat.com>, <kprateek.nayak@amd.com>,
	<linux-kernel@vger.kernel.org>, <zhanxusheng@xiaomi.com>
Subject: Re: sched/fair: which tasks should nr_pref_llc_running be compared against?
Date: Wed, 9 Sep 2026 19:40:43 +0800	[thread overview]
Message-ID: <aqFFu1Xo52cQV3iy@fengwei-dev> (raw)
In-Reply-To: <2b0a35122ee615c6fa51076e5d79330e633755ac.camel@linux.intel.com>

On Fri, Sep 04, 2026 at 01:53:10PM -0700, Tim Chen wrote:
> Good idea - the four sites really are one operation ("if the task is
> queued on its preferred LLC and runnable, move the counter"), and
> folding the two conditions into one place is what keeps them from
> drifting apart later. I've adopted it in v3; account_llc_delayed() and
> account_llc_requeue_delayed() are gone.
> 
> I split it slightly differently: a membership predicate
> 
>         static bool task_pref_llc_runnable(struct task_struct *p)
>         {
>                 return p->pref_llc_queued && !p->se.sched_delayed;
>         }
> 
> with pref_llc_running_inc()/pref_llc_running_dec() wrappers over it, so
> the call sites read as inc/dec rather than passing a +1/-1 delta.
> 
> Two things to note:
> 
> 1) I kept the call site comments of pref_llc_running_inc/dec().
>    The helper name says *what* happens,
>    but not *why* it is safe across the delay-dequeue transition - that
>    set_delayed() already did the decrement, so account_llc_dequeue()
>    must skip it, and that clearing pref_llc_queued there is what
>    neutralizes the following clear_delayed(). 
> 
> 2) The pref_llc_running_dec() placement in set_delayed()
>    is subtle. It has to be before se->sched_delayed = 1 or
>    decrement would not happen. That deserves a comment so no
>    one would move sched_delayed = 1 before the decrement.
>
> 
> alb_break_llc() decides whether to break LLC preference during active
> load balance. It does so by testing that every runnable fair task on the
> source rq prefers its LLC:
> 
> 	env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
> 
> But the two counters cover different sets. nr_pref_llc_running is updated
> in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
> so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
> clear_delayed() and drops delay-dequeued tasks.
> 
> So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
> in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
> alb_break_llc() returns false, and active balance is free to pull a task
> off its preferred LLC. Active balance only moves runnable tasks, and this
> is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
> skips the per-task test in can_migrate_task(). The runnable set is the one
> we want.
> 
> Fix it on the counter side. A task should be counted in
> nr_pref_llc_running exactly while it is both queued on its preferred LLC
> (pref_llc_queued) and runnable (!sched_delayed). Define that membership
> once in task_pref_llc_runnable(), and adjust the counter only through
> pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
> change either input: account_llc_enqueue(), account_llc_dequeue(),
> set_delayed() and clear_delayed(). Gating every update on the same
> predicate keeps the delay, wake and dequeue paths from double-counting
> or underflowing; see the comments at those sites for the ordering.
> 
> nr_llc_running and sd->llc_counts are not touched and stay on queued
> semantics.
> 
> Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
> Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
> Suggested-by: Chen Yu <yu.c.chen@intel.com>
> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
> ---
> Based on v7.3-rc1.
> 
> Changes in v3:
> - Route every nr_pref_llc_running adjustment through a single membership
>   predicate task_pref_llc_runnable(), with pref_llc_running_inc()/
>   pref_llc_running_dec() wrappers, instead of four open-coded sites
>   (Chen Yu). Keep the per-site comments that explain the delay-dequeue
>   interaction, and note that set_delayed() must adjust the counter
>   before setting se->sched_delayed.
> 
> Changes in v2:
> - Prevent a delay-dequeued task from being counted as running in the
>   enqueue path (Chen Yu).
> 
>  kernel/sched/fair.c | 52 ++++++++++++++++++++++++++++++++++++++++++++++++++--
>  1 file changed, 50 insertions(+), 2 deletions(-)
> 
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8dff37059faf..72aae7a50b8b 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1538,6 +1538,28 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
>  			(scale * per_cpu(sd_llc_size, cpu)));
>  }
>  
> +/*
> + * A task counts in nr_pref_llc_running while it is queued on its preferred
> + * LLC (pref_llc_queued) and runnable (!sched_delayed), keeping the counter in
> + * the runnable domain so alb_break_llc() can compare it with h_nr_runnable.
> + */
> +static bool task_pref_llc_runnable(struct task_struct *p)
> +{
> +	return p->pref_llc_queued && !p->se.sched_delayed;
> +}
> +
> +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p)
> +{
> +	if (task_pref_llc_runnable(p))
> +		rq->nr_pref_llc_running++;
> +}
> +
> +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p)
> +{
> +	if (task_pref_llc_runnable(p))
> +		rq->nr_pref_llc_running--;
> +}
> +
>  static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
>  {
>  	int pref_llc, pref_llc_queued;
> @@ -1549,7 +1571,6 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
>  
>  	pref_llc_queued = (pref_llc == task_llc(p));
>  	rq->nr_llc_running++;
> -	rq->nr_pref_llc_running += pref_llc_queued;
>  
>  	/*
>  	 * Record whether p is enqueued on its preferred
> @@ -1567,6 +1588,9 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
>  	 */
>  	p->pref_llc_queued = pref_llc_queued;
>  
> +	/* Skipped while delayed; clear_delayed() adds it back on wake. */
> +	pref_llc_running_inc(rq, p);
> +
>  	sd = rcu_dereference_all(rq->sd);
>  	if (sd && (unsigned int)pref_llc < sd->llc_max)
>  		sd->llc_counts[pref_llc]++;
> @@ -1583,7 +1607,12 @@ static void account_llc_dequeue(struct rq *rq, struct task_struct *p)
>  
>  	rq->nr_llc_running--;
>  	if (p->pref_llc_queued) {
> -		rq->nr_pref_llc_running--;
> +		/*
> +		 * Skipped if still delayed (set_delayed() already removed it);
> +		 * clearing pref_llc_queued below also stops clear_delayed()
> +		 * from re-adding it.
> +		 */
> +		pref_llc_running_dec(rq, p);
>  		/*
>  		 * Update the status in case
>  		 * other logic might query
> @@ -2008,6 +2037,10 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p) {}
>  
>  static void account_llc_dequeue(struct rq *rq, struct task_struct *p) {}
>  
> +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p) {}
> +
> +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p) {}
> +
>  #endif /* CONFIG_SCHED_CACHE */
>  
>  /*
> @@ -6382,6 +6415,14 @@ static __always_inline void return_cfs_rq_runtime(struct cfs_rq *cfs_rq);
>  
>  static void set_delayed(struct sched_entity *se)
>  {
> +	/*
> +	 * Drop a task leaving the runnable set. Must run before sched_delayed
> +	 * is set, or task_pref_llc_runnable() would already exclude it;
> +	 * clear_delayed() mirrors this after clearing the flag.
> +	 */
> +	if (entity_is_task(se))
> +		pref_llc_running_dec(rq_of(cfs_rq_of(se)), task_of(se));
> +
>  	se->sched_delayed = 1;
>  
>  	/*
> @@ -6412,6 +6453,13 @@ static void clear_delayed(struct sched_entity *se)
>  	if (!entity_is_task(se))
>  		return;
>  
> +	/*
> +	 * Re-add on wake, after sched_delayed is cleared. On a final delayed
> +	 * dequeue account_llc_dequeue() already cleared pref_llc_queued, so
> +	 * this does nothing.
> +	 */
> +	pref_llc_running_inc(rq_of(cfs_rq_of(se)), task_of(se));
> +
>  	for_each_sched_entity(se) {
>  		struct cfs_rq *cfs_rq = cfs_rq_of(se);
>  
> -- 
> 2.32.0
> 
> 
>

Yes, I think this version looks good now. While looking back at Xusheng's proposal,
I noticed there is another option:
if (env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_queued) {
	...
}
May I know why we did not choose this approach, is it because of the following
scenario?
Suppose there are 3 queued tasks: p1 and p2 prefer the src_rq, while p3 is a delayed
task that also prefers src_rq. In the current implementation, nr_pref_llc_running is 3
and h_nr_runnable is 2, so alb_break_llc() might return false. As a result, active load
balance would be triggered, and p1 or p2 might be migrated away, which is undesirable.
However, would this still be a problem after Lu Wang's active load balance guard patch
has been applied?
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@gmail.com/

thanks,
Chenyu

  reply	other threads:[~2026-09-09 11:54 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-27 13:50 Zhan Xusheng
2026-08-27 20:57 ` Tim Chen
2026-08-28  2:20   ` [PATCH] sched/cache: Keep nr_pref_llc_running in the runnable domain Zhan Xusheng
2026-08-28 17:08     ` Tim Chen
2026-08-30  8:17   ` sched/fair: which tasks should nr_pref_llc_running be compared against? Chen Yu
2026-09-01 20:42     ` Tim Chen
2026-09-03 15:21       ` Chen Yu
2026-09-04 20:53         ` Tim Chen
2026-09-09 11:40           ` Chen Yu [this message]
2026-09-09 17:47             ` Tim Chen
2026-09-10 10:47               ` Chen Yu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqFFu1Xo52cQV3iy@fengwei-dev \
    --to=yu.c.chen@intel.com \
    --cc=bsegall@google.com \
    --cc=chen.yu@linux.dev \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tim.c.chen@linux.intel.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=zhanxusheng1024@gmail.com \
    --cc=zhanxusheng@xiaomi.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®