From: Chen Yu <yu.c.chen@intel.com>
To: Tim Chen <tim.c.chen@linux.intel.com>
Cc: Chen Yu <chen.yu@linux.dev>,
Zhan Xusheng <zhanxusheng1024@gmail.com>, <peterz@infradead.org>,
<mingo@redhat.com>, <juri.lelli@redhat.com>,
<vincent.guittot@linaro.org>, <dietmar.eggemann@arm.com>,
<rostedt@goodmis.org>, <bsegall@google.com>, <mgorman@suse.de>,
<vschneid@redhat.com>, <kprateek.nayak@amd.com>,
<linux-kernel@vger.kernel.org>, <zhanxusheng@xiaomi.com>
Subject: Re: sched/fair: which tasks should nr_pref_llc_running be compared against?
Date: Wed, 9 Sep 2026 19:40:43 +0800 [thread overview]
Message-ID: <aqFFu1Xo52cQV3iy@fengwei-dev> (raw)
In-Reply-To: <2b0a35122ee615c6fa51076e5d79330e633755ac.camel@linux.intel.com>
On Fri, Sep 04, 2026 at 01:53:10PM -0700, Tim Chen wrote:
> Good idea - the four sites really are one operation ("if the task is
> queued on its preferred LLC and runnable, move the counter"), and
> folding the two conditions into one place is what keeps them from
> drifting apart later. I've adopted it in v3; account_llc_delayed() and
> account_llc_requeue_delayed() are gone.
>
> I split it slightly differently: a membership predicate
>
> static bool task_pref_llc_runnable(struct task_struct *p)
> {
> return p->pref_llc_queued && !p->se.sched_delayed;
> }
>
> with pref_llc_running_inc()/pref_llc_running_dec() wrappers over it, so
> the call sites read as inc/dec rather than passing a +1/-1 delta.
>
> Two things to note:
>
> 1) I kept the call site comments of pref_llc_running_inc/dec().
> The helper name says *what* happens,
> but not *why* it is safe across the delay-dequeue transition - that
> set_delayed() already did the decrement, so account_llc_dequeue()
> must skip it, and that clearing pref_llc_queued there is what
> neutralizes the following clear_delayed().
>
> 2) The pref_llc_running_dec() placement in set_delayed()
> is subtle. It has to be before se->sched_delayed = 1 or
> decrement would not happen. That deserves a comment so no
> one would move sched_delayed = 1 before the decrement.
>
>
> alb_break_llc() decides whether to break LLC preference during active
> load balance. It does so by testing that every runnable fair task on the
> source rq prefers its LLC:
>
> env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
>
> But the two counters cover different sets. nr_pref_llc_running is updated
> in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
> so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
> clear_delayed() and drops delay-dequeued tasks.
>
> So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
> in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
> alb_break_llc() returns false, and active balance is free to pull a task
> off its preferred LLC. Active balance only moves runnable tasks, and this
> is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
> skips the per-task test in can_migrate_task(). The runnable set is the one
> we want.
>
> Fix it on the counter side. A task should be counted in
> nr_pref_llc_running exactly while it is both queued on its preferred LLC
> (pref_llc_queued) and runnable (!sched_delayed). Define that membership
> once in task_pref_llc_runnable(), and adjust the counter only through
> pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
> change either input: account_llc_enqueue(), account_llc_dequeue(),
> set_delayed() and clear_delayed(). Gating every update on the same
> predicate keeps the delay, wake and dequeue paths from double-counting
> or underflowing; see the comments at those sites for the ordering.
>
> nr_llc_running and sd->llc_counts are not touched and stay on queued
> semantics.
>
> Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
> Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
> Suggested-by: Chen Yu <yu.c.chen@intel.com>
> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
> ---
> Based on v7.3-rc1.
>
> Changes in v3:
> - Route every nr_pref_llc_running adjustment through a single membership
> predicate task_pref_llc_runnable(), with pref_llc_running_inc()/
> pref_llc_running_dec() wrappers, instead of four open-coded sites
> (Chen Yu). Keep the per-site comments that explain the delay-dequeue
> interaction, and note that set_delayed() must adjust the counter
> before setting se->sched_delayed.
>
> Changes in v2:
> - Prevent a delay-dequeued task from being counted as running in the
> enqueue path (Chen Yu).
>
> kernel/sched/fair.c | 52 ++++++++++++++++++++++++++++++++++++++++++++++++++--
> 1 file changed, 50 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8dff37059faf..72aae7a50b8b 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1538,6 +1538,28 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
> (scale * per_cpu(sd_llc_size, cpu)));
> }
>
> +/*
> + * A task counts in nr_pref_llc_running while it is queued on its preferred
> + * LLC (pref_llc_queued) and runnable (!sched_delayed), keeping the counter in
> + * the runnable domain so alb_break_llc() can compare it with h_nr_runnable.
> + */
> +static bool task_pref_llc_runnable(struct task_struct *p)
> +{
> + return p->pref_llc_queued && !p->se.sched_delayed;
> +}
> +
> +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p)
> +{
> + if (task_pref_llc_runnable(p))
> + rq->nr_pref_llc_running++;
> +}
> +
> +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p)
> +{
> + if (task_pref_llc_runnable(p))
> + rq->nr_pref_llc_running--;
> +}
> +
> static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
> {
> int pref_llc, pref_llc_queued;
> @@ -1549,7 +1571,6 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
>
> pref_llc_queued = (pref_llc == task_llc(p));
> rq->nr_llc_running++;
> - rq->nr_pref_llc_running += pref_llc_queued;
>
> /*
> * Record whether p is enqueued on its preferred
> @@ -1567,6 +1588,9 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
> */
> p->pref_llc_queued = pref_llc_queued;
>
> + /* Skipped while delayed; clear_delayed() adds it back on wake. */
> + pref_llc_running_inc(rq, p);
> +
> sd = rcu_dereference_all(rq->sd);
> if (sd && (unsigned int)pref_llc < sd->llc_max)
> sd->llc_counts[pref_llc]++;
> @@ -1583,7 +1607,12 @@ static void account_llc_dequeue(struct rq *rq, struct task_struct *p)
>
> rq->nr_llc_running--;
> if (p->pref_llc_queued) {
> - rq->nr_pref_llc_running--;
> + /*
> + * Skipped if still delayed (set_delayed() already removed it);
> + * clearing pref_llc_queued below also stops clear_delayed()
> + * from re-adding it.
> + */
> + pref_llc_running_dec(rq, p);
> /*
> * Update the status in case
> * other logic might query
> @@ -2008,6 +2037,10 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p) {}
>
> static void account_llc_dequeue(struct rq *rq, struct task_struct *p) {}
>
> +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p) {}
> +
> +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p) {}
> +
> #endif /* CONFIG_SCHED_CACHE */
>
> /*
> @@ -6382,6 +6415,14 @@ static __always_inline void return_cfs_rq_runtime(struct cfs_rq *cfs_rq);
>
> static void set_delayed(struct sched_entity *se)
> {
> + /*
> + * Drop a task leaving the runnable set. Must run before sched_delayed
> + * is set, or task_pref_llc_runnable() would already exclude it;
> + * clear_delayed() mirrors this after clearing the flag.
> + */
> + if (entity_is_task(se))
> + pref_llc_running_dec(rq_of(cfs_rq_of(se)), task_of(se));
> +
> se->sched_delayed = 1;
>
> /*
> @@ -6412,6 +6453,13 @@ static void clear_delayed(struct sched_entity *se)
> if (!entity_is_task(se))
> return;
>
> + /*
> + * Re-add on wake, after sched_delayed is cleared. On a final delayed
> + * dequeue account_llc_dequeue() already cleared pref_llc_queued, so
> + * this does nothing.
> + */
> + pref_llc_running_inc(rq_of(cfs_rq_of(se)), task_of(se));
> +
> for_each_sched_entity(se) {
> struct cfs_rq *cfs_rq = cfs_rq_of(se);
>
> --
> 2.32.0
>
>
>
Yes, I think this version looks good now. While looking back at Xusheng's proposal,
I noticed there is another option:
if (env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_queued) {
...
}
May I know why we did not choose this approach, is it because of the following
scenario?
Suppose there are 3 queued tasks: p1 and p2 prefer the src_rq, while p3 is a delayed
task that also prefers src_rq. In the current implementation, nr_pref_llc_running is 3
and h_nr_runnable is 2, so alb_break_llc() might return false. As a result, active load
balance would be triggered, and p1 or p2 might be migrated away, which is undesirable.
However, would this still be a problem after Lu Wang's active load balance guard patch
has been applied?
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@gmail.com/
thanks,
Chenyu
next prev parent reply other threads:[~2026-09-09 11:54 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 13:50 Zhan Xusheng
2026-08-27 20:57 ` Tim Chen
2026-08-28 2:20 ` [PATCH] sched/cache: Keep nr_pref_llc_running in the runnable domain Zhan Xusheng
2026-08-28 17:08 ` Tim Chen
2026-08-30 8:17 ` sched/fair: which tasks should nr_pref_llc_running be compared against? Chen Yu
2026-09-01 20:42 ` Tim Chen
2026-09-03 15:21 ` Chen Yu
2026-09-04 20:53 ` Tim Chen
2026-09-09 11:40 ` Chen Yu [this message]
2026-09-09 17:47 ` Tim Chen
2026-09-10 10:47 ` Chen Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqFFu1Xo52cQV3iy@fengwei-dev \
--to=yu.c.chen@intel.com \
--cc=bsegall@google.com \
--cc=chen.yu@linux.dev \
--cc=dietmar.eggemann@arm.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=tim.c.chen@linux.intel.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=zhanxusheng1024@gmail.com \
--cc=zhanxusheng@xiaomi.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®