mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Zhang Qiao <zhangqiao22@huawei.com>
To: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: <tanghui20@huawei.com>, Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	<linux-kernel@vger.kernel.org>
Subject: Re: [RFC PATCH] sched/fair: scale wake_wide() threshold by SMT width
Date: Thu, 16 Apr 2026 15:41:47 +0800	[thread overview]
Message-ID: <45712e75-f27e-804a-940b-79e932452e6e@huawei.com> (raw)
In-Reply-To: <4276a8b7-141a-4e0b-bc35-47cb0a341801@linux.ibm.com>

Hi Shrikanth,

在 2026/4/8 1:58, Shrikanth Hegde 写道:
> Hi.
> 
> On 4/7/26 12:09 PM, Zhang Qiao wrote:
>> wake_wide() uses sd_llc_size as the spreading threshold to detect wide
>> waker/wakee relationships and to disable wake_affine() for those cases.
>>
>> On SMT systems, sd_llc_size counts logical CPUs rather than physical
>> cores. This inflates the wake_wide() threshold, allowing wake_affine()
>> to pack more tasks into one LLC domain than the actual compute capacity
>> of its physical cores can sustain. The resulting SMT interference may
>> cost more than the cache-locality benefit wake_affine() intends to gain.
>>
> 
> Isn't load balance to move it out? What does the workload do?

The workload is a producer-consumer model: one producer wakes up ~50
different consumers, with roughly 10+ consumers running concurrently.
The total number of tasks is well below the CPU count.

In this scenario, load balancing is largely ineffective. Each consumer
spends most of its time sleeping, gets woken by the producer, runs
briefly to process the message, then goes back to sleep. There is
almost no window where a consumer sits on a CPU runqueue in the runnable
state waiting to be pulled. Since load balancing can only migrate
runnable tasks, it simply has no target to act on here.

> 
>> Scale the factor by the SMT width of the current CPU so that it
>> approximates the number of independent physical cores in the LLC domain,
>> making wake_wide() more likely to kick in before SMT interference
>> becomes significant. On non-SMT systems the SMT width is 1 and behaviour
>> is unchanged.
>>
> 
> There are systems where LLC_SIZE == SMT_SIZE. i.e one core in the LLC.
> This would effectively disable wake_affine feature in such systems.
> 
> Power10 being a major example.
> 
>> Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com>
>> ---
>>   kernel/sched/fair.c | 5 +++++
>>   1 file changed, 5 insertions(+)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index f07df8987a5ef..4896582c6e904 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -7334,6 +7334,11 @@ static int wake_wide(struct task_struct *p)
>>       unsigned int slave = p->wakee_flips;
>>       int factor = __this_cpu_read(sd_llc_size);
>>   +    /* Scale factor to physical-core count to account for SMT interference. */
>> +    if (sched_smt_active())
>> +        factor = DIV_ROUND_UP(factor,
>> +                cpumask_weight(cpu_smt_mask(smp_processor_id())));
>> +
>>       if (master < slave)
>>           swap(master, slave);
>>       if (slave < factor || master < slave * factor)
> 
> .

  reply	other threads:[~2026-04-16  7:41 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-04-07  6:39 Zhang Qiao
2026-04-07 14:38 ` Dietmar Eggemann
2026-04-07 18:16   ` Shrikanth Hegde
2026-04-22 13:24     ` Dietmar Eggemann
2026-04-07 17:58 ` Shrikanth Hegde
2026-04-16  7:41   ` Zhang Qiao [this message]
2026-04-22 13:26     ` Dietmar Eggemann
2026-04-29  2:43       ` Zhang Qiao
2026-05-11 15:54         ` Dietmar Eggemann
2026-05-18  7:58           ` Zhang Qiao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=45712e75-f27e-804a-940b-79e932452e6e@huawei.com \
    --to=zhangqiao22@huawei.com \
    --cc=bsegall@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sshegde@linux.ibm.com \
    --cc=tanghui20@huawei.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®