From: Chen Yu <yu.c.chen@intel.com>
To: Mike Galbraith <efault@gmx.de>, <kprateek.nayak@amd.com>
Cc: Aishwarya Rambhadran <aishwarya.rambhadran@arm.com>,
Peter Zijlstra <peterz@infradead.org>, <mingo@kernel.org>,
<longman@redhat.com>, <chenridong@huaweicloud.com>,
<juri.lelli@redhat.com>, <vincent.guittot@linaro.org>,
<dietmar.eggemann@arm.com>, <rostedt@goodmis.org>,
<bsegall@google.com>, <mgorman@suse.de>, <vschneid@redhat.com>,
<tj@kernel.org>, <hannes@cmpxchg.org>, <mkoutny@suse.com>,
<cgroups@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<jstultz@google.com>, <kprateek.nayak@amd.com>,
<qyousef@layalina.io>, Ryan Roberts <ryan.roberts@arm.com>,
<chen.yu@linux.dev>
Subject: Re: [REGRESSION] [PATCH v3 7/7] sched/eevdf: Move to a single runqueue
Date: Mon, 28 Sep 2026 11:12:40 +0800 [thread overview]
Message-ID: <arnbKNWNIIEcQ0a-@chenyu-dev> (raw)
In-Reply-To: <128e93ca1424630546e39f9e9c1377f8185f7644.camel@gmx.de>
Hi Mike, Prateek,
On Thu, Sep 24, 2026 at 10:19:35AM +0200, Mike Galbraith wrote:
> On Thu, 2026-09-24 at 00:48 +0800, Chen Yu wrote:
> >
> > Would disabling WA_WEIGHT help? I happened to find that this workaround works
> > for me to restore netperf performance (though it is a throughput score rather
> > than a latency one), it seems that task_h_load() becomes much
> > larger(also on sched/task_h_load):
> > echo NO_WA_WEIGHT > /sys/kernel/debug/sched/features
> > https://lore.kernel.org/lkml/aoxah90s0bQ4tcUW@three-body/
>
> A long while back, I caught WA_WEIGHT causing firefox's event thread
> pool to stack on a single CPU.
I just realized that what I suggested - disabling WA_WEIGHT - does the opposite.
Actually, WA_WEIGHT does help in the netperf scenario. Disabling WA_WEIGHT makes
the performance of both "smp" and "concur" modes drop to the same poor level, which
makes it look like the regression disappears, but actually it does not.
According to my test last week, stacking the waker and wakee on the same CPU brings
a big improvement on my machine, iff the memory bandwidth is saturated. By comparison,
if the memory bandwidth is low, stacking the wakee on top of the waker causes harm.
The machine I used has 2 NUMA nodes; each node has 96 cores, and these 96 cores share
L3, while every 4 cores share L2. The L2 miss rate is much lower if the waker and wakee
are on the same CPU than if the wakee is put on other idle CPUs in the same L2 domain.
Yes, the same L2 - which goes against intuition. How could the L2 miss rate rise even
if the waker and wakee are in the same L2 domain?
I guess this is because stacking the waker and wakee on the same CPU and letting them
run alternately would be the most efficient way to consume the data without evicting
L2 cache lines too much, when the memory bandwidth is saturated.
> In that particular case, the victims
> (32 of 'em in 8 rq box!) did nothing but go back to sleep, so no real
> harm was done. However, I also measured modest real stacking impact, so
> gave task_h_load() a floor on GP.
>
> Perhaps weightless burst stacking is a bigger deal for some than it was
> here? That firefox event thread burst showed potential.
>
It depends on the workload and the memory bandwidth of the system, I suppose.
If the workload is a producer-consumer type, and the system is under memory bandwidth
pressure, stacking the producer and consumer would bring good data locality - in terms
of L2 locality, per my test.
In summary, the "regression" of flat cgroup task-pick on netperf (or, producer-consumer
type workloads) seems to be the expected consequence. Maybe we need to make WA_WEIGHT
consider more factors here.
I'll collect more data and provide the information later.
thanks,
Chenyu
> -Mike
next prev parent reply other threads:[~2026-09-28 3:25 UTC|newest]
Thread overview: 40+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-05 12:40 [PATCH v3 0/7] sched: Flatten the pick Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 1/7] sched/fair: Add cgroup_mode switch Peter Zijlstra
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 2/7] sched/fair: Add cgroup_mode: up Peter Zijlstra
2026-06-05 15:07 ` Peter Zijlstra
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 3/7] sched/fair: Add cgroup_mode: max Peter Zijlstra
2026-06-10 15:09 ` Waiman Long
2026-06-10 15:42 ` Waiman Long
2026-06-11 13:49 ` Peter Zijlstra
2026-06-11 13:47 ` Peter Zijlstra
2026-06-11 20:57 ` Waiman Long
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 4/7] sched/fair: Add cgroup_mode: concur Peter Zijlstra
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 5/7] sched/fair: Add cgroup_mode: tasks Peter Zijlstra
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 6/7] sched/fair: Change the default cgroup_mode to concur Peter Zijlstra
2026-06-30 9:03 ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 7/7] sched/eevdf: Move to a single runqueue Peter Zijlstra
2026-06-20 3:54 ` Chen, Yu C
2026-06-26 11:40 ` Peter Zijlstra
2026-06-29 14:02 ` Vincent Guittot
2026-06-30 9:03 ` [tip: sched/core] sched/fair: Fix overflow in update_tg_cfs_runnable() tip-bot2 for Chen, Yu C
2026-06-30 9:03 ` [tip: sched/core] sched/eevdf: Move to a single runqueue tip-bot2 for Peter Zijlstra (Intel)
2026-09-23 14:21 ` [REGRESSION] [PATCH v3 7/7] " Aishwarya Rambhadran
2026-09-23 16:48 ` Chen Yu
2026-09-24 7:05 ` K Prateek Nayak
2026-09-24 8:19 ` Mike Galbraith
2026-09-28 3:12 ` Chen Yu [this message]
2026-09-28 13:00 ` Mike Galbraith
2026-06-09 5:37 ` [PATCH v3 0/7] sched: Flatten the pick K Prateek Nayak
2026-06-12 2:29 ` Shubhang Kaushik
2026-08-17 16:05 ` Szabina Korbai
2026-08-17 16:35 ` K Prateek Nayak
2026-08-18 9:04 ` Szabina Korbai
2026-08-18 9:16 ` Peter Zijlstra
2026-08-21 10:37 ` Szabina Korbai
2026-08-24 14:51 ` Chen Yu
2026-08-24 13:53 ` Peter Zijlstra
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arnbKNWNIIEcQ0a-@chenyu-dev \
--to=yu.c.chen@intel.com \
--cc=aishwarya.rambhadran@arm.com \
--cc=bsegall@google.com \
--cc=cgroups@vger.kernel.org \
--cc=chen.yu@linux.dev \
--cc=chenridong@huaweicloud.com \
--cc=dietmar.eggemann@arm.com \
--cc=efault@gmx.de \
--cc=hannes@cmpxchg.org \
--cc=jstultz@google.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=longman@redhat.com \
--cc=mgorman@suse.de \
--cc=mingo@kernel.org \
--cc=mkoutny@suse.com \
--cc=peterz@infradead.org \
--cc=qyousef@layalina.io \
--cc=rostedt@goodmis.org \
--cc=ryan.roberts@arm.com \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®