From: K Prateek Nayak <kprateek.nayak@amd.com>
To: Chen Yu <chen.yu@linux.dev>
Cc: Peter Zijlstra <peterz@infradead.org>,
Chen Yu <yu.c.chen@intel.com>,
Tim Chen <tim.c.chen@linux.intel.com>,
Ingo Molnar <mingo@redhat.com>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Andrew Morton <akpm@linux-foundation.org>,
Arnd Bergmann <arnd@arndb.de>, <linux-kernel@vger.kernel.org>,
<linux-arch@vger.kernel.org>, <linux-s390@vger.kernel.org>,
<linuxppc-dev@lists.ozlabs.org>, <linux-mips@vger.kernel.org>,
<loongarch@lists.linux.dev>, <driver-core@lists.linux.dev>,
Sudeep Holla <sudeep.holla@kernel.org>,
Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
"Rafael J. Wysocki" <rafael@kernel.org>,
Danilo Krummrich <dakr@kernel.org>,
Huacai Chen <chenhuacai@kernel.org>,
Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
Jiaxun Yang <jiaxun.yang@flygoat.com>,
Madhavan Srinivasan <maddy@linux.ibm.com>,
Heiko Carstens <hca@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Alexander Gordeev <agordeev@linux.ibm.com>,
"David S. Miller" <davem@davemloft.net>,
Andreas Larsson <andreas@gaisler.com>,
Thomas Gleixner <tglx@kernel.org>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>, <x86@kernel.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
WANG Xuerui <kernel@xen0n.name>,
Michael Ellerman <mpe@ellerman.id.au>,
Nicholas Piggin <npiggin@gmail.com>,
Christophe Leroy <chleroy@kernel.org>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Sven Schnelle <svens@linux.ibm.com>,
"H. Peter Anvin" <hpa@zytor.com>
Subject: Re: [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
Date: Sun, 4 Oct 2026 11:43:02 +0530 [thread overview]
Message-ID: <08df1d81-7671-425b-85bb-51ee79c2eac7@amd.com> (raw)
In-Reply-To: <asDGbBEHkDouWOwC@three-body>
Hello Chenyu,
On 10/3/2026 2:40 PM, Chen Yu wrote:
> Hello Prateek!
>
> Thanks for bringing this interesting topic,
>
> On Thu, Oct 01, 2026 at 07:28:36PM +0000, K Prateek Nayak wrote:
>>
>> Problem
>> =======
>>
>> %cycles vs global mask operation
>>
>> global mask : 100.0000% (var: 3.28%)
>> per-NUMA mask : 32.9209% (var: 7.77%)
>> per-LLC mask : 1.2977% (var: 4.85%)
>> per-LLC mask (u8 operation; no LOCK prefix) : 0.4930% (var: 0.83%)
>>
>
> This shows a significant latency improvement, especially in the per-LLC (u64, u8) case.
> May I know if schbench / sched-messaging were used?
No, this was a custom benchmark with two threads per
CPU - one setting the CPU on bitmask, and other clearing it,
continuously yielding to each other.
It gives an idea of what the worst case looks like when there
may be short idling followed by a short runtime going in
cycles.
>
>>
>> Future work
>> ===========
>>
>> o Interoperability with cpumaks since sbm lose the crucial optimizations
>> that come naturally from for_each_cpu_and() iterations.
>>
>> o Different data representation - using the u8 variant for updates and
>> then perform a "gather" operation to build a dense mask.
>>
>> o Extending sbm work to help in wakeup (and possibly resurrect Mel's
>> optimization from [4] in some form). The current sbm is still far away
>> from being used for wakeups since updates to sbm leaf, even on a
>> 16CPUs per LLC system is visible in benchmark performance (~8-10%).
>>
>
> If we leverage sbm for the wakeup path, it is a per-LLC mask, there seems to be
> no much difference from Mel Gorman's proposal of allocating per sd_share
> unsigned long idle_cpus_span[]?
Ack! But the current form is still pretty expensive. I see about a
10% overhead of just maintaining that mask which is what I'm trying
to reduce.
> The frequent update to this mask might still cause
> c2c latency within 1 LLC. A wild guess is that maybe the u8 version is more suitable,
> because it has only max-to-8 CPUs touching the mask at the same time?
With a 64B cacheline, single cacheline can contain data for up to
64 CPUs - unlike current sbm that uses first 8-bytes, this used
the whole 64 bytes.
P.S. All versions were tested with 16 CPUs per LLC on my systems.
Going from atomic u64 to plain u8 writes probably avoids an
expensive atomic path in the H/W making them faster.
> My understanding is that the major case that sbm could fit is turning the global bitmask
> into a per-LLC bitmask, because it mainly avoids CPUs on different LLC/node writing the
> same cache line frequently (nohz.idle_cpus_mask set via nohz_balance_enter_idle() on
> many CPUs, etc), which might cause a costly cache RFO event storm. Meanwhile, with sbm,
> at the reader side, _nohz_idle_balance() could start scanning from the current CPU to find
> an idle CPU, so as to avoid the costly HITM event - the reader is on LLC1, while the writer
> is on LLC0 - so maybe:
>
> for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1)
>
> could start from this_cpu's LLC sibling first, rather than this_cpu + 1, because this_cpu+1
> might not always be the LLC sibling of this_cpu.
Ah! Good point. Let me see if wrapping within the bitmask leaf
first and then going out makes any difference.
>
> I found that in the current code, there are also other global mask:
> rd->rto_mask(mentioned by Pan Deng when running ffmpeg[1])
> rd->dlo_mask
> tick_broadcast_**mask
> maybe they can also be converted into sbm.
Ack! I was juts getting started somewhere to see if there is an
appetite for sbm :-)
>
> [1] https://lore.kernel.org/lkml/a3207ebf537bbe5605ff5454f63b5604d83a04a0.1753076363.git.pan.deng@intel.com/
>
> thanks,
> Chenyu
--
Thanks and Regards,
Prateek
prev parent reply other threads:[~2026-10-04 6:13 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 19:28 K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-03 8:27 ` Chen Yu
2026-10-04 6:17 ` K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
2026-10-03 9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
2026-10-04 6:13 ` K Prateek Nayak [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=08df1d81-7671-425b-85bb-51ee79c2eac7@amd.com \
--to=kprateek.nayak@amd.com \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=andreas@gaisler.com \
--cc=arnd@arndb.de \
--cc=borntraeger@linux.ibm.com \
--cc=bp@alien8.de \
--cc=bsegall@google.com \
--cc=chen.yu@linux.dev \
--cc=chenhuacai@kernel.org \
--cc=chleroy@kernel.org \
--cc=dakr@kernel.org \
--cc=dave.hansen@linux.intel.com \
--cc=davem@davemloft.net \
--cc=dietmar.eggemann@arm.com \
--cc=driver-core@lists.linux.dev \
--cc=gor@linux.ibm.com \
--cc=gregkh@linuxfoundation.org \
--cc=hca@linux.ibm.com \
--cc=hpa@zytor.com \
--cc=jiaxun.yang@flygoat.com \
--cc=juri.lelli@redhat.com \
--cc=kernel@xen0n.name \
--cc=linux-arch@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mips@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=loongarch@lists.linux.dev \
--cc=maddy@linux.ibm.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=peterz@infradead.org \
--cc=rafael@kernel.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=sudeep.holla@kernel.org \
--cc=svens@linux.ibm.com \
--cc=tglx@kernel.org \
--cc=tim.c.chen@linux.intel.com \
--cc=tsbogend@alpha.franken.de \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=x86@kernel.org \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®