From: Shrikanth Hegde <sshegde@linux.ibm.com>
To: K Prateek Nayak <kprateek.nayak@amd.com>,
Peter Zijlstra <peterz@infradead.org>,
Michael Ellerman <mpe@ellerman.id.au>,
linuxppc-dev@lists.ozlabs.org,
Madhavan Srinivasan <maddy@linux.ibm.com>,
Thomas Gleixner <tglx@linutronix.de>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Nicholas Piggin <npiggin@gmail.com>,
Christophe Leroy <chleroy@kernel.org>,
Chen Yu <yu.c.chen@intel.com>,
Tim Chen <tim.c.chen@linux.intel.com>,
Ingo Molnar <mingo@redhat.com>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Andrew Morton <akpm@linux-foundation.org>,
Arnd Bergmann <arnd@arndb.de>,
linux-kernel@vger.kernel.org, linux-arch@vger.kernel.org,
linux-s390@vger.kernel.org, linux-mips@vger.kernel.org,
loongarch@lists.linux.dev, driver-core@lists.linux.dev,
Ritesh Harjani <ritesh.list@gmail.com>,
Srikar Dronamraju <srikar@linux.ibm.com>
Subject: Re: [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
Date: Wed, 7 Oct 2026 20:13:28 +0530 [thread overview]
Message-ID: <61b01870-ddcf-4c4f-8604-dd63aa4a07a3@linux.ibm.com> (raw)
In-Reply-To: <20261001192849.74788-7-kprateek.nayak@amd.com>
Hi Prateek.
On 10/2/26 12:58 AM, K Prateek Nayak wrote:
> Initialize sparsebitmap (sbm) topology based on the coregroup
> information. Each coregorup gets its own sparsemask leaf.
>
> pSeries and memory hotplug are interesting since a hotplug can place a
> newly added CPU on any online node. This requires special care to allow
> estimating bitmask size considering the worst case scenarios - each CPU
> onlined is on a separate node, and this is a new N_CPU node.
>
> Platform may enforce a stricter standards for the CPUs being online and
> what nodes they can be mapped to but the current implementations makes
> no assumptions and considers each CPU can be onlined on a unique node.
>
I think this suffers the same fate as structures which are allocated at boot
time such as runqueues.
So your fallback option of putting all the disabled into singleton node may be
sensible option. (If you are not doing that already)
But yhea, will see more into it, this changing node stuff is new for me too.
Also i need to read your patch series too :)
> pSeries systems that can hotplug CPUs (detected using smp_ops) use the
> NUMA topology instead for sbm initialization. arch_sbm_cpu_instance_id()
> on these systems use cpu_to_node() mappings to match CPUs to sbm
> instances.
>
> XXX: This requires further optimizations to shorten sparsemask
> traversals by keeping the number of leaf nodes to a minimum. If there
> are nuances I'm not aware of, please reach out.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> Tested on ppc64le_defconfig with:
>
> qemu-system-ppc64 \
> -M pseries \
> -cpu power10 \
> -smp sockets=2,cores=2,threads=4 \
> -m 10G -nographic \
> -kernel vmlinux \
> -append "root=/dev/ram sched_debug"
>
> and also on ppce500 VM based on instructions in
> https://www.qemu.org/docs/master/system/ppc/ppce500.html
> ---
> arch/powerpc/kernel/setup-common.c | 88 ++++++++++++++++++++
> arch/powerpc/platforms/pseries/hotplug-cpu.c | 10 +++
> 2 files changed, 98 insertions(+)
>
> diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
> index 4afaba19b586..4b57ad553172 100644
> --- a/arch/powerpc/kernel/setup-common.c
> +++ b/arch/powerpc/kernel/setup-common.c
> @@ -11,6 +11,7 @@
> #include <linux/export.h>
> #include <linux/panic_notifier.h>
> #include <linux/string.h>
> +#include <linux/sbm.h>
> #include <linux/sched.h>
> #include <linux/init.h>
> #include <linux/kernel.h>
> @@ -602,6 +603,87 @@ static __init int add_pcspkr(void)
> device_initcall(add_pcspkr);
> #endif /* CONFIG_PCSPKR_PLATFORM */
>
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + /*
> + * In case of pSeries processors, sbm masks are
> + * grouped by nodes where the cpuhotplug
> + * operations can remove and re-add same logical
> + * CPUs on different nodes.
> + *
> + * See comment in pseries_cpu_hotplug_init().
> + */
> + if (smp_ops->cpu_disable)
> + return cpu_to_node(cpu);
> +
> + return cpu_to_coregroup_id(cpu);
> +}
> +
> +static void __init setup_sbm_topology(void)
> +{
> + int num_sbm_instances, max_threads_per_instance = 1;
> + struct cpumask *cpu_sbm_setup_map;
> + int i, *__node_thread_count;
> + int disabled_cpus = 0;
> +
> + cpu_sbm_setup_map = memblock_alloc_or_panic(cpumask_size(), __alignof__(long));
> + __node_thread_count = memblock_alloc_or_panic(nr_cpu_ids * sizeof(int),
> + __alignof__(int));
> +
> + memset(__node_thread_count, 0, nr_cpu_ids * sizeof(int));
> + memset(cpu_sbm_setup_map, 0, cpumask_size());
> +
> + for_each_possible_cpu(i) {
> + bool found = false;
> + int j;
> +
> + if (!cpu_present(i)) {
> + disabled_cpus += 1;
> + continue;
> + }
> +
> + for_each_cpu(j, cpu_sbm_setup_map) {
> + if (cpu_to_coregroup_id(i) == cpu_to_coregroup_id(j)) {
> + found = true;
> + break;
> + }
> + }
> +
> + if (!found) {
> + cpumask_set_cpu(i, cpu_sbm_setup_map);
> + __node_thread_count[i] = 1;
> + continue;
> + }
> +
> + __node_thread_count[j] += 1;
> + max_threads_per_instance = max(max_threads_per_instance,
> + __node_thread_count[j]);
> + }
> +
> + /*
> + * If CPUs are disabled, they may pop up on any online node.
> + *
> + * XXX: Any implementation nuances that can help this?
> + * pSeries says only online nodes can be extended.
> + */
> + if (disabled_cpus) {
> + num_sbm_instances = num_sbm_instances + disabled_cpus;
> + } else {
> + num_sbm_instances = cpumask_weight(cpu_sbm_setup_map);
> + }
> +
> + /*
> + * If disabled threads exists, assume the maximum threads per
> + * instance can extend by the number of disabled threads if they
> + * are all added to the same node.
> + */
> + sbm_set_topology(num_sbm_instances,
> + max_threads_per_instance + disabled_cpus);
> +
> + memblock_free(__node_thread_count, nr_cpu_ids * sizeof(int));
> + memblock_free(cpu_sbm_setup_map, cpumask_size());
> +}
> +
> static char ppc_hw_desc_buf[128] __initdata;
>
> struct seq_buf ppc_hw_desc __initdata = {
> @@ -1006,6 +1088,12 @@ void __init setup_arch(char **cmdline_p)
>
> early_memtest(min_low_pfn << PAGE_SHIFT, max_low_pfn << PAGE_SHIFT);
>
> + /*
> + * setup_arch() below can override topology for
> + * pSeries platforms as a result of hotplug nuances.
> + */
> + setup_sbm_topology();
> +
> if (ppc_md.setup_arch)
> ppc_md.setup_arch();
>
> diff --git a/arch/powerpc/platforms/pseries/hotplug-cpu.c b/arch/powerpc/platforms/pseries/hotplug-cpu.c
> index bc6926dbf148..7c1c1ac3efde 100644
> --- a/arch/powerpc/platforms/pseries/hotplug-cpu.c
> +++ b/arch/powerpc/platforms/pseries/hotplug-cpu.c
> @@ -19,6 +19,7 @@
> #include <linux/kernel.h>
> #include <linux/interrupt.h>
> #include <linux/delay.h>
> +#include <linux/sbm.h>
> #include <linux/sched.h> /* for idle_task_exit */
> #include <linux/sched/hotplug.h>
> #include <linux/cpu.h>
> @@ -870,6 +871,15 @@ void __init pseries_cpu_hotplug_init(void)
> return;
> }
>
> + /*
> + * find_cpu_id_range() only looks at online nodes.
> + *
> + * XXX: Is it possible for a CPU attached memory node to come
> + * online after this point? May need num_possbile_nodes() then
> + * unless there are platform nuances that can help optimize.
> + */
> + sbm_set_topology(num_online_nodes(), num_possible_cpus());
> +
> smp_ops->cpu_offline_self = pseries_cpu_offline_self;
> smp_ops->cpu_disable = pseries_cpu_disable;
> smp_ops->cpu_die = pseries_cpu_die;
next prev parent reply other threads:[~2026-10-07 14:44 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-07 5:43 ` Shrikanth Hegde
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
2026-10-07 3:51 ` [RFC PATCH v3.1 " K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
2026-10-07 14:43 ` Shrikanth Hegde [this message]
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
2026-10-07 10:19 ` Mete Durlu
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-03 8:27 ` Chen Yu
2026-10-04 6:17 ` K Prateek Nayak
2026-10-07 3:52 ` [RFC PATCH v3.1 " K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
2026-10-03 9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
2026-10-04 6:13 ` K Prateek Nayak
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=61b01870-ddcf-4c4f-8604-dd63aa4a07a3@linux.ibm.com \
--to=sshegde@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=arnd@arndb.de \
--cc=bsegall@google.com \
--cc=chleroy@kernel.org \
--cc=dietmar.eggemann@arm.com \
--cc=driver-core@lists.linux.dev \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-arch@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mips@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=loongarch@lists.linux.dev \
--cc=maddy@linux.ibm.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=peterz@infradead.org \
--cc=ritesh.list@gmail.com \
--cc=rostedt@goodmis.org \
--cc=srikar@linux.ibm.com \
--cc=tglx@linutronix.de \
--cc=tim.c.chen@linux.intel.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®