mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
@ 2026-10-01 19:28 K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
                   ` (13 more replies)
  0 siblings, 14 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui, Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

Problem
=======

Global cpumasks on a large multi-node systems experience an abundance of
C2C ping ponging, especially the ones that are updated very frequently
like the ones that track scheduler and timer states.


Solution
========

One great solution to this problem is to have a separate cpumask
instance per node to avoid C2C ping poinging for updates.

Replicating cpumasks in its entirity solves the challenges of updates by
keeping the writes local to the LLC domain but introduces burden on
traversal.

Steve Sistare had (almost a decade back) introduced a concept called
sparsemask [1] where each cacheline-aligned bitmap word can only
represent a small number of CPUs with 8 CPUs per bitmap word being
chosen during that time.

Although it worked, the initial sparsemask implementation had no
knowledge of topology. That changed early this year when Peter provided
a minimal implementation of sparsebitmap (sparsebitmask?) aka sbm in
[2].

This builds on Peter's idea to make the sbm implementation generic and
available to all architectures.


Aren't sbitmap built for that purpose?
======================================

It is true sbitmap (scalable bitmap) in lib/sbitmap.c was build for a
similar problem in the block layer but sbitmap is heavy handed with each
sbitmap leaf occupying two cachelines with added semantics of "cleared"
range that don't contend with a set operation on the adjacent cacheline.

sbm is to sbitmap what cpumask is to plain bitmap - they are essentially
the same concept albeit executed slightly differently.

If there is enough interest, I don't mind adapting my implementation to
fit with the sbitmap scheme with perhaps a sbitmap-lite variant :-)


Implementation
==============

Similar to how cpumaks allows architectures to define nr_cpumask_bit and
later adapt a simple bitmap to cover that range, sbm allows
architectures to define "num_instances" (leaves) and
"max_threads_per_instance" (max CPUs that can be represented in one
instance) and uses them to compute the worst-case size for a sparse
mapping of CPUs.

sbm_alloc() essentially hands a large array based on the set topology
that cba cover any CPU mapping as long as they fall within the
architecture imposed constraints.

When CPUs are brought active, a sbm index is assigned to them. An sbm
index essentailly encodes two parts:

    +----------------+---------------------------+
    | index in array | bit in that array element |
    +----------------+---------------------------+

The size of each encoding is determined during sbm initialization based
on the topology provided. If arch/ did not provide a topology, a backup
topology is used that divdes entire system in BITS_PER_LONG chunks.

Using the topology two attributes are derived:

  __sbm_shift: Bits to shift right to obtain just index in array
  __sbm_mask:  Lower bits to mask to get the position in the bitfield

Metadata is tracked separately by core and idx <-> cpu mappings are
cached in separate arrays.

This builds on Peter's ideas in [2] which used the x86 topology parsing
nuances to derive the mask and shift during early boot.


Topology parsing nuances
========================

Each architecture has a unique approach to describing system topology
but most offer a way to get the entire system topology during
smp_prepare() with one exception of PowerPC.

PowerPC and specifically pSeries hotplug seems to indicate a CPU can be
removed and re-added with the logical CPU <-> node relation completely
changing.

This creates a unique challenge since estimating the number of sbm
leaves and bits per leaf is hard and the worst case scenario is taken to
construct them.


Shameless plug
==============

If you find yourself at LPC'26, I have a small talk at Scheduling and
Real-Time Microconference.

If I have got some topology nuances for your architecture horribly
wrong, you have a chance to punch me in the face directly :-) (and maybe
help direct me in the right direction after).


Performance
===========

To measure any performance impact, the global nohz.idle_cpus_mask which
tracks the set of CPUs in NOHZ idle state have been converted to sbm
based distributed mask.

Performance is comparable to base with sbm variant performing slightly
better (~ 1-2%) on average. On worst case scenario (one event per
context switch), the distributed masks take 1/100th the cost of update
when compared to global cpumask. Quoting data from [3]:

                                            %cycles vs global mask operation

global mask                                     : 100.0000%  (var: 3.28%)
per-NUMA mask                                   :  32.9209%  (var: 7.77%)
per-LLC mask                                    :   1.2977%  (var: 4.85%)
per-LLC mask (u8 operation; no LOCK prefix)     :   0.4930%  (var: 0.83%)


Future work
===========

o Interoperability with cpumaks since sbm lose the crucial optimizations
  that come naturally from for_each_cpu_and() iterations.

o Different data representation - using the u8 variant for updates and
  then perform a "gather" operation to build a dense mask.

o Extending sbm work to help in wakeup (and possibly resurrect Mel's
  optimization from [4] in some form). The current sbm is still far away
  from being used for wakeups since  updates  to sbm leaf, even on a
  16CPUs per LLC system is visible in benchmark performance (~8-10%).


References
==========

[1] https://lore.kernel.org/lkml/1541767840-93588-2-git-send-email-steven.sistare@oracle.com/
[2] https://lore.kernel.org/lkml/20260324120008.GB3738010@noisy.programming.kicks-ass.net/
[3] https://lore.kernel.org/lkml/e093d930-79df-4285-a492-cc6d40b3cd51@amd.com/
[4] https://lore.kernel.org/lkml/20210726102247.21437-1-mgorman@techsingularity.net/

Patches are based on:

  git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core

at commit 1fb28c664a19 ("virt/steal_governor: Enable the driver").

Respective arch/ maintainers have been Cc'd on the arch sepcific
changes, everyone is Cc'd on cover letter and sbm bits. Scheduler
and lib folks along with the lists will get the entire series.


Changelog
=========

v2..v3:

o Added enablement for other architectures apart from x86.
o Dynamically establish CPU <-> sbm index relations.

This is a spiritual successor to Chenyu's v2 at
https://lore.kernel.org/lkml/20260510155920.2587431-1-yu.c.chen@intel.com/

---
K Prateek Nayak (10):
  lib/sbm: Introduce helpers for architectures to configure LLC
    properties
  drivers/base/arch_topology: Add support for initializing sbm topology
  LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT
  LoongArch: Configure sbm topology during SMP preparation
  MIPS: Initialize sbm topology on multi-node systems
  powerpc/setup: Initialize sbm topology based on coregroup / NUMA
    topology
  s390/topology: Initialize sbm topology during topology_init_early()
  sparc64: Initialize sbm topology on multi-LLC system
  lib/sbm: Dynamically allocate sbm index when CPU is activated
  sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp()

Peter Zijlstra (3):
  x86/cpu/topology: Initialize sbm topology after topology parsing
  lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on
    sbm
  sched/fair: Switch nohz.idle_cpus to use sbm

 arch/loongarch/kernel/acpi.c                 |  20 +-
 arch/loongarch/kernel/smp.c                  |  36 +++
 arch/mips/include/asm/topology.h             |   6 +
 arch/mips/kernel/topology.c                  |  43 ++++
 arch/mips/loongson64/smp.c                   |   3 +
 arch/mips/sgi-ip27/ip27-smp.c                |   3 +
 arch/powerpc/kernel/setup-common.c           |  88 +++++++
 arch/powerpc/platforms/pseries/hotplug-cpu.c |  10 +
 arch/s390/kernel/topology.c                  |  52 ++++
 arch/sparc/kernel/setup_64.c                 |  45 ++++
 arch/x86/kernel/cpu/topology.c               |  40 ++-
 drivers/base/arch_topology.c                 | 122 ++++++++-
 include/asm-generic/vmlinux.lds.h            |   6 +-
 include/linux/sbm.h                          | 107 ++++++++
 init/main.c                                  |   6 +
 kernel/sched/core.c                          |  19 ++
 kernel/sched/fair.c                          |  72 +++---
 kernel/sched/sched.h                         |   1 +
 lib/Makefile                                 |   2 +-
 lib/sbm.c                                    | 253 +++++++++++++++++++
 20 files changed, 883 insertions(+), 51 deletions(-)
 create mode 100644 include/linux/sbm.h
 create mode 100644 lib/sbm.c


base-commit: 1fb28c664a19df8d45a6afa04d28d102b04ea680
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
                   ` (12 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui, Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

Subsequent commits will add arch/ hooks to configure the sparsebitmask
(sbm) topology - number of sbm instance and the maximum CPUs that a
single instance can represent.

Introduce sbm_set_topology() that helps architectures configure the sbm
topology. Also add arch_sbm_cpu_instance_id() which will be used to map
CPUs to the sbm instances during init to keep CPUs on the same instance
on the same sbm leaf.

Individual arch/ users will be added one at a time in the subsequent
commits.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 include/linux/sbm.h |  8 ++++++++
 lib/Makefile        |  2 +-
 lib/sbm.c           | 27 +++++++++++++++++++++++++++
 3 files changed, 36 insertions(+), 1 deletion(-)
 create mode 100644 include/linux/sbm.h
 create mode 100644 lib/sbm.c

diff --git a/include/linux/sbm.h b/include/linux/sbm.h
new file mode 100644
index 000000000000..adac12ed233a
--- /dev/null
+++ b/include/linux/sbm.h
@@ -0,0 +1,8 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_SBM_H
+#define _LINUX_SBM_H
+
+int arch_sbm_cpu_instance_id(int cpu);
+void sbm_set_topology(int num_instances, int max_threads_per_instance);
+
+#endif /* _LINUX_SBM_H */
diff --git a/lib/Makefile b/lib/Makefile
index dfab958327c5..7776330f5b25 100644
--- a/lib/Makefile
+++ b/lib/Makefile
@@ -40,7 +40,7 @@ lib-y := ctype.o string.o vsprintf.o cmdline.o \
 	 is_single_threaded.o plist.o decompress.o kobject_uevent.o \
 	 earlycpio.o seq_buf.o siphash.o dec_and_lock.o \
 	 nmi_backtrace.o win_minmax.o memcat_p.o \
-	 buildid.o objpool.o iomem_copy.o sys_info.o
+	 buildid.o objpool.o iomem_copy.o sys_info.o sbm.o
 
 lib-$(CONFIG_UNION_FIND) += union_find.o
 lib-$(CONFIG_PRINTK) += dump_stack.o
diff --git a/lib/sbm.c b/lib/sbm.c
new file mode 100644
index 000000000000..82280e3306de
--- /dev/null
+++ b/lib/sbm.c
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#include <linux/sbm.h>
+#include <linux/init.h>
+#include <linux/printk.h>
+
+static int sbm_max_threads_per_instance = -1;
+static int sbm_num_instance = -1;
+
+/*
+ * In absence of an arch definition, consider all CPUs to
+ * belong to the same sbm instance.
+ */
+int __weak arch_sbm_cpu_instance_id(int cpu)
+{
+	return 0;
+}
+
+void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
+{
+	sbm_max_threads_per_instance = max_threads_per_instance;
+	sbm_num_instance = num_instances;
+
+	pr_info("sbm topology set to %d instance(s) with a maximum of %d thread(s) per instance\n",
+		sbm_num_instance,
+		sbm_max_threads_per_instance);
+}
+
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
                   ` (11 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak

Add (fragile) support for initializing sparsebitmap topology for
architectures that support GENERIC_ARCH_TOPOLOGY.

Similar to setting cpu_smt_set_num_threads(), count the number of unique
LLCs (if last_level_cache_is_valid() true, otherwise) / packages, and
the maximum threads that exist in those instance and set
sbm_set_proc_config().

Any failure on the path will default to considering the entire processor
as a single sbm domain and the default logic will divide the system into
BITS_PER_LONG chunks. In case of a failure, sbm core will skip using
arch_sbm_cpu_instance_id() and assume all CPUs belong to the same
instance with the same instance id.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 drivers/base/arch_topology.c | 122 ++++++++++++++++++++++++++++++++++-
 1 file changed, 121 insertions(+), 1 deletion(-)

diff --git a/drivers/base/arch_topology.c b/drivers/base/arch_topology.c
index 8c5e47c28d9a..f55745a93298 100644
--- a/drivers/base/arch_topology.c
+++ b/drivers/base/arch_topology.c
@@ -20,6 +20,7 @@
 #include <linux/cpumask.h>
 #include <linux/init.h>
 #include <linux/rcupdate.h>
+#include <linux/sbm.h>
 #include <linux/sched.h>
 #include <linux/units.h>
 
@@ -930,6 +931,114 @@ __weak int __init parse_acpi_topology(void)
 	return 0;
 }
 
+/*
+ * Note: arch_sbm_cpu_instance_id() is only used if init_sbm_topology()
+ * below succeeds at setting the sbm topology. Either the LLC
+ * information is valid or package_id is dependable.
+ */
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	if (last_level_cache_is_valid(cpu)) {
+		struct cacheinfo *llc_info = get_cpu_cacheinfo_llc(cpu);
+
+		if (!llc_info)
+			goto out;
+
+		if (llc_info->attributes & CACHE_ID)
+			return llc_info->id;
+
+		/*
+		 * XXX: fw_token be truncated from cast
+		 * when the value is returned as an int.
+		 */
+		return (int)((long)llc_info->fw_token);
+	}
+out:
+	return cpu_topology[cpu].package_id;
+}
+
+static void __init init_sbm_topology(void)
+{
+	int num_sbm_instances = 0, max_threads_per_instance = -1;
+	bool has_cache = false, has_package = false;
+	cpumask_var_t unique_cpus;
+	struct xarray instances;
+	unsigned long cpu, *count;
+
+	/*
+	 * If the allocation fails, the sbm core will use
+	 * the default logic of splitting CPUs evenly in
+	 * BITS_PER_LONG chunk.
+	 */
+	if (!zalloc_cpumask_var(&unique_cpus, GFP_KERNEL))
+		return;
+
+	xa_init(&instances);
+
+	for_each_possible_cpu(cpu) {
+		bool found = false;
+		int unique_cpu;
+
+		for_each_cpu(unique_cpu, unique_cpus) {
+			/*
+			 * XXX: Assumes last_level_cache_is_valid() is uniformly true
+			 * across the entire system if it is true for one CPU.
+			 */
+			if (last_level_cache_is_valid(cpu)) {
+				has_cache = true;
+				if (last_level_cache_is_shared(cpu, unique_cpu)) {
+					found = true;
+					break;
+				}
+			} else {
+				/* Go by package_id if no LLC information is found. */
+				has_package = true;
+				if (cpu_topology[unique_cpu].package_id ==
+				    cpu_topology[cpu].package_id) {
+					found = true;
+					break;
+				}
+			}
+		}
+
+		/* XXX: CPUs should not suddenly switch IDs mid way. */
+		if (has_cache && has_package)
+			goto out;
+
+		if (!found) {
+			count = kzalloc_obj(*count);
+			if (!count)
+				goto out;
+
+			cpumask_set_cpu(cpu, unique_cpus);
+			*count += 1;
+
+			xa_store(&instances, cpu, count, GFP_KERNEL);
+			continue;
+		}
+
+		count = xa_load(&instances, unique_cpu);
+		if (!count)
+			goto out;
+
+		*count += 1;
+	}
+
+	xa_for_each(&instances, cpu, count) {
+		max_threads_per_instance = max_t(int, max_threads_per_instance, *count);
+		num_sbm_instances++;
+	}
+
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+out:
+	xa_for_each(&instances, cpu, count) {
+		xa_erase(&instances, cpu);
+		kfree(count);
+	}
+	xa_destroy(&instances);
+	free_cpumask_var(unique_cpus);
+}
+
 void __init init_cpu_topology(void)
 {
 	int cpu, ret;
@@ -954,8 +1063,19 @@ void __init init_cpu_topology(void)
 			continue;
 		else if (ret != -ENOENT)
 			pr_err("Early cacheinfo failed, ret = %d\n", ret);
-		return;
+		break;
 	}
+
+	/*
+	 * If fetch_cache_info() fails for first CPU,
+	 * init_cpu_sbm_topology() will use pacakge_id instead.
+	 *
+	 * Uniform cache topology is a necessary since implementation
+	 * assumes last_level_cache_is_valid() gives same result for
+	 * all possible CPUs
+	 */
+	if (!ret || (cpu == 0 && ret == -ENOENT))
+		init_sbm_topology();
 }
 
 void store_cpu_topology(unsigned int cpuid)
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
                   ` (10 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Huacai Chen
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui

Use the CPU _PXM (Node) relation for disabled CPUs too when parsing SRAT
table. This detail will be used for early topology parsing to count the
LLC domains and set sparsebitmask (sbm) properties correctly.

In addition to sbm enablement, this also helps mapping per-CPU areas
correctly for disabled CPUs by initializing the node relations via
set_cpuid_to_node() correctly and avoiding the "rr_node" based fallback
in smp_prepare_boot_cpu().

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 arch/loongarch/kernel/acpi.c | 20 ++++++++++++--------
 1 file changed, 12 insertions(+), 8 deletions(-)

diff --git a/arch/loongarch/kernel/acpi.c b/arch/loongarch/kernel/acpi.c
index cb454ee92b20..1f15f78713b7 100644
--- a/arch/loongarch/kernel/acpi.c
+++ b/arch/loongarch/kernel/acpi.c
@@ -311,8 +311,7 @@ acpi_numa_processor_affinity_init(struct acpi_srat_cpu_affinity *pa)
 		bad_srat();
 		return;
 	}
-	if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
-		return;
+
 	pxm = pa->proximity_domain_lo;
 	if (acpi_srat_revision >= 2) {
 		pxm |= (pa->proximity_domain_hi[0] << 8);
@@ -332,9 +331,12 @@ acpi_numa_processor_affinity_init(struct acpi_srat_cpu_affinity *pa)
 		return;
 	}
 
-	early_numa_add_cpu(pa->apic_id, node);
-
 	set_cpuid_to_node(pa->apic_id, node);
+
+	if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
+		return;
+
+	early_numa_add_cpu(pa->apic_id, node);
 	node_set(node, numa_nodes_parsed);
 	pr_info("SRAT: PXM %u -> CPU 0x%02x -> Node %u\n", pxm, pa->apic_id, node);
 }
@@ -350,8 +352,7 @@ acpi_numa_x2apic_affinity_init(struct acpi_srat_x2apic_cpu_affinity *pa)
 		bad_srat();
 		return;
 	}
-	if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
-		return;
+
 	pxm = pa->proximity_domain;
 	node = acpi_map_pxm_to_node(pxm);
 	if (node < 0) {
@@ -366,9 +367,12 @@ acpi_numa_x2apic_affinity_init(struct acpi_srat_x2apic_cpu_affinity *pa)
 		return;
 	}
 
-	early_numa_add_cpu(pa->apic_id, node);
-
 	set_cpuid_to_node(pa->apic_id, node);
+
+	if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
+		return;
+
+	early_numa_add_cpu(pa->apic_id, node);
 	node_set(node, numa_nodes_parsed);
 	pr_info("SRAT: PXM %u -> CPU 0x%02x -> Node %u\n", pxm, pa->apic_id, node);
 }
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (2 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
                   ` (9 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Huacai Chen
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui

Configure the sparsebitmask (sbm) topology for multi-node processors
once SRAT is parsed and cpu_to_node() mappings are stable.

Former commit enabled parsing node mappings for disabled CPUs during
SRAT parsing and every CPU will have a valid mapping before the sbm
topology parsing is attempted ensuring correct bounds for
max_threads_per_instance.

XXX: If _PXM mappings cannot be trusted for disabled CPUs, an alternate
logic needs to be added. See PowerPC enablement later in the series.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Only build tested with config from
https://lore.kernel.org/lkml/202609190905.OZTcO07l-lkp@intel.com/
---
 arch/loongarch/kernel/smp.c | 36 ++++++++++++++++++++++++++++++++++++
 1 file changed, 36 insertions(+)

diff --git a/arch/loongarch/kernel/smp.c b/arch/loongarch/kernel/smp.c
index d4b5d1b6bb01..d8155b26c4df 100644
--- a/arch/loongarch/kernel/smp.c
+++ b/arch/loongarch/kernel/smp.c
@@ -15,6 +15,7 @@
 #include <linux/interrupt.h>
 #include <linux/irq_work.h>
 #include <linux/profile.h>
+#include <linux/sbm.h>
 #include <linux/seq_file.h>
 #include <linux/smp.h>
 #include <linux/threads.h>
@@ -73,6 +74,10 @@ static cpumask_t cpu_llc_shared_setup_map;
 /* representing cpus for which core maps can be computed */
 static cpumask_t cpu_core_setup_map;
 
+/* sbm setup data - only needed during init */
+static cpumask_t cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
 struct secondary_data cpuboot_data;
 static DEFINE_PER_CPU(int, cpu_state);
 
@@ -362,10 +367,16 @@ void __init loongson_smp_setup(void)
 	pr_info("Detected %i available CPU(s)\n", loongson_sysconf.nr_cpus);
 }
 
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	return cpu_to_node(cpu);
+}
+
 void __init loongson_prepare_cpus(unsigned int max_cpus)
 {
 	int i = 0;
 	int threads_per_core = 0;
+	int num_sbm_instances, max_threads_per_instance = 1;
 
 	parse_acpi_topology();
 	cpu_data[0].global_id = cpu_logical_map(0);
@@ -388,6 +399,31 @@ void __init loongson_prepare_cpus(unsigned int max_cpus)
 
 	per_cpu(cpu_state, smp_processor_id()) = CPU_ONLINE;
 	cpu_smt_set_num_threads(threads_per_core, threads_per_core);
+
+	for_each_possible_cpu(i) {
+		bool found = false;
+		int j;
+
+		for_each_cpu(j, &cpu_sbm_setup_map) {
+			if (cpu_to_node(i) == cpu_to_node(j)) {
+				found = true;
+				break;
+			}
+		}
+
+		if (!found) {
+			cpumask_set_cpu(i, &cpu_sbm_setup_map);
+			__node_thread_count[i] = 1;
+			continue;
+		}
+
+		__node_thread_count[j] += 1;
+		max_threads_per_instance = max(max_threads_per_instance,
+					       __node_thread_count[j]);
+	}
+
+	num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
 }
 
 /*
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (3 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
                   ` (8 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Thomas Bogendoerfer, Huacai Chen, Jiaxun Yang
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak

For MIPS systems that support NUMA, configure the sparsebitmap (sbm)
topology.

The two systems that toggle NUMA - loongson64 and sgi-ip27, find the
number of nodes and the maximum threads present in an instance. The
data structures used to parse the topology are marked __initdata to
allow reclaim post init.

In both cases - loongson64 and sgi-ip27, loongson3_smp_setup() and
node_scan_cpus() sets node IDs for all possible CPUs ensuring
configure_sbm_topology() captures the full possible system.

arch_sbm_cpu_instance_id() returns the cpu_to_node() mapping to derive
the instance ID and group the CPUs sharing the node on the same sbm
instance (leaf).

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested on QEMU with loongson3_defconfig from
https://gitlab.postmarketos.org/royka1/linux/-/blob/apollo-7.1.0-r3/arch/mips/configs/loongson3_defconfig?ref_type=tags

QEMU cmdline:

  qemu-system-mips64el \
  -M loongson3-virt,accel=tcg \
  -smp sockets=2,cores=4 \
  -cpu Loongson-3A1000 \
  -kernel vmlinux \
  -append "console=ttyS0 root=/dev/ram" \
  -nographic
---
 arch/mips/include/asm/topology.h |  6 +++++
 arch/mips/kernel/topology.c      | 43 ++++++++++++++++++++++++++++++++
 arch/mips/loongson64/smp.c       |  3 +++
 arch/mips/sgi-ip27/ip27-smp.c    |  3 +++
 4 files changed, 55 insertions(+)

diff --git a/arch/mips/include/asm/topology.h b/arch/mips/include/asm/topology.h
index 5158c802eb65..7a57cbd59819 100644
--- a/arch/mips/include/asm/topology.h
+++ b/arch/mips/include/asm/topology.h
@@ -21,4 +21,10 @@ extern struct cpumask __cpu_primary_thread_mask;
 #define cpu_primary_thread_mask ((const struct cpumask *)&__cpu_primary_thread_mask)
 #endif
 
+#ifdef CONFIG_NUMA
+extern void configure_sbm_topology(void);
+#else /* !CONFIG_NUMA */
+static void __maybe_unused configure_sbm_topology(void) { }
+#endif /* CONFIG_NUMA */
+
 #endif /* __ASM_TOPOLOGY_H */
diff --git a/arch/mips/kernel/topology.c b/arch/mips/kernel/topology.c
index 9429d85a4703..729a811e8b99 100644
--- a/arch/mips/kernel/topology.c
+++ b/arch/mips/kernel/topology.c
@@ -5,9 +5,52 @@
 #include <linux/node.h>
 #include <linux/nodemask.h>
 #include <linux/percpu.h>
+#include <linux/sbm.h>
 
 static DEFINE_PER_CPU(struct cpu, cpu_devices);
 
+#ifdef CONFIG_NUMA
+/* sbm setup data - only needed during init */
+static struct cpumask cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	return cpu_to_node(cpu);
+}
+
+void __init configure_sbm_topology(void)
+{
+	int num_sbm_instances, max_threads_per_instance = 1;
+	int i;
+
+	for_each_possible_cpu(i) {
+		bool found = false;
+		int j;
+
+		for_each_cpu(j, &cpu_sbm_setup_map) {
+			if (cpu_to_node(i) == cpu_to_node(j)) {
+				found = true;
+				break;
+			}
+		}
+
+		if (!found) {
+			cpumask_set_cpu(i, &cpu_sbm_setup_map);
+			__node_thread_count[i] = 1;
+			continue;
+		}
+
+		__node_thread_count[j] += 1;
+		max_threads_per_instance = max(max_threads_per_instance,
+					       __node_thread_count[j]);
+	}
+
+	num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+#endif /* CONFIG_NUMA */
+
 static int __init topology_init(void)
 {
 	int i, ret;
diff --git a/arch/mips/loongson64/smp.c b/arch/mips/loongson64/smp.c
index e584299d0fde..22174ae2b574 100644
--- a/arch/mips/loongson64/smp.c
+++ b/arch/mips/loongson64/smp.c
@@ -17,6 +17,7 @@
 #include <asm/smp.h>
 #include <asm/time.h>
 #include <asm/tlbflush.h>
+#include <asm/topology.h>
 #include <asm/cacheflush.h>
 #include <loongson.h>
 #include <loongson_regs.h>
@@ -493,6 +494,8 @@ static void __init loongson3_smp_setup(void)
 	if (smp_group[0])
 		ipi_write_enable(0);
 
+	configure_sbm_topology();
+
 	cpu_set_core(&cpu_data[0],
 		     cpu_logical_map(0) % loongson_sysconf.cores_per_package);
 	cpu_data[0].package = cpu_logical_map(0) / loongson_sysconf.cores_per_package;
diff --git a/arch/mips/sgi-ip27/ip27-smp.c b/arch/mips/sgi-ip27/ip27-smp.c
index 62733e049570..244f8d987f13 100644
--- a/arch/mips/sgi-ip27/ip27-smp.c
+++ b/arch/mips/sgi-ip27/ip27-smp.c
@@ -15,6 +15,7 @@
 #include <asm/page.h>
 #include <asm/processor.h>
 #include <asm/ptrace.h>
+#include <asm/topology.h>
 #include <asm/sn/agent.h>
 #include <asm/sn/arch.h>
 #include <asm/sn/gda.h>
@@ -80,6 +81,8 @@ void cpu_node_probe(void)
 		highest = node_scan_cpus(nasid, highest);
 	}
 
+	configure_sbm_topology();
+
 	printk("Discovered %d cpus on %d nodes\n", highest + 1, num_online_nodes());
 }
 
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (4 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
                   ` (7 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Madhavan Srinivasan
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	Michael Ellerman, Nicholas Piggin, Christophe Leroy

Initialize sparsebitmap (sbm) topology based on the coregroup
information. Each coregorup gets its own sparsemask leaf.

pSeries and memory hotplug are interesting since a hotplug can place a
newly added CPU on any online node. This requires special care to allow
estimating bitmask size considering the worst case scenarios - each CPU
onlined is on a separate node, and this is a new N_CPU node.

Platform may enforce a stricter standards for the CPUs being online and
what nodes they can be mapped to but the current implementations makes
no assumptions and considers each CPU can be onlined on a unique node.

pSeries systems that can hotplug CPUs (detected using smp_ops) use the
NUMA topology instead for sbm initialization. arch_sbm_cpu_instance_id()
on these systems use cpu_to_node() mappings to match CPUs to sbm
instances.

XXX: This requires further optimizations to shorten sparsemask
traversals by keeping the number of leaf nodes to a minimum. If there
are nuances I'm not aware of, please reach out.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested on ppc64le_defconfig with:

  qemu-system-ppc64 \
  -M pseries \
  -cpu power10 \
  -smp sockets=2,cores=2,threads=4 \
  -m 10G -nographic \
  -kernel vmlinux \
  -append "root=/dev/ram sched_debug"

and also on ppce500 VM based on instructions in
https://www.qemu.org/docs/master/system/ppc/ppce500.html
---
 arch/powerpc/kernel/setup-common.c           | 88 ++++++++++++++++++++
 arch/powerpc/platforms/pseries/hotplug-cpu.c | 10 +++
 2 files changed, 98 insertions(+)

diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
index 4afaba19b586..4b57ad553172 100644
--- a/arch/powerpc/kernel/setup-common.c
+++ b/arch/powerpc/kernel/setup-common.c
@@ -11,6 +11,7 @@
 #include <linux/export.h>
 #include <linux/panic_notifier.h>
 #include <linux/string.h>
+#include <linux/sbm.h>
 #include <linux/sched.h>
 #include <linux/init.h>
 #include <linux/kernel.h>
@@ -602,6 +603,87 @@ static __init int add_pcspkr(void)
 device_initcall(add_pcspkr);
 #endif	/* CONFIG_PCSPKR_PLATFORM */
 
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	/*
+	 * In case of pSeries processors, sbm masks are
+	 * grouped by nodes where the cpuhotplug
+	 * operations can remove and re-add same logical
+	 * CPUs on different nodes.
+	 *
+	 * See comment in pseries_cpu_hotplug_init().
+	 */
+	if (smp_ops->cpu_disable)
+		return cpu_to_node(cpu);
+
+	return cpu_to_coregroup_id(cpu);
+}
+
+static void __init setup_sbm_topology(void)
+{
+	int num_sbm_instances, max_threads_per_instance = 1;
+	struct cpumask *cpu_sbm_setup_map;
+	int i, *__node_thread_count;
+	int disabled_cpus = 0;
+
+	cpu_sbm_setup_map = memblock_alloc_or_panic(cpumask_size(), __alignof__(long));
+	__node_thread_count = memblock_alloc_or_panic(nr_cpu_ids * sizeof(int),
+						      __alignof__(int));
+
+	memset(__node_thread_count, 0, nr_cpu_ids * sizeof(int));
+	memset(cpu_sbm_setup_map, 0, cpumask_size());
+
+	for_each_possible_cpu(i) {
+		bool found = false;
+		int j;
+
+		if (!cpu_present(i)) {
+			disabled_cpus += 1;
+			continue;
+		}
+
+		for_each_cpu(j, cpu_sbm_setup_map) {
+			if (cpu_to_coregroup_id(i) == cpu_to_coregroup_id(j)) {
+				found = true;
+				break;
+			}
+		}
+
+		if (!found) {
+			cpumask_set_cpu(i, cpu_sbm_setup_map);
+			__node_thread_count[i] = 1;
+			continue;
+		}
+
+		__node_thread_count[j] += 1;
+		max_threads_per_instance = max(max_threads_per_instance,
+					       __node_thread_count[j]);
+	}
+
+	/*
+	 * If CPUs are disabled, they may pop up on any online node.
+	 *
+	 * XXX: Any implementation nuances that can help this?
+	 * pSeries says only online nodes can be extended.
+	 */
+	if (disabled_cpus) {
+		num_sbm_instances = num_sbm_instances + disabled_cpus;
+	} else {
+		num_sbm_instances = cpumask_weight(cpu_sbm_setup_map);
+	}
+
+	/*
+	 * If disabled threads exists, assume the maximum threads per
+	 * instance can extend by the number of disabled threads if they
+	 * are all added to the same node.
+	 */
+	sbm_set_topology(num_sbm_instances,
+			 max_threads_per_instance + disabled_cpus);
+
+	memblock_free(__node_thread_count, nr_cpu_ids * sizeof(int));
+	memblock_free(cpu_sbm_setup_map, cpumask_size());
+}
+
 static char ppc_hw_desc_buf[128] __initdata;
 
 struct seq_buf ppc_hw_desc __initdata = {
@@ -1006,6 +1088,12 @@ void __init setup_arch(char **cmdline_p)
 
 	early_memtest(min_low_pfn << PAGE_SHIFT, max_low_pfn << PAGE_SHIFT);
 
+	/*
+	 * setup_arch() below can override topology for
+	 * pSeries platforms as a result of hotplug nuances.
+	 */
+	setup_sbm_topology();
+
 	if (ppc_md.setup_arch)
 		ppc_md.setup_arch();
 
diff --git a/arch/powerpc/platforms/pseries/hotplug-cpu.c b/arch/powerpc/platforms/pseries/hotplug-cpu.c
index bc6926dbf148..7c1c1ac3efde 100644
--- a/arch/powerpc/platforms/pseries/hotplug-cpu.c
+++ b/arch/powerpc/platforms/pseries/hotplug-cpu.c
@@ -19,6 +19,7 @@
 #include <linux/kernel.h>
 #include <linux/interrupt.h>
 #include <linux/delay.h>
+#include <linux/sbm.h>
 #include <linux/sched.h>	/* for idle_task_exit */
 #include <linux/sched/hotplug.h>
 #include <linux/cpu.h>
@@ -870,6 +871,15 @@ void __init pseries_cpu_hotplug_init(void)
 		return;
 	}
 
+	/*
+	 * find_cpu_id_range() only looks at online nodes.
+	 *
+	 * XXX: Is it possible for a CPU attached memory node to come
+	 * online after this point? May need num_possbile_nodes() then
+	 * unless there are platform nuances that can help optimize.
+	 */
+	sbm_set_topology(num_online_nodes(), num_possible_cpus());
+
 	smp_ops->cpu_offline_self = pseries_cpu_offline_self;
 	smp_ops->cpu_disable = pseries_cpu_disable;
 	smp_ops->cpu_die = pseries_cpu_die;
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early()
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (5 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
                   ` (6 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Heiko Carstens, Vasily Gorbik, Alexander Gordeev
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	Christian Borntraeger, Sven Schnelle

When the topology mode is TOPOLOGY_MODE_HW, topology_init_early()
already parses the set of socket present on the system.

Use the socket_info to configure the sparsebitmask (sbm) properties for
the system - namely the number of sockets and maximum threads in a
socket instance.

arch_sbm_cpu_instance_id() traverses all socket instances to find
a matching CPU which is not optimal but instance ID is only checked
when the CPU is coming online. Since this is a rare event, the
inefficiency is tolerable.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Build tested with LKP's s390 randconfig with QEMU cmdline:

  qemu-system-s390x \
  -cpu max \
  -smp 2 \
  -m 2048 \
  -kernel ./arch/s390/boot/vmlinux  \
  -append "console=ttyAMA0 earlycon ignore_loglevel log_buf_len=10M print_fatal_signals=1 LOGLEVEL=8 sched_debug" \
  -nographic

Note: Since ctop needs KVM acceleration to emulate TOPOLOGY_MODE_HW, the
actual spasemask setting is only build tested at the moment.

XXX: Any way around this using cross-compile and QEMU?
---
 arch/s390/kernel/topology.c | 52 +++++++++++++++++++++++++++++++++++++
 1 file changed, 52 insertions(+)

diff --git a/arch/s390/kernel/topology.c b/arch/s390/kernel/topology.c
index 1377c6f3f670..80dc51f5ce6f 100644
--- a/arch/s390/kernel/topology.c
+++ b/arch/s390/kernel/topology.c
@@ -20,6 +20,7 @@
 #include <linux/init.h>
 #include <linux/slab.h>
 #include <linux/cpu.h>
+#include <linux/sbm.h>
 #include <linux/smp.h>
 #include <linux/mm.h>
 #include <linux/nodemask.h>
@@ -561,6 +562,55 @@ static int __init detect_polarization(union topology_entry *tle)
 	return tl_core->pp != POLARIZATION_HRZ;
 }
 
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	struct mask_info *info = &socket_info;
+
+	/*
+	 * sbm core should not call this if sbm_set_topology() was
+	 * skipped below due to lack of topology information.
+	 */
+	if (WARN_ON_ONCE(topology_mode != TOPOLOGY_MODE_HW))
+		return -1;
+
+	while (info) {
+		if (cpumask_test_cpu(cpu, &info->mask))
+			return info->id;
+		info = info->next;
+	}
+
+	pr_warn_once("socket mapping for CPU%d not found! Mapping to ID 0", cpu);
+	return 0;
+}
+
+static void __configure_sbm_topology(void)
+{
+	int num_sbm_instances = 0, max_threads_per_instance = -1;
+	struct mask_info *info = &socket_info;
+
+	/*
+	 * Consider single package if topology
+	 * information is unavailable.
+	 *
+	 * sbm core will handle the initialization.
+	 */
+	if (topology_mode != TOPOLOGY_MODE_HW)
+		return;
+
+	while (info) {
+		num_sbm_instances += 1;
+		max_threads_per_instance = max_t(int,
+						 cpumask_weight(&info->mask),
+						 max_threads_per_instance);
+		info = info->next;
+	}
+
+	if (WARN_ON_ONCE(!num_sbm_instances || max_threads_per_instance < 1))
+		return;
+
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
 void __init topology_init_early(void)
 {
 	struct sysinfo_15_1_x *info;
@@ -588,6 +638,8 @@ void __init topology_init_early(void)
 	cpumask_set_cpu(0, &cpu_setup_mask);
 	__arch_update_cpu_topology();
 	__arch_update_dedicated_flag(NULL);
+
+	__configure_sbm_topology();
 }
 
 static inline int topology_get_mode(int enabled)
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (6 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
                   ` (5 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, David S. Miller, Andreas Larsson
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak

Once the CPU topology is parsed, initialize the sparsebitmap (sbm)
topology by counting the number of LLC / NUMA instances in the system
and the maximum number of threads in these instance.

Similar to alloc_irqstack_bootmem(), init_sbm_topology() is done after
setup_arch() has finished mapping the cache, socket, and NUMA IDs for
all the present CPUs.

arch_sbm_cpu_instance_id() uses a combination of cpu_to_node() and
max_cache_id to determine the sparsemask instances.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested with sparc64_defconfig with

  qemu-system-sparc64 \
  -M sun4u \
  -m 512 \
  -kernel vmlinux \
  -nographic -append "console=ttyS0"

Similar to some of my previous attempts, I was not able to boot a SMP
configuration with qemu-system-sparc64. For all purposes, sbm
initialization is only build tested.
---
 arch/sparc/kernel/setup_64.c | 45 ++++++++++++++++++++++++++++++++++++
 1 file changed, 45 insertions(+)

diff --git a/arch/sparc/kernel/setup_64.c b/arch/sparc/kernel/setup_64.c
index 63615f5c99b4..7ab52d17e46f 100644
--- a/arch/sparc/kernel/setup_64.c
+++ b/arch/sparc/kernel/setup_64.c
@@ -28,6 +28,7 @@
 #include <linux/root_dev.h>
 #include <linux/interrupt.h>
 #include <linux/cpu.h>
+#include <linux/sbm.h>
 #include <linux/initrd.h>
 #include <linux/module.h>
 #include <linux/start_kernel.h>
@@ -53,6 +54,7 @@
 #include <asm/cacheflush.h>
 #include <asm/dma.h>
 #include <asm/irq.h>
+#include <asm/cpudata.h>
 
 #ifdef CONFIG_IP_PNP
 #include <net/ipconfig.h>
@@ -619,6 +621,48 @@ static void __init alloc_irqstack_bootmem(void)
 	}
 }
 
+static struct cpumask cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	/* Combine the node and cache identifier. */
+	return ((u32)cpu_to_node(cpu) << 16) |
+	       cpu_data(cpu).max_cache_id;
+}
+
+static void __init init_sbm_topology(void)
+{
+	int num_sbm_instances, max_threads_per_instance = 1;
+	int i;
+
+	for_each_possible_cpu(i) {
+		bool found = false;
+		int j;
+
+		for_each_cpu(j, &cpu_sbm_setup_map) {
+			if (cpu_to_node(i) == cpu_to_node(j) &&
+			    cpu_data(i).max_cache_id == cpu_data(j).max_cache_id) {
+				found = true;
+				break;
+			}
+		}
+
+		if (!found) {
+			cpumask_set_cpu(i, &cpu_sbm_setup_map);
+			__node_thread_count[i] = 1;
+			continue;
+		}
+
+		__node_thread_count[j] += 1;
+		max_threads_per_instance = max(max_threads_per_instance,
+					       __node_thread_count[j]);
+	}
+
+	num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
 void __init setup_arch(char **cmdline_p)
 {
 	/* Initialize PROM console and command line. */
@@ -677,6 +721,7 @@ void __init setup_arch(char **cmdline_p)
 	 * allocate the IRQ stacks.
 	 */
 	alloc_irqstack_bootmem();
+	init_sbm_topology();
 }
 
 extern int stop_a_enabled;
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (7 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-03  8:27   ` Chen Yu
  2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
                   ` (4 subsequent siblings)
  13 siblings, 1 reply; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	H. Peter Anvin

From: Peter Zijlstra <peterz@infradead.org>

Once the topology is parsed, the APICIDs of all present CPUs and topology
shifts are available for use.

Use the max APICID encountered during traversal and TOPO_DIE_DOMAIN /
TOPO_TILE_DOMAIN shift to determine the maximum number of CPUs that can
be associated with a single LLC instance and the maximum number of LLC
instances that can exist on a the processor to initialize sparsebitmap
(sbm) topology.

On AMD heterogeneous processors, the Cache Identifier leaf CPUID
0x8000001d field "NumSharingCache" give true value of threads sharing
the cache instance and not the number of bits reserved for a cache
instance in APICID.

To avoid incorrect calculations, shifts from extended topology leaf
0x80000026 on AMD and Hygon systems are used to derive the domain
shifts.

  [ tim.c.chen: TOPO_DIE_DOMAIN / TOPO_TILE_DOMAIN based
    initialization. ]

  [ kprateek: Commit message. Adapting to new sbm helper. ]

(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
(Not-yet-)Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 arch/x86/kernel/cpu/topology.c | 40 +++++++++++++++++++++++++++++++++-
 1 file changed, 39 insertions(+), 1 deletion(-)

diff --git a/arch/x86/kernel/cpu/topology.c b/arch/x86/kernel/cpu/topology.c
index 4913b64ec592..150e098f0ef8 100644
--- a/arch/x86/kernel/cpu/topology.c
+++ b/arch/x86/kernel/cpu/topology.c
@@ -23,6 +23,7 @@
  */
 #define pr_fmt(fmt) "CPU topo: " fmt
 #include <linux/cpu.h>
+#include <linux/sbm.h>
 
 #include <xen/xen.h>
 
@@ -449,13 +450,45 @@ static __init bool restrict_to_up(void)
 	return apic_is_disabled;
 }
 
+int arch_sbm_cpu_instance_id(int cpu)
+{
+	u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
+	u32 apicid = cpuid_to_apicid[cpu];
+
+	if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+	    boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+		sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
+
+	return (apicid >> sbm_shift);
+}
+
+static __init void init_sbm_topology(u32 max_apicid)
+{
+	u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
+	int num_sbm_instances, max_threads_per_instance;
+
+	/*
+	 * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
+	 * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
+	 * that instead.
+	 */
+	if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+	    boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+		sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
+
+	num_sbm_instances = 1 + (max_apicid >> sbm_shift);
+	max_threads_per_instance = (1 << sbm_shift);
+
+	sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
 void __init topology_init_possible_cpus(void)
 {
 	unsigned int assigned = topo_info.nr_assigned_cpus;
 	unsigned int disabled = topo_info.nr_disabled_cpus;
 	unsigned int cnta, cntb, cpu, allowed = 1;
 	unsigned int total = assigned + disabled;
-	u32 apicid, firstid;
+	u32 apicid, firstid, maxid;
 
 	/*
 	 * If there was no APIC registered, then fake one so that the
@@ -536,10 +569,13 @@ void __init topology_init_possible_cpus(void)
 		cpuid_to_apicid[topo_info.nr_assigned_cpus++] = apicid;
 	}
 
+	maxid = topo_info.boot_cpu_apic_id;
+
 	for (cpu = 0; cpu < allowed; cpu++) {
 		apicid = cpuid_to_apicid[cpu];
 
 		set_cpu_possible(cpu, true);
+		maxid = max(maxid, apicid);
 
 		if (apicid == BAD_APICID)
 			continue;
@@ -547,6 +583,8 @@ void __init topology_init_possible_cpus(void)
 		cpu_mark_primary_thread(cpu, apicid);
 		set_cpu_present(cpu, test_bit(apicid, phys_cpu_present_map));
 	}
+
+	init_sbm_topology(maxid);
 }
 
 /*
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (8 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
                   ` (3 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui, Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

Add infrastructure to establish CPU to sparsebitmap (sbm) index relation
before CPU is turned active. CPU coming online looks for a free slot
based on its instance ID and acquires a free slot.

If slots are exhausted, a new leaf is allocated for an instance ID. New
CPUs activating with same instance ID can claim free slots on the same
leaf but must never exceed max_threads_per_instance.

For architectures that have not initialized the sbm topology, sbm core
overrides the arch_sbm_cpu_instance_id() in sbm_cpu_instance_id()
wrapper to always return 0 keeping all CPUs on same instance.

Data structures that require fast access have been runtime constified to
enable faster access in kernel hot paths.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 include/asm-generic/vmlinux.lds.h |   6 +-
 include/linux/sbm.h               |  14 +++
 init/main.c                       |   6 +
 kernel/sched/core.c               |  17 +++
 lib/sbm.c                         | 175 +++++++++++++++++++++++++++++-
 5 files changed, 215 insertions(+), 3 deletions(-)

diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index b2988aa12f66..b0346d382cea 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -981,7 +981,11 @@
 		RUNTIME_CONST(ptr, __bfilp_cache)			\
 		RUNTIME_CONST(shift, __futex_shift)			\
 		RUNTIME_CONST(mask,  __futex_mask)			\
-		RUNTIME_CONST(ptr,   __futex_queues)
+		RUNTIME_CONST(ptr,   __futex_queues)			\
+		RUNTIME_CONST(shift, __sbm_shift)			\
+		RUNTIME_CONST(mask,  __sbm_mask)			\
+		RUNTIME_CONST(ptr,   __sbm_cpu_to_idx)			\
+		RUNTIME_CONST(ptr,   __sbm_idx_to_cpu)
 
 /* Alignment must be consistent with (kunit_suite *) in include/kunit/test.h */
 #define KUNIT_TABLE()							\
diff --git a/include/linux/sbm.h b/include/linux/sbm.h
index adac12ed233a..232b0076bb3f 100644
--- a/include/linux/sbm.h
+++ b/include/linux/sbm.h
@@ -2,7 +2,21 @@
 #ifndef _LINUX_SBM_H
 #define _LINUX_SBM_H
 
+/*
+ * Masks and shifts for sbm index to translate
+ * a sbm leaf to CPU.
+ */
+extern int __sbm_shift;
+extern int __sbm_mask;
+
 int arch_sbm_cpu_instance_id(int cpu);
 void sbm_set_topology(int num_instances, int max_threads_per_instance);
 
+int sbm_cpu_to_idx(int cpu);
+int sbm_idx_to_cpu(int idx);
+
+int alloc_sbm_index(int cpu);
+void free_sbm_index(int cpu);
+int sbm_init(void);
+
 #endif /* _LINUX_SBM_H */
diff --git a/init/main.c b/init/main.c
index 2613d3f9b3ce..b7406bd3acc8 100644
--- a/init/main.c
+++ b/init/main.c
@@ -72,6 +72,7 @@
 #include <linux/pid_namespace.h>
 #include <linux/device/driver.h>
 #include <linux/kthread.h>
+#include <linux/sbm.h>
 #include <linux/sched.h>
 #include <linux/sched/init.h>
 #include <linux/signal.h>
@@ -1652,6 +1653,11 @@ static noinline void __init kernel_init_freeable(void)
 
 	smp_prepare_cpus(setup_max_cpus);
 
+	sbm_init();
+
+	/* Finish initializing boot CPU since it is already active. */
+	alloc_sbm_index(smp_processor_id());
+
 	workqueue_init();
 
 	init_mm_internals();
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0bb86a43a592..977f579da410 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -61,6 +61,7 @@
 #include <linux/rcuwait_api.h>
 #include <linux/rseq.h>
 #include <linux/sched/wake_q.h>
+#include <linux/sbm.h>
 #include <linux/scs.h>
 #include <linux/slab.h>
 #include <linux/syscalls.h>
@@ -8636,6 +8637,14 @@ int sched_cpu_activate(unsigned int cpu)
 	 */
 	balance_push_set(cpu, false);
 
+	alloc_sbm_index(cpu);
+
+	/*
+	 * Make sure sbm mappings are visible
+	 * before CPU is toggled active.
+	 */
+	smp_mb();
+
 	/*
 	 * When going up, increment the number of cores with SMT present.
 	 */
@@ -8688,6 +8697,14 @@ int sched_cpu_deactivate(unsigned int cpu)
 
 	set_cpu_active(cpu, false);
 
+	/*
+	 * Make sure CPU is inactive before
+	 * sbm indices are reclaimed.
+	 */
+	smp_mb();
+
+	free_sbm_index(cpu);
+
 	/*
 	 * From this point forward, this CPU will refuse to run any task that
 	 * is not: migrate_disable() or KTHREAD_IS_PER_CPU, and will actively
diff --git a/lib/sbm.c b/lib/sbm.c
index 82280e3306de..e5b0508b6825 100644
--- a/lib/sbm.c
+++ b/lib/sbm.c
@@ -1,10 +1,58 @@
 /* SPDX-License-Identifier: GPL-2.0 */
 #include <linux/sbm.h>
 #include <linux/init.h>
+#include <linux/log2.h>
+#include <linux/slab.h>
+#include <linux/cache.h>
 #include <linux/printk.h>
+#include <linux/cpumask.h>
+#include <linux/jump_label.h>
 
-static int sbm_max_threads_per_instance = -1;
-static int sbm_num_instance = -1;
+#include <asm/runtime-const.h>
+
+static int sbm_max_threads_per_instance __ro_after_init = -1;
+static int sbm_num_instance __ro_after_init = -1;
+
+int __sbm_shift __ro_after_init;
+int __sbm_mask __ro_after_init;
+
+static struct {
+	int		instance_id;	/* Instance ID linked to the leaf. */
+	unsigned long	allocated_mask;	/* Set of IDs that have been allocated. */
+} *__sbm_idx_metadata __ro_after_init;
+
+/* Translations between cpu <-> sbm leaf */
+static int *__sbm_cpu_to_idx __ro_after_init;
+static int *__sbm_idx_to_cpu __ro_after_init;
+
+static __always_inline int *_sbm_cpu_to_idx(void)
+{
+	return runtime_const_ptr(__sbm_cpu_to_idx);
+}
+
+static __always_inline int *_sbm_idx_to_cpu(void)
+{
+	return runtime_const_ptr(__sbm_idx_to_cpu);
+}
+
+int sbm_cpu_to_idx(int cpu)
+{
+	return _sbm_cpu_to_idx()[cpu];
+}
+
+int sbm_idx_to_cpu(int idx)
+{
+	return _sbm_idx_to_cpu()[idx];
+}
+
+/*
+ * Certain architectures may skip initializing sbm propoerties
+ * while having an arch_sbm_cpu_instance_id() definition.
+ *
+ * In such cases, don't trust the arch/ side redefine and use
+ * the default single instance mapping.
+ */
+static DEFINE_STATIC_KEY_FALSE(sbm_arch_initialized);
 
 /*
  * In absence of an arch definition, consider all CPUs to
@@ -15,6 +63,72 @@ int __weak arch_sbm_cpu_instance_id(int cpu)
 	return 0;
 }
 
+static int sbm_cpu_to_instance(int cpu)
+{
+	if (static_branch_likely(&sbm_arch_initialized))
+		return arch_sbm_cpu_instance_id(cpu);
+
+	return 0;
+}
+
+int alloc_sbm_index(int cpu)
+{
+	int cpu_instance = sbm_cpu_to_instance(cpu);
+	int i, idx = BITS_PER_LONG, free_index = -1;
+
+	for (i = 0; i < sbm_num_instance; ++i) {
+		if (__sbm_idx_metadata[i].instance_id == cpu_instance) {
+			idx = find_first_zero_bit(&__sbm_idx_metadata[i].allocated_mask,
+						  BITS_PER_LONG);
+
+			if (idx < BITS_PER_LONG)
+				break;
+		}
+		if (free_index == -1 && __sbm_idx_metadata[i].instance_id == -1)
+			free_index = i;
+	}
+
+	if (i == sbm_num_instance && free_index == -1)
+		return -ENOENT;
+
+	if (i == sbm_num_instance) {
+		__sbm_idx_metadata[free_index].instance_id = cpu_instance;
+		i = free_index;
+		idx = 0;
+	}
+
+	WARN_ON_ONCE(idx >= sbm_max_threads_per_instance);
+
+	__set_bit(idx, &__sbm_idx_metadata[i].allocated_mask);
+
+	idx = (i << __sbm_shift) + idx;
+	_sbm_idx_to_cpu()[idx] = cpu;
+	_sbm_cpu_to_idx()[cpu] = idx;
+
+	return 0;
+}
+
+void free_sbm_index(int cpu)
+{
+	int idx = sbm_cpu_to_idx(cpu);
+	u32 leaf;
+
+	if (idx < 0)
+		return;
+
+	_sbm_idx_to_cpu()[idx] = -1;
+	_sbm_cpu_to_idx()[cpu] = -1;
+
+	leaf = runtime_const_shift_right_32(idx, __sbm_shift);
+	idx = runtime_const_mask_32(idx, __sbm_mask);
+
+	__clear_bit(idx, &__sbm_idx_metadata[leaf].allocated_mask);
+
+	if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
+	    BITS_PER_LONG)
+		__sbm_idx_metadata[leaf].instance_id = -1;
+}
+
 void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
 {
 	sbm_max_threads_per_instance = max_threads_per_instance;
@@ -25,3 +139,60 @@ void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
 		sbm_max_threads_per_instance);
 }
 
+int __init sbm_init(void)
+{
+	int i;
+
+	if (sbm_max_threads_per_instance > 0 && sbm_num_instance > 0) {
+		static_branch_enable(&sbm_arch_initialized);
+		goto init_properties;
+	}
+
+	sbm_max_threads_per_instance = BITS_PER_LONG;
+	sbm_num_instance = (nr_cpumask_bits / BITS_PER_LONG) + 1;
+
+init_properties:
+	/*
+	 * If the number of CPUs per instance cross bitmask word boundary,
+	 * split the instances into samller chunks on BITS_PER_LONG and
+	 * increase the order of leaves.
+	 */
+	if (sbm_max_threads_per_instance > BITS_PER_LONG) {
+		int split = (sbm_max_threads_per_instance + BITS_PER_LONG - 1) / BITS_PER_LONG;
+
+		sbm_max_threads_per_instance = BITS_PER_LONG;
+		sbm_num_instance *= split;
+	}
+
+	sbm_max_threads_per_instance = roundup_pow_of_two(sbm_max_threads_per_instance);
+
+	__sbm_shift = ilog2(sbm_max_threads_per_instance);
+	__sbm_mask = sbm_max_threads_per_instance - 1;
+
+	__sbm_idx_metadata = kzalloc_objs(*__sbm_idx_metadata, sbm_num_instance);
+	__sbm_cpu_to_idx = kzalloc_objs(*__sbm_cpu_to_idx, nr_cpumask_bits);
+	__sbm_idx_to_cpu = kzalloc_objs(*__sbm_idx_to_cpu,
+					sbm_max_threads_per_instance * sbm_num_instance);
+
+	BUG_ON(!__sbm_cpu_to_idx || !__sbm_idx_to_cpu || !__sbm_idx_metadata);
+
+	for (i = 0; i < nr_cpumask_bits; ++i)
+		__sbm_cpu_to_idx[i] = -1;
+
+	/* Set all instance id to -1 to allow for future allocations to claim them. */
+	for (i = 0; i < sbm_num_instance; ++i)
+		__sbm_idx_metadata[i].instance_id = -1;
+
+	runtime_const_init(shift, __sbm_shift);
+	runtime_const_init(mask,  __sbm_mask);
+	runtime_const_init(ptr,   __sbm_cpu_to_idx);
+	runtime_const_init(ptr,   __sbm_idx_to_cpu);
+
+	barrier();
+
+	pr_info("sbm instance count: %d (maximum threads per instance: %d)\n",
+		sbm_num_instance,
+		sbm_max_threads_per_instance);
+
+	return 0;
+}
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (9 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
                   ` (2 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
	WANG Xuerui, Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

From: Peter Zijlstra <peterz@infradead.org>

Introduce helpers to allocate a sparsebitmap (sbm) of arch configured
length, set a bit on the sbm, clear a bit from the sbm, and iterate all
the set indices on a sbm structure.

  [ yu.c.chen: Fixes for sbm implementation. ]
  [ kprateek: Adapting sbm implementation to a flat array implementation. ]

(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
(Not-yet-)Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 include/linux/sbm.h | 85 +++++++++++++++++++++++++++++++++++++++++++++
 lib/sbm.c           | 55 +++++++++++++++++++++++++++++
 2 files changed, 140 insertions(+)

diff --git a/include/linux/sbm.h b/include/linux/sbm.h
index 232b0076bb3f..63b116e52e6c 100644
--- a/include/linux/sbm.h
+++ b/include/linux/sbm.h
@@ -2,6 +2,8 @@
 #ifndef _LINUX_SBM_H
 #define _LINUX_SBM_H
 
+#include <linux/bitmap.h>
+
 /*
  * Masks and shifts for sbm index to translate
  * a sbm leaf to CPU.
@@ -9,12 +11,95 @@
 extern int __sbm_shift;
 extern int __sbm_mask;
 
+struct sbm {
+	unsigned long   bitmap;
+} ____cacheline_aligned;
+
 int arch_sbm_cpu_instance_id(int cpu);
 void sbm_set_topology(int num_instances, int max_threads_per_instance);
 
 int sbm_cpu_to_idx(int cpu);
 int sbm_idx_to_cpu(int idx);
 
+struct sbm *sbm_alloc(void);
+bool sbm_empty(struct sbm *sbm);
+int sbm_find_next_bit(struct sbm *sbm, int start);
+
+#define __sbm_op(sbm, func)				\
+({							\
+	int idx = sbm_cpu_to_idx(cpu);			\
+	int nr = idx >> __sbm_shift;			\
+	int bit = idx & __sbm_mask;			\
+							\
+	func(bit, &sbm[nr].bitmap);			\
+})
+
+static inline void sbm_cpu_set(struct sbm *sbm, int cpu)
+{
+	__sbm_op(sbm, set_bit);
+}
+
+static inline void sbm_cpu_clear(struct sbm *sbm, int cpu)
+{
+	__sbm_op(sbm, clear_bit);
+}
+
+static inline void __sbm_cpu_set(struct sbm *sbm, int cpu)
+{
+	__sbm_op(sbm, __set_bit);
+}
+
+static inline void __sbm_cpu_clear(struct sbm *sbm, int cpu)
+{
+	__sbm_op(sbm, __clear_bit);
+}
+
+static inline bool sbm_cpu_test(struct sbm *sbm, int cpu)
+{
+	return __sbm_op(sbm, test_bit);
+}
+
+static __always_inline
+unsigned int sbm_find_next_bit_wrap(struct sbm *sbm, int start)
+{
+	int bit = sbm_find_next_bit(sbm, start);
+
+	if (bit >= 0 || start == 0)
+		return bit;
+
+	bit = sbm_find_next_bit(sbm, 0);
+	return bit < start ? bit : -1;
+}
+
+static __always_inline
+unsigned int __sbm_for_each_wrap(struct sbm *sbm, int start, int n)
+{
+	int bit;
+
+	/* If not wrapped around */
+	if (n > start) {
+		/* and have a bit, just return it. */
+		bit = sbm_find_next_bit(sbm, n);
+		if (bit >= 0)
+			return bit;
+
+		/* Otherwise, wrap around and ... */
+		n = 0;
+	}
+
+	/* Search the other part. */
+	bit = sbm_find_next_bit(sbm, n);
+	return bit < start ? bit : -1;
+}
+
+#define sbm_for_each_set_bit(sbm, idx) \
+	for (int idx = sbm_find_next_bit(sbm, 0); \
+	     idx >= 0; idx = sbm_find_next_bit(sbm, idx+1))
+
+#define sbm_for_each_set_bit_wrap(sbm, idx, start) \
+	for (int idx = sbm_find_next_bit_wrap(sbm, start); \
+	     idx >= 0; idx = __sbm_for_each_wrap(sbm, start, idx+1))
+
 int alloc_sbm_index(int cpu);
 void free_sbm_index(int cpu);
 int sbm_init(void);
diff --git a/lib/sbm.c b/lib/sbm.c
index e5b0508b6825..eeca5ce06d50 100644
--- a/lib/sbm.c
+++ b/lib/sbm.c
@@ -12,6 +12,7 @@
 
 static int sbm_max_threads_per_instance __ro_after_init = -1;
 static int sbm_num_instance __ro_after_init = -1;
+static int sbm_max_populated_index;
 
 int __sbm_shift __ro_after_init;
 int __sbm_mask __ro_after_init;
@@ -35,6 +36,11 @@ static __always_inline int *_sbm_idx_to_cpu(void)
 	return runtime_const_ptr(__sbm_idx_to_cpu);
 }
 
+static int sbm_max_index(void)
+{
+	return READ_ONCE(sbm_max_populated_index);
+}
+
 int sbm_cpu_to_idx(int cpu)
 {
 	return _sbm_cpu_to_idx()[cpu];
@@ -45,6 +51,44 @@ int sbm_idx_to_cpu(int idx)
 	return _sbm_idx_to_cpu()[idx];
 }
 
+struct sbm *sbm_alloc(void)
+{
+	return kzalloc_objs(struct sbm, sbm_max_threads_per_instance * sbm_num_instance);
+}
+
+bool sbm_empty(struct sbm *sbm)
+{
+	int i;
+
+	for (i = 0; i <= sbm_max_index(); ++i) {
+		if (sbm[i].bitmap)
+			return false;
+	}
+
+	return true;
+}
+
+int sbm_find_next_bit(struct sbm *sbm, int start)
+{
+	u32 nr = runtime_const_shift_right_32(start, __sbm_shift);
+	u32 bit = runtime_const_mask_32(start, __sbm_mask);
+	unsigned long tmp = 0, mask = (~0UL) << bit;
+
+	for (; nr <= sbm_max_index(); nr++) {
+		tmp = sbm[nr].bitmap & mask;
+		if (tmp)
+			break;
+		/*
+		 * Consider full bitmask from
+		 * second iteration.
+		 */
+		mask = ~0UL;
+	}
+	if (!tmp)
+		return -1;
+	return (nr << __sbm_shift) | __ffs(tmp);
+}
+
 /*
  * Certain architectures may skip initializing sbm propoerties
  * while having an arch_sbm_cpu_instance_id() definition.
@@ -105,6 +149,8 @@ int alloc_sbm_index(int cpu)
 	_sbm_idx_to_cpu()[idx] = cpu;
 	_sbm_cpu_to_idx()[cpu] = idx;
 
+	WRITE_ONCE(sbm_max_populated_index, max(sbm_max_populated_index, i));
+
 	return 0;
 }
 
@@ -127,6 +173,15 @@ void free_sbm_index(int cpu)
 	if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
 	    BITS_PER_LONG)
 		__sbm_idx_metadata[leaf].instance_id = -1;
+
+	if (leaf == sbm_max_populated_index) {
+		for (idx = leaf - 1; idx > -1; idx--) {
+			if (__sbm_idx_metadata[idx].instance_id != -1)
+				break;
+		}
+
+		WRITE_ONCE(sbm_max_populated_index, max(idx, 0));
+	}
 }
 
 void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp()
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (10 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
  2026-10-03  9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak

nohz.idle_cpus_mask is not required soon after early initialization
since the boot CPU cannot go tickless idle before CPU is activated and
tick is started.

Defer allocation of nohz.idle_cpus_mask to sched_init_smp() via
init_sched_fair_class_smp(). This is preparation to convert
nohz.idle_cpus_mask to sparsebitmap (sbm) which requires architectures
to finish smp_prepare before initializing the core bits.

No functional changes intended.

Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 kernel/sched/core.c  | 2 ++
 kernel/sched/fair.c  | 8 +++++++-
 kernel/sched/sched.h | 1 +
 3 files changed, 10 insertions(+), 1 deletion(-)

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 977f579da410..127aed81d316 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -8849,6 +8849,8 @@ int sched_cpu_dying(unsigned int cpu)
 
 void __init sched_init_smp(void)
 {
+	init_sched_fair_class_smp();
+
 	sched_init_numa(NUMA_NO_NODE);
 
 	prandom_init_once(&sched_rnd_state);
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 1c687c3c70f1..a0a659f4c3be 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -15645,6 +15645,13 @@ void show_numa_stats(struct task_struct *p, struct seq_file *m)
 }
 #endif /* CONFIG_NUMA_BALANCING */
 
+__init void init_sched_fair_class_smp(void)
+{
+#ifdef CONFIG_NO_HZ_COMMON
+	zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
+#endif
+}
+
 __init void init_sched_fair_class(void)
 {
 	int i;
@@ -15666,6 +15673,5 @@ __init void init_sched_fair_class(void)
 #ifdef CONFIG_NO_HZ_COMMON
 	nohz.next_balance = jiffies;
 	nohz.next_blocked = jiffies;
-	zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
 #endif
 }
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 7d2ec527b8a2..0726e8fb5967 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -3024,6 +3024,7 @@ extern void update_max_interval(void);
 extern void init_sched_dl_class(void);
 extern void init_sched_rt_class(void);
 extern void init_sched_fair_class(void);
+extern void init_sched_fair_class_smp(void);
 
 extern void resched_curr(struct rq *rq);
 extern void resched_curr_lazy(struct rq *rq);
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (11 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
  2026-10-03  9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
  13 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
  To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, K Prateek Nayak

From: Peter Zijlstra <peterz@infradead.org>

With sbm infrastructure in place, convert the global nohz.idle_cpus
cpumask to use sparsebitmap (sbm).

  [ prateek: Used sbm_for_each_bit_wrap(), and adapted find_new_ilb() to
    the current SMT aware scheme. ]

(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
 kernel/sched/fair.c | 66 +++++++++++++++++++--------------------------
 1 file changed, 27 insertions(+), 39 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index a0a659f4c3be..0b1458c4360e 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -51,8 +51,9 @@
 #include <linux/profile.h>
 #include <linux/psi.h>
 #include <linux/ratelimit.h>
-#include <linux/task_work.h>
 #include <linux/rbtree_augmented.h>
+#include <linux/sbm.h>
+#include <linux/task_work.h>
 
 #include <asm/switch_to.h>
 
@@ -8210,7 +8211,7 @@ static DEFINE_PER_CPU(cpumask_var_t, should_we_balance_tmpmask);
 #ifdef CONFIG_NO_HZ_COMMON
 
 static struct {
-	cpumask_var_t idle_cpus_mask;
+	struct sbm *sbm;
 	int has_blocked_load;		/* Idle CPUS has blocked load */
 	int needs_update;		/* Newly idle CPUs need their next_balance collated */
 	unsigned long next_balance;     /* in jiffy units */
@@ -14030,7 +14031,7 @@ static inline int on_null_domain(struct rq *rq)
  */
 static inline int find_new_ilb(void)
 {
-	struct cpumask *ilb_cpus;
+	struct cpumask *skip_cpus;
 	int ilb_cpu, fallback = -1;
 
 	lockdep_assert_irqs_disabled();
@@ -14039,18 +14040,22 @@ static inline int find_new_ilb(void)
 	 * Reuse the per-CPU select_rq_mask, which is protected from concurrent
 	 * use on this CPU by having interrupts disabled.
 	 */
-	ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
-	cpumask_and(ilb_cpus, nohz.idle_cpus_mask,
-		    housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+	skip_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
+	cpumask_clear(skip_cpus);
+
+	sbm_for_each_set_bit(nohz.sbm, idx) {
+		ilb_cpu = sbm_idx_to_cpu(idx);
+
+		if (cpumask_test_cpu(ilb_cpu, skip_cpus))
+			continue;
 
-	for_each_cpu(ilb_cpu, ilb_cpus) {
 		if (!idle_cpu(ilb_cpu)) {
 			/*
 			 * Once an idle fallback exists, a busy CPU proves that
 			 * this core cannot be fully idle. Skip its siblings.
 			 */
 			if (sched_smt_active() && fallback >= 0)
-				cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu));
+				cpumask_or(skip_cpus, skip_cpus, cpu_smt_mask(ilb_cpu));
 			continue;
 		}
 
@@ -14069,8 +14074,7 @@ static inline int find_new_ilb(void)
 			 * The core is not idle, so there is no need to check
 			 * any of its other SMT siblings.
 			 */
-			cpumask_andnot(ilb_cpus, ilb_cpus,
-				       cpu_smt_mask(ilb_cpu));
+			cpumask_or(skip_cpus, skip_cpus, cpu_smt_mask(ilb_cpu));
 			continue;
 		}
 
@@ -14134,7 +14138,7 @@ static void nohz_balancer_kick(struct rq *rq)
 	unsigned long now = jiffies;
 	struct sched_domain_shared *sds;
 	struct sched_domain *sd;
-	int nr_busy, i, cpu = rq->cpu;
+	int nr_busy, cpu = rq->cpu;
 	unsigned int flags = 0;
 
 	if (unlikely(rq->idle_balance))
@@ -14162,11 +14166,7 @@ static void nohz_balancer_kick(struct rq *rq)
 	if (time_before(now, nohz.next_balance))
 		goto out;
 
-	/*
-	 * None are in tickless mode and hence no need for NOHZ idle load
-	 * balancing
-	 */
-	if (unlikely(cpumask_empty(nohz.idle_cpus_mask)))
+	if (unlikely(sbm_empty(nohz.sbm)))
 		return;
 
 	if (rq->nr_running >= 2) {
@@ -14186,24 +14186,6 @@ static void nohz_balancer_kick(struct rq *rq)
 		}
 	}
 
-	sd = rcu_dereference_all(per_cpu(sd_asym_packing, cpu));
-	if (sd) {
-		/*
-		 * When ASYM_PACKING; see if there's a more preferred CPU
-		 * currently idle; in which case, kick the ILB to move tasks
-		 * around.
-		 *
-		 * When balancing between cores, all the SMT siblings of the
-		 * preferred CPU must be idle.
-		 */
-		for_each_cpu_and(i, sched_domain_span(sd), nohz.idle_cpus_mask) {
-			if (sched_asym(sd, i, cpu)) {
-				flags |= NOHZ_STATS_KICK | NOHZ_BALANCE_KICK;
-				goto out;
-			}
-		}
-	}
-
 	sd = rcu_dereference_all(per_cpu(sd_asym_cpucapacity, cpu));
 	if (sd) {
 		/*
@@ -14270,7 +14252,8 @@ void nohz_balance_exit_idle(struct rq *rq)
 		return;
 
 	rq->nohz_tick_stopped = 0;
-	cpumask_clear_cpu(rq->cpu, nohz.idle_cpus_mask);
+	if (cpumask_test_cpu(rq->cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)))
+		sbm_cpu_clear(nohz.sbm, rq->cpu);
 
 	set_cpu_sd_state_busy(rq->cpu);
 }
@@ -14324,7 +14307,8 @@ void nohz_balance_enter_idle(int cpu)
 
 	rq->nohz_tick_stopped = 1;
 
-	cpumask_set_cpu(cpu, nohz.idle_cpus_mask);
+	if (cpumask_test_cpu(rq->cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)))
+		sbm_cpu_set(nohz.sbm, rq->cpu);
 
 	/*
 	 * Ensures that if nohz_idle_balance() fails to observe our
@@ -14351,7 +14335,7 @@ static bool update_nohz_stats(struct rq *rq)
 	if (!rq->has_blocked_load)
 		return false;
 
-	if (!cpumask_test_cpu(cpu, nohz.idle_cpus_mask))
+	if (!sbm_cpu_test(nohz.sbm, cpu))
 		return false;
 
 	if (!time_after(jiffies, READ_ONCE(rq->last_blocked_load_update_tick)))
@@ -14377,6 +14361,7 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
 	int this_cpu = this_rq->cpu;
 	int balance_cpu;
 	struct rq *rq;
+	int start;
 
 	WARN_ON_ONCE((flags & NOHZ_KICK_MASK) == NOHZ_BALANCE_KICK);
 
@@ -14405,7 +14390,10 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
 	 * Start with the next CPU after this_cpu so we will end with this_cpu and let a
 	 * chance for other idle cpu to pull load.
 	 */
-	for_each_cpu_wrap(balance_cpu,  nohz.idle_cpus_mask, this_cpu+1) {
+	start = sbm_cpu_to_idx(cpumask_next_wrap(this_cpu, cpu_online_mask));
+	sbm_for_each_set_bit_wrap(nohz.sbm, idx, start) {
+		balance_cpu = sbm_idx_to_cpu(idx);
+
 		if (!idle_cpu(balance_cpu))
 			continue;
 
@@ -15648,7 +15636,7 @@ void show_numa_stats(struct task_struct *p, struct seq_file *m)
 __init void init_sched_fair_class_smp(void)
 {
 #ifdef CONFIG_NO_HZ_COMMON
-	zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
+	nohz.sbm = sbm_alloc();
 #endif
 }
 
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
  2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
@ 2026-10-03  8:27   ` Chen Yu
  2026-10-04  6:17     ` K Prateek Nayak
  0 siblings, 1 reply; 18+ messages in thread
From: Chen Yu @ 2026-10-03  8:27 UTC (permalink / raw)
  To: K Prateek Nayak
  Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, H. Peter Anvin

Hello Prateek,

On Thu, Oct 01, 2026 at 07:28:45PM +0000, K Prateek Nayak wrote:

[ snip ]

> +static __init void init_sbm_topology(u32 max_apicid)
> +{
> +	u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;

As sashiko reported, and also mentioned here:
https://lore.kernel.org/lkml/20260510155920.2587431-2-yu.c.chen@intel.com/
Maybe x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1] ?

> +	int num_sbm_instances, max_threads_per_instance;
> +
> +	/*
> +	 * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
> +	 * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
> +	 * that instead.
> +	 */
> +	if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
> +	    boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
> +		sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;

Ditto.

thanks,
Chenyu

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
  2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
                   ` (12 preceding siblings ...)
  2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
@ 2026-10-03  9:10 ` Chen Yu
  2026-10-04  6:13   ` K Prateek Nayak
  13 siblings, 1 reply; 18+ messages in thread
From: Chen Yu @ 2026-10-03  9:10 UTC (permalink / raw)
  To: K Prateek Nayak
  Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, WANG Xuerui,
	Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

Hello Prateek!

Thanks for bringing this interesting topic,

On Thu, Oct 01, 2026 at 07:28:36PM +0000, K Prateek Nayak wrote:
> 
> Problem
> =======
> 
>                                             %cycles vs global mask operation
> 
> global mask                                     : 100.0000%  (var: 3.28%)
> per-NUMA mask                                   :  32.9209%  (var: 7.77%)
> per-LLC mask                                    :   1.2977%  (var: 4.85%)
> per-LLC mask (u8 operation; no LOCK prefix)     :   0.4930%  (var: 0.83%)
> 

This shows a significant latency improvement, especially in the per-LLC (u64, u8) case.
May I know if schbench / sched-messaging were used?

> 
> Future work
> ===========
> 
> o Interoperability with cpumaks since sbm lose the crucial optimizations
>   that come naturally from for_each_cpu_and() iterations.
>
> o Different data representation - using the u8 variant for updates and
>   then perform a "gather" operation to build a dense mask.
> 
> o Extending sbm work to help in wakeup (and possibly resurrect Mel's
>   optimization from [4] in some form). The current sbm is still far away
>   from being used for wakeups since  updates  to sbm leaf, even on a
>   16CPUs per LLC system is visible in benchmark performance (~8-10%).
> 

If we leverage sbm for the wakeup path, it is a per-LLC mask, there seems to be
no much difference from Mel Gorman's proposal of allocating per sd_share
unsigned long idle_cpus_span[]? The frequent update to this mask might still cause
c2c latency within 1 LLC. A wild guess is that maybe the u8 version is more suitable,
because it has only max-to-8 CPUs touching the mask at the same time?

My understanding is that the major case that sbm could fit is turning the global bitmask
into a per-LLC bitmask, because it mainly avoids CPUs on different LLC/node writing the
same cache line frequently (nohz.idle_cpus_mask set via nohz_balance_enter_idle() on
many CPUs, etc), which might cause a costly cache RFO event storm. Meanwhile, with sbm,
at the reader side, _nohz_idle_balance() could start scanning from the current CPU to find
an idle CPU, so as to avoid the costly HITM event - the reader is on LLC1, while the writer
is on LLC0 - so maybe:

for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1)

could start from this_cpu's LLC sibling first, rather than this_cpu + 1, because this_cpu+1
might not always be the LLC sibling of this_cpu.

I found that in the current code, there are also other global mask:
rd->rto_mask(mentioned by Pan Deng when running ffmpeg[1])
rd->dlo_mask
tick_broadcast_**mask
maybe they can also be converted into sbm.

[1] https://lore.kernel.org/lkml/a3207ebf537bbe5605ff5454f63b5604d83a04a0.1753076363.git.pan.deng@intel.com/

thanks,
Chenyu 

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
  2026-10-03  9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
@ 2026-10-04  6:13   ` K Prateek Nayak
  0 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-04  6:13 UTC (permalink / raw)
  To: Chen Yu
  Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
	Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
	Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
	Alexander Gordeev, David S. Miller, Andreas Larsson,
	Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, WANG Xuerui,
	Michael Ellerman, Nicholas Piggin, Christophe Leroy,
	Christian Borntraeger, Sven Schnelle, H. Peter Anvin

Hello Chenyu,

On 10/3/2026 2:40 PM, Chen Yu wrote:
> Hello Prateek!
> 
> Thanks for bringing this interesting topic,
> 
> On Thu, Oct 01, 2026 at 07:28:36PM +0000, K Prateek Nayak wrote:
>>
>> Problem
>> =======
>>
>>                                             %cycles vs global mask operation
>>
>> global mask                                     : 100.0000%  (var: 3.28%)
>> per-NUMA mask                                   :  32.9209%  (var: 7.77%)
>> per-LLC mask                                    :   1.2977%  (var: 4.85%)
>> per-LLC mask (u8 operation; no LOCK prefix)     :   0.4930%  (var: 0.83%)
>>
> 
> This shows a significant latency improvement, especially in the per-LLC (u64, u8) case.
> May I know if schbench / sched-messaging were used?

No, this was a custom benchmark with two threads per
CPU - one setting the CPU on bitmask, and other clearing it,
continuously yielding to each other.

It gives an idea of what the worst case looks like when there
may be short idling followed by a short runtime going in
cycles.

> 
>>
>> Future work
>> ===========
>>
>> o Interoperability with cpumaks since sbm lose the crucial optimizations
>>   that come naturally from for_each_cpu_and() iterations.
>>
>> o Different data representation - using the u8 variant for updates and
>>   then perform a "gather" operation to build a dense mask.
>>
>> o Extending sbm work to help in wakeup (and possibly resurrect Mel's
>>   optimization from [4] in some form). The current sbm is still far away
>>   from being used for wakeups since  updates  to sbm leaf, even on a
>>   16CPUs per LLC system is visible in benchmark performance (~8-10%).
>>
> 
> If we leverage sbm for the wakeup path, it is a per-LLC mask, there seems to be
> no much difference from Mel Gorman's proposal of allocating per sd_share
> unsigned long idle_cpus_span[]?

Ack! But the current form is still pretty expensive. I see about a
10% overhead of just maintaining that mask which is what I'm trying
to reduce.

> The frequent update to this mask might still cause
> c2c latency within 1 LLC. A wild guess is that maybe the u8 version is more suitable,
> because it has only max-to-8 CPUs touching the mask at the same time?

With a 64B cacheline, single cacheline can contain data for up to
64 CPUs - unlike current sbm that uses first 8-bytes, this used
the whole 64 bytes.

P.S. All versions were tested with 16 CPUs per LLC on my systems.

Going from atomic u64 to plain u8 writes probably avoids an
expensive atomic path in the H/W making them faster.

> My understanding is that the major case that sbm could fit is turning the global bitmask
> into a per-LLC bitmask, because it mainly avoids CPUs on different LLC/node writing the
> same cache line frequently (nohz.idle_cpus_mask set via nohz_balance_enter_idle() on
> many CPUs, etc), which might cause a costly cache RFO event storm. Meanwhile, with sbm,
> at the reader side, _nohz_idle_balance() could start scanning from the current CPU to find
> an idle CPU, so as to avoid the costly HITM event - the reader is on LLC1, while the writer
> is on LLC0 - so maybe:
> 
> for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1)
> 
> could start from this_cpu's LLC sibling first, rather than this_cpu + 1, because this_cpu+1
> might not always be the LLC sibling of this_cpu.

Ah! Good point. Let me see if wrapping within the bitmask leaf
first and then going out makes any difference. 

> 
> I found that in the current code, there are also other global mask:
> rd->rto_mask(mentioned by Pan Deng when running ffmpeg[1])
> rd->dlo_mask
> tick_broadcast_**mask
> maybe they can also be converted into sbm.

Ack! I was juts getting started somewhere to see if there is an
appetite for sbm :-)

> 
> [1] https://lore.kernel.org/lkml/a3207ebf537bbe5605ff5454f63b5604d83a04a0.1753076363.git.pan.deng@intel.com/
> 
> thanks,
> Chenyu 

-- 
Thanks and Regards,
Prateek


^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
  2026-10-03  8:27   ` Chen Yu
@ 2026-10-04  6:17     ` K Prateek Nayak
  0 siblings, 0 replies; 18+ messages in thread
From: K Prateek Nayak @ 2026-10-04  6:17 UTC (permalink / raw)
  To: Chen Yu
  Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
	Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
	linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
	driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, Shrikanth Hegde, H. Peter Anvin

Hello Chenyu,

On 10/3/2026 1:57 PM, Chen Yu wrote:
> Hello Prateek,
> 
> On Thu, Oct 01, 2026 at 07:28:45PM +0000, K Prateek Nayak wrote:
> 
> [ snip ]
> 
>> +static __init void init_sbm_topology(u32 max_apicid)
>> +{
>> +	u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
> 
> As sashiko reported, and also mentioned here:
> https://lore.kernel.org/lkml/20260510155920.2587431-2-yu.c.chen@intel.com/
> Maybe x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1] ?

Ack!

I was testing this on a 2CCX (TILE) = 1 CCD (DIE) machine and I
failed to notice my mistake since:

  x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1

and

  x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1

were the same.

> 
>> +	int num_sbm_instances, max_threads_per_instance;
>> +
>> +	/*
>> +	 * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
>> +	 * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
>> +	 * that instead.
>> +	 */
>> +	if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
>> +	    boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
>> +		sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
> 
> Ditto.

Thank you! Will make sure I pull a VM with weird topology next time
during my testing to see if everything is correct.

-- 
Thanks and Regards,
Prateek


^ permalink raw reply	[flat|nested] 18+ messages in thread

end of thread, other threads:[~2026-10-04  6:17 UTC | newest]

Thread overview: 18+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-03  8:27   ` Chen Yu
2026-10-04  6:17     ` K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
2026-10-03  9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
2026-10-04  6:13   ` K Prateek Nayak

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®