mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v7 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus
@ 2026-09-29 16:54 Andrea Righi
  2026-09-29 16:54 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
  2026-09-29 16:54 ` [PATCH 2/2] sched/topology: Add asymmetric SMT packing override Andrea Righi
  0 siblings, 2 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 16:54 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot, Will Deacon
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Christian Loehle,
	Srikar Dronamraju, Shrikanth Hegde, Phil Auld, Breno Leitao,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, Kayra Cizmeci, linux-doc, linux-kernel

NVIDIA Olympus implements SMT with two symmetric processing elements (PEs).
When only one PE is active, the core operates in single-thread mode and that
PE can use the full core resources. When both PEs are active, the core
operates in two-thread mode and the PEs share those resources. Olympus is
particularly sensitive to brief sibling activations because returning to
single-thread mode after a sibling becomes idle is not immediate. As
described by commit 293f9611ae735 ("sched/fair: Prefer fully idle cores
for NOHZ balancing"):

  Briefly activating an otherwise idle sibling can reduce the
  performance available to the other sibling and this effect does not
  necessarily end once the activated sibling becomes idle: after the ILB
  finishes and its CPU enters WFI, full single-thread performance is
  restored only after the sibling has remained idle for a qualification
  interval (10 Ki cycles on the tested Vera system).

That change prevents the NOHZ idle load balancer from unnecessarily waking
a sibling of a busy PE. However, ordinary task placement can still select
either sibling of an idle core. Repeatedly changing the active PE can keep
Olympus cores in two-thread mode despite little useful overlap between the
siblings.

The first patch teaches the fair scheduler's idle-selection paths to honor
SD_ASYM_PACKING at the shared-capacity SMT level. The scheduler first
selects a candidate CPU and core according to its existing placement and
capacity rules, then chooses the highest-priority available sibling within
that core. POWER7 already enables SD_ASYM_PACKING on its SMT domain through
the PowerPC CPU_FTR_ASYM_SMT feature selected for that processor. The first
patch also applies POWER7's existing hardware-thread order during idle
selection.

Olympus firmware does not currently provide an interface to describe the
preferred SMT sibling. Adding such a firmware or ACPI interface will take
time and will not help systems with existing firmware. Inferring the policy
from MIDR would encode a platform-specific decision in the kernel and make
it harder to replace with a firmware ABI.

The second patch adds sched_smt_asym_packing=on as an explicit fallback for
systems that need SMT asymmetric packing but cannot describe the preference
through firmware. The option adds SD_ASYM_PACKING to the SMT scheduling
domain; omitting it preserves the topology supplied by the architecture or
firmware. Priority remains defined by arch_asym_cpu_priority(). The weak
default prefers lower-numbered logical CPUs, while architecture overrides
remain authoritative and equal-priority siblings remain unordered.

On the tested Olympus system, the lower-numbered logical CPU of each core
is PE0. The preference does not identify a faster PE: PE0 and PE1 have
equal steady-state capacity. Consistently selecting the same sibling avoids
alternating the active PE across wakeups, gives the other sibling more time
idle, and allows more cores to remain in, or return to, full-resource
single-thread mode.

The sched_smt_asym_packing=on path, unchanged from v6, was tested on a
two-node Vera system using 88-thread single-precision GEMM workloads on
the 88 physical cores of NUMA node 0, with both siblings of each core
eligible. Each result below covers five runs. OpenBLAS was evaluated with
its public benchmark/sgemm.goto single-GEMM benchmark. NVIDIA Performance
Libraries (NVPL) was evaluated with a timed benchblas GEMM run. Both used
M=N=K=16384; NVPL used non-transposed inputs, alpha=1, and beta=0.

OpenBLAS throughput increased from 7.11876 +/- 0.06734 TFLOP/s on the
baseline kernel to 7.34669 +/- 0.01936 TFLOP/s with this series (+3.20%).
NVPL throughput increased from 9.64742 +/- 0.17311 TFLOP/s to
10.29695 +/- 0.01786 TFLOP/s (+6.73%). The lower standard deviation also
shows that the results became more predictable. With the series applied,
the workloads consistently settled on the lower-numbered sibling.

Changes in v7:
 - Accept only sched_smt_asym_packing=on, drop the redundant auto and off
   modes (Vincent Guittot, Shrikanth Hegde).
 - Link to v6: https://lore.kernel.org/r/20260917140707.3807229-1-arighi@nvidia.com

Changes in v6:
 - Drop the arm64 MIDR-based enablement and arch_asym_cpu_priority()
   override (Will Deacon).
 - Add the generic sched_smt_asym_packing={auto,on,off} boot option.
 - Use the default -cpu priority ordering instead of interpreting MPIDR.
 - Drop the SMT-specific asymmetric-packing static key and use
   sched_smt_active() (Vincent Guittot).
 - Link to v5: https://lore.kernel.org/r/20260909062649.469633-1-arighi@nvidia.com

Changes in v5:
 - Remove the redundant olympus_prefer_pe0 state (K Prateek Nayak).
 - Link to v4: https://lore.kernel.org/r/20260908082345.103087-1-arighi@nvidia.com

Changes in v4:
 - Honor the SMT sibling priority in the slow path (Srikar Dronamraju).
 - Rename the consolidated helper to select_idle_smt_cpu()
   (Srikar Dronamraju).
 - Link to v3: https://lore.kernel.org/r/20260907163513.4172411-1-arighi@nvidia.com

Changes in v3:
 - Consolidate the SMT-priority adjustment in select_idle_sibling()
   after an idle candidate has been selected (K Prateek Nayak).
 - Fold the asym SMT checks into select_idle_smt_priority() and scan the
   scheduling-domain span directly (K Prateek Nayak).
 - Link to v2: https://lore.kernel.org/r/20260904091838.3617894-1-arighi@nvidia.com

Changes in v2:
 - Clarify that the generic scheduler change also covers POWER7
   (Dietmar Eggemann).
 - Simplify sched_smt_asym_prefer() by inspecting the lowest scheduling
   domain directly (Dietmar Eggemann).
 - Link to v1: https://lore.kernel.org/r/20260831181800.1668646-1-arighi@nvidia.com

Andrea Righi (2):
      sched/fair: Honor asymmetric SMT priority in idle selection
      sched/topology: Add asymmetric SMT packing override

 Documentation/admin-guide/kernel-parameters.txt | 12 ++++
 kernel/sched/fair.c                             | 85 ++++++++++++++++++++-----
 kernel/sched/topology.c                         | 18 ++++++
 3 files changed, 98 insertions(+), 17 deletions(-)

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-29 16:54 [PATCH v7 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus Andrea Righi
@ 2026-09-29 16:54 ` Andrea Righi
  2026-09-29 16:54 ` [PATCH 2/2] sched/topology: Add asymmetric SMT packing override Andrea Righi
  1 sibling, 0 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 16:54 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot, Will Deacon
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Christian Loehle,
	Srikar Dronamraju, Shrikanth Hegde, Phil Auld, Breno Leitao,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, Kayra Cizmeci, linux-doc, linux-kernel

POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
CPU selection does not consult that order, so a task can wake on an
arbitrary sibling and remain there until load balancing corrects the
placement. On these systems, that initial choice can prevent the core
from entering its preferred lower-thread resource mode and cause a large
and persistent performance loss.

When idle selection finds an available CPU in an SMT core, choose the
highest-priority available sibling. On SMT2 Olympus this only changes
selection on fully idle cores. A partially idle core has only one
available CPU. On wider SMT systems such as POWER7, it also fills
available siblings in priority order while the core is partially busy.

Apply the preference to idle-core and idle-CPU scans,
asymmetric-capacity scans, target, previous, recently-used CPU fast
paths and the slow path. Inspect the lowest scheduling domain directly,
but require both CPUs to share its span because isolcpus can split
hardware siblings across scheduling domains.

Keep physical-core capacity selection independent from SMT sibling
ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
maximum capacities, then SD_ASYM_PACKING selects the preferred available
sibling inside the chosen core, whose siblings continue to share equal
capacity.

Reviewed-by: Srikar Dronamraju <srikar@linux.ibm.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Reviewed-by: Vincent Guittot <vincent.guittot@linaro.org>
Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Tested-by: Breno Leitao <leitao@debian.org>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 kernel/sched/fair.c | 85 ++++++++++++++++++++++++++++++++++++---------
 1 file changed, 68 insertions(+), 17 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index e707da7177dfe..5c98d8dfce5d8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8799,6 +8799,35 @@ static inline bool test_idle_cores(int cpu)
 	return false;
 }
 
+/*
+ * Redirect a CPU to a higher-priority available sibling in its SMT domain,
+ * subject to task affinity.
+ */
+static inline int select_idle_smt_cpu(struct task_struct *p, int cpu)
+{
+	struct sched_domain *sd;
+	int best = cpu;
+	int sibling;
+
+	if (!sched_smt_active())
+		return cpu;
+
+	sd = rcu_dereference_all(cpu_rq(cpu)->sd);
+	if (!sd || !(sd->flags & SD_SHARE_CPUCAPACITY) ||
+	    !(sd->flags & SD_ASYM_PACKING))
+		return cpu;
+
+	for_each_cpu_and(sibling, sched_domain_span(sd), p->cpus_ptr) {
+		if (sibling == best || !choose_idle_cpu(sibling, p))
+			continue;
+
+		if (sched_asym_prefer(sibling, best))
+			best = sibling;
+	}
+
+	return best;
+}
+
 /*
  * Scans the local SMT mask to see if the entire core is idle, and records this
  * information in sd_balance_shared->has_idle_cores.
@@ -9183,7 +9212,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 
 	if (choose_idle_cpu(target, p) &&
 	    asym_fits_cpu(task_util, util_min, util_max, target))
-		return target;
+		goto select_smt_priority;
 
 	/*
 	 * If the previous CPU is cache affine and idle, don't be stupid:
@@ -9193,8 +9222,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    asym_fits_cpu(task_util, util_min, util_max, prev)) {
 
 		if (!static_branch_unlikely(&sched_cluster_active) ||
-		    cpus_share_resources(prev, target))
-			return prev;
+		    cpus_share_resources(prev, target)) {
+			target = prev;
+			goto select_smt_priority;
+		}
 
 		prev_aff = prev;
 	}
@@ -9212,7 +9243,8 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    prev == smp_processor_id() &&
 	    this_rq()->nr_running <= 1 &&
 	    asym_fits_cpu(task_util, util_min, util_max, prev)) {
-		return prev;
+		target = prev;
+		goto select_smt_priority;
 	}
 
 	/* Check a recently used CPU as a potential idle candidate: */
@@ -9226,8 +9258,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    asym_fits_cpu(task_util, util_min, util_max, recent_used_cpu)) {
 
 		if (!static_branch_unlikely(&sched_cluster_active) ||
-		    cpus_share_resources(recent_used_cpu, target))
-			return recent_used_cpu;
+		    cpus_share_resources(recent_used_cpu, target)) {
+			target = recent_used_cpu;
+			goto select_smt_priority;
+		}
 
 	} else {
 		recent_used_cpu = -1;
@@ -9249,7 +9283,11 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 		 */
 		if (sd) {
 			i = select_idle_capacity(p, sd, target);
-			return ((unsigned)i < nr_cpumask_bits) ? i : target;
+			if ((unsigned int)i < nr_cpumask_bits) {
+				target = i;
+				goto select_smt_priority;
+			}
+			return target;
 		}
 	}
 
@@ -9262,14 +9300,18 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 
 		if (!has_idle_core && cpus_share_cache(prev, target)) {
 			i = select_idle_smt(p, sd, prev);
-			if ((unsigned int)i < nr_cpumask_bits)
-				return i;
+			if ((unsigned int)i < nr_cpumask_bits) {
+				target = i;
+				goto select_smt_priority;
+			}
 		}
 	}
 
 	i = select_idle_cpu(p, sd, has_idle_core, target);
-	if ((unsigned)i < nr_cpumask_bits)
-		return i;
+	if ((unsigned int)i < nr_cpumask_bits) {
+		target = i;
+		goto select_smt_priority;
+	}
 
 	/*
 	 * For cluster machines which have lower sharing cache like L2 or
@@ -9277,12 +9319,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	 * first. But prev_cpu or recent_used_cpu may also be a good candidate,
 	 * use them if possible when no idle CPU found in select_idle_cpu().
 	 */
-	if ((unsigned int)prev_aff < nr_cpumask_bits)
-		return prev_aff;
-	if ((unsigned int)recent_used_cpu < nr_cpumask_bits)
-		return recent_used_cpu;
+	if ((unsigned int)prev_aff < nr_cpumask_bits) {
+		target = prev_aff;
+		goto select_smt_priority;
+	}
+	if ((unsigned int)recent_used_cpu < nr_cpumask_bits) {
+		target = recent_used_cpu;
+		goto select_smt_priority;
+	}
 
 	return target;
+
+select_smt_priority:
+	return select_idle_smt_cpu(p, target);
 }
 
 /**
@@ -9959,8 +10008,10 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
 	}
 
 	/* Slow path */
-	if (unlikely(sd))
-		return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
+	if (unlikely(sd)) {
+		new_cpu = sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
+		return select_idle_smt_cpu(p, new_cpu);
+	}
 
 	/* Fast path */
 	if (wake_flags & WF_TTWU)
-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 2/2] sched/topology: Add asymmetric SMT packing override
  2026-09-29 16:54 [PATCH v7 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus Andrea Righi
  2026-09-29 16:54 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
@ 2026-09-29 16:54 ` Andrea Righi
  1 sibling, 0 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-29 16:54 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot, Will Deacon
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Christian Loehle,
	Srikar Dronamraju, Shrikanth Hegde, Phil Auld, Breno Leitao,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, Kayra Cizmeci, linux-doc, linux-kernel

Architectures can use SD_ASYM_PACKING to describe preferred CPU ordering
at the SMT scheduling domain. Some systems benefit from this policy, but
their firmware cannot describe the preference and inferring it from the
CPU model would embed a platform-specific policy in the kernel.

Add sched_smt_asym_packing=on boot option to force SD_ASYM_PACKING at
the SMT level. Omitting the option preserves the topology provided by
the architecture or firmware.

Apply the override centrally to domains with SD_SHARE_CPUCAPACITY so it
also covers architectures with custom SMT topology callbacks, including
powerpc.

When enabled, priority remains defined by arch_asym_cpu_priority(). The
weak default prefers lower-numbered logical CPUs, while architecture
overrides remain authoritative. Siblings with equal priorities remain
unordered. In particular, x86 with CONFIG_SCHED_MC_PRIO normally gives
both SMT siblings the same core priority, so enabling the option there
does not prioritize a sibling.

Tested-by: Breno Leitao <leitao@debian.org>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 .../admin-guide/kernel-parameters.txt          | 12 ++++++++++++
 kernel/sched/topology.c                        | 18 ++++++++++++++++++
 2 files changed, 30 insertions(+)

diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index e75344f4e0cde..8157ff434cb5e 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -6800,6 +6800,18 @@ Kernel parameters
 			solution to mutex-based priority inversion.
 			Format: <bool>
 
+	sched_smt_asym_packing= [KNL,SMP]
+			Format: on
+			Force asymmetric packing at the SMT scheduling domain.
+			Idle CPU selection prefers siblings with a higher
+			arch_asym_cpu_priority(). The default implementation
+			prefers lower-numbered logical CPUs, but architecture
+			overrides remain authoritative. Equal priorities do
+			not establish a sibling preference. For example, x86
+			with CONFIG_SCHED_MC_PRIO normally assigns the same
+			core priority to both SMT siblings, so this option
+			does not favor either sibling.
+
 	sched_verbose	[KNL,EARLY] Enables verbose scheduler debug messages.
 
 	schedstats=	[KNL,X86] Enable or disable scheduled statistics.
diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 3dab0253976fb..919d0fb00bd9f 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -32,6 +32,20 @@ static int __init sched_debug_setup(char *str)
 }
 early_param("sched_verbose", sched_debug_setup);
 
+#ifdef CONFIG_SCHED_SMT
+static bool sched_smt_asym_packing __read_mostly;
+
+static int __init setup_sched_smt_asym_packing(char *str)
+{
+	if (strcmp(str, "on"))
+		return 0;
+
+	sched_smt_asym_packing = true;
+	return 1;
+}
+__setup("sched_smt_asym_packing=", setup_sched_smt_asym_packing);
+#endif
+
 static inline bool sched_debug(void)
 {
 	return sched_debug_verbose;
@@ -1954,6 +1968,10 @@ sd_init(struct sched_domain_topology_level *tl,
 	if (WARN_ONCE(sd_flags & ~TOPOLOGY_SD_FLAGS,
 		      "wrong sd_flags in topology description\n"))
 		sd_flags &= TOPOLOGY_SD_FLAGS;
+#ifdef CONFIG_SCHED_SMT
+	if (sched_smt_asym_packing && (sd_flags & SD_SHARE_CPUCAPACITY))
+		sd_flags |= SD_ASYM_PACKING;
+#endif
 	sd_flags |= asym_cpu_capacity_classify(sd_span, cpu_map);
 
 	*sd = (struct sched_domain){
-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-17 14:05 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
  2026-09-17 16:55   ` Kayra Cizmeci
  2026-09-18 15:58   ` Vincent Guittot
@ 2026-09-21 14:40   ` Breno Leitao
  2 siblings, 0 replies; 8+ messages in thread
From: Breno Leitao @ 2026-09-21 14:40 UTC (permalink / raw)
  To: Andrea Righi
  Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Will Deacon, Dietmar Eggemann, Steven Rostedt, Ben Segall,
	Mel Gorman, Valentin Schneider, K Prateek Nayak,
	Christian Loehle, Srikar Dronamraju, Shrikanth Hegde, Phil Auld,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, linux-doc, linux-kernel

On Thu, Sep 17, 2026 at 04:05:20PM +0200, Andrea Righi wrote:
> POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
> hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
> CPU selection does not consult that order, so a task can wake on an
> arbitrary sibling and remain there until load balancing corrects the
> placement. On these systems, that initial choice can prevent the core
> from entering its preferred lower-thread resource mode and cause a large
> and persistent performance loss.
> 
> When idle selection finds an available CPU in an SMT core, choose the
> highest-priority available sibling. On SMT2 Olympus this only changes
> selection on fully idle cores. A partially idle core has only one
> available CPU. On wider SMT systems such as POWER7, it also fills
> available siblings in priority order while the core is partially busy.
> 
> Apply the preference to idle-core and idle-CPU scans,
> asymmetric-capacity scans, target, previous, recently-used CPU fast
> paths and the slow path. Inspect the lowest scheduling domain directly,
> but require both CPUs to share its span because isolcpus can split
> hardware siblings across scheduling domains.
> 
> Keep physical-core capacity selection independent from SMT sibling
> ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
> maximum capacities, then SD_ASYM_PACKING selects the preferred available
> sibling inside the chosen core, whose siblings continue to share equal
> capacity.
> 
> Reviewed-by: Srikar Dronamraju <srikar@linux.ibm.com>
> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Signed-off-by: Andrea Righi <arighi@nvidia.com>

Tested-by: Breno Leitao <leitao@debian.org>

I tested v6 on top of linux-next 20260918 with a KASAN + PROVE_LOCKING
arm64 kernel and etst on my own 2-socket NVIDIA Olympus system (88 cores
per socket, SMT2, 352 CPUs, 36 NUMA nodes).

With sched_smt_asym_packing=on, SD_ASYM_PACKING appears on the SMT
domain only; MC and NUMA are untouched, and the override is reapplied
on every domain rebuild (~400 hotplug operations, including offlining
and re-onlining all 176 second siblings). A task woken on the
higher-numbered sibling is redirected to the lower-numbered one on
800/800 wakeups, against 13/800 without the series. auto, off and an
invalid value behave as documented. No WARN/BUG/KASAN/lockdep under
stress-ng and perf bench sched.

I did not measure throughput.

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-17 14:05 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
  2026-09-17 16:55   ` Kayra Cizmeci
@ 2026-09-18 15:58   ` Vincent Guittot
  2026-09-21 14:40   ` Breno Leitao
  2 siblings, 0 replies; 8+ messages in thread
From: Vincent Guittot @ 2026-09-18 15:58 UTC (permalink / raw)
  To: Andrea Righi
  Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Will Deacon,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Christian Loehle,
	Srikar Dronamraju, Shrikanth Hegde, Phil Auld, Breno Leitao,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, linux-doc, linux-kernel

On Thu, 17 Sept 2026 at 16:08, Andrea Righi <arighi@nvidia.com> wrote:
>
> POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
> hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
> CPU selection does not consult that order, so a task can wake on an
> arbitrary sibling and remain there until load balancing corrects the
> placement. On these systems, that initial choice can prevent the core
> from entering its preferred lower-thread resource mode and cause a large
> and persistent performance loss.
>
> When idle selection finds an available CPU in an SMT core, choose the
> highest-priority available sibling. On SMT2 Olympus this only changes
> selection on fully idle cores. A partially idle core has only one
> available CPU. On wider SMT systems such as POWER7, it also fills
> available siblings in priority order while the core is partially busy.
>
> Apply the preference to idle-core and idle-CPU scans,
> asymmetric-capacity scans, target, previous, recently-used CPU fast
> paths and the slow path. Inspect the lowest scheduling domain directly,
> but require both CPUs to share its span because isolcpus can split
> hardware siblings across scheduling domains.
>
> Keep physical-core capacity selection independent from SMT sibling
> ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
> maximum capacities, then SD_ASYM_PACKING selects the preferred available
> sibling inside the chosen core, whose siblings continue to share equal
> capacity.
>
> Reviewed-by: Srikar Dronamraju <srikar@linux.ibm.com>
> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Signed-off-by: Andrea Righi <arighi@nvidia.com>

Reviewed-by: Vincent Guittot <vincent.guittot@linaro.org>


> ---
>  kernel/sched/fair.c | 85 ++++++++++++++++++++++++++++++++++++---------
>  1 file changed, 68 insertions(+), 17 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 4d0b94465d19e..8e6dc3a657cca 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -8598,6 +8598,35 @@ static inline bool test_idle_cores(int cpu)
>         return false;
>  }
>
> +/*
> + * Redirect a CPU to a higher-priority available sibling in its SMT domain,
> + * subject to task affinity.
> + */
> +static inline int select_idle_smt_cpu(struct task_struct *p, int cpu)
> +{
> +       struct sched_domain *sd;
> +       int best = cpu;
> +       int sibling;
> +
> +       if (!sched_smt_active())
> +               return cpu;
> +
> +       sd = rcu_dereference_all(cpu_rq(cpu)->sd);
> +       if (!sd || !(sd->flags & SD_SHARE_CPUCAPACITY) ||
> +           !(sd->flags & SD_ASYM_PACKING))
> +               return cpu;
> +
> +       for_each_cpu_and(sibling, sched_domain_span(sd), p->cpus_ptr) {
> +               if (sibling == best || !choose_idle_cpu(sibling, p))
> +                       continue;
> +
> +               if (sched_asym_prefer(sibling, best))
> +                       best = sibling;
> +       }
> +
> +       return best;
> +}
> +
>  /*
>   * Scans the local SMT mask to see if the entire core is idle, and records this
>   * information in sd_balance_shared->has_idle_cores.
> @@ -8982,7 +9011,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>
>         if (choose_idle_cpu(target, p) &&
>             asym_fits_cpu(task_util, util_min, util_max, target))
> -               return target;
> +               goto select_smt_priority;
>
>         /*
>          * If the previous CPU is cache affine and idle, don't be stupid:
> @@ -8992,8 +9021,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>             asym_fits_cpu(task_util, util_min, util_max, prev)) {
>
>                 if (!static_branch_unlikely(&sched_cluster_active) ||
> -                   cpus_share_resources(prev, target))
> -                       return prev;
> +                   cpus_share_resources(prev, target)) {
> +                       target = prev;
> +                       goto select_smt_priority;
> +               }
>
>                 prev_aff = prev;
>         }
> @@ -9011,7 +9042,8 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>             prev == smp_processor_id() &&
>             this_rq()->nr_running <= 1 &&
>             asym_fits_cpu(task_util, util_min, util_max, prev)) {
> -               return prev;
> +               target = prev;
> +               goto select_smt_priority;
>         }
>
>         /* Check a recently used CPU as a potential idle candidate: */
> @@ -9025,8 +9057,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>             asym_fits_cpu(task_util, util_min, util_max, recent_used_cpu)) {
>
>                 if (!static_branch_unlikely(&sched_cluster_active) ||
> -                   cpus_share_resources(recent_used_cpu, target))
> -                       return recent_used_cpu;
> +                   cpus_share_resources(recent_used_cpu, target)) {
> +                       target = recent_used_cpu;
> +                       goto select_smt_priority;
> +               }
>
>         } else {
>                 recent_used_cpu = -1;
> @@ -9048,7 +9082,11 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>                  */
>                 if (sd) {
>                         i = select_idle_capacity(p, sd, target);
> -                       return ((unsigned)i < nr_cpumask_bits) ? i : target;
> +                       if ((unsigned int)i < nr_cpumask_bits) {
> +                               target = i;
> +                               goto select_smt_priority;
> +                       }
> +                       return target;
>                 }
>         }
>
> @@ -9061,14 +9099,18 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>
>                 if (!has_idle_core && cpus_share_cache(prev, target)) {
>                         i = select_idle_smt(p, sd, prev);
> -                       if ((unsigned int)i < nr_cpumask_bits)
> -                               return i;
> +                       if ((unsigned int)i < nr_cpumask_bits) {
> +                               target = i;
> +                               goto select_smt_priority;
> +                       }
>                 }
>         }
>
>         i = select_idle_cpu(p, sd, has_idle_core, target);
> -       if ((unsigned)i < nr_cpumask_bits)
> -               return i;
> +       if ((unsigned int)i < nr_cpumask_bits) {
> +               target = i;
> +               goto select_smt_priority;
> +       }
>
>         /*
>          * For cluster machines which have lower sharing cache like L2 or
> @@ -9076,12 +9118,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>          * first. But prev_cpu or recent_used_cpu may also be a good candidate,
>          * use them if possible when no idle CPU found in select_idle_cpu().
>          */
> -       if ((unsigned int)prev_aff < nr_cpumask_bits)
> -               return prev_aff;
> -       if ((unsigned int)recent_used_cpu < nr_cpumask_bits)
> -               return recent_used_cpu;
> +       if ((unsigned int)prev_aff < nr_cpumask_bits) {
> +               target = prev_aff;
> +               goto select_smt_priority;
> +       }
> +       if ((unsigned int)recent_used_cpu < nr_cpumask_bits) {
> +               target = recent_used_cpu;
> +               goto select_smt_priority;
> +       }
>
>         return target;
> +
> +select_smt_priority:
> +       return select_idle_smt_cpu(p, target);
>  }
>
>  /**
> @@ -9758,8 +9807,10 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
>         }
>
>         /* Slow path */
> -       if (unlikely(sd))
> -               return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
> +       if (unlikely(sd)) {
> +               new_cpu = sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
> +               return select_idle_smt_cpu(p, new_cpu);
> +       }
>
>         /* Fast path */
>         if (wake_flags & WF_TTWU)
> --
> 2.55.0
>

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-17 16:55   ` Kayra Cizmeci
@ 2026-09-18  6:28     ` Kayra Cizmeci
  0 siblings, 0 replies; 8+ messages in thread
From: Kayra Cizmeci @ 2026-09-18  6:28 UTC (permalink / raw)
  To: kayracizmeci
  Cc: arighi, bsegall, christian.loehle, corbet, dietmar.eggemann,
	juri.lelli, kprateek.nayak, leitao, linux-doc, linux-kernel,
	ltrager, mgorman, mingo, pauld, peterz, rdunlap, rostedt, skhan,
	srikar, sshegde, vincent.guittot, vschneid, vsethi, will


> Hi Andrea,

> Hope I ain't got anything wrong, I'm a bit sick.


> The sched_smt_active() check on select_idle_smt_cpu seems to be redundant on here.

> We could remove this
> By moving the check on select_idle_smt_cpu() to goto block and
> adding one check before the select_idle_smt_cpu() on select_task_rq_fair(). And just calling
> select_idle_smt_cpu() on here.

> I also have a question, do we need select_idle_smt() call on here? If yes then why? I was really confused while reading.

> Thanks,
> Kayra :_:

After I don't know how much times of reading the same functions, I finally understand.
Because of has_idle_core is false, we search if there are any idle SMP's on 
any of the cores.

I normally was going to send this alot earlier, but I fall asleep after
realizing this.

Also now, I don't think we need to remove the checks for one place. It complicates things,
when I tried.

I also booted the patch, it did not worked on me since I just have a regular SMT2.
So, that was the waited behavior.

So, here's this:
Tested-by: Kayra Cizmeci <kayracizmeci@gmail.com>

Annnddddddddddddddddddddddddddddddd....

This: 
Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com>

Thanks,
Kayra :-)








^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-17 14:05 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
@ 2026-09-17 16:55   ` Kayra Cizmeci
  2026-09-18  6:28     ` Kayra Cizmeci
  2026-09-18 15:58   ` Vincent Guittot
  2026-09-21 14:40   ` Breno Leitao
  2 siblings, 1 reply; 8+ messages in thread
From: Kayra Cizmeci @ 2026-09-17 16:55 UTC (permalink / raw)
  To: arighi
  Cc: bsegall, christian.loehle, corbet, dietmar.eggemann, juri.lelli,
	kprateek.nayak, leitao, linux-doc, linux-kernel, ltrager,
	mgorman, mingo, pauld, peterz, rdunlap, rostedt, skhan, srikar,
	sshegde, vincent.guittot, vschneid, vsethi, will


Hi Andrea,

Hope I ain't got anything wrong, I'm a bit sick.

> POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
> hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
> CPU selection does not consult that order, so a task can wake on an
> arbitrary sibling and remain there until load balancing corrects the
> placement. On these systems, that initial choice can prevent the core
> from entering its preferred lower-thread resource mode and cause a large
> and persistent performance loss.

> When idle selection finds an available CPU in an SMT core, choose the
> highest-priority available sibling. On SMT2 Olympus this only changes
> selection on fully idle cores. A partially idle core has only one
> available CPU. On wider SMT systems such as POWER7, it also fills
> available siblings in priority order while the core is partially busy.

> Apply the preference to idle-core and idle-CPU scans,
> asymmetric-capacity scans, target, previous, recently-used CPU fast
> paths and the slow path. Inspect the lowest scheduling domain directly,
> but require both CPUs to share its span because isolcpus can split
> hardware siblings across scheduling domains.

> Keep physical-core capacity selection independent from SMT sibling
> ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
> maximum capacities, then SD_ASYM_PACKING selects the preferred available
> sibling inside the chosen core, whose siblings continue to share equal
> capacity.


> +/*
> + * Redirect a CPU to a higher-priority available sibling in its SMT domain,
> + * subject to task affinity.
> + */
> +static inline int select_idle_smt_cpu(struct task_struct *p, int cpu)
> +{
> +	struct sched_domain *sd;
> +	int best = cpu;
> +	int sibling;
> +
> +	if (!sched_smt_active())
> +		return cpu;
> +
> +	sd = rcu_dereference_all(cpu_rq(cpu)->sd);
> +	if (!sd || !(sd->flags & SD_SHARE_CPUCAPACITY) ||
> +	    !(sd->flags & SD_ASYM_PACKING))
> +		return cpu;
> +
> +	for_each_cpu_and(sibling, sched_domain_span(sd), p->cpus_ptr) {
> +		if (sibling == best || !choose_idle_cpu(sibling, p))
> +			continue;
> +
> +		if (sched_asym_prefer(sibling, best))
> +			best = sibling;
> +	}
> +
> +	return best;
> +}

> @@ -9061,14 +9099,18 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
>  
>  		if (!has_idle_core && cpus_share_cache(prev, target)) {
>  			i = select_idle_smt(p, sd, prev);
> -			if ((unsigned int)i < nr_cpumask_bits)
> -				return i;
> +			if ((unsigned int)i < nr_cpumask_bits) {
> +				target = i;
> +				goto select_smt_priority;
> +			}
>  		}
>  	}

The sched_smt_active() check on select_idle_smt_cpu seems to be redundant on here.

We could remove this
By moving the check on select_idle_smt_cpu() to goto block and
adding one check before the select_idle_smt_cpu() on select_task_rq_fair(). And just calling
select_idle_smt_cpu() on here.

I also have a question, do we need select_idle_smt() call on here? If yes then why? I was really confused while reading.

Thanks,
Kayra :_:



^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection
  2026-09-17 14:05 [PATCH v6 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus Andrea Righi
@ 2026-09-17 14:05 ` Andrea Righi
  2026-09-17 16:55   ` Kayra Cizmeci
                     ` (2 more replies)
  0 siblings, 3 replies; 8+ messages in thread
From: Andrea Righi @ 2026-09-17 14:05 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot, Will Deacon
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Christian Loehle,
	Srikar Dronamraju, Shrikanth Hegde, Phil Auld, Breno Leitao,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Lee Trager,
	Vikram Sethi, linux-doc, linux-kernel

POWER7 uses SD_ASYM_PACKING at the shared-capacity SMT level to order
hardware threads, and NVIDIA Olympus benefits from the same policy. Idle
CPU selection does not consult that order, so a task can wake on an
arbitrary sibling and remain there until load balancing corrects the
placement. On these systems, that initial choice can prevent the core
from entering its preferred lower-thread resource mode and cause a large
and persistent performance loss.

When idle selection finds an available CPU in an SMT core, choose the
highest-priority available sibling. On SMT2 Olympus this only changes
selection on fully idle cores. A partially idle core has only one
available CPU. On wider SMT systems such as POWER7, it also fills
available siblings in priority order while the core is partially busy.

Apply the preference to idle-core and idle-CPU scans,
asymmetric-capacity scans, target, previous, recently-used CPU fast
paths and the slow path. Inspect the lowest scheduling domain directly,
but require both CPUs to share its span because isolcpus can split
hardware siblings across scheduling domains.

Keep physical-core capacity selection independent from SMT sibling
ordering. SD_ASYM_CPUCAPACITY first selects among cores with different
maximum capacities, then SD_ASYM_PACKING selects the preferred available
sibling inside the chosen core, whose siblings continue to share equal
capacity.

Reviewed-by: Srikar Dronamraju <srikar@linux.ibm.com>
Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com>
Tested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 kernel/sched/fair.c | 85 ++++++++++++++++++++++++++++++++++++---------
 1 file changed, 68 insertions(+), 17 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 4d0b94465d19e..8e6dc3a657cca 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8598,6 +8598,35 @@ static inline bool test_idle_cores(int cpu)
 	return false;
 }
 
+/*
+ * Redirect a CPU to a higher-priority available sibling in its SMT domain,
+ * subject to task affinity.
+ */
+static inline int select_idle_smt_cpu(struct task_struct *p, int cpu)
+{
+	struct sched_domain *sd;
+	int best = cpu;
+	int sibling;
+
+	if (!sched_smt_active())
+		return cpu;
+
+	sd = rcu_dereference_all(cpu_rq(cpu)->sd);
+	if (!sd || !(sd->flags & SD_SHARE_CPUCAPACITY) ||
+	    !(sd->flags & SD_ASYM_PACKING))
+		return cpu;
+
+	for_each_cpu_and(sibling, sched_domain_span(sd), p->cpus_ptr) {
+		if (sibling == best || !choose_idle_cpu(sibling, p))
+			continue;
+
+		if (sched_asym_prefer(sibling, best))
+			best = sibling;
+	}
+
+	return best;
+}
+
 /*
  * Scans the local SMT mask to see if the entire core is idle, and records this
  * information in sd_balance_shared->has_idle_cores.
@@ -8982,7 +9011,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 
 	if (choose_idle_cpu(target, p) &&
 	    asym_fits_cpu(task_util, util_min, util_max, target))
-		return target;
+		goto select_smt_priority;
 
 	/*
 	 * If the previous CPU is cache affine and idle, don't be stupid:
@@ -8992,8 +9021,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    asym_fits_cpu(task_util, util_min, util_max, prev)) {
 
 		if (!static_branch_unlikely(&sched_cluster_active) ||
-		    cpus_share_resources(prev, target))
-			return prev;
+		    cpus_share_resources(prev, target)) {
+			target = prev;
+			goto select_smt_priority;
+		}
 
 		prev_aff = prev;
 	}
@@ -9011,7 +9042,8 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    prev == smp_processor_id() &&
 	    this_rq()->nr_running <= 1 &&
 	    asym_fits_cpu(task_util, util_min, util_max, prev)) {
-		return prev;
+		target = prev;
+		goto select_smt_priority;
 	}
 
 	/* Check a recently used CPU as a potential idle candidate: */
@@ -9025,8 +9057,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	    asym_fits_cpu(task_util, util_min, util_max, recent_used_cpu)) {
 
 		if (!static_branch_unlikely(&sched_cluster_active) ||
-		    cpus_share_resources(recent_used_cpu, target))
-			return recent_used_cpu;
+		    cpus_share_resources(recent_used_cpu, target)) {
+			target = recent_used_cpu;
+			goto select_smt_priority;
+		}
 
 	} else {
 		recent_used_cpu = -1;
@@ -9048,7 +9082,11 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 		 */
 		if (sd) {
 			i = select_idle_capacity(p, sd, target);
-			return ((unsigned)i < nr_cpumask_bits) ? i : target;
+			if ((unsigned int)i < nr_cpumask_bits) {
+				target = i;
+				goto select_smt_priority;
+			}
+			return target;
 		}
 	}
 
@@ -9061,14 +9099,18 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 
 		if (!has_idle_core && cpus_share_cache(prev, target)) {
 			i = select_idle_smt(p, sd, prev);
-			if ((unsigned int)i < nr_cpumask_bits)
-				return i;
+			if ((unsigned int)i < nr_cpumask_bits) {
+				target = i;
+				goto select_smt_priority;
+			}
 		}
 	}
 
 	i = select_idle_cpu(p, sd, has_idle_core, target);
-	if ((unsigned)i < nr_cpumask_bits)
-		return i;
+	if ((unsigned int)i < nr_cpumask_bits) {
+		target = i;
+		goto select_smt_priority;
+	}
 
 	/*
 	 * For cluster machines which have lower sharing cache like L2 or
@@ -9076,12 +9118,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
 	 * first. But prev_cpu or recent_used_cpu may also be a good candidate,
 	 * use them if possible when no idle CPU found in select_idle_cpu().
 	 */
-	if ((unsigned int)prev_aff < nr_cpumask_bits)
-		return prev_aff;
-	if ((unsigned int)recent_used_cpu < nr_cpumask_bits)
-		return recent_used_cpu;
+	if ((unsigned int)prev_aff < nr_cpumask_bits) {
+		target = prev_aff;
+		goto select_smt_priority;
+	}
+	if ((unsigned int)recent_used_cpu < nr_cpumask_bits) {
+		target = recent_used_cpu;
+		goto select_smt_priority;
+	}
 
 	return target;
+
+select_smt_priority:
+	return select_idle_smt_cpu(p, target);
 }
 
 /**
@@ -9758,8 +9807,10 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
 	}
 
 	/* Slow path */
-	if (unlikely(sd))
-		return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
+	if (unlikely(sd)) {
+		new_cpu = sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag);
+		return select_idle_smt_cpu(p, new_cpu);
+	}
 
 	/* Fast path */
 	if (wake_flags & WF_TTWU)
-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-09-29 16:56 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 16:54 [PATCH v7 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus Andrea Righi
2026-09-29 16:54 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
2026-09-29 16:54 ` [PATCH 2/2] sched/topology: Add asymmetric SMT packing override Andrea Righi
  -- strict thread matches above, loose matches on Subject: below --
2026-09-17 14:05 [PATCH v6 0/2] sched: Enable preferred SMT siblings on NVIDIA Olympus Andrea Righi
2026-09-17 14:05 ` [PATCH 1/2] sched/fair: Honor asymmetric SMT priority in idle selection Andrea Righi
2026-09-17 16:55   ` Kayra Cizmeci
2026-09-18  6:28     ` Kayra Cizmeci
2026-09-18 15:58   ` Vincent Guittot
2026-09-21 14:40   ` Breno Leitao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®