* [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority
2026-10-08 8:30 [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Mete Durlu
@ 2026-10-08 8:30 ` Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO Mete Durlu
` (2 subsequent siblings)
3 siblings, 0 replies; 12+ messages in thread
From: Mete Durlu @ 2026-10-08 8:30 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Shrikanth Hegde,
Heiko Carstens, Vasily Gorbik, Alexander Gordeev,
Christian Borntraeger, Sven Schnelle
Cc: Tim Chen, Chen Yu, Ilya Leoshkevich, Andrea Righi, linux-kernel,
linux-s390, Mete Durlu
On systems with asymmetric CPU capacities and SMT, the scheduler
currently prefers fully idle cores over idle SMT siblings of busy
cores. This leads tasks onto lower capacity idle cores when a
higher-capacity core has an idle sibling thread available. This hurts
performance in cases where an SMT sibling of a busy but high capacity
core can outperform a fully idle but low capacity core.
Introduce SCHED_IDLE_SMT_PRIO and the sched_idle_smt_prio static branch
to allow architectures to override this preference. When enabled, idle
SMT siblings of busy cores are treated as regular candidates.
The static branch defaults to false and is opt-in per architecture, so
systems that do not select ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO see no
change in behavior.
Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
implementing arch_needs_idle_smt_prio(), which is called during each
asym_cpu_capacity_scan(). In case there is no asymmetric capacities
static branches get disabled.
Signed-off-by: Mete Durlu <meted@linux.ibm.com>
---
arch/Kconfig | 14 +++++++++
include/linux/sched/topology.h | 9 ++++++
kernel/sched/fair.c | 67 ++++++++++++++++++++++++++++++------------
kernel/sched/sched.h | 10 +++++++
kernel/sched/topology.c | 17 +++++++++++
5 files changed, 99 insertions(+), 18 deletions(-)
diff --git a/arch/Kconfig b/arch/Kconfig
index 45c657772362..716c2134e81e 100644
--- a/arch/Kconfig
+++ b/arch/Kconfig
@@ -47,6 +47,9 @@ config ARCH_SUPPORTS_SCHED_SMT
config ARCH_SUPPORTS_SCHED_CLUSTER
bool
+config ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO
+ bool
+
config ARCH_SUPPORTS_SCHED_MC
bool
@@ -79,6 +82,17 @@ config SCHED_MC
making when dealing with multi-core CPU chips at a cost of slightly
increased overhead in some places. If unsure say N here.
+config SCHED_IDLE_SMT_PRIO
+ bool "Idle SMT priority support in presence of asymmetric capacities"
+ depends on SCHED_SMT
+ depends on ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO
+ default y
+ help
+ On asymmetric CPU capacity systems, an idle SMT sibling of a
+ busy high-capacity core can outperform a fully idle low-capacity
+ core. Allows architectures to override the schedulers default
+ idle core preference accordingly.
+
# Selected by HOTPLUG_CORE_SYNC_DEAD or HOTPLUG_CORE_SYNC_FULL
config HOTPLUG_CORE_SYNC
bool
diff --git a/include/linux/sched/topology.h b/include/linux/sched/topology.h
index b5d9d7c2b8ad..ae51e7c467f8 100644
--- a/include/linux/sched/topology.h
+++ b/include/linux/sched/topology.h
@@ -275,6 +275,15 @@ unsigned int arch_scale_freq_ref(int cpu)
}
#endif
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+#ifndef arch_needs_idle_smt_prio
+static __always_inline bool arch_needs_idle_smt_prio(void)
+{
+ return true;
+}
+#endif
+#endif
+
static inline int task_node(const struct task_struct *p)
{
return cpu_to_node(task_cpu(p));
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 7455a83a6a99..3e21233ee2ce 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8813,11 +8813,14 @@ static int
select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
{
/*
- * On !SMT systems, has_idle_core is always false and preferred_core
- * is always true (CPU == core), so the SMT preference logic below
- * collapses to the plain capacity scan.
- */
- bool has_idle_core = sched_smt_active() && test_idle_cores(target);
+ * On !SMT systems or when idle SMT thread priority is active,
+ * has_idle_core is always false and preferred_core is always true
+ * (CPU == core), so the SMT preference logic below collapses to the
+ * plain capacity scan.
+ */
+ bool has_idle_core = sched_smt_active() &&
+ !sched_idle_smt_prio_active() &&
+ test_idle_cores(target);
unsigned long task_util, util_min, util_max, best_cap = 0;
int fits, best_fits = ASYM_IDLE_THREAD_MISFIT;
int cpu, best_cpu = -1;
@@ -8860,8 +8863,8 @@ select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
/*
* Perfect fit: capacity satisfies util + uclamp and the CPU
- * sits on a fully-idle SMT core, this is a !SMT system, or
- * there is no idle core to find.
+ * sits on a fully-idle SMT core, this is a !SMT system, idle
+ * smt priority is active, or there is no idle core to find.
* Short-circuit the rank-based selection and return
* immediately.
*/
@@ -8936,15 +8939,17 @@ static inline bool asym_fits_cpu(unsigned long util,
* Return true only if the cpu fully fits the task requirements
* which include the utilization and the performance hints.
*
- * When SMT is active, also require that the core has no busy
- * siblings.
+ * When SMT is active or idle SMT priority is disabled, also
+ * require that the core has no busy siblings.
*
* Note: gating on is_core_idle() also makes the early-bailout
* candidates in select_idle_sibling() (target, prev,
* recent_used_cpu) idle-core-aware on ASYM+SMT, which the
* NO_ASYM path does not do.
*/
- return (!sched_smt_active() || is_core_idle(cpu)) &&
+ return (!sched_smt_active() ||
+ sched_idle_smt_prio_active() ||
+ is_core_idle(cpu)) &&
(util_fits_cpu(util, util_min, util_max, cpu) > 0);
}
@@ -11793,6 +11798,12 @@ static inline bool smt_balance(struct lb_env *env, struct sg_lb_stats *sgs,
if (!env->idle)
return false;
+ /* Skip group_smt_balance case when idle SMT priority is active */
+ if (sched_idle_smt_prio_active()) {
+ if (env->sd->flags & SD_ASYM_CPUCAPACITY)
+ return false;
+ }
+
/*
* For SMT source group, it is better to move a task
* to a CPU that doesn't have multiple tasks sharing its CPU capacity.
@@ -12121,10 +12132,10 @@ static bool update_sd_pick_busiest(struct lb_env *env,
* CPUs in the group should either be possible to resolve
* internally or be covered by avg_load imbalance (eventually).
*
- * When SMT is active, only pull a misfit to dst_cpu if it is on a
- * fully idle core; otherwise the effective capacity of the core is
- * reduced and we may not actually provide more capacity than the
- * source.
+ * When SMT is active or idle smt priority disabled, only pull a misfit
+ * to dst_cpu if it is on a fully idle core; otherwise the effective
+ * capacity of the core is reduced and we may not actually provide more
+ * capacity than the source.
*/
if ((env->sd->flags & SD_ASYM_CPUCAPACITY) &&
(sgs->group_type == group_misfit_task) &&
@@ -12429,6 +12440,17 @@ static bool update_pick_idlest(struct sched_group *idlest,
break;
case group_has_spare:
+ /*
+ * Select group with higher max capacity to prioritize groups
+ * with high capacity SMT siblings
+ */
+ if (sched_idle_smt_prio_active()) {
+ if (idlest->sgc->max_capacity > group->sgc->max_capacity)
+ return false;
+ if (idlest->sgc->max_capacity < group->sgc->max_capacity)
+ break;
+ }
+
/* Select group with most idle CPUs */
if (idlest_sgs->idle_cpus > sgs->idle_cpus)
return false;
@@ -12699,7 +12721,9 @@ static inline void update_sd_lb_stats(struct lb_env *env, struct sd_lb_stats *sd
unsigned long sum_util = 0;
bool sg_overloaded = 0, sg_overutilized = 0;
- env->dst_core_idle = !sched_smt_active() || is_core_idle(env->dst_cpu);
+ env->dst_core_idle = !sched_smt_active() ||
+ is_core_idle(env->dst_cpu) ||
+ sched_idle_smt_prio_active();
do {
struct sg_lb_stats *sgs = &tmp_sgs;
@@ -13171,11 +13195,14 @@ static struct rq *sched_balance_find_src_rq(struct lb_env *env,
bool cluster_equal_cap = static_branch_unlikely(&sched_cluster_active) &&
(get_actual_cpu_capacity(env->dst_cpu) ==
get_actual_cpu_capacity(i));
- bool smt_degraded_cap = sched_smt_active() && !is_core_idle(i);
+ bool smt_degraded_cap = sched_smt_active() &&
+ !is_core_idle(i) &&
+ !sched_idle_smt_prio_active();
/*
- * Busy SMT siblings reduce the capacity of CPU @i. Do
- * not skip it in this case.
+ * Busy SMT siblings reduce the capacity of CPU @i
+ * unless idle SMT priority is active. Do not skip it
+ * in this case.
*
* CONFIG_SCHED_CLUSTER requires balancing load across
* clusters of identical capacity, accounting for
@@ -13394,6 +13421,10 @@ static int should_we_balance(struct lb_env *env)
for_each_cpu_and(cpu, swb_cpus, env->cpus) {
if (!idle_cpu(cpu))
continue;
+ if (sched_idle_smt_prio_active()) {
+ if (env->sd->flags & SD_ASYM_CPUCAPACITY)
+ return cpu == env->dst_cpu;
+ }
/*
* Don't balance to idle SMT in busy core right away when
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..2d168e8565fc 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2241,12 +2241,22 @@ DECLARE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
extern struct static_key_false sched_asym_cpucapacity;
extern struct static_key_false sched_cluster_active;
+DECLARE_STATIC_KEY_FALSE(sched_idle_smt_prio);
static __always_inline bool sched_asym_cpucap_active(void)
{
return static_branch_unlikely(&sched_asym_cpucapacity);
}
+static __always_inline bool sched_idle_smt_prio_active(void)
+{
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+ return static_branch_unlikely(&sched_idle_smt_prio);
+#else
+ return false;
+#endif
+}
+
struct sched_group_capacity {
atomic_t ref;
/*
diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 0248227d983a..a66ea34a1dfd 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -684,6 +684,7 @@ DEFINE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
DEFINE_STATIC_KEY_FALSE(sched_asym_cpucapacity);
DEFINE_STATIC_KEY_FALSE(sched_cluster_active);
+DEFINE_STATIC_KEY_FALSE(sched_idle_smt_prio);
static void update_top_cache_domain(int cpu)
{
@@ -1559,6 +1560,21 @@ build_sched_groups(struct sched_domain *sd, int cpu)
return 0;
}
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+static void sched_set_idle_smt_prio(int asym_detected)
+{
+ bool smt_prio;
+
+ smt_prio = arch_needs_idle_smt_prio() && asym_detected;
+ if (smt_prio && !sched_idle_smt_prio_active())
+ static_branch_enable_cpuslocked(&sched_idle_smt_prio);
+ else if (!smt_prio && sched_idle_smt_prio_active())
+ static_branch_disable_cpuslocked(&sched_idle_smt_prio);
+}
+#else
+static void sched_set_idle_smt_prio(int asym_detected) {}
+#endif
+
/*
* Initialize sched groups cpu_capacity.
*
@@ -1769,6 +1785,7 @@ static void asym_cpu_capacity_scan(void)
}
}
+ sched_set_idle_smt_prio(!list_is_singular(&asym_cap_list));
/*
* Only one capacity value has been detected i.e. this system is symmetric.
* No need to keep this data around.
--
2.55.0
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO
2026-10-08 8:30 [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority Mete Durlu
@ 2026-10-08 8:30 ` Mete Durlu
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
2026-10-08 21:29 ` Tim Chen
3 siblings, 0 replies; 12+ messages in thread
From: Mete Durlu @ 2026-10-08 8:30 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Shrikanth Hegde,
Heiko Carstens, Vasily Gorbik, Alexander Gordeev,
Christian Borntraeger, Sven Schnelle
Cc: Tim Chen, Chen Yu, Ilya Leoshkevich, Andrea Righi, linux-kernel,
linux-s390, Mete Durlu
On s390, vertical polarization configuration assigns CPUs different
polarization values and capacities. These capacities represent the
virtual runtime being assigned to the shared CPUs by the underlying
hypervisor. Consequently running load on high capacity cores, even if
the SMT sibling (SMT=2 on s390) of a given core is busy is more
favorable compared to a fully idle low capacity CPU.
Implement arch_needs_idle_smt_prio() to enable the sched_idle_smt_prio
static branch when the machine has hardware backed topology.
Signed-off-by: Mete Durlu <meted@linux.ibm.com>
---
arch/s390/Kconfig | 1 +
arch/s390/include/asm/topology.h | 5 +++++
arch/s390/kernel/topology.c | 9 +++++++++
3 files changed, 15 insertions(+)
diff --git a/arch/s390/Kconfig b/arch/s390/Kconfig
index 4b51bc6e8948..bd6c826a24af 100644
--- a/arch/s390/Kconfig
+++ b/arch/s390/Kconfig
@@ -564,6 +564,7 @@ config NODES_SHIFT
config SCHED_TOPOLOGY
def_bool y
prompt "Topology scheduler support"
+ select ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO
select ARCH_SUPPORTS_SCHED_SMT
select ARCH_SUPPORTS_SCHED_MC
select SCHED_SMT
diff --git a/arch/s390/include/asm/topology.h b/arch/s390/include/asm/topology.h
index 44110847342a..f0f5477565db 100644
--- a/arch/s390/include/asm/topology.h
+++ b/arch/s390/include/asm/topology.h
@@ -49,6 +49,11 @@ void update_cpu_masks(void);
void topology_expect_change(void);
const struct cpumask *cpu_coregroup_mask(int cpu);
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+#define arch_needs_idle_smt_prio arch_needs_idle_smt_prio
+bool arch_needs_idle_smt_prio(void);
+#endif
+
#else /* CONFIG_SCHED_TOPOLOGY */
static inline void topology_init_early(void) { }
diff --git a/arch/s390/kernel/topology.c b/arch/s390/kernel/topology.c
index 42fc0294f543..6f28e2f1afec 100644
--- a/arch/s390/kernel/topology.c
+++ b/arch/s390/kernel/topology.c
@@ -344,6 +344,15 @@ int arch_update_cpu_topology(void)
return rc;
}
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+bool arch_needs_idle_smt_prio(void)
+{
+ if (topology_mode != TOPOLOGY_MODE_HW)
+ return false;
+ return true;
+}
+#endif
+
static void topology_work_fn(struct work_struct *work)
{
rebuild_sched_domains();
--
2.55.0
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 8:30 [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO Mete Durlu
@ 2026-10-08 12:10 ` Andrea Righi
2026-10-08 14:58 ` Shrikanth Hegde
2026-10-08 21:52 ` Tim Chen
2026-10-08 21:29 ` Tim Chen
3 siblings, 2 replies; 12+ messages in thread
From: Andrea Righi @ 2026-10-08 12:10 UTC (permalink / raw)
To: Mete Durlu
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Shrikanth Hegde,
Heiko Carstens, Vasily Gorbik, Alexander Gordeev,
Christian Borntraeger, Sven Schnelle, Tim Chen, Chen Yu,
Ilya Leoshkevich, linux-kernel, linux-s390
Hi Mete,
On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> Summary
> ===========================================================================
> On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> cores over idle SMT siblings of busy cores. This works generally well but
> virtualized platforms where low capacity cores should be avoided are not
> considered. Introduce a new config option and arch hook to prioritize
> SMT utilization.
>
> Background
> ===========================================================================
> The scheduler has been moving toward better utilization of fully idle cores
> and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> broke the behavior s390 is relying on to concentrate workloads on its
> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> entitlement assigned by the hypervisor. The idle SMT siblings of busy
> high-capacity cores start to perform better than fully idle low-capacity
> cores as the whole machine(containing the logical partitions) starts
> approaching to a {fully,over}loaded state. Grouping load on high-capacity
> cores keeps shared low-capacity cores(which are shared more aggressively)
> idle longer, reducing noise to neighboring partitions and improving
> overall performance.
>
> Approach
> ===========================================================================
> This series introduces SCHED_IDLE_SMT_PRIO config option and the
> sched_idle_smt_prio static branch, allowing architectures to treat idle
> SMT siblings of busy cores as equal candidates during asymmetric capacity
> load balancing.
> Static branch checks are placed at paths considering fully idle cores
> over idle SMT threads in presence of asymmetric CPU capacities within
> scheduling groups. Inserted checks mostly override hints for idle core
> selection or cause early exits favoring idle SMT siblings.
> One more static branch check is added to new task path in order to favor
> high capacity SMT siblings during initial task placement, therefore the
> tasks are immediately placed in high capacity SMT siblings instead of
> considering most idle but low capacity cores.
>
> Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> implementing arch_needs_idle_smt_prio(), which is evaluated during each
> asym_cpu_capacity_scan() to track the state as runtime topology changes.
>
> The second patch enables the feature for s390 when running on
> hardware-backed topology in an LPAR with vertical polarization active
> and system is approaching to a target state.
> Otherwise the branch stays disabled and the scheduler falls back to the
> standard idle-core preference.
>
> No functional change on architectures that do not select
> ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
>
> Performance Results
> ===========================================================================
> Since this change effects the {fully,over}loaded state it is difficult
> to test at the moment. Therefore there are no concrete numbers or
> metrics for now for s390.
> For architectures which do not opt in to this feature, no performance
> change is expected.
>
> Considerations and Open Questions
> ===========================================================================
> The main goal for this series is adding a simple mechanism for
> architectures to switch between idle core and idle SMT priority while
> keeping the introduced footprint as small as possible. But there are
> some ideas and questions to consider;
>
> 1. Should arch_needs_idle_smt_prio() hook be removed?
> The architecture hook is there to allow for any sort of logic to
> dynamically decide when to flip the mechanism, but it can also be
> removed if everyone agrees that this behaviour is not something that
> should be dynamically flipped. It can be simply tied to detection of
> asymmetric capacities and selection of Kconfig option.
>
> 2. Is there a simpler way to implement this mechanism?
> The proposed approach is chosen as the other features effecting the
> scheduler's behavior, use the same method. If there is a more
> efficient way to implement the same mechanism I'd be glad to use
> that instead.
>
> 3. There are no performance measurements *yet*.
> On s390 the ideal conditions for this mechanism to be beneficial
> usually surface when the whole system is fully loaded and resource
> sharing between the logical partitions starts to get expensive.
> Creating such an environment requires time, therefore no benchmark
> results are available yet.
I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
mask's current behavior, soft preference would allow tasks to spill onto other
CPUs once the preferred CPUs have no idle threads available (we have something
like this in sched_ext's scx_cosmos).
That might let us preserve the general preference for fully idle SMT cores while
searching in this order: fully idle preferred cores, idle SMT threads on
preferred cores, then non-preferred cores. This current patch series seems to
drop the first distinction, allowing a partially idle preferred core to win even
when another preferred core is fully idle.
This would need to coexist with the steal governor's stronger use of
cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
rather than a drop-in replacement.
What do you think (Mete / Shrikanth)?
Thanks,
-Andrea
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
@ 2026-10-08 14:58 ` Shrikanth Hegde
2026-10-08 15:24 ` Mete Durlu
2026-10-08 21:52 ` Tim Chen
1 sibling, 1 reply; 12+ messages in thread
From: Shrikanth Hegde @ 2026-10-08 14:58 UTC (permalink / raw)
To: Andrea Righi, Mete Durlu
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
Hi Meter/Andrea.
On 10/8/26 5:40 PM, Andrea Righi wrote:
> Hi Mete,
>
> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
>> Summary
>> ===========================================================================
>> On systems with asymmetric CPU capacities, the scheduler prefers fully idle
>> cores over idle SMT siblings of busy cores. This works generally well but
>> virtualized platforms where low capacity cores should be avoided are not
>> considered. Introduce a new config option and arch hook to prioritize
>> SMT utilization.
That's true only under physical CPU contention right? or is it always?
>>
>> Background
>> ===========================================================================
>> The scheduler has been moving toward better utilization of fully idle cores
>> and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
>> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
>> broke the behavior s390 is relying on to concentrate workloads on its
>> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
>> entitlement assigned by the hypervisor. The idle SMT siblings of busy
>> high-capacity cores start to perform better than fully idle low-capacity
>> cores as the whole machine(containing the logical partitions) starts
>> approaching to a {fully,over}loaded state. Grouping load on high-capacity
>> cores keeps shared low-capacity cores(which are shared more aggressively)
>> idle longer, reducing noise to neighboring partitions and improving
>> overall performance.
>>
>> Approach
>> ===========================================================================
>> This series introduces SCHED_IDLE_SMT_PRIO config option and the
>> sched_idle_smt_prio static branch, allowing architectures to treat idle
>> SMT siblings of busy cores as equal candidates during asymmetric capacity
>> load balancing.
>> Static branch checks are placed at paths considering fully idle cores
>> over idle SMT threads in presence of asymmetric CPU capacities within
>> scheduling groups. Inserted checks mostly override hints for idle core
>> selection or cause early exits favoring idle SMT siblings.
>> One more static branch check is added to new task path in order to favor
>> high capacity SMT siblings during initial task placement, therefore the
>> tasks are immediately placed in high capacity SMT siblings instead of
>> considering most idle but low capacity cores.
>>
>> Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
>> implementing arch_needs_idle_smt_prio(), which is evaluated during each
>> asym_cpu_capacity_scan() to track the state as runtime topology changes.
>>
>> The second patch enables the feature for s390 when running on
>> hardware-backed topology in an LPAR with vertical polarization active
>> and system is approaching to a target state.
>> Otherwise the branch stays disabled and the scheduler falls back to the
>> standard idle-core preference.
>>
>> No functional change on architectures that do not select
>> ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
>>
>> Performance Results
>> ===========================================================================
>> Since this change effects the {fully,over}loaded state it is difficult
>> to test at the moment. Therefore there are no concrete numbers or
>> metrics for now for s390.
>> For architectures which do not opt in to this feature, no performance
>> change is expected.
>>
>> Considerations and Open Questions
>> ===========================================================================
>> The main goal for this series is adding a simple mechanism for
>> architectures to switch between idle core and idle SMT priority while
>> keeping the introduced footprint as small as possible. But there are
>> some ideas and questions to consider;
>>
>> 1. Should arch_needs_idle_smt_prio() hook be removed?
>> The architecture hook is there to allow for any sort of logic to
>> dynamically decide when to flip the mechanism, but it can also be
>> removed if everyone agrees that this behaviour is not something that
>> should be dynamically flipped. It can be simply tied to detection of
>> asymmetric capacities and selection of Kconfig option.
>>
>> 2. Is there a simpler way to implement this mechanism?
>> The proposed approach is chosen as the other features effecting the
>> scheduler's behavior, use the same method. If there is a more
>> efficient way to implement the same mechanism I'd be glad to use
>> that instead.
>>
>> 3. There are no performance measurements *yet*.
>> On s390 the ideal conditions for this mechanism to be beneficial
>> usually surface when the whole system is fully loaded and resource
>> sharing between the logical partitions starts to get expensive.
>> Creating such an environment requires time, therefore no benchmark
>> results are available yet.
If it only under contention, then you probably don't want to use low cores for
anything right? If above is yes, then using steal governor would help to avoid low
core if you mark them as non-preferred. (which i guess you guys are exploring already)
But, find_new_ilb isn't aware of preferred CPU state as of initial patches.
I had thought of making a change to pick a idle preferred CPU instead of a idle non-preferred
core for idle load balancing. It is not done due to below reasons.
- Performance numbers didn't show any major improvements with real life workloads.
- Code becomes quite complex for find_new_ilb.
>
> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
> mask's current behavior, soft preference would allow tasks to spill onto other
> CPUs once the preferred CPUs have no idle threads available (we have something
> like this in sched_ext's scx_cosmos).
>
cpu_preferred infra won't allow to spill over if preferred CPUs don't have an idle CPU.
It rather enforces packing onto preferred CPUs even if it means rq has more than 1 task.
> That might let us preserve the general preference for fully idle SMT cores while
> searching in this order: fully idle preferred cores, idle SMT threads on
> preferred cores, then non-preferred cores. This current patch series seems to
> drop the first distinction, allowing a partially idle preferred core to win even
> when another preferred core is fully idle.
>
> This would need to coexist with the steal governor's stronger use of
> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> rather than a drop-in replacement.
>
> What do you think (Mete / Shrikanth)?
>
> Thanks,
> -Andrea
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 14:58 ` Shrikanth Hegde
@ 2026-10-08 15:24 ` Mete Durlu
2026-10-08 15:44 ` Shrikanth Hegde
2026-10-09 15:32 ` Vincent Guittot
0 siblings, 2 replies; 12+ messages in thread
From: Mete Durlu @ 2026-10-08 15:24 UTC (permalink / raw)
To: Shrikanth Hegde, Andrea Righi
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
On 08/10/2026 16:58, Shrikanth Hegde wrote:
> Hi Meter/Andrea.
>
> On 10/8/26 5:40 PM, Andrea Righi wrote:
>> Hi Mete,
>>
>> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
>>> Summary
>>> ===========================================================================
>>> On systems with asymmetric CPU capacities, the scheduler prefers
>>> fully idle
>>> cores over idle SMT siblings of busy cores. This works generally well
>>> but
>>> virtualized platforms where low capacity cores should be avoided are not
>>> considered. Introduce a new config option and arch hook to prioritize
>>> SMT utilization.
>
> That's true only under physical CPU contention right? or is it always?
Right, because of that s390 treats all CPUs as equal and starts to
assign lower capacities to CPUs with low entitlement once a certain
steal time threshold is crossed (aka physical CPU contention).
>>>
>>> Background
>>> ===========================================================================
>>> The scheduler has been moving toward better utilization of fully idle
>>> cores
>>> and avoiding SMT penalties. This trend (notably after commit
>>> 25a32e400a14
>>> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle
>>> selection")
>>> broke the behavior s390 is relying on to concentrate workloads on its
>>> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
>>> entitlement assigned by the hypervisor. The idle SMT siblings of busy
>>> high-capacity cores start to perform better than fully idle low-capacity
>>> cores as the whole machine(containing the logical partitions) starts
>>> approaching to a {fully,over}loaded state. Grouping load on high-
>>> capacity
>>> cores keeps shared low-capacity cores(which are shared more
>>> aggressively)
>>> idle longer, reducing noise to neighboring partitions and improving
>>> overall performance.
>>>
[...snip...]
>>> Considerations and Open Questions
>>> ===========================================================================
>>> The main goal for this series is adding a simple mechanism for
>>> architectures to switch between idle core and idle SMT priority while
>>> keeping the introduced footprint as small as possible. But there are
>>> some ideas and questions to consider;
>>>
>>> 1. Should arch_needs_idle_smt_prio() hook be removed?
>>> The architecture hook is there to allow for any sort of logic to
>>> dynamically decide when to flip the mechanism, but it can also be
>>> removed if everyone agrees that this behaviour is not something that
>>> should be dynamically flipped. It can be simply tied to detection of
>>> asymmetric capacities and selection of Kconfig option.
>>>
>>> 2. Is there a simpler way to implement this mechanism?
>>> The proposed approach is chosen as the other features effecting the
>>> scheduler's behavior, use the same method. If there is a more
>>> efficient way to implement the same mechanism I'd be glad to use
>>> that instead.
>>>
>>> 3. There are no performance measurements *yet*.
>>> On s390 the ideal conditions for this mechanism to be beneficial
>>> usually surface when the whole system is fully loaded and resource
>>> sharing between the logical partitions starts to get expensive.
>>> Creating such an environment requires time, therefore no benchmark
>>> results are available yet.
>
> If it only under contention, then you probably don't want to use low
> cores for
> anything right? If above is yes, then using steal governor would help to
> avoid low
> core if you mark them as non-preferred. (which i guess you guys are
> exploring already)
IMO it should be dynamically adjustable depending on the contention
level. But yes, on worst case low capacity cores should be avoided by
marking them as non-preferred.
> But, find_new_ilb isn't aware of preferred CPU state as of initial patches.
> I had thought of making a change to pick a idle preferred CPU instead of
> a idle non-preferred
> core for idle load balancing. It is not done due to below reasons.
> - Performance numbers didn't show any major improvements with real life
> workloads.
> - Code becomes quite complex for find_new_ilb.
IIUC, find_new_ilb() finds an idle CPU that can do load balancing
for other idle CPUs. As long as the tasks don't land on non-preferred
CPUs, where ilb occurs should not matter. (At least for s390)
>>
>> I'm wondering if we could build on Shrikanth's cpu_preferred_mask
>> infrastructure
>> to express a "soft preference" for the s390 higher-entitlement CPUs.
>> Unlike the
>> mask's current behavior, soft preference would allow tasks to spill
>> onto other
>> CPUs once the preferred CPUs have no idle threads available (we have
>> something
>> like this in sched_ext's scx_cosmos).
>>
>
> cpu_preferred infra won't allow to spill over if preferred CPUs don't
> have an idle CPU.
> It rather enforces packing onto preferred CPUs even if it means rq has
> more than 1 task.
I think what Andrea has in mind is something similar to what he did with
select_idle_capacity(). Depending on the state of the non-preferred mask
and other factors cores and SMT siblings can be evaluated as ideal
candidate somehow. Wouldn't it be possible to extend the
infrastructure you introduced to achieve this?
>> That might let us preserve the general preference for fully idle SMT
>> cores while
>> searching in this order: fully idle preferred cores, idle SMT threads on
>> preferred cores, then non-preferred cores. This current patch series
>> seems to
>> drop the first distinction, allowing a partially idle preferred core
>> to win even
>> when another preferred core is fully idle.
>>
>> This would need to coexist with the steal governor's stronger use of
>> cpu_preferred_mask, so I'm suggesting it as a potential direction to
>> explore
>> rather than a drop-in replacement.
I think this makes a lot of sense. An approach to first fill preferred
CPUs before spilling to the non-preferred cores until contention
forces the evacuation of non-prefered CPUs.
Shrikanth, would it make sense if I try to find a way to extend your
current implementation in this way?
Thanks!
-Mete
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 15:24 ` Mete Durlu
@ 2026-10-08 15:44 ` Shrikanth Hegde
2026-10-09 15:12 ` Andrea Righi
2026-10-09 15:32 ` Vincent Guittot
1 sibling, 1 reply; 12+ messages in thread
From: Shrikanth Hegde @ 2026-10-08 15:44 UTC (permalink / raw)
To: Mete Durlu, Andrea Righi
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
Hi Mete.
On 10/8/26 8:54 PM, Mete Durlu wrote:
> On 08/10/2026 16:58, Shrikanth Hegde wrote:
>> Hi Meter/Andrea.
>>
>> On 10/8/26 5:40 PM, Andrea Righi wrote:
>>> Hi Mete,
>>>
>>> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
>>>> Summary
>>>> ===========================================================================
>>>> On systems with asymmetric CPU capacities, the scheduler prefers fully idle
>>>> cores over idle SMT siblings of busy cores. This works generally well but
>>>> virtualized platforms where low capacity cores should be avoided are not
>>>> considered. Introduce a new config option and arch hook to prioritize
>>>> SMT utilization.
>>
>> That's true only under physical CPU contention right? or is it always?
>
> Right, because of that s390 treats all CPUs as equal and starts to
> assign lower capacities to CPUs with low entitlement once a certain
> steal time threshold is crossed (aka physical CPU contention).
>
>>>>
>>>> Background
>>>> ===========================================================================
>>>> The scheduler has been moving toward better utilization of fully idle cores
>>>> and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
>>>> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
>>>> broke the behavior s390 is relying on to concentrate workloads on its
>>>> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
>>>> entitlement assigned by the hypervisor. The idle SMT siblings of busy
>>>> high-capacity cores start to perform better than fully idle low-capacity
>>>> cores as the whole machine(containing the logical partitions) starts
>>>> approaching to a {fully,over}loaded state. Grouping load on high- capacity
>>>> cores keeps shared low-capacity cores(which are shared more aggressively)
>>>> idle longer, reducing noise to neighboring partitions and improving
>>>> overall performance.
>>>>
>
> [...snip...]
>
>>>> Considerations and Open Questions
>>>> ===========================================================================
>>>> The main goal for this series is adding a simple mechanism for
>>>> architectures to switch between idle core and idle SMT priority while
>>>> keeping the introduced footprint as small as possible. But there are
>>>> some ideas and questions to consider;
>>>>
>>>> 1. Should arch_needs_idle_smt_prio() hook be removed?
>>>> The architecture hook is there to allow for any sort of logic to
>>>> dynamically decide when to flip the mechanism, but it can also be
>>>> removed if everyone agrees that this behaviour is not something that
>>>> should be dynamically flipped. It can be simply tied to detection of
>>>> asymmetric capacities and selection of Kconfig option.
>>>>
>>>> 2. Is there a simpler way to implement this mechanism?
>>>> The proposed approach is chosen as the other features effecting the
>>>> scheduler's behavior, use the same method. If there is a more
>>>> efficient way to implement the same mechanism I'd be glad to use
>>>> that instead.
>>>>
>>>> 3. There are no performance measurements *yet*.
>>>> On s390 the ideal conditions for this mechanism to be beneficial
>>>> usually surface when the whole system is fully loaded and resource
>>>> sharing between the logical partitions starts to get expensive.
>>>> Creating such an environment requires time, therefore no benchmark
>>>> results are available yet.
>>
>> If it only under contention, then you probably don't want to use low cores for
>> anything right? If above is yes, then using steal governor would help to avoid low
>> core if you mark them as non-preferred. (which i guess you guys are exploring already)
>
> IMO it should be dynamically adjustable depending on the contention
> level. But yes, on worst case low capacity cores should be avoided by
> marking them as non-preferred.
>
>> But, find_new_ilb isn't aware of preferred CPU state as of initial patches.
>> I had thought of making a change to pick a idle preferred CPU instead of a idle non-preferred
>> core for idle load balancing. It is not done due to below reasons.
>> - Performance numbers didn't show any major improvements with real life workloads.
>> - Code becomes quite complex for find_new_ilb.
>
> IIUC, find_new_ilb() finds an idle CPU that can do load balancing
> for other idle CPUs. As long as the tasks don't land on non-preferred
> CPUs, where ilb occurs should not matter. (At least for s390)
>
If that is what you want, then steal governor is the right choice.
Even if idle load balancing runs on a non-preferred idle CPU, it is going
to do the idle load balancing for all the idle CPUs. Including preferred and
non-preferred. And actual load balancing bails out early and does not pull any
task towards if it is a non-preferred CPU.
>>>
>>> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
>>> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
>>> mask's current behavior, soft preference would allow tasks to spill onto other
>>> CPUs once the preferred CPUs have no idle threads available (we have something
>>> like this in sched_ext's scx_cosmos).
>>>
>>
>> cpu_preferred infra won't allow to spill over if preferred CPUs don't have an idle CPU.
>> It rather enforces packing onto preferred CPUs even if it means rq has more than 1 task.
>
> I think what Andrea has in mind is something similar to what he did with
> select_idle_capacity(). Depending on the state of the non-preferred mask
> and other factors cores and SMT siblings can be evaluated as ideal
> candidate somehow. Wouldn't it be possible to extend the
> infrastructure you introduced to achieve this?
>
>>> That might let us preserve the general preference for fully idle SMT cores while
>>> searching in this order: fully idle preferred cores, idle SMT threads on
>>> preferred cores, then non-preferred cores. This current patch series seems to
>>> drop the first distinction, allowing a partially idle preferred core to win even
>>> when another preferred core is fully idle.
>>>
>>> This would need to coexist with the steal governor's stronger use of
>>> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
>>> rather than a drop-in replacement.
>
> I think this makes a lot of sense. An approach to first fill preferred
> CPUs before spilling to the non-preferred cores until contention
> forces the evacuation of non-prefered CPUs.
Well, when we see steal time, it usually means there is high contention and
we are using more vCPUs at this moment than possible.
So we mark them as non-preferred. Note that they are idle cpus. But non-preferred.
That means preferred CPUs may have more than 1 task. If we spill over if there
is idle non-preferred CPUs, we will be back to square one. No?
>
> Shrikanth, would it make sense if I try to find a way to extend your
> current implementation in this way?
>
> Thanks!
> -Mete
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 15:44 ` Shrikanth Hegde
@ 2026-10-09 15:12 ` Andrea Righi
2026-10-09 15:35 ` Shrikanth Hegde
0 siblings, 1 reply; 12+ messages in thread
From: Andrea Righi @ 2026-10-09 15:12 UTC (permalink / raw)
To: Shrikanth Hegde
Cc: Mete Durlu, Ingo Molnar, Peter Zijlstra, Juri Lelli,
Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
On Thu, Oct 08, 2026 at 09:14:49PM +0530, Shrikanth Hegde wrote:
...
> > > > That might let us preserve the general preference for fully idle SMT cores while
> > > > searching in this order: fully idle preferred cores, idle SMT threads on
> > > > preferred cores, then non-preferred cores. This current patch series seems to
> > > > drop the first distinction, allowing a partially idle preferred core to win even
> > > > when another preferred core is fully idle.
> > > >
> > > > This would need to coexist with the steal governor's stronger use of
> > > > cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> > > > rather than a drop-in replacement.
> >
> > I think this makes a lot of sense. An approach to first fill preferred
> > CPUs before spilling to the non-preferred cores until contention
> > forces the evacuation of non-prefered CPUs.
>
> Well, when we see steal time, it usually means there is high contention and
> we are using more vCPUs at this moment than possible.
> So we mark them as non-preferred. Note that they are idle cpus. But non-preferred.
>
> That means preferred CPUs may have more than 1 task. If we spill over if there
> is idle non-preferred CPUs, we will be back to square one. No?
I agree that CPUs excluded by the steal governor should remain avoided even when
all preferred CPUs are busy.
I'm wondering if we should have two levels: cpu_preferred_mask acting as a hard
affinity constraint and an entitlement-based soft affinity within that mask? The
soft affinity would prefer high-entitlement CPUs: first fully idle cores, then
idle SMT siblings. Once those CPUs are busy, tasks could spill onto
lower-entitlement CPUs still inside cpu_preferred_mask, again preferring fully
idle cores over idle SMT siblings. If all CPUs inside cpu_preferred_mask are
busy, tasks would queue there. The soft affinity would never override the
governor's restriction by spilling onto CPUs it excluded due to contention.
Thanks,
-Andrea
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-09 15:12 ` Andrea Righi
@ 2026-10-09 15:35 ` Shrikanth Hegde
0 siblings, 0 replies; 12+ messages in thread
From: Shrikanth Hegde @ 2026-10-09 15:35 UTC (permalink / raw)
To: Andrea Righi, Mete Durlu
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
On 10/9/26 8:42 PM, Andrea Righi wrote:
> On Thu, Oct 08, 2026 at 09:14:49PM +0530, Shrikanth Hegde wrote:
> ...
>>>>> That might let us preserve the general preference for fully idle SMT cores while
>>>>> searching in this order: fully idle preferred cores, idle SMT threads on
>>>>> preferred cores, then non-preferred cores. This current patch series seems to
>>>>> drop the first distinction, allowing a partially idle preferred core to win even
>>>>> when another preferred core is fully idle.
>>>>>
>>>>> This would need to coexist with the steal governor's stronger use of
>>>>> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
>>>>> rather than a drop-in replacement.
>>>
>>> I think this makes a lot of sense. An approach to first fill preferred
>>> CPUs before spilling to the non-preferred cores until contention
>>> forces the evacuation of non-prefered CPUs.
>>
>> Well, when we see steal time, it usually means there is high contention and
>> we are using more vCPUs at this moment than possible.
>> So we mark them as non-preferred. Note that they are idle cpus. But non-preferred.
>>
>> That means preferred CPUs may have more than 1 task. If we spill over if there
>> is idle non-preferred CPUs, we will be back to square one. No?
>
> I agree that CPUs excluded by the steal governor should remain avoided even when
> all preferred CPUs are busy.
>
> I'm wondering if we should have two levels: cpu_preferred_mask acting as a hard
> affinity constraint and an entitlement-based soft affinity within that mask? The
That not what entitlement usually means. At least not PowerVM. when one creates VM,
they assign VP=Virtual Cores and EC=Entitled Cores. and Hypervisor ensure at least EC
worth of cores are always assigned to that VM. There is no need of priority within that
entitlement.
> soft affinity would prefer high-entitlement CPUs: first fully idle cores, then
> idle SMT siblings. Once those CPUs are busy, tasks could spill onto
I think when one has high steal time and hence reduced preferred CPUs, there will
likely be no fully idle cores within preferred CPUs.
Note that current logic takes out last set of CPUs and find_new_ilb iterates from
the beginning. So if there is idle core within the preferred CPUs, it will be chosen.
> lower-entitlement CPUs still inside cpu_preferred_mask, again preferring fully
> idle cores over idle SMT siblings. If all CPUs inside cpu_preferred_mask are
> busy, tasks would queue there. The soft affinity would never override the
> governor's restriction by spilling onto CPUs it excluded due to contention.
>
As we were discussing, as long as threads don't end up on non-preferred CPUs,
we might be okay. With the above, find_new_ilb will becomes very complex IMHO
and I am not sure if it is worth it.
likely what we need is,
- Fully idle preferred Core.
- idle preferred CPU.
- any idle CPU.
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 15:24 ` Mete Durlu
2026-10-08 15:44 ` Shrikanth Hegde
@ 2026-10-09 15:32 ` Vincent Guittot
1 sibling, 0 replies; 12+ messages in thread
From: Vincent Guittot @ 2026-10-09 15:32 UTC (permalink / raw)
To: Mete Durlu
Cc: Shrikanth Hegde, Andrea Righi, Ingo Molnar, Peter Zijlstra,
Juri Lelli, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, Heiko Carstens,
Vasily Gorbik, Alexander Gordeev, Christian Borntraeger,
Sven Schnelle, Tim Chen, Chen Yu, Ilya Leoshkevich, linux-kernel,
linux-s390
On Thu, 8 Oct 2026 at 17:25, Mete Durlu <meted@linux.ibm.com> wrote:
>
> On 08/10/2026 16:58, Shrikanth Hegde wrote:
> > Hi Meter/Andrea.
> >
> > On 10/8/26 5:40 PM, Andrea Righi wrote:
> >> Hi Mete,
> >>
> >> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> >>> Summary
> >>> ===========================================================================
> >>> On systems with asymmetric CPU capacities, the scheduler prefers
> >>> fully idle
> >>> cores over idle SMT siblings of busy cores. This works generally well
> >>> but
> >>> virtualized platforms where low capacity cores should be avoided are not
> >>> considered. Introduce a new config option and arch hook to prioritize
> >>> SMT utilization.
> >
> > That's true only under physical CPU contention right? or is it always?
>
> Right, because of that s390 treats all CPUs as equal and starts to
> assign lower capacities to CPUs with low entitlement once a certain
> steal time threshold is crossed (aka physical CPU contention).
I'm not sure to get what your topology is because "if s390 treats all
CPUs as equal" then you should never use select_idle_capacity()
because their capacity should be fixed and equal and
SD_ASYM_CPUCAPACITY_FULL never set.
If your hypervisor lowers the capacity once a certain threshold is
crossed, we already account for the steal time in the scheduler with
irq and remove it from the capacity available for tasks but the
original capacity doesn't change.
>
> >>>
> >>> Background
> >>> ===========================================================================
> >>> The scheduler has been moving toward better utilization of fully idle
> >>> cores
> >>> and avoiding SMT penalties. This trend (notably after commit
> >>> 25a32e400a14
> >>> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle
> >>> selection")
> >>> broke the behavior s390 is relying on to concentrate workloads on its
> >>> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> >>> entitlement assigned by the hypervisor. The idle SMT siblings of busy
> >>> high-capacity cores start to perform better than fully idle low-capacity
> >>> cores as the whole machine(containing the logical partitions) starts
> >>> approaching to a {fully,over}loaded state. Grouping load on high-
> >>> capacity
> >>> cores keeps shared low-capacity cores(which are shared more
> >>> aggressively)
> >>> idle longer, reducing noise to neighboring partitions and improving
> >>> overall performance.
> >>>
>
> [...snip...]
>
> >>> Considerations and Open Questions
> >>> ===========================================================================
> >>> The main goal for this series is adding a simple mechanism for
> >>> architectures to switch between idle core and idle SMT priority while
> >>> keeping the introduced footprint as small as possible. But there are
> >>> some ideas and questions to consider;
> >>>
> >>> 1. Should arch_needs_idle_smt_prio() hook be removed?
> >>> The architecture hook is there to allow for any sort of logic to
> >>> dynamically decide when to flip the mechanism, but it can also be
> >>> removed if everyone agrees that this behaviour is not something that
> >>> should be dynamically flipped. It can be simply tied to detection of
> >>> asymmetric capacities and selection of Kconfig option.
> >>>
> >>> 2. Is there a simpler way to implement this mechanism?
> >>> The proposed approach is chosen as the other features effecting the
> >>> scheduler's behavior, use the same method. If there is a more
> >>> efficient way to implement the same mechanism I'd be glad to use
> >>> that instead.
> >>>
> >>> 3. There are no performance measurements *yet*.
> >>> On s390 the ideal conditions for this mechanism to be beneficial
> >>> usually surface when the whole system is fully loaded and resource
> >>> sharing between the logical partitions starts to get expensive.
> >>> Creating such an environment requires time, therefore no benchmark
> >>> results are available yet.
> >
> > If it only under contention, then you probably don't want to use low
> > cores for
> > anything right? If above is yes, then using steal governor would help to
> > avoid low
> > core if you mark them as non-preferred. (which i guess you guys are
> > exploring already)
>
> IMO it should be dynamically adjustable depending on the contention
> level. But yes, on worst case low capacity cores should be avoided by
> marking them as non-preferred.
>
> > But, find_new_ilb isn't aware of preferred CPU state as of initial patches.
> > I had thought of making a change to pick a idle preferred CPU instead of
> > a idle non-preferred
> > core for idle load balancing. It is not done due to below reasons.
> > - Performance numbers didn't show any major improvements with real life
> > workloads.
> > - Code becomes quite complex for find_new_ilb.
>
> IIUC, find_new_ilb() finds an idle CPU that can do load balancing
> for other idle CPUs. As long as the tasks don't land on non-preferred
> CPUs, where ilb occurs should not matter. (At least for s390)
>
> >>
> >> I'm wondering if we could build on Shrikanth's cpu_preferred_mask
> >> infrastructure
> >> to express a "soft preference" for the s390 higher-entitlement CPUs.
> >> Unlike the
> >> mask's current behavior, soft preference would allow tasks to spill
> >> onto other
> >> CPUs once the preferred CPUs have no idle threads available (we have
> >> something
> >> like this in sched_ext's scx_cosmos).
> >>
> >
> > cpu_preferred infra won't allow to spill over if preferred CPUs don't
> > have an idle CPU.
> > It rather enforces packing onto preferred CPUs even if it means rq has
> > more than 1 task.
>
> I think what Andrea has in mind is something similar to what he did with
> select_idle_capacity(). Depending on the state of the non-preferred mask
> and other factors cores and SMT siblings can be evaluated as ideal
> candidate somehow. Wouldn't it be possible to extend the
> infrastructure you introduced to achieve this?
>
> >> That might let us preserve the general preference for fully idle SMT
> >> cores while
> >> searching in this order: fully idle preferred cores, idle SMT threads on
> >> preferred cores, then non-preferred cores. This current patch series
> >> seems to
> >> drop the first distinction, allowing a partially idle preferred core
> >> to win even
> >> when another preferred core is fully idle.
> >>
> >> This would need to coexist with the steal governor's stronger use of
> >> cpu_preferred_mask, so I'm suggesting it as a potential direction to
> >> explore
> >> rather than a drop-in replacement.
>
> I think this makes a lot of sense. An approach to first fill preferred
> CPUs before spilling to the non-preferred cores until contention
> forces the evacuation of non-prefered CPUs.
>
> Shrikanth, would it make sense if I try to find a way to extend your
> current implementation in this way?
>
> Thanks!
> -Mete
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
2026-10-08 14:58 ` Shrikanth Hegde
@ 2026-10-08 21:52 ` Tim Chen
1 sibling, 0 replies; 12+ messages in thread
From: Tim Chen @ 2026-10-08 21:52 UTC (permalink / raw)
To: Andrea Righi, Mete Durlu
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Shrikanth Hegde,
Heiko Carstens, Vasily Gorbik, Alexander Gordeev,
Christian Borntraeger, Sven Schnelle, Chen Yu, Ilya Leoshkevich,
linux-kernel, linux-s390
On Thu, 2026-10-08 at 14:10 +0200, Andrea Righi wrote:
> Hi Mete,
>
> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> > Summary
> > ===========================================================================
> > On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> > cores over idle SMT siblings of busy cores. This works generally well but
> > virtualized platforms where low capacity cores should be avoided are not
> > considered. Introduce a new config option and arch hook to prioritize
> > SMT utilization.
> >
> > Background
> > ===========================================================================
> > The scheduler has been moving toward better utilization of fully idle cores
> > and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> > ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> > broke the behavior s390 is relying on to concentrate workloads on its
> > high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> > entitlement assigned by the hypervisor. The idle SMT siblings of busy
> > high-capacity cores start to perform better than fully idle low-capacity
> > cores as the whole machine(containing the logical partitions) starts
> > approaching to a {fully,over}loaded state. Grouping load on high-capacity
> > cores keeps shared low-capacity cores(which are shared more aggressively)
> > idle longer, reducing noise to neighboring partitions and improving
> > overall performance.
> >
> > Approach
> > ===========================================================================
> > This series introduces SCHED_IDLE_SMT_PRIO config option and the
> > sched_idle_smt_prio static branch, allowing architectures to treat idle
> > SMT siblings of busy cores as equal candidates during asymmetric capacity
> > load balancing.
> > Static branch checks are placed at paths considering fully idle cores
> > over idle SMT threads in presence of asymmetric CPU capacities within
> > scheduling groups. Inserted checks mostly override hints for idle core
> > selection or cause early exits favoring idle SMT siblings.
> > One more static branch check is added to new task path in order to favor
> > high capacity SMT siblings during initial task placement, therefore the
> > tasks are immediately placed in high capacity SMT siblings instead of
> > considering most idle but low capacity cores.
> >
> > Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> > implementing arch_needs_idle_smt_prio(), which is evaluated during each
> > asym_cpu_capacity_scan() to track the state as runtime topology changes.
> >
> > The second patch enables the feature for s390 when running on
> > hardware-backed topology in an LPAR with vertical polarization active
> > and system is approaching to a target state.
> > Otherwise the branch stays disabled and the scheduler falls back to the
> > standard idle-core preference.
> >
> > No functional change on architectures that do not select
> > ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
> >
> > Performance Results
> > ===========================================================================
> > Since this change effects the {fully,over}loaded state it is difficult
> > to test at the moment. Therefore there are no concrete numbers or
> > metrics for now for s390.
> > For architectures which do not opt in to this feature, no performance
> > change is expected.
> >
> > Considerations and Open Questions
> > ===========================================================================
> > The main goal for this series is adding a simple mechanism for
> > architectures to switch between idle core and idle SMT priority while
> > keeping the introduced footprint as small as possible. But there are
> > some ideas and questions to consider;
> >
> > 1. Should arch_needs_idle_smt_prio() hook be removed?
> > The architecture hook is there to allow for any sort of logic to
> > dynamically decide when to flip the mechanism, but it can also be
> > removed if everyone agrees that this behaviour is not something that
> > should be dynamically flipped. It can be simply tied to detection of
> > asymmetric capacities and selection of Kconfig option.
> >
> > 2. Is there a simpler way to implement this mechanism?
> > The proposed approach is chosen as the other features effecting the
> > scheduler's behavior, use the same method. If there is a more
> > efficient way to implement the same mechanism I'd be glad to use
> > that instead.
> >
> > 3. There are no performance measurements *yet*.
> > On s390 the ideal conditions for this mechanism to be beneficial
> > usually surface when the whole system is fully loaded and resource
> > sharing between the logical partitions starts to get expensive.
> > Creating such an environment requires time, therefore no benchmark
> > results are available yet.
>
> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
> mask's current behavior, soft preference would allow tasks to spill onto other
> CPUs once the preferred CPUs have no idle threads available (we have something
> like this in sched_ext's scx_cosmos).
Soft preference is okay if contention is low.
But if steal% remains high, we may still need to transition to hard
boundaries/preferences on the cores to run on to bring steal% down.
Tim
>
> That might let us preserve the general preference for fully idle SMT cores while
> searching in this order: fully idle preferred cores, idle SMT threads on
> preferred cores, then non-preferred cores. This current patch series seems to
> drop the first distinction, allowing a partially idle preferred core to win even
> when another preferred core is fully idle.
>
> This would need to coexist with the steal governor's stronger use of
> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> rather than a drop-in replacement.
>
> What do you think (Mete / Shrikanth)?
>
> Thanks,
> -Andrea
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
2026-10-08 8:30 [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Mete Durlu
` (2 preceding siblings ...)
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
@ 2026-10-08 21:29 ` Tim Chen
3 siblings, 0 replies; 12+ messages in thread
From: Tim Chen @ 2026-10-08 21:29 UTC (permalink / raw)
To: Mete Durlu, Ingo Molnar, Peter Zijlstra, Juri Lelli,
Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, Shrikanth Hegde,
Heiko Carstens, Vasily Gorbik, Alexander Gordeev,
Christian Borntraeger, Sven Schnelle
Cc: Chen Yu, Ilya Leoshkevich, Andrea Righi, linux-kernel, linux-s390
On Thu, 2026-10-08 at 10:30 +0200, Mete Durlu wrote:
> Summary
> ===========================================================================
> On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> cores over idle SMT siblings of busy cores. This works generally well but
> virtualized platforms where low capacity cores should be avoided are not
> considered. Introduce a new config option and arch hook to prioritize
> SMT utilization.
>
> Background
> ===========================================================================
> The scheduler has been moving toward better utilization of fully idle cores
> and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> broke the behavior s390 is relying on to concentrate workloads on its
> high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> entitlement assigned by the hypervisor. The idle SMT siblings of busy
> high-capacity cores start to perform better than fully idle low-capacity
> cores as the whole machine(containing the logical partitions) starts
> approaching to a {fully,over}loaded state. Grouping load on high-capacity
> cores keeps shared low-capacity cores(which are shared more aggressively)
> idle longer, reducing noise to neighboring partitions and improving
> overall performance.
>
> Approach
> ===========================================================================
> This series introduces SCHED_IDLE_SMT_PRIO config option and the
> sched_idle_smt_prio static branch, allowing architectures to treat idle
> SMT siblings of busy cores as equal candidates during asymmetric capacity
> load balancing.
> Static branch checks are placed at paths considering fully idle cores
> over idle SMT threads in presence of asymmetric CPU capacities within
> scheduling groups. Inserted checks mostly override hints for idle core
> selection or cause early exits favoring idle SMT siblings.
> One more static branch check is added to new task path in order to favor
> high capacity SMT siblings during initial task placement, therefore the
> tasks are immediately placed in high capacity SMT siblings instead of
> considering most idle but low capacity cores.
>
> Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> implementing arch_needs_idle_smt_prio(), which is evaluated during each
> asym_cpu_capacity_scan() to track the state as runtime topology changes.
>
> The second patch enables the feature for s390 when running on
> hardware-backed topology in an LPAR with vertical polarization active
> and system is approaching to a target state.
> Otherwise the branch stays disabled and the scheduler falls back to the
> standard idle-core preference.
>
> No functional change on architectures that do not select
> ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
>
> Performance Results
> ===========================================================================
> Since this change effects the {fully,over}loaded state it is difficult
> to test at the moment. Therefore there are no concrete numbers or
> metrics for now for s390.
> For architectures which do not opt in to this feature, no performance
> change is expected.
>
> Considerations and Open Questions
> ===========================================================================
> The main goal for this series is adding a simple mechanism for
> architectures to switch between idle core and idle SMT priority while
> keeping the introduced footprint as small as possible. But there are
> some ideas and questions to consider;
>
> 1. Should arch_needs_idle_smt_prio() hook be removed?
Hi Mete,
One assumption in the series should be made explicit, and I think
it answers your question 1.
The idea only works if two CPUs that the guest sees as SMT siblings
really run on the same physical core. "An idle sibling of a busy
high-capacity core" is cheaper than a fully idle low-capacity core
only because it uses a physical core the partition already has. That
holds on an LPAR, where hypervisor dispatches whole cores. It does not
hold in general for a guest whose hypervisor schedules vCPUs one at a
time, such as a typical KVM guest with "threads=2". There the guest's
sibling masks are just a description, and packing onto "siblings"
gives no benefit.
Nothing in the series checks this, and it is not documented:
- In patch 1, the default arch_needs_idle_smt_prio() returns true,
and the Kconfig help does not mention that siblings must share a
physical core. Any arch that selects
ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO takes this on without being told.
- In patch 2, the hook checks only the topology mode:
> +bool arch_needs_idle_smt_prio(void)
> +{
> + if (topology_mode != TOPOLOGY_MODE_HW)
> + return false;
> + return true;
> +}
Today s390 seems to be safe because sibling masks are
only built in HW mode. It is the only place where this
precondition can be enforced.
So I would keep the hook. It is the only place where this
precondition can be enforced. Removing it and tying the feature to
"Kconfig + asymmetric capacities" would drop the check entirely.
Concretely:
1) Document the contract in the Kconfig help and above the hook,
for example: "Only select this if CPUs reported as SMT siblings
are always dispatched on the same physical core."
2) Make the s390 hook test the conditions from the cover letter,
something like:
return machine_is_lpar() &&
topology_mode == TOPOLOGY_MODE_HW &&
smp_cpu_mtid;
Tim
> The architecture hook is there to allow for any sort of logic to
> dynamically decide when to flip the mechanism, but it can also be
> removed if everyone agrees that this behaviour is not something that
> should be dynamically flipped. It can be simply tied to detection of
> asymmetric capacities and selection of Kconfig option.
>
> 2. Is there a simpler way to implement this mechanism?
> The proposed approach is chosen as the other features effecting the
> scheduler's behavior, use the same method. If there is a more
> efficient way to implement the same mechanism I'd be glad to use
> that instead.
>
> 3. There are no performance measurements *yet*.
> On s390 the ideal conditions for this mechanism to be beneficial
> usually surface when the whole system is fully loaded and resource
> sharing between the logical partitions starts to get expensive.
> Creating such an environment requires time, therefore no benchmark
> results are available yet.
>
> ---
> base-commit: 587858367581b9c55c3690f4e63382ad622719d4
>
> Mete Durlu (2):
> kernel/sched: Introduce idle SMT priority
> s390/topology: Enable SCHED_IDLE_SMT_PRIO
>
> arch/Kconfig | 14 +++++++++
> arch/s390/Kconfig | 1 +
> arch/s390/include/asm/topology.h | 5 +++
> arch/s390/kernel/topology.c | 9 ++++++
> include/linux/sched/topology.h | 9 ++++++
> kernel/sched/fair.c | 67 +++++++++++++++++++++++++++++-----------
> kernel/sched/sched.h | 10 ++++++
> kernel/sched/topology.c | 17 ++++++++++
> 8 files changed, 114 insertions(+), 18 deletions(-)
^ permalink raw reply [flat|nested] 12+ messages in thread