From: Mete Durlu <meted@linux.ibm.com>
To: Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Heiko Carstens <hca@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Alexander Gordeev <agordeev@linux.ibm.com>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Sven Schnelle <svens@linux.ibm.com>
Cc: Tim Chen <tim.c.chen@linux.intel.com>,
Chen Yu <yu.c.chen@intel.com>,
Ilya Leoshkevich <iii@linux.ibm.com>,
Andrea Righi <arighi@nvidia.com>,
linux-kernel@vger.kernel.org, linux-s390@vger.kernel.org,
Mete Durlu <meted@linux.ibm.com>
Subject: [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority
Date: Thu, 08 Oct 2026 10:30:47 +0200 [thread overview]
Message-ID: <20261008-hiperdispatchfix-v1-1-73fe41081070@linux.ibm.com> (raw)
In-Reply-To: <20261008-hiperdispatchfix-v1-0-73fe41081070@linux.ibm.com>
On systems with asymmetric CPU capacities and SMT, the scheduler
currently prefers fully idle cores over idle SMT siblings of busy
cores. This leads tasks onto lower capacity idle cores when a
higher-capacity core has an idle sibling thread available. This hurts
performance in cases where an SMT sibling of a busy but high capacity
core can outperform a fully idle but low capacity core.
Introduce SCHED_IDLE_SMT_PRIO and the sched_idle_smt_prio static branch
to allow architectures to override this preference. When enabled, idle
SMT siblings of busy cores are treated as regular candidates.
The static branch defaults to false and is opt-in per architecture, so
systems that do not select ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO see no
change in behavior.
Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
implementing arch_needs_idle_smt_prio(), which is called during each
asym_cpu_capacity_scan(). In case there is no asymmetric capacities
static branches get disabled.
Signed-off-by: Mete Durlu <meted@linux.ibm.com>
---
arch/Kconfig | 14 +++++++++
include/linux/sched/topology.h | 9 ++++++
kernel/sched/fair.c | 67 ++++++++++++++++++++++++++++++------------
kernel/sched/sched.h | 10 +++++++
kernel/sched/topology.c | 17 +++++++++++
5 files changed, 99 insertions(+), 18 deletions(-)
diff --git a/arch/Kconfig b/arch/Kconfig
index 45c657772362..716c2134e81e 100644
--- a/arch/Kconfig
+++ b/arch/Kconfig
@@ -47,6 +47,9 @@ config ARCH_SUPPORTS_SCHED_SMT
config ARCH_SUPPORTS_SCHED_CLUSTER
bool
+config ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO
+ bool
+
config ARCH_SUPPORTS_SCHED_MC
bool
@@ -79,6 +82,17 @@ config SCHED_MC
making when dealing with multi-core CPU chips at a cost of slightly
increased overhead in some places. If unsure say N here.
+config SCHED_IDLE_SMT_PRIO
+ bool "Idle SMT priority support in presence of asymmetric capacities"
+ depends on SCHED_SMT
+ depends on ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO
+ default y
+ help
+ On asymmetric CPU capacity systems, an idle SMT sibling of a
+ busy high-capacity core can outperform a fully idle low-capacity
+ core. Allows architectures to override the schedulers default
+ idle core preference accordingly.
+
# Selected by HOTPLUG_CORE_SYNC_DEAD or HOTPLUG_CORE_SYNC_FULL
config HOTPLUG_CORE_SYNC
bool
diff --git a/include/linux/sched/topology.h b/include/linux/sched/topology.h
index b5d9d7c2b8ad..ae51e7c467f8 100644
--- a/include/linux/sched/topology.h
+++ b/include/linux/sched/topology.h
@@ -275,6 +275,15 @@ unsigned int arch_scale_freq_ref(int cpu)
}
#endif
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+#ifndef arch_needs_idle_smt_prio
+static __always_inline bool arch_needs_idle_smt_prio(void)
+{
+ return true;
+}
+#endif
+#endif
+
static inline int task_node(const struct task_struct *p)
{
return cpu_to_node(task_cpu(p));
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 7455a83a6a99..3e21233ee2ce 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8813,11 +8813,14 @@ static int
select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
{
/*
- * On !SMT systems, has_idle_core is always false and preferred_core
- * is always true (CPU == core), so the SMT preference logic below
- * collapses to the plain capacity scan.
- */
- bool has_idle_core = sched_smt_active() && test_idle_cores(target);
+ * On !SMT systems or when idle SMT thread priority is active,
+ * has_idle_core is always false and preferred_core is always true
+ * (CPU == core), so the SMT preference logic below collapses to the
+ * plain capacity scan.
+ */
+ bool has_idle_core = sched_smt_active() &&
+ !sched_idle_smt_prio_active() &&
+ test_idle_cores(target);
unsigned long task_util, util_min, util_max, best_cap = 0;
int fits, best_fits = ASYM_IDLE_THREAD_MISFIT;
int cpu, best_cpu = -1;
@@ -8860,8 +8863,8 @@ select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
/*
* Perfect fit: capacity satisfies util + uclamp and the CPU
- * sits on a fully-idle SMT core, this is a !SMT system, or
- * there is no idle core to find.
+ * sits on a fully-idle SMT core, this is a !SMT system, idle
+ * smt priority is active, or there is no idle core to find.
* Short-circuit the rank-based selection and return
* immediately.
*/
@@ -8936,15 +8939,17 @@ static inline bool asym_fits_cpu(unsigned long util,
* Return true only if the cpu fully fits the task requirements
* which include the utilization and the performance hints.
*
- * When SMT is active, also require that the core has no busy
- * siblings.
+ * When SMT is active or idle SMT priority is disabled, also
+ * require that the core has no busy siblings.
*
* Note: gating on is_core_idle() also makes the early-bailout
* candidates in select_idle_sibling() (target, prev,
* recent_used_cpu) idle-core-aware on ASYM+SMT, which the
* NO_ASYM path does not do.
*/
- return (!sched_smt_active() || is_core_idle(cpu)) &&
+ return (!sched_smt_active() ||
+ sched_idle_smt_prio_active() ||
+ is_core_idle(cpu)) &&
(util_fits_cpu(util, util_min, util_max, cpu) > 0);
}
@@ -11793,6 +11798,12 @@ static inline bool smt_balance(struct lb_env *env, struct sg_lb_stats *sgs,
if (!env->idle)
return false;
+ /* Skip group_smt_balance case when idle SMT priority is active */
+ if (sched_idle_smt_prio_active()) {
+ if (env->sd->flags & SD_ASYM_CPUCAPACITY)
+ return false;
+ }
+
/*
* For SMT source group, it is better to move a task
* to a CPU that doesn't have multiple tasks sharing its CPU capacity.
@@ -12121,10 +12132,10 @@ static bool update_sd_pick_busiest(struct lb_env *env,
* CPUs in the group should either be possible to resolve
* internally or be covered by avg_load imbalance (eventually).
*
- * When SMT is active, only pull a misfit to dst_cpu if it is on a
- * fully idle core; otherwise the effective capacity of the core is
- * reduced and we may not actually provide more capacity than the
- * source.
+ * When SMT is active or idle smt priority disabled, only pull a misfit
+ * to dst_cpu if it is on a fully idle core; otherwise the effective
+ * capacity of the core is reduced and we may not actually provide more
+ * capacity than the source.
*/
if ((env->sd->flags & SD_ASYM_CPUCAPACITY) &&
(sgs->group_type == group_misfit_task) &&
@@ -12429,6 +12440,17 @@ static bool update_pick_idlest(struct sched_group *idlest,
break;
case group_has_spare:
+ /*
+ * Select group with higher max capacity to prioritize groups
+ * with high capacity SMT siblings
+ */
+ if (sched_idle_smt_prio_active()) {
+ if (idlest->sgc->max_capacity > group->sgc->max_capacity)
+ return false;
+ if (idlest->sgc->max_capacity < group->sgc->max_capacity)
+ break;
+ }
+
/* Select group with most idle CPUs */
if (idlest_sgs->idle_cpus > sgs->idle_cpus)
return false;
@@ -12699,7 +12721,9 @@ static inline void update_sd_lb_stats(struct lb_env *env, struct sd_lb_stats *sd
unsigned long sum_util = 0;
bool sg_overloaded = 0, sg_overutilized = 0;
- env->dst_core_idle = !sched_smt_active() || is_core_idle(env->dst_cpu);
+ env->dst_core_idle = !sched_smt_active() ||
+ is_core_idle(env->dst_cpu) ||
+ sched_idle_smt_prio_active();
do {
struct sg_lb_stats *sgs = &tmp_sgs;
@@ -13171,11 +13195,14 @@ static struct rq *sched_balance_find_src_rq(struct lb_env *env,
bool cluster_equal_cap = static_branch_unlikely(&sched_cluster_active) &&
(get_actual_cpu_capacity(env->dst_cpu) ==
get_actual_cpu_capacity(i));
- bool smt_degraded_cap = sched_smt_active() && !is_core_idle(i);
+ bool smt_degraded_cap = sched_smt_active() &&
+ !is_core_idle(i) &&
+ !sched_idle_smt_prio_active();
/*
- * Busy SMT siblings reduce the capacity of CPU @i. Do
- * not skip it in this case.
+ * Busy SMT siblings reduce the capacity of CPU @i
+ * unless idle SMT priority is active. Do not skip it
+ * in this case.
*
* CONFIG_SCHED_CLUSTER requires balancing load across
* clusters of identical capacity, accounting for
@@ -13394,6 +13421,10 @@ static int should_we_balance(struct lb_env *env)
for_each_cpu_and(cpu, swb_cpus, env->cpus) {
if (!idle_cpu(cpu))
continue;
+ if (sched_idle_smt_prio_active()) {
+ if (env->sd->flags & SD_ASYM_CPUCAPACITY)
+ return cpu == env->dst_cpu;
+ }
/*
* Don't balance to idle SMT in busy core right away when
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..2d168e8565fc 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2241,12 +2241,22 @@ DECLARE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
extern struct static_key_false sched_asym_cpucapacity;
extern struct static_key_false sched_cluster_active;
+DECLARE_STATIC_KEY_FALSE(sched_idle_smt_prio);
static __always_inline bool sched_asym_cpucap_active(void)
{
return static_branch_unlikely(&sched_asym_cpucapacity);
}
+static __always_inline bool sched_idle_smt_prio_active(void)
+{
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+ return static_branch_unlikely(&sched_idle_smt_prio);
+#else
+ return false;
+#endif
+}
+
struct sched_group_capacity {
atomic_t ref;
/*
diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 0248227d983a..a66ea34a1dfd 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -684,6 +684,7 @@ DEFINE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
DEFINE_STATIC_KEY_FALSE(sched_asym_cpucapacity);
DEFINE_STATIC_KEY_FALSE(sched_cluster_active);
+DEFINE_STATIC_KEY_FALSE(sched_idle_smt_prio);
static void update_top_cache_domain(int cpu)
{
@@ -1559,6 +1560,21 @@ build_sched_groups(struct sched_domain *sd, int cpu)
return 0;
}
+#ifdef CONFIG_SCHED_IDLE_SMT_PRIO
+static void sched_set_idle_smt_prio(int asym_detected)
+{
+ bool smt_prio;
+
+ smt_prio = arch_needs_idle_smt_prio() && asym_detected;
+ if (smt_prio && !sched_idle_smt_prio_active())
+ static_branch_enable_cpuslocked(&sched_idle_smt_prio);
+ else if (!smt_prio && sched_idle_smt_prio_active())
+ static_branch_disable_cpuslocked(&sched_idle_smt_prio);
+}
+#else
+static void sched_set_idle_smt_prio(int asym_detected) {}
+#endif
+
/*
* Initialize sched groups cpu_capacity.
*
@@ -1769,6 +1785,7 @@ static void asym_cpu_capacity_scan(void)
}
}
+ sched_set_idle_smt_prio(!list_is_singular(&asym_cap_list));
/*
* Only one capacity value has been detected i.e. this system is symmetric.
* No need to keep this data around.
--
2.55.0
next prev parent reply other threads:[~2026-10-08 8:31 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 8:30 [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Mete Durlu
2026-10-08 8:30 ` Mete Durlu [this message]
2026-10-08 8:30 ` [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO Mete Durlu
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
2026-10-08 14:58 ` Shrikanth Hegde
2026-10-08 15:24 ` Mete Durlu
2026-10-08 15:44 ` Shrikanth Hegde
2026-10-09 15:12 ` Andrea Righi
2026-10-09 15:35 ` Shrikanth Hegde
2026-10-09 15:32 ` Vincent Guittot
2026-10-08 21:52 ` Tim Chen
2026-10-08 21:29 ` Tim Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261008-hiperdispatchfix-v1-1-73fe41081070@linux.ibm.com \
--to=meted@linux.ibm.com \
--cc=agordeev@linux.ibm.com \
--cc=arighi@nvidia.com \
--cc=borntraeger@linux.ibm.com \
--cc=bsegall@google.com \
--cc=dietmar.eggemann@arm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=iii@linux.ibm.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=svens@linux.ibm.com \
--cc=tim.c.chen@linux.intel.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®