From: Andrea Righi <arighi@nvidia.com>
To: Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Vincent Guittot <vincent.guittot@linaro.org>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Joel Fernandes <joelagnelf@nvidia.com>,
linux-kernel@vger.kernel.org
Subject: [PATCH] sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection
Date: Wed, 18 Mar 2026 10:22:14 +0100 [thread overview]
Message-ID: <20260318092214.130908-1-arighi@nvidia.com> (raw)
On systems with asymmetric CPU capacity (e.g., ACPI/CPPC reporting
different per-core frequencies), the wakeup path uses
select_idle_capacity() and prioritizes idle CPUs with higher capacity
for better task placement. However, when those CPUs belong to SMT cores,
their effective capacity can be much lower than the nominal capacity
when the sibling thread is busy: SMT siblings compete for shared
resources, so a "high capacity" CPU that is idle but whose sibling is
busy does not deliver its full capacity. This effective capacity
reduction cannot be modeled by the static capacity value alone.
Introduce SMT awareness in the asym-capacity idle selection policy: when
SMT is active prefer fully-idle SMT cores over partially-idle ones. A
two-phase selection first tries only CPUs on fully idle cores, then
falls back to any idle CPU if none fit.
Prioritizing fully-idle SMT cores yields better task placement because
the effective capacity of partially-idle SMT cores is reduced; always
preferring them when available leads to more accurate capacity usage on
task wakeup.
On an SMT system with asymmetric CPU capacities, SMT-aware idle
selection has been shown to improve throughput by around 15-18% for
CPU-bound workloads, running an amount of tasks equal to the amount of
SMT cores.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/fair.c | 24 +++++++++++++++++++++---
1 file changed, 21 insertions(+), 3 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 0a35a82e47920..0f97c44d4606b 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -7945,9 +7945,13 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool
* Scan the asym_capacity domain for idle CPUs; pick the first idle one on which
* the task fits. If no CPU is big enough, but there are idle ones, try to
* maximize capacity.
+ *
+ * When @smt_idle_only is true (asym + SMT), only consider CPUs on cores whose
+ * SMT siblings are all idle, to avoid stacking and sharing SMT resources.
*/
static int
-select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
+select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target,
+ bool smt_idle_only)
{
unsigned long task_util, util_min, util_max, best_cap = 0;
int fits, best_fits = 0;
@@ -7967,6 +7971,9 @@ select_idle_capacity(struct task_struct *p, struct sched_domain *sd, int target)
if (!choose_idle_cpu(cpu, p))
continue;
+ if (smt_idle_only && !is_core_idle(cpu))
+ continue;
+
fits = util_fits_cpu(task_util, util_min, util_max, cpu);
/* This CPU fits with all requirements */
@@ -8102,8 +8109,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
* capacity path.
*/
if (sd) {
- i = select_idle_capacity(p, sd, target);
- return ((unsigned)i < nr_cpumask_bits) ? i : target;
+ /*
+ * When asym + SMT and the hint says idle cores exist,
+ * try idle cores first to avoid stacking on SMT; else
+ * scan all idle CPUs.
+ */
+ if (sched_smt_active() && test_idle_cores(target)) {
+ i = select_idle_capacity(p, sd, target, true);
+ if ((unsigned int)i >= nr_cpumask_bits)
+ i = select_idle_capacity(p, sd, target, false);
+ } else {
+ i = select_idle_capacity(p, sd, target, false);
+ }
+ return ((unsigned int)i < nr_cpumask_bits) ? i : target;
}
}
--
2.53.0
next reply other threads:[~2026-03-18 9:22 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-03-18 9:22 Andrea Righi [this message]
2026-03-18 9:41 ` Vincent Guittot
2026-03-18 10:31 ` Andrea Righi
2026-03-18 15:43 ` Christian Loehle
2026-03-18 17:09 ` Andrea Righi
2026-03-19 7:20 ` Vincent Guittot
2026-03-19 8:45 ` Andrea Righi
2026-03-19 11:58 ` Christian Loehle
2026-03-19 14:00 ` Andrea Righi
2026-03-19 7:17 ` Vincent Guittot
2026-03-19 11:11 ` Andrea Righi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260318092214.130908-1-arighi@nvidia.com \
--to=arighi@nvidia.com \
--cc=bsegall@google.com \
--cc=dietmar.eggemann@arm.com \
--cc=joelagnelf@nvidia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®