From: Tim Chen <tim.c.chen@linux.intel.com>
To: Andrea Righi <arighi@nvidia.com>, Mete Durlu <meted@linux.ibm.com>
Cc: Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Heiko Carstens <hca@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Alexander Gordeev <agordeev@linux.ibm.com>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Sven Schnelle <svens@linux.ibm.com>,
Chen Yu <yu.c.chen@intel.com>,
Ilya Leoshkevich <iii@linux.ibm.com>,
linux-kernel@vger.kernel.org, linux-s390@vger.kernel.org
Subject: Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
Date: Thu, 08 Oct 2026 14:52:14 -0700 [thread overview]
Message-ID: <30e586ab643a8911f6c3cd930796292a7da460ef.camel@linux.intel.com> (raw)
In-Reply-To: <aseIPhrDGCYFtJNE@gpd4>
On Thu, 2026-10-08 at 14:10 +0200, Andrea Righi wrote:
> Hi Mete,
>
> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> > Summary
> > ===========================================================================
> > On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> > cores over idle SMT siblings of busy cores. This works generally well but
> > virtualized platforms where low capacity cores should be avoided are not
> > considered. Introduce a new config option and arch hook to prioritize
> > SMT utilization.
> >
> > Background
> > ===========================================================================
> > The scheduler has been moving toward better utilization of fully idle cores
> > and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> > ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> > broke the behavior s390 is relying on to concentrate workloads on its
> > high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> > entitlement assigned by the hypervisor. The idle SMT siblings of busy
> > high-capacity cores start to perform better than fully idle low-capacity
> > cores as the whole machine(containing the logical partitions) starts
> > approaching to a {fully,over}loaded state. Grouping load on high-capacity
> > cores keeps shared low-capacity cores(which are shared more aggressively)
> > idle longer, reducing noise to neighboring partitions and improving
> > overall performance.
> >
> > Approach
> > ===========================================================================
> > This series introduces SCHED_IDLE_SMT_PRIO config option and the
> > sched_idle_smt_prio static branch, allowing architectures to treat idle
> > SMT siblings of busy cores as equal candidates during asymmetric capacity
> > load balancing.
> > Static branch checks are placed at paths considering fully idle cores
> > over idle SMT threads in presence of asymmetric CPU capacities within
> > scheduling groups. Inserted checks mostly override hints for idle core
> > selection or cause early exits favoring idle SMT siblings.
> > One more static branch check is added to new task path in order to favor
> > high capacity SMT siblings during initial task placement, therefore the
> > tasks are immediately placed in high capacity SMT siblings instead of
> > considering most idle but low capacity cores.
> >
> > Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> > implementing arch_needs_idle_smt_prio(), which is evaluated during each
> > asym_cpu_capacity_scan() to track the state as runtime topology changes.
> >
> > The second patch enables the feature for s390 when running on
> > hardware-backed topology in an LPAR with vertical polarization active
> > and system is approaching to a target state.
> > Otherwise the branch stays disabled and the scheduler falls back to the
> > standard idle-core preference.
> >
> > No functional change on architectures that do not select
> > ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
> >
> > Performance Results
> > ===========================================================================
> > Since this change effects the {fully,over}loaded state it is difficult
> > to test at the moment. Therefore there are no concrete numbers or
> > metrics for now for s390.
> > For architectures which do not opt in to this feature, no performance
> > change is expected.
> >
> > Considerations and Open Questions
> > ===========================================================================
> > The main goal for this series is adding a simple mechanism for
> > architectures to switch between idle core and idle SMT priority while
> > keeping the introduced footprint as small as possible. But there are
> > some ideas and questions to consider;
> >
> > 1. Should arch_needs_idle_smt_prio() hook be removed?
> > The architecture hook is there to allow for any sort of logic to
> > dynamically decide when to flip the mechanism, but it can also be
> > removed if everyone agrees that this behaviour is not something that
> > should be dynamically flipped. It can be simply tied to detection of
> > asymmetric capacities and selection of Kconfig option.
> >
> > 2. Is there a simpler way to implement this mechanism?
> > The proposed approach is chosen as the other features effecting the
> > scheduler's behavior, use the same method. If there is a more
> > efficient way to implement the same mechanism I'd be glad to use
> > that instead.
> >
> > 3. There are no performance measurements *yet*.
> > On s390 the ideal conditions for this mechanism to be beneficial
> > usually surface when the whole system is fully loaded and resource
> > sharing between the logical partitions starts to get expensive.
> > Creating such an environment requires time, therefore no benchmark
> > results are available yet.
>
> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
> mask's current behavior, soft preference would allow tasks to spill onto other
> CPUs once the preferred CPUs have no idle threads available (we have something
> like this in sched_ext's scx_cosmos).
Soft preference is okay if contention is low.
But if steal% remains high, we may still need to transition to hard
boundaries/preferences on the cores to run on to bring steal% down.
Tim
>
> That might let us preserve the general preference for fully idle SMT cores while
> searching in this order: fully idle preferred cores, idle SMT threads on
> preferred cores, then non-preferred cores. This current patch series seems to
> drop the first distinction, allowing a partially idle preferred core to win even
> when another preferred core is fully idle.
>
> This would need to coexist with the steal governor's stronger use of
> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> rather than a drop-in replacement.
>
> What do you think (Mete / Shrikanth)?
>
> Thanks,
> -Andrea
next prev parent reply other threads:[~2026-10-08 21:52 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 8:30 Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority Mete Durlu
2026-10-08 8:30 ` [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO Mete Durlu
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
2026-10-08 14:58 ` Shrikanth Hegde
2026-10-08 15:24 ` Mete Durlu
2026-10-08 15:44 ` Shrikanth Hegde
2026-10-08 21:52 ` Tim Chen [this message]
2026-10-08 21:29 ` Tim Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=30e586ab643a8911f6c3cd930796292a7da460ef.camel@linux.intel.com \
--to=tim.c.chen@linux.intel.com \
--cc=agordeev@linux.ibm.com \
--cc=arighi@nvidia.com \
--cc=borntraeger@linux.ibm.com \
--cc=bsegall@google.com \
--cc=dietmar.eggemann@arm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=iii@linux.ibm.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=meted@linux.ibm.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=svens@linux.ibm.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®