mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Tim Chen <tim.c.chen@linux.intel.com>
To: Andrea Righi <arighi@nvidia.com>, Mete Durlu <meted@linux.ibm.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	 Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann	 <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall	 <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider	 <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Shrikanth Hegde <sshegde@linux.ibm.com>,
	Heiko Carstens <hca@linux.ibm.com>,
	Vasily Gorbik	 <gor@linux.ibm.com>,
	Alexander Gordeev <agordeev@linux.ibm.com>,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	Sven Schnelle <svens@linux.ibm.com>,
	Chen Yu <yu.c.chen@intel.com>,
	 Ilya Leoshkevich	 <iii@linux.ibm.com>,
	linux-kernel@vger.kernel.org, linux-s390@vger.kernel.org
Subject: Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems
Date: Thu, 08 Oct 2026 14:52:14 -0700	[thread overview]
Message-ID: <30e586ab643a8911f6c3cd930796292a7da460ef.camel@linux.intel.com> (raw)
In-Reply-To: <aseIPhrDGCYFtJNE@gpd4>

On Thu, 2026-10-08 at 14:10 +0200, Andrea Righi wrote:
> Hi Mete,
> 
> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> > Summary
> > ===========================================================================
> > On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> > cores over idle SMT siblings of busy cores. This works generally well but
> > virtualized platforms where low capacity cores should be avoided are not
> > considered. Introduce a new config option and arch hook to prioritize
> > SMT utilization.
> > 
> > Background
> > ===========================================================================
> > The scheduler has been moving toward better utilization of fully idle cores
> > and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> > ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> > broke the behavior s390 is relying on to concentrate workloads on its
> > high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> > entitlement assigned by the hypervisor. The idle SMT siblings of busy
> > high-capacity cores start to perform better than fully idle low-capacity
> > cores as the whole machine(containing the logical partitions) starts
> > approaching to a {fully,over}loaded state. Grouping load on high-capacity
> > cores keeps shared low-capacity cores(which are shared more aggressively)
> > idle longer, reducing noise to neighboring partitions and improving
> > overall performance.
> > 
> > Approach
> > ===========================================================================
> > This series introduces SCHED_IDLE_SMT_PRIO config option and the
> > sched_idle_smt_prio static branch, allowing architectures to treat idle
> > SMT siblings of busy cores as equal candidates during asymmetric capacity
> > load balancing.
> > Static branch checks are placed at paths considering fully idle cores
> > over idle SMT threads in presence of asymmetric CPU capacities within
> > scheduling groups. Inserted checks mostly override hints for idle core
> > selection or cause early exits favoring idle SMT siblings.
> > One more static branch check is added to new task path in order to favor
> > high capacity SMT siblings during initial task placement, therefore the
> > tasks are immediately placed in high capacity SMT siblings instead of
> > considering most idle but low capacity cores.
> > 
> > Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> > implementing arch_needs_idle_smt_prio(), which is evaluated during each
> > asym_cpu_capacity_scan() to track the state as runtime topology changes.
> > 
> > The second patch enables the feature for s390 when running on
> > hardware-backed topology in an LPAR with vertical polarization active
> > and system is approaching to a target state.
> > Otherwise the branch stays disabled and the scheduler falls back to the
> > standard idle-core preference.
> > 
> > No functional change on architectures that do not select
> > ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
> > 
> > Performance Results
> > ===========================================================================
> > Since this change effects the {fully,over}loaded state it is difficult
> > to test at the moment. Therefore there are no concrete numbers or
> > metrics for now for s390.
> > For architectures which do not opt in to this feature, no performance
> > change is expected.
> > 
> > Considerations and Open Questions
> > ===========================================================================
> > The main goal for this series is adding a simple mechanism for
> > architectures to switch between idle core and idle SMT priority while
> > keeping the introduced footprint as small as possible. But there are
> > some ideas and questions to consider;
> > 
> > 1. Should arch_needs_idle_smt_prio() hook be removed?
> >    The architecture hook is there to allow for any sort of logic to
> >    dynamically decide when to flip the mechanism, but it can also be
> >    removed if everyone agrees that this behaviour is not something that
> >    should be dynamically flipped. It can be simply tied to detection of
> >    asymmetric capacities and selection of Kconfig option.
> > 
> > 2. Is there a simpler way to implement this mechanism?
> >    The proposed approach is chosen as the other features effecting the
> >    scheduler's behavior, use the same method. If there is a more
> >    efficient way to implement the same mechanism I'd be glad to use
> >    that instead.
> > 
> > 3. There are no performance measurements *yet*.
> >    On s390 the ideal conditions for this mechanism to be beneficial
> >    usually surface when the whole system is fully loaded and resource
> >    sharing between the logical partitions starts to get expensive.
> >    Creating such an environment requires time, therefore no benchmark
> >    results are available yet.
> 
> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
> mask's current behavior, soft preference would allow tasks to spill onto other
> CPUs once the preferred CPUs have no idle threads available (we have something
> like this in sched_ext's scx_cosmos).

Soft preference is okay if contention is low.
But if steal% remains high, we may still need to transition to hard
boundaries/preferences on the cores to run on to bring steal% down.

Tim

> 
> That might let us preserve the general preference for fully idle SMT cores while
> searching in this order: fully idle preferred cores, idle SMT threads on
> preferred cores, then non-preferred cores. This current patch series seems to
> drop the first distinction, allowing a partially idle preferred core to win even
> when another preferred core is fully idle.
> 
> This would need to coexist with the steal governor's stronger use of
> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> rather than a drop-in replacement.
> 
> What do you think (Mete / Shrikanth)?
> 
> Thanks,
> -Andrea

  parent reply	other threads:[~2026-10-08 21:52 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08  8:30 Mete Durlu
2026-10-08  8:30 ` [PATCH RFC 1/2] kernel/sched: Introduce idle SMT priority Mete Durlu
2026-10-08  8:30 ` [PATCH RFC 2/2] s390/topology: Enable SCHED_IDLE_SMT_PRIO Mete Durlu
2026-10-08 12:10 ` [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems Andrea Righi
2026-10-08 14:58   ` Shrikanth Hegde
2026-10-08 15:24     ` Mete Durlu
2026-10-08 15:44       ` Shrikanth Hegde
2026-10-08 21:52   ` Tim Chen [this message]
2026-10-08 21:29 ` Tim Chen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=30e586ab643a8911f6c3cd930796292a7da460ef.camel@linux.intel.com \
    --to=tim.c.chen@linux.intel.com \
    --cc=agordeev@linux.ibm.com \
    --cc=arighi@nvidia.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=bsegall@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=iii@linux.ibm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=meted@linux.ibm.com \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sshegde@linux.ibm.com \
    --cc=svens@linux.ibm.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®