mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system
@ 2026-09-28 18:37 Tim Chen
  2026-09-28 19:52 ` Kayra Cizmeci
  0 siblings, 1 reply; 5+ messages in thread
From: Tim Chen @ 2026-09-28 18:37 UTC (permalink / raw)
  To: Peter Zijlstra, Ingo Molnar
  Cc: Tim Chen, Chen Yu, Mario Limonciello, Badole, Vishal,
	linux-kernel, maintainer : X86 ARCHITECTURE, platform-driver-x86,
	K Prateek Nayak, Ricardo Neri, stable, Klaus Kusche

A regression was reported on an AMD Ryzen AI HX 370 running a cache
intensive Clang full-LTO link. The little cores run at a much lower
frequency (3.3 GHz vs 5.1 GHz) and have only half of the L3 cache
(8 MB vs 16 MB), so pinning such a task to the little-core LLC hurts
twice, and full-LTO builds slow down dramatically compared to
pre-cache-aware-scheduling kernels.

Asym packing and cache aware scheduling express conflicting placement
strategy. Asym packing wants a task to run on the highest priority CPU,
whereas cache aware scheduling wants to co-locate the tasks of a process
on one LLC regardless of the priority of CPUs in that LLC.

When asym packing tries to migrate task to an empty
core that has higher priority than source cpu, let asym packing win.
Moving tasks to a higher performing idle core will buy more
performance than cache co-location.

Fixes: 23b2b5ccc45c ("sched/cache: Introduce helper functions to enforce LLC migration policy")
Reported-by: Klaus Kusche <klaus.kusche@computerix.info>
Closes: https://lore.kernel.org/lkml/2180ea5a-eb28-4152-8d4d-cd00b0c24b2e@computerix.info/
Tested-by: Klaus Kusche <klaus.kusche@computerix.info>
Tested-by: Ricardo Neri <ricardo.neri@intel.com>
Cc: stable@vger.kernel.org # 7.2.x
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
 kernel/sched/fair.c | 21 ++++++++++++++++++---
 1 file changed, 18 insertions(+), 3 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f265de8721dbd..31eb45efd063b 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10823,6 +10823,8 @@ static inline bool task_misfits_asym_cpu(struct lb_env *env, struct task_struct
 	return false;
 }
 
+static inline bool sched_asym(struct sched_domain *sd, int dst_cpu, int src_cpu);
+
 /*
  * Check if task p can migrate from source LLC to
  * destination LLC in terms of cache aware load balance.
@@ -10847,6 +10849,10 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
 	if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
 		return mig_unrestricted;
 
+	/* Prioritize asym packing over cache awareness */
+	if (sched_asym(env->sd, dst_cpu, src_cpu))
+		return mig_unrestricted;
+
 	/* skip cache aware load balance for too many threads */
 	if (invalid_llc_nr(grp, p, dst_cpu) ||
 	    exceed_llc_capacity(grp, dst_cpu)) {
@@ -12043,6 +12049,15 @@ static inline bool llc_balance(struct lb_env *env, struct sg_lb_stats *sgs,
 	    sgs->group_misfit_task_load)
 		return false;
 
+	/*
+	 * On asym packing domains, if the destination CPU
+	 * has higher priority than all CPUs in the source group,
+	 * prioritize asym packing.
+	 */
+	if ((env->sd->flags & SD_ASYM_PACKING) &&
+	    sgs->group_asym_packing)
+		return false;
+
 	/*
 	 * Skip cache aware tagging if nr_balanced_failed is sufficiently high.
 	 * Threshold of cache_nice_tries is set to 1 higher than nr_balance_failed
@@ -13458,12 +13473,12 @@ static int need_active_balance(struct lb_env *env)
 {
 	struct sched_domain *sd = env->sd;
 
-	if (alb_break_llc(env))
-		return 0;
-
 	if (asym_active_balance(env))
 		return 1;
 
+	if (alb_break_llc(env))
+		return 0;
+
 	if (imbalanced_active_balance(env))
 		return 1;
 
-- 
2.32.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system
  2026-09-28 18:37 [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system Tim Chen
@ 2026-09-28 19:52 ` Kayra Cizmeci
  2026-09-29 17:51   ` Tim Chen
  0 siblings, 1 reply; 5+ messages in thread
From: Kayra Cizmeci @ 2026-09-28 19:52 UTC (permalink / raw)
  To: tim.c.chen, Nathan Chancellor, Nick Desaulniers, Bill Wendling,
	Justin Stitt
  Cc: KPrateek.Nayak, Vishal.Badole, klaus.kusche, linux-kernel,
	mario.limonciello, mingo, peterz, platform-driver-x86,
	ricardo.neri, stable, x86, yu.c.chen, llvm

Helloooo Tim,

> A regression was reported on an AMD Ryzen AI HX 370 running a cache
> intensive Clang full-LTO link. The little cores run at a much lower
> frequency (3.3 GHz vs 5.1 GHz) and have only half of the L3 cache
> (8 MB vs 16 MB), so pinning such a task to the little-core LLC hurts
> twice, and full-LTO builds slow down dramatically compared to
> pre-cache-aware-scheduling kernels.

> Asym packing and cache aware scheduling express conflicting placement
> strategy. Asym packing wants a task to run on the highest priority CPU,
> whereas cache aware scheduling wants to co-locate the tasks of a process
> on one LLC regardless of the priority of CPUs in that LLC.

> When asym packing tries to migrate task to an empty
> core that has higher priority than source cpu, let asym packing win.
> Moving tasks to a higher performing idle core will buy more
> performance than cache co-location.


> +static inline bool sched_asym(struct sched_domain *sd, int dst_cpu, int src_cpu);
> +
>  /*
>   * Check if task p can migrate from source LLC to
>   * destination LLC in terms of cache aware load balance.
> @@ -10847,6 +10849,10 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
>  	if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
>  		return mig_unrestricted;
>  
> +	/* Prioritize asym packing over cache awareness */
> +	if (sched_asym(env->sd, dst_cpu, src_cpu))
> +		return mig_unrestricted;
> +
>  	/* skip cache aware load balance for too many threads */
>  	if (invalid_llc_nr(grp, p, dst_cpu) ||
>  	    exceed_llc_capacity(grp, dst_cpu)) {
> @@ -12043,6 +12049,15 @@ static inline bool llc_balance(struct lb_env *env, struct sg_lb_stats *sgs,
>  	    sgs->group_misfit_task_load)
>  		return false;
>  
> +	/*
> +	 * On asym packing domains, if the destination CPU
> +	 * has higher priority than all CPUs in the source group,
> +	 * prioritize asym packing.
> +	 */
> +	if ((env->sd->flags & SD_ASYM_PACKING) &&
> +	    sgs->group_asym_packing)
> +		return false;
> +
>  	/*
>  	 * Skip cache aware tagging if nr_balanced_failed is sufficiently high.
>  	 * Threshold of cache_nice_tries is set to 1 higher than nr_balance_failed
> @@ -13458,12 +13473,12 @@ static int need_active_balance(struct lb_env *env)
>  {
>  	struct sched_domain *sd = env->sd;
>  
> -	if (alb_break_llc(env))
> -		return 0;
> -
>  	if (asym_active_balance(env))
>  		return 1;
>  
> +	if (alb_break_llc(env))
> +		return 0;
> +
>  	if (imbalanced_active_balance(env))
>  		return 1;
 
Hope I could test this. But I don't really have hardware for it :-(.

Regardless tho.

Let's say when entering can_migrate_llc_task(), dst_cpu is CPU0, while src_cpu is CPU1.
And CPU0 has a bigger asym_prio than CPU1. No SMT. When entering can_migrate_llc_task() and sched_asym()
from there sched_use_asym_prio() returns true without checking if the core is fully idle or not.
And sched_asym_prefer() comes back true too, so sched_asym() returns true and, we just returned mig_unrestricted.

I could be missing something, If I'm not tho is that on purpose? If it is, the last paragraph needs to change
since it says that "asym packing tries to migrate task to an empty core" and after that "higher performing idle core"

Thanks,
Kayra


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system
  2026-09-28 19:52 ` Kayra Cizmeci
@ 2026-09-29 17:51   ` Tim Chen
  2026-09-29 18:21     ` Kayra Cizmeci
  0 siblings, 1 reply; 5+ messages in thread
From: Tim Chen @ 2026-09-29 17:51 UTC (permalink / raw)
  To: Kayra Cizmeci, Nathan Chancellor, Nick Desaulniers,
	Bill Wendling, Justin Stitt
  Cc: KPrateek.Nayak, Vishal.Badole, klaus.kusche, linux-kernel,
	mario.limonciello, mingo, peterz, platform-driver-x86,
	ricardo.neri, stable, x86, yu.c.chen, llvm

On Mon, 2026-09-28 at 22:52 +0300, Kayra Cizmeci wrote:
> Helloooo Tim,
> 
> > A regression was reported on an AMD Ryzen AI HX 370 running a cache
> > intensive Clang full-LTO link. The little cores run at a much lower
> > frequency (3.3 GHz vs 5.1 GHz) and have only half of the L3 cache
> > (8 MB vs 16 MB), so pinning such a task to the little-core LLC hurts
> > twice, and full-LTO builds slow down dramatically compared to
> > pre-cache-aware-scheduling kernels.
> 
> > Asym packing and cache aware scheduling express conflicting placement
> > strategy. Asym packing wants a task to run on the highest priority CPU,
> > whereas cache aware scheduling wants to co-locate the tasks of a process
> > on one LLC regardless of the priority of CPUs in that LLC.
> 
> > When asym packing tries to migrate task to an empty
> > core that has higher priority than source cpu, let asym packing win.
> > Moving tasks to a higher performing idle core will buy more
> > performance than cache co-location.
> 
> 
> > +static inline bool sched_asym(struct sched_domain *sd, int dst_cpu, int src_cpu);
> > +
> >  /*
> >   * Check if task p can migrate from source LLC to
> >   * destination LLC in terms of cache aware load balance.
> > @@ -10847,6 +10849,10 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
> >  	if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
> >  		return mig_unrestricted;
> >  
> > +	/* Prioritize asym packing over cache awareness */
> > +	if (sched_asym(env->sd, dst_cpu, src_cpu))
> > +		return mig_unrestricted;
> > +
> >  	/* skip cache aware load balance for too many threads */
> >  	if (invalid_llc_nr(grp, p, dst_cpu) ||
> >  	    exceed_llc_capacity(grp, dst_cpu)) {
> > @@ -12043,6 +12049,15 @@ static inline bool llc_balance(struct lb_env *env, struct sg_lb_stats *sgs,
> >  	    sgs->group_misfit_task_load)
> >  		return false;
> >  
> > +	/*
> > +	 * On asym packing domains, if the destination CPU
> > +	 * has higher priority than all CPUs in the source group,
> > +	 * prioritize asym packing.
> > +	 */
> > +	if ((env->sd->flags & SD_ASYM_PACKING) &&
> > +	    sgs->group_asym_packing)
> > +		return false;
> > +
> >  	/*
> >  	 * Skip cache aware tagging if nr_balanced_failed is sufficiently high.
> >  	 * Threshold of cache_nice_tries is set to 1 higher than nr_balance_failed
> > @@ -13458,12 +13473,12 @@ static int need_active_balance(struct lb_env *env)
> >  {
> >  	struct sched_domain *sd = env->sd;
> >  
> > -	if (alb_break_llc(env))
> > -		return 0;
> > -
> >  	if (asym_active_balance(env))
> >  		return 1;
> >  
> > +	if (alb_break_llc(env))
> > +		return 0;
> > +
> >  	if (imbalanced_active_balance(env))
> >  		return 1;
>  
> Hope I could test this. But I don't really have hardware for it :-(.
> 
> Regardless tho.
> 
> Let's say when entering can_migrate_llc_task(), dst_cpu is CPU0, while src_cpu is CPU1.
> And CPU0 has a bigger asym_prio than CPU1. No SMT. When entering can_migrate_llc_task() and sched_asym()
> from there sched_use_asym_prio() returns true without checking if the core is fully idle or not.

sched_asym() does check whether the destination core is idle in sched_use_asym_prio() for non SMT domain.

                  (false for non-SMT sd)           (check idle core)
       return sd->flags & SD_SHARE_CPUCAPACITY || is_core_idle(cpu);

That is also a pre-condition for setting group_asym_packing.

> And sched_asym_prefer() comes back true too, so sched_asym() returns true and, we just returned mig_unrestricted.
> 
> I could be missing something, If I'm not tho is that on purpose? If it is, the last paragraph needs to change
> since it says that "asym packing tries to migrate task to an empty core" and after that "higher performing idle core"

I think I did try to point out the idle core aspect in my commit log:

"When asym packing tries to migrate task to an empty
core that has higher priority than source cpu, let asym packing win.
Moving tasks to a higher performing idle core will buy more
performance than cache co-location."

I am not sure how you would like it changed.

Thanks.

Tim

> 
> Thanks,
> Kayra

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system
  2026-09-29 17:51   ` Tim Chen
@ 2026-09-29 18:21     ` Kayra Cizmeci
  2026-09-29 20:08       ` Tim Chen
  0 siblings, 1 reply; 5+ messages in thread
From: Kayra Cizmeci @ 2026-09-29 18:21 UTC (permalink / raw)
  To: tim.c.chen
  Cc: KPrateek.Nayak, Vishal.Badole, justinstitt, kayracizmeci,
	klaus.kusche, linux-kernel, llvm, mario.limonciello, mingo,
	morbo, nathan, ndesaulniers, peterz, platform-driver-x86,
	ricardo.neri, stable, x86, yu.c.chen

Hi Tim :>,

> > Let's say when entering can_migrate_llc_task(), dst_cpu is CPU0, while src_cpu is CPU1.
> > And CPU0 has a bigger asym_prio than CPU1. No SMT. When entering can_migrate_llc_task() and sched_asym()
> > from there sched_use_asym_prio() returns true without checking if the core is fully idle or not.

> sched_asym() does check whether the destination core is idle in sched_use_asym_prio() for non SMT domain.
> 
>                   (false for non-SMT sd)           (check idle core)
>        return sd->flags & SD_SHARE_CPUCAPACITY || is_core_idle(cpu);
> 
> That is also a pre-condition for setting group_asym_packing.

Sorry for not showing the code earlier, here it is: 

static inline bool is_core_idle(int cpu)
{
	int sibling;

	for_each_cpu(sibling, cpu_smt_mask(cpu)) {
		if (cpu == sibling)
			continue;

		if (!idle_cpu(sibling))
			return false;
	}

	return true;
}
static bool sched_use_asym_prio(struct sched_domain *sd, int cpu)
{
	if (!(sd->flags & SD_ASYM_PACKING))
		return false;

	if (!sched_smt_active())
		return true;

	return sd->flags & SD_SHARE_CPUCAPACITY || is_core_idle(cpu);
}

On sched_use_asym_prio(), before the idle check a CPU without SMT returns true.
But I'm actually wrong on that one because on that one we're laying
on the idle protection outside to the CPU. If there are not any
other brothers, and I'm idle then my brother-family
is idle... Ah I messed up describing this.

But is there are any idle protection outside? I couldn't
find any. I could be missing something tho.

> > And sched_asym_prefer() comes back true too, so sched_asym() returns true and, we just returned mig_unrestricted.
> >
> > I could be missing something, If I'm not tho is that on purpose? If it is, the last paragraph needs to change
> > since it says that "asym packing tries to migrate task to an empty core" and after that "higher performing idle core"

> I think I did try to point out the idle core aspect in my commit log:

> "When asym packing tries to migrate task to an empty
> core that has higher priority than source cpu, let asym packing win.
> Moving tasks to a higher performing idle core will buy more
> performance than cache co-location."

I know. I was trying to say that if we're not choosing an idle
core this needs to change.

Thanks,
Kayra


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system
  2026-09-29 18:21     ` Kayra Cizmeci
@ 2026-09-29 20:08       ` Tim Chen
  0 siblings, 0 replies; 5+ messages in thread
From: Tim Chen @ 2026-09-29 20:08 UTC (permalink / raw)
  To: Kayra Cizmeci
  Cc: KPrateek.Nayak, Vishal.Badole, justinstitt, klaus.kusche,
	linux-kernel, llvm, mario.limonciello, mingo, morbo, nathan,
	ndesaulniers, peterz, platform-driver-x86, ricardo.neri, stable,
	x86, yu.c.chen

On Tue, 2026-09-29 at 21:21 +0300, Kayra Cizmeci wrote:
> Hi Tim :>,
> 
> > > Let's say when entering can_migrate_llc_task(), dst_cpu is CPU0, while src_cpu is CPU1.
> > > And CPU0 has a bigger asym_prio than CPU1. No SMT. When entering can_migrate_llc_task() and sched_asym()
> > > from there sched_use_asym_prio() returns true without checking if the core is fully idle or not.
> 
> > sched_asym() does check whether the destination core is idle in sched_use_asym_prio() for non SMT domain.
> > 
> >                   (false for non-SMT sd)           (check idle core)
> >        return sd->flags & SD_SHARE_CPUCAPACITY || is_core_idle(cpu);
> > 
> > That is also a pre-condition for setting group_asym_packing.
> 
> Sorry for not showing the code earlier, here it is: 
> 
> static inline bool is_core_idle(int cpu)
> {
> 	int sibling;
> 
> 	for_each_cpu(sibling, cpu_smt_mask(cpu)) {
> 		if (cpu == sibling)
> 			continue;
> 
> 		if (!idle_cpu(sibling))
> 			return false;
> 	}
> 
> 	return true;
> }
> static bool sched_use_asym_prio(struct sched_domain *sd, int cpu)
> {
> 	if (!(sd->flags & SD_ASYM_PACKING))
> 		return false;
> 
> 	if (!sched_smt_active())
> 		return true;
> 
> 	return sd->flags & SD_SHARE_CPUCAPACITY || is_core_idle(cpu);
> }
> 
> On sched_use_asym_prio(), before the idle check a CPU without SMT returns true.
> But I'm actually wrong on that one because on that one we're laying
> on the idle protection outside to the CPU. If there are not any
> other brothers, and I'm idle then my brother-family
> is idle... Ah I messed up describing this.
> 
> But is there are any idle protection outside? I couldn't
> find any. I could be missing something tho.

Actually idle cpu is being checked as a pre-condition for setting
asym_packing.

               /* Check if dst CPU is idle and preferred to this group */
                if (env->idle && sgs->sum_h_nr_running &&
                    sched_group_asym(env, sgs, group))
                        sgs->group_asym_packing = 1;

Are you saying that we should update the first chunk to add an env->idle check
to match the asym_packing migration pre-requisite? Like below?

@@ -10847,6 +10849,10 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
 	if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
 		return mig_unrestricted;
 
+	/* Prioritize asym packing over cache awareness */
+	if (env->idle && sched_asym(env->sd, dst_cpu, src_cpu))
+		return mig_unrestricted;
+  

I think this is a valid point.

Tim

> 
> > > And sched_asym_prefer() comes back true too, so sched_asym() returns true and, we just returned mig_unrestricted.
> > > 
> > > I could be missing something, If I'm not tho is that on purpose? If it is, the last paragraph needs to change
> > > since it says that "asym packing tries to migrate task to an empty core" and after that "higher performing idle core"
> 
> > I think I did try to point out the idle core aspect in my commit log:
> 
> > "When asym packing tries to migrate task to an empty
> > core that has higher priority than source cpu, let asym packing win.
> > Moving tasks to a higher performing idle core will buy more
> > performance than cache co-location."
> 
> I know. I was trying to say that if we're not choosing an idle
> core this needs to change.
> 
> Thanks,
> Kayra
> 

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-29 20:08 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 18:37 [PATCH] sched/cache: Honor asym packing over cache aware scheduling on hybrid system Tim Chen
2026-09-28 19:52 ` Kayra Cizmeci
2026-09-29 17:51   ` Tim Chen
2026-09-29 18:21     ` Kayra Cizmeci
2026-09-29 20:08       ` Tim Chen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®