* [PATCH v2 0/3] sched/fair: Update decision matrix and rename pinned-task group state
@ 2026-10-02 15:46 Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
` (2 more replies)
0 siblings, 3 replies; 5+ messages in thread
From: Jemmy Wong @ 2026-10-02 15:46 UTC (permalink / raw)
To: Tim Chen, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
Patch 1 updates the group-type decision matrix above
sched_balance_find_src_group(), which was never extended for
group_smt_balance (fee1759e4f04) and group_llc_balance (f38cc2f0d8a3).
Patches 2-3 rename the state that records balancing failures caused by
cpus_ptr-pinned tasks. Since commit 6263322c5e8f ("sched/fair: Rewrite
group_imb trigger") it is an affinity flag, not a load-skew measure,
yet it kept the names group_imbalanced and sgc->imbalance, which
suggest a load skew and are easily confused with env->imbalance. They
become group_pinned_imb and sgc->pinned_imb.
Patch 1 does not depend on patches 2-3. No functional change.
Changes in v2:
- Patch 1: llc_balance vs. local has_spare is "force", not "nr_idle"
(Tim Chen); reword the changelog accordingly.
- Patch 2: Rename to group_pinned_imb instead of group_pinned_task
(Tim Chen).
- Patch 3: Rename sgc->pinned_task to sgc->pinned_imb and the reader to
sg_pinned_imb(), to match.
v1: https://lore.kernel.org/all/20260928042018.10618-1-jemmywong512@gmail.com/
Jemmy Wong (3):
sched/fair: Add smt_balance and llc_balance to the decision matrix
sched/fair: Rename group_imbalanced to group_pinned_imb
sched/fair: Rename sgc->imbalance to sgc->pinned_imb
kernel/sched/fair.c | 66 +++++++++++++++++++++++---------------------
kernel/sched/sched.h | 6 +++-
2 files changed, 39 insertions(+), 33 deletions(-)
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
2026-10-02 15:46 [PATCH v2 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
@ 2026-10-02 15:46 ` Jemmy Wong
2026-10-02 21:36 ` Tim Chen
2026-10-02 15:46 ` [PATCH v2 2/3] sched/fair: Rename group_imbalanced to group_pinned_imb Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_imb Jemmy Wong
2 siblings, 1 reply; 5+ messages in thread
From: Jemmy Wong @ 2026-10-02 15:46 UTC (permalink / raw)
To: Tim Chen, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
The group-type matrix was introduced in commit 0b0695f2b34a
("sched/fair: Rework load_balance()"). When commit fee1759e4f04
("sched/fair: Determine active load balance for SMT sched groups")
added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
Prioritize tasks preferring destination LLC during balancing")
added group_llc_balance, neither commit updated the matrix table.
Both types are only tagged on non-local groups in update_sg_lb_stats(),
so their local columns are N/A.
As busiest, group_smt_balance is only set when dst_cpu is idle and the
SMT group runs more than one task. Against a local has_spare or
fully_busy group it goes through the nr_idle checks, where a non-SMT
dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
local imbalanced or overloaded group the local group is busier and the
pair is balanced.
As busiest, group_llc_balance is not an unconditional force. Against a
local has_spare group it forces the pull when prefer_sibling is set,
because the group_llc_balance test comes before sibling_imbalance() in
sched_balance_find_src_group(). SD_PREFER_SIBLING is only cleared for
NUMA domains. Against a local fully_busy or imbalanced group the nr_idle
checks apply, and against a local overloaded group the local group is
busier and the pair is balanced.
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 16 +++++++++-------
1 file changed, 9 insertions(+), 7 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5cdde4..bcb9987952b9 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -13184,13 +13184,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
/*
* Decision matrix according to the local and busiest group type:
*
- * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
- * has_spare nr_idle balanced N/A N/A balanced balanced
- * fully_busy nr_idle nr_idle N/A N/A balanced balanced
- * misfit_task force N/A N/A N/A N/A N/A
- * asym_packing force force N/A N/A force force
- * imbalanced force force N/A N/A force force
- * overloaded force force N/A N/A force avg_load
+ * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
+ * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
+ * misfit_task force N/A N/A N/A N/A N/A N/A N/A
+ * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
+ * asym_packing force force N/A N/A N/A force N/A force
+ * imbalanced force force N/A N/A N/A force N/A force
+ * llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
+ * overloaded force force N/A N/A N/A force N/A avg_load
*
* N/A : Not Applicable because already filtered while updating
* statistics.
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v2 2/3] sched/fair: Rename group_imbalanced to group_pinned_imb
2026-10-02 15:46 [PATCH v2 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
@ 2026-10-02 15:46 ` Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_imb Jemmy Wong
2 siblings, 0 replies; 5+ messages in thread
From: Jemmy Wong @ 2026-10-02 15:46 UTC (permalink / raw)
To: Tim Chen, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
The name dates back to the original group_imb heuristic, which flagged
a group when the load difference between its busiest and idlest CPU
exceeded the average task load. Commit 6263322c5e8f ("sched/fair:
Rewrite group_imb trigger") replaced that heuristic: the flag is now
set only when a lower domain fails to balance because tasks are pinned
by cpus_ptr (LBF_SOME_PINNED), and kept while all tasks are pinned
(LBF_ALL_PINNED). The name was carried over unchanged and later became
group_imbalanced in commit 0b0695f2b34a ("sched/fair: Rework
load_balance()").
Today the name no longer matches the condition:
- group_classify() checks group_overloaded first, so a group whose
load really is skewed is usually not classified as imbalanced,
while a flagged group may carry a single extra task.
- calculate_imbalance() does not measure any imbalance for this type;
it moves one task (migrate_task, imbalance = 1).
- Elsewhere in fair.c "imbalance" consistently means the amount of
load to move (env->imbalance, imbalance_pct, calculate_imbalance(),
the lb_imbalance_* schedstats), and imbalanced_active_balance() uses
"imbalanced" for repeated balance failures, unrelated to this flag.
Rename it to group_pinned_imb. The "pinned" prefix names the cause that
raises it, an imbalance the balancer could not fix because tasks are
pinned by cpus_ptr, and distinguishes it from the unqualified
"imbalanced" uses above. The enum is local to fair.c, so no tracepoint,
schedstat or other user-visible interface is affected.
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 30 +++++++++++++++---------------
1 file changed, 15 insertions(+), 15 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index bcb9987952b9..49de871ee7ab 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10579,7 +10579,7 @@ enum group_type {
* The tasks' affinity constraints previously prevented the scheduler
* from balancing the load across the system.
*/
- group_imbalanced,
+ group_pinned_imb,
/*
* There are tasks running on non-preferred LLC, possible to move
* them to their preferred LLC without creating too much imbalance.
@@ -11938,7 +11938,7 @@ group_type group_classify(unsigned int imbalance_pct,
return group_llc_balance;
if (sg_imbalanced(group))
- return group_imbalanced;
+ return group_pinned_imb;
if (sgs->group_asym_packing)
return group_asym_packing;
@@ -12401,10 +12401,10 @@ static bool update_sd_pick_busiest(struct lb_env *env,
/* Select the group with most tasks preferring dst LLC */
return update_llc_busiest(env, busiest, sgs);
- case group_imbalanced:
+ case group_pinned_imb:
/*
- * Select the 1st imbalanced group as we don't have any way to
- * choose one more than another.
+ * Select the 1st group with pinned tasks as we don't
+ * have any way to choose one more than another.
*/
return false;
@@ -12653,7 +12653,7 @@ static bool update_pick_idlest(struct sched_group *idlest,
break;
case group_llc_balance:
- case group_imbalanced:
+ case group_pinned_imb:
case group_asym_packing:
case group_smt_balance:
/* Those types are not used in the slow wakeup path */
@@ -12786,7 +12786,7 @@ sched_balance_find_dst_group(struct sched_domain *sd, struct task_struct *p, int
break;
case group_llc_balance:
- case group_imbalanced:
+ case group_pinned_imb:
case group_asym_packing:
case group_smt_balance:
/* Those type are not used in the slow wakeup path */
@@ -13049,11 +13049,11 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
}
#endif
- if (busiest->group_type == group_imbalanced) {
+ if (busiest->group_type == group_pinned_imb) {
/*
- * In the group_imb case we cannot rely on group-wide averages
- * to ensure CPU-load equilibrium, try to move any task to fix
- * the imbalance. The next load balance will take care of
+ * In the group_pinned_imb case we cannot rely on group-wide
+ * averages to ensure CPU-load equilibrium, try to move any task
+ * to fix the imbalance. The next load balance will take care of
* balancing back the system.
*/
env->migration_type = migrate_task;
@@ -13184,13 +13184,13 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
/*
* Decision matrix according to the local and busiest group type:
*
- * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * busiest \ local has_spare fully_busy misfit smt asym pinned llc overloaded
* has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
* fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
* misfit_task force N/A N/A N/A N/A N/A N/A N/A
* smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
* asym_packing force force N/A N/A N/A force N/A force
- * imbalanced force force N/A N/A N/A force N/A force
+ * pinned_imb force force N/A N/A N/A force N/A force
* llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
* overloaded force force N/A N/A N/A force N/A avg_load
*
@@ -13245,11 +13245,11 @@ static struct sched_group *sched_balance_find_src_group(struct lb_env *env)
goto force_balance;
/*
- * If the busiest group is imbalanced the below checks don't
+ * If the busiest group has pinned tasks the below checks don't
* work because they assume all things are equal, which typically
* isn't true due to cpus_ptr constraints and the like.
*/
- if (busiest->group_type == group_imbalanced)
+ if (busiest->group_type == group_pinned_imb)
goto force_balance;
local = &sds.local_stat;
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v2 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_imb
2026-10-02 15:46 [PATCH v2 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 2/3] sched/fair: Rename group_imbalanced to group_pinned_imb Jemmy Wong
@ 2026-10-02 15:46 ` Jemmy Wong
2 siblings, 0 replies; 5+ messages in thread
From: Jemmy Wong @ 2026-10-02 15:46 UTC (permalink / raw)
To: Tim Chen, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
sgc->imbalance does not describe a load skew either. It is set only
when a child domain failed to balance because some tasks were pinned by
cpus_ptr, and the parent then classifies the group as group_pinned_imb
so that it can force a migration.
The unqualified name is easily confused with env->imbalance, the amount
of load still to move, and with imbalance_pct. And after the previous
rename, sg_imbalanced() returns group_pinned_imb, so the helper and the
group type no longer agree.
Rename the field to pinned_imb and its reader to sg_pinned_imb().
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 24 ++++++++++++------------
kernel/sched/sched.h | 6 +++++-
2 files changed, 17 insertions(+), 13 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 49de871ee7ab..0f773e1af9f5 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -11832,7 +11832,7 @@ static inline bool check_misfit_status(struct rq *rq)
}
/*
- * Group imbalance indicates (and tries to solve) the problem where balancing
+ * sgc->pinned_imb indicates (and tries to solve) the problem where balancing
* groups is inadequate due to ->cpus_ptr constraints.
*
* Imagine a situation of two groups of 4 CPUs each and 4 tasks each with a
@@ -11860,9 +11860,9 @@ static inline bool check_misfit_status(struct rq *rq)
* subtle and fragile situation.
*/
-static inline int sg_imbalanced(struct sched_group *group)
+static inline int sg_pinned_imb(struct sched_group *group)
{
- return group->sgc->imbalance;
+ return group->sgc->pinned_imb;
}
/*
@@ -11937,7 +11937,7 @@ group_type group_classify(unsigned int imbalance_pct,
if (sgs->group_llc_balance)
return group_llc_balance;
- if (sg_imbalanced(group))
+ if (sg_pinned_imb(group))
return group_pinned_imb;
if (sgs->group_asym_packing)
@@ -13870,10 +13870,10 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
* We failed to reach balance because of affinity.
*/
if (sd_parent) {
- int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+ int *pinned_imb = &sd_parent->groups->sgc->pinned_imb;
if ((env.flags & LBF_SOME_PINNED) && env.imbalance > 0)
- *group_imbalance = 1;
+ *pinned_imb = 1;
}
/* All tasks on this runqueue were pinned by CPU affinity */
@@ -13973,21 +13973,21 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
out_balanced:
/*
* We reach balance although we may have faced some affinity
- * constraints. Clear the imbalance flag only if other tasks got
+ * constraints. Clear the pinned flag only if other tasks got
* a chance to move and fix the imbalance.
*/
if (sd_parent && !(env.flags & LBF_ALL_PINNED)) {
- int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+ int *pinned_imb = &sd_parent->groups->sgc->pinned_imb;
- if (*group_imbalance)
- *group_imbalance = 0;
+ if (*pinned_imb)
+ *pinned_imb = 0;
}
out_all_pinned:
/*
* We reach balance because all tasks are pinned at this level so
- * we can't migrate them. Let the imbalance flag set so parent level
- * can try to migrate them.
+ * we can't migrate them. Let the pinned flag stay set so the parent
+ * level can try to migrate them.
*/
schedstat_inc(sd->lb_balanced[idle]);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..14ce0f6e6027 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2257,7 +2257,11 @@ struct sched_group_capacity {
unsigned long min_capacity; /* Min per-CPU capacity in group */
unsigned long max_capacity; /* Max per-CPU capacity in group */
unsigned long next_update;
- int imbalance; /* XXX unrelated to capacity but shared group state */
+ /*
+ * Set when a child domain cannot balance pinned tasks.
+ * XXX unrelated to capacity, but it is shared group state.
+ */
+ int pinned_imb;
int id;
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
2026-10-02 15:46 ` [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
@ 2026-10-02 21:36 ` Tim Chen
0 siblings, 0 replies; 5+ messages in thread
From: Tim Chen @ 2026-10-02 21:36 UTC (permalink / raw)
To: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
On Fri, 2026-10-02 at 23:46 +0800, Jemmy Wong wrote:
> The group-type matrix was introduced in commit 0b0695f2b34a
> ("sched/fair: Rework load_balance()"). When commit fee1759e4f04
> ("sched/fair: Determine active load balance for SMT sched groups")
> added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
> Prioritize tasks preferring destination LLC during balancing")
> added group_llc_balance, neither commit updated the matrix table.
>
> Both types are only tagged on non-local groups in update_sg_lb_stats(),
> so their local columns are N/A.
>
> As busiest, group_smt_balance is only set when dst_cpu is idle and the
> SMT group runs more than one task. Against a local has_spare or
> fully_busy group it goes through the nr_idle checks, where a non-SMT
> dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
> local imbalanced or overloaded group the local group is busier and the
> pair is balanced.
>
> As busiest, group_llc_balance is not an unconditional force. Against a
> local has_spare group it forces the pull when prefer_sibling is set,
> because the group_llc_balance test comes before sibling_imbalance() in
> sched_balance_find_src_group(). SD_PREFER_SIBLING is only cleared for
> NUMA domains. Against a local fully_busy or imbalanced group the nr_idle
> checks apply, and against a local overloaded group the local group is
> busier and the pair is balanced.
>
> No functional change.
>
> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
> ---
> kernel/sched/fair.c | 16 +++++++++-------
> 1 file changed, 9 insertions(+), 7 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 57360f5cdde4..bcb9987952b9 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -13184,13 +13184,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
> /*
> * Decision matrix according to the local and busiest group type:
> *
> - * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
> - * has_spare nr_idle balanced N/A N/A balanced balanced
> - * fully_busy nr_idle nr_idle N/A N/A balanced balanced
> - * misfit_task force N/A N/A N/A N/A N/A
> - * asym_packing force force N/A N/A force force
> - * imbalanced force force N/A N/A force force
> - * overloaded force force N/A N/A force avg_load
> + * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
> + * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
> + * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
> + * misfit_task force N/A N/A N/A N/A N/A N/A N/A
> + * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
> + * asym_packing force force N/A N/A N/A force N/A force
> + * imbalanced force force N/A N/A N/A force N/A force
> + * llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
> + * overloaded force force N/A N/A N/A force N/A avg_load
> *
> * N/A : Not Applicable because already filtered while updating
> * statistics.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-02 21:36 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 15:46 [PATCH v2 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
2026-10-02 21:36 ` Tim Chen
2026-10-02 15:46 ` [PATCH v2 2/3] sched/fair: Rename group_imbalanced to group_pinned_imb Jemmy Wong
2026-10-02 15:46 ` [PATCH v2 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_imb Jemmy Wong
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®