* [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state
@ 2026-09-28 4:20 Jemmy Wong
2026-09-28 4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Jemmy Wong @ 2026-09-28 4:20 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel
Patch 1 brings the group-type decision matrix above
sched_balance_find_src_group() up to date. The matrix was introduced
by Vincent Guittot in commit 0b0695f2b34a ("sched/fair: Rework
load_balance()"). When commit fee1759e4f04 ("sched/fair: Determine
active load balance for SMT sched groups") added group_smt_balance
and commit f38cc2f0d8a3 ("sched/cache: Prioritize tasks preferring
destination LLC during balancing") added group_llc_balance, neither
author updated the matrix table.
Both are only tagged on non-local groups, so their columns are N/A.
As busiest, both go through the nr_idle checks rather than an
unconditional force, and both are balanced against a busier local
group.
Patches 2-3 rename the state that tracks load balancing failures caused
by cpus_ptr-pinned tasks. The names come from the original group_imb
heuristic, which did detect per-CPU load skew; commit 6263322c5e8f
("sched/fair: Rewrite group_imb trigger") turned it into an affinity
failure flag but kept the name. Today group_imbalanced is not what
group_classify() returns for skewed load, calculate_imbalance()
computes no imbalance for it, and sgc->imbalance collides with
env->imbalance and imbalance_pct. They become group_pinned_task and
sgc->pinned_task, matching the LBF_*_PINNED flags that set them.
Patch 1 does not depend on patches 2-3 and can be applied on its own.
No functional change.
Jemmy Wong (3):
sched/fair: Add smt_balance and llc_balance to the decision matrix
sched/fair: Rename group_imbalanced to group_pinned_task
sched/fair: Rename sgc->imbalance to sgc->pinned_task
kernel/sched/fair.c | 66 +++++++++++++++++++++++---------------------
kernel/sched/sched.h | 6 +++-
2 files changed, 39 insertions(+), 33 deletions(-)
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
2026-09-28 4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
@ 2026-09-28 4:20 ` Jemmy Wong
2026-10-01 21:34 ` Tim Chen
2026-09-28 4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
2026-09-28 4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong
2 siblings, 1 reply; 8+ messages in thread
From: Jemmy Wong @ 2026-09-28 4:20 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel
The group-type matrix was introduced in commit 0b0695f2b34a
("sched/fair: Rework load_balance()"). When commit fee1759e4f04
("sched/fair: Determine active load balance for SMT sched groups")
added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
Prioritize tasks preferring destination LLC during balancing")
added group_llc_balance, neither commit updated the matrix table.
Both types are only tagged on non-local groups in update_sg_lb_stats(),
so their local columns are N/A.
As busiest, group_smt_balance is only set when dst_cpu is idle and the
SMT group runs more than one task. Against a local has_spare or
fully_busy group it goes through the nr_idle checks, where a non-SMT
dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
local imbalanced or overloaded group the local group is busier and the
pair is balanced.
As busiest, group_llc_balance is not an unconditional force. A local
overloaded group is busier and the pair is balanced; otherwise the
nr_idle checks apply, with prefer_sibling still able to force the pull
when the local group has spare capacity.
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 16 +++++++++-------
1 file changed, 9 insertions(+), 7 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 7455a83a6a99..f9ddcecfd19d 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -12947,13 +12947,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
/*
* Decision matrix according to the local and busiest group type:
*
- * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
- * has_spare nr_idle balanced N/A N/A balanced balanced
- * fully_busy nr_idle nr_idle N/A N/A balanced balanced
- * misfit_task force N/A N/A N/A N/A N/A
- * asym_packing force force N/A N/A force force
- * imbalanced force force N/A N/A force force
- * overloaded force force N/A N/A force avg_load
+ * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
+ * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
+ * misfit_task force N/A N/A N/A N/A N/A N/A N/A
+ * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
+ * asym_packing force force N/A N/A N/A force N/A force
+ * imbalanced force force N/A N/A N/A force N/A force
+ * llc_balance nr_idle nr_idle N/A N/A N/A nr_idle N/A balanced
+ * overloaded force force N/A N/A N/A force N/A avg_load
*
* N/A : Not Applicable because already filtered while updating
* statistics.
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task
2026-09-28 4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-09-28 4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
@ 2026-09-28 4:20 ` Jemmy Wong
2026-10-01 21:54 ` Tim Chen
2026-09-28 4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong
2 siblings, 1 reply; 8+ messages in thread
From: Jemmy Wong @ 2026-09-28 4:20 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel
The name dates back to the original group_imb heuristic, which flagged
a group when the load difference between its busiest and idlest CPU
exceeded the average task load. Commit 6263322c5e8f ("sched/fair:
Rewrite group_imb trigger") replaced that heuristic: the flag is now
set only when a lower domain fails to balance because tasks are pinned
by cpus_ptr (LBF_SOME_PINNED), and kept while all tasks are pinned
(LBF_ALL_PINNED). The name was carried over unchanged and later became
group_imbalanced in commit 0b0695f2b34a ("sched/fair: Rework
load_balance()").
Today the name no longer matches the condition:
- group_classify() checks group_overloaded first, so a group whose
load really is skewed is usually not classified as imbalanced,
while a flagged group may carry a single extra task.
- calculate_imbalance() does not measure any imbalance for this type;
it moves one task (migrate_task, imbalance = 1).
- Elsewhere in fair.c "imbalance" consistently means the amount of
load to move (env->imbalance, imbalance_pct, calculate_imbalance(),
the lb_imbalance_* schedstats), and imbalanced_active_balance() uses
"imbalanced" for repeated balance failures, unrelated to this flag.
Rename it to group_pinned_task, which describes the condition that
raises it and follows the adjective_noun pattern of group_misfit_task.
The enum is local to fair.c, so no tracepoint, schedstat or other
user-visible interface is affected.
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 30 +++++++++++++++---------------
1 file changed, 15 insertions(+), 15 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f9ddcecfd19d..aa63950976aa 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10363,7 +10363,7 @@ enum group_type {
* The tasks' affinity constraints previously prevented the scheduler
* from balancing the load across the system.
*/
- group_imbalanced,
+ group_pinned_task,
/*
* There are tasks running on non-preferred LLC, possible to move
* them to their preferred LLC without creating too much imbalance.
@@ -11701,7 +11701,7 @@ group_type group_classify(unsigned int imbalance_pct,
return group_llc_balance;
if (sg_imbalanced(group))
- return group_imbalanced;
+ return group_pinned_task;
if (sgs->group_asym_packing)
return group_asym_packing;
@@ -12164,10 +12164,10 @@ static bool update_sd_pick_busiest(struct lb_env *env,
/* Select the group with most tasks preferring dst LLC */
return update_llc_busiest(env, busiest, sgs);
- case group_imbalanced:
+ case group_pinned_task:
/*
- * Select the 1st imbalanced group as we don't have any way to
- * choose one more than another.
+ * Select the 1st group with pinned tasks as we don't
+ * have any way to choose one more than another.
*/
return false;
@@ -12416,7 +12416,7 @@ static bool update_pick_idlest(struct sched_group *idlest,
break;
case group_llc_balance:
- case group_imbalanced:
+ case group_pinned_task:
case group_asym_packing:
case group_smt_balance:
/* Those types are not used in the slow wakeup path */
@@ -12549,7 +12549,7 @@ sched_balance_find_dst_group(struct sched_domain *sd, struct task_struct *p, int
break;
case group_llc_balance:
- case group_imbalanced:
+ case group_pinned_task:
case group_asym_packing:
case group_smt_balance:
/* Those type are not used in the slow wakeup path */
@@ -12812,11 +12812,11 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
}
#endif
- if (busiest->group_type == group_imbalanced) {
+ if (busiest->group_type == group_pinned_task) {
/*
- * In the group_imb case we cannot rely on group-wide averages
- * to ensure CPU-load equilibrium, try to move any task to fix
- * the imbalance. The next load balance will take care of
+ * In the group_pinned_task case we cannot rely on group-wide
+ * averages to ensure CPU-load equilibrium, try to move any task
+ * to fix the imbalance. The next load balance will take care of
* balancing back the system.
*/
env->migration_type = migrate_task;
@@ -12947,13 +12947,13 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
/*
* Decision matrix according to the local and busiest group type:
*
- * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * busiest \ local has_spare fully_busy misfit smt asym pinned llc overloaded
* has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
* fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
* misfit_task force N/A N/A N/A N/A N/A N/A N/A
* smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
* asym_packing force force N/A N/A N/A force N/A force
- * imbalanced force force N/A N/A N/A force N/A force
+ * pinned_task force force N/A N/A N/A force N/A force
* llc_balance nr_idle nr_idle N/A N/A N/A nr_idle N/A balanced
* overloaded force force N/A N/A N/A force N/A avg_load
*
@@ -13008,11 +13008,11 @@ static struct sched_group *sched_balance_find_src_group(struct lb_env *env)
goto force_balance;
/*
- * If the busiest group is imbalanced the below checks don't
+ * If the busiest group has pinned tasks the below checks don't
* work because they assume all things are equal, which typically
* isn't true due to cpus_ptr constraints and the like.
*/
- if (busiest->group_type == group_imbalanced)
+ if (busiest->group_type == group_pinned_task)
goto force_balance;
local = &sds.local_stat;
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task
2026-09-28 4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-09-28 4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
2026-09-28 4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
@ 2026-09-28 4:20 ` Jemmy Wong
2 siblings, 0 replies; 8+ messages in thread
From: Jemmy Wong @ 2026-09-28 4:20 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel
sgc->imbalance does not describe a load skew either. It is set only
when a child domain failed to balance because some tasks were pinned by
cpus_ptr, and the parent then classifies the group as group_pinned_task
so that it can force a migration.
The old name also collides with env->imbalance, the amount of load
still to move, and with imbalance_pct. And after the previous rename,
sg_imbalanced() returns group_pinned_task, so the helper and the group
type no longer agree.
Rename the field to pinned_task and its reader to sg_pinned_task().
No functional change.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 24 ++++++++++++------------
kernel/sched/sched.h | 6 +++++-
2 files changed, 17 insertions(+), 13 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index aa63950976aa..43bdca636461 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -11595,7 +11595,7 @@ static inline bool check_misfit_status(struct rq *rq)
}
/*
- * Group imbalance indicates (and tries to solve) the problem where balancing
+ * sgc->pinned_task indicates (and tries to solve) the problem where balancing
* groups is inadequate due to ->cpus_ptr constraints.
*
* Imagine a situation of two groups of 4 CPUs each and 4 tasks each with a
@@ -11623,9 +11623,9 @@ static inline bool check_misfit_status(struct rq *rq)
* subtle and fragile situation.
*/
-static inline int sg_imbalanced(struct sched_group *group)
+static inline int sg_pinned_task(struct sched_group *group)
{
- return group->sgc->imbalance;
+ return group->sgc->pinned_task;
}
/*
@@ -11700,7 +11700,7 @@ group_type group_classify(unsigned int imbalance_pct,
if (sgs->group_llc_balance)
return group_llc_balance;
- if (sg_imbalanced(group))
+ if (sg_pinned_task(group))
return group_pinned_task;
if (sgs->group_asym_packing)
@@ -13619,10 +13619,10 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
* We failed to reach balance because of affinity.
*/
if (sd_parent) {
- int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+ int *pinned_task = &sd_parent->groups->sgc->pinned_task;
if ((env.flags & LBF_SOME_PINNED) && env.imbalance > 0)
- *group_imbalance = 1;
+ *pinned_task = 1;
}
/* All tasks on this runqueue were pinned by CPU affinity */
@@ -13722,21 +13722,21 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
out_balanced:
/*
* We reach balance although we may have faced some affinity
- * constraints. Clear the imbalance flag only if other tasks got
+ * constraints. Clear the pinned flag only if other tasks got
* a chance to move and fix the imbalance.
*/
if (sd_parent && !(env.flags & LBF_ALL_PINNED)) {
- int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+ int *pinned_task = &sd_parent->groups->sgc->pinned_task;
- if (*group_imbalance)
- *group_imbalance = 0;
+ if (*pinned_task)
+ *pinned_task = 0;
}
out_all_pinned:
/*
* We reach balance because all tasks are pinned at this level so
- * we can't migrate them. Let the imbalance flag set so parent level
- * can try to migrate them.
+ * we can't migrate them. Let the pinned flag stay set so the parent
+ * level can try to migrate them.
*/
schedstat_inc(sd->lb_balanced[idle]);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..b64966dde395 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2257,7 +2257,11 @@ struct sched_group_capacity {
unsigned long min_capacity; /* Min per-CPU capacity in group */
unsigned long max_capacity; /* Max per-CPU capacity in group */
unsigned long next_update;
- int imbalance; /* XXX unrelated to capacity but shared group state */
+ /*
+ * Set when a child domain cannot balance pinned tasks.
+ * XXX unrelated to capacity, but it is shared group state.
+ */
+ int pinned_task;
int id;
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
2026-09-28 4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
@ 2026-10-01 21:34 ` Tim Chen
2026-10-02 11:02 ` Jemmy Wong
0 siblings, 1 reply; 8+ messages in thread
From: Tim Chen @ 2026-10-01 21:34 UTC (permalink / raw)
To: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
On Mon, 2026-09-28 at 12:20 +0800, Jemmy Wong wrote:
> The group-type matrix was introduced in commit 0b0695f2b34a
> ("sched/fair: Rework load_balance()"). When commit fee1759e4f04
> ("sched/fair: Determine active load balance for SMT sched groups")
> added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
> Prioritize tasks preferring destination LLC during balancing")
> added group_llc_balance, neither commit updated the matrix table.
>
> Both types are only tagged on non-local groups in update_sg_lb_stats(),
> so their local columns are N/A.
>
> As busiest, group_smt_balance is only set when dst_cpu is idle and the
> SMT group runs more than one task. Against a local has_spare or
> fully_busy group it goes through the nr_idle checks, where a non-SMT
> dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
> local imbalanced or overloaded group the local group is busier and the
> pair is balanced.
>
> As busiest, group_llc_balance is not an unconditional force. A local
> overloaded group is busier and the pair is balanced; otherwise the
> nr_idle checks apply, with prefer_sibling still able to force the pull
> when the local group has spare capacity.
>
> No functional change.
>
> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
> ---
> kernel/sched/fair.c | 16 +++++++++-------
> 1 file changed, 9 insertions(+), 7 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 7455a83a6a99..f9ddcecfd19d 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -12947,13 +12947,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
> /*
> * Decision matrix according to the local and busiest group type:
> *
> - * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
> - * has_spare nr_idle balanced N/A N/A balanced balanced
> - * fully_busy nr_idle nr_idle N/A N/A balanced balanced
> - * misfit_task force N/A N/A N/A N/A N/A
> - * asym_packing force force N/A N/A force force
> - * imbalanced force force N/A N/A force force
> - * overloaded force force N/A N/A force avg_load
> + * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
> + * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
> + * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
> + * misfit_task force N/A N/A N/A N/A N/A N/A N/A
> + * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
> + * asym_packing force force N/A N/A N/A force N/A force
> + * imbalanced force force N/A N/A N/A force N/A force
Thanks for picking this up - nice to have the table match the code
again. I walked the new rows and columns against
sched_balance_find_src_group(), and they line up, with one exception:
> + * llc_balance nr_idle nr_idle N/A N/A N/A nr_idle N/A balanced
I think local=has_spare, busiest=llc_balance is "force", not "nr_idle":
if (sds.prefer_sibling && local->group_type == group_has_spare &&
(busiest->group_type == group_llc_balance ||
sibling_imbalance(env, &sds, busiest, local) > 1))
goto force_balance;
The group_llc_balance clause short-circuits the sibling_imbalance()
test, so this pair forces unconditionally once prefer_sibling is set,
and prefer_sibling is set for LLC-vs-LLC balancing: the groups span
per-LLC (SD_SHARE_LLC) domains, which keep SD_PREFER_SIBLING (only
SD_NUMA strips it). That also matches your changelog ("prefer_sibling
still able to force the pull...") - the table just reads nr_idle where
the prose says force.
Could you flip that one cell?
* llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
Thanks.
Tim
> + * overloaded force force N/A N/A N/A force N/A avg_load
> *
> * N/A : Not Applicable because already filtered while updating
> * statistics.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task
2026-09-28 4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
@ 2026-10-01 21:54 ` Tim Chen
2026-10-02 11:03 ` Jemmy Wong
0 siblings, 1 reply; 8+ messages in thread
From: Tim Chen @ 2026-10-01 21:54 UTC (permalink / raw)
To: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, linux-kernel
On Mon, 2026-09-28 at 12:20 +0800, Jemmy Wong wrote:
> The name dates back to the original group_imb heuristic, which flagged
> a group when the load difference between its busiest and idlest CPU
> exceeded the average task load. Commit 6263322c5e8f ("sched/fair:
> Rewrite group_imb trigger") replaced that heuristic: the flag is now
> set only when a lower domain fails to balance because tasks are pinned
> by cpus_ptr (LBF_SOME_PINNED), and kept while all tasks are pinned
> (LBF_ALL_PINNED). The name was carried over unchanged and later became
> group_imbalanced in commit 0b0695f2b34a ("sched/fair: Rework
> load_balance()").
>
> Today the name no longer matches the condition:
>
> - group_classify() checks group_overloaded first, so a group whose
> load really is skewed is usually not classified as imbalanced,
> while a flagged group may carry a single extra task.
>
> - calculate_imbalance() does not measure any imbalance for this type;
> it moves one task (migrate_task, imbalance = 1).
>
> - Elsewhere in fair.c "imbalance" consistently means the amount of
> load to move (env->imbalance, imbalance_pct, calculate_imbalance(),
> the lb_imbalance_* schedstats), and imbalanced_active_balance() uses
> "imbalanced" for repeated balance failures, unrelated to this flag.
>
> Rename it to group_pinned_task, which describes the condition that
> raises it and follows the adjective_noun pattern of group_misfit_task.
> The enum is local to fair.c, so no tracepoint, schedstat or other
> user-visible interface is affected.
>
> No functional change.
>
> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
> ---
> kernel/sched/fair.c | 30 +++++++++++++++---------------
> 1 file changed, 15 insertions(+), 15 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index f9ddcecfd19d..aa63950976aa 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -10363,7 +10363,7 @@ enum group_type {
> * The tasks' affinity constraints previously prevented the scheduler
> * from balancing the load across the system.
> */
> - group_imbalanced,
> + group_pinned_task,
I believe we are trying to fix imbalance due to pinned tasks.
Perhaps group_pinned_imbalance is more descriptive.
Tim
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
2026-10-01 21:34 ` Tim Chen
@ 2026-10-02 11:02 ` Jemmy Wong
0 siblings, 0 replies; 8+ messages in thread
From: Jemmy Wong @ 2026-10-02 11:02 UTC (permalink / raw)
To: Tim Chen
Cc: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli,
Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, linux-kernel
Hi Tim,
Thanks for the review.
> On Oct 2, 2026, at 5:34 AM, Tim Chen <tim.c.chen@linux.intel.com> wrote:
>
> On Mon, 2026-09-28 at 12:20 +0800, Jemmy Wong wrote:
>> The group-type matrix was introduced in commit 0b0695f2b34a
>> ("sched/fair: Rework load_balance()"). When commit fee1759e4f04
>> ("sched/fair: Determine active load balance for SMT sched groups")
>> added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
>> Prioritize tasks preferring destination LLC during balancing")
>> added group_llc_balance, neither commit updated the matrix table.
>>
>> Both types are only tagged on non-local groups in update_sg_lb_stats(),
>> so their local columns are N/A.
>>
>> As busiest, group_smt_balance is only set when dst_cpu is idle and the
>> SMT group runs more than one task. Against a local has_spare or
>> fully_busy group it goes through the nr_idle checks, where a non-SMT
>> dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
>> local imbalanced or overloaded group the local group is busier and the
>> pair is balanced.
>>
>> As busiest, group_llc_balance is not an unconditional force. A local
>> overloaded group is busier and the pair is balanced; otherwise the
>> nr_idle checks apply, with prefer_sibling still able to force the pull
>> when the local group has spare capacity.
>>
>> No functional change.
>>
>> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
>> ---
>> kernel/sched/fair.c | 16 +++++++++-------
>> 1 file changed, 9 insertions(+), 7 deletions(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 7455a83a6a99..f9ddcecfd19d 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -12947,13 +12947,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
>> /*
>> * Decision matrix according to the local and busiest group type:
>> *
>> - * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
>> - * has_spare nr_idle balanced N/A N/A balanced balanced
>> - * fully_busy nr_idle nr_idle N/A N/A balanced balanced
>> - * misfit_task force N/A N/A N/A N/A N/A
>> - * asym_packing force force N/A N/A force force
>> - * imbalanced force force N/A N/A force force
>> - * overloaded force force N/A N/A force avg_load
>> + * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
>> + * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
>> + * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
>> + * misfit_task force N/A N/A N/A N/A N/A N/A N/A
>> + * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
>> + * asym_packing force force N/A N/A N/A force N/A force
>> + * imbalanced force force N/A N/A N/A force N/A force
>
> Thanks for picking this up - nice to have the table match the code
> again. I walked the new rows and columns against
> sched_balance_find_src_group(), and they line up, with one exception:
>
>> + * llc_balance nr_idle nr_idle N/A N/A N/A nr_idle N/A balanced
>
> I think local=has_spare, busiest=llc_balance is "force", not "nr_idle":
>
> if (sds.prefer_sibling && local->group_type == group_has_spare &&
> (busiest->group_type == group_llc_balance ||
> sibling_imbalance(env, &sds, busiest, local) > 1))
> goto force_balance;
>
> The group_llc_balance clause short-circuits the sibling_imbalance()
> test, so this pair forces unconditionally once prefer_sibling is set,
> and prefer_sibling is set for LLC-vs-LLC balancing: the groups span
> per-LLC (SD_SHARE_LLC) domains, which keep SD_PREFER_SIBLING (only
> SD_NUMA strips it). That also matches your changelog ("prefer_sibling
> still able to force the pull...") - the table just reads nr_idle where
> the prose says force.
>
> Could you flip that one cell?
>
> * llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
You're right. The group_llc_balance clause short-circuits
sibling_imbalance(), so when the local group is has_spare and
prefer_sibling is set, a busiest llc_balance group is always pulled
from. I'll flip that cell to "force" in v2.
I also checked when prefer_sibling is set. It comes from the busiest
group's flags, i.e. the child domain's SD_PREFER_SIBLING (see
update_sd_lb_stats()). sd_init() sets that flag on every level, so SMT,
CLUSTER, MC and PKG all have it, and only SD_NUMA domains clear it. So
the pull is forced unless the child domain is a NUMA one. I'll reword
the changelog and cover letter to match.
> Thanks.
>
> Tim
>
>> + * overloaded force force N/A N/A N/A force N/A avg_load
>> *
>> * N/A : Not Applicable because already filtered while updating
>> * statistics.
Thanks,
Jemmy
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task
2026-10-01 21:54 ` Tim Chen
@ 2026-10-02 11:03 ` Jemmy Wong
0 siblings, 0 replies; 8+ messages in thread
From: Jemmy Wong @ 2026-10-02 11:03 UTC (permalink / raw)
To: Tim Chen
Cc: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli,
Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, linux-kernel
Hi Tim,
Thanks for the review.
> On Oct 2, 2026, at 5:54 AM, Tim Chen <tim.c.chen@linux.intel.com> wrote:
>
> On Mon, 2026-09-28 at 12:20 +0800, Jemmy Wong wrote:
>> The name dates back to the original group_imb heuristic, which flagged
>> a group when the load difference between its busiest and idlest CPU
>> exceeded the average task load. Commit 6263322c5e8f ("sched/fair:
>> Rewrite group_imb trigger") replaced that heuristic: the flag is now
>> set only when a lower domain fails to balance because tasks are pinned
>> by cpus_ptr (LBF_SOME_PINNED), and kept while all tasks are pinned
>> (LBF_ALL_PINNED). The name was carried over unchanged and later became
>> group_imbalanced in commit 0b0695f2b34a ("sched/fair: Rework
>> load_balance()").
>>
>> Today the name no longer matches the condition:
>>
>> - group_classify() checks group_overloaded first, so a group whose
>> load really is skewed is usually not classified as imbalanced,
>> while a flagged group may carry a single extra task.
>>
>> - calculate_imbalance() does not measure any imbalance for this type;
>> it moves one task (migrate_task, imbalance = 1).
>>
>> - Elsewhere in fair.c "imbalance" consistently means the amount of
>> load to move (env->imbalance, imbalance_pct, calculate_imbalance(),
>> the lb_imbalance_* schedstats), and imbalanced_active_balance() uses
>> "imbalanced" for repeated balance failures, unrelated to this flag.
>>
>> Rename it to group_pinned_task, which describes the condition that
>> raises it and follows the adjective_noun pattern of group_misfit_task.
>> The enum is local to fair.c, so no tracepoint, schedstat or other
>> user-visible interface is affected.
>>
>> No functional change.
>>
>> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
>> ---
>> kernel/sched/fair.c | 30 +++++++++++++++---------------
>> 1 file changed, 15 insertions(+), 15 deletions(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index f9ddcecfd19d..aa63950976aa 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -10363,7 +10363,7 @@ enum group_type {
>> * The tasks' affinity constraints previously prevented the scheduler
>> * from balancing the load across the system.
>> */
>> - group_imbalanced,
>> + group_pinned_task,
>
> I believe we are trying to fix imbalance due to pinned tasks.
> Perhaps group_pinned_imbalance is more descriptive.
>
> Tim
Agreed. I'll use group_pinned_imb in v2, which keeps the "pinned"
qualifier and matches the short "imb" naming. I'll also rename
sgc->imbalance to sgc->pinned_imb and its reader to sg_pinned_imb() in
patch 3 to match.
Thanks,
Jemmy
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-10-02 11:03 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-09-28 4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
2026-10-01 21:34 ` Tim Chen
2026-10-02 11:02 ` Jemmy Wong
2026-09-28 4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
2026-10-01 21:54 ` Tim Chen
2026-10-02 11:03 ` Jemmy Wong
2026-09-28 4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®