mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state
@ 2026-09-28  4:20 Jemmy Wong
  2026-09-28  4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Jemmy Wong @ 2026-09-28  4:20 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel

Patch 1 brings the group-type decision matrix above
sched_balance_find_src_group() up to date. The matrix was introduced
by Vincent Guittot in commit 0b0695f2b34a ("sched/fair: Rework
load_balance()"). When commit fee1759e4f04 ("sched/fair: Determine
active load balance for SMT sched groups") added group_smt_balance
and commit f38cc2f0d8a3 ("sched/cache: Prioritize tasks preferring
destination LLC during balancing") added group_llc_balance, neither
author updated the matrix table.

Both are only tagged on non-local groups, so their columns are N/A.
As busiest, both go through the nr_idle checks rather than an
unconditional force, and both are balanced against a busier local
group.

Patches 2-3 rename the state that tracks load balancing failures caused
by cpus_ptr-pinned tasks. The names come from the original group_imb
heuristic, which did detect per-CPU load skew; commit 6263322c5e8f
("sched/fair: Rewrite group_imb trigger") turned it into an affinity
failure flag but kept the name. Today group_imbalanced is not what
group_classify() returns for skewed load, calculate_imbalance()
computes no imbalance for it, and sgc->imbalance collides with
env->imbalance and imbalance_pct. They become group_pinned_task and
sgc->pinned_task, matching the LBF_*_PINNED flags that set them.

Patch 1 does not depend on patches 2-3 and can be applied on its own.

No functional change.

Jemmy Wong (3):
  sched/fair: Add smt_balance and llc_balance to the decision matrix
  sched/fair: Rename group_imbalanced to group_pinned_task
  sched/fair: Rename sgc->imbalance to sgc->pinned_task

 kernel/sched/fair.c  | 66 +++++++++++++++++++++++---------------------
 kernel/sched/sched.h |  6 +++-
 2 files changed, 39 insertions(+), 33 deletions(-)

--
2.54.0 (Apple Git-157)

^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
  2026-09-28  4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
@ 2026-09-28  4:20 ` Jemmy Wong
  2026-09-28  4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
  2026-09-28  4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong
  2 siblings, 0 replies; 4+ messages in thread
From: Jemmy Wong @ 2026-09-28  4:20 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel

The group-type matrix was introduced in commit 0b0695f2b34a
("sched/fair: Rework load_balance()"). When commit fee1759e4f04
("sched/fair: Determine active load balance for SMT sched groups")
added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
Prioritize tasks preferring destination LLC during balancing")
added group_llc_balance, neither commit updated the matrix table.

Both types are only tagged on non-local groups in update_sg_lb_stats(),
so their local columns are N/A.

As busiest, group_smt_balance is only set when dst_cpu is idle and the
SMT group runs more than one task. Against a local has_spare or
fully_busy group it goes through the nr_idle checks, where a non-SMT
dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
local imbalanced or overloaded group the local group is busier and the
pair is balanced.

As busiest, group_llc_balance is not an unconditional force. A local
overloaded group is busier and the pair is balanced; otherwise the
nr_idle checks apply, with prefer_sibling still able to force the pull
when the local group has spare capacity.

No functional change.

Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
 kernel/sched/fair.c | 16 +++++++++-------
 1 file changed, 9 insertions(+), 7 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 7455a83a6a99..f9ddcecfd19d 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -12947,13 +12947,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
 /*
  * Decision matrix according to the local and busiest group type:
  *
- * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
- * has_spare        nr_idle   balanced   N/A    N/A  balanced   balanced
- * fully_busy       nr_idle   nr_idle    N/A    N/A  balanced   balanced
- * misfit_task      force     N/A        N/A    N/A  N/A        N/A
- * asym_packing     force     force      N/A    N/A  force      force
- * imbalanced       force     force      N/A    N/A  force      force
- * overloaded       force     force      N/A    N/A  force      avg_load
+ * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * has_spare        nr_idle   balanced   N/A   N/A  N/A  balanced  N/A  balanced
+ * fully_busy       nr_idle   nr_idle    N/A   N/A  N/A  balanced  N/A  balanced
+ * misfit_task      force     N/A        N/A   N/A  N/A  N/A       N/A  N/A
+ * smt_balance      nr_idle   nr_idle    N/A   N/A  N/A  balanced  N/A  balanced
+ * asym_packing     force     force      N/A   N/A  N/A  force     N/A  force
+ * imbalanced       force     force      N/A   N/A  N/A  force     N/A  force
+ * llc_balance      nr_idle   nr_idle    N/A   N/A  N/A  nr_idle   N/A  balanced
+ * overloaded       force     force      N/A   N/A  N/A  force     N/A  avg_load
  *
  * N/A :      Not Applicable because already filtered while updating
  *            statistics.
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task
  2026-09-28  4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
  2026-09-28  4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
@ 2026-09-28  4:20 ` Jemmy Wong
  2026-09-28  4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong
  2 siblings, 0 replies; 4+ messages in thread
From: Jemmy Wong @ 2026-09-28  4:20 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel

The name dates back to the original group_imb heuristic, which flagged
a group when the load difference between its busiest and idlest CPU
exceeded the average task load. Commit 6263322c5e8f ("sched/fair:
Rewrite group_imb trigger") replaced that heuristic: the flag is now
set only when a lower domain fails to balance because tasks are pinned
by cpus_ptr (LBF_SOME_PINNED), and kept while all tasks are pinned
(LBF_ALL_PINNED). The name was carried over unchanged and later became
group_imbalanced in commit 0b0695f2b34a ("sched/fair: Rework
load_balance()").

Today the name no longer matches the condition:

 - group_classify() checks group_overloaded first, so a group whose
   load really is skewed is usually not classified as imbalanced,
   while a flagged group may carry a single extra task.

 - calculate_imbalance() does not measure any imbalance for this type;
   it moves one task (migrate_task, imbalance = 1).

 - Elsewhere in fair.c "imbalance" consistently means the amount of
   load to move (env->imbalance, imbalance_pct, calculate_imbalance(),
   the lb_imbalance_* schedstats), and imbalanced_active_balance() uses
   "imbalanced" for repeated balance failures, unrelated to this flag.

Rename it to group_pinned_task, which describes the condition that
raises it and follows the adjective_noun pattern of group_misfit_task.
The enum is local to fair.c, so no tracepoint, schedstat or other
user-visible interface is affected.

No functional change.

Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
 kernel/sched/fair.c | 30 +++++++++++++++---------------
 1 file changed, 15 insertions(+), 15 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f9ddcecfd19d..aa63950976aa 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10363,7 +10363,7 @@ enum group_type {
 	 * The tasks' affinity constraints previously prevented the scheduler
 	 * from balancing the load across the system.
 	 */
-	group_imbalanced,
+	group_pinned_task,
 	/*
 	 * There are tasks running on non-preferred LLC, possible to move
 	 * them to their preferred LLC without creating too much imbalance.
@@ -11701,7 +11701,7 @@ group_type group_classify(unsigned int imbalance_pct,
 		return group_llc_balance;
 
 	if (sg_imbalanced(group))
-		return group_imbalanced;
+		return group_pinned_task;
 
 	if (sgs->group_asym_packing)
 		return group_asym_packing;
@@ -12164,10 +12164,10 @@ static bool update_sd_pick_busiest(struct lb_env *env,
 		/* Select the group with most tasks preferring dst LLC */
 		return update_llc_busiest(env, busiest, sgs);
 
-	case group_imbalanced:
+	case group_pinned_task:
 		/*
-		 * Select the 1st imbalanced group as we don't have any way to
-		 * choose one more than another.
+		 * Select the 1st group with pinned tasks as we don't
+		 * have any way to choose one more than another.
 		 */
 		return false;
 
@@ -12416,7 +12416,7 @@ static bool update_pick_idlest(struct sched_group *idlest,
 		break;
 
 	case group_llc_balance:
-	case group_imbalanced:
+	case group_pinned_task:
 	case group_asym_packing:
 	case group_smt_balance:
 		/* Those types are not used in the slow wakeup path */
@@ -12549,7 +12549,7 @@ sched_balance_find_dst_group(struct sched_domain *sd, struct task_struct *p, int
 		break;
 
 	case group_llc_balance:
-	case group_imbalanced:
+	case group_pinned_task:
 	case group_asym_packing:
 	case group_smt_balance:
 		/* Those type are not used in the slow wakeup path */
@@ -12812,11 +12812,11 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
 	}
 #endif
 
-	if (busiest->group_type == group_imbalanced) {
+	if (busiest->group_type == group_pinned_task) {
 		/*
-		 * In the group_imb case we cannot rely on group-wide averages
-		 * to ensure CPU-load equilibrium, try to move any task to fix
-		 * the imbalance. The next load balance will take care of
+		 * In the group_pinned_task case we cannot rely on group-wide
+		 * averages to ensure CPU-load equilibrium, try to move any task
+		 * to fix the imbalance. The next load balance will take care of
 		 * balancing back the system.
 		 */
 		env->migration_type = migrate_task;
@@ -12947,13 +12947,13 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
 /*
  * Decision matrix according to the local and busiest group type:
  *
- * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
+ * busiest \ local has_spare fully_busy misfit smt asym   pinned   llc overloaded
  * has_spare        nr_idle   balanced   N/A   N/A  N/A  balanced  N/A  balanced
  * fully_busy       nr_idle   nr_idle    N/A   N/A  N/A  balanced  N/A  balanced
  * misfit_task      force     N/A        N/A   N/A  N/A  N/A       N/A  N/A
  * smt_balance      nr_idle   nr_idle    N/A   N/A  N/A  balanced  N/A  balanced
  * asym_packing     force     force      N/A   N/A  N/A  force     N/A  force
- * imbalanced       force     force      N/A   N/A  N/A  force     N/A  force
+ * pinned_task      force     force      N/A   N/A  N/A  force     N/A  force
  * llc_balance      nr_idle   nr_idle    N/A   N/A  N/A  nr_idle   N/A  balanced
  * overloaded       force     force      N/A   N/A  N/A  force     N/A  avg_load
  *
@@ -13008,11 +13008,11 @@ static struct sched_group *sched_balance_find_src_group(struct lb_env *env)
 		goto force_balance;
 
 	/*
-	 * If the busiest group is imbalanced the below checks don't
+	 * If the busiest group has pinned tasks the below checks don't
 	 * work because they assume all things are equal, which typically
 	 * isn't true due to cpus_ptr constraints and the like.
 	 */
-	if (busiest->group_type == group_imbalanced)
+	if (busiest->group_type == group_pinned_task)
 		goto force_balance;
 
 	local = &sds.local_stat;
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task
  2026-09-28  4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
  2026-09-28  4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
  2026-09-28  4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
@ 2026-09-28  4:20 ` Jemmy Wong
  2 siblings, 0 replies; 4+ messages in thread
From: Jemmy Wong @ 2026-09-28  4:20 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Tim Chen, linux-kernel

sgc->imbalance does not describe a load skew either. It is set only
when a child domain failed to balance because some tasks were pinned by
cpus_ptr, and the parent then classifies the group as group_pinned_task
so that it can force a migration.

The old name also collides with env->imbalance, the amount of load
still to move, and with imbalance_pct. And after the previous rename,
sg_imbalanced() returns group_pinned_task, so the helper and the group
type no longer agree.

Rename the field to pinned_task and its reader to sg_pinned_task().

No functional change.

Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
 kernel/sched/fair.c  | 24 ++++++++++++------------
 kernel/sched/sched.h |  6 +++++-
 2 files changed, 17 insertions(+), 13 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index aa63950976aa..43bdca636461 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -11595,7 +11595,7 @@ static inline bool check_misfit_status(struct rq *rq)
 }
 
 /*
- * Group imbalance indicates (and tries to solve) the problem where balancing
+ * sgc->pinned_task indicates (and tries to solve) the problem where balancing
  * groups is inadequate due to ->cpus_ptr constraints.
  *
  * Imagine a situation of two groups of 4 CPUs each and 4 tasks each with a
@@ -11623,9 +11623,9 @@ static inline bool check_misfit_status(struct rq *rq)
  * subtle and fragile situation.
  */
 
-static inline int sg_imbalanced(struct sched_group *group)
+static inline int sg_pinned_task(struct sched_group *group)
 {
-	return group->sgc->imbalance;
+	return group->sgc->pinned_task;
 }
 
 /*
@@ -11700,7 +11700,7 @@ group_type group_classify(unsigned int imbalance_pct,
 	if (sgs->group_llc_balance)
 		return group_llc_balance;
 
-	if (sg_imbalanced(group))
+	if (sg_pinned_task(group))
 		return group_pinned_task;
 
 	if (sgs->group_asym_packing)
@@ -13619,10 +13619,10 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
 		 * We failed to reach balance because of affinity.
 		 */
 		if (sd_parent) {
-			int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+			int *pinned_task = &sd_parent->groups->sgc->pinned_task;
 
 			if ((env.flags & LBF_SOME_PINNED) && env.imbalance > 0)
-				*group_imbalance = 1;
+				*pinned_task = 1;
 		}
 
 		/* All tasks on this runqueue were pinned by CPU affinity */
@@ -13722,21 +13722,21 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
 out_balanced:
 	/*
 	 * We reach balance although we may have faced some affinity
-	 * constraints. Clear the imbalance flag only if other tasks got
+	 * constraints. Clear the pinned flag only if other tasks got
 	 * a chance to move and fix the imbalance.
 	 */
 	if (sd_parent && !(env.flags & LBF_ALL_PINNED)) {
-		int *group_imbalance = &sd_parent->groups->sgc->imbalance;
+		int *pinned_task = &sd_parent->groups->sgc->pinned_task;
 
-		if (*group_imbalance)
-			*group_imbalance = 0;
+		if (*pinned_task)
+			*pinned_task = 0;
 	}
 
 out_all_pinned:
 	/*
 	 * We reach balance because all tasks are pinned at this level so
-	 * we can't migrate them. Let the imbalance flag set so parent level
-	 * can try to migrate them.
+	 * we can't migrate them. Let the pinned flag stay set so the parent
+	 * level can try to migrate them.
 	 */
 	schedstat_inc(sd->lb_balanced[idle]);
 
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..b64966dde395 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2257,7 +2257,11 @@ struct sched_group_capacity {
 	unsigned long		min_capacity;		/* Min per-CPU capacity in group */
 	unsigned long		max_capacity;		/* Max per-CPU capacity in group */
 	unsigned long		next_update;
-	int			imbalance;		/* XXX unrelated to capacity but shared group state */
+	/*
+	 * Set when a child domain cannot balance pinned tasks.
+	 * XXX unrelated to capacity, but it is shared group state.
+	 */
+	int			pinned_task;
 
 	int			id;
 
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-28  4:21 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28  4:20 [PATCH 0/3] sched/fair: Update decision matrix and rename pinned-task group state Jemmy Wong
2026-09-28  4:20 ` [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix Jemmy Wong
2026-09-28  4:20 ` [PATCH 2/3] sched/fair: Rename group_imbalanced to group_pinned_task Jemmy Wong
2026-09-28  4:20 ` [PATCH 3/3] sched/fair: Rename sgc->imbalance to sgc->pinned_task Jemmy Wong

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®