mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
@ 2026-09-27 16:34 Waiman Long
  2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
  2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long
  0 siblings, 2 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
  To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
  Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long

When testing some upstream cpuset patch, it was found that the current
SCHED_DEADLINE shrink test in validate_change() can incorrectly get
triggered due to the fact that the exclusive flag may get set on a
cpuset that is not a valid partition root. This series fixes this
problem by making sure that the test will only be triggered on a valid
partition root and remove the now unnecessary code to handle the
exclusive flag in the v2 code.

Waiman Long (2):
  cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root
    only
  cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2

 kernel/cgroup/cpuset-v1.c | 20 ++++++++++--
 kernel/cgroup/cpuset.c    | 69 +++++----------------------------------
 2 files changed, 27 insertions(+), 62 deletions(-)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH-next 1/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
  2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
@ 2026-09-27 16:34 ` Waiman Long
  2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long
  1 sibling, 0 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
  To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
  Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long

Commit f82f80426f7a ("sched/deadline: Ensure that updates to exclusive
cpusets don't break AC") adds a check in validate_change() to make
sure that there is enough bandwidth for SCHED_DEADLINE tasks if we
shrink a v1 exclusive cpuset that has CS_CPU_EXCLUSIVE flag set.

With the introduction of cpuset partition in cgroup v2, we keep setting
the CS_CPU_EXCLUSIVE flag for a partition root so that the SCHED_DEADLINE
check will continue to work as intended. However it turns out that the
current code isn't perfect and there are cases where a cpuset isn't
a valid partition root, but the exclusive flag is still incorrectly
set. This can leads to SCHED_DEADLINE check being incorrectly triggered
when there are deadline tasks in the system. This can result in unexpected
-EBUSY failure when making changes to cpuset control files. Fix that
by checking for a valid partition root in the case of v2 and break out
the v1 specific check back into cpuset1_validate_change(). It is far
easier and less cumbersome than to make sure that the exclusive flag
is only set for valid partition roots.

The is_in_v2_mode() check guarding the call to cpuset1_validate_change()
is also moved to the appropriate place inside cpuset1_validate_change()
and a cpuset_v2() is now being used as a guard which should be more
accurate for the a v1 system with v2 mode enabled.

Even though commit a86ce68078b2 ("cgroup/cpuset: Extract out
CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling") is marked as a
commit to be fixed, the problem may exist before that.

Fixes: a86ce68078b2 ("cgroup/cpuset: Extract out CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling")
Signed-off-by: Waiman Long <longman@redhat.com>
---
 kernel/cgroup/cpuset-v1.c | 20 ++++++++++++++++++--
 kernel/cgroup/cpuset.c    | 17 +++++------------
 2 files changed, 23 insertions(+), 14 deletions(-)

diff --git a/kernel/cgroup/cpuset-v1.c b/kernel/cgroup/cpuset-v1.c
index 562ad35f00d0..c03ae8aac03a 100644
--- a/kernel/cgroup/cpuset-v1.c
+++ b/kernel/cgroup/cpuset-v1.c
@@ -61,6 +61,11 @@ struct cpuset_remove_tasks_struct {
 #define FM_MAXCNT 1000000	/* limit cnt to avoid overflow */
 #define FM_SCALE 1000		/* faux fixed point scale */
 
+static inline bool is_in_v2_mode(void)
+{
+	return cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE;
+}
+
 /* Initialize a frequency meter */
 static void fmeter_init(struct fmeter *fmp)
 {
@@ -357,10 +362,21 @@ int cpuset1_validate_change(struct cpuset *cur, struct cpuset *trial)
 		if (!is_cpuset_subset(c, trial))
 			goto out;
 
-	/* On legacy hierarchy, we must be a subset of our parent cpuset. */
+	/*
+	 * We can't shrink if we won't have enough room for SCHED_DEADLINE
+	 * tasks in a scheduling partition.
+	 */
+	if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+	    !cpuset_cpumask_can_shrink(cur->cpus_allowed, trial->cpus_allowed))
+		goto out;
+
+	/*
+	 * On legacy hierarchy with v2 mode off, we must be a subset of our
+	 * parent cpuset.
+	 */
 	ret = -EACCES;
 	par = parent_cs(cur);
-	if (par && !is_cpuset_subset(trial, par))
+	if (par && !is_in_v2_mode() && !is_cpuset_subset(trial, par))
 		goto out;
 
 	/*
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 753aa65afcd7..5639c486c967 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -752,7 +752,7 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
 
 	rcu_read_lock();
 
-	if (!is_in_v2_mode())
+	if (!cpuset_v2())
 		ret = cpuset1_validate_change(cur, trial);
 	if (ret)
 		goto out;
@@ -765,23 +765,16 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
 
 	/*
 	 * We can't shrink if we won't have enough room for SCHED_DEADLINE
-	 * tasks. This check is not done when scheduling is disabled as the
-	 * users should know what they are doing.
-	 *
-	 * For v1, effective_cpus == cpus_allowed & user_xcpus() returns
-	 * cpus_allowed.
-	 *
-	 * For v2, is_cpu_exclusive() & is_sched_load_balance() are true only
-	 * for non-isolated partition root. At this point, the target
-	 * effective_cpus isn't computed yet. user_xcpus() is the best
-	 * approximation.
+	 * tasks. This check is only done on non-isolated partition root.
+	 * At this point, the target effective_cpus isn't computed yet.
+	 * user_xcpus() is the best approximation.
 	 *
 	 * TBD: May need to precompute the real effective_cpus here in case
 	 * incorrect scheduling of SCHED_DEADLINE tasks in a partition
 	 * becomes an issue.
 	 */
 	ret = -EBUSY;
-	if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+	if (is_partition_valid(cur) && is_sched_load_balance(cur) &&
 	    !cpuset_cpumask_can_shrink(cur->effective_cpus, user_xcpus(trial)))
 		goto out;
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2
  2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
  2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
@ 2026-09-27 16:34 ` Waiman Long
  1 sibling, 0 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
  To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
  Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long

With the move of is_cpu_exclusive() DEADLINE shrink test into
cpuset1_validate_change() in the previous patch, there is no need to
have the CS_CPU_EXCLUSIVE flag set on valid partition root anymore.
Remove update_partition_exclusive_flag() and other CS_CPU_EXCLUSIVE
flag handling code in cpuset.c.

Signed-off-by: Waiman Long <longman@redhat.com>
---
 kernel/cgroup/cpuset.c | 52 ++++--------------------------------------
 1 file changed, 4 insertions(+), 48 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 5639c486c967..b07a7e8cef16 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1166,25 +1166,6 @@ enum partition_cmd {
 static void update_sibling_cpumasks(struct cpuset *parent, struct cpuset *cs,
 				    struct tmpmasks *tmp);
 
-/*
- * Update partition exclusive flag
- *
- * Return: 0 if successful, an error code otherwise
- */
-static int update_partition_exclusive_flag(struct cpuset *cs, int new_prs)
-{
-	bool exclusive = (new_prs > PRS_MEMBER);
-
-	if (exclusive && !is_cpu_exclusive(cs)) {
-		if (cpuset_update_flag(CS_CPU_EXCLUSIVE, cs, 1))
-			return PERR_NOTEXCL;
-	} else if (!exclusive && is_cpu_exclusive(cs)) {
-		/* Turning off CS_CPU_EXCLUSIVE will not return error */
-		cpuset_update_flag(CS_CPU_EXCLUSIVE, cs, 0);
-	}
-	return 0;
-}
-
 /*
  * Update partition load balance flag and/or rebuild sched domain
  *
@@ -1240,11 +1221,9 @@ static void reset_partition_data(struct cpuset *cs)
 
 	lockdep_assert_held(&callback_lock);
 
-	if (cpumask_empty(cs->exclusive_cpus)) {
+	if (cpumask_empty(cs->exclusive_cpus))
 		cpumask_clear(cs->effective_xcpus);
-		if (is_cpu_exclusive(cs))
-			clear_bit(CS_CPU_EXCLUSIVE, &cs->flags);
-	}
+
 	if (!cpumask_and(cs->effective_cpus, parent->effective_cpus, cs->cpus_allowed))
 		cpumask_copy(cs->effective_cpus, parent->effective_cpus);
 }
@@ -2020,19 +1999,6 @@ static int update_parent_effective_cpumask(struct cpuset *cs, int cmd,
 	if (!adding && !deleting && (new_prs == old_prs))
 		return 0;
 
-	/*
-	 * Transitioning between invalid to valid or vice versa may require
-	 * changing CS_CPU_EXCLUSIVE. In the case of partcmd_update,
-	 * validate_change() has already been successfully called and
-	 * CPU lists in cs haven't been updated yet. So defer it to later.
-	 */
-	if ((old_prs != new_prs) && (cmd != partcmd_update))  {
-		int err = update_partition_exclusive_flag(cs, new_prs);
-
-		if (err)
-			return err;
-	}
-
 	/*
 	 * Change the parent's effective_cpus & effective_xcpus (top cpuset
 	 * only).
@@ -2055,9 +2021,6 @@ static int update_parent_effective_cpumask(struct cpuset *cs, int cmd,
 
 	spin_unlock_irq(&callback_lock);
 
-	if ((old_prs != new_prs) && (cmd == partcmd_update))
-		update_partition_exclusive_flag(cs, new_prs);
-
 	if (adding || deleting) {
 		cpuset_update_tasks_cpumask(parent, tmp->addmask);
 		update_sibling_cpumasks(parent, cs, tmp);
@@ -2934,10 +2897,6 @@ static int update_prstate(struct cpuset *cs, int new_prs)
 	if (alloc_tmpmasks(&tmpmask))
 		return -ENOMEM;
 
-	err = update_partition_exclusive_flag(cs, new_prs);
-	if (err)
-		goto out;
-
 	if (!old_prs) {
 		/*
 		 * cpus_allowed and exclusive_cpus cannot be both empty.
@@ -3001,13 +2960,10 @@ static int update_prstate(struct cpuset *cs, int new_prs)
 	}
 out:
 	/*
-	 * Make partition invalid & disable CS_CPU_EXCLUSIVE if an error
-	 * happens.
+	 * Make partition invalid if an error happens.
 	 */
-	if (err) {
+	if (err)
 		new_prs = -new_prs;
-		update_partition_exclusive_flag(cs, new_prs);
-	}
 
 	spin_lock_irq(&callback_lock);
 	cs->partition_root_state = new_prs;
-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-27 16:35 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®