* [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
@ 2026-09-27 16:34 Waiman Long
2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long
0 siblings, 2 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long
When testing some upstream cpuset patch, it was found that the current
SCHED_DEADLINE shrink test in validate_change() can incorrectly get
triggered due to the fact that the exclusive flag may get set on a
cpuset that is not a valid partition root. This series fixes this
problem by making sure that the test will only be triggered on a valid
partition root and remove the now unnecessary code to handle the
exclusive flag in the v2 code.
Waiman Long (2):
cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root
only
cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2
kernel/cgroup/cpuset-v1.c | 20 ++++++++++--
kernel/cgroup/cpuset.c | 69 +++++----------------------------------
2 files changed, 27 insertions(+), 62 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH-next 1/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
@ 2026-09-27 16:34 ` Waiman Long
2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long
1 sibling, 0 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long
Commit f82f80426f7a ("sched/deadline: Ensure that updates to exclusive
cpusets don't break AC") adds a check in validate_change() to make
sure that there is enough bandwidth for SCHED_DEADLINE tasks if we
shrink a v1 exclusive cpuset that has CS_CPU_EXCLUSIVE flag set.
With the introduction of cpuset partition in cgroup v2, we keep setting
the CS_CPU_EXCLUSIVE flag for a partition root so that the SCHED_DEADLINE
check will continue to work as intended. However it turns out that the
current code isn't perfect and there are cases where a cpuset isn't
a valid partition root, but the exclusive flag is still incorrectly
set. This can leads to SCHED_DEADLINE check being incorrectly triggered
when there are deadline tasks in the system. This can result in unexpected
-EBUSY failure when making changes to cpuset control files. Fix that
by checking for a valid partition root in the case of v2 and break out
the v1 specific check back into cpuset1_validate_change(). It is far
easier and less cumbersome than to make sure that the exclusive flag
is only set for valid partition roots.
The is_in_v2_mode() check guarding the call to cpuset1_validate_change()
is also moved to the appropriate place inside cpuset1_validate_change()
and a cpuset_v2() is now being used as a guard which should be more
accurate for the a v1 system with v2 mode enabled.
Even though commit a86ce68078b2 ("cgroup/cpuset: Extract out
CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling") is marked as a
commit to be fixed, the problem may exist before that.
Fixes: a86ce68078b2 ("cgroup/cpuset: Extract out CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling")
Signed-off-by: Waiman Long <longman@redhat.com>
---
kernel/cgroup/cpuset-v1.c | 20 ++++++++++++++++++--
kernel/cgroup/cpuset.c | 17 +++++------------
2 files changed, 23 insertions(+), 14 deletions(-)
diff --git a/kernel/cgroup/cpuset-v1.c b/kernel/cgroup/cpuset-v1.c
index 562ad35f00d0..c03ae8aac03a 100644
--- a/kernel/cgroup/cpuset-v1.c
+++ b/kernel/cgroup/cpuset-v1.c
@@ -61,6 +61,11 @@ struct cpuset_remove_tasks_struct {
#define FM_MAXCNT 1000000 /* limit cnt to avoid overflow */
#define FM_SCALE 1000 /* faux fixed point scale */
+static inline bool is_in_v2_mode(void)
+{
+ return cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE;
+}
+
/* Initialize a frequency meter */
static void fmeter_init(struct fmeter *fmp)
{
@@ -357,10 +362,21 @@ int cpuset1_validate_change(struct cpuset *cur, struct cpuset *trial)
if (!is_cpuset_subset(c, trial))
goto out;
- /* On legacy hierarchy, we must be a subset of our parent cpuset. */
+ /*
+ * We can't shrink if we won't have enough room for SCHED_DEADLINE
+ * tasks in a scheduling partition.
+ */
+ if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+ !cpuset_cpumask_can_shrink(cur->cpus_allowed, trial->cpus_allowed))
+ goto out;
+
+ /*
+ * On legacy hierarchy with v2 mode off, we must be a subset of our
+ * parent cpuset.
+ */
ret = -EACCES;
par = parent_cs(cur);
- if (par && !is_cpuset_subset(trial, par))
+ if (par && !is_in_v2_mode() && !is_cpuset_subset(trial, par))
goto out;
/*
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 753aa65afcd7..5639c486c967 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -752,7 +752,7 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
rcu_read_lock();
- if (!is_in_v2_mode())
+ if (!cpuset_v2())
ret = cpuset1_validate_change(cur, trial);
if (ret)
goto out;
@@ -765,23 +765,16 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
/*
* We can't shrink if we won't have enough room for SCHED_DEADLINE
- * tasks. This check is not done when scheduling is disabled as the
- * users should know what they are doing.
- *
- * For v1, effective_cpus == cpus_allowed & user_xcpus() returns
- * cpus_allowed.
- *
- * For v2, is_cpu_exclusive() & is_sched_load_balance() are true only
- * for non-isolated partition root. At this point, the target
- * effective_cpus isn't computed yet. user_xcpus() is the best
- * approximation.
+ * tasks. This check is only done on non-isolated partition root.
+ * At this point, the target effective_cpus isn't computed yet.
+ * user_xcpus() is the best approximation.
*
* TBD: May need to precompute the real effective_cpus here in case
* incorrect scheduling of SCHED_DEADLINE tasks in a partition
* becomes an issue.
*/
ret = -EBUSY;
- if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+ if (is_partition_valid(cur) && is_sched_load_balance(cur) &&
!cpuset_cpumask_can_shrink(cur->effective_cpus, user_xcpus(trial)))
goto out;
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2
2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
@ 2026-09-27 16:34 ` Waiman Long
1 sibling, 0 replies; 3+ messages in thread
From: Waiman Long @ 2026-09-27 16:34 UTC (permalink / raw)
To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long
With the move of is_cpu_exclusive() DEADLINE shrink test into
cpuset1_validate_change() in the previous patch, there is no need to
have the CS_CPU_EXCLUSIVE flag set on valid partition root anymore.
Remove update_partition_exclusive_flag() and other CS_CPU_EXCLUSIVE
flag handling code in cpuset.c.
Signed-off-by: Waiman Long <longman@redhat.com>
---
kernel/cgroup/cpuset.c | 52 ++++--------------------------------------
1 file changed, 4 insertions(+), 48 deletions(-)
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 5639c486c967..b07a7e8cef16 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1166,25 +1166,6 @@ enum partition_cmd {
static void update_sibling_cpumasks(struct cpuset *parent, struct cpuset *cs,
struct tmpmasks *tmp);
-/*
- * Update partition exclusive flag
- *
- * Return: 0 if successful, an error code otherwise
- */
-static int update_partition_exclusive_flag(struct cpuset *cs, int new_prs)
-{
- bool exclusive = (new_prs > PRS_MEMBER);
-
- if (exclusive && !is_cpu_exclusive(cs)) {
- if (cpuset_update_flag(CS_CPU_EXCLUSIVE, cs, 1))
- return PERR_NOTEXCL;
- } else if (!exclusive && is_cpu_exclusive(cs)) {
- /* Turning off CS_CPU_EXCLUSIVE will not return error */
- cpuset_update_flag(CS_CPU_EXCLUSIVE, cs, 0);
- }
- return 0;
-}
-
/*
* Update partition load balance flag and/or rebuild sched domain
*
@@ -1240,11 +1221,9 @@ static void reset_partition_data(struct cpuset *cs)
lockdep_assert_held(&callback_lock);
- if (cpumask_empty(cs->exclusive_cpus)) {
+ if (cpumask_empty(cs->exclusive_cpus))
cpumask_clear(cs->effective_xcpus);
- if (is_cpu_exclusive(cs))
- clear_bit(CS_CPU_EXCLUSIVE, &cs->flags);
- }
+
if (!cpumask_and(cs->effective_cpus, parent->effective_cpus, cs->cpus_allowed))
cpumask_copy(cs->effective_cpus, parent->effective_cpus);
}
@@ -2020,19 +1999,6 @@ static int update_parent_effective_cpumask(struct cpuset *cs, int cmd,
if (!adding && !deleting && (new_prs == old_prs))
return 0;
- /*
- * Transitioning between invalid to valid or vice versa may require
- * changing CS_CPU_EXCLUSIVE. In the case of partcmd_update,
- * validate_change() has already been successfully called and
- * CPU lists in cs haven't been updated yet. So defer it to later.
- */
- if ((old_prs != new_prs) && (cmd != partcmd_update)) {
- int err = update_partition_exclusive_flag(cs, new_prs);
-
- if (err)
- return err;
- }
-
/*
* Change the parent's effective_cpus & effective_xcpus (top cpuset
* only).
@@ -2055,9 +2021,6 @@ static int update_parent_effective_cpumask(struct cpuset *cs, int cmd,
spin_unlock_irq(&callback_lock);
- if ((old_prs != new_prs) && (cmd == partcmd_update))
- update_partition_exclusive_flag(cs, new_prs);
-
if (adding || deleting) {
cpuset_update_tasks_cpumask(parent, tmp->addmask);
update_sibling_cpumasks(parent, cs, tmp);
@@ -2934,10 +2897,6 @@ static int update_prstate(struct cpuset *cs, int new_prs)
if (alloc_tmpmasks(&tmpmask))
return -ENOMEM;
- err = update_partition_exclusive_flag(cs, new_prs);
- if (err)
- goto out;
-
if (!old_prs) {
/*
* cpus_allowed and exclusive_cpus cannot be both empty.
@@ -3001,13 +2960,10 @@ static int update_prstate(struct cpuset *cs, int new_prs)
}
out:
/*
- * Make partition invalid & disable CS_CPU_EXCLUSIVE if an error
- * happens.
+ * Make partition invalid if an error happens.
*/
- if (err) {
+ if (err)
new_prs = -new_prs;
- update_partition_exclusive_flag(cs, new_prs);
- }
spin_lock_irq(&callback_lock);
cs->partition_root_state = new_prs;
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-27 16:35 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 16:34 [PATCH-next 0/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only Waiman Long
2026-09-27 16:34 ` [PATCH-next 1/2] " Waiman Long
2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®