mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Waiman Long <longman@redhat.com>
To: "Ridong Chen" <ridong.chen@linux.dev>,
	"Tejun Heo" <tj@kernel.org>,
	"Johannes Weiner" <hannes@cmpxchg.org>,
	"Michal Koutný" <mkoutny@suse.com>
Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
	Hui Peng <benquike@gmail.com>,
	Guopeng Zhang <guopeng.zhang@linux.dev>,
	Waiman Long <longman@redhat.com>
Subject: [PATCH-next 1/2] cgroup/cpuset: Run SCHED_DEADLINE shrink test on valid partition root only
Date: Sun, 27 Sep 2026 12:34:53 -0400	[thread overview]
Message-ID: <20260927163454.345463-2-longman@redhat.com> (raw)
In-Reply-To: <20260927163454.345463-1-longman@redhat.com>

Commit f82f80426f7a ("sched/deadline: Ensure that updates to exclusive
cpusets don't break AC") adds a check in validate_change() to make
sure that there is enough bandwidth for SCHED_DEADLINE tasks if we
shrink a v1 exclusive cpuset that has CS_CPU_EXCLUSIVE flag set.

With the introduction of cpuset partition in cgroup v2, we keep setting
the CS_CPU_EXCLUSIVE flag for a partition root so that the SCHED_DEADLINE
check will continue to work as intended. However it turns out that the
current code isn't perfect and there are cases where a cpuset isn't
a valid partition root, but the exclusive flag is still incorrectly
set. This can leads to SCHED_DEADLINE check being incorrectly triggered
when there are deadline tasks in the system. This can result in unexpected
-EBUSY failure when making changes to cpuset control files. Fix that
by checking for a valid partition root in the case of v2 and break out
the v1 specific check back into cpuset1_validate_change(). It is far
easier and less cumbersome than to make sure that the exclusive flag
is only set for valid partition roots.

The is_in_v2_mode() check guarding the call to cpuset1_validate_change()
is also moved to the appropriate place inside cpuset1_validate_change()
and a cpuset_v2() is now being used as a guard which should be more
accurate for the a v1 system with v2 mode enabled.

Even though commit a86ce68078b2 ("cgroup/cpuset: Extract out
CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling") is marked as a
commit to be fixed, the problem may exist before that.

Fixes: a86ce68078b2 ("cgroup/cpuset: Extract out CS_CPU_EXCLUSIVE & CS_SCHED_LOAD_BALANCE handling")
Signed-off-by: Waiman Long <longman@redhat.com>
---
 kernel/cgroup/cpuset-v1.c | 20 ++++++++++++++++++--
 kernel/cgroup/cpuset.c    | 17 +++++------------
 2 files changed, 23 insertions(+), 14 deletions(-)

diff --git a/kernel/cgroup/cpuset-v1.c b/kernel/cgroup/cpuset-v1.c
index 562ad35f00d0..c03ae8aac03a 100644
--- a/kernel/cgroup/cpuset-v1.c
+++ b/kernel/cgroup/cpuset-v1.c
@@ -61,6 +61,11 @@ struct cpuset_remove_tasks_struct {
 #define FM_MAXCNT 1000000	/* limit cnt to avoid overflow */
 #define FM_SCALE 1000		/* faux fixed point scale */
 
+static inline bool is_in_v2_mode(void)
+{
+	return cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE;
+}
+
 /* Initialize a frequency meter */
 static void fmeter_init(struct fmeter *fmp)
 {
@@ -357,10 +362,21 @@ int cpuset1_validate_change(struct cpuset *cur, struct cpuset *trial)
 		if (!is_cpuset_subset(c, trial))
 			goto out;
 
-	/* On legacy hierarchy, we must be a subset of our parent cpuset. */
+	/*
+	 * We can't shrink if we won't have enough room for SCHED_DEADLINE
+	 * tasks in a scheduling partition.
+	 */
+	if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+	    !cpuset_cpumask_can_shrink(cur->cpus_allowed, trial->cpus_allowed))
+		goto out;
+
+	/*
+	 * On legacy hierarchy with v2 mode off, we must be a subset of our
+	 * parent cpuset.
+	 */
 	ret = -EACCES;
 	par = parent_cs(cur);
-	if (par && !is_cpuset_subset(trial, par))
+	if (par && !is_in_v2_mode() && !is_cpuset_subset(trial, par))
 		goto out;
 
 	/*
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 753aa65afcd7..5639c486c967 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -752,7 +752,7 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
 
 	rcu_read_lock();
 
-	if (!is_in_v2_mode())
+	if (!cpuset_v2())
 		ret = cpuset1_validate_change(cur, trial);
 	if (ret)
 		goto out;
@@ -765,23 +765,16 @@ static int validate_change(struct cpuset *cur, struct cpuset *trial)
 
 	/*
 	 * We can't shrink if we won't have enough room for SCHED_DEADLINE
-	 * tasks. This check is not done when scheduling is disabled as the
-	 * users should know what they are doing.
-	 *
-	 * For v1, effective_cpus == cpus_allowed & user_xcpus() returns
-	 * cpus_allowed.
-	 *
-	 * For v2, is_cpu_exclusive() & is_sched_load_balance() are true only
-	 * for non-isolated partition root. At this point, the target
-	 * effective_cpus isn't computed yet. user_xcpus() is the best
-	 * approximation.
+	 * tasks. This check is only done on non-isolated partition root.
+	 * At this point, the target effective_cpus isn't computed yet.
+	 * user_xcpus() is the best approximation.
 	 *
 	 * TBD: May need to precompute the real effective_cpus here in case
 	 * incorrect scheduling of SCHED_DEADLINE tasks in a partition
 	 * becomes an issue.
 	 */
 	ret = -EBUSY;
-	if (is_cpu_exclusive(cur) && is_sched_load_balance(cur) &&
+	if (is_partition_valid(cur) && is_sched_load_balance(cur) &&
 	    !cpuset_cpumask_can_shrink(cur->effective_cpus, user_xcpus(trial)))
 		goto out;
 
-- 
2.55.0


  reply	other threads:[~2026-09-27 16:35 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-27 16:34 [PATCH-next 0/2] " Waiman Long
2026-09-27 16:34 ` Waiman Long [this message]
2026-09-27 16:34 ` [PATCH-next 2/2] cgroup/cpuset: Remove CS_CPU_EXCLUSIVE handling code from v2 Waiman Long

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260927163454.345463-2-longman@redhat.com \
    --to=longman@redhat.com \
    --cc=benquike@gmail.com \
    --cc=cgroups@vger.kernel.org \
    --cc=guopeng.zhang@linux.dev \
    --cc=hannes@cmpxchg.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mkoutny@suse.com \
    --cc=ridong.chen@linux.dev \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®