From: Waiman Long <longman@redhat.com>
To: "Ridong Chen" <ridong.chen@linux.dev>,
"Tejun Heo" <tj@kernel.org>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Michal Koutný" <mkoutny@suse.com>
Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
Hui Peng <benquike@gmail.com>,
Guopeng Zhang <guopeng.zhang@linux.dev>,
Waiman Long <longman@redhat.com>
Subject: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails
Date: Mon, 28 Sep 2026 19:51:59 -0400 [thread overview]
Message-ID: <20260928235159.515654-1-longman@redhat.com> (raw)
Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full
don't leave any housekeeping"), the partition state can be freely switched
from "root" to "isolated" and vice versa. After that commit, the switch
from "root" to "isolated" can fail if it exhausts all the housekeeping
CPUs. Later on, the switch from "isolated" to "root" can also fail if
some of the partition CPUs are boot-time isolated by "isolcpus".
The partition is made invalid when the switch fails. However, the
remote_partition flag for a remote partition can remain set and the CPUs
from the invalidated partition aren't cleared from subpartitions_cpus.
Fix this by properly disable the partition in this case.
In the case of remote_partition flag, it should be cleared for a
invalidated remote partition. To be safe, the reset_partition_data() is
now enhanced to always clear the remote_partition flag. So there is no
need to explicitly clear remote_partition in remote_partition_disable().
On a x86 test system with boot option "isolcpus=10 cgroup_debug" set
and more than 16 cores, the following commands was executed after boot.
# cd /sys/fs/cgroup
# echo +cpuset > cgroup.subtree_control
# cat cpuset.cpus.isolated
10
# mkdir A
# echo 10-12 > A/cpuset.cpus
# echo isolated > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
isolated
Before the patch:
# echo root > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
root invalid (partition config conflicts with housekeeping setup)
# cat cpuset.cpus.isolated .__DEBUG__.cpuset.cpus.subpartitions
10-12
10-12
After the patch:
# echo root > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
root invalid (partition config conflicts with housekeeping setup)
# cat cpuset.cpus.isolated .__DEBUG__.cpuset.cpus.subpartitions
10
Reported-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/lkml/7f4c57b26ad120ab30adf35653f1dd94@kernel.org/
Fixes: 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full don't leave any housekeeping")
Signed-off-by: Waiman Long <longman@redhat.com>
---
kernel/cgroup/cpuset.c | 18 ++++++++++++------
1 file changed, 12 insertions(+), 6 deletions(-)
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 4a915da37bb3..9874c25eb83c 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1229,6 +1229,7 @@ static void reset_partition_data(struct cpuset *cs)
lockdep_assert_held(&callback_lock);
+ cs->remote_partition = false;
if (cpumask_empty(cs->exclusive_cpus))
cpumask_clear(cs->effective_xcpus);
@@ -1615,7 +1616,6 @@ static void remote_partition_disable(struct cpuset *cs, struct tmpmasks *tmp)
!cpumask_empty(subpartitions_cpus));
spin_lock_irq(&callback_lock);
- cs->remote_partition = false;
partition_xcpus_del(cs->partition_root_state, NULL, cs->effective_xcpus);
if (cs->prs_err)
cs->partition_root_state = -cs->partition_root_state;
@@ -2892,6 +2892,7 @@ static int update_prstate(struct cpuset *cs, int new_prs)
struct cpuset *parent = parent_cs(cs);
struct tmpmasks tmpmask;
bool isolcpus_updated = false;
+ bool disable_partition = false;
if (old_prs == new_prs)
return 0;
@@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs)
*/
if (((new_prs == PRS_ISOLATED) &&
!isolated_cpus_can_update(cs->effective_xcpus, NULL)) ||
- prstate_housekeeping_conflict(new_prs, cs->effective_xcpus))
+ prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) {
err = PERR_HKEEPING;
- else
+ disable_partition = true;
+ } else {
isolcpus_updated = true;
+ }
} else {
/*
* Switching back to member is always allowed even if it
* disables child partitions.
*/
+ disable_partition = true;
+ }
+out:
+ if (disable_partition) {
if (is_remote_partition(cs))
remote_partition_disable(cs, &tmpmask);
else
update_parent_effective_cpumask(cs, partcmd_disable,
NULL, &tmpmask);
-
/*
* Invalidation of child partitions will be done in
- * update_cpumasks_hier().
+ * update_cpumasks_hier() below.
*/
}
-out:
+
/*
* Make partition invalid if an error happens.
*/
--
2.55.0
next reply other threads:[~2026-09-28 23:52 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 23:51 Waiman Long [this message]
2026-09-29 2:07 ` Tejun Heo
2026-09-29 3:16 ` Waiman Long
2026-09-29 7:40 ` Guopeng Zhang
2026-09-29 15:49 ` Waiman Long
2026-09-29 16:42 ` Tejun Heo
2026-09-29 18:21 ` Waiman Long
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928235159.515654-1-longman@redhat.com \
--to=longman@redhat.com \
--cc=benquike@gmail.com \
--cc=cgroups@vger.kernel.org \
--cc=guopeng.zhang@linux.dev \
--cc=hannes@cmpxchg.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mkoutny@suse.com \
--cc=ridong.chen@linux.dev \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®