* [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails
@ 2026-09-28 23:51 Waiman Long
2026-09-29 2:07 ` Tejun Heo
0 siblings, 1 reply; 7+ messages in thread
From: Waiman Long @ 2026-09-28 23:51 UTC (permalink / raw)
To: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný
Cc: cgroups, linux-kernel, Hui Peng, Guopeng Zhang, Waiman Long
Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full
don't leave any housekeeping"), the partition state can be freely switched
from "root" to "isolated" and vice versa. After that commit, the switch
from "root" to "isolated" can fail if it exhausts all the housekeeping
CPUs. Later on, the switch from "isolated" to "root" can also fail if
some of the partition CPUs are boot-time isolated by "isolcpus".
The partition is made invalid when the switch fails. However, the
remote_partition flag for a remote partition can remain set and the CPUs
from the invalidated partition aren't cleared from subpartitions_cpus.
Fix this by properly disable the partition in this case.
In the case of remote_partition flag, it should be cleared for a
invalidated remote partition. To be safe, the reset_partition_data() is
now enhanced to always clear the remote_partition flag. So there is no
need to explicitly clear remote_partition in remote_partition_disable().
On a x86 test system with boot option "isolcpus=10 cgroup_debug" set
and more than 16 cores, the following commands was executed after boot.
# cd /sys/fs/cgroup
# echo +cpuset > cgroup.subtree_control
# cat cpuset.cpus.isolated
10
# mkdir A
# echo 10-12 > A/cpuset.cpus
# echo isolated > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
isolated
Before the patch:
# echo root > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
root invalid (partition config conflicts with housekeeping setup)
# cat cpuset.cpus.isolated .__DEBUG__.cpuset.cpus.subpartitions
10-12
10-12
After the patch:
# echo root > A/cpuset.cpus.partition
# cat A/cpuset.cpus.partition
root invalid (partition config conflicts with housekeeping setup)
# cat cpuset.cpus.isolated .__DEBUG__.cpuset.cpus.subpartitions
10
Reported-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/lkml/7f4c57b26ad120ab30adf35653f1dd94@kernel.org/
Fixes: 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full don't leave any housekeeping")
Signed-off-by: Waiman Long <longman@redhat.com>
---
kernel/cgroup/cpuset.c | 18 ++++++++++++------
1 file changed, 12 insertions(+), 6 deletions(-)
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 4a915da37bb3..9874c25eb83c 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1229,6 +1229,7 @@ static void reset_partition_data(struct cpuset *cs)
lockdep_assert_held(&callback_lock);
+ cs->remote_partition = false;
if (cpumask_empty(cs->exclusive_cpus))
cpumask_clear(cs->effective_xcpus);
@@ -1615,7 +1616,6 @@ static void remote_partition_disable(struct cpuset *cs, struct tmpmasks *tmp)
!cpumask_empty(subpartitions_cpus));
spin_lock_irq(&callback_lock);
- cs->remote_partition = false;
partition_xcpus_del(cs->partition_root_state, NULL, cs->effective_xcpus);
if (cs->prs_err)
cs->partition_root_state = -cs->partition_root_state;
@@ -2892,6 +2892,7 @@ static int update_prstate(struct cpuset *cs, int new_prs)
struct cpuset *parent = parent_cs(cs);
struct tmpmasks tmpmask;
bool isolcpus_updated = false;
+ bool disable_partition = false;
if (old_prs == new_prs)
return 0;
@@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs)
*/
if (((new_prs == PRS_ISOLATED) &&
!isolated_cpus_can_update(cs->effective_xcpus, NULL)) ||
- prstate_housekeeping_conflict(new_prs, cs->effective_xcpus))
+ prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) {
err = PERR_HKEEPING;
- else
+ disable_partition = true;
+ } else {
isolcpus_updated = true;
+ }
} else {
/*
* Switching back to member is always allowed even if it
* disables child partitions.
*/
+ disable_partition = true;
+ }
+out:
+ if (disable_partition) {
if (is_remote_partition(cs))
remote_partition_disable(cs, &tmpmask);
else
update_parent_effective_cpumask(cs, partcmd_disable,
NULL, &tmpmask);
-
/*
* Invalidation of child partitions will be done in
- * update_cpumasks_hier().
+ * update_cpumasks_hier() below.
*/
}
-out:
+
/*
* Make partition invalid if an error happens.
*/
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-28 23:51 [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails Waiman Long @ 2026-09-29 2:07 ` Tejun Heo 2026-09-29 3:16 ` Waiman Long 0 siblings, 1 reply; 7+ messages in thread From: Tejun Heo @ 2026-09-29 2:07 UTC (permalink / raw) To: Waiman Long Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, Guopeng Zhang, cgroups, linux-kernel Hello, Waiman. The following is a Claude-generated review. On Mon, Sep 28, 2026 at 07:51:59PM -0400, Waiman Long wrote: > Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full > don't leave any housekeeping"), the partition state can be freely switched > from "root" to "isolated" and vice versa. After that commit, the switch > from "root" to "isolated" can fail if it exhausts all the housekeeping > CPUs. Later on, the switch from "isolated" to "root" can also fail if > some of the partition CPUs are boot-time isolated by "isolcpus". The isolated -> root failure came from b1034a690129 ("cgroup/cpuset: Ensure domain isolated CPUs stay in root or isolated partition"). Maybe add a Fixes: tag for it too? > The partition is made invalid when the switch fails. However, the > remote_partition flag for a remote partition can remain set and the CPUs > from the invalidated partition aren't cleared from subpartitions_cpus. > Fix this by properly disable the partition in this case. The reproducer below is a local partition and the fix covers that case too. The CPUs weren't given back to the parent, which is what leaves them in subpartitions_cpus and isolated_cpus in the example. Maybe describe both? Also, "properly disable" -> "disabling". > In the case of remote_partition flag, it should be cleared for a > invalidated remote partition. To be safe, the reset_partition_data() is > now enhanced to always clear the remote_partition flag. So there is no > need to explicitly clear remote_partition in remote_partition_disable(). After the update_prstate() change, every path that invalidates a remote partition goes through remote_partition_disable(), so the other reset_partition_data() callers never see the flag set. If one did, clearing only the flag would leave its CPUs in subpartitions_cpus and turn the WARN_ON_ONCE() in partition_xcpus_del() into a silent leak. Maybe drop this part? > On a x86 test system with boot option "isolcpus=10 cgroup_debug" set > and more than 16 cores, the following commands was executed after boot. "a x86" -> "an x86", "commands was" -> "commands were". > @@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs) > */ > if (((new_prs == PRS_ISOLATED) && > !isolated_cpus_can_update(cs->effective_xcpus, NULL)) || > - prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) > + prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) { > err = PERR_HKEEPING; > - else > + disable_partition = true; If a root -> isolated switch fails isolated_cpus_can_update() under an isolated parent, partcmd_disable hands the CPUs back to the parent and partition_xcpus_del() adds them to isolated_cpus, which is the state the check just rejected. Switching to member ends up in the same place, so this may be fine as is. > + } > +out: > + if (disable_partition) { The early goto out paths never need the disable. Maybe put this block before out: instead? Also, update_cpumasks_hier() below gets force only when switching to member, to update effective_xcpus. Now that the failure path disables the partition too, should it pass disable_partition? Separately, a partition invalidated with PERR_HKEEPING can become valid again through partcmd_update without newmask (hotplug, or update_cpumasks_hier() from an ancestor), which doesn't check housekeeping. A failed member -> root enable has the same problem, so it isn't from this patch. Thanks. -- tejun ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-29 2:07 ` Tejun Heo @ 2026-09-29 3:16 ` Waiman Long 2026-09-29 7:40 ` Guopeng Zhang 0 siblings, 1 reply; 7+ messages in thread From: Waiman Long @ 2026-09-29 3:16 UTC (permalink / raw) To: Tejun Heo Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, Guopeng Zhang, cgroups, linux-kernel On 9/28/26 10:07 PM, Tejun Heo wrote: > Hello, Waiman. > > The following is a Claude-generated review. > > On Mon, Sep 28, 2026 at 07:51:59PM -0400, Waiman Long wrote: >> Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full >> don't leave any housekeeping"), the partition state can be freely switched >> from "root" to "isolated" and vice versa. After that commit, the switch >> from "root" to "isolated" can fail if it exhausts all the housekeeping >> CPUs. Later on, the switch from "isolated" to "root" can also fail if >> some of the partition CPUs are boot-time isolated by "isolcpus". > The isolated -> root failure came from b1034a690129 ("cgroup/cpuset: > Ensure domain isolated CPUs stay in root or isolated partition"). Maybe > add a Fixes: tag for it too? Commit b1034a690129 ("cgroup/cpuset: Ensure domain isolated CPUs stay in root or isolated partition") comes after 103b08709e8a. From my point of view, the first commit is the one that introduces the bug. The second commit comes along assuming the existing code is right. That is why I didn't tag the second one even though the second one is easier to trigger than the first one. >> The partition is made invalid when the switch fails. However, the >> remote_partition flag for a remote partition can remain set and the CPUs >> from the invalidated partition aren't cleared from subpartitions_cpus. >> Fix this by properly disable the partition in this case. > The reproducer below is a local partition and the fix covers that case > too. The CPUs weren't given back to the parent, which is what leaves them > in subpartitions_cpus and isolated_cpus in the example. Maybe describe > both? Also, "properly disable" -> "disabling". OK, will update the commit log. > >> In the case of remote_partition flag, it should be cleared for a >> invalidated remote partition. To be safe, the reset_partition_data() is >> now enhanced to always clear the remote_partition flag. So there is no >> need to explicitly clear remote_partition in remote_partition_disable(). > After the update_prstate() change, every path that invalidates a remote > partition goes through remote_partition_disable(), so the other > reset_partition_data() callers never see the flag set. If one did, > clearing only the flag would leave its CPUs in subpartitions_cpus and turn > the WARN_ON_ONCE() in partition_xcpus_del() into a silent leak. Maybe drop > this part? Yes, I can drop this part. >> On a x86 test system with boot option "isolcpus=10 cgroup_debug" set >> and more than 16 cores, the following commands was executed after boot. > "a x86" -> "an x86", "commands was" -> "commands were". > >> @@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs) >> */ >> if (((new_prs == PRS_ISOLATED) && >> !isolated_cpus_can_update(cs->effective_xcpus, NULL)) || >> - prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) >> + prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) { >> err = PERR_HKEEPING; >> - else >> + disable_partition = true; > If a root -> isolated switch fails isolated_cpus_can_update() under an > isolated parent, partcmd_disable hands the CPUs back to the parent and > partition_xcpus_del() adds them to isolated_cpus, which is the state the > check just rejected. Switching to member ends up in the same place, so > this may be fine as is. If the parent is a valid isolated partition before the child partition is created, it had passed the isolated_cpus_can_update() check. So returning the CPUs back to its parent is fine. > >> + } >> +out: >> + if (disable_partition) { > The early goto out paths never need the disable. Maybe put this block > before out: instead? Yes, that made sense though the current placement should cause any problem. > > Also, update_cpumasks_hier() below gets force only when switching to > member, to update effective_xcpus. Now that the failure path disables the > partition too, should it pass disable_partition? Yes, it should. > > Separately, a partition invalidated with PERR_HKEEPING can become valid > again through partcmd_update without newmask (hotplug, or > update_cpumasks_hier() from an ancestor), which doesn't check > housekeeping. A failed member -> root enable has the same problem, so it > isn't from this patch. I need to take a further look at that. Cheers, Longman > > Thanks. > > -- > tejun ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-29 3:16 ` Waiman Long @ 2026-09-29 7:40 ` Guopeng Zhang 2026-09-29 15:49 ` Waiman Long 2026-09-29 16:42 ` Tejun Heo 0 siblings, 2 replies; 7+ messages in thread From: Guopeng Zhang @ 2026-09-29 7:40 UTC (permalink / raw) To: Waiman Long, Tejun Heo Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, cgroups, linux-kernel Hi Waiman, Tejun, 在 2026/9/29 11:16, Waiman Long 写道: > On 9/28/26 10:07 PM, Tejun Heo wrote: >> Hello, Waiman. >> >> The following is a Claude-generated review. >> >> On Mon, Sep 28, 2026 at 07:51:59PM -0400, Waiman Long wrote: >>> Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full >>> don't leave any housekeeping"), the partition state can be freely switched >>> from "root" to "isolated" and vice versa. After that commit, the switch >>> from "root" to "isolated" can fail if it exhausts all the housekeeping >>> CPUs. Later on, the switch from "isolated" to "root" can also fail if >>> some of the partition CPUs are boot-time isolated by "isolcpus". >> The isolated -> root failure came from b1034a690129 ("cgroup/cpuset: >> Ensure domain isolated CPUs stay in root or isolated partition"). Maybe >> add a Fixes: tag for it too? > > Commit b1034a690129 ("cgroup/cpuset: Ensure domain isolated CPUs stay in root or isolated partition") comes after 103b08709e8a. From my point of view, the first commit is the one that introduces the bug. The second commit comes along assuming the existing code is right. That is why I didn't tag the second one even though the second one is easier to trigger than the first one. > >>> The partition is made invalid when the switch fails. However, the >>> remote_partition flag for a remote partition can remain set and the CPUs >>> from the invalidated partition aren't cleared from subpartitions_cpus. >>> Fix this by properly disable the partition in this case. >> The reproducer below is a local partition and the fix covers that case >> too. The CPUs weren't given back to the parent, which is what leaves them >> in subpartitions_cpus and isolated_cpus in the example. Maybe describe >> both? Also, "properly disable" -> "disabling". > OK, will update the commit log. >> >>> In the case of remote_partition flag, it should be cleared for a >>> invalidated remote partition. To be safe, the reset_partition_data() is >>> now enhanced to always clear the remote_partition flag. So there is no >>> need to explicitly clear remote_partition in remote_partition_disable(). >> After the update_prstate() change, every path that invalidates a remote >> partition goes through remote_partition_disable(), so the other >> reset_partition_data() callers never see the flag set. If one did, >> clearing only the flag would leave its CPUs in subpartitions_cpus and turn >> the WARN_ON_ONCE() in partition_xcpus_del() into a silent leak. Maybe drop >> this part? > Yes, I can drop this part. >>> On a x86 test system with boot option "isolcpus=10 cgroup_debug" set >>> and more than 16 cores, the following commands was executed after boot. >> "a x86" -> "an x86", "commands was" -> "commands were". >> >>> @@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs) >>> */ >>> if (((new_prs == PRS_ISOLATED) && >>> !isolated_cpus_can_update(cs->effective_xcpus, NULL)) || >>> - prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) >>> + prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) { >>> err = PERR_HKEEPING; >>> - else >>> + disable_partition = true; >> If a root -> isolated switch fails isolated_cpus_can_update() under an >> isolated parent, partcmd_disable hands the CPUs back to the parent and >> partition_xcpus_del() adds them to isolated_cpus, which is the state the >> check just rejected. Switching to member ends up in the same place, so >> this may be fine as is. > If the parent is a valid isolated partition before the child partition is created, it had passed the isolated_cpus_can_update() check. So returning the CPUs back to its parent is fine. However, after the parent is created, the ownership of full housekeeping CPUs can still change, so passing the check at creation time doesn't seem to guarantee that returning CPUs later is still safe. I tested with this patch applied on a 32 vCPU VM booted with: nohz_full=1,3 maxcpus=4 CPUs 0-3 are online, and CPU0 and CPU2 are initially available as full housekeeping CPUs: cd /sys/fs/cgroup echo +cpuset > cgroup.subtree_control mkdir A C echo +cpuset > A/cgroup.subtree_control mkdir A/B echo 0-1 > A/cpuset.cpus echo isolated > A/cpuset.cpus.partition # isolated={0,1}, full HK={2} echo 0 > A/B/cpuset.cpus echo root > A/B/cpuset.cpus.partition # isolated={1}, full HK={0,2} echo 2 > C/cpuset.cpus echo isolated > C/cpuset.cpus.partition # isolated={1,2}, full HK={0} At this point, CPU0 is the last CPU that is neither domain-isolated nor in nohz_full. Then: echo isolated > A/B/cpuset.cpus.partition cat A/B/cpuset.cpus.partition isolated invalid (partition config conflicts with housekeeping setup) cat cpuset.cpus.isolated 0-2 The root -> isolated switch of A/B is rejected as expected, but the new disable path then hands CPU0 back to the isolated parent A, and partition_xcpus_del() puts it right back into isolated_cpus. In the end the domain-isolated CPUs are {0,1,2} and the nohz_full CPUs are {1,3}, so no full housekeeping CPU is left. Also, as Tejun pointed out, switching A/B to member leads to the same problem. What I tried in patches 2-3 of [1] was: if returning CPUs to an isolated parent would consume the last full housekeeping CPU, invalidate the outermost isolated ancestor and return its CPUs instead. Maybe this part can serve as a reference. The implementation there may not be good enough - it was an earlier idea and I am not sure whether it helps :) [1] https://lore.kernel.org/all/20260910094546.5852-1-guopeng.zhang@linux.dev/ Thanks, Guopeng ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-29 7:40 ` Guopeng Zhang @ 2026-09-29 15:49 ` Waiman Long 2026-09-29 16:42 ` Tejun Heo 1 sibling, 0 replies; 7+ messages in thread From: Waiman Long @ 2026-09-29 15:49 UTC (permalink / raw) To: Guopeng Zhang, Tejun Heo Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, cgroups, linux-kernel On 9/29/26 3:40 AM, Guopeng Zhang wrote: > Hi Waiman, Tejun, > 在 2026/9/29 11:16, Waiman Long 写道: >> On 9/28/26 10:07 PM, Tejun Heo wrote: >>> Hello, Waiman. >>> >>> The following is a Claude-generated review. >>> >>> On Mon, Sep 28, 2026 at 07:51:59PM -0400, Waiman Long wrote: >>>> Before commit 103b08709e8a ("cgroup/cpuset: Fail if isolated and nohz_full >>>> don't leave any housekeeping"), the partition state can be freely switched >>>> from "root" to "isolated" and vice versa. After that commit, the switch >>>> from "root" to "isolated" can fail if it exhausts all the housekeeping >>>> CPUs. Later on, the switch from "isolated" to "root" can also fail if >>>> some of the partition CPUs are boot-time isolated by "isolcpus". >>> The isolated -> root failure came from b1034a690129 ("cgroup/cpuset: >>> Ensure domain isolated CPUs stay in root or isolated partition"). Maybe >>> add a Fixes: tag for it too? >> Commit b1034a690129 ("cgroup/cpuset: Ensure domain isolated CPUs stay in root or isolated partition") comes after 103b08709e8a. From my point of view, the first commit is the one that introduces the bug. The second commit comes along assuming the existing code is right. That is why I didn't tag the second one even though the second one is easier to trigger than the first one. >> >>>> The partition is made invalid when the switch fails. However, the >>>> remote_partition flag for a remote partition can remain set and the CPUs >>>> from the invalidated partition aren't cleared from subpartitions_cpus. >>>> Fix this by properly disable the partition in this case. >>> The reproducer below is a local partition and the fix covers that case >>> too. The CPUs weren't given back to the parent, which is what leaves them >>> in subpartitions_cpus and isolated_cpus in the example. Maybe describe >>> both? Also, "properly disable" -> "disabling". >> OK, will update the commit log. >>>> In the case of remote_partition flag, it should be cleared for a >>>> invalidated remote partition. To be safe, the reset_partition_data() is >>>> now enhanced to always clear the remote_partition flag. So there is no >>>> need to explicitly clear remote_partition in remote_partition_disable(). >>> After the update_prstate() change, every path that invalidates a remote >>> partition goes through remote_partition_disable(), so the other >>> reset_partition_data() callers never see the flag set. If one did, >>> clearing only the flag would leave its CPUs in subpartitions_cpus and turn >>> the WARN_ON_ONCE() in partition_xcpus_del() into a silent leak. Maybe drop >>> this part? >> Yes, I can drop this part. >>>> On a x86 test system with boot option "isolcpus=10 cgroup_debug" set >>>> and more than 16 cores, the following commands was executed after boot. >>> "a x86" -> "an x86", "commands was" -> "commands were". >>> >>>> @@ -2946,27 +2947,32 @@ static int update_prstate(struct cpuset *cs, int new_prs) >>>> */ >>>> if (((new_prs == PRS_ISOLATED) && >>>> !isolated_cpus_can_update(cs->effective_xcpus, NULL)) || >>>> - prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) >>>> + prstate_housekeeping_conflict(new_prs, cs->effective_xcpus)) { >>>> err = PERR_HKEEPING; >>>> - else >>>> + disable_partition = true; >>> If a root -> isolated switch fails isolated_cpus_can_update() under an >>> isolated parent, partcmd_disable hands the CPUs back to the parent and >>> partition_xcpus_del() adds them to isolated_cpus, which is the state the >>> check just rejected. Switching to member ends up in the same place, so >>> this may be fine as is. >> If the parent is a valid isolated partition before the child partition is created, it had passed the isolated_cpus_can_update() check. So returning the CPUs back to its parent is fine. > However, after the parent is created, the ownership of full housekeeping > CPUs can still change, so passing the check at creation time doesn't seem > to guarantee that returning CPUs later is still safe. With additional thought after my email, I realized that my statement isn't quite correct. So invalidating the partition handing the CPUs back to its rightful parent may not be right. A simpler way to handle this is to keep the status quo and return an error to the users instead. Cheers, Longman > I tested with this patch applied on a 32 vCPU VM booted with: > > nohz_full=1,3 maxcpus=4 > > CPUs 0-3 are online, and CPU0 and CPU2 are initially available as full > housekeeping CPUs: > > cd /sys/fs/cgroup > echo +cpuset > cgroup.subtree_control > > mkdir A C > echo +cpuset > A/cgroup.subtree_control > mkdir A/B > > echo 0-1 > A/cpuset.cpus > echo isolated > A/cpuset.cpus.partition # isolated={0,1}, full HK={2} > > echo 0 > A/B/cpuset.cpus > echo root > A/B/cpuset.cpus.partition # isolated={1}, full HK={0,2} > > echo 2 > C/cpuset.cpus > echo isolated > C/cpuset.cpus.partition # isolated={1,2}, full HK={0} > > At this point, CPU0 is the last CPU that is neither domain-isolated nor in > nohz_full. Then: > > echo isolated > A/B/cpuset.cpus.partition > > cat A/B/cpuset.cpus.partition > isolated invalid (partition config conflicts with housekeeping setup) > > cat cpuset.cpus.isolated > 0-2 > > The root -> isolated switch of A/B is rejected as expected, but the new > disable path then hands CPU0 back to the isolated parent A, and > partition_xcpus_del() puts it right back into isolated_cpus. > > In the end the domain-isolated CPUs are {0,1,2} and the nohz_full CPUs are > {1,3}, so no full housekeeping CPU is left. > > Also, as Tejun pointed out, switching A/B to member leads to the same > problem. What I tried in patches 2-3 of [1] was: if returning CPUs to an > isolated parent would consume the last full housekeeping CPU, invalidate > the outermost isolated ancestor and return its CPUs instead. Maybe this > part can serve as a reference. The implementation there may not be good > enough - it was an earlier idea and I am not sure whether it helps :) > > [1] https://lore.kernel.org/all/20260910094546.5852-1-guopeng.zhang@linux.dev/ > > Thanks, > Guopeng ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-29 7:40 ` Guopeng Zhang 2026-09-29 15:49 ` Waiman Long @ 2026-09-29 16:42 ` Tejun Heo 2026-09-29 18:21 ` Waiman Long 1 sibling, 1 reply; 7+ messages in thread From: Tejun Heo @ 2026-09-29 16:42 UTC (permalink / raw) To: Guopeng Zhang, Waiman Long Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, cgroups, linux-kernel Hello, On Tue, Sep 29, 2026 at 03:40:27PM +0800, Guopeng Zhang wrote: > However, after the parent is created, the ownership of full housekeeping > CPUs can still change, so passing the check at creation time doesn't seem > to guarantee that returning CPUs later is still safe. I think any CPU under an isolated partition should count as isolated when testing whether the system has a housekeeping CPU left, whether child partitions carved some out or not. Those CPUs go back to the isolated parent whenever the child partition is disabled or invalidated, so they can't be counted on for housekeeping. In your example, making C isolated would then fail. Thanks. -- tejun ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails 2026-09-29 16:42 ` Tejun Heo @ 2026-09-29 18:21 ` Waiman Long 0 siblings, 0 replies; 7+ messages in thread From: Waiman Long @ 2026-09-29 18:21 UTC (permalink / raw) To: Tejun Heo, Guopeng Zhang Cc: Ridong Chen, Johannes Weiner, Michal Koutný, Hui Peng, cgroups, linux-kernel On 9/29/26 12:42 PM, Tejun Heo wrote: > Hello, > > On Tue, Sep 29, 2026 at 03:40:27PM +0800, Guopeng Zhang wrote: >> However, after the parent is created, the ownership of full housekeeping >> CPUs can still change, so passing the check at creation time doesn't seem >> to guarantee that returning CPUs later is still safe. > I think any CPU under an isolated partition should count as isolated when > testing whether the system has a housekeeping CPU left, whether child > partitions carved some out or not. Those CPUs go back to the isolated > parent whenever the child partition is disabled or invalidated, so they > can't be counted on for housekeeping. In your example, making C isolated > would then fail. I am now thinking about treating the whole user defined exclusive CPUs (exclusive_cpus or cpus) as wholly owned by the cpuset in question from the housekeeping check perspective. That will require reworking the current housekeeping checking code and will require more extensive changes. So it will take some time before a patchset will be ready for review. This will prevent housekeeping check failure when a child partition is disabled and its CPU are moved back to its parent. Cheers, Longman ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-29 18:21 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-28 23:51 [PATCH] cgroup/cpuset: Properly disable partition when partition state switching fails Waiman Long 2026-09-29 2:07 ` Tejun Heo 2026-09-29 3:16 ` Waiman Long 2026-09-29 7:40 ` Guopeng Zhang 2026-09-29 15:49 ` Waiman Long 2026-09-29 16:42 ` Tejun Heo 2026-09-29 18:21 ` Waiman Long
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®