mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
@ 2026-09-27  9:57 Guopeng Zhang
  2026-09-28  0:00 ` Waiman Long
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Guopeng Zhang @ 2026-09-27  9:57 UTC (permalink / raw)
  To: Waiman Long, Ridong Chen
  Cc: Tejun Heo, Johannes Weiner, Michal Koutný,
	cgroups, linux-kernel, Guopeng Zhang

From: Guopeng Zhang <zhangguopeng@kylinos.cn>

Widening an ancestor's exclusive CPU mask can add a boot-isolated CPU
to a valid remote partition root without touching the partition's own
control files. The partition then load balances that CPU, silently
defeating isolcpus=domain for it.

This can be reproduced on a 32-CPU system booted with
isolcpus=domain,4:

    cd /sys/fs/cgroup
    echo +cpuset > cgroup.subtree_control
    mkdir -p A/B
    echo +cpuset > A/cgroup.subtree_control
    echo 2-4 > A/cpuset.cpus
    echo 2-3 > A/cpuset.cpus.exclusive
    echo 2-4 > A/B/cpuset.cpus
    echo 2-4 > A/B/cpuset.cpus.exclusive
    echo root > A/B/cpuset.cpus.partition
    cat A/B/cpuset.cpus.effective               # 2-3
    echo 2-4 > A/cpuset.cpus.exclusive
    cat A/B/cpuset.cpus.partition               # root
    cat A/B/cpuset.cpus.effective               # 2-4

The last write returns 0 and leaves the hierarchy in this state:

    root (cpuset.cpus.effective=0-1,5-31)
    |
    \-- A (member):           cpuset.cpus=2-4
        |                     cpuset.cpus.exclusive=2-4
        \-- B (root, remote): cpuset.cpus=2-4
                              cpuset.cpus.effective=2-4

B is a remote partition: it takes its CPUs directly from the root
cpuset, and A only passes its exclusive list down. Before the last
write, that list is 2-3, so B holds 2-3 and CPU 4 stays in the
root cpuset as a boot-isolated CPU. The write widens A's exclusive
list to 2-4, which additionally grants CPU 4 to B. Nothing rejects
the grant: B remains a valid root partition, and CPU 4 is still
listed in cpuset.cpus.isolated while sitting in a load-balanced
partition.

remote_partition_enable() already rejects such grants through
prstate_housekeeping_conflict(). remote_cpus_update(), which applies
ancestor changes to a remote partition, does not.

Both paths also open-code remote partition validation. Move those
checks into validate_remote_partition(), and check the resulting
effective exclusive CPU mask for a housekeeping conflict there. The
existing prs_err path then invalidates the remote partition instead of
assigning it a boot-isolated CPU.

Fixes: f62a5d39368e ("cgroup/cpuset: Remove remote_partition_check() & make update_cpumasks_hier() handle remote partition")
Suggested-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
---
Changes since v1:
- Consolidate remote partition validation in a common helper, as
  suggested by Ridong.
- Check housekeeping conflicts against the resulting effective exclusive
  CPU mask.
- Rebase onto cgroup/for-7.3-fixes.

 kernel/cgroup/cpuset.c | 71 +++++++++++++++++++++++++++++-------------
 1 file changed, 49 insertions(+), 22 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 1fcec89a28b9..5adf47217e59 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -1561,6 +1561,45 @@ static inline bool is_local_partition(struct cpuset *cs)
 	return is_partition_valid(cs) && !is_remote_partition(cs);
 }
 
+/**
+ * validate_remote_partition - Validate a remote partition CPU change
+ * @cs: cpuset being enabled or updated
+ * @prs: partition root state to validate
+ * @excpus: resulting effective exclusive CPU mask
+ * @addcpus: exclusive CPUs to be added
+ * @delcpus: exclusive CPUs to be deleted, can be NULL
+ *
+ * Return: PERR_NONE if valid, otherwise an appropriate error code
+ */
+static enum prs_errcode
+validate_remote_partition(struct cpuset *cs, int prs,
+			  struct cpumask *excpus,
+			  struct cpumask *addcpus,
+			  struct cpumask *delcpus)
+{
+	bool updating = is_remote_partition(cs);
+
+	if (!capable(CAP_SYS_ADMIN))
+		return PERR_ACCESS;
+
+	if (!updating &&
+	    (!cpumask_intersects(excpus, cpu_active_mask) ||
+	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
+		return PERR_INVCPUS;
+
+	if (cpumask_intersects(addcpus, subpartitions_cpus) ||
+	    (updating &&
+	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
+		return PERR_NOCPUS;
+
+	if ((prs == PRS_ISOLATED &&
+	     !isolated_cpus_can_update(addcpus, delcpus)) ||
+	    prstate_housekeeping_conflict(prs, excpus))
+		return PERR_HKEEPING;
+
+	return PERR_NONE;
+}
+
 /*
  * remote_partition_enable - Enable current cpuset as a remote partition root
  * @cs: the cpuset to update
@@ -1574,11 +1613,7 @@ static inline bool is_local_partition(struct cpuset *cs)
 static int remote_partition_enable(struct cpuset *cs, int new_prs,
 				   struct tmpmasks *tmp)
 {
-	/*
-	 * The user must have sysadmin privilege.
-	 */
-	if (!capable(CAP_SYS_ADMIN))
-		return PERR_ACCESS;
+	enum prs_errcode err;
 
 	/*
 	 * The requested exclusive_cpus must not be allocated to other
@@ -1591,15 +1626,10 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs,
 	 * above it or remote partition root underneath it is not allowed.
 	 */
 	compute_excpus(cs, tmp->new_cpus);
-	if (!cpumask_intersects(tmp->new_cpus, cpu_active_mask) ||
-	    cpumask_subset(top_cpuset.effective_cpus, tmp->new_cpus))
-		return PERR_INVCPUS;
-	if (cpumask_intersects(tmp->new_cpus, subpartitions_cpus))
-		return PERR_NOCPUS;
-	if (((new_prs == PRS_ISOLATED) &&
-	     !isolated_cpus_can_update(tmp->new_cpus, NULL)) ||
-	    prstate_housekeeping_conflict(new_prs, tmp->new_cpus))
-		return PERR_HKEEPING;
+	err = validate_remote_partition(cs, new_prs,
+					tmp->new_cpus, tmp->new_cpus, NULL);
+	if (err)
+		return err;
 
 	spin_lock_irq(&callback_lock);
 	partition_xcpus_add(new_prs, NULL, tmp->new_cpus);
@@ -1672,6 +1702,7 @@ static void remote_partition_disable(struct cpuset *cs, struct tmpmasks *tmp)
 static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
 			       struct cpumask *excpus, struct tmpmasks *tmp)
 {
+	enum prs_errcode err;
 	bool adding, deleting;
 	int prs = cs->partition_root_state;
 
@@ -1695,14 +1726,10 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
 	 */
 	if (adding) {
 		WARN_ON_ONCE(cpumask_intersects(tmp->addmask, subpartitions_cpus));
-		if (!capable(CAP_SYS_ADMIN))
-			WRITE_ONCE(cs->prs_err, PERR_ACCESS);
-		else if (cpumask_intersects(tmp->addmask, subpartitions_cpus) ||
-			 cpumask_subset(top_cpuset.effective_cpus, tmp->addmask))
-			WRITE_ONCE(cs->prs_err, PERR_NOCPUS);
-		else if ((prs == PRS_ISOLATED) &&
-			 !isolated_cpus_can_update(tmp->addmask, tmp->delmask))
-			WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
+		err = validate_remote_partition(cs, prs, excpus,
+						tmp->addmask, tmp->delmask);
+		if (err)
+			WRITE_ONCE(cs->prs_err, err);
 		if (cs->prs_err)
 			goto invalidate;
 	}

base-commit: 31c88350b7dd1522792f726f79607f31bb55c50f
-- 
2.43.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
  2026-09-27  9:57 [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict Guopeng Zhang
@ 2026-09-28  0:00 ` Waiman Long
  2026-09-28  1:29 ` Ridong Chen
  2026-09-28 18:44 ` Tejun Heo
  2 siblings, 0 replies; 6+ messages in thread
From: Waiman Long @ 2026-09-28  0:00 UTC (permalink / raw)
  To: Guopeng Zhang, Ridong Chen
  Cc: Tejun Heo, Johannes Weiner, Michal Koutný,
	cgroups, linux-kernel, Guopeng Zhang

On 9/27/26 5:57 AM, Guopeng Zhang wrote:
> From: Guopeng Zhang <zhangguopeng@kylinos.cn>
>
> Widening an ancestor's exclusive CPU mask can add a boot-isolated CPU
> to a valid remote partition root without touching the partition's own
> control files. The partition then load balances that CPU, silently
> defeating isolcpus=domain for it.
>
> This can be reproduced on a 32-CPU system booted with
> isolcpus=domain,4:
>
>      cd /sys/fs/cgroup
>      echo +cpuset > cgroup.subtree_control
>      mkdir -p A/B
>      echo +cpuset > A/cgroup.subtree_control
>      echo 2-4 > A/cpuset.cpus
>      echo 2-3 > A/cpuset.cpus.exclusive
>      echo 2-4 > A/B/cpuset.cpus
>      echo 2-4 > A/B/cpuset.cpus.exclusive
>      echo root > A/B/cpuset.cpus.partition
>      cat A/B/cpuset.cpus.effective               # 2-3
>      echo 2-4 > A/cpuset.cpus.exclusive
>      cat A/B/cpuset.cpus.partition               # root
>      cat A/B/cpuset.cpus.effective               # 2-4
>
> The last write returns 0 and leaves the hierarchy in this state:
>
>      root (cpuset.cpus.effective=0-1,5-31)
>      |
>      \-- A (member):           cpuset.cpus=2-4
>          |                     cpuset.cpus.exclusive=2-4
>          \-- B (root, remote): cpuset.cpus=2-4
>                                cpuset.cpus.effective=2-4
>
> B is a remote partition: it takes its CPUs directly from the root
> cpuset, and A only passes its exclusive list down. Before the last
> write, that list is 2-3, so B holds 2-3 and CPU 4 stays in the
> root cpuset as a boot-isolated CPU. The write widens A's exclusive
> list to 2-4, which additionally grants CPU 4 to B. Nothing rejects
> the grant: B remains a valid root partition, and CPU 4 is still
> listed in cpuset.cpus.isolated while sitting in a load-balanced
> partition.
>
> remote_partition_enable() already rejects such grants through
> prstate_housekeeping_conflict(). remote_cpus_update(), which applies
> ancestor changes to a remote partition, does not.
>
> Both paths also open-code remote partition validation. Move those
> checks into validate_remote_partition(), and check the resulting
> effective exclusive CPU mask for a housekeeping conflict there. The
> existing prs_err path then invalidates the remote partition instead of
> assigning it a boot-isolated CPU.
>
> Fixes: f62a5d39368e ("cgroup/cpuset: Remove remote_partition_check() & make update_cpumasks_hier() handle remote partition")
> Suggested-by: Ridong Chen <ridong.chen@linux.dev>
> Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
> ---
> Changes since v1:
> - Consolidate remote partition validation in a common helper, as
>    suggested by Ridong.
> - Check housekeeping conflicts against the resulting effective exclusive
>    CPU mask.
> - Rebase onto cgroup/for-7.3-fixes.
>
>   kernel/cgroup/cpuset.c | 71 +++++++++++++++++++++++++++++-------------
>   1 file changed, 49 insertions(+), 22 deletions(-)
>
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index 1fcec89a28b9..5adf47217e59 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -1561,6 +1561,45 @@ static inline bool is_local_partition(struct cpuset *cs)
>   	return is_partition_valid(cs) && !is_remote_partition(cs);
>   }
>   
> +/**
> + * validate_remote_partition - Validate a remote partition CPU change
> + * @cs: cpuset being enabled or updated
> + * @prs: partition root state to validate
> + * @excpus: resulting effective exclusive CPU mask
> + * @addcpus: exclusive CPUs to be added
> + * @delcpus: exclusive CPUs to be deleted, can be NULL
> + *
> + * Return: PERR_NONE if valid, otherwise an appropriate error code
> + */
> +static enum prs_errcode
> +validate_remote_partition(struct cpuset *cs, int prs,
> +			  struct cpumask *excpus,
> +			  struct cpumask *addcpus,
> +			  struct cpumask *delcpus)
> +{
> +	bool updating = is_remote_partition(cs);
> +
> +	if (!capable(CAP_SYS_ADMIN))
> +		return PERR_ACCESS;
> +
> +	if (!updating &&
> +	    (!cpumask_intersects(excpus, cpu_active_mask) ||
> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
> +		return PERR_INVCPUS;
> +
> +	if (cpumask_intersects(addcpus, subpartitions_cpus) ||
> +	    (updating &&
> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
> +		return PERR_NOCPUS;
> +
> +	if ((prs == PRS_ISOLATED &&
> +	     !isolated_cpus_can_update(addcpus, delcpus)) ||
> +	    prstate_housekeeping_conflict(prs, excpus))
> +		return PERR_HKEEPING;
> +
> +	return PERR_NONE;
> +}
> +
>   /*
>    * remote_partition_enable - Enable current cpuset as a remote partition root
>    * @cs: the cpuset to update
> @@ -1574,11 +1613,7 @@ static inline bool is_local_partition(struct cpuset *cs)
>   static int remote_partition_enable(struct cpuset *cs, int new_prs,
>   				   struct tmpmasks *tmp)
>   {
> -	/*
> -	 * The user must have sysadmin privilege.
> -	 */
> -	if (!capable(CAP_SYS_ADMIN))
> -		return PERR_ACCESS;
> +	enum prs_errcode err;
>   
>   	/*
>   	 * The requested exclusive_cpus must not be allocated to other
> @@ -1591,15 +1626,10 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs,
>   	 * above it or remote partition root underneath it is not allowed.
>   	 */
>   	compute_excpus(cs, tmp->new_cpus);
> -	if (!cpumask_intersects(tmp->new_cpus, cpu_active_mask) ||
> -	    cpumask_subset(top_cpuset.effective_cpus, tmp->new_cpus))
> -		return PERR_INVCPUS;
> -	if (cpumask_intersects(tmp->new_cpus, subpartitions_cpus))
> -		return PERR_NOCPUS;
> -	if (((new_prs == PRS_ISOLATED) &&
> -	     !isolated_cpus_can_update(tmp->new_cpus, NULL)) ||
> -	    prstate_housekeeping_conflict(new_prs, tmp->new_cpus))
> -		return PERR_HKEEPING;
> +	err = validate_remote_partition(cs, new_prs,
> +					tmp->new_cpus, tmp->new_cpus, NULL);
> +	if (err)
> +		return err;
>   
>   	spin_lock_irq(&callback_lock);
>   	partition_xcpus_add(new_prs, NULL, tmp->new_cpus);
> @@ -1672,6 +1702,7 @@ static void remote_partition_disable(struct cpuset *cs, struct tmpmasks *tmp)
>   static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
>   			       struct cpumask *excpus, struct tmpmasks *tmp)
>   {
> +	enum prs_errcode err;
>   	bool adding, deleting;
>   	int prs = cs->partition_root_state;
>   
> @@ -1695,14 +1726,10 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
>   	 */
>   	if (adding) {
>   		WARN_ON_ONCE(cpumask_intersects(tmp->addmask, subpartitions_cpus));
> -		if (!capable(CAP_SYS_ADMIN))
> -			WRITE_ONCE(cs->prs_err, PERR_ACCESS);
> -		else if (cpumask_intersects(tmp->addmask, subpartitions_cpus) ||
> -			 cpumask_subset(top_cpuset.effective_cpus, tmp->addmask))
> -			WRITE_ONCE(cs->prs_err, PERR_NOCPUS);
> -		else if ((prs == PRS_ISOLATED) &&
> -			 !isolated_cpus_can_update(tmp->addmask, tmp->delmask))
> -			WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
> +		err = validate_remote_partition(cs, prs, excpus,
> +						tmp->addmask, tmp->delmask);
> +		if (err)
> +			WRITE_ONCE(cs->prs_err, err);
>   		if (cs->prs_err)
>   			goto invalidate;
>   	}
>
> base-commit: 31c88350b7dd1522792f726f79607f31bb55c50f

Thanks for fixing this.

Reviewed-by: Waiman Long <longman@redhat.com>


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
  2026-09-27  9:57 [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict Guopeng Zhang
  2026-09-28  0:00 ` Waiman Long
@ 2026-09-28  1:29 ` Ridong Chen
  2026-09-28 18:44 ` Tejun Heo
  2 siblings, 0 replies; 6+ messages in thread
From: Ridong Chen @ 2026-09-28  1:29 UTC (permalink / raw)
  To: Guopeng Zhang, Waiman Long
  Cc: Tejun Heo, Johannes Weiner, Michal Koutný,
	cgroups, linux-kernel, Guopeng Zhang



On 9/27/2026 5:57 PM, Guopeng Zhang wrote:
> From: Guopeng Zhang <zhangguopeng@kylinos.cn>
> 
> Widening an ancestor's exclusive CPU mask can add a boot-isolated CPU
> to a valid remote partition root without touching the partition's own
> control files. The partition then load balances that CPU, silently
> defeating isolcpus=domain for it.
> 
> This can be reproduced on a 32-CPU system booted with
> isolcpus=domain,4:
> 
>      cd /sys/fs/cgroup
>      echo +cpuset > cgroup.subtree_control
>      mkdir -p A/B
>      echo +cpuset > A/cgroup.subtree_control
>      echo 2-4 > A/cpuset.cpus
>      echo 2-3 > A/cpuset.cpus.exclusive
>      echo 2-4 > A/B/cpuset.cpus
>      echo 2-4 > A/B/cpuset.cpus.exclusive
>      echo root > A/B/cpuset.cpus.partition
>      cat A/B/cpuset.cpus.effective               # 2-3
>      echo 2-4 > A/cpuset.cpus.exclusive
>      cat A/B/cpuset.cpus.partition               # root
>      cat A/B/cpuset.cpus.effective               # 2-4
> 
> The last write returns 0 and leaves the hierarchy in this state:
> 
>      root (cpuset.cpus.effective=0-1,5-31)
>      |
>      \-- A (member):           cpuset.cpus=2-4
>          |                     cpuset.cpus.exclusive=2-4
>          \-- B (root, remote): cpuset.cpus=2-4
>                                cpuset.cpus.effective=2-4
> 
> B is a remote partition: it takes its CPUs directly from the root
> cpuset, and A only passes its exclusive list down. Before the last
> write, that list is 2-3, so B holds 2-3 and CPU 4 stays in the
> root cpuset as a boot-isolated CPU. The write widens A's exclusive
> list to 2-4, which additionally grants CPU 4 to B. Nothing rejects
> the grant: B remains a valid root partition, and CPU 4 is still
> listed in cpuset.cpus.isolated while sitting in a load-balanced
> partition.
> 
> remote_partition_enable() already rejects such grants through
> prstate_housekeeping_conflict(). remote_cpus_update(), which applies
> ancestor changes to a remote partition, does not.
> 
> Both paths also open-code remote partition validation. Move those
> checks into validate_remote_partition(), and check the resulting
> effective exclusive CPU mask for a housekeeping conflict there. The
> existing prs_err path then invalidates the remote partition instead of
> assigning it a boot-isolated CPU.
> 
> Fixes: f62a5d39368e ("cgroup/cpuset: Remove remote_partition_check() & make update_cpumasks_hier() handle remote partition")
> Suggested-by: Ridong Chen <ridong.chen@linux.dev>
> Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
> ---
> Changes since v1:
> - Consolidate remote partition validation in a common helper, as
>    suggested by Ridong.
> - Check housekeeping conflicts against the resulting effective exclusive
>    CPU mask.
> - Rebase onto cgroup/for-7.3-fixes.
> 
>   kernel/cgroup/cpuset.c | 71 +++++++++++++++++++++++++++++-------------
>   1 file changed, 49 insertions(+), 22 deletions(-)
> 
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index 1fcec89a28b9..5adf47217e59 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -1561,6 +1561,45 @@ static inline bool is_local_partition(struct cpuset *cs)
>   	return is_partition_valid(cs) && !is_remote_partition(cs);
>   }
>   
> +/**
> + * validate_remote_partition - Validate a remote partition CPU change
> + * @cs: cpuset being enabled or updated
> + * @prs: partition root state to validate
> + * @excpus: resulting effective exclusive CPU mask
> + * @addcpus: exclusive CPUs to be added
> + * @delcpus: exclusive CPUs to be deleted, can be NULL
> + *
> + * Return: PERR_NONE if valid, otherwise an appropriate error code
> + */
> +static enum prs_errcode
> +validate_remote_partition(struct cpuset *cs, int prs,
> +			  struct cpumask *excpus,
> +			  struct cpumask *addcpus,
> +			  struct cpumask *delcpus)
> +{
> +	bool updating = is_remote_partition(cs);
> +
> +	if (!capable(CAP_SYS_ADMIN))
> +		return PERR_ACCESS;
> +
> +	if (!updating &&
> +	    (!cpumask_intersects(excpus, cpu_active_mask) ||
> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
> +		return PERR_INVCPUS;
> +
> +	if (cpumask_intersects(addcpus, subpartitions_cpus) ||
> +	    (updating &&
> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
> +		return PERR_NOCPUS;
> +
> +	if ((prs == PRS_ISOLATED &&
> +	     !isolated_cpus_can_update(addcpus, delcpus)) ||
> +	    prstate_housekeeping_conflict(prs, excpus))
> +		return PERR_HKEEPING;
> +
> +	return PERR_NONE;
> +}
> +

validate_partition is used for both local and remote partitions, so I think this 
check may be added here in the future. For now, it looks good to me.

>   /*
>    * remote_partition_enable - Enable current cpuset as a remote partition root
>    * @cs: the cpuset to update
> @@ -1574,11 +1613,7 @@ static inline bool is_local_partition(struct cpuset *cs)
>   static int remote_partition_enable(struct cpuset *cs, int new_prs,
>   				   struct tmpmasks *tmp)
>   {
> -	/*
> -	 * The user must have sysadmin privilege.
> -	 */
> -	if (!capable(CAP_SYS_ADMIN))
> -		return PERR_ACCESS;
> +	enum prs_errcode err;
>   
>   	/*
>   	 * The requested exclusive_cpus must not be allocated to other
> @@ -1591,15 +1626,10 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs,
>   	 * above it or remote partition root underneath it is not allowed.
>   	 */
>   	compute_excpus(cs, tmp->new_cpus);
> -	if (!cpumask_intersects(tmp->new_cpus, cpu_active_mask) ||
> -	    cpumask_subset(top_cpuset.effective_cpus, tmp->new_cpus))
> -		return PERR_INVCPUS;
> -	if (cpumask_intersects(tmp->new_cpus, subpartitions_cpus))
> -		return PERR_NOCPUS;
> -	if (((new_prs == PRS_ISOLATED) &&
> -	     !isolated_cpus_can_update(tmp->new_cpus, NULL)) ||
> -	    prstate_housekeeping_conflict(new_prs, tmp->new_cpus))
> -		return PERR_HKEEPING;
> +	err = validate_remote_partition(cs, new_prs,
> +					tmp->new_cpus, tmp->new_cpus, NULL);
> +	if (err)
> +		return err;
>   
>   	spin_lock_irq(&callback_lock);
>   	partition_xcpus_add(new_prs, NULL, tmp->new_cpus);
> @@ -1672,6 +1702,7 @@ static void remote_partition_disable(struct cpuset *cs, struct tmpmasks *tmp)
>   static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
>   			       struct cpumask *excpus, struct tmpmasks *tmp)
>   {
> +	enum prs_errcode err;
>   	bool adding, deleting;
>   	int prs = cs->partition_root_state;
>   
> @@ -1695,14 +1726,10 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
>   	 */
>   	if (adding) {
>   		WARN_ON_ONCE(cpumask_intersects(tmp->addmask, subpartitions_cpus));
> -		if (!capable(CAP_SYS_ADMIN))
> -			WRITE_ONCE(cs->prs_err, PERR_ACCESS);
> -		else if (cpumask_intersects(tmp->addmask, subpartitions_cpus) ||
> -			 cpumask_subset(top_cpuset.effective_cpus, tmp->addmask))
> -			WRITE_ONCE(cs->prs_err, PERR_NOCPUS);
> -		else if ((prs == PRS_ISOLATED) &&
> -			 !isolated_cpus_can_update(tmp->addmask, tmp->delmask))
> -			WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
> +		err = validate_remote_partition(cs, prs, excpus,
> +						tmp->addmask, tmp->delmask);
> +		if (err)
> +			WRITE_ONCE(cs->prs_err, err);
>   		if (cs->prs_err)
>   			goto invalidate;
>   	}
> 
> base-commit: 31c88350b7dd1522792f726f79607f31bb55c50f

Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Thanks.
-- 
Best regards
Ridong


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
  2026-09-27  9:57 [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict Guopeng Zhang
  2026-09-28  0:00 ` Waiman Long
  2026-09-28  1:29 ` Ridong Chen
@ 2026-09-28 18:44 ` Tejun Heo
  2026-09-28 23:56   ` Waiman Long
  2026-09-29  2:33   ` Guopeng Zhang
  2 siblings, 2 replies; 6+ messages in thread
From: Tejun Heo @ 2026-09-28 18:44 UTC (permalink / raw)
  To: Guopeng Zhang, Waiman Long
  Cc: Ridong Chen, Johannes Weiner, Michal Koutný,
	Guopeng Zhang, cgroups, linux-kernel

Hello, Guopeng.

The following is a Claude-generated review.

On Sun, Sep 27, 2026 at 05:57:21PM +0800, Guopeng Zhang wrote:
> +	bool updating = is_remote_partition(cs);
> +
> +	if (!capable(CAP_SYS_ADMIN))
> +		return PERR_ACCESS;
> +
> +	if (!updating &&
> +	    (!cpumask_intersects(excpus, cpu_active_mask) ||
> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
> +		return PERR_INVCPUS;

cs->remote_partition can be stale when remote_partition_enable() is
called. If a root <-> isolated switch of a valid remote partition fails
in update_prstate(), the partition is made invalid without
remote_partition_disable(), so the flag stays set and its CPUs stay in
subpartitions_cpus. A later enable then skips the PERR_INVCPUS checks.
With isolcpus=domain,4:

1. A is a member with cpuset.cpus 2-6 and cpuset.cpus.exclusive 2-4. B
   has cpuset.cpus 2-4 and no cpuset.cpus.exclusive.
2. "isolated" to B's cpuset.cpus.partition. B becomes a valid remote
   partition with effective_xcpus 2-4.
3. "root" to B's cpuset.cpus.partition fails with PERR_HKEEPING.
   effective_xcpus is cleared but remote_partition stays set.
4. 5-6 to A's cpuset.cpus.exclusive. B's excpus becomes empty, which
   matches the cleared effective_xcpus, so update_cpumasks_hier() skips
   B.
5. "isolated" to B's cpuset.cpus.partition. This used to fail with
   PERR_INVCPUS. Now B becomes a valid isolated partition with empty
   effective_xcpus and the WARN_ON_ONCE() at the end of update_prstate()
   triggers.

This is from reading the code, not reproduced. Maybe have the callers
pass whether it's an enable or an update instead of deriving it from
cs->remote_partition?

The stale flag itself is a separate, pre-existing bug. Switching B back
to member after step 3 trips the WARN_ON_ONCE(old_prs < 0) in
partition_xcpus_del().

Thanks.

--
tejun

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
  2026-09-28 18:44 ` Tejun Heo
@ 2026-09-28 23:56   ` Waiman Long
  2026-09-29  2:33   ` Guopeng Zhang
  1 sibling, 0 replies; 6+ messages in thread
From: Waiman Long @ 2026-09-28 23:56 UTC (permalink / raw)
  To: Tejun Heo, Guopeng Zhang
  Cc: Ridong Chen, Johannes Weiner, Michal Koutný,
	Guopeng Zhang, cgroups, linux-kernel

On 9/28/26 2:44 PM, Tejun Heo wrote:
> Hello, Guopeng.
>
> The following is a Claude-generated review.
>
> On Sun, Sep 27, 2026 at 05:57:21PM +0800, Guopeng Zhang wrote:
>> +	bool updating = is_remote_partition(cs);
>> +
>> +	if (!capable(CAP_SYS_ADMIN))
>> +		return PERR_ACCESS;
>> +
>> +	if (!updating &&
>> +	    (!cpumask_intersects(excpus, cpu_active_mask) ||
>> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
>> +		return PERR_INVCPUS;
> cs->remote_partition can be stale when remote_partition_enable() is
> called. If a root <-> isolated switch of a valid remote partition fails
> in update_prstate(), the partition is made invalid without
> remote_partition_disable(), so the flag stays set and its CPUs stay in
> subpartitions_cpus. A later enable then skips the PERR_INVCPUS checks.
> With isolcpus=domain,4:
>
> 1. A is a member with cpuset.cpus 2-6 and cpuset.cpus.exclusive 2-4. B
>     has cpuset.cpus 2-4 and no cpuset.cpus.exclusive.
> 2. "isolated" to B's cpuset.cpus.partition. B becomes a valid remote
>     partition with effective_xcpus 2-4.
> 3. "root" to B's cpuset.cpus.partition fails with PERR_HKEEPING.
>     effective_xcpus is cleared but remote_partition stays set.
> 4. 5-6 to A's cpuset.cpus.exclusive. B's excpus becomes empty, which
>     matches the cleared effective_xcpus, so update_cpumasks_hier() skips
>     B.
> 5. "isolated" to B's cpuset.cpus.partition. This used to fail with
>     PERR_INVCPUS. Now B becomes a valid isolated partition with empty
>     effective_xcpus and the WARN_ON_ONCE() at the end of update_prstate()
>     triggers.
>
> This is from reading the code, not reproduced. Maybe have the callers
> pass whether it's an enable or an update instead of deriving it from
> cs->remote_partition?
>
> The stale flag itself is a separate, pre-existing bug. Switching B back
> to member after step 3 trips the WARN_ON_ONCE(old_prs < 0) in
> partition_xcpus_del().

You are right. It is a bug that has to be fixed. I have posted a cpuset 
patch to fix that partition state switch bug.

Cheers,
Longman

>
> Thanks.
>
> --
> tejun


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict
  2026-09-28 18:44 ` Tejun Heo
  2026-09-28 23:56   ` Waiman Long
@ 2026-09-29  2:33   ` Guopeng Zhang
  1 sibling, 0 replies; 6+ messages in thread
From: Guopeng Zhang @ 2026-09-29  2:33 UTC (permalink / raw)
  To: Tejun Heo, Waiman Long
  Cc: Ridong Chen, Johannes Weiner, Michal Koutný,
	Guopeng Zhang, cgroups, linux-kernel



在 2026/9/29 02:44, Tejun Heo 写道:
Hi,Tejun

> Hello, Guopeng.
> 
> The following is a Claude-generated review.
> 
> On Sun, Sep 27, 2026 at 05:57:21PM +0800, Guopeng Zhang wrote:
>> +	bool updating = is_remote_partition(cs);
>> +
>> +	if (!capable(CAP_SYS_ADMIN))
>> +		return PERR_ACCESS;
>> +
>> +	if (!updating &&
>> +	    (!cpumask_intersects(excpus, cpu_active_mask) ||
>> +	     cpumask_subset(top_cpuset.effective_cpus, addcpus)))
>> +		return PERR_INVCPUS;
> 
> cs->remote_partition can be stale when remote_partition_enable() is
> called. If a root <-> isolated switch of a valid remote partition fails
> in update_prstate(), the partition is made invalid without
> remote_partition_disable(), so the flag stays set and its CPUs stay in
> subpartitions_cpus. A later enable then skips the PERR_INVCPUS checks.
> With isolcpus=domain,4:
> 
> 1. A is a member with cpuset.cpus 2-6 and cpuset.cpus.exclusive 2-4. B
>    has cpuset.cpus 2-4 and no cpuset.cpus.exclusive.
> 2. "isolated" to B's cpuset.cpus.partition. B becomes a valid remote
>    partition with effective_xcpus 2-4.
> 3. "root" to B's cpuset.cpus.partition fails with PERR_HKEEPING.
>    effective_xcpus is cleared but remote_partition stays set.
> 4. 5-6 to A's cpuset.cpus.exclusive. B's excpus becomes empty, which
>    matches the cleared effective_xcpus, so update_cpumasks_hier() skips
>    B.
> 5. "isolated" to B's cpuset.cpus.partition. This used to fail with
>    PERR_INVCPUS. Now B becomes a valid isolated partition with empty
>    effective_xcpus and the WARN_ON_ONCE() at the end of update_prstate()
>    triggers.
> 
> This is from reading the code, not reproduced. Maybe have the callers
> pass whether it's an enable or an update instead of deriving it from
> cs->remote_partition?

You're right. I'll fix it in v3.

Thanks,
Guopeng
> 
> The stale flag itself is a separate, pre-existing bug. Switching B back
> to member after step 3 trips the WARN_ON_ONCE(old_prs < 0) in
> partition_xcpus_del().
> 
> Thanks.
> 
> --
> tejun


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-29  2:33 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27  9:57 [PATCH v2] cgroup/cpuset: Invalidate remote partition on housekeeping conflict Guopeng Zhang
2026-09-28  0:00 ` Waiman Long
2026-09-28  1:29 ` Ridong Chen
2026-09-28 18:44 ` Tejun Heo
2026-09-28 23:56   ` Waiman Long
2026-09-29  2:33   ` Guopeng Zhang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®