* Re: [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
2026-09-30 3:18 [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode() Waiman Long
@ 2026-09-30 4:13 ` Ridong Chen
2026-09-30 5:38 ` Andrea Righi
` (2 subsequent siblings)
3 siblings, 0 replies; 6+ messages in thread
From: Ridong Chen @ 2026-09-30 4:13 UTC (permalink / raw)
To: Waiman Long, Tejun Heo, Johannes Weiner, Michal Koutný
Cc: cgroups, linux-kernel, Andrea Righi
On 9/30/2026 11:18 AM, Waiman Long wrote:
> After seeing the patch [1] to guard is_in_v2_mode() with RCU, it makes
> me realize that is_in_v2_mode() may be called in a context where a new
> cgroup filesystem is being rebound with stale cpuset_cgrp_subsys.root
> pointer. Avoid this potential UaF situation by adding a new cpuset_v2_mode
> flag which is set when the cpuset_v2_mode mount option is used. This
> flag is written into only when cpuset_bind() is being called with a
> stable cpuset_cgrp_subsys.root value. The is_in_v2_mode() helper is
> modified to read the new cpuset_v2_mode flag instead of accessing
> cpuset_cgrp_subsys.root directly.
>
> [1] https://lore.kernel.org/lkml/20260929084124.626693-2-arighi@nvidia.com
>
> Fixes: b8d1b8ee93df ("cpuset: Allow v2 behavior in v1 cgroup")
> Signed-off-by: Waiman Long <longman@redhat.com>
> ---
> kernel/cgroup/cpuset.c | 11 +++++++++--
> 1 file changed, 9 insertions(+), 2 deletions(-)
>
> [v2] Update the comment as suggested by Ridong
>
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index 19661df6244f..4a2b6f154609 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -152,6 +152,12 @@ static cpumask_var_t isolated_cpus; /* CSCB */
> */
> static bool update_housekeeping; /* RWCS */
>
> +/*
> + * Set if "cpuset_v2_mode" mount option is used
> + * Cached at bind time and not lock protected; accessed via {READ,WRITE}_ONCE
> + */
> +static bool cpuset_v2_mode;
> +
> /*
> * Copy of isolated_cpus to be passed to housekeeping_update()
> */
> @@ -439,8 +445,7 @@ static inline bool cpuset_v2(void)
> */
> static inline bool is_in_v2_mode(void)
> {
> - return cpuset_v2() ||
> - (cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE);
> + return cpuset_v2() || READ_ONCE(cpuset_v2_mode);
> }
>
> /**
> @@ -3664,6 +3669,8 @@ static void cpuset_bind(struct cgroup_subsys_state *root_css)
> mutex_lock(&cpuset_mutex);
> spin_lock_irq(&callback_lock);
>
> + WRITE_ONCE(cpuset_v2_mode,
> + !!(cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE));
> if (is_in_v2_mode()) {
> cpumask_copy(top_cpuset.cpus_allowed, cpu_possible_mask);
> cpumask_copy(top_cpuset.effective_xcpus, cpu_possible_mask);
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Thanks.
--
Best regards
Ridong
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
2026-09-30 3:18 [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode() Waiman Long
2026-09-30 4:13 ` Ridong Chen
@ 2026-09-30 5:38 ` Andrea Righi
2026-09-30 17:45 ` Tejun Heo
2026-09-30 17:57 ` Michal Koutný
3 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-09-30 5:38 UTC (permalink / raw)
To: Waiman Long
Cc: Ridong Chen, Tejun Heo, Johannes Weiner, Michal Koutný,
cgroups, linux-kernel
Hi Waiman,
On Tue, Sep 29, 2026 at 11:18:33PM -0400, Waiman Long wrote:
> After seeing the patch [1] to guard is_in_v2_mode() with RCU, it makes
> me realize that is_in_v2_mode() may be called in a context where a new
> cgroup filesystem is being rebound with stale cpuset_cgrp_subsys.root
> pointer. Avoid this potential UaF situation by adding a new cpuset_v2_mode
> flag which is set when the cpuset_v2_mode mount option is used. This
> flag is written into only when cpuset_bind() is being called with a
> stable cpuset_cgrp_subsys.root value. The is_in_v2_mode() helper is
> modified to read the new cpuset_v2_mode flag instead of accessing
> cpuset_cgrp_subsys.root directly.
>
> [1] https://lore.kernel.org/lkml/20260929084124.626693-2-arighi@nvidia.com
>
> Fixes: b8d1b8ee93df ("cpuset: Allow v2 behavior in v1 cgroup")
> Signed-off-by: Waiman Long <longman@redhat.com>
This works for me.
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Thanks,
-Andrea
> ---
> kernel/cgroup/cpuset.c | 11 +++++++++--
> 1 file changed, 9 insertions(+), 2 deletions(-)
>
> [v2] Update the comment as suggested by Ridong
>
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index 19661df6244f..4a2b6f154609 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -152,6 +152,12 @@ static cpumask_var_t isolated_cpus; /* CSCB */
> */
> static bool update_housekeeping; /* RWCS */
>
> +/*
> + * Set if "cpuset_v2_mode" mount option is used
> + * Cached at bind time and not lock protected; accessed via {READ,WRITE}_ONCE
> + */
> +static bool cpuset_v2_mode;
> +
> /*
> * Copy of isolated_cpus to be passed to housekeeping_update()
> */
> @@ -439,8 +445,7 @@ static inline bool cpuset_v2(void)
> */
> static inline bool is_in_v2_mode(void)
> {
> - return cpuset_v2() ||
> - (cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE);
> + return cpuset_v2() || READ_ONCE(cpuset_v2_mode);
> }
>
> /**
> @@ -3664,6 +3669,8 @@ static void cpuset_bind(struct cgroup_subsys_state *root_css)
> mutex_lock(&cpuset_mutex);
> spin_lock_irq(&callback_lock);
>
> + WRITE_ONCE(cpuset_v2_mode,
> + !!(cpuset_cgrp_subsys.root->flags & CGRP_ROOT_CPUSET_V2_MODE));
> if (is_in_v2_mode()) {
> cpumask_copy(top_cpuset.cpus_allowed, cpu_possible_mask);
> cpumask_copy(top_cpuset.effective_xcpus, cpu_possible_mask);
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
2026-09-30 3:18 [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode() Waiman Long
2026-09-30 4:13 ` Ridong Chen
2026-09-30 5:38 ` Andrea Righi
@ 2026-09-30 17:45 ` Tejun Heo
2026-09-30 17:57 ` Michal Koutný
3 siblings, 0 replies; 6+ messages in thread
From: Tejun Heo @ 2026-09-30 17:45 UTC (permalink / raw)
To: Waiman Long
Cc: Ridong Chen, Johannes Weiner, Michal Koutný,
Andrea Righi, cgroups, linux-kernel
Applied to cgroup/for-7.3-fixes with the following tags added:
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
2026-09-30 3:18 [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode() Waiman Long
` (2 preceding siblings ...)
2026-09-30 17:45 ` Tejun Heo
@ 2026-09-30 17:57 ` Michal Koutný
2026-09-30 19:03 ` Waiman Long
3 siblings, 1 reply; 6+ messages in thread
From: Michal Koutný @ 2026-09-30 17:57 UTC (permalink / raw)
To: Waiman Long
Cc: Ridong Chen, Tejun Heo, Johannes Weiner, cgroups, linux-kernel,
Andrea Righi
[-- Attachment #1: Type: text/plain, Size: 1372 bytes --]
On Tue, Sep 29, 2026 at 11:18:33PM -0400, Waiman Long <longman@redhat.com> wrote:
> After seeing the patch [1] to guard is_in_v2_mode() with RCU, it makes
> me realize that is_in_v2_mode() may be called in a context where a new
> cgroup filesystem is being rebound with stale cpuset_cgrp_subsys.root
> pointer.
add: which is a privileged operation.
> Avoid this potential UaF situation by adding a new cpuset_v2_mode
> flag which is set when the cpuset_v2_mode mount option is used. This
> flag is written into only when cpuset_bind() is being called with a
> stable cpuset_cgrp_subsys.root value. The is_in_v2_mode() helper is
> modified to read the new cpuset_v2_mode flag instead of accessing
> cpuset_cgrp_subsys.root directly.
You write about "potential" situation. But are there any such callers
(after the cpuset_v2() conversion in cpuset_num_cpus())?
> Fixes: b8d1b8ee93df ("cpuset: Allow v2 behavior in v1 cgroup")
I'd even consider (pruning to)
Fixes: d23b5c5777158 ("cgroup: Make operations on the cgroup root_list RCU safe")
(Admittedly, is_in_v2_mode() wasn't always called under cgroup_mutex, but
I'd argue (without comprehensive analysis) that predicate's callers were
synchronized via cpuset_mutex in cpuset_bind() (and cgroup_mutex)
anyways.)
The caching you proposed looks safe. (I'm mainly writing because of the
commit message wrt UaF.)
Michal
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 265 bytes --]
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2] cgroup/cpuset: Don't access cpuset_cgrp_subsys.root in is_in_v2_mode()
2026-09-30 17:57 ` Michal Koutný
@ 2026-09-30 19:03 ` Waiman Long
0 siblings, 0 replies; 6+ messages in thread
From: Waiman Long @ 2026-09-30 19:03 UTC (permalink / raw)
To: Michal Koutný
Cc: Ridong Chen, Tejun Heo, Johannes Weiner, cgroups, linux-kernel,
Andrea Righi
On 9/30/26 1:57 PM, Michal Koutný wrote:
> On Tue, Sep 29, 2026 at 11:18:33PM -0400, Waiman Long <longman@redhat.com> wrote:
>> After seeing the patch [1] to guard is_in_v2_mode() with RCU, it makes
>> me realize that is_in_v2_mode() may be called in a context where a new
>> cgroup filesystem is being rebound with stale cpuset_cgrp_subsys.root
>> pointer.
> add: which is a privileged operation.
>
>> Avoid this potential UaF situation by adding a new cpuset_v2_mode
>> flag which is set when the cpuset_v2_mode mount option is used. This
>> flag is written into only when cpuset_bind() is being called with a
>> stable cpuset_cgrp_subsys.root value. The is_in_v2_mode() helper is
>> modified to read the new cpuset_v2_mode flag instead of accessing
>> cpuset_cgrp_subsys.root directly.
> You write about "potential" situation. But are there any such callers
> (after the cpuset_v2() conversion in cpuset_num_cpus())?
Other than the cpuset_num_cpus(), is_in_v2_mode() should only be used
with either cpuset_mutex or callback_lock held. Rebinding a cgroup
subsystem is protected by holding cgroup_mutex. I am not aware of any
situation where rebinding is happening while cpuset code is being called
into, but I can't rule out that possibility. Also it is possible that
future cpuset code extension may make it possible that this race
condition can happen. For safety, it is better to make the code safe.
That is the reason why I said it is a potential situation.
>
>> Fixes: b8d1b8ee93df ("cpuset: Allow v2 behavior in v1 cgroup")
> I'd even consider (pruning to)
> Fixes: d23b5c5777158 ("cgroup: Make operations on the cgroup root_list RCU safe")
>
> (Admittedly, is_in_v2_mode() wasn't always called under cgroup_mutex, but
> I'd argue (without comprehensive analysis) that predicate's callers were
> synchronized via cpuset_mutex in cpuset_bind() (and cgroup_mutex)
> anyways.)
>
> The caching you proposed looks safe. (I'm mainly writing because of the
> commit message wrt UaF.)
I said UaF because it is what Andrea patch is implying. As I said
before, I don't think that can happen, but I can't rule it out.
Cheers,
Longman
>
> Michal
^ permalink raw reply [flat|nested] 6+ messages in thread