* [PATCH] sched/cache: Remove the old cache group footprint on exec
@ 2026-10-09 15:12 Jemmy Wong
2026-10-09 21:01 ` Tim Chen
0 siblings, 1 reply; 3+ messages in thread
From: Jemmy Wong @ 2026-10-09 15:12 UTC (permalink / raw)
To: Ingo Molnar, Peter Zijlstra, Tim Chen
Cc: Juri Lelli, Vincent Guittot, Dietmar Eggemann, Steven Rostedt,
Ben Segall, Mel Gorman, Valentin Schneider, K Prateek Nayak,
Chen Yu, Jemmy Wong, linux-kernel
Exec replaces the task's cache group before resetting its NUMA fault
statistics, but only the exit path subtracts the task's contribution
from the old group's footprint.
Although de_thread() removes other members of the executing task's
thread group, tasks created with CLONE_VM without CLONE_THREAD can
retain the old mm and cache group. The executing task's contribution
then remains in that group without further updates or decay, potentially
suppressing cache-aware aggregation through the LLC capacity check.
Factor the existing footprint subtraction into a helper and use it when
leaving a cache group on both exec and exit. Subtract before dropping the
old group reference, preserving the existing underflow protection.
Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
---
kernel/sched/fair.c | 51 ++++++++++++++++++++++++++-------------------
1 file changed, 29 insertions(+), 22 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5cdde4..af182c9fcde7 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1771,6 +1771,25 @@ void sched_cache_fork_cleanup(struct task_struct *p)
RCU_INIT_POINTER(p->sched_cache_grp, NULL);
}
+static void sched_cache_footprint_sub(struct sched_cache_group *grp,
+ struct task_struct *p)
+{
+#ifdef CONFIG_NUMA_BALANCING
+ /*
+ * Remove this task's contribution when it leaves the group, either
+ * through exit or exec. Other thread groups sharing the old mm can
+ * keep the group alive after exec resets this task's NUMA statistics.
+ * Unlocked for performance; clamp to avoid underflow.
+ */
+ if (grp && p->total_numa_faults) {
+ unsigned long fp = READ_ONCE(grp->footprint);
+ unsigned long sub = min(fp, p->total_numa_faults);
+
+ WRITE_ONCE(grp->footprint, fp - sub);
+ }
+#endif
+}
+
void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
{
struct sched_cache_group *old;
@@ -1780,6 +1799,7 @@ void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
* the old one. @p is current and the only writer of its own pointer.
*/
old = sched_cache_replace_grp(p, sched_cache_group_get(mm->sched_cache_grp));
+ sched_cache_footprint_sub(old, p);
sched_cache_group_put(old);
}
@@ -1787,19 +1807,7 @@ void sched_cache_exit_mm(struct task_struct *p)
{
struct sched_cache_group *grp = sched_cache_replace_grp(p, NULL);
-#ifdef CONFIG_NUMA_BALANCING
- /*
- * Subtract this task's footprint from the group before dropping the
- * reference, so the group footprint converges as its threads exit.
- * Unlocked for performance; clamp to avoid underflow.
- */
- if (grp && p->total_numa_faults) {
- unsigned long fp = READ_ONCE(grp->footprint);
- unsigned long sub = min(fp, p->total_numa_faults);
-
- WRITE_ONCE(grp->footprint, fp - sub);
- }
-#endif
+ sched_cache_footprint_sub(grp, p);
sched_cache_group_put(grp);
}
@@ -3975,16 +3983,15 @@ static void task_numa_placement(struct task_struct *p)
* sharing this mm. Acceptable since footprint is a
* heuristic and occasional lost updates are tolerable.
*
- * If a task exits, its corresponding footprint must
- * be subtracted from p->sched_cache_grp->footprint,
- * otherwise the footprint will not converge: the
- * exiting thread's footprint remains unchanged/undecayed.
- * See exit_mm().
+ * If a task leaves its cache group through exit or exec,
+ * its contribution must be subtracted from the old group's
+ * footprint. Otherwise, that contribution remains
+ * unchanged/undecayed while other tasks keep the group alive.
+ * See sched_cache_footprint_sub().
*
- * Lost updates and unsynchronized subtraction
- * in exit_mm() can cause footprint + diff to
- * go negative. Clamp to zero to prevent the
- * unsigned footprint from wrapping.
+ * Lost updates and unsynchronized subtraction on exit or
+ * exec can cause footprint + diff to go negative. Clamp
+ * to zero to prevent the unsigned footprint from wrapping.
*/
scoped_guard(rcu) {
grp = rcu_dereference(p->sched_cache_grp);
--
2.54.0 (Apple Git-157)
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] sched/cache: Remove the old cache group footprint on exec
2026-10-09 15:12 [PATCH] sched/cache: Remove the old cache group footprint on exec Jemmy Wong
@ 2026-10-09 21:01 ` Tim Chen
2026-10-10 3:13 ` Jemmy Wong
0 siblings, 1 reply; 3+ messages in thread
From: Tim Chen @ 2026-10-09 21:01 UTC (permalink / raw)
To: Jemmy Wong, Ingo Molnar, Peter Zijlstra
Cc: Juri Lelli, Vincent Guittot, Dietmar Eggemann, Steven Rostedt,
Ben Segall, Mel Gorman, Valentin Schneider, K Prateek Nayak,
Chen Yu, linux-kernel
On Fri, 2026-10-09 at 23:12 +0800, Jemmy Wong wrote:
> Exec replaces the task's cache group before resetting its NUMA fault
> statistics, but only the exit path subtracts the task's contribution
> from the old group's footprint.
>
> Although de_thread() removes other members of the executing task's
> thread group, tasks created with CLONE_VM without CLONE_THREAD can
> retain the old mm and cache group. The executing task's contribution
> then remains in that group without further updates or decay, potentially
> suppressing cache-aware aggregation through the LLC capacity check.
Thanks for catching this.
It might help to mention vfork() explicitly, where the parent keeps the old mm
alive while the child execs. In practice a vfork child rarely builds
up NUMA faults before exec, since NUMA scanning starts late, so the
leak mostly matters for longer-lived CLONE_VM tasks that later exec. If
you have a workload where you saw this, please mention it. Otherwise,
saying it was found by code inspection is fine.
Please also add:
Fixes: b636fef85bda ("sched/cache: Introduce task_struct->sched_cache_grp to fix UAF")
>
> Factor the existing footprint subtraction into a helper and use it when
> leaving a cache group on both exec and exit. Subtract before dropping the
> old group reference, preserving the existing underflow protection.
>
> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
> ---
> kernel/sched/fair.c | 51 ++++++++++++++++++++++++++-------------------
> 1 file changed, 29 insertions(+), 22 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 57360f5cdde4..af182c9fcde7 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1771,6 +1771,25 @@ void sched_cache_fork_cleanup(struct task_struct *p)
> RCU_INIT_POINTER(p->sched_cache_grp, NULL);
> }
>
> +static void sched_cache_footprint_sub(struct sched_cache_group *grp,
> + struct task_struct *p)
> +{
> +#ifdef CONFIG_NUMA_BALANCING
> + /*
> + * Remove this task's contribution when it leaves the group, either
> + * through exit or exec. Other thread groups sharing the old mm can
> + * keep the group alive after exec resets this task's NUMA statistics.
> + * Unlocked for performance; clamp to avoid underflow.
> + */
> + if (grp && p->total_numa_faults) {
> + unsigned long fp = READ_ONCE(grp->footprint);
> + unsigned long sub = min(fp, p->total_numa_faults);
> +
> + WRITE_ONCE(grp->footprint, fp - sub);
> + }
> +#endif
> +}
> +
> void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
> {
> struct sched_cache_group *old;
> @@ -1780,6 +1799,7 @@ void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
> * the old one. @p is current and the only writer of its own pointer.
> */
> old = sched_cache_replace_grp(p, sched_cache_group_get(mm->sched_cache_grp));
> + sched_cache_footprint_sub(old, p);
> sched_cache_group_put(old);
> }
This relies on running before task_numa_free() in bprm_execve()
clears p->total_numa_faults, and nothing enforces that ordering. A
short comment here noting the dependency would keep a future exec
rework from quietly bringing the leak back.
With the changelog and comment tweaks above:
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
>
> @@ -1787,19 +1807,7 @@ void sched_cache_exit_mm(struct task_struct *p)
> {
> struct sched_cache_group *grp = sched_cache_replace_grp(p, NULL);
>
> -#ifdef CONFIG_NUMA_BALANCING
> - /*
> - * Subtract this task's footprint from the group before dropping the
> - * reference, so the group footprint converges as its threads exit.
> - * Unlocked for performance; clamp to avoid underflow.
> - */
> - if (grp && p->total_numa_faults) {
> - unsigned long fp = READ_ONCE(grp->footprint);
> - unsigned long sub = min(fp, p->total_numa_faults);
> -
> - WRITE_ONCE(grp->footprint, fp - sub);
> - }
> -#endif
> + sched_cache_footprint_sub(grp, p);
> sched_cache_group_put(grp);
> }
>
> @@ -3975,16 +3983,15 @@ static void task_numa_placement(struct task_struct *p)
> * sharing this mm. Acceptable since footprint is a
> * heuristic and occasional lost updates are tolerable.
> *
> - * If a task exits, its corresponding footprint must
> - * be subtracted from p->sched_cache_grp->footprint,
> - * otherwise the footprint will not converge: the
> - * exiting thread's footprint remains unchanged/undecayed.
> - * See exit_mm().
> + * If a task leaves its cache group through exit or exec,
> + * its contribution must be subtracted from the old group's
> + * footprint. Otherwise, that contribution remains
> + * unchanged/undecayed while other tasks keep the group alive.
> + * See sched_cache_footprint_sub().
> *
> - * Lost updates and unsynchronized subtraction
> - * in exit_mm() can cause footprint + diff to
> - * go negative. Clamp to zero to prevent the
> - * unsigned footprint from wrapping.
> + * Lost updates and unsynchronized subtraction on exit or
> + * exec can cause footprint + diff to go negative. Clamp
> + * to zero to prevent the unsigned footprint from wrapping.
> */
> scoped_guard(rcu) {
> grp = rcu_dereference(p->sched_cache_grp);
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] sched/cache: Remove the old cache group footprint on exec
2026-10-09 21:01 ` Tim Chen
@ 2026-10-10 3:13 ` Jemmy Wong
0 siblings, 0 replies; 3+ messages in thread
From: Jemmy Wong @ 2026-10-10 3:13 UTC (permalink / raw)
To: Tim Chen
Cc: Jemmy Wong, Ingo Molnar, Peter Zijlstra, Juri Lelli,
Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, Chen Yu,
linux-kernel
Hi Tim,
Thanks for the review.
This was found by code inspection, not a workload reproducer.
I'll mention vfork() and clarify that the issue mainly concerns
longer-lived CLONE_VM tasks that accumulate NUMA faults before exec.
I'll also add the Fixes tag and document that footprint subtraction
must precede task_numa_free() resetting p->total_numa_faults.
I'll send v2 with these changes and your Reviewed-by tag.
Thanks,
Jemmy
> On Oct 10, 2026, at 5:01 AM, Tim Chen <tim.c.chen@linux.intel.com> wrote:
>
> On Fri, 2026-10-09 at 23:12 +0800, Jemmy Wong wrote:
>> Exec replaces the task's cache group before resetting its NUMA fault
>> statistics, but only the exit path subtracts the task's contribution
>> from the old group's footprint.
>>
>> Although de_thread() removes other members of the executing task's
>> thread group, tasks created with CLONE_VM without CLONE_THREAD can
>> retain the old mm and cache group. The executing task's contribution
>> then remains in that group without further updates or decay, potentially
>> suppressing cache-aware aggregation through the LLC capacity check.
>
> Thanks for catching this.
>
> It might help to mention vfork() explicitly, where the parent keeps the old mm
> alive while the child execs. In practice a vfork child rarely builds
> up NUMA faults before exec, since NUMA scanning starts late, so the
> leak mostly matters for longer-lived CLONE_VM tasks that later exec. If
> you have a workload where you saw this, please mention it. Otherwise,
> saying it was found by code inspection is fine.
>
> Please also add:
>
> Fixes: b636fef85bda ("sched/cache: Introduce task_struct->sched_cache_grp to fix UAF")
>
>
>>
>> Factor the existing footprint subtraction into a helper and use it when
>> leaving a cache group on both exec and exit. Subtract before dropping the
>> old group reference, preserving the existing underflow protection.
>>
>> Signed-off-by: Jemmy Wong <jemmywong512@gmail.com>
>> ---
>> kernel/sched/fair.c | 51 ++++++++++++++++++++++++++-------------------
>> 1 file changed, 29 insertions(+), 22 deletions(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 57360f5cdde4..af182c9fcde7 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -1771,6 +1771,25 @@ void sched_cache_fork_cleanup(struct task_struct *p)
>> RCU_INIT_POINTER(p->sched_cache_grp, NULL);
>> }
>>
>> +static void sched_cache_footprint_sub(struct sched_cache_group *grp,
>> + struct task_struct *p)
>> +{
>> +#ifdef CONFIG_NUMA_BALANCING
>> + /*
>> + * Remove this task's contribution when it leaves the group, either
>> + * through exit or exec. Other thread groups sharing the old mm can
>> + * keep the group alive after exec resets this task's NUMA statistics.
>> + * Unlocked for performance; clamp to avoid underflow.
>> + */
>> + if (grp && p->total_numa_faults) {
>> + unsigned long fp = READ_ONCE(grp->footprint);
>> + unsigned long sub = min(fp, p->total_numa_faults);
>> +
>> + WRITE_ONCE(grp->footprint, fp - sub);
>> + }
>> +#endif
>> +}
>> +
>> void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
>> {
>> struct sched_cache_group *old;
>> @@ -1780,6 +1799,7 @@ void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm)
>> * the old one. @p is current and the only writer of its own pointer.
>> */
>> old = sched_cache_replace_grp(p, sched_cache_group_get(mm->sched_cache_grp));
>> + sched_cache_footprint_sub(old, p);
>> sched_cache_group_put(old);
>> }
>
> This relies on running before task_numa_free() in bprm_execve()
> clears p->total_numa_faults, and nothing enforces that ordering. A
> short comment here noting the dependency would keep a future exec
> rework from quietly bringing the leak back.
>
> With the changelog and comment tweaks above:
>
> Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
>
>
>
>>
>> @@ -1787,19 +1807,7 @@ void sched_cache_exit_mm(struct task_struct *p)
>> {
>> struct sched_cache_group *grp = sched_cache_replace_grp(p, NULL);
>>
>> -#ifdef CONFIG_NUMA_BALANCING
>> - /*
>> - * Subtract this task's footprint from the group before dropping the
>> - * reference, so the group footprint converges as its threads exit.
>> - * Unlocked for performance; clamp to avoid underflow.
>> - */
>> - if (grp && p->total_numa_faults) {
>> - unsigned long fp = READ_ONCE(grp->footprint);
>> - unsigned long sub = min(fp, p->total_numa_faults);
>> -
>> - WRITE_ONCE(grp->footprint, fp - sub);
>> - }
>> -#endif
>> + sched_cache_footprint_sub(grp, p);
>> sched_cache_group_put(grp);
>> }
>>
>> @@ -3975,16 +3983,15 @@ static void task_numa_placement(struct task_struct *p)
>> * sharing this mm. Acceptable since footprint is a
>> * heuristic and occasional lost updates are tolerable.
>> *
>> - * If a task exits, its corresponding footprint must
>> - * be subtracted from p->sched_cache_grp->footprint,
>> - * otherwise the footprint will not converge: the
>> - * exiting thread's footprint remains unchanged/undecayed.
>> - * See exit_mm().
>> + * If a task leaves its cache group through exit or exec,
>> + * its contribution must be subtracted from the old group's
>> + * footprint. Otherwise, that contribution remains
>> + * unchanged/undecayed while other tasks keep the group alive.
>> + * See sched_cache_footprint_sub().
>> *
>> - * Lost updates and unsynchronized subtraction
>> - * in exit_mm() can cause footprint + diff to
>> - * go negative. Clamp to zero to prevent the
>> - * unsigned footprint from wrapping.
>> + * Lost updates and unsynchronized subtraction on exit or
>> + * exec can cause footprint + diff to go negative. Clamp
>> + * to zero to prevent the unsigned footprint from wrapping.
>> */
>> scoped_guard(rcu) {
>> grp = rcu_dereference(p->sched_cache_grp);
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-10 3:13 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 15:12 [PATCH] sched/cache: Remove the old cache group footprint on exec Jemmy Wong
2026-10-09 21:01 ` Tim Chen
2026-10-10 3:13 ` Jemmy Wong
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®