* [PATCH 0/4] sched/cache: Fixes for cache aware scheduling
@ 2026-09-10 17:46 Tim Chen
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
` (3 more replies)
0 siblings, 4 replies; 11+ messages in thread
From: Tim Chen @ 2026-09-10 17:46 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar
Cc: Tim Chen, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
Hi all,
Cache aware scheduling went in for v7.2 and people have found a few
things wrong with it since. We collect the fixes in this series
so it is easier to track. Two keep tasks from
being stranded outside, or yanked away from, their preferred LLC; two
fix a use after free. Patches 1-2 stand alone, 3 and 4 go together.
Patch 1: alb_break_llc() compares nr_pref_llc_running with
cfs.h_nr_runnable, but those count different sets - one follows queued
tasks, the other drops delay-dequeued ones. With DELAY_DEQUEUE the
equality stops holding and active balance pulls a task off its preferred
LLC. So fix the counter. Reported by Zhan Xusheng:
https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
Patch 2 (Lu Wang): the stopper doing active load balance builds a fresh
lb_env that doesn't inherit migration_type, so can_migrate_task() can
move a task *out* of its preferred LLC. A new LBF_ACTIVE_LB_LLC flag
and picking the stopper callback at kick time keep the intent; passing
migration_type through the stopper would muddy delayed dequeue. v4:
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@gmail.com/
Patches 3-4 are for the use after free Hyunwoo Kim caught with KASAN:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
account_mm_sched() reaches the stats via p->mm->sc_stat, but a task can
be switching mm on one CPU while another is inside account_mm_sched(),
so the mm and the stats inside it can go away underneath. Locking the
rq in the mm free path felt like the wrong trade, so patch 4 pulls
sched_cache_stat out of mm_struct into a refcounted, RCU freed
sched_cache_group - just moving code - and patch 5 does the real fix:
each task takes its own reference (copy_mm(), exec_mmap(), dropped in
exit_mm()), so the group outlives any mm switch. Same Fixes: tag and
Hyunwoo's Tested-by on both; they want to go in together.
Nice side effect: the group no longer follows the address space,
so a user defined group, or cgroup or numa_group could own it later.
These are also the grouping by prctl RFC's first two patches, sent here
so the fix isn't held up by that discussion.
BTW, there are two other issues in discussion currently and need
a bit more work:
1. Incorrect donor context being passed to task_tick_cache().
https://lore.kernel.org/lkml/20260909092901.2989564-1-sh_def@163.com/
It is currently under discussion and is not included in this series.
2. Cache aware scheduling interfering with ITMT.
https://lore.kernel.org/lkml/20260810033742.1688718-1-yu.c.chen@intel.com/
https://lore.kernel.org/lkml/2fe2c681-b748-41fa-8b56-1169c86cefbc@intel.com/
Applies on sched/urgent branch.
Tim Chen and Chen Yu
Lu Wang (1):
sched/cache: Honor migrate_llc_task semantics in active load balance
Tim Chen (3):
sched/cache: Keep nr_pref_llc_running in the runnable domain
sched/cache: Decouple sched_cache_group from mm
sched/cache: Introduce task_struct->sched_cache_grp
fs/exec.c | 14 ++
include/linux/mm_types.h | 15 +-
include/linux/sched.h | 11 +-
kernel/exit.c | 28 +++-
kernel/fork.c | 23 +++
kernel/sched/build_utility.c | 4 +
kernel/sched/cache_sched.c | 39 +++++
kernel/sched/fair.c | 297 ++++++++++++++++++++++++++---------
kernel/sched/sched.h | 3 +
9 files changed, 340 insertions(+), 94 deletions(-)
create mode 100644 kernel/sched/cache_sched.c
--
2.32.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
2026-09-10 17:46 [PATCH 0/4] sched/cache: Fixes for cache aware scheduling Tim Chen
@ 2026-09-10 17:46 ` Tim Chen
2026-09-10 18:33 ` Kayra Cizmeci
2026-09-10 22:48 ` Kayra Cizmeci
2026-09-10 17:46 ` [PATCH 2/4] sched/cache: Honor migrate_llc_task semantics in active load balance Tim Chen
` (2 subsequent siblings)
3 siblings, 2 replies; 11+ messages in thread
From: Tim Chen @ 2026-09-10 17:46 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar
Cc: Tim Chen, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
alb_break_llc() decides whether to break LLC preference during active
load balance. It does so by testing that every runnable fair task on the
source rq prefers its LLC:
env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
But the two counters cover different sets. nr_pref_llc_running is updated
in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
clear_delayed() and drops delay-dequeued tasks.
So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
alb_break_llc() returns false, and active balance is free to pull a task
off its preferred LLC. Active balance only moves runnable tasks, and this
is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
skips the per-task test in can_migrate_task(). The runnable set is the one
we want.
Fix it on the counter side. A task should be counted in
nr_pref_llc_running exactly while it is both queued on its preferred LLC
(pref_llc_queued) and runnable (!sched_delayed). Define that membership
once in task_pref_llc_runnable(), and adjust the counter only through
pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
change either input: account_llc_enqueue(), account_llc_dequeue(),
set_delayed() and clear_delayed(). Gating every update on the same
predicate keeps the delay, wake and dequeue paths from double-counting
or underflowing; see the comments at those sites for the ordering.
nr_llc_running and sd->llc_counts are not touched and stay on queued
semantics.
Reported-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/
Suggested-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
kernel/sched/fair.c | 52 +++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 50 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index d5989b53adef..b1ef013b0342 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1551,6 +1551,28 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
(scale * per_cpu(sd_llc_size, cpu)));
}
+/*
+ * A task counts in nr_pref_llc_running while it is queued on its preferred
+ * LLC (pref_llc_queued) and runnable (!sched_delayed), keeping the counter in
+ * the runnable domain so alb_break_llc() can compare it with h_nr_runnable.
+ */
+static bool task_pref_llc_runnable(struct task_struct *p)
+{
+ return p->pref_llc_queued && !p->se.sched_delayed;
+}
+
+static void pref_llc_running_inc(struct rq *rq, struct task_struct *p)
+{
+ if (task_pref_llc_runnable(p))
+ rq->nr_pref_llc_running++;
+}
+
+static void pref_llc_running_dec(struct rq *rq, struct task_struct *p)
+{
+ if (task_pref_llc_runnable(p))
+ rq->nr_pref_llc_running--;
+}
+
static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
{
int pref_llc, pref_llc_queued;
@@ -1562,7 +1584,6 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
pref_llc_queued = (pref_llc == task_llc(p));
rq->nr_llc_running++;
- rq->nr_pref_llc_running += pref_llc_queued;
/*
* Record whether p is enqueued on its preferred
@@ -1580,6 +1601,9 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p)
*/
p->pref_llc_queued = pref_llc_queued;
+ /* Skipped while delayed; clear_delayed() adds it back on wake. */
+ pref_llc_running_inc(rq, p);
+
sd = rcu_dereference_all(rq->sd);
if (sd && (unsigned int)pref_llc < sd->llc_max)
sd->llc_counts[pref_llc]++;
@@ -1596,7 +1620,12 @@ static void account_llc_dequeue(struct rq *rq, struct task_struct *p)
rq->nr_llc_running--;
if (p->pref_llc_queued) {
- rq->nr_pref_llc_running--;
+ /*
+ * Skipped if still delayed (set_delayed() already removed it);
+ * clearing pref_llc_queued below also stops clear_delayed()
+ * from re-adding it.
+ */
+ pref_llc_running_dec(rq, p);
/*
* Update the status in case
* other logic might query
@@ -2021,6 +2050,10 @@ static void account_llc_enqueue(struct rq *rq, struct task_struct *p) {}
static void account_llc_dequeue(struct rq *rq, struct task_struct *p) {}
+static void pref_llc_running_inc(struct rq *rq, struct task_struct *p) {}
+
+static void pref_llc_running_dec(struct rq *rq, struct task_struct *p) {}
+
#endif /* CONFIG_SCHED_CACHE */
/*
@@ -6395,6 +6428,14 @@ static __always_inline void return_cfs_rq_runtime(struct cfs_rq *cfs_rq);
static void set_delayed(struct sched_entity *se)
{
+ /*
+ * Drop a task leaving the runnable set. Must run before sched_delayed
+ * is set, or task_pref_llc_runnable() would already exclude it;
+ * clear_delayed() mirrors this after clearing the flag.
+ */
+ if (entity_is_task(se))
+ pref_llc_running_dec(rq_of(cfs_rq_of(se)), task_of(se));
+
se->sched_delayed = 1;
/*
@@ -6425,6 +6466,13 @@ static void clear_delayed(struct sched_entity *se)
if (!entity_is_task(se))
return;
+ /*
+ * Re-add on wake, after sched_delayed is cleared. On a final delayed
+ * dequeue account_llc_dequeue() already cleared pref_llc_queued, so
+ * this does nothing.
+ */
+ pref_llc_running_inc(rq_of(cfs_rq_of(se)), task_of(se));
+
for_each_sched_entity(se) {
struct cfs_rq *cfs_rq = cfs_rq_of(se);
--
2.32.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH 2/4] sched/cache: Honor migrate_llc_task semantics in active load balance
2026-09-10 17:46 [PATCH 0/4] sched/cache: Fixes for cache aware scheduling Tim Chen
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
@ 2026-09-10 17:46 ` Tim Chen
2026-09-10 17:46 ` [PATCH 3/4] sched/cache: Decouple sched_cache_group from mm Tim Chen
2026-09-10 17:46 ` [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp Tim Chen
3 siblings, 0 replies; 11+ messages in thread
From: Tim Chen @ 2026-09-10 17:46 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar
Cc: Lu Wang, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng, Yi Lai,
Tim Chen, linux-kernel, linux-mm, linux-fsdevel
From: Lu Wang <wanglu.priv@gmail.com>
CAS introduced the migrate_llc_task migration type to direct tasks
toward their preferred LLC, but its semantics can be lost when passive
load balance falls back to active load balance. This may allow ALB to
select a candidate whose preferred LLC does not match the destination,
moving it away from its preferred LLC.
Example scenario:
src_rq has two runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc),
while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is
set because src_rq has at least one task, p1, that wants to migrate to
dst_rq. In ALB, can_migrate_task() finds p2 and returns true for it,
thus moving p2 out of its preferred LLC.
Solution:
The CPU stopper in ALB constructs a fresh lb_env that does not inherit
migration_type from the passive load-balance pass. Two approaches are
possible:
(a) Add a new member to struct rq so ALB can inherit migrate_llc_task
from the passive LB that triggered it.
(b) Define a new flag LBF_ACTIVE_LB_LLC and select the stopper callback
at kick time to preserve the migration semantics across the
asynchronous boundary.
We choose (b) because it avoids passing migration_type through the
stopper, which would affect the meaning of migration_type for
delayed-dequeue tasks.
Fixes: e4c9a4cb244a ("sched/cache: Add migrate_llc_task migration type for cache-aware balancing")
Suggested-by: "Chen, Yu C" <yu.c.chen@intel.com>
Reviewed-by: Tim Chen <tim.c.chen@linux.intel.com>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Lu Wang <wanglu.priv@gmail.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
kernel/sched/fair.c | 57 ++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 51 insertions(+), 6 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b1ef013b0342..32213801ea39 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10448,6 +10448,7 @@ enum migration_type {
#define LBF_SOME_PINNED 0x08
#define LBF_ACTIVE_LB 0x10
#define LBF_LLC_PINNED 0x20
+#define LBF_ACTIVE_LB_LLC 0x40
struct lb_env {
struct sched_domain *sd;
@@ -10866,6 +10867,21 @@ alb_break_llc(struct lb_env *env)
return false;
}
+/*
+ * Returns true if p's preferred LLC does not match the destination CPU
+ * under migrate_llc_task semantics. Passive LB passes migrate_llc_task
+ * in env->migration_type, while active LB carries LBF_ACTIVE_LB_LLC in
+ * env->flags to avoid overwriting env->migration_type.
+ */
+static inline bool
+migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env)
+{
+ return sched_cache_enabled() &&
+ (env->migration_type == migrate_llc_task ||
+ env->flags & LBF_ACTIVE_LB_LLC) &&
+ READ_ONCE(p->preferred_llc) != llc_id(env->dst_cpu);
+}
+
/*
* Check if migrating task p from env->src_cpu to
* env->dst_cpu breaks LLC localiy.
@@ -10894,8 +10910,7 @@ static bool migrate_degrades_llc(struct task_struct *p, struct lb_env *env)
* run on env->dst_cpu, skip the tasks do not prefer
* env->dst_cpu, and find the one that prefers.
*/
- if (env->migration_type == migrate_llc_task &&
- READ_ONCE(p->preferred_llc) != llc_id(env->dst_cpu))
+ if (migrate_llc_task_wrong_dst(p, env))
return true;
if (can_migrate_llc_task(env, p) != mig_forbid)
@@ -10917,6 +10932,12 @@ alb_break_llc(struct lb_env *env)
return false;
}
+static inline bool
+migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env)
+{
+ return false;
+}
+
static inline bool
migrate_degrades_llc(struct task_struct *p, struct lb_env *env)
{
@@ -11016,7 +11037,7 @@ int can_migrate_task(struct task_struct *p, struct lb_env *env)
* 4) too many balance attempts have failed.
*/
if (env->flags & LBF_ACTIVE_LB)
- return 1;
+ return !migrate_llc_task_wrong_dst(p, env);
degrades = migrate_degrades_locality(p, env);
if (!degrades) {
@@ -13415,6 +13436,20 @@ static int need_active_balance(struct lb_env *env)
}
static int active_load_balance_cpu_stop(void *data);
+static int active_load_balance_llc_cpu_stop(void *data);
+
+/*
+ * migration_type is checked elsewhere to decide migration policy, so
+ * it shouldn't be repurposed just to flag an LLC-directed active
+ * balance across the stopper. Pick the callback here instead.
+ */
+static inline cpu_stop_fn_t alb_stop_fn(struct lb_env *env)
+{
+ if (env->migration_type == migrate_llc_task)
+ return active_load_balance_llc_cpu_stop;
+
+ return active_load_balance_cpu_stop;
+}
static int should_we_balance(struct lb_env *env)
{
@@ -13760,7 +13795,7 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
}
if (active_balance) {
stop_one_cpu_nowait(cpu_of(busiest),
- active_load_balance_cpu_stop, busiest,
+ alb_stop_fn(&env), busiest,
&busiest->active_balance_work);
}
preempt_enable();
@@ -13865,7 +13900,7 @@ update_next_balance(struct sched_domain *sd, unsigned long *next_balance)
* least 1 task to be running on each physical CPU where possible, and
* avoids physical / logical imbalances.
*/
-static int active_load_balance_cpu_stop(void *data)
+static int __active_load_balance_cpu_stop(void *data, unsigned int lb_flags)
{
struct rq *busiest_rq = data;
int busiest_cpu = cpu_of(busiest_rq);
@@ -13915,7 +13950,7 @@ static int active_load_balance_cpu_stop(void *data)
.src_cpu = busiest_rq->cpu,
.src_rq = busiest_rq,
.idle = CPU_IDLE,
- .flags = LBF_ACTIVE_LB,
+ .flags = LBF_ACTIVE_LB | lb_flags,
};
schedstat_inc(sd->alb_count);
@@ -13943,6 +13978,16 @@ static int active_load_balance_cpu_stop(void *data)
return 0;
}
+static int active_load_balance_cpu_stop(void *data)
+{
+ return __active_load_balance_cpu_stop(data, 0);
+}
+
+static int active_load_balance_llc_cpu_stop(void *data)
+{
+ return __active_load_balance_cpu_stop(data, LBF_ACTIVE_LB_LLC);
+}
+
/*
* Scale the max sched_balance_rq interval with the number of CPUs in the system.
* This trades load-balance latency on larger machines for less cross talk.
--
2.32.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH 3/4] sched/cache: Decouple sched_cache_group from mm
2026-09-10 17:46 [PATCH 0/4] sched/cache: Fixes for cache aware scheduling Tim Chen
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
2026-09-10 17:46 ` [PATCH 2/4] sched/cache: Honor migrate_llc_task semantics in active load balance Tim Chen
@ 2026-09-10 17:46 ` Tim Chen
2026-09-10 17:46 ` [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp Tim Chen
3 siblings, 0 replies; 11+ messages in thread
From: Tim Chen @ 2026-09-10 17:46 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar
Cc: Tim Chen, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
Currently the sched cache grouping is by mm and the scheduling statistics
sched_cache_stat lives in the mm structure. This ties the life cycle
of scheduling stats with mm.
In account_mm_sched(), the scheduling stats are accessed by
task->mm->sc_stat. However, a task may be switching mm on one CPU when
another CPU is running account_mm_sched(), and possibly accessing the
old mm that was freed. This problem was found when running tests with
KASAN by Hyunwoo. https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Instead of serializing the mm access by introducing extra acquisition of
rq lock in the mm free path, extract sched_cache_stat from mm_struct,
rename it as sched_cache_group and manage its life cycle apart from
mm_struct with its own ref counting. This allows us in the next patch
access sched_cache_group directly from task, and add a refcount
on sched_cache_group when a task links to it. This prevents the use
after free issue when accessing stale and released old mm and its
sched cache stat a task switches to a new mm while account_mm_sched()
is done elsewhere.
The other benefit of this restructure is in the future, the grouping of
tasks to a LLC would have the flexibility to be associated with a user
defined grouping, or cgroup, cookie group, numa_group or others instead
of just with a single mm address space.
Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
object allocated from mm_struct. The mm_struct now holds a pointer
(sched_cache_grp) to this object instead of embedding it.
Introduce kernel/sched/cache_sched.c to host the cache aware scheduling
helpers and define sched_cache_group_put() there.
Meanwhile skip kthreads in account_mm_sched(), consistent with
task_tick_cache().
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Tested-by: Hyunwoo Kim <imv4bel@gmail.com>
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
include/linux/mm_types.h | 15 ++---
include/linux/sched.h | 8 ++-
kernel/exit.c | 6 +-
kernel/sched/build_utility.c | 4 ++
kernel/sched/cache_sched.c | 20 ++++++
kernel/sched/fair.c | 118 ++++++++++++++++++++++-------------
6 files changed, 113 insertions(+), 58 deletions(-)
create mode 100644 kernel/sched/cache_sched.c
diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
index 6d815f6440c9..f3e5a2fadbe5 100644
--- a/include/linux/mm_types.h
+++ b/include/linux/mm_types.h
@@ -1226,7 +1226,7 @@ struct mm_struct {
struct mm_mm_cid mm_cid;
/* sched_cache related statistics */
- struct sched_cache_stat sc_stat;
+ struct sched_cache_group *sched_cache_grp;
#ifdef CONFIG_MMU
atomic_long_t pgtables_bytes; /* size of all page tables */
#endif
@@ -1624,8 +1624,9 @@ static inline unsigned int mm_cid_size(void)
#endif /* CONFIG_SCHED_MM_CID */
#ifdef CONFIG_SCHED_CACHE
-void mm_init_sched(struct mm_struct *mm,
- struct sched_cache_time __percpu *pcpu_sched);
+int mm_init_sched(struct mm_struct *mm,
+ struct sched_cache_time __percpu *pcpu_sched);
+void mm_destroy_sched(struct mm_struct *mm);
static inline int mm_alloc_sched_noprof(struct mm_struct *mm)
{
@@ -1635,17 +1636,11 @@ static inline int mm_alloc_sched_noprof(struct mm_struct *mm)
if (!pcpu_sched)
return -ENOMEM;
- mm_init_sched(mm, pcpu_sched);
- return 0;
+ return mm_init_sched(mm, pcpu_sched);
}
#define mm_alloc_sched(...) alloc_hooks(mm_alloc_sched_noprof(__VA_ARGS__))
-static inline void mm_destroy_sched(struct mm_struct *mm)
-{
- free_percpu(mm->sc_stat.pcpu_sched);
- mm->sc_stat.pcpu_sched = NULL;
-}
#else /* !CONFIG_SCHED_CACHE */
static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; }
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 8b3d47a325cc..1f254364f216 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -2405,7 +2405,7 @@ struct sched_cache_time {
unsigned long epoch;
};
-struct sched_cache_stat {
+struct sched_cache_group {
struct sched_cache_time __percpu *pcpu_sched;
raw_spinlock_t lock;
unsigned long epoch;
@@ -2413,11 +2413,15 @@ struct sched_cache_stat {
unsigned long next_scan;
unsigned long footprint;
int cpu;
+ refcount_t refcnt;
+ struct rcu_head rcu;
} ____cacheline_aligned_in_smp;
+void sched_cache_group_put(struct sched_cache_group *grp);
+
#else
-struct sched_cache_stat { };
+struct sched_cache_group { };
#endif
diff --git a/kernel/exit.c b/kernel/exit.c
index 97686af89501..006edcc0c2c5 100644
--- a/kernel/exit.c
+++ b/kernel/exit.c
@@ -560,12 +560,12 @@ static void exit_mm_sched_cache(struct mm_struct *mm)
return;
/*
* No lock protection due to performance considerations.
- * Make sure mm->sc_stat.footprint does not become
+ * Make sure the group footprint does not become
* negative.
*/
- fp = READ_ONCE(mm->sc_stat.footprint);
+ fp = READ_ONCE(mm->sched_cache_grp->footprint);
sub = min(fp, current->total_numa_faults);
- WRITE_ONCE(mm->sc_stat.footprint, fp - sub);
+ WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub);
}
#else
static inline void exit_mm_sched_cache(struct mm_struct *mm)
diff --git a/kernel/sched/build_utility.c b/kernel/sched/build_utility.c
index e2cf3b08d4e9..24202893b262 100644
--- a/kernel/sched/build_utility.c
+++ b/kernel/sched/build_utility.c
@@ -89,6 +89,10 @@
# include "core_sched.c"
#endif
+#ifdef CONFIG_SCHED_CACHE
+# include "cache_sched.c"
+#endif
+
#ifdef CONFIG_PSI
# include "psi.c"
#endif
diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c
new file mode 100644
index 000000000000..d492df55f9d5
--- /dev/null
+++ b/kernel/sched/cache_sched.c
@@ -0,0 +1,20 @@
+// SPDX-License-Identifier: GPL-2.0-only
+#include "sched.h"
+
+static void sched_cache_group_free_rcu(struct rcu_head *rcu)
+{
+ struct sched_cache_group *grp =
+ container_of(rcu, struct sched_cache_group, rcu);
+
+ /* free_percpu() may be called from atomic context. */
+ free_percpu(grp->pcpu_sched);
+ kfree(grp);
+}
+
+void sched_cache_group_put(struct sched_cache_group *grp)
+{
+ if (!grp || !refcount_dec_and_test(&grp->refcnt))
+ return;
+
+ call_rcu(&grp->rcu, sched_cache_group_free_rcu);
+}
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 32213801ea39..b5a823f0a622 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1502,7 +1502,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
* excluded.
*/
llc = sd->llc_bytes;
- footprint = READ_ONCE(mm->sc_stat.footprint);
+ footprint = READ_ONCE(mm->sched_cache_grp->footprint);
/*
* Scale the LLC size by 256*llc_aggr_tolerance
@@ -1547,7 +1547,7 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
if (scale == INT_MAX)
return false;
- return !fits_capacity((mm->sc_stat.nr_running_avg * cpu_smt_num_threads),
+ return !fits_capacity((mm->sched_cache_grp->nr_running_avg * cpu_smt_num_threads),
(scale * per_cpu(sd_llc_size, cpu)));
}
@@ -1653,12 +1653,20 @@ static void account_llc_dequeue(struct rq *rq, struct task_struct *p)
}
}
-void mm_init_sched(struct mm_struct *mm,
- struct sched_cache_time __percpu *_pcpu_sched)
+int mm_init_sched(struct mm_struct *mm,
+ struct sched_cache_time __percpu *_pcpu_sched)
{
+ struct sched_cache_group *grp;
unsigned long epoch = 0;
int i;
+ grp = kzalloc_obj(*grp);
+ if (!grp) {
+ free_percpu(_pcpu_sched);
+ mm->sched_cache_grp = NULL;
+ return -ENOMEM;
+ }
+
for_each_possible_cpu(i) {
struct sched_cache_time *pcpu_sched = per_cpu_ptr(_pcpu_sched, i);
struct rq *rq = cpu_rq(i);
@@ -1669,18 +1677,35 @@ void mm_init_sched(struct mm_struct *mm,
epoch = rq->cpu_epoch;
}
- raw_spin_lock_init(&mm->sc_stat.lock);
- mm->sc_stat.epoch = epoch;
- mm->sc_stat.cpu = -1;
- mm->sc_stat.next_scan = jiffies;
- mm->sc_stat.nr_running_avg = 0;
- mm->sc_stat.footprint = 0;
+ raw_spin_lock_init(&grp->lock);
+ grp->epoch = epoch;
+ grp->cpu = -1;
+ grp->next_scan = jiffies;
+ grp->nr_running_avg = 0;
+ grp->footprint = 0;
+ refcount_set(&grp->refcnt, 1);
/*
- * The update to mm->sc_stat should not be reordered
- * before initialization to mm's other fields, in case
+ * The update to grp->pcpu_sched should not be reordered
+ * before initialization to grp's other fields, in case
* the readers may get invalid mm_sched_epoch, etc.
*/
- smp_store_release(&mm->sc_stat.pcpu_sched, _pcpu_sched);
+ smp_store_release(&grp->pcpu_sched, _pcpu_sched);
+ /*
+ * Publish the group last. Not every reader qualifies it by
+ * grp->pcpu_sched - can_migrate_llc_task() only checks that the
+ * pointer is non-NULL before reading grp->footprint and
+ * grp->nr_running_avg - so a reachable group must already be
+ * fully initialized.
+ */
+ mm->sched_cache_grp = grp;
+ return 0;
+}
+
+void mm_destroy_sched(struct mm_struct *mm)
+{
+ if (mm->sched_cache_grp)
+ sched_cache_group_put(mm->sched_cache_grp);
+ mm->sched_cache_grp = NULL;
}
/* because why would C be fully specified */
@@ -1738,7 +1763,7 @@ static int get_pref_llc(struct task_struct *p, struct mm_struct *mm)
if (!mm)
return -1;
- mm_sched_cpu = READ_ONCE(mm->sc_stat.cpu);
+ mm_sched_cpu = READ_ONCE(mm->sched_cache_grp->cpu);
if (mm_sched_cpu != -1) {
mm_sched_llc = llc_id(mm_sched_cpu);
@@ -1781,11 +1806,15 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
/*
* init_task, kthreads and user thread created
* by user_mode_thread() don't have mm.
+ *
+ * A kthread can temporarily adopt an mm via kthread_use_mm(),
+ * so p->mm alone does not imply a user task.
*/
- if (!mm || !mm->sc_stat.pcpu_sched)
+ if (!mm || p->flags & PF_KTHREAD || !mm->sched_cache_grp ||
+ !mm->sched_cache_grp->pcpu_sched)
return;
- pcpu_sched = per_cpu_ptr(mm->sc_stat.pcpu_sched, cpu_of(rq));
+ pcpu_sched = per_cpu_ptr(mm->sched_cache_grp->pcpu_sched, cpu_of(rq));
scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) {
__update_mm_sched(rq, pcpu_sched);
@@ -1798,11 +1827,11 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
* If this process hasn't hit task_cache_work() for a while invalidate
* its preferred state.
*/
- if ((long)(epoch - READ_ONCE(mm->sc_stat.epoch)) > llc_epoch_affinity_timeout ||
+ if ((long)(epoch - READ_ONCE(mm->sched_cache_grp->epoch)) > llc_epoch_affinity_timeout ||
invalid_llc_nr(mm, p, cpu_of(rq)) ||
exceed_llc_capacity(mm, cpu_of(rq))) {
- if (READ_ONCE(mm->sc_stat.cpu) != -1)
- WRITE_ONCE(mm->sc_stat.cpu, -1);
+ if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
+ WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
}
mm_sched_llc = get_pref_llc(p, mm);
@@ -1826,19 +1855,19 @@ static void task_tick_cache(struct rq *rq, struct task_struct *p)
return;
if (!mm || p->flags & PF_KTHREAD ||
- !mm->sc_stat.pcpu_sched)
+ !mm->sched_cache_grp->pcpu_sched)
return;
epoch = rq->cpu_epoch;
/* avoid moving backwards */
- if (time_after_eq(mm->sc_stat.epoch, epoch))
+ if (time_after_eq(mm->sched_cache_grp->epoch, epoch))
return;
- guard(raw_spinlock)(&mm->sc_stat.lock);
+ guard(raw_spinlock)(&mm->sched_cache_grp->lock);
if (work->next == work) {
task_work_add(p, work, TWA_RESUME);
- WRITE_ONCE(mm->sc_stat.epoch, epoch);
+ WRITE_ONCE(mm->sched_cache_grp->epoch, epoch);
}
}
@@ -1850,7 +1879,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p)
if (!static_branch_likely(&sched_numa_balancing))
goto out;
- cpu = READ_ONCE(p->mm->sc_stat.cpu);
+ cpu = READ_ONCE(p->mm->sched_cache_grp->cpu);
if (cpu != -1)
nid = cpu_to_node(cpu);
curr_cpu = task_cpu(p);
@@ -1922,12 +1951,12 @@ static void task_cache_work(struct callback_head *work)
if (p->flags & PF_EXITING)
return;
- next_scan = READ_ONCE(mm->sc_stat.next_scan);
+ next_scan = READ_ONCE(mm->sched_cache_grp->next_scan);
if (time_before(now, next_scan))
return;
/* only 1 thread is allowed to scan */
- if (!try_cmpxchg(&mm->sc_stat.next_scan, &next_scan,
+ if (!try_cmpxchg(&mm->sched_cache_grp->next_scan, &next_scan,
now + max_t(unsigned long,
READ_ONCE(llc_epoch_period), 1)))
return;
@@ -1935,8 +1964,8 @@ static void task_cache_work(struct callback_head *work)
curr_cpu = task_cpu(p);
if (invalid_llc_nr(mm, p, curr_cpu) ||
exceed_llc_capacity(mm, curr_cpu)) {
- if (READ_ONCE(mm->sc_stat.cpu) != -1)
- WRITE_ONCE(mm->sc_stat.cpu, -1);
+ if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
+ WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
return;
}
@@ -1959,8 +1988,10 @@ static void task_cache_work(struct callback_head *work)
continue;
for_each_cpu(i, sched_domain_span(sd)) {
+ struct sched_cache_group *grp = mm->sched_cache_grp;
+
occ = fraction_mm_sched(cpu_rq(i),
- per_cpu_ptr(mm->sc_stat.pcpu_sched, i));
+ per_cpu_ptr(grp->pcpu_sched, i));
a_occ += occ;
if (occ > m_occ) {
m_occ = occ;
@@ -1993,7 +2024,7 @@ static void task_cache_work(struct callback_head *work)
m_a_cpu = m_cpu;
}
- if (llc_id(cpu) == llc_id(READ_ONCE(mm->sc_stat.cpu)))
+ if (llc_id(cpu) == llc_id(READ_ONCE(mm->sched_cache_grp->cpu)))
curr_m_a_occ = a_occ;
cpumask_andnot(cpus, cpus, sched_domain_span(sd));
@@ -2002,7 +2033,7 @@ static void task_cache_work(struct callback_head *work)
if (m_a_occ > (2 * curr_m_a_occ)) {
/*
- * Avoid switching sc_stat.cpu too fast.
+ * Avoid switching sched_cache_grp->cpu too fast.
* The reason to choose 2X is because:
* 1. It is better to keep the preferred LLC stable,
* rather than changing it frequently and cause migrations
@@ -2011,10 +2042,10 @@ static void task_cache_work(struct callback_head *work)
* 3. 2X is chosen based on test results, as it delivers
* the optimal performance gain so far.
*/
- WRITE_ONCE(mm->sc_stat.cpu, m_a_cpu);
+ WRITE_ONCE(mm->sched_cache_grp->cpu, m_a_cpu);
}
- update_avg_scale(&mm->sc_stat.nr_running_avg, nr_running);
+ update_avg_scale(&mm->sched_cache_grp->nr_running_avg, nr_running);
free_cpumask_var(cpus);
}
@@ -3822,18 +3853,19 @@ static void task_numa_placement(struct task_struct *p)
* heuristic and occasional lost updates are tolerable.
*
* If a task exits, its corresponding footprint must
- * be subtracted from the mm->sc_stat.footprint, otherwise
- * the mm->sc_stat.footprint will not converge:
- * the exiting thread's footprint remains unchanged/undecayed
- * in mm->sc_stat.footprint. See exit_mm().
+ * be subtracted from the mm->sched_cache_grp->footprint,
+ * otherwise the mm->sched_cache_grp->footprint will not
+ * converge: the exiting thread's footprint remains
+ * unchanged/undecayed in mm->sched_cache_grp->footprint.
+ * See exit_mm().
*
* Lost updates and unsynchronized subtraction
* in exit_mm() can cause footprint + diff to
* go negative. Clamp to zero to prevent the
* unsigned footprint from wrapping.
*/
- new_fp = (long)READ_ONCE(p->mm->sc_stat.footprint) + diff;
- WRITE_ONCE(p->mm->sc_stat.footprint,
+ new_fp = (long)READ_ONCE(p->mm->sched_cache_grp->footprint) + diff;
+ WRITE_ONCE(p->mm->sched_cache_grp->footprint,
max(new_fp, 0L));
#endif
}
@@ -10788,18 +10820,18 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
src_cpu = env->src_cpu;
dst_cpu = env->dst_cpu;
mm = p->mm;
- if (!mm)
+ if (!mm || !mm->sched_cache_grp)
return mig_unrestricted;
- cpu = READ_ONCE(mm->sc_stat.cpu);
+ cpu = READ_ONCE(mm->sched_cache_grp->cpu);
if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
return mig_unrestricted;
/* skip cache aware load balance for too many threads */
if (invalid_llc_nr(mm, p, dst_cpu) ||
exceed_llc_capacity(mm, dst_cpu)) {
- if (READ_ONCE(mm->sc_stat.cpu) != -1)
- WRITE_ONCE(mm->sc_stat.cpu, -1);
+ if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
+ WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
return mig_unrestricted;
}
--
2.32.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp
2026-09-10 17:46 [PATCH 0/4] sched/cache: Fixes for cache aware scheduling Tim Chen
` (2 preceding siblings ...)
2026-09-10 17:46 ` [PATCH 3/4] sched/cache: Decouple sched_cache_group from mm Tim Chen
@ 2026-09-10 17:46 ` Tim Chen
2026-09-10 19:19 ` Peter Zijlstra
3 siblings, 1 reply; 11+ messages in thread
From: Tim Chen @ 2026-09-10 17:46 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar
Cc: Tim Chen, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
Add a sched_cache_grp pointer to task_struct so that scheduler code
can access the cache group directly via the task, without going
through mm->sched_cache_grp. This decouples the scheduler's hot-path
accesses from the mm_struct.
Each task holds its own refcount on the sched_cache_group, separate
from the reference held by its mm_struct. The reference is acquired
in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm().
This fixes use after free problem when accessing sched_cache_grp
in account_mm_sched() via mm as reported in
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Convert all scheduler code in fair.c and exit.c to use
p->sched_cache_grp instead of p->mm->sched_cache_grp.
Add sched_cache_group_get() to kernel/sched/cache_sched.c.
Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
Tested-by: Hyunwoo Kim <imv4bel@gmail.com>
Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
Co-developed-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
fs/exec.c | 14 ++++
include/linux/sched.h | 3 +
kernel/exit.c | 26 +++++--
kernel/fork.c | 23 ++++++
kernel/sched/cache_sched.c | 19 +++++
kernel/sched/fair.c | 142 +++++++++++++++++++++----------------
kernel/sched/sched.h | 3 +
7 files changed, 164 insertions(+), 66 deletions(-)
diff --git a/fs/exec.c b/fs/exec.c
index 745f6eb5279e..7a8a9954343e 100644
--- a/fs/exec.c
+++ b/fs/exec.c
@@ -882,6 +882,20 @@ static int exec_mmap(struct linux_binprm *bprm)
active_mm = tsk->active_mm;
tsk->active_mm = mm;
tsk->mm = mm;
+#ifdef CONFIG_SCHED_CACHE
+ {
+ struct sched_cache_group *old_grp, *new_grp;
+
+ old_grp = rcu_dereference_protected(tsk->sched_cache_grp, true);
+
+ /* Acquire the reference before publishing the pointer. */
+ new_grp = sched_cache_group_get(mm->sched_cache_grp);
+
+ rcu_assign_pointer(tsk->sched_cache_grp, new_grp);
+ if (old_grp)
+ sched_cache_group_put(old_grp);
+ }
+#endif
mm_init_cid(mm, tsk);
exec_state = task_exec_state_replace(tsk, exec_state);
/*
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 1f254364f216..cab8e89b1462 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1434,6 +1434,7 @@ struct task_struct {
#ifdef CONFIG_SCHED_CACHE
struct callback_head cache_work;
int preferred_llc;
+ struct sched_cache_group __rcu *sched_cache_grp;
/* 1: task was enqueued to its preferred LLC, 0 otherwise */
int pref_llc_queued;
#endif
@@ -2418,6 +2419,8 @@ struct sched_cache_group {
} ____cacheline_aligned_in_smp;
void sched_cache_group_put(struct sched_cache_group *grp);
+struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *grp);
+struct sched_cache_group *task_cache_group_get(struct task_struct *p);
#else
diff --git a/kernel/exit.c b/kernel/exit.c
index 006edcc0c2c5..442535778ce1 100644
--- a/kernel/exit.c
+++ b/kernel/exit.c
@@ -552,23 +552,25 @@ void mm_update_next_owner(struct mm_struct *mm)
* Subtract the memory footprint of the current task from
* mm.
*/
-static void exit_mm_sched_cache(struct mm_struct *mm)
+static void exit_mm_sched_cache(void)
{
+ struct sched_cache_group *grp =
+ rcu_dereference_protected(current->sched_cache_grp, true);
unsigned long fp, sub;
- if (!current->total_numa_faults)
+ if (!grp || !current->total_numa_faults)
return;
/*
* No lock protection due to performance considerations.
* Make sure the group footprint does not become
* negative.
*/
- fp = READ_ONCE(mm->sched_cache_grp->footprint);
+ fp = READ_ONCE(grp->footprint);
sub = min(fp, current->total_numa_faults);
- WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub);
+ WRITE_ONCE(grp->footprint, fp - sub);
}
#else
-static inline void exit_mm_sched_cache(struct mm_struct *mm)
+static inline void exit_mm_sched_cache(void)
{
}
#endif /* CONFIG_SCHED_CACHE CONFIG_NUMA_BALANCING */
@@ -585,7 +587,19 @@ static void exit_mm(void)
if (!mm)
return;
- exit_mm_sched_cache(mm);
+ exit_mm_sched_cache();
+
+#ifdef CONFIG_SCHED_CACHE
+ {
+ struct sched_cache_group *grp =
+ rcu_dereference_protected(current->sched_cache_grp, true);
+
+ rcu_assign_pointer(current->sched_cache_grp, NULL);
+
+ if (grp)
+ sched_cache_group_put(grp);
+ }
+#endif
mmap_read_lock(mm);
mmgrab_lazy_tlb(mm);
diff --git a/kernel/fork.c b/kernel/fork.c
index 416758c8a3d4..2e79548cb7c1 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -1599,6 +1599,19 @@ static int copy_mm(u64 clone_flags, struct task_struct *tsk)
tsk->mm = mm;
tsk->active_mm = mm;
+#ifdef CONFIG_SCHED_CACHE
+ {
+ /*
+ * A task holds its own reference on the group, separate from
+ * the reference held by its mm_struct. Acquire it before
+ * publishing the pointer.
+ */
+ struct sched_cache_group *grp =
+ sched_cache_group_get(mm->sched_cache_grp);
+
+ rcu_assign_pointer(tsk->sched_cache_grp, grp);
+ }
+#endif
return 0;
}
@@ -2599,6 +2612,16 @@ __latent_entropy struct task_struct *copy_process(
bad_fork_cleanup_namespaces:
exit_nsproxy_namespaces(p);
bad_fork_cleanup_mm:
+#ifdef CONFIG_SCHED_CACHE
+ /*
+ * copy_mm() took a task reference on the cache group; a failed fork
+ * never reaches exit_mm(), so release it here to avoid leaking the
+ * group and its per-CPU buffer.
+ */
+ sched_cache_group_put(rcu_dereference_protected(p->sched_cache_grp, true));
+ RCU_INIT_POINTER(p->sched_cache_grp, NULL);
+#endif
+
if (p->mm) {
mm_clear_owner(p->mm, p);
mmput(p->mm);
diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c
index d492df55f9d5..99d07e1e067c 100644
--- a/kernel/sched/cache_sched.c
+++ b/kernel/sched/cache_sched.c
@@ -1,6 +1,25 @@
// SPDX-License-Identifier: GPL-2.0-only
#include "sched.h"
+struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *grp)
+{
+ /*
+ * refcount_inc_not_zero() is the acquire primitive for lockless
+ * (RCU) lookups; plain refcount_inc() would scribble the count if
+ * it already reached zero. Return NULL in that case.
+ */
+ if (grp && !refcount_inc_not_zero(&grp->refcnt))
+ grp = NULL;
+
+ return grp;
+}
+
+struct sched_cache_group *task_cache_group_get(struct task_struct *p)
+{
+ guard(rcu)();
+ return sched_cache_group_get(rcu_dereference(p->sched_cache_grp));
+}
+
static void sched_cache_group_free_rcu(struct rcu_head *rcu)
{
struct sched_cache_group *grp =
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b5a823f0a622..272dce2baf32 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1483,7 +1483,7 @@ static inline int get_sched_cache_scale(int mul)
return (1 + (tol - 1) * mul);
}
-static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
+static bool exceed_llc_capacity(struct sched_cache_group *grp, int cpu)
{
#ifdef CONFIG_NUMA_BALANCING
unsigned long llc, footprint;
@@ -1502,7 +1502,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
* excluded.
*/
llc = sd->llc_bytes;
- footprint = READ_ONCE(mm->sched_cache_grp->footprint);
+ footprint = READ_ONCE(grp->footprint);
/*
* Scale the LLC size by 256*llc_aggr_tolerance
@@ -1531,7 +1531,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
return false;
}
-static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
+static bool invalid_llc_nr(struct sched_cache_group *grp, struct task_struct *p,
int cpu)
{
int scale;
@@ -1547,7 +1547,7 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p,
if (scale == INT_MAX)
return false;
- return !fits_capacity((mm->sched_cache_grp->nr_running_avg * cpu_smt_num_threads),
+ return !fits_capacity((grp->nr_running_avg * cpu_smt_num_threads),
(scale * per_cpu(sd_llc_size, cpu)));
}
@@ -1756,14 +1756,14 @@ static unsigned long fraction_mm_sched(struct rq *rq,
return div64_u64(NICE_0_LOAD * pcpu_sched->runtime, rq->cpu_runtime + 1);
}
-static int get_pref_llc(struct task_struct *p, struct mm_struct *mm)
+static int get_pref_llc(struct task_struct *p, struct sched_cache_group *grp)
{
int mm_sched_llc = -1, mm_sched_cpu;
- if (!mm)
+ if (!grp)
return -1;
- mm_sched_cpu = READ_ONCE(mm->sched_cache_grp->cpu);
+ mm_sched_cpu = READ_ONCE(grp->cpu);
if (mm_sched_cpu != -1) {
mm_sched_llc = llc_id(mm_sched_cpu);
@@ -1793,8 +1793,8 @@ static unsigned int task_running_on_cpu(int cpu, struct task_struct *p);
static inline
void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
{
+ struct sched_cache_group *grp = rcu_dereference_all(p->sched_cache_grp);
struct sched_cache_time *pcpu_sched;
- struct mm_struct *mm = p->mm;
int mm_sched_llc = -1;
unsigned long epoch;
@@ -1805,16 +1805,12 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
return;
/*
* init_task, kthreads and user thread created
- * by user_mode_thread() don't have mm.
- *
- * A kthread can temporarily adopt an mm via kthread_use_mm(),
- * so p->mm alone does not imply a user task.
+ * by user_mode_thread() don't have a cache group.
*/
- if (!mm || p->flags & PF_KTHREAD || !mm->sched_cache_grp ||
- !mm->sched_cache_grp->pcpu_sched)
+ if (!grp || p->flags & PF_KTHREAD || !grp->pcpu_sched)
return;
- pcpu_sched = per_cpu_ptr(mm->sched_cache_grp->pcpu_sched, cpu_of(rq));
+ pcpu_sched = per_cpu_ptr(grp->pcpu_sched, cpu_of(rq));
scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) {
__update_mm_sched(rq, pcpu_sched);
@@ -1827,14 +1823,14 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
* If this process hasn't hit task_cache_work() for a while invalidate
* its preferred state.
*/
- if ((long)(epoch - READ_ONCE(mm->sched_cache_grp->epoch)) > llc_epoch_affinity_timeout ||
- invalid_llc_nr(mm, p, cpu_of(rq)) ||
- exceed_llc_capacity(mm, cpu_of(rq))) {
- if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
- WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
+ if ((long)(epoch - READ_ONCE(grp->epoch)) > llc_epoch_affinity_timeout ||
+ invalid_llc_nr(grp, p, cpu_of(rq)) ||
+ exceed_llc_capacity(grp, cpu_of(rq))) {
+ if (READ_ONCE(grp->cpu) != -1)
+ WRITE_ONCE(grp->cpu, -1);
}
- mm_sched_llc = get_pref_llc(p, mm);
+ mm_sched_llc = get_pref_llc(p, grp);
/* task not on rq accounted later in account_entity_enqueue() */
if (task_running_on_cpu(rq->cpu, p) &&
@@ -1847,31 +1843,32 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
static void task_tick_cache(struct rq *rq, struct task_struct *p)
{
+ struct sched_cache_group *grp = rcu_dereference_all(p->sched_cache_grp);
struct callback_head *work = &p->cache_work;
- struct mm_struct *mm = p->mm;
unsigned long epoch;
if (!sched_cache_enabled())
return;
- if (!mm || p->flags & PF_KTHREAD ||
- !mm->sched_cache_grp->pcpu_sched)
+ if (!grp || p->flags & PF_KTHREAD ||
+ !grp->pcpu_sched)
return;
epoch = rq->cpu_epoch;
/* avoid moving backwards */
- if (time_after_eq(mm->sched_cache_grp->epoch, epoch))
+ if (time_after_eq(grp->epoch, epoch))
return;
- guard(raw_spinlock)(&mm->sched_cache_grp->lock);
+ guard(raw_spinlock)(&grp->lock);
if (work->next == work) {
task_work_add(p, work, TWA_RESUME);
- WRITE_ONCE(mm->sched_cache_grp->epoch, epoch);
+ WRITE_ONCE(grp->epoch, epoch);
}
}
-static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p)
+static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p,
+ struct sched_cache_group *grp)
{
#ifdef CONFIG_NUMA_BALANCING
int cpu, curr_cpu, nid, pref_nid;
@@ -1879,7 +1876,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p)
if (!static_branch_likely(&sched_numa_balancing))
goto out;
- cpu = READ_ONCE(p->mm->sched_cache_grp->cpu);
+ cpu = READ_ONCE(grp->cpu);
if (cpu != -1)
nid = cpu_to_node(cpu);
curr_cpu = task_cpu(p);
@@ -1940,9 +1937,7 @@ static void task_cache_work(struct callback_head *work)
unsigned long next_scan, now = jiffies;
struct task_struct *p = current, *cur;
unsigned long curr_m_a_occ = 0;
- struct mm_struct *mm = p->mm;
unsigned long m_a_occ = 0;
- cpumask_var_t cpus;
WARN_ON_ONCE(work != &p->cache_work);
@@ -1951,32 +1946,44 @@ static void task_cache_work(struct callback_head *work)
if (p->flags & PF_EXITING)
return;
- next_scan = READ_ONCE(mm->sched_cache_grp->next_scan);
+ /*
+ * A reference makes sure grp is not released by others. The rcu
+ * lock can not be held till after zalloc_cpumask_var() below,
+ * because the latter might sleep.
+ */
+ struct sched_cache_group *grp __free(sched_cache_group_put) =
+ task_cache_group_get(p);
+ if (!grp)
+ return;
+
+ next_scan = READ_ONCE(grp->next_scan);
if (time_before(now, next_scan))
return;
/* only 1 thread is allowed to scan */
- if (!try_cmpxchg(&mm->sched_cache_grp->next_scan, &next_scan,
+ if (!try_cmpxchg(&grp->next_scan, &next_scan,
now + max_t(unsigned long,
READ_ONCE(llc_epoch_period), 1)))
return;
curr_cpu = task_cpu(p);
- if (invalid_llc_nr(mm, p, curr_cpu) ||
- exceed_llc_capacity(mm, curr_cpu)) {
- if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
- WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
+ if (invalid_llc_nr(grp, p, curr_cpu) ||
+ exceed_llc_capacity(grp, curr_cpu)) {
+ if (READ_ONCE(grp->cpu) != -1)
+ WRITE_ONCE(grp->cpu, -1);
return;
}
+ cpumask_var_t cpus __free(free_cpumask_var) = CPUMASK_VAR_NULL;
+
if (!zalloc_cpumask_var(&cpus, GFP_KERNEL))
return;
scoped_guard (cpus_read_lock) {
guard(rcu)();
- get_scan_cpumasks(cpus, p);
+ get_scan_cpumasks(cpus, p, grp);
for_each_cpu(cpu, cpus) {
/* XXX sched_cluster_active */
@@ -1988,8 +1995,6 @@ static void task_cache_work(struct callback_head *work)
continue;
for_each_cpu(i, sched_domain_span(sd)) {
- struct sched_cache_group *grp = mm->sched_cache_grp;
-
occ = fraction_mm_sched(cpu_rq(i),
per_cpu_ptr(grp->pcpu_sched, i));
a_occ += occ;
@@ -1998,9 +2003,13 @@ static void task_cache_work(struct callback_head *work)
m_cpu = i;
}
+ /*
+ * rcu_access_pointer() is used because the
+ * pointer is only compared, never dereferenced.
+ */
cur = rcu_dereference_all(cpu_rq(i)->curr);
if (cur && !(cur->flags & (PF_EXITING | PF_KTHREAD)) &&
- cur->mm == mm)
+ rcu_access_pointer(cur->sched_cache_grp) == grp)
nr_running++;
}
@@ -2024,7 +2033,7 @@ static void task_cache_work(struct callback_head *work)
m_a_cpu = m_cpu;
}
- if (llc_id(cpu) == llc_id(READ_ONCE(mm->sched_cache_grp->cpu)))
+ if (llc_id(cpu) == llc_id(READ_ONCE(grp->cpu)))
curr_m_a_occ = a_occ;
cpumask_andnot(cpus, cpus, sched_domain_span(sd));
@@ -2042,11 +2051,10 @@ static void task_cache_work(struct callback_head *work)
* 3. 2X is chosen based on test results, as it delivers
* the optimal performance gain so far.
*/
- WRITE_ONCE(mm->sched_cache_grp->cpu, m_a_cpu);
+ WRITE_ONCE(grp->cpu, m_a_cpu);
}
- update_avg_scale(&mm->sched_cache_grp->nr_running_avg, nr_running);
- free_cpumask_var(cpus);
+ update_avg_scale(&grp->nr_running_avg, nr_running);
}
void init_sched_mm(struct task_struct *p)
@@ -2055,6 +2063,13 @@ void init_sched_mm(struct task_struct *p)
init_task_work(work, task_cache_work);
work->next = work;
+ /*
+ * dup_task_struct() copies the parent's task_struct, including its
+ * sched_cache_grp, for which the child holds no reference. Clear it
+ * here - before copy_mm() runs - so the child never carries a
+ * borrowed pointer that the fork error path would put.
+ */
+ RCU_INIT_POINTER(p->sched_cache_grp, NULL);
/*
* Reset new task's preference to avoid
* polluting account_llc_enqueue().
@@ -3853,10 +3868,9 @@ static void task_numa_placement(struct task_struct *p)
* heuristic and occasional lost updates are tolerable.
*
* If a task exits, its corresponding footprint must
- * be subtracted from the mm->sched_cache_grp->footprint,
- * otherwise the mm->sched_cache_grp->footprint will not
- * converge: the exiting thread's footprint remains
- * unchanged/undecayed in mm->sched_cache_grp->footprint.
+ * be subtracted from p->sched_cache_grp->footprint,
+ * otherwise the footprint will not converge: the
+ * exiting thread's footprint remains unchanged/undecayed.
* See exit_mm().
*
* Lost updates and unsynchronized subtraction
@@ -3864,9 +3878,17 @@ static void task_numa_placement(struct task_struct *p)
* go negative. Clamp to zero to prevent the
* unsigned footprint from wrapping.
*/
- new_fp = (long)READ_ONCE(p->mm->sched_cache_grp->footprint) + diff;
- WRITE_ONCE(p->mm->sched_cache_grp->footprint,
- max(new_fp, 0L));
+ {
+ struct sched_cache_group *grp;
+
+ guard(rcu)();
+ grp = rcu_dereference(p->sched_cache_grp);
+
+ if (grp) {
+ new_fp = (long)READ_ONCE(grp->footprint) + diff;
+ WRITE_ONCE(grp->footprint, max(new_fp, 0L));
+ }
+ }
#endif
}
@@ -10810,7 +10832,7 @@ static inline bool task_misfits_asym_cpu(struct lb_env *env, struct task_struct
static enum llc_mig can_migrate_llc_task(struct lb_env *env,
struct task_struct *p)
{
- struct mm_struct *mm;
+ struct sched_cache_group *grp;
bool to_pref;
int cpu, src_cpu, dst_cpu;
@@ -10819,19 +10841,19 @@ static enum llc_mig can_migrate_llc_task(struct lb_env *env,
src_cpu = env->src_cpu;
dst_cpu = env->dst_cpu;
- mm = p->mm;
- if (!mm || !mm->sched_cache_grp)
+ grp = rcu_dereference_all(p->sched_cache_grp);
+ if (!grp)
return mig_unrestricted;
- cpu = READ_ONCE(mm->sched_cache_grp->cpu);
+ cpu = READ_ONCE(grp->cpu);
if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
return mig_unrestricted;
/* skip cache aware load balance for too many threads */
- if (invalid_llc_nr(mm, p, dst_cpu) ||
- exceed_llc_capacity(mm, dst_cpu)) {
- if (READ_ONCE(mm->sched_cache_grp->cpu) != -1)
- WRITE_ONCE(mm->sched_cache_grp->cpu, -1);
+ if (invalid_llc_nr(grp, p, dst_cpu) ||
+ exceed_llc_capacity(grp, dst_cpu)) {
+ if (READ_ONCE(grp->cpu) != -1)
+ WRITE_ONCE(grp->cpu, -1);
return mig_unrestricted;
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..8b67af28a471 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -4145,6 +4145,9 @@ static inline bool sched_cache_enabled(void)
return static_branch_unlikely(&sched_cache_active);
}
+DEFINE_FREE(sched_cache_group_put, struct sched_cache_group *,
+ sched_cache_group_put(_T));
+
extern void sched_cache_active_set(void);
#endif
--
2.32.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
@ 2026-09-10 18:33 ` Kayra Cizmeci
2026-09-10 20:46 ` Tim Chen
2026-09-10 22:48 ` Kayra Cizmeci
1 sibling, 1 reply; 11+ messages in thread
From: Kayra Cizmeci @ 2026-09-10 18:33 UTC (permalink / raw)
To: tim.c.chen
Cc: brauner, bsegall, dietmar.eggemann, imv4bel, jack, juri.lelli,
kees, kprateek.nayak, linux-fsdevel, linux-kernel, linux-mm,
mgorman, mingo, peterz, qyousef, ricardo.neri-calderon, rostedt,
srikar, sshegde, vincent.guittot, vineethr, viro, vschneid,
wanglu.priv, yi1.lai, yu.c.chen, zhanxusheng1024, zhanxusheng,
ziqianlu
Hello :>,
> alb_break_llc() decides whether to break LLC preference during active
> load balance. It does so by testing that every runnable fair task on the
> source rq prefers its LLC:
>
> env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
>
> But the two counters cover different sets. nr_pref_llc_running is updated
> in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
> so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
> clear_delayed() and drops delay-dequeued tasks.
> So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
> in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
> alb_break_llc() returns false, and active balance is free to pull a task
> off its preferred LLC. Active balance only moves runnable tasks, and this
> is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
> skips the per-task test in can_migrate_task(). The runnable set is the one
> we want.
> Fix it on the counter side. A task should be counted in
> nr_pref_llc_running exactly while it is both queued on its preferred LLC
> (pref_llc_queued) and runnable (!sched_delayed). Define that membership
> once in task_pref_llc_runnable(), and adjust the counter only through
> pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
> change either input: account_llc_enqueue(), account_llc_dequeue(),
> set_delayed() and clear_delayed(). Gating every update on the same
> predicate keeps the delay, wake and dequeue paths from double-counting
> or underflowing; see the comments at those sites for the ordering.
> nr_llc_running and sd->llc_counts are not touched and stay on queued
> semantics.
I have one question tho, can't we combine the checks with h_nr_runnable? On the paper
if we are updating h_nr_runnable we could check if the nr_pref_llc_running can be
updated and update it if the condition is right. Because, every nr_pref_llc_running enters
h_nr_runnable while not every h_nr_runnable enters nr_pref_llc_running.
Why instead we just check the nr_pref_llc_running's conditions on task_pref_llc_runnable()
and call these dec and inc functions after the h_nr_runnable updates. Wouldn't it be clear that way?
If possible?
Thanks,
Kayra :_:
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp
2026-09-10 17:46 ` [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp Tim Chen
@ 2026-09-10 19:19 ` Peter Zijlstra
2026-09-10 22:50 ` Tim Chen
0 siblings, 1 reply; 11+ messages in thread
From: Peter Zijlstra @ 2026-09-10 19:19 UTC (permalink / raw)
To: Tim Chen
Cc: Ingo Molnar, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
On Thu, Sep 10, 2026 at 10:46:12AM -0700, Tim Chen wrote:
> Co-developed-by: Chen Yu <yu.c.chen@intel.com>
> Signed-off-by: Chen Yu <yu.c.chen@intel.com>
> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
:-(
> ---
> fs/exec.c | 14 ++++
> include/linux/sched.h | 3 +
> kernel/exit.c | 26 +++++--
> kernel/fork.c | 23 ++++++
> kernel/sched/cache_sched.c | 19 +++++
> kernel/sched/fair.c | 142 +++++++++++++++++++++----------------
> kernel/sched/sched.h | 3 +
> 7 files changed, 164 insertions(+), 66 deletions(-)
>
> diff --git a/fs/exec.c b/fs/exec.c
> index 745f6eb5279e..7a8a9954343e 100644
> --- a/fs/exec.c
> +++ b/fs/exec.c
> @@ -882,6 +882,20 @@ static int exec_mmap(struct linux_binprm *bprm)
> active_mm = tsk->active_mm;
> tsk->active_mm = mm;
> tsk->mm = mm;
> +#ifdef CONFIG_SCHED_CACHE
> + {
> + struct sched_cache_group *old_grp, *new_grp;
> +
> + old_grp = rcu_dereference_protected(tsk->sched_cache_grp, true);
> +
> + /* Acquire the reference before publishing the pointer. */
> + new_grp = sched_cache_group_get(mm->sched_cache_grp);
> +
> + rcu_assign_pointer(tsk->sched_cache_grp, new_grp);
> + if (old_grp)
> + sched_cache_group_put(old_grp);
> + }
> +#endif
Guys no! This is horrific crap. This is not how we do things and I would
have expected you all to know this.
Have you heard of this new fangled thing called a function?
Imagine all of those being just:
sched_cache_exec_mmap(tsk, mm);
Also: rcu_dereference_protected(.c = true) is another offence, that's
just wrong.
> diff --git a/kernel/exit.c b/kernel/exit.c
> index 006edcc0c2c5..442535778ce1 100644
> --- a/kernel/exit.c
> +++ b/kernel/exit.c
> @@ -552,23 +552,25 @@ void mm_update_next_owner(struct mm_struct *mm)
> * Subtract the memory footprint of the current task from
> * mm.
> */
> -static void exit_mm_sched_cache(struct mm_struct *mm)
> +static void exit_mm_sched_cache(void)
> {
> + struct sched_cache_group *grp =
> + rcu_dereference_protected(current->sched_cache_grp, true);
> unsigned long fp, sub;
>
> - if (!current->total_numa_faults)
> + if (!grp || !current->total_numa_faults)
> return;
> /*
> * No lock protection due to performance considerations.
> * Make sure the group footprint does not become
> * negative.
> */
> - fp = READ_ONCE(mm->sched_cache_grp->footprint);
> + fp = READ_ONCE(grp->footprint);
> sub = min(fp, current->total_numa_faults);
> - WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub);
> + WRITE_ONCE(grp->footprint, fp - sub);
> }
> #else
> -static inline void exit_mm_sched_cache(struct mm_struct *mm)
> +static inline void exit_mm_sched_cache(void)
> {
> }
> #endif /* CONFIG_SCHED_CACHE CONFIG_NUMA_BALANCING */
> @@ -585,7 +587,19 @@ static void exit_mm(void)
> if (!mm)
> return;
>
> - exit_mm_sched_cache(mm);
> + exit_mm_sched_cache();
> +
> +#ifdef CONFIG_SCHED_CACHE
> + {
> + struct sched_cache_group *grp =
> + rcu_dereference_protected(current->sched_cache_grp, true);
> +
> + rcu_assign_pointer(current->sched_cache_grp, NULL);
> +
> + if (grp)
> + sched_cache_group_put(grp);
> + }
> +#endif
Seriously, WTF ?!
>
> mmap_read_lock(mm);
> mmgrab_lazy_tlb(mm);
> diff --git a/kernel/fork.c b/kernel/fork.c
> index 416758c8a3d4..2e79548cb7c1 100644
> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -1599,6 +1599,19 @@ static int copy_mm(u64 clone_flags, struct task_struct *tsk)
>
> tsk->mm = mm;
> tsk->active_mm = mm;
> +#ifdef CONFIG_SCHED_CACHE
> + {
> + /*
> + * A task holds its own reference on the group, separate from
> + * the reference held by its mm_struct. Acquire it before
> + * publishing the pointer.
> + */
> + struct sched_cache_group *grp =
> + sched_cache_group_get(mm->sched_cache_grp);
> +
> + rcu_assign_pointer(tsk->sched_cache_grp, grp);
> + }
> +#endif
And again.
> return 0;
> }
>
> @@ -2599,6 +2612,16 @@ __latent_entropy struct task_struct *copy_process(
> bad_fork_cleanup_namespaces:
> exit_nsproxy_namespaces(p);
> bad_fork_cleanup_mm:
> +#ifdef CONFIG_SCHED_CACHE
> + /*
> + * copy_mm() took a task reference on the cache group; a failed fork
> + * never reaches exit_mm(), so release it here to avoid leaking the
> + * group and its per-CPU buffer.
> + */
> + sched_cache_group_put(rcu_dereference_protected(p->sched_cache_grp, true));
> + RCU_INIT_POINTER(p->sched_cache_grp, NULL);
> +#endif
> +
> if (p->mm) {
> mm_clear_owner(p->mm, p);
> mmput(p->mm);
Drugs, it must be drugs and lots of it :-(
Please, try again.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
2026-09-10 18:33 ` Kayra Cizmeci
@ 2026-09-10 20:46 ` Tim Chen
2026-09-10 22:03 ` Kayra Cizmeci
0 siblings, 1 reply; 11+ messages in thread
From: Tim Chen @ 2026-09-10 20:46 UTC (permalink / raw)
To: Kayra Cizmeci
Cc: brauner, bsegall, dietmar.eggemann, imv4bel, jack, juri.lelli,
kees, kprateek.nayak, linux-fsdevel, linux-kernel, linux-mm,
mgorman, mingo, peterz, qyousef, ricardo.neri-calderon, rostedt,
srikar, sshegde, vincent.guittot, vineethr, viro, vschneid,
wanglu.priv, yi1.lai, yu.c.chen, zhanxusheng1024, zhanxusheng,
ziqianlu
On Thu, 2026-09-10 at 21:33 +0300, Kayra Cizmeci wrote:
> Hello :>,
>
> > alb_break_llc() decides whether to break LLC preference during active
> > load balance. It does so by testing that every runnable fair task on the
> > source rq prefers its LLC:
> >
> > env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
> >
> > But the two counters cover different sets. nr_pref_llc_running is updated
> > in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
> > so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
> > clear_delayed() and drops delay-dequeued tasks.
>
> > So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
> > in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
> > alb_break_llc() returns false, and active balance is free to pull a task
> > off its preferred LLC. Active balance only moves runnable tasks, and this
> > is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
> > skips the per-task test in can_migrate_task(). The runnable set is the one
> > we want.
>
> > Fix it on the counter side. A task should be counted in
> > nr_pref_llc_running exactly while it is both queued on its preferred LLC
> > (pref_llc_queued) and runnable (!sched_delayed). Define that membership
> > once in task_pref_llc_runnable(), and adjust the counter only through
> > pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
> > change either input: account_llc_enqueue(), account_llc_dequeue(),
> > set_delayed() and clear_delayed(). Gating every update on the same
> > predicate keeps the delay, wake and dequeue paths from double-counting
> > or underflowing; see the comments at those sites for the ordering.
>
> > nr_llc_running and sd->llc_counts are not touched and stay on queued
> > semantics.
>
> I have one question tho, can't we combine the checks with h_nr_runnable? On the paper
> if we are updating h_nr_runnable we could check if the nr_pref_llc_running can be
> updated and update it if the condition is right. Because, every nr_pref_llc_running enters
> h_nr_runnable while not every h_nr_runnable enters nr_pref_llc_running.
>
> Why instead we just check the nr_pref_llc_running's conditions on task_pref_llc_runnable()
> and call these dec and inc functions after the h_nr_runnable updates. Wouldn't it be clear that way?
> If possible?
Yes, nr_pref_llc_running is a subset of h_nr_runnable.
We have to keep nr_pref_llc_running accounting apart from h_nr_runnable in set_delayed().
Note that in set_delayed(), pref_llc_running_dec() has to run while the task still looks runnable,
that is before se->sched_delayed = 1, because task_pref_llc_runnable()
gates on !sched_delayed. h_nr_runnable is decremented after the flag is
set:
if (entity_is_task(se))
pref_llc_running_dec(...); /* sched_delayed still 0 */
se->sched_delayed = 1;
...
for_each_sched_entity(se)
cfs_rq->h_nr_runnable--; /* sched_delayed already 1 */
So moving the accounting next to (or after) the h_nr_runnable update
would make task_pref_llc_runnable() return false and skip the
decrement, leaving nr_pref_llc_running too high.
clear_delayed() happens to be safe either way, since it clears
sched_delayed first, but keeping the two symmetric and calling inc/dec
explicitly at each site is what lets the single task_pref_llc_runnable()
predicate stay the one source of truth.
There is also a scope difference: h_nr_runnable is per-cfs_rq and
updated at every level of the hierarchy in the for_each_sched_entity()
loop, while nr_pref_llc_running is a per-rq scalar updated once per
task - which is why the dec sits before the loop, not inside it.
Thanks.
Tim
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
2026-09-10 20:46 ` Tim Chen
@ 2026-09-10 22:03 ` Kayra Cizmeci
0 siblings, 0 replies; 11+ messages in thread
From: Kayra Cizmeci @ 2026-09-10 22:03 UTC (permalink / raw)
To: tim.c.chen
Cc: brauner, bsegall, dietmar.eggemann, imv4bel, jack, juri.lelli,
kayracizmeci, kees, kprateek.nayak, linux-fsdevel, linux-kernel,
linux-mm, mgorman, mingo, peterz, qyousef, ricardo.neri-calderon,
rostedt, srikar, sshegde, vincent.guittot, vineethr, viro,
vschneid, wanglu.priv, yi1.lai, yu.c.chen, zhanxusheng1024,
zhanxusheng, ziqianlu
Hello Tim,
> So moving the accounting next to (or after) the h_nr_runnable update
> would make task_pref_llc_runnable() return false and skip the
> decrement, leaving nr_pref_llc_running too high.
What I really wanted wasn't getting the accounting next to or after the h_nr_runnable.
If we are updating h_nr_runnable in some way that means we don't need
its check since it's already getting updated. And if it's getting updated
that means on that branch we know how our check should behave since we
are a subset of it. We can skip the delayed check on that way since we are
trying to behave as h_nr_runnable's subset.
> if (entity_is_task(se))
> pref_llc_running_dec(...); /* sched_delayed still 0 */
> se->sched_delayed = 1;
> ...
> for_each_sched_entity(se)
> cfs_rq->h_nr_runnable--; /* sched_delayed already 1 */
For example:
In this code the h_nr_runnable is updated the same way regarding what is sched_delayed.
That means if we want to behave as a subset of it, we don't need the check delayed,
since we check the delayed to be a subset but if h_nr_runnable is decreasing/increasing
we should look into our checks.
My head hurts. Please notify if I'm wrong.
Thanks,
Kayra
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
2026-09-10 18:33 ` Kayra Cizmeci
@ 2026-09-10 22:48 ` Kayra Cizmeci
1 sibling, 0 replies; 11+ messages in thread
From: Kayra Cizmeci @ 2026-09-10 22:48 UTC (permalink / raw)
To: tim.c.chen
Cc: brauner, bsegall, dietmar.eggemann, imv4bel, jack, juri.lelli,
kees, kprateek.nayak, linux-fsdevel, linux-kernel, linux-mm,
mgorman, mingo, peterz, qyousef, ricardo.neri-calderon, rostedt,
srikar, sshegde, vincent.guittot, vineethr, viro, vschneid,
wanglu.priv, yi1.lai, yu.c.chen, zhanxusheng1024, zhanxusheng,
ziqianlu, Kayra Cizmeci
Hello Tim,
I thought about the older messages and I think it's really not worth it.
I mean, even if I'm right on 1 call site, the others won't behave the same.
And I think this version is cleaner anyways.
Only thing I could see was that on 1/4.
Feel free to add:
Reviewed-by: Kayra Cizmeci <kayracizmeci@gmail.com>
Now I'm going to sleep. Ah..
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp
2026-09-10 19:19 ` Peter Zijlstra
@ 2026-09-10 22:50 ` Tim Chen
0 siblings, 0 replies; 11+ messages in thread
From: Tim Chen @ 2026-09-10 22:50 UTC (permalink / raw)
To: Peter Zijlstra
Cc: Ingo Molnar, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
Yi Lai, linux-kernel, linux-mm, linux-fsdevel
On Thu, 2026-09-10 at 21:19 +0200, Peter Zijlstra wrote:
> On Thu, Sep 10, 2026 at 10:46:12AM -0700, Tim Chen wrote:
>
> > Co-developed-by: Chen Yu <yu.c.chen@intel.com>
> > Signed-off-by: Chen Yu <yu.c.chen@intel.com>
> > Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
>
> :-(
>
> > ---
> > fs/exec.c | 14 ++++
> > include/linux/sched.h | 3 +
> > kernel/exit.c | 26 +++++--
> > kernel/fork.c | 23 ++++++
> > kernel/sched/cache_sched.c | 19 +++++
> > kernel/sched/fair.c | 142 +++++++++++++++++++++----------------
> > kernel/sched/sched.h | 3 +
> > 7 files changed, 164 insertions(+), 66 deletions(-)
> >
> > diff --git a/fs/exec.c b/fs/exec.c
> > index 745f6eb5279e..7a8a9954343e 100644
> > --- a/fs/exec.c
> > +++ b/fs/exec.c
> > @@ -882,6 +882,20 @@ static int exec_mmap(struct linux_binprm *bprm)
> > active_mm = tsk->active_mm;
> > tsk->active_mm = mm;
> > tsk->mm = mm;
> > +#ifdef CONFIG_SCHED_CACHE
> > + {
> > + struct sched_cache_group *old_grp, *new_grp;
> > +
> > + old_grp = rcu_dereference_protected(tsk->sched_cache_grp, true);
> > +
> > + /* Acquire the reference before publishing the pointer. */
> > + new_grp = sched_cache_group_get(mm->sched_cache_grp);
> > +
> > + rcu_assign_pointer(tsk->sched_cache_grp, new_grp);
> > + if (old_grp)
> > + sched_cache_group_put(old_grp);
> > + }
> > +#endif
>
> Guys no! This is horrific crap. This is not how we do things and I would
> have expected you all to know this.
>
> Have you heard of this new fangled thing called a function?
>
> Imagine all of those being just:
>
> sched_cache_exec_mmap(tsk, mm);
>
>
> Also: rcu_dereference_protected(.c = true) is another offence, that's
> just wrong.
>
>
> > diff --git a/kernel/exit.c b/kernel/exit.c
> > index 006edcc0c2c5..442535778ce1 100644
> > --- a/kernel/exit.c
> > +++ b/kernel/exit.c
> > @@ -552,23 +552,25 @@ void mm_update_next_owner(struct mm_struct *mm)
> > * Subtract the memory footprint of the current task from
> > * mm.
> > */
> > -static void exit_mm_sched_cache(struct mm_struct *mm)
> > +static void exit_mm_sched_cache(void)
> > {
> > + struct sched_cache_group *grp =
> > + rcu_dereference_protected(current->sched_cache_grp, true);
> > unsigned long fp, sub;
> >
> > - if (!current->total_numa_faults)
> > + if (!grp || !current->total_numa_faults)
> > return;
> > /*
> > * No lock protection due to performance considerations.
> > * Make sure the group footprint does not become
> > * negative.
> > */
> > - fp = READ_ONCE(mm->sched_cache_grp->footprint);
> > + fp = READ_ONCE(grp->footprint);
> > sub = min(fp, current->total_numa_faults);
> > - WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub);
> > + WRITE_ONCE(grp->footprint, fp - sub);
> > }
> > #else
> > -static inline void exit_mm_sched_cache(struct mm_struct *mm)
> > +static inline void exit_mm_sched_cache(void)
> > {
> > }
> > #endif /* CONFIG_SCHED_CACHE CONFIG_NUMA_BALANCING */
> > @@ -585,7 +587,19 @@ static void exit_mm(void)
> > if (!mm)
> > return;
> >
> > - exit_mm_sched_cache(mm);
> > + exit_mm_sched_cache();
> > +
> > +#ifdef CONFIG_SCHED_CACHE
> > + {
> > + struct sched_cache_group *grp =
> > + rcu_dereference_protected(current->sched_cache_grp, true);
> > +
> > + rcu_assign_pointer(current->sched_cache_grp, NULL);
> > +
> > + if (grp)
> > + sched_cache_group_put(grp);
> > + }
> > +#endif
>
> Seriously, WTF ?!
>
> >
> > mmap_read_lock(mm);
> > mmgrab_lazy_tlb(mm);
> > diff --git a/kernel/fork.c b/kernel/fork.c
> > index 416758c8a3d4..2e79548cb7c1 100644
> > --- a/kernel/fork.c
> > +++ b/kernel/fork.c
> > @@ -1599,6 +1599,19 @@ static int copy_mm(u64 clone_flags, struct task_struct *tsk)
> >
> > tsk->mm = mm;
> > tsk->active_mm = mm;
> > +#ifdef CONFIG_SCHED_CACHE
> > + {
> > + /*
> > + * A task holds its own reference on the group, separate from
> > + * the reference held by its mm_struct. Acquire it before
> > + * publishing the pointer.
> > + */
> > + struct sched_cache_group *grp =
> > + sched_cache_group_get(mm->sched_cache_grp);
> > +
> > + rcu_assign_pointer(tsk->sched_cache_grp, grp);
> > + }
> > +#endif
>
> And again.
>
> > return 0;
> > }
> >
> > @@ -2599,6 +2612,16 @@ __latent_entropy struct task_struct *copy_process(
> > bad_fork_cleanup_namespaces:
> > exit_nsproxy_namespaces(p);
> > bad_fork_cleanup_mm:
> > +#ifdef CONFIG_SCHED_CACHE
> > + /*
> > + * copy_mm() took a task reference on the cache group; a failed fork
> > + * never reaches exit_mm(), so release it here to avoid leaking the
> > + * group and its per-CPU buffer.
> > + */
> > + sched_cache_group_put(rcu_dereference_protected(p->sched_cache_grp, true));
> > + RCU_INIT_POINTER(p->sched_cache_grp, NULL);
> > +#endif
> > +
> > if (p->mm) {
> > mm_clear_owner(p->mm, p);
> > mmput(p->mm);
>
> Drugs, it must be drugs and lots of it :-(
>
>
Sorry for the warts in this version. Will clean it up and send
an update.
Tim
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-09-10 22:50 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-10 17:46 [PATCH 0/4] sched/cache: Fixes for cache aware scheduling Tim Chen
2026-09-10 17:46 ` [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
2026-09-10 18:33 ` Kayra Cizmeci
2026-09-10 20:46 ` Tim Chen
2026-09-10 22:03 ` Kayra Cizmeci
2026-09-10 22:48 ` Kayra Cizmeci
2026-09-10 17:46 ` [PATCH 2/4] sched/cache: Honor migrate_llc_task semantics in active load balance Tim Chen
2026-09-10 17:46 ` [PATCH 3/4] sched/cache: Decouple sched_cache_group from mm Tim Chen
2026-09-10 17:46 ` [PATCH 4/4] sched/cache: Introduce task_struct->sched_cache_grp Tim Chen
2026-09-10 19:19 ` Peter Zijlstra
2026-09-10 22:50 ` Tim Chen
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®