mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/6] sched/cache: Fixes for cache aware scheduling
@ 2026-09-22  0:37 Tim Chen
  2026-09-22  0:37 ` [PATCH v2 1/6] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
                   ` (5 more replies)
  0 siblings, 6 replies; 12+ messages in thread
From: Tim Chen @ 2026-09-22  0:37 UTC (permalink / raw)
  To: Peter Zijlstra, Ingo Molnar
  Cc: Tim Chen, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
	Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
	K Prateek Nayak, Kees Cook, Christian Brauner, Alexander Viro,
	Jan Kara, Shrikanth Hegde, Qais Yousef, Aaron Lu,
	Srikar Dronamraju, Vineeth Remanan Pillai, Ricardo Neri-Calderon,
	Chen Yu, Lu Wang, Hyunwoo Kim, Zhan Xusheng, Zhan Xusheng,
	Yi Lai, Rafael J . Wysocki, Greg Kroah-Hartman, Danilo Krummrich,
	Zenghui Yu, linux-kernel, linux-mm, linux-fsdevel

Hi all,

This is an update of the patches to fix cache aware scheduling issues
found in v7.2.

We collect the fixes in this series so it is easier to track. We have
added two new fixes for issues found since v1 of this series.

Patch 1: alb_break_llc() compares nr_pref_llc_running with
cfs.h_nr_runnable, but those count different sets - one follows queued
tasks, the other drops delay-dequeued ones.  With DELAY_DEQUEUE the
equality stops holding and active balance pulls a task off its preferred
LLC.  So fix the counter.  Reported by Zhan Xusheng:
https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xiaomi.com/

v1->v2: Minor code rearrangements in set_delayed().

Patch 2 (Lu Wang): the stopper doing active load balance builds a fresh
lb_env that doesn't inherit migration_type, so can_migrate_task() can
move a task *out* of its preferred LLC.  A new LBF_ACTIVE_LB_LLC flag
and picking the stopper callback at kick time keep the intent; passing
migration_type through the stopper would muddy delayed dequeue.  v4:
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@gmail.com/

v1->v2: No change.

Patches 3-4 are for the use after free Hyunwoo Kim caught with KASAN:
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
account_mm_sched() reaches the stats via p->mm->sc_stat, but a task can
be switching mm on one CPU while another is inside account_mm_sched(),
so the mm and the stats inside it can go away underneath.  Locking the
rq in the mm free path felt like the wrong trade, so patch 4 pulls
sched_cache_stat out of mm_struct into a refcounted, RCU freed
sched_cache_group - just moving code - and patch 4 does the real fix:
each task takes its own reference (copy_mm(), exec_mmap(), dropped in
exit_mm()), and leverage call_rcu() to to protect
against UAF in account_mm_sched() so the group outlives any mm switch.
Zehnghui Yu also independentaly found this issue with memory poison.
https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@linux.dev/

@Hyunwoo and @Zhenhui, will appreciated you can test these patches
and add your Tested-by

Nice side effect: the group no longer follows the address space,
so a user defined group, or cgroup or numa_group could own it later.
These are also the grouping by prctl RFC's first two patches, sent here
so the fix isn't held up by that discussion.

v1->v2: Put cache aware related code in process exit/fork/copy in
its own functions

Patch 5 This patch makes sure kernel threads are excluded from cache
aware scheduling consideration.

v1->v2: Split from previous patch 3 as suggested by Peter Z.

Patch 6 is a new patch to fix an issue of undercomputing the LLC size
during CPU hot plug events.  The LLC size was used for estimating
if a process's memory footprint will fit a LLC. 
https://lore.kernel.org/all/20260916134432.11767-1-davichazbh@gmail.com/

BTW, there are two other issues in discussion currently and need
a bit more work:
1. Incorrect donor context being passed to task_tick_cache().  
https://lore.kernel.org/lkml/20260909092901.2989564-1-sh_def@163.com/
It is currently under discussion and is not included in this series.
2. Cache aware scheduling interfering with ITMT.
https://lore.kernel.org/lkml/20260810033742.1688718-1-yu.c.chen@intel.com/
https://lore.kernel.org/lkml/2fe2c681-b748-41fa-8b56-1169c86cefbc@intel.com/ 

Applies on sched/urgent branch. 

Tim Chen and Chen Yu

Chen Yu (1):
  sched/cache: Skip kernel thread for cache aware scheduling

Davi Chaves Azevedo (1):
  sched/cache: Refresh LLC capacity across CPU hotplug

Lu Wang (1):
  sched/cache: Honor migrate_llc_task semantics in active load balance

Tim Chen (3):
  sched/cache: Keep nr_pref_llc_running in the runnable domain
  sched/cache: Decouple sched_cache_group from mm
  sched/cache: Introduce task_struct->sched_cache_grp

 drivers/base/cacheinfo.c       |  11 +-
 fs/exec.c                      |   1 +
 include/linux/mm_types.h       |  15 +-
 include/linux/sched.h          |  21 ++-
 include/linux/sched/topology.h |   4 +-
 kernel/exit.c                  |  28 +--
 kernel/fork.c                  |   2 +
 kernel/sched/build_utility.c   |   4 +
 kernel/sched/cache_sched.c     | 106 +++++++++++
 kernel/sched/fair.c            | 310 ++++++++++++++++++++++++---------
 kernel/sched/sched.h           |   3 +
 kernel/sched/topology.c        |  22 ++-
 12 files changed, 394 insertions(+), 133 deletions(-)
 create mode 100644 kernel/sched/cache_sched.c

^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-09-22  8:30 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-22  0:37 [PATCH v2 0/6] sched/cache: Fixes for cache aware scheduling Tim Chen
2026-09-22  0:37 ` [PATCH v2 1/6] sched/cache: Keep nr_pref_llc_running in the runnable domain Tim Chen
2026-09-22  7:22   ` Peter Zijlstra
2026-09-22  8:25     ` Chen, Yu C
2026-09-22  8:28       ` Peter Zijlstra
2026-09-22  0:37 ` [PATCH v2 2/6] sched/cache: Honor migrate_llc_task semantics in active load balance Tim Chen
2026-09-22  7:23   ` Peter Zijlstra
2026-09-22  8:30     ` Chen, Yu C
2026-09-22  0:37 ` [PATCH v2 3/6] sched/cache: Decouple sched_cache_group from mm Tim Chen
2026-09-22  0:37 ` [PATCH v2 4/6] sched/cache: Introduce task_struct->sched_cache_grp Tim Chen
2026-09-22  0:37 ` [PATCH v2 5/6] sched/cache: Skip kernel thread for cache aware scheduling Tim Chen
2026-09-22  0:37 ` [PATCH v2 6/6] sched/cache: Refresh LLC capacity across CPU hotplug Tim Chen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®