From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f173.google.com (mail-pl1-f173.google.com [209.85.214.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6406D30DD1B for ; Wed, 15 Apr 2026 02:07:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.173 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776218829; cv=none; b=H1y81lE8u6DNN0ER5fSRoTmkvFuFdKpWbaJhsHo5jpvlwMYYTxF9u4pNCf0+8rCBOFL6U30iryJ7ZK0XI/KFZkVsvBLBu7iQe2x6HqgP6OVlZP+luBawpaHe41QVEFwrYWT6m5UmsGzi4PSf3snh2YVCgnBsPZMmJRd5HDjnEBM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776218829; c=relaxed/simple; bh=9iVP9E/OfR6TOv9AY0Oaj7vHy/72RW30bN5As/nMjHs=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=eV7xTaAUaJ5M9j6pLKFV6zKQiXOgNrMpFhuRDfOKOuga4KZceMGHvYqdwoe48YgkB7XGKl30Qu8jcH2S7Iq/P/pP+IrMq8d9ikiM7JivZJz6LK+aiBd481hcXq588kePneV8OC6LgFwg3CtvWF7d7c0Tiqs3CxPRRmTB1oN96DU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=VjHgTUKC; arc=none smtp.client-ip=209.85.214.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="VjHgTUKC" Received: by mail-pl1-f173.google.com with SMTP id d9443c01a7336-2aaf59c4f7cso28998745ad.1 for ; Tue, 14 Apr 2026 19:07:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1776218827; x=1776823627; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=8PcxP68CMtYJqUfAQasBqtVmx2hydVy8qT+a8gk8Z7c=; b=VjHgTUKCZgTHBGz6Kznmw3Ap74KEkC15UReVcHrm4t3G0l7mstX9cuLWWTdOprYLih vfkIKYblo96nmu5XSNVVKkrE9sh8nEe4HP7OTQJI88N+ZJQJyIxOJHMZjWJ9naEGbel+ LFQzVyDEXkztZJr/n/ilS8fyMtbH4nVEdUg5iTVLIu1A2bE+Q4ySPsBGzvBaSY3NhOqq FTntWPyDNJ0rkFcX2xF31xXyy+ypcwbASH+/PWyI4Dadwk9IvnXXxXP+J6yqdgAp8SL0 4mXCeLIIFdGfk4ymw+ozYONnolkJwoAEfM52ScmA5QqPgTQQIyHmpPMIEQrIITbvfzJf /xKg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1776218827; x=1776823627; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=8PcxP68CMtYJqUfAQasBqtVmx2hydVy8qT+a8gk8Z7c=; b=ff+B+e9mfzJ5UNTtiEVaBCdCdONgIiSZH93jvcpeydtT5iNWP54CSF4FlDfPs68aR1 cpIi1ul5o2fMXg2DdjMrb/rM9dZCpkPdNs5JdmwLGQcXT16DUJVdQlODkmGojFAWpY+1 eFNtld5PJfQMVfUJp0+ySIbEo910Q//1aJ0q4CX1L7f6mSSf3BffZmU8kL7l8ui7DOye AFZbfbRHxa1TbQ9xKajGvwEP/Et3mQYzJjYtoC4xD3ZIYTrui0y6jHLe7X72Eji9pOTY syhWccPYLnIkfQO9RecarlKLt3dgbn1/kdnvIDKbw0p4n1f2gijkjHedQkQunEThMJ7s zBYA== X-Forwarded-Encrypted: i=1; AFNElJ/EKmHTqJQSrEZtSyDFEgZUQveRmAxQFOQDs5iBMJ8UfYNNazDn2i39VPnJH3rTCGkG+SRdqNzekqDCzqk=@vger.kernel.org X-Gm-Message-State: AOJu0Yw4vCeXGNYm9YzsUnWMkVJv55EaZOuteMeLxn0sz4M3loflRSCe iqNN4wgB2ALx0UoNvHg7bv9U5+pOu8jhuwuV6WhmobFna4LzwaQjfMt5 X-Gm-Gg: AeBDiesJf0f2dQoM6l9e/wIOdmxsM/UO0GaFW8hfW7WptTcW7jzFScoTXENqBps0SSE iLjMBtpP/4SyseYti+PwhPQvAQ9f5mTAg7yWrkx2fY02rMqLYTKT8SP6MFoY5wv913a1hTXpTgM CLWXobjTgTCD43m7WFbWTpszx0OMna9av81Ck/j8+B7VvCkkKjfGiEf/ObySx8Mqp8vz0dNPSic iF5o9Pn4wpCJCVAII6+FPT+HpuA/C5/LHKOTyOAuNyzmew+7Yw3/qu1x5RKlsBxEoQwEivbKcl5 tOZHk7vCfx+FJ+xdWnugK4Y0jHbRgN4DBUDTjJP2Dtwr/6M1TqTZjTTiaejOvlKvJh6XJT+Konn ZIJzjareHdmrAmudj7OT6rg9rr4VWTH/8i/qXwbRy2fV4Yjt38+6kMtSUyDnDsB8ts7RS/OpR2T CE1XIFy0rJEFBhcPOMDkkg12OvdkCh/PGSkAoYWGFDqOkLcC193g== X-Received: by 2002:a17:903:acf:b0:2b2:539b:d29a with SMTP id d9443c01a7336-2b2d5a1525amr201295565ad.23.1776218826307; Tue, 14 Apr 2026 19:07:06 -0700 (PDT) Received: from [192.168.255.10] ([43.132.141.20]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2b4782c17e4sm2838565ad.77.2026.04.14.19.06.59 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 14 Apr 2026 19:07:05 -0700 (PDT) Message-ID: <09cf7ee3-6e27-4505-9692-4b4a4707c8b2@gmail.com> Date: Wed, 15 Apr 2026 10:06:56 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [Patch v4 01/22] sched/cache: Introduce infrastructure for cache-aware load balancing To: Tim Chen , Peter Zijlstra , Ingo Molnar , K Prateek Nayak , "Gautham R . Shenoy" , Vincent Guittot , Vern Hao Cc: Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , Josh Don , Gavin Guo , Qais Yousef , Libo Chen , linux-kernel@vger.kernel.org References: <6269a53221b9439b9ca00d18a9d1946fb64d8cff.1775065312.git.tim.c.chen@linux.intel.com> From: Vern Hao In-Reply-To: <6269a53221b9439b9ca00d18a9d1946fb64d8cff.1775065312.git.tim.c.chen@linux.intel.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit Hi Tim, On 2026/4/2 05:52, Tim Chen wrote: > From: "Peter Zijlstra (Intel)" > > Adds infrastructure to enable cache-aware load balancing, > which improves cache locality by grouping tasks that share resources > within the same cache domain. This reduces cache misses and improves > overall data access efficiency. > > In this initial implementation, threads belonging to the same process > are treated as entities that likely share working sets. The mechanism > tracks per-process CPU occupancy across cache domains and attempts to > migrate threads toward cache-hot domains where their process already > has active threads, thereby enhancing locality. > > This provides a basic model for cache affinity. While the current code > targets the last-level cache (LLC), the approach could be extended to > other domain types such as clusters (L2) or node-internal groupings. > > At present, the mechanism selects the CPU within an LLC that has the > highest recent runtime. Subsequent patches in this series will use this > information in the load-balancing path to guide task placement toward > preferred LLCs. > > In the future, more advanced policies could be integrated through NUMA > balancing-for example, migrating a task to its preferred LLC when spare > capacity exists, or swapping tasks across LLCs to improve cache affinity. > Grouping of tasks could also be generalized from that of a process > to be that of a NUMA group, or be user configurable. > > Signed-off-by: Peter Zijlstra (Intel) > Signed-off-by: Chen Yu > Signed-off-by: Tim Chen > --- > > Notes: > v3->v4: > No change. > > include/linux/mm_types.h | 32 +++++ > include/linux/sched.h | 24 ++++ > init/Kconfig | 11 ++ > kernel/fork.c | 6 + > kernel/sched/core.c | 6 + > kernel/sched/fair.c | 266 +++++++++++++++++++++++++++++++++++++++ > kernel/sched/sched.h | 14 +++ > 7 files changed, 359 insertions(+) > > diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h > index 3cc8ae722886..67b2dfcc71ea 100644 > --- a/include/linux/mm_types.h > +++ b/include/linux/mm_types.h > @@ -1173,6 +1173,8 @@ struct mm_struct { > /* MM CID related storage */ > struct mm_mm_cid mm_cid; > > + /* sched_cache related statistics */ > + struct sched_cache_stat sc_stat; > #ifdef CONFIG_MMU > atomic_long_t pgtables_bytes; /* size of all page tables */ > #endif > @@ -1575,6 +1577,36 @@ static inline unsigned int mm_cid_size(void) > # define MM_CID_STATIC_SIZE 0 > #endif /* CONFIG_SCHED_MM_CID */ > > +#ifdef CONFIG_SCHED_CACHE > +void mm_init_sched(struct mm_struct *mm, > + struct sched_cache_time __percpu *pcpu_sched); > + > +static inline int mm_alloc_sched_noprof(struct mm_struct *mm) > +{ > + struct sched_cache_time __percpu *pcpu_sched = > + alloc_percpu_noprof(struct sched_cache_time); > + > + if (!pcpu_sched) > + return -ENOMEM; > + > + mm_init_sched(mm, pcpu_sched); > + return 0; > +} > + > +#define mm_alloc_sched(...) alloc_hooks(mm_alloc_sched_noprof(__VA_ARGS__)) > + > +static inline void mm_destroy_sched(struct mm_struct *mm) > +{ > + free_percpu(mm->sc_stat.pcpu_sched); > + mm->sc_stat.pcpu_sched = NULL; > +} > +#else /* !CONFIG_SCHED_CACHE */ > + > +static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; } > +static inline void mm_destroy_sched(struct mm_struct *mm) { } > + > +#endif /* CONFIG_SCHED_CACHE */ > + > struct mmu_gather; > extern void tlb_gather_mmu(struct mmu_gather *tlb, struct mm_struct *mm); > extern void tlb_gather_mmu_fullmm(struct mmu_gather *tlb, struct mm_struct *mm); > diff --git a/include/linux/sched.h b/include/linux/sched.h > index a7b4a980eb2f..bd33f5b9096b 100644 > --- a/include/linux/sched.h > +++ b/include/linux/sched.h > @@ -1406,6 +1406,10 @@ struct task_struct { > unsigned long numa_pages_migrated; > #endif /* CONFIG_NUMA_BALANCING */ > > +#ifdef CONFIG_SCHED_CACHE > + struct callback_head cache_work; > +#endif > + > struct rseq_data rseq; > struct sched_mm_cid mm_cid; > > @@ -2376,6 +2380,26 @@ static __always_inline int task_mm_cid(struct task_struct *t) > } > #endif > > +#ifdef CONFIG_SCHED_CACHE > + > +struct sched_cache_time { > + u64 runtime; > + unsigned long epoch; > +}; > + > +struct sched_cache_stat { > + struct sched_cache_time __percpu *pcpu_sched; > + raw_spinlock_t lock; > + unsigned long epoch; > + int cpu; > +} ____cacheline_aligned_in_smp; > + > +#else > + > +struct sched_cache_stat { }; > + > +#endif > + > #ifndef MODULE > #ifndef COMPILE_OFFSETS > > diff --git a/init/Kconfig b/init/Kconfig > index 444ce811ea67..d1f3579d6ea4 100644 > --- a/init/Kconfig > +++ b/init/Kconfig > @@ -1005,6 +1005,17 @@ config NUMA_BALANCING > > This system will be inactive on UMA systems. > > +config SCHED_CACHE > + bool "Cache aware load balance" > + default y > + depends on SMP > + help > + When enabled, the scheduler will attempt to aggregate tasks from > + the same process onto a single Last Level Cache (LLC) domain when > + possible. This improves cache locality by keeping tasks that share > + resources within the same cache domain, reducing cache misses and > + lowering data access latency. > + > config NUMA_BALANCING_DEFAULT_ENABLED > bool "Automatically enable NUMA aware memory/task placement" > default y > diff --git a/kernel/fork.c b/kernel/fork.c > index 65113a304518..98ef5c997cc3 100644 > --- a/kernel/fork.c > +++ b/kernel/fork.c > @@ -724,6 +724,7 @@ void __mmdrop(struct mm_struct *mm) > cleanup_lazy_tlbs(mm); > > WARN_ON_ONCE(mm == current->active_mm); > + mm_destroy_sched(mm); > mm_free_pgd(mm); > mm_free_id(mm); > destroy_context(mm); > @@ -1124,6 +1125,9 @@ static struct mm_struct *mm_init(struct mm_struct *mm, struct task_struct *p, > if (mm_alloc_cid(mm, p)) > goto fail_cid; > > + if (mm_alloc_sched(mm)) > + goto fail_sched; > + > if (percpu_counter_init_many(mm->rss_stat, 0, GFP_KERNEL_ACCOUNT, > NR_MM_COUNTERS)) > goto fail_pcpu; > @@ -1133,6 +1137,8 @@ static struct mm_struct *mm_init(struct mm_struct *mm, struct task_struct *p, > return mm; > > fail_pcpu: > + mm_destroy_sched(mm); > +fail_sched: > mm_destroy_cid(mm); > fail_cid: > destroy_context(mm); > diff --git a/kernel/sched/core.c b/kernel/sched/core.c > index b7f77c165a6e..eff8695000e7 100644 > --- a/kernel/sched/core.c > +++ b/kernel/sched/core.c > @@ -4437,6 +4437,7 @@ static void __sched_fork(u64 clone_flags, struct task_struct *p) > init_numa_balancing(clone_flags, p); > p->wake_entry.u_flags = CSD_TYPE_TTWU; > p->migration_pending = NULL; > + init_sched_mm(p); > } > > DEFINE_STATIC_KEY_FALSE(sched_numa_balancing); > @@ -8749,6 +8750,11 @@ void __init sched_init(void) > > rq->core_cookie = 0UL; > #endif > +#ifdef CONFIG_SCHED_CACHE > + raw_spin_lock_init(&rq->cpu_epoch_lock); > + rq->cpu_epoch_next = jiffies; > +#endif > + > zalloc_cpumask_var_node(&rq->scratch_mask, GFP_KERNEL, cpu_to_node(i)); > } > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index bf948db905ed..eb3cfb852a93 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -1228,6 +1228,8 @@ void post_init_entity_util_avg(struct task_struct *p) > sa->runnable_avg = sa->util_avg; > } > > +static inline void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec); > + > static s64 update_se(struct rq *rq, struct sched_entity *se) > { > u64 now = rq_clock_task(rq); > @@ -1250,6 +1252,7 @@ static s64 update_se(struct rq *rq, struct sched_entity *se) > > trace_sched_stat_runtime(running, delta_exec); > account_group_exec_runtime(running, delta_exec); > + account_mm_sched(rq, running, delta_exec); > > /* cgroup time is always accounted against the donor */ > cgroup_account_cputime(donor, delta_exec); > @@ -1271,6 +1274,267 @@ static s64 update_se(struct rq *rq, struct sched_entity *se) > > static void set_next_buddy(struct sched_entity *se); > > +#ifdef CONFIG_SCHED_CACHE > + > +/* > + * XXX numbers come from a place the sun don't shine -- probably wants to be SD > + * tunable or so. > + */ > +#define EPOCH_PERIOD (HZ / 100) /* 10 ms */ > +#define EPOCH_LLC_AFFINITY_TIMEOUT 5 /* 50 ms */ > + > +static int llc_id(int cpu) > +{ > + if (cpu < 0) > + return -1; > + > + return per_cpu(sd_llc_id, cpu); > +} > + > +void mm_init_sched(struct mm_struct *mm, > + struct sched_cache_time __percpu *_pcpu_sched) > +{ > + unsigned long epoch = 0; > + int i; > + > + for_each_possible_cpu(i) { > + struct sched_cache_time *pcpu_sched = per_cpu_ptr(_pcpu_sched, i); > + struct rq *rq = cpu_rq(i); > + > + pcpu_sched->runtime = 0; > + /* a slightly stale cpu epoch is acceptible */ > + pcpu_sched->epoch = rq->cpu_epoch; > + epoch = rq->cpu_epoch; > + } > + > + raw_spin_lock_init(&mm->sc_stat.lock); > + mm->sc_stat.epoch = epoch; > + mm->sc_stat.cpu = -1; > + > + /* > + * The update to mm->sc_stat should not be reordered > + * before initialization to mm's other fields, in case > + * the readers may get invalid mm_sched_epoch, etc. > + */ > + smp_store_release(&mm->sc_stat.pcpu_sched, _pcpu_sched); > +} > + > +/* because why would C be fully specified */ > +static __always_inline void __shr_u64(u64 *val, unsigned int n) > +{ > + if (n >= 64) { > + *val = 0; > + return; > + } > + *val >>= n; > +} > + > +static inline void __update_mm_sched(struct rq *rq, > + struct sched_cache_time *pcpu_sched) > +{ > + lockdep_assert_held(&rq->cpu_epoch_lock); > + > + unsigned long n, now = jiffies; > + long delta = now - rq->cpu_epoch_next; > + > + if (delta > 0) { > + n = (delta + EPOCH_PERIOD - 1) / EPOCH_PERIOD; > + rq->cpu_epoch += n; > + rq->cpu_epoch_next += n * EPOCH_PERIOD; > + __shr_u64(&rq->cpu_runtime, n); > + } > + > + n = rq->cpu_epoch - pcpu_sched->epoch; > + if (n) { > + pcpu_sched->epoch += n; > + __shr_u64(&pcpu_sched->runtime, n); > + } > +} > + > +static unsigned long fraction_mm_sched(struct rq *rq, > + struct sched_cache_time *pcpu_sched) > +{ > + guard(raw_spinlock_irqsave)(&rq->cpu_epoch_lock); > + > + __update_mm_sched(rq, pcpu_sched); > + > + /* > + * Runtime is a geometric series (r=0.5) and as such will sum to twice > + * the accumulation period, this means the multiplcation here should > + * not overflow. > + */ > + return div64_u64(NICE_0_LOAD * pcpu_sched->runtime, rq->cpu_runtime + 1); > +} > + > +static inline > +void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) > +{ > + struct sched_cache_time *pcpu_sched; > + struct mm_struct *mm = p->mm; > + unsigned long epoch; > + > + if (!sched_cache_enabled()) > + return; > + > + if (p->sched_class != &fair_sched_class) > + return; > + /* > + * init_task, kthreads and user thread created > + * by user_mode_thread() don't have mm. > + */ > + if (!mm || !mm->sc_stat.pcpu_sched) > + return; > + > + pcpu_sched = per_cpu_ptr(p->mm->sc_stat.pcpu_sched, cpu_of(rq)); I do some test,  when i see many pids schedstat,  such as  cat proc/xxx/schedstat, there will cause a panic,maybe the 'p' is exit or others reason,  sorry, i dont record these msg.  how about fix it like below. pcpu_sched = per_cpu_ptr(mm->sc_stat.pcpu_sched, cpu_of(rq)); and "[Patch v4 21/22] -- DO NOT APPLY!!! -- sched/cache/debug: Display the per LLC occupancy for each process via proc fs" has the same problem. > + > + scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) { > + __update_mm_sched(rq, pcpu_sched); > + pcpu_sched->runtime += delta_exec; > + rq->cpu_runtime += delta_exec; > + epoch = rq->cpu_epoch; > + } > + > + /* > + * If this process hasn't hit task_cache_work() for a while, or it > + * has only 1 thread, invalidate its preferred state. > + */ > + if (time_after(epoch, > + READ_ONCE(mm->sc_stat.epoch) + EPOCH_LLC_AFFINITY_TIMEOUT) || > + get_nr_threads(p) <= 1) { > + if (mm->sc_stat.cpu != -1) > + mm->sc_stat.cpu = -1; > + } > +} > + > +static void task_tick_cache(struct rq *rq, struct task_struct *p) > +{ > + struct callback_head *work = &p->cache_work; > + struct mm_struct *mm = p->mm; > + unsigned long epoch; > + > + if (!sched_cache_enabled()) > + return; > + > + if (!mm || !mm->sc_stat.pcpu_sched) > + return; > + > + epoch = rq->cpu_epoch; > + /* avoid moving backwards */ > + if (time_after_eq(mm->sc_stat.epoch, epoch)) > + return; > + > + guard(raw_spinlock)(&mm->sc_stat.lock); > + > + if (work->next == work) { > + task_work_add(p, work, TWA_RESUME); > + WRITE_ONCE(mm->sc_stat.epoch, epoch); > + } > +} > + > +static void task_cache_work(struct callback_head *work) > +{ > + struct task_struct *p = current; > + struct mm_struct *mm = p->mm; > + unsigned long m_a_occ = 0; > + unsigned long curr_m_a_occ = 0; > + int cpu, m_a_cpu = -1; > + cpumask_var_t cpus; > + > + WARN_ON_ONCE(work != &p->cache_work); > + > + work->next = work; > + > + if (p->flags & PF_EXITING) > + return; > + > + if (!zalloc_cpumask_var(&cpus, GFP_KERNEL)) > + return; > + > + scoped_guard (cpus_read_lock) { > + cpumask_copy(cpus, cpu_online_mask); > + > + for_each_cpu(cpu, cpus) { > + /* XXX sched_cluster_active */ > + struct sched_domain *sd = per_cpu(sd_llc, cpu); > + unsigned long occ, m_occ = 0, a_occ = 0; > + int m_cpu = -1, i; > + > + if (!sd) > + continue; > + > + for_each_cpu(i, sched_domain_span(sd)) { > + occ = fraction_mm_sched(cpu_rq(i), > + per_cpu_ptr(mm->sc_stat.pcpu_sched, i)); > + a_occ += occ; > + if (occ > m_occ) { > + m_occ = occ; > + m_cpu = i; > + } > + } > + > + /* > + * Compare the accumulated occupancy of each LLC. The > + * reason for using accumulated occupancy rather than average > + * per CPU occupancy is that it works better in asymmetric LLC > + * scenarios. > + * For example, if there are 2 threads in a 4CPU LLC and 3 > + * threads in an 8CPU LLC, it might be better to choose the one > + * with 3 threads. However, this would not be the case if the > + * occupancy is divided by the number of CPUs in an LLC (i.e., > + * if average per CPU occupancy is used). > + * Besides, NUMA balancing fault statistics behave similarly: > + * the total number of faults per node is compared rather than > + * the average number of faults per CPU. This strategy is also > + * followed here. > + */ > + if (a_occ > m_a_occ) { > + m_a_occ = a_occ; > + m_a_cpu = m_cpu; > + } > + > + if (llc_id(cpu) == llc_id(mm->sc_stat.cpu)) > + curr_m_a_occ = a_occ; > + > + cpumask_andnot(cpus, cpus, sched_domain_span(sd)); > + } > + } > + > + if (m_a_occ > (2 * curr_m_a_occ)) { > + /* > + * Avoid switching sc_stat.cpu too fast. > + * The reason to choose 2X is because: > + * 1. It is better to keep the preferred LLC stable, > + * rather than changing it frequently and cause migrations > + * 2. 2X means the new preferred LLC has at least 1 more > + * busy CPU than the old one(200% vs 100%, eg) > + * 3. 2X is chosen based on test results, as it delivers > + * the optimal performance gain so far. > + */ > + mm->sc_stat.cpu = m_a_cpu; > + } > + > + free_cpumask_var(cpus); > +} > + > +void init_sched_mm(struct task_struct *p) > +{ > + struct callback_head *work = &p->cache_work; > + > + init_task_work(work, task_cache_work); > + work->next = work; > +} > + > +#else /* CONFIG_SCHED_CACHE */ > + > +static inline void account_mm_sched(struct rq *rq, struct task_struct *p, > + s64 delta_exec) { } > + > +void init_sched_mm(struct task_struct *p) { } > + > +static void task_tick_cache(struct rq *rq, struct task_struct *p) { } > + > +#endif /* CONFIG_SCHED_CACHE */ > + > /* > * Used by other classes to account runtime. > */ > @@ -13448,6 +13712,8 @@ static void task_tick_fair(struct rq *rq, struct task_struct *curr, int queued) > if (static_branch_unlikely(&sched_numa_balancing)) > task_tick_numa(rq, curr); > > + task_tick_cache(rq, curr); > + > update_misfit_status(curr, rq); > check_update_overutilized_status(task_rq(curr)); > > diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h > index 43bbf0693cca..0a38bfc704a4 100644 > --- a/kernel/sched/sched.h > +++ b/kernel/sched/sched.h > @@ -1177,6 +1177,12 @@ struct rq { > struct scx_rq scx; > struct sched_dl_entity ext_server; > #endif > +#ifdef CONFIG_SCHED_CACHE > + raw_spinlock_t cpu_epoch_lock ____cacheline_aligned; > + u64 cpu_runtime; > + unsigned long cpu_epoch; > + unsigned long cpu_epoch_next; > +#endif > > struct sched_dl_entity fair_server; > > @@ -4007,6 +4013,14 @@ static inline void mm_cid_switch_to(struct task_struct *prev, struct task_struct > static inline void mm_cid_switch_to(struct task_struct *prev, struct task_struct *next) { } > #endif /* !CONFIG_SCHED_MM_CID */ > > +#ifdef CONFIG_SCHED_CACHE > +static inline bool sched_cache_enabled(void) > +{ > + return false; > +} > +#endif > +extern void init_sched_mm(struct task_struct *p); > + > extern u64 avg_vruntime(struct cfs_rq *cfs_rq); > extern int entity_eligible(struct cfs_rq *cfs_rq, struct sched_entity *se); > static inline