From: Tim Chen <tim.c.chen@linux.intel.com>
To: "Chen, Yu C" <yu.c.chen@intel.com>,
Madadi Vineeth Reddy <vineethr@linux.ibm.com>
Cc: Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Hillf Danton <hdanton@sina.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Jianyong Wu <jianyong.wu@outlook.com>,
Yangyu Chen <cyy@cyyself.name>,
Tingyin Duan <tingyin.duan@gmail.com>,
Vern Hao <vernhao@tencent.com>, Len Brown <len.brown@intel.com>,
Aubrey Li <aubrey.li@intel.com>, Zhao Liu <zhao1.liu@intel.com>,
Chen Yu <yu.chen.surf@gmail.com>,
Adam Li <adamli@os.amperecomputing.com>,
Tim Chen <tim.c.chen@intel.com>,
linux-kernel@vger.kernel.org, haoxing990@gmail.com
Subject: Re: [PATCH 01/19] sched/fair: Add infrastructure for cache-aware load balancing
Date: Wed, 15 Oct 2025 12:32:40 -0700 [thread overview]
Message-ID: <da4d350862807bcf18626009b6fae248475acb1e.camel@linux.intel.com> (raw)
In-Reply-To: <5f140e59-23f9-46dd-bf5e-7bef0d897cd0@intel.com>
On Wed, 2025-10-15 at 12:54 +0800, Chen, Yu C wrote:
> On 10/15/2025 3:12 AM, Madadi Vineeth Reddy wrote:
> > On 11/10/25 23:54, Tim Chen wrote:
> > > From: "Peter Zijlstra (Intel)" <peterz@infradead.org>
> > >
> > > Cache-aware load balancing aims to aggregate tasks with potential
> > > shared resources into the same cache domain. This approach enhances
> > > cache locality, thereby optimizing system performance by reducing
> > > cache misses and improving data access efficiency.
> > >
>
> [snip]
>
> > > +static void __no_profile task_cache_work(struct callback_head *work)
> > > +{
> > > + struct task_struct *p = current;
> > > + struct mm_struct *mm = p->mm;
> > > + unsigned long m_a_occ = 0;
> > > + unsigned long curr_m_a_occ = 0;
> > > + int cpu, m_a_cpu = -1, cache_cpu,
> > > + pref_nid = NUMA_NO_NODE, curr_cpu;
> > > + cpumask_var_t cpus;
> > > +
> > > + WARN_ON_ONCE(work != &p->cache_work);
> > > +
> > > + work->next = work;
> > > +
> > > + if (p->flags & PF_EXITING)
> > > + return;
> > > +
> > > + if (!zalloc_cpumask_var(&cpus, GFP_KERNEL))
> > > + return;
> > > +
> > > + curr_cpu = task_cpu(p);
> > > + cache_cpu = mm->mm_sched_cpu;
> > > +#ifdef CONFIG_NUMA_BALANCING
> > > + if (static_branch_likely(&sched_numa_balancing))
> > > + pref_nid = p->numa_preferred_nid;
> > > +#endif
> > > +
> > > + scoped_guard (cpus_read_lock) {
> > > + get_scan_cpumasks(cpus, cache_cpu,
> > > + pref_nid, curr_cpu);
> > > +
> >
> > IIUC, `get_scan_cpumasks` ORs together the preferred NUMA node, cache CPU's node,
> > and current CPU's node. This could result in scanning multiple nodes, not preferring
> > the NUMA preferred node.
> >
>
> Yes, it is possible, please see comments below.
>
> > > + for_each_cpu(cpu, cpus) {
> > > + /* XXX sched_cluster_active */
> > > + struct sched_domain *sd = per_cpu(sd_llc, cpu);
> > > + unsigned long occ, m_occ = 0, a_occ = 0;
> > > + int m_cpu = -1, i;
> > > +
> > > + if (!sd)
> > > + continue;
> > > +
> > > + for_each_cpu(i, sched_domain_span(sd)) {
> > > + occ = fraction_mm_sched(cpu_rq(i),
> > > + per_cpu_ptr(mm->pcpu_sched, i));
> > > + a_occ += occ;
> > > + if (occ > m_occ) {
> > > + m_occ = occ;
> > > + m_cpu = i;
> > > + }
> > > + }
> > > +
> > > + /*
> > > + * Compare the accumulated occupancy of each LLC. The
> > > + * reason for using accumulated occupancy rather than average
> > > + * per CPU occupancy is that it works better in asymmetric LLC
> > > + * scenarios.
> > > + * For example, if there are 2 threads in a 4CPU LLC and 3
> > > + * threads in an 8CPU LLC, it might be better to choose the one
> > > + * with 3 threads. However, this would not be the case if the
> > > + * occupancy is divided by the number of CPUs in an LLC (i.e.,
> > > + * if average per CPU occupancy is used).
> > > + * Besides, NUMA balancing fault statistics behave similarly:
> > > + * the total number of faults per node is compared rather than
> > > + * the average number of faults per CPU. This strategy is also
> > > + * followed here.
> > > + */
> > > + if (a_occ > m_a_occ) {
> > > + m_a_occ = a_occ;
> > > + m_a_cpu = m_cpu;
> > > + }
> > > +
> > > + if (llc_id(cpu) == llc_id(mm->mm_sched_cpu))
> > > + curr_m_a_occ = a_occ;
> > > +
> > > + cpumask_andnot(cpus, cpus, sched_domain_span(sd));
> > > + }
> >
> > This means NUMA preference has no effect on the selection, except in the
> > unlikely case of exactly equal occupancy across LLCs on different nodes
> > (where iteration order determines the winner).
> >
> > How does it handle when cache locality and memory locality conflict?
> > Shouldn't numa preferred node get preference? Also scanning multiple
> > nodes add overhead, so can restricting it to numa preferred node be
> > better and scan others only when there is no numa preferred node?
> >
>
> Basically, yes, you're right. Ideally, we should prioritize the NUMA
> preferred node as the top priority. There's one case I find hard to
> handle: the NUMA preferred node is per task rather than per process.
> It's possible that different threads of the same process have different
> preferred nodes; as a result, the process-wide preferred LLC could bounce
> between different nodes, which might cause costly task migrations across
> nodes. As a workaround, we tried to keep the scan CPU mask covering the
> process's current preferred LLC to ensure the old preferred LLC is included
> in the candidates. After all, we have a 2X threshold for switching the
> preferred LLC.
If tasks in a process had different preferred nodes, they would
belong to different numa_groups, and majority of their data would
be from different NUMA nodes.
To resolve such conflict, we'll need to change the aggregation of tasks by
process, to aggregation of tasks by numa_group when NUMA balancing is
enabled. This probably makes more sense as tasks in a numa_group
have more shared data and would benefit from co-locating in the
same cache.
Thanks.
Tim
next prev parent reply other threads:[~2025-10-15 19:32 UTC|newest]
Thread overview: 116+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-10-11 18:24 [PATCH 00/19] Cache Aware Scheduling Tim Chen
2025-10-11 18:24 ` [PATCH 01/19] sched/fair: Add infrastructure for cache-aware load balancing Tim Chen
2025-10-14 19:12 ` Madadi Vineeth Reddy
2025-10-15 4:54 ` Chen, Yu C
2025-10-15 19:32 ` Tim Chen [this message]
2025-10-16 3:11 ` Chen, Yu C
2025-10-15 11:54 ` Peter Zijlstra
2025-10-15 16:07 ` Chen, Yu C
2025-10-23 7:26 ` kernel test robot
2025-10-27 4:47 ` K Prateek Nayak
2025-10-27 13:35 ` Chen, Yu C
2025-10-11 18:24 ` [PATCH 02/19] sched/fair: Record per-LLC utilization to guide cache-aware scheduling decisions Tim Chen
2025-10-15 10:15 ` Peter Zijlstra
2025-10-15 16:27 ` Chen, Yu C
2025-10-27 5:01 ` K Prateek Nayak
2025-10-27 14:07 ` Chen, Yu C
2025-10-28 2:50 ` K Prateek Nayak
2025-10-11 18:24 ` [PATCH 03/19] sched/fair: Introduce helper functions to enforce LLC migration policy Tim Chen
2025-10-11 18:24 ` [PATCH 04/19] sched/fair: Introduce a static key to enable cache aware only for multi LLCs Tim Chen
2025-10-15 11:04 ` Peter Zijlstra
2025-10-15 16:25 ` Chen, Yu C
2025-10-15 16:36 ` Shrikanth Hegde
2025-10-15 17:01 ` Chen, Yu C
2025-10-16 7:42 ` Peter Zijlstra
2025-10-17 2:08 ` Chen, Yu C
2025-10-16 7:40 ` Peter Zijlstra
2025-10-27 5:42 ` K Prateek Nayak
2025-10-27 12:56 ` Chen, Yu C
2025-10-27 23:36 ` Tim Chen
2025-10-29 12:36 ` Chen, Yu C
2025-10-28 2:46 ` K Prateek Nayak
2025-10-11 18:24 ` [PATCH 05/19] sched/fair: Add LLC index mapping for CPUs Tim Chen
2025-10-15 11:08 ` Peter Zijlstra
2025-10-15 11:58 ` Peter Zijlstra
2025-10-15 20:12 ` Tim Chen
2025-10-11 18:24 ` [PATCH 06/19] sched/fair: Assign preferred LLC ID to processes Tim Chen
2025-10-14 5:16 ` Chen, Yu C
2025-10-15 11:15 ` Peter Zijlstra
2025-10-16 3:13 ` Chen, Yu C
2025-10-17 4:50 ` Chen, Yu C
2025-10-20 9:41 ` Vern Hao
2025-10-11 18:24 ` [PATCH 07/19] sched/fair: Track LLC-preferred tasks per runqueue Tim Chen
2025-10-15 12:05 ` Peter Zijlstra
2025-10-15 20:03 ` Tim Chen
2025-10-16 7:44 ` Peter Zijlstra
2025-10-16 20:06 ` Tim Chen
2025-10-27 6:04 ` K Prateek Nayak
2025-10-28 15:15 ` Chen, Yu C
2025-10-28 15:46 ` Tim Chen
2025-10-29 4:32 ` K Prateek Nayak
2025-10-29 12:48 ` Chen, Yu C
2025-10-29 4:00 ` K Prateek Nayak
2025-10-28 17:06 ` Tim Chen
2025-10-11 18:24 ` [PATCH 08/19] sched/fair: Introduce per runqueue task LLC preference counter Tim Chen
2025-10-15 12:21 ` Peter Zijlstra
2025-10-15 20:41 ` Tim Chen
2025-10-16 7:49 ` Peter Zijlstra
2025-10-21 8:28 ` Madadi Vineeth Reddy
2025-10-23 6:07 ` Chen, Yu C
2025-10-11 18:24 ` [PATCH 09/19] sched/fair: Count tasks prefering each LLC in a sched group Tim Chen
2025-10-15 12:22 ` Peter Zijlstra
2025-10-15 20:42 ` Tim Chen
2025-10-15 12:25 ` Peter Zijlstra
2025-10-15 20:43 ` Tim Chen
2025-10-27 8:33 ` K Prateek Nayak
2025-10-27 23:19 ` Tim Chen
2025-10-11 18:24 ` [PATCH 10/19] sched/fair: Prioritize tasks preferring destination LLC during balancing Tim Chen
2025-10-15 7:23 ` kernel test robot
2025-10-15 15:08 ` Peter Zijlstra
2025-10-15 21:28 ` Tim Chen
2025-10-15 15:10 ` Peter Zijlstra
2025-10-15 16:03 ` Chen, Yu C
2025-10-24 9:32 ` Aaron Lu
2025-10-27 2:00 ` Chen, Yu C
2025-10-29 9:51 ` Aaron Lu
2025-10-29 13:19 ` Chen, Yu C
2025-10-27 6:29 ` K Prateek Nayak
2025-10-28 12:11 ` Chen, Yu C
2025-10-11 18:24 ` [PATCH 11/19] sched/fair: Identify busiest sched_group for LLC-aware load balancing Tim Chen
2025-10-15 15:24 ` Peter Zijlstra
2025-10-15 21:18 ` Tim Chen
2025-10-11 18:24 ` [PATCH 12/19] sched/fair: Add migrate_llc_task migration type for cache-aware balancing Tim Chen
2025-10-27 9:04 ` K Prateek Nayak
2025-10-27 22:59 ` Tim Chen
2025-10-11 18:24 ` [PATCH 13/19] sched/fair: Handle moving single tasks to/from their preferred LLC Tim Chen
2025-10-11 18:24 ` [PATCH 14/19] sched/fair: Consider LLC preference when selecting tasks for load balancing Tim Chen
2025-10-11 18:24 ` [PATCH 15/19] sched/fair: Respect LLC preference in task migration and detach Tim Chen
2025-10-28 6:02 ` K Prateek Nayak
2025-10-28 11:58 ` Chen, Yu C
2025-10-28 15:30 ` Tim Chen
2025-10-29 4:15 ` K Prateek Nayak
2025-10-29 3:54 ` K Prateek Nayak
2025-10-29 14:23 ` Chen, Yu C
2025-10-29 21:09 ` Tim Chen
2025-10-30 4:19 ` K Prateek Nayak
2025-10-30 20:07 ` Tim Chen
2025-10-31 3:32 ` K Prateek Nayak
2025-10-31 15:17 ` Chen, Yu C
2025-11-03 21:41 ` Tim Chen
2025-11-03 22:07 ` Tim Chen
2025-10-11 18:24 ` [PATCH 16/19] sched/fair: Exclude processes with many threads from cache-aware scheduling Tim Chen
2025-10-23 7:22 ` kernel test robot
2025-10-11 18:24 ` [PATCH 17/19] sched/fair: Disable cache aware scheduling for processes with high thread counts Tim Chen
2025-10-22 17:21 ` Madadi Vineeth Reddy
2025-10-23 6:55 ` Chen, Yu C
2025-10-11 18:24 ` [PATCH 18/19] sched/fair: Avoid cache-aware scheduling for memory-heavy processes Tim Chen
2025-10-15 6:57 ` kernel test robot
2025-10-16 4:44 ` Chen, Yu C
2025-10-11 18:24 ` [PATCH 19/19] sched/fair: Add user control to adjust the tolerance of cache-aware scheduling Tim Chen
2025-10-29 8:07 ` Aaron Lu
2025-10-29 12:54 ` Chen, Yu C
2025-10-14 12:13 ` [PATCH 00/19] Cache Aware Scheduling Madadi Vineeth Reddy
2025-10-14 21:48 ` Tim Chen
2025-10-15 5:38 ` Chen, Yu C
2025-10-15 18:26 ` Madadi Vineeth Reddy
2025-10-16 4:57 ` Chen, Yu C
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=da4d350862807bcf18626009b6fae248475acb1e.camel@linux.intel.com \
--to=tim.c.chen@linux.intel.com \
--cc=adamli@os.amperecomputing.com \
--cc=aubrey.li@intel.com \
--cc=bsegall@google.com \
--cc=cyy@cyyself.name \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=haoxing990@gmail.com \
--cc=hdanton@sina.com \
--cc=jianyong.wu@outlook.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=len.brown@intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@intel.com \
--cc=tingyin.duan@gmail.com \
--cc=vernhao@tencent.com \
--cc=vincent.guittot@linaro.org \
--cc=vineethr@linux.ibm.com \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
--cc=yu.chen.surf@gmail.com \
--cc=zhao1.liu@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome