From: Qais Yousef <qyousef@layalina.io>
To: Tim Chen <tim.c.chen@linux.intel.com>
Cc: Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Madadi Vineeth Reddy <vineethr@linux.ibm.com>,
Hillf Danton <hdanton@sina.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Jianyong Wu <jianyong.wu@outlook.com>,
Yangyu Chen <cyy@cyyself.name>,
Tingyin Duan <tingyin.duan@gmail.com>,
Vern Hao <vernhao@tencent.com>, Vern Hao <haoxing990@gmail.com>,
Len Brown <len.brown@intel.com>, Aubrey Li <aubrey.li@intel.com>,
Zhao Liu <zhao1.liu@intel.com>, Chen Yu <yu.chen.surf@gmail.com>,
Chen Yu <yu.c.chen@intel.com>,
Adam Li <adamli@os.amperecomputing.com>,
Aaron Lu <ziqianlu@bytedance.com>,
Tim Chen <tim.c.chen@intel.com>, Josh Don <joshdon@google.com>,
Gavin Guo <gavinguo@igalia.com>,
Libo Chen <libchen@purestorage.com>,
linux-kernel@vger.kernel.org
Subject: Re: [Patch v4 00/22] Cache aware scheduling
Date: Thu, 16 Apr 2026 01:27:49 +0100 [thread overview]
Message-ID: <20260416002749.muyrcycmtabksav4@airbuntu> (raw)
In-Reply-To: <cover.1775065312.git.tim.c.chen@linux.intel.com>
On 04/01/26 14:52, Tim Chen wrote:
> This patch series introduces infrastructure for cache-aware load
> balancing, with the goal of co-locating tasks that share data within
> the same Last Level Cache (LLC) domain. By improving cache locality,
> the scheduler can reduce cache bouncing and cache misses, ultimately
> improving data access efficiency. The design builds on the initial
> prototype from Peter [1].
>
> This initial implementation treats threads within the same process
> as entities that are likely to share data. During load balancing, the
> scheduler attempts to aggregate such threads onto the same LLC domain
> whenever possible.
>
> Most of the feedback received on v3 has been addressed. Some aspects
> could be enhanced later after the basic cache-aware portion has landed:
>
> There were discussions around grouping tasks using mechanisms other
> than process membership. While we agree that more flexible grouping
> is desirable, this series intentionally focuses on establishing basic
> process-based grouping first, with alternative grouping mechanisms to
> be explored in a follow-on series.
>
> There was also discussion in v3 that the task wakeup path should be used
> to perform cache-aware scheduling. According to previous test results,
> performing task aggregation in the wakeup path introduced task migration
> bouncing. Primarily that was due to the wake up path not having the up
> to date LLC load information. That led to over-aggregation that needed
> to be corrected later in load balancing. Load balancing path was chosen
> as the conservative path to perform task aggregation. The task wakeup
> path will be investigated as a future enhancement.
I posted schedqos announcement yesterday, which I think (hope) would be the
right way to address these concerns about tagging tasks.
https://lore.kernel.org/lkml/20260415000910.2h5misvwc45bdumu@airbuntu/
It would be trivial to add experimental branch to add new QoS flavour to say
NUMA_SENSITIVE etc. I am still trying to think of a generic description to
address a number of use cases (see Execution Profiles in README.md), not just
this particular numa sensitive one, but the experimental branch should help
iterate and drive the kernel development for wake up path + push lb instead of
using load balance which I really doubt will work well in practice since this
is slow to react, and you're relying on overcommitting the system by default by
making every task of every process data dependent and require it to be
co-located. I think in practice admins will care about specific applications to
be kept within a single LLC, and if they are willing to spend the effort, they
can tag specific tasks of a specific application.
We are trying to make sure we have one coherent story to define these type of
QoS requirements, and delegate to userspace to make these decisions/policies.
The current line of thinking is that wakeup + push lb should be generally good
to address the different needs for various placement requirements. I understood
you believe the same. If not, it would be good so we can think how to further
generalize.
Also QoS IMHO should be viewed as a scarce resource. For best effort delivery
(which is the best we can do in reality, this is not hard real time system), it
is easier to provide good best effort when the average noise level is low, ie:
few tasks are required to be kept within the same LLC. If we overcommit often,
we will crumble often. So IMHO the key is to delegate to userspace to tag, and
make them take responsibility of handling potential overcommit and decide which
workload is really important to tag and which one they can let go of or move to
another machine to get their desired perf/latencies.
For the kernel interface to tag tasks and set a cookie, I plan to rebase and
repost [1] as I need it for rampup multiplier to help counter DVFS related
latencies and slow migration in HMP systems. The idea was for it to be generic
and flexible to add whatever; in this case I think we need to add a QoS to tell
scheduler these tasks are data co-dependent with a unique cookie, which implies
that they need to stay within the same LLC, and can be extended to help keep
these tasks within the same L2 or L1 if it makes sense (ie: they are small and
can be packed on the same CPU).
I am not opposed to merging this first, but I think the load balanced based
approach is wrong and if merged must be removed later in favour of wakeup
+ push lb one with user space based tagging.
Based on what Peter said the wake up path is trivial to add, push lb is almost
ready, and I hope we have the tools to auto tag processes/tasks now to
potentially try to work on this approach first instead.
[1] https://lore.kernel.org/lkml/20240820163512.1096301-11-qyousef@layalina.io/
next prev parent reply other threads:[~2026-04-16 0:27 UTC|newest]
Thread overview: 72+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-04-01 21:52 Tim Chen
2026-04-01 21:52 ` [Patch v4 01/22] sched/cache: Introduce infrastructure for cache-aware load balancing Tim Chen
2026-04-09 12:41 ` Peter Zijlstra
2026-04-09 19:21 ` Tim Chen
2026-04-09 23:00 ` Peter Zijlstra
2026-04-10 6:30 ` Chen, Yu C
2026-04-15 2:06 ` Vern Hao
2026-04-15 3:34 ` Chen, Yu C
2026-09-14 17:50 ` Zenghui Yu
2026-09-14 23:12 ` Tim Chen
2026-09-18 7:14 ` Zenghui Yu
2026-04-01 21:52 ` [Patch v4 02/22] sched/cache: Limit the scan number of CPUs when calculating task occupancy Tim Chen
2026-04-09 13:17 ` Luo Gengkun
2026-04-09 13:41 ` Peter Zijlstra
2026-04-10 10:12 ` Luo Gengkun
2026-04-10 7:29 ` Chen, Yu C
2026-04-10 10:20 ` Luo Gengkun
2026-04-10 17:12 ` Tim Chen
2026-04-10 17:27 ` Chen, Yu C
2026-04-13 7:23 ` [RFC PATCH] sched/fair: dynamically scale the period of cache work Jianyong Wu
2026-04-13 8:38 ` Chen, Yu C
2026-04-13 11:27 ` Jianyong Wu
2026-04-15 3:31 ` Chen, Yu C
2026-04-16 3:39 ` Jianyong Wu
2026-04-15 17:22 ` Tim Chen
2026-04-16 6:50 ` Jianyong Wu
2026-04-14 15:07 ` [PATCH v2] sched/cache: Reduce the overhead of task_cache_work by only scan the visisted cpus Luo Gengkun
2026-04-15 3:10 ` Chen, Yu C
2026-04-18 9:01 ` Luo Gengkun
2026-04-20 7:53 ` Chen, Yu C
2026-04-23 8:54 ` [PATCH v3] " Luo Gengkun
2026-04-01 21:52 ` [Patch v4 03/22] sched/cache: Record per LLC utilization to guide cache aware scheduling decisions Tim Chen
2026-04-01 21:52 ` [Patch v4 04/22] sched/cache: Introduce helper functions to enforce LLC migration policy Tim Chen
2026-04-01 21:52 ` [Patch v4 05/22] sched/cache: Make LLC id continuous Tim Chen
2026-04-01 21:52 ` [Patch v4 06/22] sched/cache: Assign preferred LLC ID to processes Tim Chen
2026-04-01 21:52 ` [Patch v4 07/22] sched/cache: Track LLC-preferred tasks per runqueue Tim Chen
2026-04-01 21:52 ` [Patch v4 08/22] sched/cache: Introduce per CPU's tasks LLC preference counter Tim Chen
2026-04-01 21:52 ` [Patch v4 09/22] sched/cache: Calculate the percpu sd task LLC preference Tim Chen
2026-04-01 21:52 ` [Patch v4 10/22] sched/cache: Count tasks prefering destination LLC in a sched group Tim Chen
2026-04-01 21:52 ` [Patch v4 11/22] sched/cache: Check local_group only once in update_sg_lb_stats() Tim Chen
2026-04-01 21:52 ` [Patch v4 12/22] sched/cache: Prioritize tasks preferring destination LLC during balancing Tim Chen
2026-04-01 21:52 ` [Patch v4 13/22] sched/cache: Add migrate_llc_task migration type for cache-aware balancing Tim Chen
2026-04-01 21:52 ` [Patch v4 14/22] sched/cache: Handle moving single tasks to/from their preferred LLC Tim Chen
2026-04-01 21:52 ` [Patch v4 15/22] sched/cache: Respect LLC preference in task migration and detach Tim Chen
2026-04-01 21:52 ` [Patch v4 16/22] sched/cache: Disable cache aware scheduling for processes with high thread counts Tim Chen
2026-04-09 12:43 ` Peter Zijlstra
2026-04-09 19:27 ` Tim Chen
2026-04-01 21:52 ` [Patch v4 17/22] sched/cache: Avoid cache-aware scheduling for memory-heavy processes Tim Chen
2026-04-09 12:46 ` Peter Zijlstra
2026-04-09 12:55 ` Peter Zijlstra
2026-04-10 8:59 ` Chen, Yu C
2026-04-10 9:20 ` Peter Zijlstra
2026-04-01 21:52 ` [Patch v4 18/22] sched/cache: Enable cache aware scheduling for multi LLCs NUMA node Tim Chen
2026-04-09 13:37 ` Peter Zijlstra
2026-04-09 19:39 ` Tim Chen
2026-04-01 21:52 ` [Patch v4 19/22] sched/cache: Allow the user space to turn on and off cache aware scheduling Tim Chen
2026-04-01 21:52 ` [Patch v4 20/22] sched/cache: Add user control to adjust the aggressiveness of cache-aware scheduling Tim Chen
2026-04-01 21:52 ` [Patch v4 21/22] -- DO NOT APPLY!!! -- sched/cache/debug: Display the per LLC occupancy for each process via proc fs Tim Chen
2026-04-01 21:52 ` [Patch v4 22/22] -- DO NOT APPLY!!! -- sched/cache/debug: Add ftrace to track the load balance statistics Tim Chen
2026-04-09 13:54 ` [Patch v4 00/22] Cache aware scheduling Peter Zijlstra
2026-04-09 20:02 ` Tim Chen
2026-04-14 3:20 ` Duan Tingyin
2026-04-15 17:35 ` Tim Chen
2026-04-16 0:27 ` Qais Yousef [this message]
2026-04-20 9:01 ` Chen, Yu C
2026-04-21 0:34 ` Qais Yousef
2026-04-21 20:57 ` Tim Chen
2026-04-23 15:06 ` Qais Yousef
2026-04-23 16:48 ` Chen, Yu C
2026-04-25 0:05 ` Qais Yousef
2026-04-23 17:17 ` Chen, Yu C
2026-04-25 0:14 ` Qais Yousef
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260416002749.muyrcycmtabksav4@airbuntu \
--to=qyousef@layalina.io \
--cc=adamli@os.amperecomputing.com \
--cc=aubrey.li@intel.com \
--cc=bsegall@google.com \
--cc=cyy@cyyself.name \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=gavinguo@igalia.com \
--cc=haoxing990@gmail.com \
--cc=hdanton@sina.com \
--cc=jianyong.wu@outlook.com \
--cc=joshdon@google.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=len.brown@intel.com \
--cc=libchen@purestorage.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@intel.com \
--cc=tim.c.chen@linux.intel.com \
--cc=tingyin.duan@gmail.com \
--cc=vernhao@tencent.com \
--cc=vincent.guittot@linaro.org \
--cc=vineethr@linux.ibm.com \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
--cc=yu.chen.surf@gmail.com \
--cc=zhao1.liu@intel.com \
--cc=ziqianlu@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®