From: "Chen, Yu C" <yu.c.chen@intel.com>
To: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Madadi Vineeth Reddy <vineethr@linux.ibm.com>,
"Hillf Danton" <hdanton@sina.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
"Jianyong Wu" <jianyong.wu@outlook.com>,
Yangyu Chen <cyy@cyyself.name>,
Tingyin Duan <tingyin.duan@gmail.com>,
Vern Hao <vernhao@tencent.com>, Vern Hao <haoxing990@gmail.com>,
Len Brown <len.brown@intel.com>, Aubrey Li <aubrey.li@intel.com>,
Zhao Liu <zhao1.liu@intel.com>, Chen Yu <yu.chen.surf@gmail.com>,
Ingo Molnar <mingo@redhat.com>,
Adam Li <adamli@os.amperecomputing.com>,
Aaron Lu <ziqianlu@bytedance.com>,
Tim Chen <tim.c.chen@intel.com>, <linux-kernel@vger.kernel.org>,
Vincent Guittot <vincent.guittot@linaro.org>,
Peter Zijlstra <peterz@infradead.org>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>,
Tim Chen <tim.c.chen@linux.intel.com>
Subject: Re: [PATCH v2 04/23] sched/cache: Make LLC id continuous
Date: Wed, 24 Dec 2025 15:08:47 +0800 [thread overview]
Message-ID: <d4110256-7f7d-45ef-88b0-01fb12d07308@intel.com> (raw)
In-Reply-To: <2f9165fe-55f2-4919-be01-e9d2cdd1f960@amd.com>
Hello Prateek,
On 12/23/2025 1:31 PM, K Prateek Nayak wrote:
> Hello Tim, Chenyu,
>
> On 12/4/2025 4:37 AM, Tim Chen wrote:
>> +/*
>> + * Assign continuous llc id for the CPU, and return
>> + * the assigned llc id.
>> + */
>> +static int update_llc_id(struct sched_domain *sd,
>> + int cpu)
>> +{
>> + int id = per_cpu(sd_llc_id, cpu), i;
>> +
>> + if (id >= 0)
>> + return id;
>> +
>> + if (sd) {
>> + /* Look for any assigned id and reuse it.*/
>> + for_each_cpu(i, sched_domain_span(sd)) {
>> + id = per_cpu(sd_llc_id, i);
>> +
>> + if (id >= 0) {
>> + per_cpu(sd_llc_id, cpu) = id;
>> + return id;
>> + }
>> + }
>> + }
>
> I don't really like tying this down to the sched_domain span since
> partition and other weirdness can cause the max_llc count to go
> unnecessarily high. The tl->mask() (from sched_domain_topology_level)
> should give the mask considering all online CPUs and not bothering
> about cpusets.
OK, using the topology_level's mask (tl's mask) should allow us to
skip the cpuset partition. I just wanted to check if your concern
is about the excessive number of sd_llc_ids caused by the cpuset?
I was under the impression that without this patch, llc_ids are
unique across different partitions.
For example, on vanilla kernel without cache_aware,
suppose 1 LLC has CPU0,1,2,3. Before partition, all
CPUs have the same llc_id 0. Then create a new partition,
mkdir -p /sys/fs/cgroup/cgroup0
echo "3" > /sys/fs/cgroup/cgroup0/cpuset.cpus
echo root > /sys/fs/cgroup/cgroup0/cpuset.cpus.partition
CPU0,1,2 share llc_id 0, and CPU3 has a dedicated llc_id 3.
Do you suggest to let CPU3 reuse llc_id 0, so as to save
more llc_id space?
>
> How about something like:
>
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 5b17d8e3cb55..c19b1c4e6472 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -8270,6 +8270,18 @@ static void cpuset_cpu_active(void)
> static void cpuset_cpu_inactive(unsigned int cpu)
> {
> if (!cpuhp_tasks_frozen) {
> + /*
> + * This is necessary since offline CPUs are
> + * taken out of the tl->mask() and a newly
> + * onlined CPU in same LLC will not realize
> + * whether it should reuse the LLC ID owned
> + * by an offline CPU without knowing the
> + * LLC association.
> + *
> + * Safe to release the reference if this is
> + * the last CPU in the LLC going offline.
> + */
> + sched_domain_free_llc_id(cpu);
I'm OK with replacing the domain based cpumask by the topology_level
mask, just wondering whether re-using the llc_id would increase
the risk of race condition - it is possible that, a CPU has different
llc_ids before/after online/offline. Can we assign/reserve a "static"
llc_id for each CPU, whether it is online or offline? In this way,
we don't need to worry about the data synchronization when using
llc_id(). For example, I can think of adjusting the data in
percpu nr_pref_llc[max_llcs] on every CPU whenever a CPU gets
offline/online.
> cpuset_update_active_cpus();
> } else {
> num_cpus_frozen++;
> diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
> index 41caa22e0680..1378a1cfad18 100644
> --- a/kernel/sched/debug.c
> +++ b/kernel/sched/debug.c
> @@ -631,6 +631,7 @@ void update_sched_domain_debugfs(void)
> i++;
> }
>
> + debugfs_create_u32("llc_id", 0444, d_cpu, (u32 *)per_cpu_ptr(&sd_llc_id, cpu));
> __cpumask_clear_cpu(cpu, sd_sysctl_cpus);
> }
> }
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index 3ceaa9dc9a9e..69fad88b57d8 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -2142,6 +2142,7 @@ extern int group_balance_cpu(struct sched_group *sg);
>
> extern void update_sched_domain_debugfs(void);
> extern void dirty_sched_domain_sysctl(int cpu);
> +void sched_domain_free_llc_id(int cpu);
>
> extern int sched_update_scaling(void);
>
> diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
> index cf643a5ddedd..d6e134767f30 100644
> --- a/kernel/sched/topology.c
> +++ b/kernel/sched/topology.c
> @@ -20,6 +20,46 @@ void sched_domains_mutex_unlock(void)
> /* Protected by sched_domains_mutex: */
> static cpumask_var_t sched_domains_tmpmask;
> static cpumask_var_t sched_domains_tmpmask2;
> +static cpumask_var_t sched_llc_id_alloc_mask;
> +DEFINE_PER_CPU(int, sd_llc_id) = -1;
> +static int max_llcs = 0;
> +
> +static inline int sched_domain_alloc_llc_id(void)
> +{
> + int llc_id;
> +
> + lockdep_assert_held(&sched_domains_mutex);
> +
> + llc_id = cpumask_first_zero(sched_llc_id_alloc_mask);
> + BUG_ON((unsigned int)llc_id >= nr_cpumask_bits);
> + cpumask_set_cpu(llc_id, sched_llc_id_alloc_mask);
> + ++max_llcs;
> +
> + return llc_id;
> +}
> +
> +void sched_domain_free_llc_id(int cpu)
> +{
> + int i, llc_id = per_cpu(sd_llc_id, cpu);
> + bool found = false;
> +
> + lockdep_assert_cpus_held(); /* For cpu_active_mask. */
> + guard(mutex)(&sched_domains_mutex);
> +
> + per_cpu(sd_llc_id, cpu) = -1;
> + for_each_cpu(i, cpu_active_mask) {
> + if (per_cpu(sd_llc_id, i) == llc_id) {
> + found = true;
> + break;
> + }
> + }
> +
> + /* Allow future hotplugs to claim this ID */
> + if (!found) {
> + cpumask_clear_cpu(llc_id, sched_llc_id_alloc_mask);
> + --max_llcs;
Maybe only allow increasing the value of max_llcs when a new LLC
is detected. That says, max_llcs represents the total number of LLCs
that have ever been detected, even if some of the corresponding
CPUs have been taken offline via runtime hotplug. In this way, the
data synchronization might be simpler, maybe trade additional memory
space for code simplicity?
> + }
> +}
>
> static int __init sched_debug_setup(char *str)
> {
> @@ -658,7 +698,6 @@ static void destroy_sched_domains(struct sched_domain *sd)
> */
> DEFINE_PER_CPU(struct sched_domain __rcu *, sd_llc);
> DEFINE_PER_CPU(int, sd_llc_size);
> -DEFINE_PER_CPU(int, sd_llc_id);
> DEFINE_PER_CPU(int, sd_share_id);
> DEFINE_PER_CPU(struct sched_domain_shared __rcu *, sd_llc_shared);
> DEFINE_PER_CPU(struct sched_domain __rcu *, sd_numa);
> @@ -684,7 +723,6 @@ static void update_top_cache_domain(int cpu)
>
> rcu_assign_pointer(per_cpu(sd_llc, cpu), sd);
> per_cpu(sd_llc_size, cpu) = size;
> - per_cpu(sd_llc_id, cpu) = id;
> rcu_assign_pointer(per_cpu(sd_llc_shared, cpu), sds);
>
> sd = lowest_flag_domain(cpu, SD_CLUSTER);
> @@ -2567,10 +2605,35 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
>
> /* Set up domains for CPUs specified by the cpu_map: */
> for_each_cpu(i, cpu_map) {
> - struct sched_domain_topology_level *tl;
> + struct sched_domain_topology_level *tl, *tl_llc = NULL;
> + bool done = false;
>
> sd = NULL;
> for_each_sd_topology(tl) {
> + int flags = 0;
> +
> + if (tl->sd_flags)
> + flags = (*tl->sd_flags)();
> +
> + if (flags & SD_SHARE_LLC) {
> + tl_llc = tl;
> +
> + /*
> + * Entire cpu_map has been covered. We are
> + * traversing only to find the highest
> + * SD_SHARE_LLC level.
> + */
> + if (done)
> + continue;
> + }
> +
> + /*
> + * Since SD_SHARE_LLC is SDF_SHARED_CHILD, we can
> + * safely break out if the entire cpu_map has been
> + * covered by a child domain.
> + */
> + if (done)
> + break;
>
> sd = build_sched_domain(tl, cpu_map, attr, sd, i);
>
> @@ -2579,7 +2642,41 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
> if (tl == sched_domain_topology)
> *per_cpu_ptr(d.sd, i) = sd;
> if (cpumask_equal(cpu_map, sched_domain_span(sd)))
> - break;
> + done = true;
> + }
> +
> + /* First time visiting this CPU. Assign the llc_id. */
> + if (per_cpu(sd_llc_id, i) == -1) {
> + int j, llc_id = -1;
> +
> + /*
> + * In case there are no SD_SHARE_LLC domains,
> + * each CPU gets its own llc_id. Find the first
> + * free bit on the mask and use it.
> + */
> + if (!tl_llc) {
> + per_cpu(sd_llc_id, i) = sched_domain_alloc_llc_id();
> + continue;
> + }
> +
> + /*
> + * Visit all the CPUs of the LLC irrespective of the
> + * partition constraints and find if any of them have
> + * a valid llc_id.
> + */
> + for_each_cpu(j, tl_llc->mask(tl, i)) {
This is doable, we can use tl rather than domain's mask to
share llc_id among partitions.
> + llc_id = per_cpu(sd_llc_id, j);
> +
> + /* Found a valid llc_id for CPU's LLC. */
> + if (llc_id != -1)
> + break;
> + }
> +
> + /* Valid llc_id not found. Allocate a new one. */
> + if (llc_id == -1)
> + llc_id = sched_domain_alloc_llc_id();
> +
> + per_cpu(sd_llc_id, i) = llc_id;
> }
> }
>
> @@ -2759,6 +2856,7 @@ int __init sched_init_domains(const struct cpumask *cpu_map)
>
> zalloc_cpumask_var(&sched_domains_tmpmask, GFP_KERNEL);
> zalloc_cpumask_var(&sched_domains_tmpmask2, GFP_KERNEL);
> + zalloc_cpumask_var(&sched_llc_id_alloc_mask, GFP_KERNEL);
> zalloc_cpumask_var(&fallback_doms, GFP_KERNEL);
>
> arch_update_cpu_topology();
> ---
>
> AFAICT, "sd_llc_id" isn't compared across different partitions so having
> the CPUs that are actually associated with same physical LLC but across
> different partitions sharing the same "sd_llc_id" shouldn't be a problem.
>
> Thoughts?
>
This means cpus_share_resources(int this_cpu, int that_cpu)
should be invoked when this_cpu and that_cpu belong to the same
partition.
In this way, we do not alter the context of cpus_share_resources(). We can
conduct an audit of the places where cpus_share_resources() is used.
Happy holidays,
Chenyu
next prev parent reply other threads:[~2025-12-24 7:09 UTC|newest]
Thread overview: 120+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-12-03 23:07 [PATCH v2 00/23] Cache aware scheduling Tim Chen
2025-12-03 23:07 ` [PATCH v2 01/23] sched/cache: Introduce infrastructure for cache-aware load balancing Tim Chen
2025-12-09 11:12 ` Peter Zijlstra
2025-12-09 21:39 ` Tim Chen
2025-12-10 9:37 ` Peter Zijlstra
2025-12-10 13:57 ` Chen, Yu C
2025-12-10 15:11 ` Peter Zijlstra
2025-12-11 9:03 ` Vern Hao
2025-12-16 6:12 ` Chen, Yu C
2025-12-17 1:17 ` Vern Hao
2026-01-15 21:47 ` Tim Chen
[not found] ` <fbf52d91-0605-4608-b9cc-e8cc56115fd5@gmail.com>
2025-12-16 22:30 ` Tim Chen
2025-12-03 23:07 ` [PATCH v2 02/23] sched/cache: Record per-LLC utilization to guide cache-aware scheduling decisions Tim Chen
2025-12-09 11:21 ` Peter Zijlstra
2025-12-10 14:02 ` Chen, Yu C
2025-12-10 15:13 ` Peter Zijlstra
2025-12-10 23:58 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 03/23] sched/cache: Introduce helper functions to enforce LLC migration policy Tim Chen
2026-01-22 18:13 ` Yangyu Chen
2026-01-22 20:43 ` Tim Chen
2025-12-03 23:07 ` [PATCH v2 04/23] sched/cache: Make LLC id continuous Tim Chen
2025-12-09 11:58 ` Peter Zijlstra
2025-12-15 20:49 ` Tim Chen
2025-12-16 5:31 ` Chen, Yu C
2025-12-16 19:53 ` Tim Chen
2025-12-17 5:25 ` Chen, Yu C
2025-12-23 5:31 ` K Prateek Nayak
2025-12-24 7:08 ` Chen, Yu C [this message]
2025-12-24 8:19 ` K Prateek Nayak
2025-12-24 9:46 ` Chen, Yu C
2025-12-26 3:17 ` K Prateek Nayak
2025-12-03 23:07 ` [PATCH v2 05/23] sched/cache: Assign preferred LLC ID to processes Tim Chen
2025-12-09 12:11 ` Peter Zijlstra
2025-12-09 22:34 ` Tim Chen
2025-12-12 3:34 ` Vern Hao
2025-12-15 19:32 ` Tim Chen
2025-12-19 4:01 ` Vern Hao
2025-12-24 10:20 ` Chen, Yu C
2026-01-07 4:49 ` Jianyong Wu
2026-01-07 8:38 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 06/23] sched/cache: Track LLC-preferred tasks per runqueue Tim Chen
2025-12-09 12:16 ` Peter Zijlstra
2025-12-09 22:55 ` Tim Chen
2025-12-10 9:42 ` Peter Zijlstra
2025-12-16 0:20 ` Chen, Yu C
2025-12-17 10:04 ` Vern Hao
2025-12-17 12:37 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 07/23] sched/cache: Introduce per runqueue task LLC preference counter Tim Chen
2025-12-09 13:06 ` Peter Zijlstra
2025-12-09 23:17 ` Tim Chen
2025-12-10 12:43 ` Peter Zijlstra
2025-12-10 18:36 ` Tim Chen
2025-12-10 12:51 ` Peter Zijlstra
2025-12-10 18:49 ` Tim Chen
2025-12-11 10:31 ` Peter Zijlstra
2025-12-15 19:21 ` Tim Chen
2025-12-16 22:45 ` Tim Chen
2025-12-03 23:07 ` [PATCH v2 08/23] sched/cache: Calculate the per runqueue task LLC preference Tim Chen
2025-12-03 23:07 ` [PATCH v2 09/23] sched/cache: Count tasks prefering destination LLC in a sched group Tim Chen
2025-12-10 12:52 ` Peter Zijlstra
2025-12-10 14:05 ` Chen, Yu C
2025-12-10 15:16 ` Peter Zijlstra
2025-12-10 19:00 ` Tim Chen
2025-12-10 23:50 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 10/23] sched/cache: Check local_group only once in update_sg_lb_stats() Tim Chen
2025-12-03 23:07 ` [PATCH v2 11/23] sched/cache: Prioritize tasks preferring destination LLC during balancing Tim Chen
2025-12-03 23:07 ` [PATCH v2 12/23] sched/cache: Add migrate_llc_task migration type for cache-aware balancing Tim Chen
2025-12-10 13:32 ` Peter Zijlstra
2025-12-16 0:52 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 13/23] sched/cache: Handle moving single tasks to/from their preferred LLC Tim Chen
2025-12-03 23:07 ` [PATCH v2 14/23] sched/cache: Consider LLC preference when selecting tasks for load balancing Tim Chen
2025-12-10 15:58 ` Peter Zijlstra
2025-12-03 23:07 ` [PATCH v2 15/23] sched/cache: Respect LLC preference in task migration and detach Tim Chen
2025-12-10 16:30 ` Peter Zijlstra
2025-12-16 7:30 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 16/23] sched/cache: Introduce sched_cache_present to enable cache aware scheduling for multi LLCs NUMA node Tim Chen
2025-12-10 16:32 ` Peter Zijlstra
2025-12-10 16:52 ` Peter Zijlstra
2025-12-16 7:36 ` Chen, Yu C
2025-12-16 7:31 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 17/23] sched/cache: Record the number of active threads per process for cache-aware scheduling Tim Chen
2025-12-10 16:51 ` Peter Zijlstra
2025-12-16 7:40 ` Chen, Yu C
2025-12-17 9:40 ` Aaron Lu
2025-12-17 12:51 ` Chen, Yu C
2025-12-19 3:32 ` Aaron Lu
2025-12-03 23:07 ` [PATCH v2 18/23] sched/cache: Disable cache aware scheduling for processes with high thread counts Tim Chen
2025-12-03 23:07 ` [PATCH v2 19/23] sched/cache: Avoid cache-aware scheduling for memory-heavy processes Tim Chen
2025-12-18 3:59 ` Vern Hao
2025-12-18 8:32 ` Chen, Yu C
2025-12-18 9:42 ` Vern Hao
2025-12-19 3:14 ` K Prateek Nayak
2025-12-19 12:55 ` Chen, Yu C
2025-12-22 2:49 ` Vern Hao
2025-12-22 2:19 ` Vern Hao
2025-12-03 23:07 ` [PATCH v2 20/23] sched/cache: Add user control to adjust the parameters of cache-aware scheduling Tim Chen
2025-12-10 17:02 ` Peter Zijlstra
2025-12-16 7:42 ` Chen, Yu C
2025-12-19 4:14 ` Vern Hao
2025-12-19 13:21 ` Chen, Yu C
2025-12-19 13:39 ` Chen, Yu C
2025-12-23 12:12 ` Yangyu Chen
2025-12-23 16:44 ` Yangyu Chen
2025-12-24 3:28 ` Yangyu Chen
2025-12-24 7:51 ` Chen, Yu C
2025-12-24 12:15 ` Yangyu Chen
2026-01-15 10:03 ` Jianyong Wu
2026-01-15 12:13 ` Chen, Yu C
2026-01-21 15:21 ` Yangyu Chen
2026-01-21 15:38 ` Chen, Yu C
2025-12-03 23:07 ` [PATCH v2 21/23] -- DO NOT APPLY!!! -- sched/cache/stats: Add schedstat for cache aware load balancing Tim Chen
2025-12-19 5:03 ` Yangyu Chen
2025-12-19 14:41 ` Chen, Yu C
2025-12-19 14:48 ` Yangyu Chen
2025-12-03 23:07 ` [PATCH v2 22/23] -- DO NOT APPLY!!! -- sched/cache/debug: Add ftrace to track the load balance statistics Tim Chen
2025-12-03 23:07 ` [PATCH v2 23/23] -- DO NOT APPLY!!! -- sched/cache/debug: Display the per LLC occupancy for each process via proc fs Tim Chen
2025-12-17 9:59 ` Aaron Lu
2025-12-17 13:01 ` Chen, Yu C
2025-12-19 3:19 ` [PATCH v2 00/23] Cache aware scheduling Aaron Lu
2025-12-19 13:04 ` Chen, Yu C
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d4110256-7f7d-45ef-88b0-01fb12d07308@intel.com \
--to=yu.c.chen@intel.com \
--cc=adamli@os.amperecomputing.com \
--cc=aubrey.li@intel.com \
--cc=bsegall@google.com \
--cc=cyy@cyyself.name \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=haoxing990@gmail.com \
--cc=hdanton@sina.com \
--cc=jianyong.wu@outlook.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=len.brown@intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@intel.com \
--cc=tim.c.chen@linux.intel.com \
--cc=tingyin.duan@gmail.com \
--cc=vernhao@tencent.com \
--cc=vincent.guittot@linaro.org \
--cc=vineethr@linux.ibm.com \
--cc=vschneid@redhat.com \
--cc=yu.chen.surf@gmail.com \
--cc=zhao1.liu@intel.com \
--cc=ziqianlu@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome