From: Adam Li <adamli@os.amperecomputing.com>
To: Chen Yu <yu.c.chen@intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>,
Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Libo Chen <libo.chen@oracle.com>,
Madadi Vineeth Reddy <vineethr@linux.ibm.com>,
Hillf Danton <hdanton@sina.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Jianyong Wu <jianyong.wu@outlook.com>,
Yangyu Chen <cyy@cyyself.name>,
Tingyin Duan <tingyin.duan@gmail.com>,
Vern Hao <vernhao@tencent.com>, Len Brown <len.brown@intel.com>,
Tim Chen <tim.c.chen@linux.intel.com>,
Aubrey Li <aubrey.li@intel.com>, Zhao Liu <zhao1.liu@intel.com>,
Chen Yu <yu.chen.surf@gmail.com>,
linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH v4 08/28] sched: Set up LLC indexing
Date: Fri, 26 Sep 2025 14:14:19 +0800 [thread overview]
Message-ID: <baa45b9a-d8ef-4652-a3c2-83596216fcb6@os.amperecomputing.com> (raw)
In-Reply-To: <959d897daadc28b8115c97df04eec2af0fd79c5d.1754712565.git.tim.c.chen@linux.intel.com>
Hi Chen Yu,
I tested the patch set on AmpereOne CPU with 192 cores.
With certain firmware setting, each core has its own L1/L2 cache.
But *no* cores share LLC (L3). So *no* schedule domain
has flag 'SD_SHARE_LLC'.
With this topology:
per_cpu(sd_llc_id, cpu) is actually the cpu id (0-191).
And kernel bug will be triggered at:
'BUG_ON(idx > MAX_LLC)'
Please see details bellow.
The bug will disappear if setting 'MAX_LLC' to 192.
But I think we might disable CAS(cache aware scheduling)
if no domain has 'SD_SHARE_LLC'.
On 8/9/2025 1:03 PM, Chen Yu wrote:
> From: Tim Chen <tim.c.chen@linux.intel.com>
>
> Prepare for indexing arrays that track in each run queue: the number
> of tasks preferring current LLC and each of the other LLC.
>
> The reason to introduce LLC index is because the per LLC-scope data
> is needed to do cache aware load balancing. However, the native lld_id
> is usually the first CPU of that LLC domain, which is not continuous,
> which might waste the space if the per LLC-scope data is stored
> in an array (in current implementation).
>
> In the future, this LLC index could be removed after
> the native llc_id is used as the key to search into xarray based
> array.
>
> Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
> ---
> include/linux/sched.h | 3 +++
> kernel/sched/fair.c | 12 ++++++++++++
> kernel/sched/sched.h | 2 ++
> kernel/sched/topology.c | 29 +++++++++++++++++++++++++++++
> 4 files changed, 46 insertions(+)
>
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index 02ff8b8be25b..81d92e8097f5 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -809,6 +809,9 @@ struct kmap_ctrl {
> #endif
> };
>
> +/* XXX need fix to not use magic number */
> +#define MAX_LLC 64
The bug will disappear if setting 'MAX_LLC' to 192.
> +
> struct task_struct {
> #ifdef CONFIG_THREAD_INFO_IN_TASK
> /*
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 3128dbcf0a36..f5075d287c51 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1183,6 +1183,18 @@ static int llc_id(int cpu)
> return per_cpu(sd_llc_id, cpu);
> }
>
> +/*
> + * continuous index.
> + * TBD: replace by xarray with key llc_id()
> + */
> +static inline int llc_idx(int cpu)
> +{
> + if (cpu < 0)
> + return -1;
> +
> + return per_cpu(sd_llc_idx, cpu);
> +}
> +
> void mm_init_sched(struct mm_struct *mm, struct mm_sched __percpu *_pcpu_sched)
> {
> unsigned long epoch;
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index 83552aab74fb..c37c74dfce25 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -2056,6 +2056,7 @@ static inline struct sched_domain *lowest_flag_domain(int cpu, int flag)
> DECLARE_PER_CPU(struct sched_domain __rcu *, sd_llc);
> DECLARE_PER_CPU(int, sd_llc_size);
> DECLARE_PER_CPU(int, sd_llc_id);
> +DECLARE_PER_CPU(int, sd_llc_idx);
> DECLARE_PER_CPU(int, sd_share_id);
> DECLARE_PER_CPU(struct sched_domain_shared __rcu *, sd_llc_shared);
> DECLARE_PER_CPU(struct sched_domain __rcu *, sd_numa);
> @@ -2064,6 +2065,7 @@ DECLARE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
>
> extern struct static_key_false sched_asym_cpucapacity;
> extern struct static_key_false sched_cluster_active;
> +extern int max_llcs;
>
> static __always_inline bool sched_asym_cpucap_active(void)
> {
> diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
> index b958fe48e020..91a2b7f65fee 100644
> --- a/kernel/sched/topology.c
> +++ b/kernel/sched/topology.c
> @@ -657,6 +657,7 @@ static void destroy_sched_domains(struct sched_domain *sd)
> DEFINE_PER_CPU(struct sched_domain __rcu *, sd_llc);
> DEFINE_PER_CPU(int, sd_llc_size);
> DEFINE_PER_CPU(int, sd_llc_id);
> +DEFINE_PER_CPU(int, sd_llc_idx);
> DEFINE_PER_CPU(int, sd_share_id);
> DEFINE_PER_CPU(struct sched_domain_shared __rcu *, sd_llc_shared);
> DEFINE_PER_CPU(struct sched_domain __rcu *, sd_numa);
> @@ -666,6 +667,25 @@ DEFINE_PER_CPU(struct sched_domain __rcu *, sd_asym_cpucapacity);
> DEFINE_STATIC_KEY_FALSE(sched_asym_cpucapacity);
> DEFINE_STATIC_KEY_FALSE(sched_cluster_active);
>
> +int max_llcs = -1;
> +
> +static void update_llc_idx(int cpu)
> +{
> +#ifdef CONFIG_SCHED_CACHE
> + int idx = -1, llc_id = -1;
> +
> + llc_id = per_cpu(sd_llc_id, cpu);
> + idx = per_cpu(sd_llc_idx, llc_id);
> +
> + if (idx < 0) {
> + idx = max_llcs++;
> + BUG_ON(idx > MAX_LLC);
[ 2.793048] ------------[ cut here ]------------
[ 2.797737] kernel BUG at kernel/sched/topology.c:682!
[ 2.802957] Internal error: Oops - BUG: 00000000f2000800 [#1] SMP
[snip]
[ 2.917450] update_top_cache_domain+0x248/0x260 (P)
[ 2.922493] cpu_attach_domain+0x170/0x250
[ 2.926653] build_sched_domains+0x640/0x7b0
[ 2.930989] sched_init_domains+0xbc/0xd8
[ 2.935062] sched_init_smp+0x48/0xd0
[ 2.938778] kernel_init_freeable+0xf4/0x148
[ 2.943115] kernel_init+0x2c/0x150
[ 2.946658] ret_from_fork+0x10/0x20
[ 2.950288] Code: d2800001 17ffffc7 d2800000 17ffffdf (d4210000)
[ 2.956481] ---[ end trace 0000000000000000 ]---
[ 2.961170] Kernel panic - not syncing: Oops - BUG: Fatal exception
[ 2.967540] SMP: stopping secondary CPUs
[ 2.971529] ---[ end Kernel panic - not syncing: Oops - BUG: Fatal exception ]---
> + per_cpu(sd_llc_idx, llc_id) = idx;
> + }
> + per_cpu(sd_llc_idx, cpu) = idx;
> +#endif
> +}
> +
> static void update_top_cache_domain(int cpu)
> {
> struct sched_domain_shared *sds = NULL;
> @@ -684,6 +704,7 @@ static void update_top_cache_domain(int cpu)
> per_cpu(sd_llc_size, cpu) = size;
> per_cpu(sd_llc_id, cpu) = id;
> rcu_assign_pointer(per_cpu(sd_llc_shared, cpu), sds);
> + update_llc_idx(cpu);
>
> sd = lowest_flag_domain(cpu, SD_CLUSTER);
> if (sd)
> @@ -2456,6 +2477,14 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
> bool has_asym = false;
> bool has_cluster = false;
>
> +#ifdef CONFIG_SCHED_CACHE
> + if (max_llcs < 0) {
> + for_each_possible_cpu(i)
> + per_cpu(sd_llc_idx, i) = -1;
> + max_llcs = 0;
> + }
> +#endif
> +
> if (WARN_ON(cpumask_empty(cpu_map)))
> goto error;
>
A draft patch like bellow can fix the kernel BUG:
1) Do not call update_llc_idx() if domain has no SD_SHARE_LLC
2) Disable CAS if domain has no SD_SHARE_LLC
diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 8483c02b4d28..cde9b6cdb1de 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -704,7 +704,8 @@ static void update_top_cache_domain(int cpu)
per_cpu(sd_llc_size, cpu) = size;
per_cpu(sd_llc_id, cpu) = id;
rcu_assign_pointer(per_cpu(sd_llc_shared, cpu), sds);
- update_llc_idx(cpu);
+ if (sd)
+ update_llc_idx(cpu);
sd = lowest_flag_domain(cpu, SD_CLUSTER);
if (sd)
@@ -2476,6 +2477,7 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
int i, ret = -ENOMEM;
bool has_asym = false;
bool has_cluster = false;
+ bool has_llc = false;
bool llc_has_parent_sd = false;
unsigned int multi_llcs_node = 1;
@@ -2621,6 +2623,9 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
if (lowest_flag_domain(i, SD_CLUSTER))
has_cluster = true;
+
+ if (highest_flag_domain(i, SD_SHARE_LLC))
+ has_llc = true;
}
rcu_read_unlock();
@@ -2631,7 +2636,8 @@ build_sched_domains(const struct cpumask *cpu_map, struct sched_domain_attr *att
static_branch_inc_cpuslocked(&sched_cluster_active);
#ifdef CONFIG_SCHED_CACHE
- if (llc_has_parent_sd && multi_llcs_node && !sched_asym_cpucap_active())
+ if (has_llc && llc_has_parent_sd && multi_llcs_node &&
+ !sched_asym_cpucap_active())
static_branch_inc_cpuslocked(&sched_cache_present);
#endif
Thanks,
-adam
next prev parent reply other threads:[~2025-09-26 6:14 UTC|newest]
Thread overview: 52+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-08-09 4:57 [RFC PATCH v4 00/28] Cache aware load-balancing Chen Yu
2025-08-09 5:00 ` [RFC PATCH v4 01/28] sched: " Chen Yu
2025-08-12 1:30 ` kernel test robot
2025-08-12 3:26 ` Chen, Yu C
2025-08-09 5:01 ` [RFC PATCH v4 02/28] sched: Several fixes for cache aware scheduling Chen Yu
2025-08-09 5:01 ` [RFC PATCH v4 03/28] sched: Avoid task migration within its preferred LLC Chen Yu
2025-08-09 5:02 ` [RFC PATCH v4 04/28] sched: Avoid calculating the cpumask if the system is overloaded Chen Yu
2025-08-09 5:02 ` [RFC PATCH v4 05/28] sched: Add hysteresis to switch a task's preferred LLC Chen Yu
2025-09-29 13:44 ` Peter Zijlstra
2025-09-30 4:31 ` Chen, Yu C
2025-08-09 5:02 ` [RFC PATCH v4 06/28] sched: Save the per LLC utilization for better cache aware scheduling Chen Yu
2025-09-29 14:09 ` Peter Zijlstra
2025-09-30 4:34 ` Chen, Yu C
2025-08-09 5:03 ` [RFC PATCH v4 07/28] sched: Add helper function to decide whether to allow " Chen Yu
2025-10-01 13:17 ` Peter Zijlstra
2025-10-02 11:31 ` Chen, Yu C
2025-10-02 11:50 ` Peter Zijlstra
2025-10-02 12:51 ` Chen, Yu C
2025-10-02 17:46 ` Tim Chen
2025-08-09 5:03 ` [RFC PATCH v4 08/28] sched: Set up LLC indexing Chen Yu
2025-09-26 6:14 ` Adam Li [this message]
2025-09-26 13:51 ` Chen, Yu C
2025-09-29 10:43 ` Adam Li
2025-09-30 2:54 ` Chen, Yu C
2025-08-09 5:03 ` [RFC PATCH v4 09/28] sched: Introduce task preferred LLC field Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 10/28] sched: Calculate the number of tasks that have LLC preference on a runqueue Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 11/28] sched: Introduce per runqueue task LLC preference counter Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 12/28] sched: Calculate the total number of preferred LLC tasks during load balance Chen Yu
2025-08-09 5:05 ` [RFC PATCH v4 13/28] sched: Tag the sched group as llc_balance if it has tasks prefer other LLC Chen Yu
2025-08-09 5:05 ` [RFC PATCH v4 14/28] sched: Introduce update_llc_busiest() to deal with groups having preferred LLC tasks Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 15/28] sched: Introduce a new migration_type to track the preferred LLC load balance Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 16/28] sched: Consider LLC locality for active balance Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 17/28] sched: Consider LLC preference when picking tasks from busiest queue Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 18/28] sched: Do not migrate task if it is moving out of its preferred LLC Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 19/28] sched: Introduce SCHED_CACHE_LB to control cache aware load balance Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 20/28] sched: Introduce SCHED_CACHE_WAKE to control LLC aggregation on wake up Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 21/28] sched: Introduce a static key to enable cache aware only for multi LLCs Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 22/28] sched: Turn EPOCH_PERIOD and EPOCH_OLD into tunnable debugfs Chen Yu
2025-08-09 5:08 ` [RFC PATCH v4 23/28] sched: Scan a task's preferred node for preferred LLC Chen Yu
2025-08-12 1:59 ` kernel test robot
2025-08-12 3:36 ` Chen, Yu C
2025-08-09 5:08 ` [RFC PATCH v4 24/28] sched: Record average number of runninhg tasks per process Chen Yu
2025-08-09 5:08 ` [RFC PATCH v4 25/28] sched: Skip cache aware scheduling if the process has many active threads Chen Yu
2025-09-02 3:52 ` Tingyin Duan
2025-09-02 5:16 ` Tingyin Duan
2025-09-02 6:14 ` Chen, Yu C
2025-09-02 7:56 ` Duan Tingyin
2025-08-09 5:08 ` [RFC PATCH v4 26/28] sched: Do not enable cache aware scheduling for process with large RSS Chen Yu
2025-09-26 8:48 ` Adam Li
2025-09-26 14:30 ` Chen, Yu C
2025-08-09 5:09 ` [RFC PATCH v4 27/28] sched: Allow the user space to tune the scale factor for RSS comparison Chen Yu
2025-08-09 5:09 ` [RFC PATCH v4 28/28] sched: Add ftrace to track cache aware load balance and hottest CPU changes Chen Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=baa45b9a-d8ef-4652-a3c2-83596216fcb6@os.amperecomputing.com \
--to=adamli@os.amperecomputing.com \
--cc=aubrey.li@intel.com \
--cc=bsegall@google.com \
--cc=cyy@cyyself.name \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=hdanton@sina.com \
--cc=jianyong.wu@outlook.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=len.brown@intel.com \
--cc=libo.chen@oracle.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@linux.intel.com \
--cc=tingyin.duan@gmail.com \
--cc=vernhao@tencent.com \
--cc=vincent.guittot@linaro.org \
--cc=vineethr@linux.ibm.com \
--cc=vschneid@redhat.com \
--cc=yu.c.chen@intel.com \
--cc=yu.chen.surf@gmail.com \
--cc=zhao1.liu@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®