From: Chen Yu <yu.c.chen@intel.com>
To: Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>
Cc: Vincent Guittot <vincent.guittot@linaro.org>,
Juri Lelli <juri.lelli@redhat.com>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Libo Chen <libo.chen@oracle.com>,
Madadi Vineeth Reddy <vineethr@linux.ibm.com>,
Hillf Danton <hdanton@sina.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Jianyong Wu <jianyong.wu@outlook.com>,
Yangyu Chen <cyy@cyyself.name>,
Tingyin Duan <tingyin.duan@gmail.com>,
Vern Hao <vernhao@tencent.com>, Len Brown <len.brown@intel.com>,
Tim Chen <tim.c.chen@linux.intel.com>,
Aubrey Li <aubrey.li@intel.com>, Zhao Liu <zhao1.liu@intel.com>,
Chen Yu <yu.chen.surf@gmail.com>, Chen Yu <yu.c.chen@intel.com>,
linux-kernel@vger.kernel.org
Subject: [RFC PATCH v4 27/28] sched: Allow the user space to tune the scale factor for RSS comparison
Date: Sat, 9 Aug 2025 13:09:02 +0800 [thread overview]
Message-ID: <81c197882b7c9f4325a5cb32f8a9d1e1fc900297.1754712565.git.tim.c.chen@linux.intel.com> (raw)
In-Reply-To: <cover.1754712565.git.tim.c.chen@linux.intel.com>
sched_cache compares the process's resident pages with
the size of the LLC to determine whether task aggregation
on the preferred LLC might cause cache contention. If the
former is larger than the latter, skip cache-aware task
aggregation. However, some workloads with large resident
pages have a small memory footprint; such workloads could
benefit from cache-aware scheduling. The kernel lacks a
efficient mechanism to track the task's memory footprint
(yes, we have resctrl, but it is for user-space query,
and not process scope), so it is up to userspace to pass
this hint to the kernel.
Introduce /sys/kernel/debug/sched/sched_cache_ignore_rss
to control the extent to which users ignore the RSS
restriction. This value ranges from 0 to 100. A value of
0 means that the user disables the cache aware scheduling.
1 means if a process's RSS is larger than the LLC size,
cache-aware scheduling will be skipped. 100 means cache
aware scheduling is alwasy enabled regardless of RSS size.
N (between 1 and 100) means turn off cache aware scheduling
when RSS is greater than (N-1) * 256 * LLC size
For example, suppose the L3 size is 32MB. If the
sysctl_sched_cache_ignore_rss is 1: When the RSS is larger
than 32MB, the process is regarded as exceeding the LLC capacity.
If the sysctl_sched_cache_ignore_rss is 99: When the RSS is
larger than 784GB, the process is regarded as exceeding the
LLC capacity(please refer to the code):
784GB = (1 + (99 - 1) * 256) * 32MB
Additionally, the number of SMTs is also considered for
sysctl_sched_cache_aggr_cap; if there are many SMTs in the core,
sysctl_llc_aggr_cap will be reduced. This inhibits task aggregation
from cache-aware scheduling on systems with a high number of SMTs,
like Power 10 and Power 11.
Reported-by: K Prateek Nayak <kprateek.nayak@amd.com>
Reported-by: Madadi Vineeth Reddy <vineethr@linux.ibm.com>
Reported-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Co-developed-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
---
kernel/sched/debug.c | 82 +++++++++++++++++++++++++++++++++++++++++++-
kernel/sched/fair.c | 10 ++++--
kernel/sched/sched.h | 3 +-
3 files changed, 90 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index 7a9ec03704b9..6676fc2a8c08 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -166,6 +166,83 @@ static const struct file_operations sched_feat_fops = {
.release = single_release,
};
+#ifdef CONFIG_SCHED_CACHE
+#define SCHED_CACHE_CREATE_CONTROL(name, val) \
+static int sysctl_sched_cache_##name = val; \
+static ssize_t sched_cache_write_##name(struct file *filp, \
+ const char __user *ubuf, \
+ size_t cnt, loff_t *ppos) \
+{ \
+ char buf[16]; \
+ unsigned int percent; \
+ if (cnt > 15) \
+ cnt = 15; \
+ if (copy_from_user(&buf, ubuf, cnt)) \
+ return -EFAULT; \
+ buf[cnt] = '\0'; \
+ if (kstrtouint(buf, 10, &percent)) \
+ return -EINVAL; \
+ if (percent > 100) \
+ return -EINVAL; \
+ sysctl_sched_cache_##name = percent; \
+ *ppos += cnt; \
+ return cnt; \
+} \
+static int sched_cache_show_##name(struct seq_file *m, void *v) \
+{ \
+ seq_printf(m, "%d\n", sysctl_sched_cache_##name); \
+ return 0; \
+} \
+static int sched_cache_open_##name(struct inode *inode, \
+ struct file *filp) \
+{ \
+ return single_open(filp, sched_cache_show_##name, NULL); \
+} \
+static const struct file_operations sched_cache_fops_##name = { \
+ .open = sched_cache_open_##name, \
+ .write = sched_cache_write_##name, \
+ .read = seq_read, \
+ .llseek = seq_lseek, \
+ .release = single_release, \
+}
+
+SCHED_CACHE_CREATE_CONTROL(ignore_rss, 1);
+int get_sched_cache_rss_scale(void)
+{
+ if (!sysctl_sched_cache_ignore_rss)
+ return 0;
+
+ if (sysctl_sched_cache_ignore_rss >= 100)
+ return INT_MAX;
+ /*
+ * Suppose the L3 size is 32MB. If the
+ * sysctl_sched_cache_ignore_rss is 1:
+ * When the RSS is larger than 32MB,
+ * the process is regarded as exceeding
+ * the LLC capacity. If the
+ * sysctl_sched_cache_ignore_rss is 99:
+ * When the RSS is larger than 784GB,
+ * the process is regarded as exceeding
+ * the LLC capacity:
+ * 784GB = (1 + (99 - 1) * 256) * 32MB
+ */
+ return (1 + (sysctl_sched_cache_ignore_rss - 1) * 256);
+}
+
+SCHED_CACHE_CREATE_CONTROL(aggr_cap, 50);
+int get_sched_cache_cap_scale(void)
+{
+ int smt_nr = 1;
+
+#ifdef CONFIG_SCHED_SMT
+ if (sched_smt_active())
+ smt_nr =
+ cpumask_weight(cpu_smt_mask(raw_smp_processor_id()));
+#endif
+ return (sysctl_sched_cache_aggr_cap / smt_nr);
+}
+#endif /* SCHED_CACHE */
+
#ifdef CONFIG_SMP
static ssize_t sched_scaling_write(struct file *filp, const char __user *ubuf,
@@ -533,10 +610,13 @@ static __init int sched_init_debug(void)
#endif
#ifdef CONFIG_SCHED_CACHE
- debugfs_create_u32("llc_aggr_cap", 0644, debugfs_sched, &sysctl_llc_aggr_cap);
debugfs_create_u32("llc_aggr_imb", 0644, debugfs_sched, &sysctl_llc_aggr_imb);
debugfs_create_u32("llc_period", 0644, debugfs_sched, &sysctl_llc_period);
debugfs_create_u32("llc_old", 0644, debugfs_sched, &sysctl_llc_old);
+ debugfs_create_file("llc_aggr_cap", 0644, debugfs_sched, NULL,
+ &sched_cache_fops_aggr_cap);
+ debugfs_create_file("llc_ignore_rss", 0644, debugfs_sched, NULL,
+ &sched_cache_fops_ignore_rss);
#endif
debugfs_create_file("debug", 0444, debugfs_sched, NULL, &sched_debug_fops);
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index cbda7dad1305..018825f04063 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1211,6 +1211,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
struct cacheinfo *l3_leaf;
unsigned long rss;
unsigned int llc;
+ int scale;
/*
* get_cpu_cacheinfo_level() can not be used
@@ -1230,7 +1231,11 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu)
rss = get_mm_counter(mm, MM_ANONPAGES) +
get_mm_counter(mm, MM_SHMEMPAGES);
- return (llc <= (rss * PAGE_SIZE));
+ scale = get_sched_cache_rss_scale();
+ if (scale == INT_MAX)
+ return false;
+
+ return ((llc * scale) <= (rss * PAGE_SIZE));
}
static bool exceed_llc_nr(struct mm_struct *mm, int cpu)
@@ -9037,7 +9042,6 @@ static int find_energy_efficient_cpu(struct task_struct *p, int prev_cpu)
static long __migrate_degrades_locality(struct task_struct *p,
int src_cpu, int dst_cpu,
bool idle);
-__read_mostly unsigned int sysctl_llc_aggr_cap = 50;
__read_mostly unsigned int sysctl_llc_aggr_imb = 20;
/*
@@ -9049,7 +9053,7 @@ __read_mostly unsigned int sysctl_llc_aggr_imb = 20;
* (default: ~50%)
*/
#define fits_llc_capacity(util, max) \
- ((util) * 100 < (max) * sysctl_llc_aggr_cap)
+ ((util) * 100 < (max) * get_sched_cache_cap_scale())
/*
* The margin used when comparing utilization.
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index d752d64d4acd..eaeca4e77ead 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2855,11 +2855,12 @@ extern unsigned int sysctl_numa_balancing_scan_size;
extern unsigned int sysctl_numa_balancing_hot_threshold;
#ifdef CONFIG_SCHED_CACHE
-extern unsigned int sysctl_llc_aggr_cap;
extern unsigned int sysctl_llc_aggr_imb;
extern struct static_key_false sched_cache_present;
extern unsigned int sysctl_llc_period;
extern unsigned int sysctl_llc_old;
+int get_sched_cache_rss_scale(void);
+int get_sched_cache_cap_scale(void);
#endif
#ifdef CONFIG_SCHED_HRTICK
--
2.25.1
next prev parent reply other threads:[~2025-08-09 5:15 UTC|newest]
Thread overview: 52+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-08-09 4:57 [RFC PATCH v4 00/28] Cache aware load-balancing Chen Yu
2025-08-09 5:00 ` [RFC PATCH v4 01/28] sched: " Chen Yu
2025-08-12 1:30 ` kernel test robot
2025-08-12 3:26 ` Chen, Yu C
2025-08-09 5:01 ` [RFC PATCH v4 02/28] sched: Several fixes for cache aware scheduling Chen Yu
2025-08-09 5:01 ` [RFC PATCH v4 03/28] sched: Avoid task migration within its preferred LLC Chen Yu
2025-08-09 5:02 ` [RFC PATCH v4 04/28] sched: Avoid calculating the cpumask if the system is overloaded Chen Yu
2025-08-09 5:02 ` [RFC PATCH v4 05/28] sched: Add hysteresis to switch a task's preferred LLC Chen Yu
2025-09-29 13:44 ` Peter Zijlstra
2025-09-30 4:31 ` Chen, Yu C
2025-08-09 5:02 ` [RFC PATCH v4 06/28] sched: Save the per LLC utilization for better cache aware scheduling Chen Yu
2025-09-29 14:09 ` Peter Zijlstra
2025-09-30 4:34 ` Chen, Yu C
2025-08-09 5:03 ` [RFC PATCH v4 07/28] sched: Add helper function to decide whether to allow " Chen Yu
2025-10-01 13:17 ` Peter Zijlstra
2025-10-02 11:31 ` Chen, Yu C
2025-10-02 11:50 ` Peter Zijlstra
2025-10-02 12:51 ` Chen, Yu C
2025-10-02 17:46 ` Tim Chen
2025-08-09 5:03 ` [RFC PATCH v4 08/28] sched: Set up LLC indexing Chen Yu
2025-09-26 6:14 ` Adam Li
2025-09-26 13:51 ` Chen, Yu C
2025-09-29 10:43 ` Adam Li
2025-09-30 2:54 ` Chen, Yu C
2025-08-09 5:03 ` [RFC PATCH v4 09/28] sched: Introduce task preferred LLC field Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 10/28] sched: Calculate the number of tasks that have LLC preference on a runqueue Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 11/28] sched: Introduce per runqueue task LLC preference counter Chen Yu
2025-08-09 5:04 ` [RFC PATCH v4 12/28] sched: Calculate the total number of preferred LLC tasks during load balance Chen Yu
2025-08-09 5:05 ` [RFC PATCH v4 13/28] sched: Tag the sched group as llc_balance if it has tasks prefer other LLC Chen Yu
2025-08-09 5:05 ` [RFC PATCH v4 14/28] sched: Introduce update_llc_busiest() to deal with groups having preferred LLC tasks Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 15/28] sched: Introduce a new migration_type to track the preferred LLC load balance Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 16/28] sched: Consider LLC locality for active balance Chen Yu
2025-08-09 5:06 ` [RFC PATCH v4 17/28] sched: Consider LLC preference when picking tasks from busiest queue Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 18/28] sched: Do not migrate task if it is moving out of its preferred LLC Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 19/28] sched: Introduce SCHED_CACHE_LB to control cache aware load balance Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 20/28] sched: Introduce SCHED_CACHE_WAKE to control LLC aggregation on wake up Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 21/28] sched: Introduce a static key to enable cache aware only for multi LLCs Chen Yu
2025-08-09 5:07 ` [RFC PATCH v4 22/28] sched: Turn EPOCH_PERIOD and EPOCH_OLD into tunnable debugfs Chen Yu
2025-08-09 5:08 ` [RFC PATCH v4 23/28] sched: Scan a task's preferred node for preferred LLC Chen Yu
2025-08-12 1:59 ` kernel test robot
2025-08-12 3:36 ` Chen, Yu C
2025-08-09 5:08 ` [RFC PATCH v4 24/28] sched: Record average number of runninhg tasks per process Chen Yu
2025-08-09 5:08 ` [RFC PATCH v4 25/28] sched: Skip cache aware scheduling if the process has many active threads Chen Yu
2025-09-02 3:52 ` Tingyin Duan
2025-09-02 5:16 ` Tingyin Duan
2025-09-02 6:14 ` Chen, Yu C
2025-09-02 7:56 ` Duan Tingyin
2025-08-09 5:08 ` [RFC PATCH v4 26/28] sched: Do not enable cache aware scheduling for process with large RSS Chen Yu
2025-09-26 8:48 ` Adam Li
2025-09-26 14:30 ` Chen, Yu C
2025-08-09 5:09 ` Chen Yu [this message]
2025-08-09 5:09 ` [RFC PATCH v4 28/28] sched: Add ftrace to track cache aware load balance and hottest CPU changes Chen Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=81c197882b7c9f4325a5cb32f8a9d1e1fc900297.1754712565.git.tim.c.chen@linux.intel.com \
--to=yu.c.chen@intel.com \
--cc=aubrey.li@intel.com \
--cc=bsegall@google.com \
--cc=cyy@cyyself.name \
--cc=dietmar.eggemann@arm.com \
--cc=gautham.shenoy@amd.com \
--cc=hdanton@sina.com \
--cc=jianyong.wu@outlook.com \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=len.brown@intel.com \
--cc=libo.chen@oracle.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@linux.intel.com \
--cc=tingyin.duan@gmail.com \
--cc=vernhao@tencent.com \
--cc=vincent.guittot@linaro.org \
--cc=vineethr@linux.ibm.com \
--cc=vschneid@redhat.com \
--cc=yu.chen.surf@gmail.com \
--cc=zhao1.liu@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®