From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from sg-1-101.ptr.blmpb.com (sg-1-101.ptr.blmpb.com [118.26.132.101]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EB6F0345CB9 for ; Wed, 17 Dec 2025 09:59:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=118.26.132.101 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765965599; cv=none; b=h+kpxflKo6+UGDy021g8ietO7bKz/ZrydB8DdixKYP0SeKp1rw3Ci7eSHBnUG0S5Sg2ZS51bpAe9K3dtPbfUVSB/BL9eNpoQGrjy6Wp2J1CE+QXuJvfUMgDLp8w80J2A542nRtJaJq8l1udu9xSrwvJwjrxxaqi7GJgrXSHbz3M= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765965599; c=relaxed/simple; bh=vQIOBEvmCxSdqnHgcOq4dTnJT39howNk7LxiUFYlP7I=; h=To:Cc:Date:Message-Id:Content-Disposition:Content-Type:From: Subject:Mime-Version:References:In-Reply-To; b=JQKGPAaPLCk7yTw3Guo/RQnYA91Y/qyChlqQu5pYC7MfLIVoj4ZAIdQaxGuzLrM3R8reAPnMzq5CYJ3p74A6tx94yC/9Gt6X7tyBSpCnEiXbK6FVXq242UpjNVVAiMyad2PgE+DD8uWp3x/Wwchg0UcWxlD2hBZWo/ASME5s/2c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=IHIMU260; arc=none smtp.client-ip=118.26.132.101 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="IHIMU260" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=2212171451; d=bytedance.com; t=1765965578; h=from:subject: mime-version:from:date:message-id:subject:to:cc:reply-to:content-type: mime-version:in-reply-to:message-id; bh=ExDaJCs8R5WuVD++6ippYwYbF5ZoDiAb/miqIwv3AbA=; b=IHIMU260X3RI7whMCJ1TNxcsi+0MDJ1P28pt1YpkTmWco8ABKJzt0QNFfM/nZ6q4ZMEwMg oi7RPMjH0Yu6zzt9I+X1Oxom0Ne1E7DMW7lUQhP5anVKuuViROMDFnC8w9W3WhFZYffnvU ZzgQ7okueJDapgx/SN26MUSTuByQVgGcABDdqftKjWdRuHLOBXwS8NQWVNWmdVU2Pg2pDY KKGgYuJ4IBT7AtqRsH24TALpCj0hgouDTdRokFM+b9ICvc8XP99YOOObOIlDIQ8McKyMEA QewDrYsBS9TGNUqDxi6YuAFQlUoqw+4vWiS1repFtxEg/PWRRoGYn9rhWTtjoA== X-Original-From: Aaron Lu X-Lms-Return-Path: To: "Chen Yu" , "Tim Chen" Cc: "Peter Zijlstra" , "Ingo Molnar" , "K Prateek Nayak" , "Gautham R . Shenoy" , "Vincent Guittot" , "Juri Lelli" , "Dietmar Eggemann" , "Steven Rostedt" , "Ben Segall" , "Mel Gorman" , "Valentin Schneider" , "Madadi Vineeth Reddy" , "Hillf Danton" , "Shrikanth Hegde" , "Jianyong Wu" , "Yangyu Chen" , "Tingyin Duan" , "Vern Hao" , "Vern Hao" , "Len Brown" , "Aubrey Li" , "Zhao Liu" , "Chen Yu" , "Adam Li" , "Tim Chen" , Date: Wed, 17 Dec 2025 17:59:12 +0800 Message-Id: <20251217095912.GB2073381@bytedance.com> Content-Disposition: inline Content-Type: text/plain; charset=UTF-8 From: "Aaron Lu" Subject: Re: [PATCH v2 23/23] -- DO NOT APPLY!!! -- sched/cache/debug: Display the per LLC occupancy for each process via proc fs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <0eaf9b9f89f0d97dbf46b760421f65aee3ffe063.1764801860.git.tim.c.chen@linux.intel.com> In-Reply-To: <0eaf9b9f89f0d97dbf46b760421f65aee3ffe063.1764801860.git.tim.c.chen@linux.intel.com> Content-Transfer-Encoding: 7bit On Wed, Dec 03, 2025 at 03:07:42PM -0800, Tim Chen wrote: > From: Chen Yu > > Debug patch only. > > Show the per-LLC occupancy in /proc/{PID}/schedstat, with each column > corresponding to one LLC. This can be used to verify if the cache-aware > load balancer works as expected by aggregating threads onto dedicated LLCs. > > Suppose there are 2 LLCs and the sampling duration is 10 seconds: > > Enable the cache aware load balance: > 0 12281 <--- LLC0 residency delta is 0, LLC1 is 12 seconds > 0 18881 > 0 16217 > > disable the cache aware load balance: > 6497 15802 > 9299 5435 > 17811 8278 > > Signed-off-by: Chen Yu > Signed-off-by: Tim Chen > --- > fs/proc/base.c | 22 ++++++++++++++++++++++ > include/linux/mm_types.h | 19 +++++++++++++++++-- > include/linux/sched.h | 3 +++ > kernel/sched/fair.c | 40 ++++++++++++++++++++++++++++++++++++++-- > 4 files changed, 80 insertions(+), 4 deletions(-) > > diff --git a/fs/proc/base.c b/fs/proc/base.c > index 6299878e3d97..f4be96f4bd01 100644 > --- a/fs/proc/base.c > +++ b/fs/proc/base.c > @@ -518,6 +518,28 @@ static int proc_pid_schedstat(struct seq_file *m, struct pid_namespace *ns, > (unsigned long long)task->se.sum_exec_runtime, > (unsigned long long)task->sched_info.run_delay, > task->sched_info.pcount); > +#ifdef CONFIG_SCHED_CACHE > + if (sched_cache_enabled()) { > + struct mm_struct *mm = task->mm; > + u64 *llc_runtime; > + > + if (!mm) > + return 0; > + > + llc_runtime = kcalloc(max_llcs, sizeof(u64), GFP_KERNEL); > + if (!llc_runtime) > + return 0; > + > + if (get_mm_per_llc_runtime(task, llc_runtime)) > + goto out; > + > + for (int i = 0; i < max_llcs; i++) > + seq_printf(m, "%llu ", llc_runtime[i]); I feel it is better to also mark the current preferred LLC of this process so that I can know how well it works. > + seq_puts(m, "\n"); > +out: > + kfree(llc_runtime); > + } > +#endif > > return 0; > } BTW, is there a way to tell if a process is being taken care of by 'cache aware scheduling' or it's blocked due to its huge rss or having too many threads? I used below debug code to get these info through schedstat, but maybe I missed something and there is a simpler method? diff --git a/fs/proc/base.c b/fs/proc/base.c index f4be96f4bd015..c709a1a1bd867 100644 --- a/fs/proc/base.c +++ b/fs/proc/base.c @@ -505,6 +505,7 @@ static int proc_pid_stack(struct seq_file *m, struct pid_namespace *ns, #endif #ifdef CONFIG_SCHED_INFO +DECLARE_PER_CPU(int, sd_llc_id); /* * Provides /proc/PID/schedstat */ @@ -522,6 +523,7 @@ static int proc_pid_schedstat(struct seq_file *m, struct pid_namespace *ns, if (sched_cache_enabled()) { struct mm_struct *mm = task->mm; u64 *llc_runtime; + int mm_sched_llc; if (!mm) return 0; @@ -533,8 +535,17 @@ static int proc_pid_schedstat(struct seq_file *m, struct pid_namespace *ns, if (get_mm_per_llc_runtime(task, llc_runtime)) goto out; + if (mm->mm_sched_cpu == -1) + mm_sched_llc = -1; + else + mm_sched_llc = per_cpu(sd_llc_id, mm->mm_sched_cpu); + + seq_printf(m, "%llu 0x%x\n", mm->nr_running_avg, mm->mm_sched_flags); for (int i = 0; i < max_llcs; i++) - seq_printf(m, "%llu ", llc_runtime[i]); + seq_printf(m, "%s%s%llu ", + i == task->preferred_llc ? "*" : "", + i == mm_sched_llc ? "?" : "", + llc_runtime[i]); seq_puts(m, "\n"); out: kfree(llc_runtime); diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 255c22be7312f..06bb106d1b724 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1048,6 +1048,7 @@ struct mm_struct { raw_spinlock_t mm_sched_lock; unsigned long mm_sched_epoch; int mm_sched_cpu; + int mm_sched_flags; u64 nr_running_avg ____cacheline_aligned_in_smp; #endif diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 205208f061bb3..ab1cdba65d389 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1237,12 +1237,20 @@ static inline int get_sched_cache_scale(int mul) return (1 + (llc_aggr_tolerance - 1) * mul); } +#define MM_SCHED_EXCEED_LLC_CAPACITY 1 +#define MM_SCHED_NO_CACHE_INFO 2 +#define MM_SCHED_EXCEED_LLC_NR 4 +#define MM_SCHED_NR_THREADS 8 + static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) { unsigned int llc, scale; struct cacheinfo *ci; unsigned long rss; + mm->mm_sched_flags &= ~MM_SCHED_NO_CACHE_INFO; + mm->mm_sched_flags &= ~MM_SCHED_EXCEED_LLC_CAPACITY; + /* * get_cpu_cacheinfo_level() can not be used * because it requires the cpu_hotplug_lock @@ -1257,8 +1265,10 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) * L2 becomes the LLC. */ ci = _get_cpu_cacheinfo_level(cpu, 2); - if (!ci) + if (!ci) { + mm->mm_sched_flags |= MM_SCHED_NO_CACHE_INFO; return true; + } } llc = ci->size; @@ -1283,13 +1293,20 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) if (scale == INT_MAX) return false; - return ((llc * scale) <= (rss * PAGE_SIZE)); + if ((llc * scale) <= (rss * PAGE_SIZE)) { + mm->mm_sched_flags |= MM_SCHED_EXCEED_LLC_CAPACITY; + return true; + } + + return false; } static bool exceed_llc_nr(struct mm_struct *mm, int cpu) { int smt_nr = 1, scale; + mm->mm_sched_flags &= ~MM_SCHED_EXCEED_LLC_NR; + #ifdef CONFIG_SCHED_SMT if (sched_smt_active()) smt_nr = cpumask_weight(cpu_smt_mask(cpu)); @@ -1313,7 +1330,12 @@ static bool exceed_llc_nr(struct mm_struct *mm, int cpu) if (scale == INT_MAX) return false; - return ((mm->nr_running_avg * smt_nr) > (scale * per_cpu(sd_llc_size, cpu))); + if ((mm->nr_running_avg * smt_nr) > (scale * per_cpu(sd_llc_size, cpu))) { + mm->mm_sched_flags |= MM_SCHED_EXCEED_LLC_NR; + return true; + } + + return false; } static void account_llc_enqueue(struct rq *rq, struct task_struct *p)