From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f171.google.com (mail-pl1-f171.google.com [209.85.214.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A31063112D0 for ; Thu, 18 Dec 2025 05:55:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.171 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766037360; cv=none; b=ULTWbnsUA1n5X/PykITyJH9Hadt1eb6osN6vt0zT5V0d9oxDVxRMacLtkfNKR+1r0yRDOlfqW6bsufHmc34ZXwf5L2UmFRYfe45Cj3eX3ou4u83aRYcxe8D/nvvMFlS3ENC+KiYrWFrqP8/BsHua7htpKM9C1JD2H6R5c3h7cBo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766037360; c=relaxed/simple; bh=YkutC7etZtjmEOxrI5zeneiXx7T2+ifwgkUmJ+G6Stc=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=K9aVNPfmHj9kZAoUDxygpmCuD2KTCDxct9xVsyH2x0fkRyxyylD/DZZG8YD6cfdyBzqieeTUgJqnvNqlUcsMh0XWJ+hY9DuTaD3RYX+4cl7AqfKJGHcZktli8SG5+Tcz48YoXmqBLLp0tIQjADQiy+vpHzqk2jE9bMHRPsBBTUQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Zz03RreF; arc=none smtp.client-ip=209.85.214.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Zz03RreF" Received: by mail-pl1-f171.google.com with SMTP id d9443c01a7336-29f0f875bc5so3652315ad.3 for ; Wed, 17 Dec 2025 21:55:58 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1766037358; x=1766642158; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=7Po1P1QnKZUtzbh+XGAJJZjx8YM9JBOzpdkGqlUoKGk=; b=Zz03RreFKBWFSaASiST8vEHSEbHQ/qLKkUvr4Yq6oeZC7YB7Z4g1atKrG2TLhTLUc0 Gshxb6WXKaIX/uBzcRV48YzIz+xgJm74qGtHoGuNgUT1sXdTdKIVs+6QHhb5HfExXlD5 kFsbPzBoraJ+A+cf2oZXDgv2lct7O7h7xgoiSL1R12720LzBq3ZVR/nF2t7rLWAilGrA gavTGc5cfiuZvAQ0lifJidGAjE5oeqJ/jtUJf9Rhu6BDrofx0heuy7240VwOVE+o3A7t FHqPabdsFn3WrDx1Fu6bxItDTWE54kQh5YgE+CvF7C3kALlmSCJCX5XkMoWxedRjrrN6 8d+Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766037358; x=1766642158; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=7Po1P1QnKZUtzbh+XGAJJZjx8YM9JBOzpdkGqlUoKGk=; b=mkXfVHTRNviek7ps4FvKfNje+oGibnM7E61L45yJx/IE9XfwcUPEFgPZeG2gPNhjIY CePDEgZwGEEhNEdd/li8yxXC1rtb3wLYkpaSymUhH6e0QSIr1IkEInGrJGhvtyTHD95o ycEkNegO7wcfAwxW/ieVyeYnVfwfx4Du60Jd6Ts+eXqTRLHMt7Dz9Gg5AimtD94M9KZH qS+5Pck+h45p2HyHQe1K6FG8h/TZQ/UbrY8Y54QAaXvDR2on0oF1kngjyrSjZK6ny/HO S0md2ZIwSbdo8vVeAt6H3VXg2VwbaOhVNPmdz/hqzHYlNxgT5Zqr9YhuP4WBZYCCKptk y5Dg== X-Forwarded-Encrypted: i=1; AJvYcCX4LS9CyYUUILU5wGvUyo1Zr/ZoxATzs6G7bfFWopSmyqEjB68ZwPMpuAaJBEDSADsUl12eI6t536MfvHc=@vger.kernel.org X-Gm-Message-State: AOJu0YzHwrWIVeGhToOZoakGHUhwEhvcttKJMyQLoz4dxQD4oEKhOaN8 hT4BqT3Gpndgdg1lz2Xh5QxGrU69Avgi8PugEBZAedTXjvHoyf7I7cl3p1EhZK38 X-Gm-Gg: AY/fxX7RNdeQ8kU1wCgOwGCXjMpf5wmFlCypF7QiW0ggEUn6T23ztJhHMRCZTmd3sUx l0EzeI3hAoh7KNKArIyTb+/6RNm8/jHP2tI2nAMXzITM5SNm+9fptYWlCoFm8LF8XZcDgIxrOsU KM7E3a4RdrrjfKAJEN5B3j/2fEsBcsvfQoUqYpan5e2BM+XaJo4m5O9VNNrYMYbyLKNtNaZa4Yj DdUb29fNiDogT8rOYrK97V7W3PVB8yezASxtljU4mViDeeH2LZmjiQVUTvLvQWdIBAcMsWaTAG5 JxMc4lNGdYrLk7eDqTxSnXaW4c7Es3KDlMf3Op3sBe7x1M9+uUYwnRt/4T1xMcFKRCN0caGCbhx TOk+fnUF8NhNUkyIx2qWOkcEJU8Euh2MDlytZWU32XJXbRxr7qWZ/6IUai+s8HlTKyq08PK2OLJ T2EvJ3B2WI5089D3rhjILzYp5Pf45PeChTs7svktXo9gmKJcKt X-Google-Smtp-Source: AGHT+IGkucIIrmtz4eGlq7TnX9I2YCl0+nFNQv38kfs4eSYaQ9VtCUWXiBaBGwOL29Tvm13zrzR+Rw== X-Received: by 2002:a17:903:2bcb:b0:28e:aacb:e702 with SMTP id d9443c01a7336-29f23dfdda5mr226786555ad.2.1766030406139; Wed, 17 Dec 2025 20:00:06 -0800 (PST) Received: from [192.168.255.10] ([43.132.141.20]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2a2d19280absm8326615ad.88.2025.12.17.20.00.00 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 17 Dec 2025 20:00:05 -0800 (PST) Message-ID: <91a7c325-5093-4417-aa98-34df694b0c39@gmail.com> Date: Thu, 18 Dec 2025 11:59:57 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 19/23] sched/cache: Avoid cache-aware scheduling for memory-heavy processes To: Tim Chen , Peter Zijlstra , Ingo Molnar , K Prateek Nayak , "Gautham R . Shenoy" , Vincent Guittot Cc: Chen Yu , Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org References: From: Vern Hao In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 2025/12/4 07:07, Tim Chen wrote: > From: Chen Yu > > Prateek and Tingyin reported that memory-intensive workloads (such as > stream) can saturate memory bandwidth and caches on the preferred LLC > when sched_cache aggregates too many threads. > > To mitigate this, estimate a process's memory footprint by comparing > its RSS (anonymous and shared pages) to the size of the LLC. If RSS > exceeds the LLC size, skip cache-aware scheduling. Restricting RSS prevents many applications from benefiting from this optimization. I believe this restriction should be lifted. For memory-intensive workloads, the optimization may simply yield no gains, but it certainly shouldn't make performance worse. We need to further refine this logic. > Note that RSS is only an approximation of the memory footprint. > By default, the comparison is strict, but a later patch will allow > users to provide a hint to adjust this threshold. > > According to the test from Adam, some systems do not have shared L3 > but with shared L2 as clusters. In this case, the L2 becomes the LLC[1]. > > Link[1]: https://lore.kernel.org/all/3cb6ebc7-a2fd-42b3-8739-b00e28a09cb6@os.amperecomputing.com/ > > Co-developed-by: Tim Chen > Signed-off-by: Chen Yu > Signed-off-by: Tim Chen > --- > > Notes: > v1->v2: Assigned curr_cpu in task_cache_work() before checking > exceed_llc_capacity(mm, curr_cpu) to avoid out-of-bound > access.(lkp/0day) > > include/linux/cacheinfo.h | 21 ++++++++++------- > kernel/sched/fair.c | 49 +++++++++++++++++++++++++++++++++++---- > 2 files changed, 57 insertions(+), 13 deletions(-) > > diff --git a/include/linux/cacheinfo.h b/include/linux/cacheinfo.h > index c8f4f0a0b874..82d0d59ca0e1 100644 > --- a/include/linux/cacheinfo.h > +++ b/include/linux/cacheinfo.h > @@ -113,18 +113,11 @@ int acpi_get_cache_info(unsigned int cpu, > > const struct attribute_group *cache_get_priv_group(struct cacheinfo *this_leaf); > > -/* > - * Get the cacheinfo structure for the cache associated with @cpu at > - * level @level. > - * cpuhp lock must be held. > - */ > -static inline struct cacheinfo *get_cpu_cacheinfo_level(int cpu, int level) > +static inline struct cacheinfo *_get_cpu_cacheinfo_level(int cpu, int level) > { > struct cpu_cacheinfo *ci = get_cpu_cacheinfo(cpu); > int i; > > - lockdep_assert_cpus_held(); > - > for (i = 0; i < ci->num_leaves; i++) { > if (ci->info_list[i].level == level) { > if (ci->info_list[i].attributes & CACHE_ID) > @@ -136,6 +129,18 @@ static inline struct cacheinfo *get_cpu_cacheinfo_level(int cpu, int level) > return NULL; > } > > +/* > + * Get the cacheinfo structure for the cache associated with @cpu at > + * level @level. > + * cpuhp lock must be held. > + */ > +static inline struct cacheinfo *get_cpu_cacheinfo_level(int cpu, int level) > +{ > + lockdep_assert_cpus_held(); > + > + return _get_cpu_cacheinfo_level(cpu, level); > +} > + > /* > * Get the id of the cache associated with @cpu at level @level. > * cpuhp lock must be held. > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index 6afa3f9a4e9b..424ec601cfdf 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -1223,6 +1223,38 @@ static int llc_id(int cpu) > return llc; > } > > +static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) > +{ > + struct cacheinfo *ci; > + unsigned long rss; > + unsigned int llc; > + > + /* > + * get_cpu_cacheinfo_level() can not be used > + * because it requires the cpu_hotplug_lock > + * to be held. Use _get_cpu_cacheinfo_level() > + * directly because the 'cpu' can not be > + * offlined at the moment. > + */ > + ci = _get_cpu_cacheinfo_level(cpu, 3); > + if (!ci) { > + /* > + * On system without L3 but with shared L2, > + * L2 becomes the LLC. > + */ > + ci = _get_cpu_cacheinfo_level(cpu, 2); > + if (!ci) > + return true; > + } Is there must call it one by one for get llc size? a static variable instead in building sched domain? > + > + llc = ci->size; > + > + rss = get_mm_counter(mm, MM_ANONPAGES) + > + get_mm_counter(mm, MM_SHMEMPAGES); > + > + return (llc <= (rss * PAGE_SIZE)); > +} > + > static bool exceed_llc_nr(struct mm_struct *mm, int cpu) > { > int smt_nr = 1; > @@ -1382,7 +1414,8 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) > */ > if (epoch - READ_ONCE(mm->mm_sched_epoch) > EPOCH_LLC_AFFINITY_TIMEOUT || > get_nr_threads(p) <= 1 || > - exceed_llc_nr(mm, cpu_of(rq))) { > + exceed_llc_nr(mm, cpu_of(rq)) || > + exceed_llc_capacity(mm, cpu_of(rq))) { > if (mm->mm_sched_cpu != -1) > mm->mm_sched_cpu = -1; > } > @@ -1439,7 +1472,7 @@ static void __no_profile task_cache_work(struct callback_head *work) > struct mm_struct *mm = p->mm; > unsigned long m_a_occ = 0; > unsigned long curr_m_a_occ = 0; > - int cpu, m_a_cpu = -1, nr_running = 0; > + int cpu, m_a_cpu = -1, nr_running = 0, curr_cpu; > cpumask_var_t cpus; > > WARN_ON_ONCE(work != &p->cache_work); > @@ -1449,7 +1482,9 @@ static void __no_profile task_cache_work(struct callback_head *work) > if (p->flags & PF_EXITING) > return; > > - if (get_nr_threads(p) <= 1) { > + curr_cpu = task_cpu(p); > + if (get_nr_threads(p) <= 1 || > + exceed_llc_capacity(mm, curr_cpu)) { > if (mm->mm_sched_cpu != -1) > mm->mm_sched_cpu = -1; > > @@ -9895,8 +9930,12 @@ static enum llc_mig can_migrate_llc_task(int src_cpu, int dst_cpu, > if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) > return mig_unrestricted; > > - /* skip cache aware load balance for single/too many threads */ > - if (get_nr_threads(p) <= 1 || exceed_llc_nr(mm, dst_cpu)) > + /* > + * Skip cache aware load balance for single/too many threads > + * or large footprint. > + */ > + if (get_nr_threads(p) <= 1 || exceed_llc_nr(mm, dst_cpu) || > + exceed_llc_capacity(mm, dst_cpu)) > return mig_unrestricted; > > if (cpus_share_cache(dst_cpu, cpu))