From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f173.google.com (mail-pf1-f173.google.com [209.85.210.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F2AE1DA62E for ; Mon, 22 Dec 2025 02:19:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.173 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766369986; cv=none; b=PX0g0gKP5FpMH9W60hlqKuBUF0+9V6LafDqaH3agkdYMFHDxjSZpvifXILQVE8oTu1YeYM+Rt7pesNSVCs5Ubw3x+uj9IgbycU7kvFtLoGgYv9Y0uPsy7TytV5qzNnTSsIpqPbVEdTeBbyJYEP8/j6LKdaASNANLn2EdLiXj1VM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766369986; c=relaxed/simple; bh=6qSk1LxkAAbC2ARnc7a8EDtiGTJ7FPX+YveXX6YDR+o=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=tRVeXL6asR2nnawcc0XC7H2sDAt/LE9VgKvf9/vCqWt/oP5ujFTvT0Wrc7airS9QKtL6rHyBvF72a5seBtdlhVFOGSCidK1KX334c+q+g00EbHZhB9EhQvKewtIcbPnHBIZn1YKbs5YiYLwDGqGfZBbPyTC1uKuo77iwunjq0CI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=GMNK44AP; arc=none smtp.client-ip=209.85.210.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="GMNK44AP" Received: by mail-pf1-f173.google.com with SMTP id d2e1a72fcca58-7ba55660769so2728788b3a.1 for ; Sun, 21 Dec 2025 18:19:44 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1766369984; x=1766974784; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=+xC8gGwzSYciG4fQIhKWCC6bP7IzVjct1WkD0coKXVA=; b=GMNK44APxPxmRQ7qn0WgfYW60xVdWGr6I51F+7GLDxvbgcDkyw7F0yR8uvIpNZpczo DdSbUKpaaHo0Ghhg8Bxzf7hcIuT6VKJDUQXWRZjROu0bunZj6QSCSdSzC7tLZ/2Wufk1 pULQQJJLZ4RQbjmg+AV6+/TgBqFKkk1zTIiiSvdvuJypft2z12I1ek/UBZTKbILGWt3A UAALtwb/nRhD1R4mV8JmqJQlIMQN/zYj32qD9IFuLxlMqDUiGY76xW3p2WHyFR2AtzDJ Ot4r+q2uqpTGjD1tfEujTeIN16fsijAZAI7ToJybLOyvp9ZwFubO8UqUgLIrcUjmSRFf 6psg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766369984; x=1766974784; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=+xC8gGwzSYciG4fQIhKWCC6bP7IzVjct1WkD0coKXVA=; b=McPOXELKPp/EW0JGpdqUP22suXrp0MXAzdItCE16WbZ2TrZJKh2RtYCUiv6pkLtLqB GMf89OJ4b6VxxkTnOZkvdKpLaHTo2EuJ8RZhnEx8+KiorLaJ2lRFj7INoxc2OLiSDrOo Ls6bC/pEVN9A9AJ+8EF9qGAseDHaCDCJyKu7R1X7jD6o/lOfu86YCPaKn62ikpI/2VdU E/d2igWUZdEL8W4VXDLwkMe3FSzxU/wpK3iAqlrp2pEXuuUFKO+/GzMdVqfh++wzB6/r JobdXcpvdaZm8ag3BrktMsvy7wQKKrbIM57DH4MvEw3Eq0cAnVhnrd+/HSvFLg36h8Ma K03w== X-Forwarded-Encrypted: i=1; AJvYcCV6pxQSAtTjOBnloHOV+ZYG/jBI3fKPLDyteaouRKwQ7N1+oP0PRpz+wN5lMhiBLTKxH/zOCqZnFHMhXqo=@vger.kernel.org X-Gm-Message-State: AOJu0YyJ29tTyFb7zsoCr2qdFPLT/CmvywMsrTVn2JqqNrXoeXZtmZ2Z s891CiOeRqHcfg6SPzX4FbfYHV/ORSqskRzsJ9XxsxBR46okJ1LayfS7 X-Gm-Gg: AY/fxX4A4+Fu/JIAK2yJ6ub8Qg+gG82cF9UbBZniIpBXyW4/b08VPizI0pEzp+r5vj5 rObpybo+HPQkWfWSSXa8br+uJFuW0sU9lAQuMrqwgIXbVMZ97FR8HNlAJiCnb2lDITwF/MpQUOs xep1uN5jvV8R2k28KgRRvWdGy18e8VFULQfxMnioAFUGDl3CxuR7xHGIDl6+ajvPuz+nMS3aVpE 6+clXM/r9ZlPTb32m5HKkkBFcWIpmL3leDliPD+N4c9m0hsEzbtEaem94qsoZAB0eVp3aGloBH5 wbuQmGOdJJurfopyo4IDrzYbcoXwxpGe5iP8IitmCqv6HsKHPjb6SxeOFCAplsIMddVcb7VOte3 ZpKSmcJLDysY8ZE8l8lhqslFMsJQvOP/0yBin5efYP3aRy8sLPXJVpuF4BxgP+pNTuCp2Q7YTVQ dAZU8anWO3F5niNMiF1A07g0iaaEvH8zxsmpxJle3bew== X-Google-Smtp-Source: AGHT+IGtKIPvMh+UidQEYVCpswR0H2VxjsJ/H3opaSrKaQoreVXhN6d9jdrUcZt+fVXVNXXvDGbSWQ== X-Received: by 2002:a05:6a20:72a5:b0:34f:68e9:da94 with SMTP id adf61e73a8af0-376a81dcc8fmr9628924637.30.1766369984165; Sun, 21 Dec 2025 18:19:44 -0800 (PST) Received: from [192.168.255.10] ([43.132.141.24]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-c1e79a17fdesm7730260a12.8.2025.12.21.18.19.37 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sun, 21 Dec 2025 18:19:43 -0800 (PST) Message-ID: Date: Mon, 22 Dec 2025 10:19:35 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 19/23] sched/cache: Avoid cache-aware scheduling for memory-heavy processes To: K Prateek Nayak , "Chen, Yu C" Cc: Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org, Tim Chen , Peter Zijlstra , Vincent Guittot , "Gautham R . Shenoy" , Ingo Molnar , Vern Hao References: <91a7c325-5093-4417-aa98-34df694b0c39@gmail.com> <61cc2b92-1b5a-4af5-9d88-96097c3b0619@intel.com> <2924fe29-e813-4dc5-8a1b-86890029e372@gmail.com> From: Vern Hao In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 2025/12/19 11:14, K Prateek Nayak wrote: > Hello Vern, > > On 12/18/2025 3:12 PM, Vern Hao wrote: >> On 2025/12/18 16:32, Chen, Yu C wrote: >>> On 12/18/2025 11:59 AM, Vern Hao wrote: >>>> On 2025/12/4 07:07, Tim Chen wrote: >>>>> From: Chen Yu >>>>> >>>>> Prateek and Tingyin reported that memory-intensive workloads (such as >>>>> stream) can saturate memory bandwidth and caches on the preferred LLC >>>>> when sched_cache aggregates too many threads. >>>>> >>>>> To mitigate this, estimate a process's memory footprint by comparing >>>>> its RSS (anonymous and shared pages) to the size of the LLC. If RSS >>>>> exceeds the LLC size, skip cache-aware scheduling. >>>> Restricting RSS prevents many applications from benefiting from this optimization. I believe this restriction should be lifted. For memory- intensive workloads, the optimization may simply yield no gains, but it certainly shouldn't make performance worse. We need to further refine this logic. >>> Memory-intensive workloads may trigger performance regressions when >>> memory bandwidth(from L3 cache to memory controller) is saturated due >> RSS size and bandwidth saturation are not necessarily linked, In my view, the optimization should be robust enough that it doesn't cause a noticeable drop in performance, no matter how large the RSS is. > Easier said than done. I agree RSS size is not a clear indication of > bandwidth saturation. With NUMA Balancing enabled, we can use the > hinting faults to estimate the working set and make decisions but for > systems that do not have NUMA, short of programming some performance > counters, there is no real way to estimate the working set. I see the challenge, but the reality is that many production workloads have large memory footprints and deserve to see performance gains as well. In my testing with Chen Yu on STREAM, it's intriguing that the performance is fine without |llc_enable| but drops significantly once it's turned on.I sincerely hope this situation can be optimized; otherwise, we won't be able to utilize these optimizations in large-memory scenarios. > > Hinting faults are known to cause overheads so enabling them without > NUMA can cause noticeable overheads with no real benefits. > >> We need to have a more profound discussion on this. > What do you have in mind? I am wondering if we could address this through alternative approaches, such as reducing the migration frequency or preventing excessive task stacking within a single LLC. Of course, defining the right metrics to evaluate these conditions remains a significant challenge. > > From where I stand, having the RSS based bailout for now won't make > things worse for these tasks with huge memory reserves and when we can > all agree on some generic method to estimate the working set of a task, > we can always add it into exceed_llc_capacity(). >