From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f48.google.com (mail-pj1-f48.google.com [209.85.216.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C52DA2FE579 for ; Mon, 22 Dec 2025 02:49:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.48 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766371763; cv=none; b=MOg4BNEvxYWJxsc3FM7YXSn1hSE8PfkmcLS7bmWj+vXYNWOOnnPdu/4C2ed8x2OBDjchMlzhdb5kOfNUvn6kvi1DvSUAGZTZ3na/K8GYuY942EPcIPp4h/nVA4Zak8epena4w5i2xZmtcLj0IivVP090KDWm207GcOuDAigPdss= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766371763; c=relaxed/simple; bh=yhedNj7pvnVSwm1qxdSfb5wajKwNbrvQdjf+zYOsh6U=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=mLdLZpUzCAee6my9/GUVPbbKpi9thudf19CO8BFRa7fCtvUPjJhi3EN20e+g4FmShazQaXMNirZ+bKUveXR1F/C1r1teu9YQ+d96DyxckpORAj5arP18tHBzNs/OW3v6LkGIbLM22bLTrXNy76hQkAIgdCc0fE36Er+CBdZksgA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=VF8IZsos; arc=none smtp.client-ip=209.85.216.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="VF8IZsos" Received: by mail-pj1-f48.google.com with SMTP id 98e67ed59e1d1-34c24f4dfb7so2640976a91.0 for ; Sun, 21 Dec 2025 18:49:21 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1766371761; x=1766976561; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=TtRfWTGXxplZbjJuJt9wEgCFEyBdrv21B5LT/YyJyZM=; b=VF8IZsosQiEN1Rww1xBXtyVwajbsnkZfK0eEmvWc2r9uj3OYUZDZtSvMGCjAuvHRe8 lzDuYyKsk3rkEDYhe/kYUkUh3TQ3uEN4Rlrzji+1PaQF8vJly1DP0nTiGbvJhcBr/l/o 9mUcN6NL33gwFuXHy8VdC5ipBEYpeg+nQdYRCeQRIfd12kzS9kmOEZ/+HAGp1aI/dYBJ NVuIdASApg6h1+r6PkZwQjIfvrX652trdadrRS3VKu3InneHmRgCLoNLleYXJ9MqTgGx 51oBK5wrg8B3o850KehUO+qvVw1hqheImKNlSEBDkZClTXlSwuzXnQU/szIy83Pzya9f 7HTA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766371761; x=1766976561; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=TtRfWTGXxplZbjJuJt9wEgCFEyBdrv21B5LT/YyJyZM=; b=AuYLDMkWRgLNBc+84ch4F+SSW3x2hTTmqGH1cLxt8+Hcw9RLAHYYQYvkysg0kLOQw8 N0Ip/AiGSSf6zzVBMJ1GOcr+Rr9WZID9dXhW0nhK4zK5Txahip5NZ967ATsVpV8rrOzn rpKJDKckT2a3RIyq8USjozisKJpiHUGf3BfOd/A7dE/YGxjJHsIdentGNdYxcTPdvAex rEhzZ1EVLc7dcE9tsW2TVcsgSF8vA0tz1LPsbO54MFv99yKJTr06cWBg1u26wX8QBtLM /rsQtV+TrP8pLkUZUJFlOr6W86Uvw68kAEtiuzmADO+x1w2EpSu0POD4CFMTfIF2GuLQ B98A== X-Forwarded-Encrypted: i=1; AJvYcCWOKPvAQ7Ym72ZDrcxKn5RFLw+Ue0SjUdr3ccSnX5NDti2Voyp4jub4q7cYApeSCO+dxQAC3AlF/pBMM1A=@vger.kernel.org X-Gm-Message-State: AOJu0YwKf64Rg73FGeK4FfOViI99znR51wxttpGB0gXiZVKI/DdHvTfz tmFo/BYzQR4ga+HrD6gwmHx5eB1i9NfNJyUgEfz6gjPtHpcRzWo8C0UU X-Gm-Gg: AY/fxX5wmBHroNjNVV6Y71C7/mqre1q6UsIMYIVBi2Uig2yre6xpS2GbQqrzv9/XujM mM2799+am0i8Yl2r5NWAODh8lpcuw78j3jyvrHJBxMQB8OB4mPMKCkjIQ82ntlPc8iI9ixfNW23 3Xu+1JPHb29meIFYEdYyJjU+qBUDwiuvb1f7RSjiYrb50fV8lvhktYPj94NBR4psGyMj36VBP1B bXwVecJ8Ozx03fvVkSscD7CIfsYcQrjtORXwLQtHo6Bv4s7yo35GQZNj38NmPJ1i23J3PJ+ap3R ACch25MUU1NkDbiuez+wb+tEbnhrh77YlgqyMFVSVW5dHPlZ82rZ2YnQI7R//G10Q63HvF4rmab OW5IsaiRX4BwDUy2QadgWh+8T+/6SS4GVnRUNt7YCvqf4Eyh6Izsgjyjf5RVIuxN5CNA99yv3sy 8WvKAV7DQDwY8hb8YkSsRVemZ+7isRRrT39u+Qeg5pyA== X-Google-Smtp-Source: AGHT+IEgMLAFizeQYH45KvdA8evc0Dh8AYGQmqJ3x2QSxBYuncNSgGsvcKj1B1ZyV4cPbLZtTgRptw== X-Received: by 2002:a17:90a:a790:b0:341:2141:df76 with SMTP id 98e67ed59e1d1-34e921448f1mr5520369a91.13.1766371760856; Sun, 21 Dec 2025 18:49:20 -0800 (PST) Received: from [192.168.255.10] ([43.132.141.25]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-34e920c9a7csm8234309a91.0.2025.12.21.18.49.14 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sun, 21 Dec 2025 18:49:20 -0800 (PST) Message-ID: <6da4333d-2f64-4b4b-8b51-8d0ca937b946@gmail.com> Date: Mon, 22 Dec 2025 10:49:11 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 19/23] sched/cache: Avoid cache-aware scheduling for memory-heavy processes To: "Chen, Yu C" , K Prateek Nayak Cc: Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org, Tim Chen , Peter Zijlstra , Vincent Guittot , "Gautham R . Shenoy" , Ingo Molnar , Vern Hao References: <91a7c325-5093-4417-aa98-34df694b0c39@gmail.com> <61cc2b92-1b5a-4af5-9d88-96097c3b0619@intel.com> <2924fe29-e813-4dc5-8a1b-86890029e372@gmail.com> <94c8c1af-d9a5-411b-bc54-a7b28d6cff29@intel.com> From: Vern Hao In-Reply-To: <94c8c1af-d9a5-411b-bc54-a7b28d6cff29@intel.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 2025/12/19 20:55, Chen, Yu C wrote: > On 12/19/2025 11:14 AM, K Prateek Nayak wrote: >> Hello Vern, >> >> On 12/18/2025 3:12 PM, Vern Hao wrote: >>> >>> On 2025/12/18 16:32, Chen, Yu C wrote: >>>> On 12/18/2025 11:59 AM, Vern Hao wrote: >>>>> >>>>> On 2025/12/4 07:07, Tim Chen wrote: >>>>>> From: Chen Yu >>>>>> >>>>>> Prateek and Tingyin reported that memory-intensive workloads >>>>>> (such as >>>>>> stream) can saturate memory bandwidth and caches on the preferred >>>>>> LLC >>>>>> when sched_cache aggregates too many threads. >>>>>> >>>>>> To mitigate this, estimate a process's memory footprint by comparing >>>>>> its RSS (anonymous and shared pages) to the size of the LLC. If RSS >>>>>> exceeds the LLC size, skip cache-aware scheduling. >>>>> Restricting RSS prevents many applications from benefiting from >>>>> this optimization. I believe this restriction should be lifted. >>>>> For memory- intensive workloads, the optimization may simply yield >>>>> no gains, but it certainly shouldn't make performance worse. We >>>>> need to further refine this logic. >>>> >>>> Memory-intensive workloads may trigger performance regressions when >>>> memory bandwidth(from L3 cache to memory controller) is saturated due >>> RSS size and bandwidth saturation are not necessarily linked, In my >>> view, the optimization should be robust enough that it doesn't cause >>> a noticeable drop in performance, no matter how large the RSS is. >> >> Easier said than done. I agree RSS size is not a clear indication of >> bandwidth saturation. With NUMA Balancing enabled, we can use the >> hinting faults to estimate the working set and make decisions but for >> systems that do not have NUMA, short of programming some performance >> counters, there is no real way to estimate the working set. >> >> Hinting faults are known to cause overheads so enabling them without >> NUMA can cause noticeable overheads with no real benefits. >> >>> We need to have a more profound discussion on this. >> >> What do you have in mind? >> >>  From where I stand, having the RSS based bailout for now won't make >> things worse for these tasks with huge memory reserves and when we can >> all agree on some generic method to estimate the working set of a task, >> we can always add it into exceed_llc_capacity(). >> > > Prateek, thanks very much for the practical callouts - using RSS seems > to be > the best trade-off we can go with for now. Vern, I get your point > about the > concern between RSS and actual memory footprint. However, detecting > the working > set doesn’t seem to be accurate or generic in kernel space - even with > NUMA fault statistics sampling. One reliable way I can think of to >  detect the working set is in user space, via resctrl (Intel RDT, AMD > QoS, > Arm MPAM). So maybe we can leverage that information to implement > fine-grained > control on a per-process or per-task basis later. OK, I agree, thanks. > > thanks, > Chenyu