From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ed1-f54.google.com (mail-ed1-f54.google.com [209.85.208.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ABEDA198A17 for ; Thu, 16 Apr 2026 00:27:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776299277; cv=none; b=t0kD3Se/HmGu04S0WW15UznXortSBpvRmEGfUVWzHvb+4PU6NEKovPyiVV+SDgTQtJlgVsE0+o8bEuGerbqIjioWspvUKH5oX29vBUXNp8laWeY3n+C0EVmWMWLNb57EUxy+0id4dewJiimFTyZpLiwkkdUeAKtliLmsHIh3GH8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776299277; c=relaxed/simple; bh=HpgSlS4JIlfOR6SBd10GHjtZuEA4+Q5QthdqV+vtT9c=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=jno89Ces/vSyYiCTj8qa7ia7KuAIjWJZZMD27Fm6+oOMJEsYQdbA5uFoWu1Fm6TvYeWalsSB6798G17w0K5r0mjQqh2+6uK6ZTjR1oDNGDpmJPNETjz2Nw5fEDCFj/GWJUH+q7ZLVsWcfvoJ7UrgazlZrRIJFWvTVxTTfU1RchA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=layalina.io; spf=pass smtp.mailfrom=layalina.io; dkim=pass (2048-bit key) header.d=layalina-io.20251104.gappssmtp.com header.i=@layalina-io.20251104.gappssmtp.com header.b=OMzuodJG; arc=none smtp.client-ip=209.85.208.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=layalina.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=layalina.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=layalina-io.20251104.gappssmtp.com header.i=@layalina-io.20251104.gappssmtp.com header.b="OMzuodJG" Received: by mail-ed1-f54.google.com with SMTP id 4fb4d7f45d1cf-66bd4e0560fso152270a12.0 for ; Wed, 15 Apr 2026 17:27:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=layalina-io.20251104.gappssmtp.com; s=20251104; t=1776299274; x=1776904074; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=BekkhlooZzjZ9OsNkfpJmZF2gtbmVKYIPu48NzrpIE8=; b=OMzuodJGRMMR0saOSaz+XfI6k0s31hE+0T8iF0YbrE2CVxi8BFJh+6VGhuzeta8KZU umxYRi1dlS/DDWN3LnLLgXBfbtFFwS8vmuP/K+LTKxaz7FsFRTZTLhxJBIPEJ+EdBbzE e23/jpfLh+0VZpLedWs18mkYYdx+r2kzUfROaJsce+xrnkEPV7SMxoiprCsi4cZSjTfR y0i8K3wuRY0bBsw2dzMf7DNspwqrpv6d1G8T9eo4k5SCI0x4KG4JnqG2NNRCt+a6t8M1 0i61I6Skv9tiCytWgpfSg8J5mxxhvPHfYy5EjyQSVqVMixNyGUCXQrQmiYf5sGiDnS7j zNOg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1776299274; x=1776904074; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=BekkhlooZzjZ9OsNkfpJmZF2gtbmVKYIPu48NzrpIE8=; b=CmlsFLIzAhxMkYpva4/mH4SRsc5kLpDpvPrvwY2O99zDXv+my1w4Mw0mm6IIwAN75M VaGs4jf1VARC6eVikqBMNkDgLos7X28tC8HRQnREXJsVe9O37usKOu9cLWnmOT73LH82 lyNNseY3HhmLmzrH9+OLCgtwhhyiBf+Fm48Ug+arUFSrfQw3vO1rxgxBHCpYD9YJxEpk 458xpxkeu+etp/HI7HhSxyOfXACVsfrzYNwqODKPe27DRMluxlqz/lHtrlF+xLTR3LX1 6uqmJhcyS3PYoxLHBWMxEtCAioF2/p6m2QNcRdffo9eF3il8MWqgZB5mnJ6LIQ7iB4IV 6jNA== X-Forwarded-Encrypted: i=1; AFNElJ8Q7R9rXcWdbDe0zse97Vb4nsV4IUvUVoID+OVOh2EshjFm5GtByQEW69HMEKmxy/1NukGKi0NBQsEDv3Q=@vger.kernel.org X-Gm-Message-State: AOJu0YyatpehfoWvaPk+8a5eA7+Op7UnP7qqm37Q8dCxzAmqorPWRddf gTomuEac58OqrpjQIx1H1BpDNCZ3kWIJ6/QOTNixZcNfhCDSYa7Q4lGH4laZNhccOik= X-Gm-Gg: AeBDieuCaJ06YIj2gTuYjvr4S7pDcw/HpIROr1O504nRHHGhS/1+ViK0MLubMUx1AG7 SYzELKY9I7x6irOb+4REcvXwwm00TUyLIkg7GWCGjL3HGGa6xJTGArvz0iMtmxPJevXZjDbQYap Az7xlTfabrPF27PPnI/5uGPr6DrABMFpYMqQFXOQBPLo+IMhMLM0yqYGYvl+aMbQtykkM73GiTz 8rHv5/q7K/1Exg2aJn4N6QjwjuCexXrpDvFjeYbMFTlp2Tbn1atLhZdyLmpWLSFMaiXvavKlNUb WpxFKS2QNAgBTsMSVppCxtjF57r2cEySaC/gZrOh4gIu7mMQGTHwxuBigL7b5J4dlS/IwSVIH7w vQCc8xiR/6N90ARkvdRSz04PGoEm5qm5Cph/vJpdw5CbKFkQsS8N8T9tdGdRonlrHJSkk1+Xj99 n+0Nu0Gpve8vvp0/RPgzl9Jpt68ndZ X-Received: by 2002:a17:907:1b1e:b0:b9d:7b9d:df02 with SMTP id a640c23a62f3a-ba2693bf42bmr79847666b.17.1776299273773; Wed, 15 Apr 2026 17:27:53 -0700 (PDT) Received: from airbuntu ([146.70.179.107]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-ba2677d6277sm39209466b.61.2026.04.15.17.27.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 15 Apr 2026 17:27:53 -0700 (PDT) Date: Thu, 16 Apr 2026 01:27:49 +0100 From: Qais Yousef To: Tim Chen Cc: Peter Zijlstra , Ingo Molnar , K Prateek Nayak , "Gautham R . Shenoy" , Vincent Guittot , Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , Josh Don , Gavin Guo , Libo Chen , linux-kernel@vger.kernel.org Subject: Re: [Patch v4 00/22] Cache aware scheduling Message-ID: <20260416002749.muyrcycmtabksav4@airbuntu> References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline In-Reply-To: On 04/01/26 14:52, Tim Chen wrote: > This patch series introduces infrastructure for cache-aware load > balancing, with the goal of co-locating tasks that share data within > the same Last Level Cache (LLC) domain. By improving cache locality, > the scheduler can reduce cache bouncing and cache misses, ultimately > improving data access efficiency. The design builds on the initial > prototype from Peter [1]. > > This initial implementation treats threads within the same process > as entities that are likely to share data. During load balancing, the > scheduler attempts to aggregate such threads onto the same LLC domain > whenever possible. > > Most of the feedback received on v3 has been addressed. Some aspects > could be enhanced later after the basic cache-aware portion has landed: > > There were discussions around grouping tasks using mechanisms other > than process membership. While we agree that more flexible grouping > is desirable, this series intentionally focuses on establishing basic > process-based grouping first, with alternative grouping mechanisms to > be explored in a follow-on series. > > There was also discussion in v3 that the task wakeup path should be used > to perform cache-aware scheduling. According to previous test results, > performing task aggregation in the wakeup path introduced task migration > bouncing. Primarily that was due to the wake up path not having the up > to date LLC load information. That led to over-aggregation that needed > to be corrected later in load balancing. Load balancing path was chosen > as the conservative path to perform task aggregation. The task wakeup > path will be investigated as a future enhancement. I posted schedqos announcement yesterday, which I think (hope) would be the right way to address these concerns about tagging tasks. https://lore.kernel.org/lkml/20260415000910.2h5misvwc45bdumu@airbuntu/ It would be trivial to add experimental branch to add new QoS flavour to say NUMA_SENSITIVE etc. I am still trying to think of a generic description to address a number of use cases (see Execution Profiles in README.md), not just this particular numa sensitive one, but the experimental branch should help iterate and drive the kernel development for wake up path + push lb instead of using load balance which I really doubt will work well in practice since this is slow to react, and you're relying on overcommitting the system by default by making every task of every process data dependent and require it to be co-located. I think in practice admins will care about specific applications to be kept within a single LLC, and if they are willing to spend the effort, they can tag specific tasks of a specific application. We are trying to make sure we have one coherent story to define these type of QoS requirements, and delegate to userspace to make these decisions/policies. The current line of thinking is that wakeup + push lb should be generally good to address the different needs for various placement requirements. I understood you believe the same. If not, it would be good so we can think how to further generalize. Also QoS IMHO should be viewed as a scarce resource. For best effort delivery (which is the best we can do in reality, this is not hard real time system), it is easier to provide good best effort when the average noise level is low, ie: few tasks are required to be kept within the same LLC. If we overcommit often, we will crumble often. So IMHO the key is to delegate to userspace to tag, and make them take responsibility of handling potential overcommit and decide which workload is really important to tag and which one they can let go of or move to another machine to get their desired perf/latencies. For the kernel interface to tag tasks and set a cookie, I plan to rebase and repost [1] as I need it for rampup multiplier to help counter DVFS related latencies and slow migration in HMP systems. The idea was for it to be generic and flexible to add whatever; in this case I think we need to add a QoS to tell scheduler these tasks are data co-dependent with a unique cookie, which implies that they need to stay within the same LLC, and can be extended to help keep these tasks within the same L2 or L1 if it makes sense (ie: they are small and can be packed on the same CPU). I am not opposed to merging this first, but I think the load balanced based approach is wrong and if merged must be removed later in favour of wakeup + push lb one with user space based tagging. Based on what Peter said the wake up path is trivial to add, push lb is almost ready, and I hope we have the tools to auto tag processes/tasks now to potentially try to work on this approach first instead. [1] https://lore.kernel.org/lkml/20240820163512.1096301-11-qyousef@layalina.io/