From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f176.google.com (mail-pf1-f176.google.com [209.85.210.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B23928F7D for ; Wed, 17 Dec 2025 01:17:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.176 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765934253; cv=none; b=mb6DFViH5HgDOUcb+x9YNA0wD/B7GWfrLjRXCcoLBW9GwdivtO4oTnWDr47MHYIj9FKI4dWQ8wsuf1VDNPp66lXpCXv/4UZNrfvhNPlcY9QEk7pi5VCDjXeNy+T3Ebsm2542fCvXD8Ha3NFDPXzas/In75Wjiy+199RO5lgdCW4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765934253; c=relaxed/simple; bh=S4BAjx3K9mbtDbysZVoQAnWcIsKJHU2A7vbc65eQ5Tg=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=B298D0XF7hNa6ZYbUWKGAkVWMGw5cFFK7FnVAKWB8JkiU9GwjyRsoeD9L8ihwBKWOx09cwjToB3V8GH6mVD0ql365rf0uN0G9Amdll43FDkT9mvHViVNk0U6rpDvzMJvF2ka3LdxR7S0RbwUFh4dEjZtUuf9/C16IH9jylKE7i8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=etWbdHHl; arc=none smtp.client-ip=209.85.210.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="etWbdHHl" Received: by mail-pf1-f176.google.com with SMTP id d2e1a72fcca58-7e1651ae0d5so4335055b3a.1 for ; Tue, 16 Dec 2025 17:17:31 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1765934251; x=1766539051; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=Mov8o4TbNylCcSi9hPiNepL3O+pfiutLqNZ3MBrseBQ=; b=etWbdHHl2RZHS+h1pJ/FyJ5R309A1T7hVVqNxDuos5Usa20xyKf0bO2FT/vEqqL6mw IYqOFGLBvqY6uXrujtWtV9TpVaEz0mXF7CHs3ZlKIX1FsAXHq7RvCYxluAf9nPiBt+Dd /ViidhDtpctI3NjesLiAFn81TdIvdymH3p/bF2x+UZUMlh2e3bKYX0DAjp9k/OsRsKsh Fzq9KR5a2irtKdt/RYO59V7qslpocydHeT1vefsE13SkieU2hrs+VmphuDDEvBvkP2uU ASddUxf4xtSNS0o46ioO0jqA0nlOLD8hgeIrvqNBSvmcrUfqpvnN/DUVTclBvBXkK6Cd Yc5Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1765934251; x=1766539051; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=Mov8o4TbNylCcSi9hPiNepL3O+pfiutLqNZ3MBrseBQ=; b=DQ8M4f52vwMyxBDYx8l7d9COT+JikmqNj55c+rqeEfemQqNcC8nyo1XX+enEktgfj9 iKPGGhlUmfRnsX7oYYcL+oymZUbMrCaXTcv7cSjO1+IbUnlozaNmemXwORb4fz2NKcx8 7lInlAVIBgAKDl1jzByNrtXEIaoJRBDrgWZiGotpZdtljBBrqPctMFrethk9tiQXPGqK C8xcKSG6igBpHqu+ZD+UAuKx8wT8Pt/8yZ3e1ubAzVPBN3mGpgs+gSr8PeZ6HmPtDa4n MsVt2iHIPwri1xyWWQ7HX7/a2PPE0FszKtllct56qRelLY3CRYNtyIAKaUEvpWfJRiTZ 0PUw== X-Forwarded-Encrypted: i=1; AJvYcCVbyQzev/FlgSA99o7SrbqUzqX/QTEgHvUhMhJFfGwkk5cUpKfXzjMShj3OsUY7neR+utSPEnTrYFQssfE=@vger.kernel.org X-Gm-Message-State: AOJu0Yyxy+0ySh6XORPdK+pTkDQl3Otbcs74QnNrNUZbS1kKSG1IYn57 Soybtz9wVCGHABFE86NeMe3RoWBUzsFuDgqWS4COn18laMqE5P/zVVJS X-Gm-Gg: AY/fxX4YTyPZydAfeulEnIJLp1G2sT6ShXfyDxx2glGpjDoM1yWYsyg2+0OgB5WcE9y cVGG0+cT394jJr0qDbEjBrxL51jhZY+oSc0DNdbA7FDmgAz/VJCWfty44+tIa4MmnwII3N7p8eX KmVc4A9hLTeaWinXdAdtLvW9hYHNo3DufmXZUFvBoQIA+JJbNxFKug1fiSySRHEjLcGWrJZOqrE UGW2TxQNBRdCHNvO7zxFOVT2Z16JJCEIMGAgv6tfwRVfs8dhJ+Dq0ylGlLW9sHTketbTkYBfuHD Z2kZ6kIuxUPXgZEHw+12Dgfi6unEyWOuJm4zKpVkQji3VEWiZB2iFkMcLqWtz0zIkcfmHGeolW5 ArDfSM7gkWuWpm9RXCDZLrtQVKwwQPTqcCOh/QSmLcXlyCwwsMJkL50EaRZMLMlwhRFCgwSG74O MNJMTP7LzjFYVbjx/ZxSSKkqZ8p4KUbr9lWTJDhA== X-Google-Smtp-Source: AGHT+IHWN5WDgCDTIkv99YxW2LCivJ2eu7FhuZFouw1+yM7UArs6uB0v8BV22Rb6eCJVXU6ko/kJnA== X-Received: by 2002:a05:6a00:8016:b0:7a9:e786:bdaf with SMTP id d2e1a72fcca58-7f66763d327mr12983845b3a.14.1765934250871; Tue, 16 Dec 2025 17:17:30 -0800 (PST) Received: from [192.168.255.10] ([43.132.141.20]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-7fcb8f101casm792094b3a.18.2025.12.16.17.17.24 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 16 Dec 2025 17:17:30 -0800 (PST) Message-ID: <7d5bb7c4-abc5-470e-84fe-72a3b1d3a2f4@gmail.com> Date: Wed, 17 Dec 2025 09:17:20 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 01/23] sched/cache: Introduce infrastructure for cache-aware load balancing To: "Chen, Yu C" Cc: Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org, Peter Zijlstra , Ingo Molnar , K Prateek Nayak , Vincent Guittot , "Gautham R . Shenoy" , Tim Chen References: <06f0d7edbc3185ec730b50b3b00d87ace44169b3.1764801860.git.tim.c.chen@linux.intel.com> <7e4640a2-f79f-4f14-b099-d97bfd842b37@intel.com> From: Vern Hao In-Reply-To: <7e4640a2-f79f-4f14-b099-d97bfd842b37@intel.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 2025/12/16 14:12, Chen, Yu C wrote: > On 12/11/2025 5:03 PM, Vern Hao wrote: >> Hi, Peter, Chen Yu and Tim: >> >> On 2025/12/4 07:07, Tim Chen wrote: >>> From: "Peter Zijlstra (Intel)" >>> >>> Adds infrastructure to enable cache-aware load balancing, >>> which improves cache locality by grouping tasks that share resources >>> within the same cache domain. This reduces cache misses and improves >>> overall data access efficiency. >>> >>> In this initial implementation, threads belonging to the same process >>> are treated as entities that likely share working sets. The mechanism >>> tracks per-process CPU occupancy across cache domains and attempts to >>> migrate threads toward cache-hot domains where their process already >>> has active threads, thereby enhancing locality. >>> >>> This provides a basic model for cache affinity. While the current code >>> targets the last-level cache (LLC), the approach could be extended to >>> other domain types such as clusters (L2) or node-internal groupings. >>> >>> At present, the mechanism selects the CPU within an LLC that has the >>> highest recent runtime. Subsequent patches in this series will use this >>> information in the load-balancing path to guide task placement toward >>> preferred LLCs. >>> >>> In the future, more advanced policies could be integrated through NUMA >>> balancing-for example, migrating a task to its preferred LLC when spare >>> capacity exists, or swapping tasks across LLCs to improve cache >>> affinity. >>> Grouping of tasks could also be generalized from that of a process >>> to be that of a NUMA group, or be user configurable. >>> >>> Originally-by: Peter Zijlstra (Intel) >>> Signed-off-by: Chen Yu >>> Signed-off-by: Tim Chen >>> --- >>> >>> Notes: >>>      v1->v2: >>>         Restore the original CPU scan to cover all online CPUs, >>>         rather than scanning within the preferred NUMA node. >>>         (Peter Zijlstra) >>>         Use rq->curr instead of rq->donor. (K Prateek Nayak) >>>         Minor fix in task_tick_cache() to use >>>         if (mm->mm_sched_epoch >= rq->cpu_epoch) >>>         to avoid mm_sched_epoch going backwards. >>> >>>   include/linux/mm_types.h |  44 +++++++ >>>   include/linux/sched.h    |  11 ++ >>>   init/Kconfig             |  11 ++ >>>   kernel/fork.c            |   6 + >>>   kernel/sched/core.c      |   6 + >>>   kernel/sched/fair.c      | 258 >>> +++++++++++++++++++++++++++++++++++++++ >>>   kernel/sched/sched.h     |   8 ++ >>>   7 files changed, 344 insertions(+) >>> >>> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h >>> index 90e5790c318f..1ea16ef90566 100644 >>> --- a/include/linux/mm_types.h >>> +++ b/include/linux/mm_types.h >>> @@ -939,6 +939,11 @@ typedef struct { >>>       DECLARE_BITMAP(__mm_flags, NUM_MM_FLAG_BITS); >>>   } __private mm_flags_t; >>> +struct mm_sched { >>> +    u64 runtime; >>> +    unsigned long epoch; >>> +}; >>> + >>>   struct kioctx_table; >>>   struct iommu_mm_data; >>>   struct mm_struct { >>> @@ -1029,6 +1034,17 @@ struct mm_struct { >>>            */ >>>           raw_spinlock_t cpus_allowed_lock; >>>   #endif >>> +#ifdef CONFIG_SCHED_CACHE >>> +        /* >>> +         * Track per-cpu-per-process occupancy as a proxy for cache >>> residency. >>> +         * See account_mm_sched() and ... >>> +         */ >>> +        struct mm_sched __percpu *pcpu_sched; >>> +        raw_spinlock_t mm_sched_lock; >>> +        unsigned long mm_sched_epoch; >>> +        int mm_sched_cpu; >> As we discussed earlier,I continue to believe that dedicating >> 'mm_sched_cpu' to handle the aggregated hotspots of all threads is >> inappropriate, as the multiple threads lack a necessary correlation >> in our real application. >> >> So, I was wondering if we could put this variable into struct >> task_struct, That allows us to better monitor the hotspot CPU of each >> thread, despite some details needing consideration. >> > > I suppose you are suggesting a fine-grained control for a set of tasks. > Process-scope aggregation could be a start as the default strategy( > conservative, benefit multi-thread workloads that share data per process, > not introduce regression). Yes, in our real-world business scenarios at Tencent, I have indeed encountered this issue where multiple threads are divided into several categories to handle different transactions, so they are not share the hot data, the 'mm_sched_cpu'  does not represent all of their task, so add a control interface such as cgroup or others will be a good idea. > > On top of that, I wonder if we could provide task-scope control like > sched_setattr(), similar to core-scheduling cookie mechanism, for > users that want aggressive aggregation. But before doing that, we need a > mechanism that that leverages a monitor system(like PMU) to figure out There will maybe a trouble, If the environment is running on a VM, We could use tags to differentiate these tasks and do some tests to verify the performance difference between unifying the |mm_sched_cpu| and not unifying. > if putting these tasks together would bring benefit(if I understand > Steven's suggestion correctly on LPC), or detection tasks that share > resource, then maybe leverage QOS interfaces to enable the cache-aware > aggregation(something Qias mentioned on the LPC). > > thanks, > Chenyu >