From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from BN1PR04CU002.outbound.protection.outlook.com (mail-eastus2azon11010062.outbound.protection.outlook.com [52.101.56.62]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E3DEC171CD for ; Mon, 19 Jan 2026 04:49:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.56.62 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768798142; cv=fail; b=klQ4gMr71HvbDreCELGisTaubVlIlauBoybPTZS6xxgg0JYZIFYJvihzjrapvtQ+eCdBEz37zstLIrNxlDR5KDBbPuy62p5XoRixaZeZX5CGnZvLeDE7J9nClExjH6Dh4K4AO72iGiJtuRhrg9Pv8JhPMiUChoQlwPrKj2q014w= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768798142; c=relaxed/simple; bh=oKlKdv/aEJWNXSgZsP4+Q+IcrEx4c6TQz8VZvfllDl0=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=cs+iB4iXk3JEZYPNOonbxKMirSDqU/w/YfL+riGZ0yrbXcT27gAjvq+UIiIN6sXqop1YaarU/RCgVA8z5XjURzVjyx90D+3RrXC09v2385ZzqcYBkUjZnMbImhoZrOtZap9rGwrME5BWMgWUclwdH9KEzo7KraJ/kFPI3QkO2Vc= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com; spf=fail smtp.mailfrom=amd.com; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b=Bew686z7; arc=fail smtp.client-ip=52.101.56.62 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=amd.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b="Bew686z7" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=O1hhQpKXf6+P29sOToV4Hdp+X3q3sAMTnSFtoTr3GkCwWm0W90Y5WbzmwskCTWi5sLzZdveGk8GHTzEay4sfyqAGG88Vw9C9I3O4JOh902yfQApUAbm155UTy0xIMjGrc4SUQGOM2k3VbfTPpPnJyhJTBtBbsy3PIJXX/Txg1dwA8TYMyKRJU+NZuHfb0Nkvq94WI8DDs1wR+UDgWpGJYFXrE9ydsGNYPy+9Qv7NLWSKmSB4V/I4Gt4ppO0IHsI6yjIMDP2r10qo0bMPTh8mkNsJdOZ/7Uw4SIPM576SFCZRdrP/2pHZd6FWwIMJ7TUhWhA7sftvBCn3rWif3Ctrpg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=ygGdUDYYx4xHvOJnzHpbk/w7umQdJFH3PrBzGaNhSY0=; b=GiJWdtRiRdRy6kn/C34734fY4htvCPf580vjgIpFlPwrrDL8XaLegc1SEbvyjvBow18ECIzOTY/PM1JOUiBgKFVSpJxcVpa1ZU5af8/5STX9KiFeKARF9lTpkDvmiZh8mS9OtQrqyzWgAB3qWiz6BUO0n7nEvMye/sr51o4gAedUC1WwN2sj0amoXmxs0JM5xOH9c+sRb6jAfG/BlQI1jdW0fdv5k7+a6YjyYV+4RfOqdS45D3g8fpkr/S+bFUKZ/M8sO1PmbblrRRTELWNlge9pHvtIlJpuxmCwnqwUdKLXk97n3k4SPMzRmy/xpl1KFHFvaxWr907qnxM1+10TZQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 165.204.84.17) smtp.rcpttodomain=gmail.com smtp.mailfrom=amd.com; dmarc=pass (p=quarantine sp=quarantine pct=100) action=none header.from=amd.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amd.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=ygGdUDYYx4xHvOJnzHpbk/w7umQdJFH3PrBzGaNhSY0=; b=Bew686z7C1+KWvXx0/kYBgoEHbxQIdegw11jA1EbdUgGMH/GayA7Ld9pRe2ZPlCP212Pq3CPT0jbhsnGj3J4AsCwtYLffwnuEgENL5IX4bFpNw4z46SjUzqfApYW0Wz9W10r9gKtvX7VxBEZni6A3u2l1jEHJ15yNB4o2vKf+vA= Received: from BL0PR02CA0002.namprd02.prod.outlook.com (2603:10b6:207:3c::15) by LV2PR12MB5726.namprd12.prod.outlook.com (2603:10b6:408:17e::9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9520.5; Mon, 19 Jan 2026 04:48:56 +0000 Received: from BL02EPF0001A0FC.namprd03.prod.outlook.com (2603:10b6:207:3c:cafe::32) by BL0PR02CA0002.outlook.office365.com (2603:10b6:207:3c::15) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.20.9520.8 via Frontend Transport; Mon, 19 Jan 2026 04:48:57 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 165.204.84.17) smtp.mailfrom=amd.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=amd.com; Received-SPF: Pass (protection.outlook.com: domain of amd.com designates 165.204.84.17 as permitted sender) receiver=protection.outlook.com; client-ip=165.204.84.17; helo=satlexmb08.amd.com; pr=C Received: from satlexmb08.amd.com (165.204.84.17) by BL02EPF0001A0FC.mail.protection.outlook.com (10.167.242.103) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.9542.4 via Frontend Transport; Mon, 19 Jan 2026 04:48:56 +0000 Received: from satlexmb07.amd.com (10.181.42.216) by satlexmb08.amd.com (10.181.42.217) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.17; Sun, 18 Jan 2026 22:48:55 -0600 Received: from [10.85.36.78] (10.180.168.240) by satlexmb07.amd.com (10.181.42.216) with Microsoft SMTP Server id 15.2.2562.17 via Frontend Transport; Sun, 18 Jan 2026 20:48:52 -0800 Message-ID: <77d096d2-bb2d-4f51-b29f-5db688acf605@amd.com> Date: Mon, 19 Jan 2026 10:18:46 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] sched/topology: delay imb_numa_nr calculation until after domain degeneration To: Tianxiang Peng , , , , , , , , , CC: , , Tianxiang Peng References: <20260119035727.2867477-1-txpeng@tencent.com> Content-Language: en-US From: K Prateek Nayak In-Reply-To: <20260119035727.2867477-1-txpeng@tencent.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 7bit X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BL02EPF0001A0FC:EE_|LV2PR12MB5726:EE_ X-MS-Office365-Filtering-Correlation-Id: 067d7304-101b-484a-7af6-08de57160d97 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|36860700013|7416014|376014|82310400026|921020; X-Microsoft-Antispam-Message-Info: =?utf-8?B?bmYxclVzcFUyYlcyTGx2MnkyZ0xqWGhvNlZqQlFFNmR0d0N4SnVvZEtuMWgw?= =?utf-8?B?ejZGVzlwUkZqTzU1NFJWUkwraHd6ZGR4TlIweUhLakJHbnc0bjZUQ29ORURM?= =?utf-8?B?alJDQzQ3dXdxZ21xZ0ZZWEJTZlJSOVRSREpOdFRDT2xmY2E3QTVDVkR1OFlX?= =?utf-8?B?RUNHNVRnTHlCZ0JBUWlUY2tvdDJVb0t3bW1DWjBYSkxaMFVZY0xBS1hNb2NF?= =?utf-8?B?QUJUVmVmZjlRVHZZbHFQdDloUDQ3bDY1V3AxN3F5TmdGMDJ4QjdTRzVXM3du?= =?utf-8?B?RTBuTmttU2hZZGlhYTNOK21zYTYxR3BkaC9pbDdXOStuNTZjZE5VM2dhVjFI?= =?utf-8?B?ZmxGeXloU25DenVtblpUQm5Ra0IwSVpKcEZxbXllNTRWOC9wbzlpelpxMndV?= =?utf-8?B?cDB2eHhhS3NEYjZnQ0I0TTdhWHRCdVB2Z0RVQm1aczRJSFBDdmRSNHBuTGx4?= =?utf-8?B?b3JqMllmYXN2aWh3eGVvcFN1N1FNczVZZGptQUVqd3A1ZFpRSGk2MG96MjRt?= =?utf-8?B?aDZlR0QySmcrRCt3Q0RGK2Qwd1haS1JqMkUydnhCTTVPWkFKVlVjcVFRWHVD?= =?utf-8?B?NXBNbW5aSXczZXFnTERtQnNodzI3S1NCVW5VSHVjeWlUZUlCZ1orcmM3eFFK?= =?utf-8?B?Qm9QN1V5MU1TaDFQN0lyZTkrY2hHd3pNRkVsblVJOEdRaW44NDZONlZpRW03?= =?utf-8?B?dXM5Z3ZkbzJ2UjNlSUl6OHVkWjVmNGJPZXhxaHIwMlFSaGRFb1p1eUdmRGF4?= =?utf-8?B?NXFYNnErVGRvc2cwQ0JjWWxWT3hJQWdUcXU2MWlySVpOTG9Zdks1eTBRL2hC?= =?utf-8?B?NHdrWVY5dVlmYXFnRStCK2Y5U2hRYnRha2NoUXZROUUwdTllVWNjNnM4SUov?= =?utf-8?B?WjZhRmtWNUZhYk1oWXZLalN0NUo5V3kvbGVzNnN2b1pHM0hmdE8vbmZUaTMv?= =?utf-8?B?dWJIYUl0QkdDa1lwcGxEbnN5aVZWWFY1MFN2M0VIa0JTOVRaT01lQ0F0dmtN?= =?utf-8?B?SzdpcE1OQnd3VUdyM3NuTnk5WlFyb0U4TE1WblFqN1BCVWFDRnNrU2JXMjJQ?= =?utf-8?B?SDdUOFltMVUzYXBxOU13T0ZXWXZqOG9XZkpkM0dpdFM4eVZjeDdmNUp1dFRN?= =?utf-8?B?ZmFkRkUrL2wvUml1ditvMXlXNXVncHV5aVBtNUg4cDhiSVhQRTFIU0VPOUY1?= =?utf-8?B?VFVEU1daZGdmNDVrTVFsdmRGa0pRc2lGTzZFMURkeTM5VS9HZ1Vxd2JURXRD?= =?utf-8?B?b215My94SWk3N0Y3YVR0L1NtTnFUOTFVUHdaREpYOE1EUnhYTVYwNWI5TnN0?= =?utf-8?B?MjNwWkVGS1VRRkJpVGlWcDkyY2ZqOWpLVWZLeXNBNXhXUHREeitTOG1YaGVS?= =?utf-8?B?TDhPY1BCNkdVTkpmUTV3M1duMlU3SVlDc3ROT3pUdUt6aDFnOHU4c1B3c0Mr?= =?utf-8?B?WUhDN2kwNGJqdFBZMVlrNUdNYVBSSTZWMlVta2oxMm1oTEJaQmJGSDR3QWFQ?= =?utf-8?B?cXZaSHBJTC9uKytSTVF0QkZMckZnL3QrUXcvN0RYN25kcCtSUkJCMlBwWm1P?= =?utf-8?B?L0lYZHJoV2p5WGhPMEdUMlE1UXUyTFN0RGJZRTI1S1Q0NGh3cDRqT3dqdk1C?= =?utf-8?B?VWhUdWFoZGptVk16cjVsTmVDNGZtcWkrWVpyR2d1M05kOTlKUU1ra3BwaVJn?= =?utf-8?B?UEl4bUNGZFJLcUFZSG1LQlVGcWZaOFpHeHlQbU5BU3NkdkhjWXdKdHBIa1BP?= =?utf-8?B?bnUyazdxTU8xTUNyV0NqSGd4eDdKQnh2QUo3YVIrRTNHaDJGUllXMERZVGtW?= =?utf-8?B?NHI4eWh6b2FXeDdDUmpOWWJVNmVSKzhBZEtHcGJYL0Vsa3BxSEx2ekY0Tnk1?= =?utf-8?B?RGxyWDdUWkxmVWhjbGp6TFkvQ3Z0dnNhRnVNNzNoZGNQTXRzSlptcXU1SDZQ?= =?utf-8?B?NGxpa01SOGVnK3dXRXp4OUJ3bnRETHUzenI0bHc3VnZ5L3NndTNrUzJCMzRR?= =?utf-8?B?Y3R1RnBWc1o4NVhwUnpHRUdVU0RJdjJGQ2doRzh2cjdkcUcwUVI5cExLVGNv?= =?utf-8?B?cTlyWC8yMkU4RHpmcnhlVW96WkxmL1RRQmJDc2JEZi9WM0tCMEZSeml2aGtG?= =?utf-8?B?NFNaNDFDMU9uZlRTSE8zKzJDQlhUU0hTV1JNeGg1TjlxTUdQS1dsSUo3blFW?= =?utf-8?B?YjlMdElaYlhNWGllQ2x1R3VNMEZpNGxNTi9KUUt1S2hsN0FINkk4cEd4bHkz?= =?utf-8?Q?nUM09SZEBGqiKFQb6lDeJjEIkabRqUxY1zx1+ssWkA=3D?= X-Forefront-Antispam-Report: CIP:165.204.84.17;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:satlexmb08.amd.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(1800799024)(36860700013)(7416014)(376014)(82310400026)(921020);DIR:OUT;SFP:1101; X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 19 Jan 2026 04:48:56.1696 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 067d7304-101b-484a-7af6-08de57160d97 X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=3dd8961f-e488-4e60-8e11-a82d994e183d;Ip=[165.204.84.17];Helo=[satlexmb08.amd.com] X-MS-Exchange-CrossTenant-AuthSource: BL02EPF0001A0FC.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: LV2PR12MB5726 Hello Tianxiang, On 1/19/2026 9:27 AM, Tianxiang Peng wrote: > Currently, imb_numa_nr is calculated before sched_domain degeneration in > cpu_attach_domain(), which might reflect a transient topology that no longer > exists. That is the intention. We want to spot systems with multiple LLCs per PKG before degenration. Before this patch, the NUMA imbalance threshold used to be 12.5% of total CPUs in the NUMA domain which was found to be sub-optimal for split-LLC architectures since B/W intensive tasks will saturate the B/W at per-LLC level. > > This is observed on our Kunpeng 920 systems (4 NUMA nodes, 80 cores per > node, 8 cores per cluster), where the initial PKG domain is redundant with > the MC domain and subsequently removed. > > Observed topology data on Kunpeng 920: > > Topology order: Child -> Parent > > [Before Patch] > (before degeneration) > Domains: CLS(8) -> MC(80) -> PKG(80) > Flags: [LLC] [LLC] [!LLC] > imb_numa_nr: 0 0 10 > > (after degeneration) > Domains: CLS(8) -> MC(80) -> NUMA(160) > Flags: [LLC] [LLC] [!LLC] > imb_numa_nr: 0 0 10 > > [After Patch] > (before degeneration) > Domains: CLS(8) -> MC(80) -> PKG(80) > Flags: [LLC] [LLC] [!LLC] > > (after degeneration) > Domains: CLS(8) -> MC(80) -> NUMA(160) > Flags: [LLC] [LLC] [!LLC] > imb_numa_nr: 0 0 2 This is too aggressive for unified LLC architecture. You start exploring the remote NUMA node just at 2 running tasks. when you have 6 other CLS domains and 78 CPUs idle on the local NUMA. See commit 23e6082a522e ("sched: Limit the amount of NUMA imbalance that can exist at fork time") which has more data to back the threshold. What workload is it helping? Probably steam with 10 threads but then unixbench spawn will regresses. If you have 10 B/W intensive long running tasks, the NUMA balancing will eventually see the local vs remote faults and set the preferred NUMA domain anyways. I'm sure a bunch of medium utilization benchmarks, especially multi-threaded ones that share data and code between threads, on other unified LLC platforms like Intel will regress. > > Move the calculation to cpu_attach_domain() after degeneration to > ensure it always reflects the effective topology. > > Signed-off-by: Tianxiang Peng > Reviewed-by: Hao Peng > --- > kernel/sched/topology.c | 115 ++++++++++++++++++++-------------------- > 1 file changed, 57 insertions(+), 58 deletions(-) > > diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c > index cf643a5ddedd..e8774e587f15 100644 > --- a/kernel/sched/topology.c > +++ b/kernel/sched/topology.c > @@ -717,6 +717,8 @@ cpu_attach_domain(struct sched_domain *sd, struct root_domain *rd, int cpu) > { > struct rq *rq = cpu_rq(cpu); > struct sched_domain *tmp; > + unsigned int imb = 0; > + unsigned int imb_span = 1; > > /* Remove the sched domains which do not contribute to scheduling. */ > for (tmp = sd; tmp; ) { > @@ -764,6 +766,61 @@ cpu_attach_domain(struct sched_domain *sd, struct root_domain *rd, int cpu) > } > } > > + /* > + * Calculate an allowed NUMA imbalance such that LLCs do not get > + * imbalanced. > + * Perform this calculation after domain degeneration so that > + * sd->imb_numa_nr reflects the final effective topology. > + */ > + for (tmp = sd; tmp; tmp = tmp->parent) { > + struct sched_domain *child = tmp->child; > + > + if (!(tmp->flags & SD_SHARE_LLC) && child && > + (child->flags & SD_SHARE_LLC)) { > + struct sched_domain __rcu *top_p; > + unsigned int nr_llcs; > + > + /* > + * For a single LLC per node, allow an > + * imbalance up to 12.5% of the node. This is > + * arbitrary cutoff based two factors -- SMT and > + * memory channels. For SMT-2, the intent is to > + * avoid premature sharing of HT resources but > + * SMT-4 or SMT-8 *may* benefit from a different > + * cutoff. For memory channels, this is a very > + * rough estimate of how many channels may be > + * active and is based on recent CPUs with > + * many cores. This entire part of comment is no longer true after you moved this logic to post degeneration. > + * > + * For multiple LLCs, allow an imbalance > + * until multiple tasks would share an LLC > + * on one node while LLCs on another node > + * remain idle. This assumes that there are > + * enough logical CPUs per LLC to avoid SMT > + * factors and that there is a correlation > + * between LLCs and memory channels. > + */ > + nr_llcs = tmp->span_weight / child->span_weight; > + if (nr_llcs == 1) > + imb = tmp->span_weight >> 3; > + else > + imb = nr_llcs; > + imb = max(1U, imb); > + tmp->imb_numa_nr = imb; > + > + /* Set span based on the first NUMA domain. */ > + top_p = tmp->parent; > + while (top_p && !(top_p->flags & SD_NUMA)) { > + top_p = top_p->parent; > + } > + imb_span = top_p ? top_p->span_weight : tmp->span_weight; > + } else { > + int factor = max(1U, (tmp->span_weight / imb_span)); > + > + tmp->imb_numa_nr = imb * factor; > + } > + } > + > sched_domain_debug(sd, cpu); > > rq_attach_root(rq, rd); -- Thanks and Regards, Prateek