From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from PH8PR06CU001.outbound.protection.outlook.com (mail-westus3azon11012052.outbound.protection.outlook.com [40.107.209.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 77A6F3B0593 for ; Thu, 30 Jul 2026 06:27:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.107.209.52 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785392851; cv=fail; b=LvjbVIjq/0NRgtygNUfWj/A12qc27qT6FnNSsTB3mY5yS1WjIKo5SPYr9lBxSAPGGsdtbaFSVMEcH4lpQJjoSw/byvzbhBkQxmR1NLE+i2GBsdaP4AF7zn7/QfjSA09FRd0X50qbeKZ/U4NlkGpdL+ab+Y5AZzcL3gZuf8BwTjE= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785392851; c=relaxed/simple; bh=Gdk6yhlPs000Hf5h0ftnQ93E3uJXptrj7OtbG+hSQcw=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=dcXqVfbqRFalum4PJmvmMc1jN1l7JxbsEIT2eAufc+Re6ta6ELKTaSkZE1lVEZJ89uOVuDbV2HeMztpR5EDe+ogVfdKFiOBfgwX0ScDfMWgTcef6GRvVe7kWTR88zdeE30LivoR4LQj1+3aE+n9fSgECI05RTfzJqrhGhOspcfo= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com; spf=fail smtp.mailfrom=amd.com; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b=ORXxMd8y; arc=fail smtp.client-ip=40.107.209.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=amd.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b="ORXxMd8y" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=KMz20eHcGTmk8RMbzF5hbUXkQfwsVCPGZqHKE+Kj70I4WI8azLR2Yr4fQJpFTuJdwwZPhZcrt0l+82wkNqUDnrijAPu5BrvJ3KBnvi9VkRVOKW0gq3fgIxd+TtAFa8yjQ8ZjoR5tPjuIqpkh3Sj4WpfF6EIsJvL3mkfpZ6Hq5LGVIdY4xj5e+CV8gn2W9UrScZBtPTFKcOhRuWSFOegUjXg8y59mp2Cy4p6gaXxj3feiw6JY7CR8fXHAyoMqQfbEgLXWtWdN3dcw2g0xJj4JMfYB+53J6eN8sPV6HQnSt5yghs0Gce/Rbo6vfeY+YCGpWtdxDINPfmxeuw4tb4/QBw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=Sad8StPYkv8GhYjexobGSCH5AGTYSGaaML/svxLUZ8c=; b=AmkTA1/SC+wk+9ZxprK/AlYKq0/k1m465kx9VMd5PjWjY/ZkzPCo+GkASeZvRG5E86WEauos5eKUyMPXBCDVwfRfDu3k0PK7nOas8xB/spZUAdn9boet9Zex5uEZL3ZEbSTUchUpChh1PV2n0vBCN7IAlFeZXvFWAoStRPI9M7i4iXBJJlux8+NC+7TLdMf5LlyxGsoA1/+N2fNHsYRZm7Gf9E5Vzh2YTotul0/XGWw9TfBcjgwEaoXfVJMMH2RpFYIl6hAtul+XDgeCGgh+gmsx7gQWWJ8LLhPHtIT0UmeDKDbA45oiyILgGfgMjBe6pV1VS7EqxJRXmsMZlXEi7Q== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 165.204.84.17) smtp.rcpttodomain=gentwo.org smtp.mailfrom=amd.com; dmarc=pass (p=quarantine sp=quarantine pct=100) action=none header.from=amd.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amd.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=Sad8StPYkv8GhYjexobGSCH5AGTYSGaaML/svxLUZ8c=; b=ORXxMd8yzMgz77z2Dz4EFW/WsfsT4jleq1nXgSJUL/RvXggPrEw1HzEvkWXmxXMR60Jzd50+LFl3IHiY0haFj0oPk0vWVHre8lyhTlxlUo9RfbW5/4ULwZzUsVjgwFpEx/MHVHp9u6sDaE7pi35wpp53wfNVVX8o9+8qvEhgP9Y= Received: from CH2PR07CA0051.namprd07.prod.outlook.com (2603:10b6:610:5b::25) by CYYPR12MB8703.namprd12.prod.outlook.com (2603:10b6:930:c4::9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.270.14; Thu, 30 Jul 2026 06:27:19 +0000 Received: from CH2PEPF0000014A.namprd02.prod.outlook.com (2603:10b6:610:5b:cafe::5a) by CH2PR07CA0051.outlook.office365.com (2603:10b6:610:5b::25) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.270.13 via Frontend Transport; Thu, 30 Jul 2026 06:27:19 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 165.204.84.17) smtp.mailfrom=amd.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=amd.com; Received-SPF: Pass (protection.outlook.com: domain of amd.com designates 165.204.84.17 as permitted sender) receiver=protection.outlook.com; client-ip=165.204.84.17; helo=satlexmb08.amd.com; pr=C Received: from satlexmb08.amd.com (165.204.84.17) by CH2PEPF0000014A.mail.protection.outlook.com (10.167.244.107) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.292.8 via Frontend Transport; Thu, 30 Jul 2026 06:27:19 +0000 Received: from Satlexmb09.amd.com (10.181.42.218) by satlexmb08.amd.com (10.181.42.217) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.41; Thu, 30 Jul 2026 01:27:19 -0500 Received: from satlexmb07.amd.com (10.181.42.216) by satlexmb09.amd.com (10.181.42.218) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.41; Wed, 29 Jul 2026 23:27:18 -0700 Received: from [10.136.32.30] (10.180.168.240) by satlexmb07.amd.com (10.181.42.216) with Microsoft SMTP Server id 15.2.2562.41 via Frontend Transport; Thu, 30 Jul 2026 01:27:14 -0500 Message-ID: Date: Thu, 30 Jul 2026 11:57:08 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3] sched/fair: Prefer waker CPU for non-SMT reciprocal sync wakeups To: "Shubhang Kaushik (Ampere)" , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Christian Loehle , Madadi Vineeth Reddy CC: "Christoph Lameter (Ampere)" , Shubhang Kaushik , References: <20260727-b4-sched-sync-wakeup-v3-1-90cf481dbd85@gentwo.org> Content-Language: en-US From: K Prateek Nayak In-Reply-To: <20260727-b4-sched-sync-wakeup-v3-1-90cf481dbd85@gentwo.org> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 7bit X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CH2PEPF0000014A:EE_|CYYPR12MB8703:EE_ X-MS-Office365-Filtering-Correlation-Id: 786005a5-2104-4cca-72d2-08deee039b9c X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|82310400026|36860700016|23010399003|1800799024|921020|6133799003|56012099006|11063799006|10067099003|18002099003|22082099003; X-Microsoft-Antispam-Message-Info: Y0avZZnCfeeQM2dDr6HICmKCFoVy6OId10zkRY1mxehw/8TMjib99JdTm4yUpmxI3BqFAISnzPKOPBg1Yn9b2EochQkTb+J2b7nbmIvQ1OKYCQCAUh6TRer991fgW6Qx8PHyhh1YkHgvcnZA2Yc5UzezcRQSi2Cests2UR9CUU8M86tvJygUdLYFj+ZBaEatIk1wj+h7ebe0wpwQ3ei5CBRlndNRmKJ7nDonzqfB24t703J1a/7ChuFzZjIoYczN/Zm0vKs9h9nDCNuDA0lQB1vJ4QGU1m5rRu4vPUYHi6lwGTjxu3C+2pHe7YTZUFUVMvhZaiIVXACpk3JpfM+BTWUTKq15X7iABW1y3YgGHWC0pBcKJon10zW5vgCWZZZi6u1zqUYiMjGtE/yw/DyoBxemaPdo6urcS6IfkyLIu8/6yaUMexssi4fZz5s13xi51FQ1epwYsxuyUZAXP1c3IuRXKUScozU/Lt/voxc+t7USdPAGwGINBSBxhoCMVZS09NWFfCiXLU1BHvaLKgbp44k2OhiQr8SqzY47cqkXuziTFsfnRQTvV9qhGPVIf+/JhkMFQbDVAJkUkKcbf12kz7UXdO8o2j5uHh8y5OVnYASajMWb+hXmoHkx/r68Gr7rn52s7dS7QOXW181Tw3uEdx1/xUlTvZWcHcZAqGEZt8Y1K1rJ0XYjjpLQWDzGoagLeOu/rCpOGYlmH/b3Pl5PTr+TWb05M+0G/sMW0BwjrShIfYy2zoHpk05k07gzogqI X-Forefront-Antispam-Report: CIP:165.204.84.17;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:satlexmb08.amd.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(376014)(7416014)(82310400026)(36860700016)(23010399003)(1800799024)(921020)(6133799003)(56012099006)(11063799006)(10067099003)(18002099003)(22082099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: BnuD3HxWL1Lt6sy1ysjDJEYTuF04aXhJrK1pxLpozXPyXwZSiyN27Ydp/u5+BAAvB0l0qpwtJ4bmKQbjwK9oqmPweOjl22BWve6zCnF3z/MDmNMZfb1QzesW3HpQ+VfqznRZ6wasXhTvKfh5ZFyxPx+AcyzMQ7JeS/FvATJqDYTdKTDrJxefth9m0QjAfsjYsscG+qTDGAUkY4ubhDVuqvL2bJvmVuFeqAa1kii4PFIzvVO6z9e0iLpQsZUD0rSeh7qAmCT+t4YnBLZqqo+zoqu0pU+ILTOxnzGTd18YOVuFo2P/X9yqEF8a2y5K04weoPBhId7h7dtAZhqA+1jAhH+zf62vLCrO2DTYo/vqrXHq9dw7mJh8Q/wWr6ISJhbptFLtj3z3+ubi0k+uGM1n7KsQwK6odSPK3AMx9I51J8NP5+lMkLr5/8JhEGCsmdjl X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 30 Jul 2026 06:27:19.5693 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 786005a5-2104-4cca-72d2-08deee039b9c X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=3dd8961f-e488-4e60-8e11-a82d994e183d;Ip=[165.204.84.17];Helo=[satlexmb08.amd.com] X-MS-Exchange-CrossTenant-AuthSource: CH2PEPF0000014A.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CYYPR12MB8703 Hello Shubhang, On 7/28/2026 5:28 AM, Shubhang Kaushik (Ampere) wrote: > Pipe-style ping-pong workloads can be dominated by handoff cost. In > such cases, placing the wakee on an idle CPU can be slower than keeping > the pair on the same runqueue. > > Use the existing last_wakee and wake_wide() state to identify narrow > reciprocal WF_SYNC wakeups: > > A wakes B > B wakes A > A wakes B > ... > > When the wake-affine domain allows SD_WAKE_AFFINE, prefer the waker CPU > for these narrow reciprocal handoffs on non-SMT systems. Do so only when > the waker CPU has no other runnable fair task and the wakee fits there on > asymmetric-capacity systems. > > SMT systems, and wakeups that do not match this pattern, continue through > the existing wake_affine() and select_idle_sibling() path. > > Signed-off-by: Shubhang Kaushik (Ampere) > --- > Tested on 80-core non-SMT Ampere Altra: perf bench sched pipe -l 1000000 > improved by about 30%, averaged over 40 runs. Hackbench, schbench and > SPECjBB showed no material regression. > > Baseline: v7.2-rc5 > --- > Changes in v3: > - Limit the direct waker-CPU preference to !sched_smt_active(); SMT > systems continue through the existing wake_affine() and > select_idle_sibling() path. Building on top of Chris' suggestion on v2 for systems with SMT, we can push that check further down into select_idle_sibling() and can take a call at the point where we know what test_idle_core() returns. This is what I tried out on top of tip:sched/core: (Lightly tested on a SMT-2 system) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index df8c9c2c7918..5821cbd930ae 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1301,7 +1301,6 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) #include "pelt.h" -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); static unsigned long task_h_load(struct task_struct *p); static unsigned long capacity_of(int cpu); @@ -8636,7 +8635,7 @@ static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpu /* * Scan the local SMT mask for idle CPUs. */ -static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int target) +static int select_idle_smt(struct task_struct *p, struct root_domain *rd, int target) { int cpu; @@ -8644,10 +8643,13 @@ static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int t if (cpu == target) continue; /* - * Check if the CPU is in the LLC scheduling domain of @target. - * Due to isolcpus, there is no guarantee that all the siblings are in the domain. + * Check if the CPU is in the scheduling domain of @target. + * Due to isolcpus, there is no guarantee that all the + * siblings are in the domain. */ - if (!cpumask_test_cpu(cpu, sched_domain_span(sd))) + if (!cpumask_test_cpu(cpu, rd->span)) + continue; + if (sched_asym_cpucap_active() && !task_fits_cpu(p, cpu)) continue; if (choose_idle_cpu(cpu, p)) return cpu; @@ -8928,12 +8930,12 @@ static inline bool asym_fits_cpu(unsigned long util, /* * Try and locate an idle core/thread in the LLC cache domain. */ -static int select_idle_sibling(struct task_struct *p, int prev, int target) +static int select_idle_sibling(struct task_struct *p, int prev, int target, int sync) { bool has_idle_core = false; struct sched_domain *sd; unsigned long task_util, util_min, util_max; - int i, recent_used_cpu, prev_aff = -1; + int i, this_cpu, recent_used_cpu, prev_aff = -1; /* * On asymmetric system, update task utilization because we will check @@ -8977,9 +8979,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) * essentially a sync wakeup. An obvious example of this * pattern is IO completions. */ + this_cpu = smp_processor_id(); if (is_per_cpu_kthread(current) && in_task() && - prev == smp_processor_id() && + prev == this_cpu && this_rq()->nr_running <= 1 && asym_fits_cpu(task_util, util_min, util_max, prev)) { return prev; @@ -9003,6 +9006,32 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) recent_used_cpu = -1; } + has_idle_core = sched_smt_active() && test_idle_cores(target); + + if (!has_idle_core) { + struct rq *target_rq = cpu_rq(target); + + /* Prefer an idle thread on same core where data is hot. */ + if (sched_smt_active() && cpus_share_cache(prev, target)) { + i = select_idle_smt(p, target_rq->rd, prev); + if ((unsigned int)i < nr_cpumask_bits) + return i; + } + + /* + * Tasks are likely a sync wakeup pair that passed WA_IDLE. + * Prefer to temporarily stack them on the same CPU since the + * waker is likely to go away soon and there are no idle cores. + */ + if (sync && + in_task() && + target == this_cpu && + p->last_wakee == current && + (target_rq->nr_running - cfs_h_nr_delayed(target_rq)) <= 1 && + asym_fits_cpu(task_util, util_min, util_max, target)) + return target; + } + /* * For asymmetric CPU capacity systems, our domain of interest is * sd_asym_cpucapacity rather than sd_llc. @@ -9027,16 +9056,6 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) if (!sd) return target; - if (sched_smt_active()) { - has_idle_core = test_idle_cores(target); - - if (!has_idle_core && cpus_share_cache(prev, target)) { - i = select_idle_smt(p, sd, prev); - if ((unsigned int)i < nr_cpumask_bits) - return i; - } - } - i = select_idle_cpu(p, sd, has_idle_core, target); if ((unsigned)i < nr_cpumask_bits) return i; @@ -9734,7 +9753,7 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) /* Fast path */ if (wake_flags & WF_TTWU) - return select_idle_sibling(p, prev_cpu, new_cpu); + return select_idle_sibling(p, prev_cpu, new_cpu, sync); return new_cpu; } --- I'm currently seeing a ~10% improvement for the workload you mentioned (perf bench sched pipe -l 1000000) on average. I haven't tried anything else yet but would love to know your thoughts. I'm using rq->rd->span to know the CPUs covered by the cpuset instead of sched_domain_span(sd_llc) in select_idle_smt() to make it work for sched_asym_cpucap_active() + sched_smt_active() where some cores may have more than one CPUs and the LLC is defined at core boundary. Basically I wanted to avoid this ugly: sd = rcu_dereference_all(per_cpu((sched_asym_cpucap_active()) ? sd_asym : sd_llc, target)); if (!sd) goto skip; pattern and rq->rd->span seemed just fine since it doesn't need a null check and gives the desired boundary. Could you please check if the improvements still persist on your system with the check pushed down into select_idle_sibling(). Thank you. > - Drop the redundant affinity check; want_affine already verifies the > waker CPU is allowed. > - Use a plain p->last_wakee read instead of READ_ONCE(). > - Rebase and refresh testing on v7.2-rc5. > > Link to v2: https://lore.kernel.org/r/20260722-b4-sched-sync-wakeup-v2-1-f1164560b24b@gentwo.org > > Changes in v2: > - Move the reciprocal handoff preference under the existing > SD_WAKE_AFFINE domain check. > - Drop futex from the changelog motivation. > - Refresh perf bench sched pipe results after rebasing. > > Link to v1: https://lore.kernel.org/r/20260721-b4-sched-sync-wakeup-v1-1-dc94f184e27f@gentwo.org > --- > kernel/sched/fair.c | 25 +++++++++++++++++++++++++ > 1 file changed, 25 insertions(+) > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index d78467ec6ee1343050fcc2794dafb38ade3599e5..e61062d20da772d29da6f5f377a150b4b5128619 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -8794,6 +8794,26 @@ static inline bool asym_fits_cpu(unsigned long util, > return true; > } > > +/* > + * For reciprocal WF_SYNC handoffs, prefer the waker CPU when it has no > + * other runnable fair task. > + */ > +static bool prefer_sync_pair_cpu(struct task_struct *p, int cpu) > +{ > + struct rq *rq = cpu_rq(cpu); > + > + if ((rq->nr_running - cfs_h_nr_delayed(rq)) != 1) > + return false; > + > + if (sched_asym_cpucap_active()) { > + sync_entity_load_avg(&p->se); > + if (!task_fits_cpu(p, cpu)) > + return false; > + } > + > + return true; > +} > + > /* > * Try and locate an idle core/thread in the LLC cache domain. > */ > @@ -9579,6 +9599,11 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) > */ > if (want_affine && (tmp->flags & SD_WAKE_AFFINE) && > cpumask_test_cpu(prev_cpu, sched_domain_span(tmp))) { > + if (sync && !sched_smt_active() && For the record, without !sched_smt_active(), the runtime for "perf bench sched pipe -l 1000000" almost doubles in my case but looks like that condition might overall be good with a bunch of defensive checks on SMT systems too. > + p->last_wakee == current && > + prefer_sync_pair_cpu(p, cpu)) > + return cpu; > + > if (cpu != prev_cpu) > new_cpu = wake_affine(tmp, p, cpu, prev_cpu, sync); > > > --- > base-commit: f5098b6bae761e346ebcd9da7f95622c04733cff > change-id: 20260721-b4-sched-sync-wakeup-04d40cbeb1da > > Best regards, -- Thanks and Regards, Prateek