From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 90DF12749DC for ; Fri, 24 Jul 2026 12:51:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784897495; cv=none; b=Mqbr2VHQrVTb3HBzr8fpKvsMXddh3pmT8F0N3WX/aPxHjdXM4xcBniNQiOcq10PxZQ2e9zpWTaWHg3ksKh+s7j2ULyH7V+zTONVTx8TnT4zhlV7AB2FYGKqtTd+af7Rd9zqaEjb0haeRyV7eOm9douRsNXGZeUPDqvD4bto0Uto= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784897495; c=relaxed/simple; bh=RVmhF6MLBnQ6Zg6H9wBlq8kNzT8SEt1Ck/LIIxuHUyg=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=Ud/3urbZXq27C6Pn3CVNsszNBwwRL3OAn4EM+YKSJibM33IT/LNnQgWI0F3ZXN1l4o9Yj9FjEzheMnAmhEpqDxoFq2N9mBTq0XoG1VyOHIE6b5WdA5735WesHvDKvlCYS8wSzd1ESfEtF/u1kHFrid+DFEfzgFfepGeEovylq1c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=M0WRMZAh; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="M0WRMZAh" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 66OCgeKS1865414; Fri, 24 Jul 2026 12:51:06 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=wCxef8 eWkKajygX88HEPKRsF6D18RwPW9EQH8mhx7yc=; b=M0WRMZAhSdWPA1YyBNIr9p 6DrdpK094h2QYuUkh1SI4zobUWMk9F3D3XDKboNnlWZEYaP8NMBxiuo+GHDUqGIw zWjsdNyJqESk+6N5E1EIcbWWkF2U6mWyXo4WqASRVIM2tyRY/DJmgrmdH5v/WYDp z6qHcnXJcH2YgGRK4u2dlWpfdayqGkAoVN8aQI7v+q54snVeAfAfIZnZciDfSXMx 8xG0Ehn5IHLLq4NQodt98jblah2EUo662+KJPmpWg2LGu1xn4/OjCjDy4dvrWeW0 N/1aPN0LGEHSkPMlL2/j4MYAtwr35f+JAUf9VlD115mBauoaS9q60jCS8Q6P9qDg == Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fg791cquc-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 24 Jul 2026 12:51:05 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 66OCoVFA004261; Fri, 24 Jul 2026 12:51:05 GMT Received: from smtprelay07.dal12v.mail.ibm.com ([172.16.1.9]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fgp1grr24-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 24 Jul 2026 12:51:05 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay07.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 66OCp4NO61669822 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 24 Jul 2026 12:51:04 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id A4CE45805A; Fri, 24 Jul 2026 12:51:04 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 96D4058052; Fri, 24 Jul 2026 12:50:59 +0000 (GMT) Received: from [9.43.38.94] (unknown [9.43.38.94]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 24 Jul 2026 12:50:59 +0000 (GMT) Message-ID: <60a584c5-25ac-4077-a725-a2f9ee74318d@linux.ibm.com> Date: Fri, 24 Jul 2026 18:20:58 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2] sched/fair: Prefer waker CPU for reciprocal sync wakeups To: "Shubhang Kaushik (Ampere)" Cc: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , "Christoph Lameter (Ampere)" , Shubhang Kaushik , linux-kernel@vger.kernel.org, Madadi Vineeth Reddy References: <20260722-b4-sched-sync-wakeup-v2-1-f1164560b24b@gentwo.org> Content-Language: en-US From: Madadi Vineeth Reddy In-Reply-To: <20260722-b4-sched-sync-wakeup-v2-1-f1164560b24b@gentwo.org> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-ORIG-GUID: U9RO8_Gs9sKsTFsYmfC-bKk0YssULoF1 X-Authority-Analysis: v=2.4 cv=V6RNF+ni c=1 sm=1 tr=0 ts=6a635fba cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=IkcTkHD0fZMA:10 a=RAioF0-LDSMA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=VwQbUJbxAAAA:8 a=PuvxfXWCAAAA:8 a=1rQMotphXtXr621WLMQA:9 a=QEXdDO2ut3YA:10 a=uAr15Ul7AJ1q7o2wzYQp:22 X-Proofpoint-Spam-Info: AW1haW4tMjYwNzI0MDExNSBTYWx0ZWRfX+mBiwUfKWXUb Ghj26PlQhcDLAxFppx5LpyBiOcXP9kB2LfwyOEl5OLXHPbtvn0RhGgKAF4Dtl90V87vY6ywjoq3 kohb6Jq798snlTBbir7HR/m2Sp8Vrh4= X-Proofpoint-GUID: VPOkZagq0Lr3qTWUBnOOL_1gOMIK-1jL X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNzI0MDExNSBTYWx0ZWRfX75UnokApgclY VboELlYz4KthvBLjMmauSNA69+JdOydfuptO2o0K3Y8nldn2dxAokxqM5ih0NpplWrjB0ldtkw9 NdH/jYjFI22NLgYISpPW3WEFD//m+AE5v79WTWy49Z7aYFZRjYwPVnTV9sfIFCBINLgmeKH9aXq wWVu6Hrkkz8hPTZv90mtYFInvR1mSIUTMDvVBZLoHLW4se+J65H/ORKDM7hEyPq/vQJ/u5NO5jO 6tQ3pHY8IbRlU3dlmDJFXuhbyXRoo1lMdLmIzMw0gh0NRd7vmFZBTp0ebbpDGFUAMwzqCFqIIv9 ISt0kb2A7L2xL2ZUNaY8MSuDHTZ+jLq0snyvWeBihuH0IkmiEMBCF5AxRWlNRVV+owZ0NFvHQle 3cU300oKbKR16A42WQgw6Wl/9r1eAAketfvb7rwNEvAtakf63fXTaPLDEHbQyBaTvla05v47qPP dhvZmYx8uOJ/L79dzQw== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-07-24_02,2026-07-24_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 malwarescore=0 adultscore=0 bulkscore=0 lowpriorityscore=0 clxscore=1011 spamscore=0 impostorscore=0 phishscore=0 priorityscore=1501 suspectscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2607240115 On 23/07/26 04:20, Shubhang Kaushik (Ampere) wrote: > Pipe-style ping-pong workloads can be dominated by handoff cost. In > such cases, placing the wakee on an idle CPU can be slower than keeping > the pair on the same runqueue. > > Use the existing last_wakee and wake_wide() state to identify narrow > reciprocal WF_SYNC wakeups: > > A wakes B > B wakes A > A wakes B > ... > I've been looking at the same underlying problem from the POWER side, where the sync hint is consumed by wake_affine() and then discarded by select_idle_sibling(). > When the wake-affine domain allows SD_WAKE_AFFINE, prefer the waker CPU > for these narrow reciprocal handoffs. Do so only when the waker CPU has no > other runnable fair task, the wakee is allowed on that CPU, and the wakee > fits there on asymmetric capacity systems. > > Wakeups that do not match this pattern continue through the existing > wake_affine() and select_idle_sibling() path. > > Signed-off-by: Shubhang Kaushik (Ampere) > --- > Tested on 80-core Ampere Altra: perf bench sched pipe -l 1000000 improved > by about 30%, averaged over 20 runs. Hackbench, schbench and SPECjBB > showed no material regression. > > Baseline ~ v7.2-rc4 (mainline origin/master at 248951ddc14d) > --- > Changes in v2: > - Move the reciprocal handoff preference under the existing > SD_WAKE_AFFINE domain check. > - Drop futex from the changelog motivation. > - Refresh perf bench sched pipe results after rebasing. > > Link to v1: https://lore.kernel.org/r/20260721-b4-sched-sync-wakeup-v1-1-dc94f184e27f@gentwo.org > --- > kernel/sched/fair.c | 28 ++++++++++++++++++++++++++++ > 1 file changed, 28 insertions(+) > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index d78467ec6ee1343050fcc2794dafb38ade3599e5..d188e91b85d74dd3ed3c8b99f171bd142e72da5e 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -8794,6 +8794,29 @@ static inline bool asym_fits_cpu(unsigned long util, > return true; > } > > +/* > + * For reciprocal WF_SYNC handoffs, prefer the waker CPU when it has no > + * other runnable fair task. > + */ > +static bool prefer_sync_pair_cpu(struct task_struct *p, int cpu) > +{ > + struct rq *rq = cpu_rq(cpu); > + > + if ((rq->nr_running - cfs_h_nr_delayed(rq)) != 1) > + return false; > + > + if (!cpumask_test_cpu(cpu, p->cpus_ptr)) > + return false; > + > + if (sched_asym_cpucap_active()) { > + sync_entity_load_avg(&p->se); > + if (!task_fits_cpu(p, cpu)) > + return false; > + } > + > + return true; > +} > + > /* > * Try and locate an idle core/thread in the LLC cache domain. > */ > @@ -9579,6 +9602,11 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) > */ > if (want_affine && (tmp->flags & SD_WAKE_AFFINE) && > cpumask_test_cpu(prev_cpu, sched_domain_span(tmp))) { > + if (sync && > + READ_ONCE(p->last_wakee) == current && > + prefer_sync_pair_cpu(p, cpu)) > + return cpu; The difference is SMT. POWER10/11 are SMT8: stacking there puts the pair on one thread while up to seven siblings on the same core sit idle, sharing the same LLC. So instead of returning the waker's CPU, I'm looking at letting the waker's *core* count as idle when the waker's runqueue has a single runnable task, so the wakee lands on a sibling thread. Like below diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 728965851842..8d1a2f62431b 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1233,7 +1233,7 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) #include "pelt.h" -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); +static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu, int want_affine); static unsigned long task_h_load(struct task_struct *p); static unsigned long capacity_of(int cpu); @@ -7840,13 +7840,18 @@ void __update_idle_core(struct rq *rq) * there are no idle cores left in the system; tracked through * sd_llc->shared->has_idle_cores and enabled through update_idle_core() above. */ -static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, int *idle_cpu) +static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, int *idle_cpu, bool want_affine) { bool idle = true; int cpu; + int this_cpu = smp_processor_id(); for_each_cpu(cpu, cpu_smt_mask(core)) { - if (!available_idle_cpu(cpu)) { + struct rq *rq = cpu_rq(cpu); + bool sync_waker = want_affine && cpu == this_cpu && + ((rq->nr_running - cfs_h_nr_delayed(rq)) == 1); + + if (!available_idle_cpu(cpu) && !sync_waker) { idle = false; if (*idle_cpu == -1) { if (choose_sched_idle_rq(cpu_rq(cpu), p) && @@ -7858,6 +7863,7 @@ static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpu } break; } + if (*idle_cpu == -1 && cpumask_test_cpu(cpu, cpus)) *idle_cpu = cpu; } @@ -7920,7 +7926,7 @@ static inline int select_idle_smt(struct task_struct *p, struct sched_domain *sd * comparing the average scan cost (tracked in sd->avg_scan_cost) against the * average idle time for this rq (as found in rq->avg_idle). */ -static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, int target) +static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, int target, int want_affine) { struct cpumask *cpus = this_cpu_cpumask_var_ptr(select_rq_mask); int i, cpu, idle_cpu = -1, nr = INT_MAX; @@ -7953,7 +7959,7 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool continue; if (has_idle_core) { - i = select_idle_core(p, cpu, cpus, &idle_cpu); + i = select_idle_core(p, cpu, cpus, &idle_cpu, want_affine); if ((unsigned int)i < nr_cpumask_bits) return i; } else { @@ -7970,7 +7976,7 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool for_each_cpu_wrap(cpu, cpus, target + 1) { if (has_idle_core) { - i = select_idle_core(p, cpu, cpus, &idle_cpu); + i = select_idle_core(p, cpu, cpus, &idle_cpu, want_affine); if ((unsigned int)i < nr_cpumask_bits) return i; @@ -8060,7 +8066,7 @@ static inline bool asym_fits_cpu(unsigned long util, /* * Try and locate an idle core/thread in the LLC cache domain. */ -static int select_idle_sibling(struct task_struct *p, int prev, int target) +static int select_idle_sibling(struct task_struct *p, int prev, int target, int want_affine) { bool has_idle_core = false; struct sched_domain *sd; @@ -8169,7 +8175,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) } } - i = select_idle_cpu(p, sd, has_idle_core, target); + i = select_idle_cpu(p, sd, has_idle_core, target, want_affine); if ((unsigned)i < nr_cpumask_bits) return i; @@ -8858,8 +8864,10 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag); /* Fast path */ - if (wake_flags & WF_TTWU) - return select_idle_sibling(p, prev_cpu, new_cpu); + if (wake_flags & WF_TTWU) { + want_affine = want_affine && sync && (new_cpu != prev_cpu); + return select_idle_sibling(p, prev_cpu, new_cpu, want_affine); + } return new_cpu; } The cache locality is the same as stacking, but the waker's remaining work after the wakeup still runs in parallel rather than being serialized behind the wakee. I plan to post this patch as part of powerpc sched-domain series. Thanks, Vineeth > + > if (cpu != prev_cpu) > new_cpu = wake_affine(tmp, p, cpu, prev_cpu, sync); > > > --- > base-commit: 248951ddc14de84de3910f9b13f51491a8cd91df > change-id: 20260721-b4-sched-sync-wakeup-04d40cbeb1da > > Best regards,