From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 20EFF237713 for ; Sat, 1 Aug 2026 04:04:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785557069; cv=none; b=gkakMmWuaxFKA5qrH57wVsYsR9RSEsWLi9AMCIgCoFH8f/VpUv7PsIvgA8OwyEzv6nMSpVTQ+ZHSdeABZaEhmzYBlqZVcAJ7ALALm/JYrbHCb4dMK6v7j4Pz0uxfk7qtTHVR/ETPZIjsIngyT24c57my8Ij58euEMsS7tKRbpA8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785557069; c=relaxed/simple; bh=nDv6/xd6/26Yucis0EdeXGvyQW8Zg7RLVqT9om+9mhU=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=AMuvNW3uaQObFVhlxMJfjdKM97BCmlpDKIHyhY56dFkketAT27EFd8DgEFvp2iKGjSS9YuARhx0/vDpzB20hcRLYyRa+tTm+j0VtwMvm++dg9l8irIQq4zmrBejmL6CRvd2ieG0OQEN8Go/elF0FZnAkTrulvqlnMHmhv2Kh9Ow= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=juTa1WSp; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="juTa1WSp" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6712O99p3512862; Sat, 1 Aug 2026 04:03:37 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=NaPzwv jqV//aLIMM7fHg6aAFZrXFRugm7aZLBxbxL2U=; b=juTa1WSpMM/rMw+xllLUfH 0m/BNX9u5aCcufJKCmn1+aB+vyee2a9s0p+E9uiSG3iNuaLS3slgjIvRYzspW3bA 9kzX1bS5QI8R55nBO2PtEoRFnBA4r2j+bKWwspXagjukXmwU30MKUc6A9rI7DdYE erUNOW/RPLGQFtNjY5kyfSG6Im9jukuHHqbgY2ifxrJJ5wz3Q7P6F/cevQYBCKXv e9lEt+LH+k9sPTua6w8//TVd8WYokjon0N36bqLUnF5zuiEdb6bAxxjPY6jkjuJT Pam1sjze9ioUMl6GMa86CXrM/R/MKeJjxf0yx+iTUJpQjWsADSJguv1b207hhJxQ == Received: from ppma21.wdc07v.mail.ibm.com (5b.69.3da9.ip4.static.sl-reverse.com [169.61.105.91]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs67h8f4w-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Sat, 01 Aug 2026 04:03:36 +0000 (GMT) Received: from pps.filterd (ppma21.wdc07v.mail.ibm.com [127.0.0.1]) by ppma21.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6713fGEd009923; Sat, 1 Aug 2026 04:03:35 GMT Received: from smtprelay03.wdc07v.mail.ibm.com ([172.16.1.70]) by ppma21.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fn8fkjrg1-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Sat, 01 Aug 2026 04:03:35 +0000 (GMT) Received: from smtpav04.wdc07v.mail.ibm.com (smtpav04.wdc07v.mail.ibm.com [10.39.53.231]) by smtprelay03.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67142wpq9831028 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Sat, 1 Aug 2026 04:02:58 GMT Received: from smtpav04.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 353E058050; Sat, 1 Aug 2026 04:03:35 +0000 (GMT) Received: from smtpav04.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 36ABE58045; Sat, 1 Aug 2026 04:03:29 +0000 (GMT) Received: from [9.43.87.78] (unknown [9.43.87.78]) by smtpav04.wdc07v.mail.ibm.com (Postfix) with ESMTP; Sat, 1 Aug 2026 04:03:28 +0000 (GMT) Message-ID: Date: Sat, 1 Aug 2026 09:33:27 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3] sched/fair: Prefer waker CPU for non-SMT reciprocal sync wakeups To: K Prateek Nayak Cc: "Shubhang Kaushik (Ampere)" , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Christian Loehle , "Christoph Lameter (Ampere)" , Shubhang Kaushik , linux-kernel@vger.kernel.org, Madadi Vineeth Reddy References: <20260727-b4-sched-sync-wakeup-v3-1-90cf481dbd85@gentwo.org> Content-Language: en-US From: Madadi Vineeth Reddy In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwODAxMDAyMiBTYWx0ZWRfXyV6/eVVRgU7I xXdbl7n43Yo+7ZRo/9m/IGSOHoxagX3FuDLku3kr3UG4COmJOc5ueV0EH/0IY5cjaEGFF3qYmxh PUUegVbQunCQ4SblEu17KylZVUxh4mU= X-Authority-Analysis: v=2.4 cv=I7VVgtgg c=1 sm=1 tr=0 ts=6a6d7018 cx=c_pps a=GFwsV6G8L6GxiO2Y/PsHdQ==:117 a=GFwsV6G8L6GxiO2Y/PsHdQ==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=PuvxfXWCAAAA:8 a=otZkiQFeCS44dqKxZt8A:9 a=QEXdDO2ut3YA:10 a=uAr15Ul7AJ1q7o2wzYQp:22 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODAxMDAyMiBTYWx0ZWRfX6hyxp13g0iTP m0TqROn3U+lsr2u/dBgL3TQ0PjScy195r8DiioReh849ncCTgrre5IhN/jXylLbJUxebUZwA/r9 oqosh0c/VHNEe/BdG7WoC2yxFr1F7K63tHtTYL47wrkVSF2xERb865V4FhBDUT/JbBHW9qHahS8 V7Cj1G1QGkGkpwvfSGWzHqUXxqSW9xlMmSaFC62uq3YmZDuAwEcS50MRwaaZGG5ZMC+fhCEBRG4 pmCxmCfd5m4AF+6+cCCArSDzDzwRark6gP0SM9aWNxN5q7G1j4fr6tRDbUsjF26ux8k8D3W1e0a 3zA3V3UcS20tU9SyA3Bsqb+onKY3fh3+IBUSBN1yR9T0TC0VQ0v0VwJn0kLF0I/sZexPZSweLcu jRz2oEmpPZUiye1V3aX/sFaszmx6YBkyKPOjycn7Ga9ILdFyYVtH/gpQ7OfKyceIkw/WVJ+K1Gm CUh98yicQc+uWE1kgzw== X-Proofpoint-ORIG-GUID: jBmJjjRdorbBE5JpJ1FYOvWGVM9lkiCD X-Proofpoint-GUID: IHyNKLUUwnngTCD_be6q6DPIlQbFA8Pt X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-07-31_07,2026-07-30_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 impostorscore=0 clxscore=1011 priorityscore=1501 suspectscore=0 malwarescore=0 adultscore=0 lowpriorityscore=0 phishscore=0 spamscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608010022 Hi Prateek, On 30/07/26 11:57, K Prateek Nayak wrote: > Hello Shubhang, > > On 7/28/2026 5:28 AM, Shubhang Kaushik (Ampere) wrote: >> Pipe-style ping-pong workloads can be dominated by handoff cost. In >> such cases, placing the wakee on an idle CPU can be slower than keeping >> the pair on the same runqueue. >> >> Use the existing last_wakee and wake_wide() state to identify narrow >> reciprocal WF_SYNC wakeups: >> >> A wakes B >> B wakes A >> A wakes B >> ... >> >> When the wake-affine domain allows SD_WAKE_AFFINE, prefer the waker CPU >> for these narrow reciprocal handoffs on non-SMT systems. Do so only when >> the waker CPU has no other runnable fair task and the wakee fits there on >> asymmetric-capacity systems. >> >> SMT systems, and wakeups that do not match this pattern, continue through >> the existing wake_affine() and select_idle_sibling() path. >> >> Signed-off-by: Shubhang Kaushik (Ampere) >> --- >> Tested on 80-core non-SMT Ampere Altra: perf bench sched pipe -l 1000000 >> improved by about 30%, averaged over 40 runs. Hackbench, schbench and >> SPECjBB showed no material regression. >> >> Baseline: v7.2-rc5 >> --- >> Changes in v3: >> - Limit the direct waker-CPU preference to !sched_smt_active(); SMT >> systems continue through the existing wake_affine() and >> select_idle_sibling() path. > > Building on top of Chris' suggestion on v2 for systems with SMT, we can > push that check further down into select_idle_sibling() and can take a > call at the point where we know what test_idle_core() returns. > > This is what I tried out on top of tip:sched/core: > > (Lightly tested on a SMT-2 system) > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index df8c9c2c7918..5821cbd930ae 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -1301,7 +1301,6 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) > > #include "pelt.h" > > -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); > static unsigned long task_h_load(struct task_struct *p); > static unsigned long capacity_of(int cpu); > > @@ -8636,7 +8635,7 @@ static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpu > /* > * Scan the local SMT mask for idle CPUs. > */ > -static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int target) > +static int select_idle_smt(struct task_struct *p, struct root_domain *rd, int target) > { > int cpu; > > @@ -8644,10 +8643,13 @@ static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int t > if (cpu == target) > continue; > /* > - * Check if the CPU is in the LLC scheduling domain of @target. > - * Due to isolcpus, there is no guarantee that all the siblings are in the domain. > + * Check if the CPU is in the scheduling domain of @target. > + * Due to isolcpus, there is no guarantee that all the > + * siblings are in the domain. > */ > - if (!cpumask_test_cpu(cpu, sched_domain_span(sd))) > + if (!cpumask_test_cpu(cpu, rd->span)) > + continue; > + if (sched_asym_cpucap_active() && !task_fits_cpu(p, cpu)) > continue; > if (choose_idle_cpu(cpu, p)) > return cpu; > @@ -8928,12 +8930,12 @@ static inline bool asym_fits_cpu(unsigned long util, > /* > * Try and locate an idle core/thread in the LLC cache domain. > */ > -static int select_idle_sibling(struct task_struct *p, int prev, int target) > +static int select_idle_sibling(struct task_struct *p, int prev, int target, int sync) > { > bool has_idle_core = false; > struct sched_domain *sd; > unsigned long task_util, util_min, util_max; > - int i, recent_used_cpu, prev_aff = -1; > + int i, this_cpu, recent_used_cpu, prev_aff = -1; > > /* > * On asymmetric system, update task utilization because we will check > @@ -8977,9 +8979,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) > * essentially a sync wakeup. An obvious example of this > * pattern is IO completions. > */ > + this_cpu = smp_processor_id(); > if (is_per_cpu_kthread(current) && > in_task() && > - prev == smp_processor_id() && > + prev == this_cpu && > this_rq()->nr_running <= 1 && > asym_fits_cpu(task_util, util_min, util_max, prev)) { > return prev; > @@ -9003,6 +9006,32 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) > recent_used_cpu = -1; > } > > + has_idle_core = sched_smt_active() && test_idle_cores(target); > + > + if (!has_idle_core) { > + struct rq *target_rq = cpu_rq(target); > + > + /* Prefer an idle thread on same core where data is hot. */ > + if (sched_smt_active() && cpus_share_cache(prev, target)) { > + i = select_idle_smt(p, target_rq->rd, prev); > + if ((unsigned int)i < nr_cpumask_bits) > + return i; > + } > + > + /* > + * Tasks are likely a sync wakeup pair that passed WA_IDLE. > + * Prefer to temporarily stack them on the same CPU since the > + * waker is likely to go away soon and there are no idle cores. > + */ > + if (sync && > + in_task() && > + target == this_cpu && > + p->last_wakee == current && > + (target_rq->nr_running - cfs_h_nr_delayed(target_rq)) <= 1 && > + asym_fits_cpu(task_util, util_min, util_max, target)) > + return target; > + } > + > /* > * For asymmetric CPU capacity systems, our domain of interest is > * sd_asym_cpucapacity rather than sd_llc. > @@ -9027,16 +9056,6 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) > if (!sd) > return target; > > - if (sched_smt_active()) { > - has_idle_core = test_idle_cores(target); > - > - if (!has_idle_core && cpus_share_cache(prev, target)) { > - i = select_idle_smt(p, sd, prev); > - if ((unsigned int)i < nr_cpumask_bits) > - return i; > - } > - } > - > i = select_idle_cpu(p, sd, has_idle_core, target); > if ((unsigned)i < nr_cpumask_bits) > return i; > @@ -9734,7 +9753,7 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) > > /* Fast path */ > if (wake_flags & WF_TTWU) > - return select_idle_sibling(p, prev_cpu, new_cpu); > + return select_idle_sibling(p, prev_cpu, new_cpu, sync); > > return new_cpu; > } I have been looking at the same problem from the SMT side which I mentioned in v2 of this patch: https://lore.kernel.org/all/60a584c5-25ac-4077-a725-a2f9ee74318d@linux.ibm.com/ Posted a patch for it today: https://lore.kernel.org/lkml/20260801035532.260625-1-vineethr@linux.ibm.com/ It lets the waker's CPU count as idle inside select_idle_core(), so the waker's core stays an idle-core candidate and the wakee lands on one of its sibling threads. On a sync wakeup the waker's core already holds the data, so this keeps the cache sharing. Thanks, Vineeth > --- > > I'm currently seeing a ~10% improvement for the workload you mentioned > (perf bench sched pipe -l 1000000) on average. I haven't tried anything > else yet but would love to know your thoughts. > > I'm using rq->rd->span to know the CPUs covered by the cpuset instead of > sched_domain_span(sd_llc) in select_idle_smt() to make it work for > sched_asym_cpucap_active() + sched_smt_active() where some cores may > have more than one CPUs and the LLC is defined at core boundary. > > Basically I wanted to avoid this ugly: > > sd = rcu_dereference_all(per_cpu((sched_asym_cpucap_active()) ? sd_asym : sd_llc, target)); > > if (!sd) > goto skip; > > pattern and rq->rd->span seemed just fine since it doesn't need a null > check and gives the desired boundary. > > Could you please check if the improvements still persist on your system > with the check pushed down into select_idle_sibling(). Thank you. > > >> - Drop the redundant affinity check; want_affine already verifies the >> waker CPU is allowed. >> - Use a plain p->last_wakee read instead of READ_ONCE(). >> - Rebase and refresh testing on v7.2-rc5. >> >> Link to v2: https://lore.kernel.org/r/20260722-b4-sched-sync-wakeup-v2-1-f1164560b24b@gentwo.org >> >> Changes in v2: >> - Move the reciprocal handoff preference under the existing >> SD_WAKE_AFFINE domain check. >> - Drop futex from the changelog motivation. >> - Refresh perf bench sched pipe results after rebasing. >> >> Link to v1: https://lore.kernel.org/r/20260721-b4-sched-sync-wakeup-v1-1-dc94f184e27f@gentwo.org >> --- >> kernel/sched/fair.c | 25 +++++++++++++++++++++++++ >> 1 file changed, 25 insertions(+) >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index d78467ec6ee1343050fcc2794dafb38ade3599e5..e61062d20da772d29da6f5f377a150b4b5128619 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -8794,6 +8794,26 @@ static inline bool asym_fits_cpu(unsigned long util, >> return true; >> } >> >> +/* >> + * For reciprocal WF_SYNC handoffs, prefer the waker CPU when it has no >> + * other runnable fair task. >> + */ >> +static bool prefer_sync_pair_cpu(struct task_struct *p, int cpu) >> +{ >> + struct rq *rq = cpu_rq(cpu); >> + >> + if ((rq->nr_running - cfs_h_nr_delayed(rq)) != 1) >> + return false; >> + >> + if (sched_asym_cpucap_active()) { >> + sync_entity_load_avg(&p->se); >> + if (!task_fits_cpu(p, cpu)) >> + return false; >> + } >> + >> + return true; >> +} >> + >> /* >> * Try and locate an idle core/thread in the LLC cache domain. >> */ >> @@ -9579,6 +9599,11 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) >> */ >> if (want_affine && (tmp->flags & SD_WAKE_AFFINE) && >> cpumask_test_cpu(prev_cpu, sched_domain_span(tmp))) { >> + if (sync && !sched_smt_active() && > > For the record, without !sched_smt_active(), the runtime for > "perf bench sched pipe -l 1000000" almost doubles in my case but > looks like that condition might overall be good with a bunch of > defensive checks on SMT systems too. > >> + p->last_wakee == current && >> + prefer_sync_pair_cpu(p, cpu)) >> + return cpu; >> + >> if (cpu != prev_cpu) >> new_cpu = wake_affine(tmp, p, cpu, prev_cpu, sync); >> >> >> --- >> base-commit: f5098b6bae761e346ebcd9da7f95622c04733cff >> change-id: 20260721-b4-sched-sync-wakeup-04d40cbeb1da >> >> Best regards, >