From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3BF3F46A5E6 for ; Wed, 29 Jul 2026 10:58:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785322716; cv=none; b=CrNWv+AhkaL+REJ2Wx2QR6YqzSF0Ozf8Jb3plgQTo4Z9WR/Ac6rpSmzyjoKbHPuNjk+56PPqYtc4ykKa4pvZFDrCn8ZE6tOzDm0l94tcCye/i7yjUZe+6u3prFGTHUc9lnycGFbK/5wQ3JpvnaXEAVG2AyZFPqi8H2oTABm/bcY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785322716; c=relaxed/simple; bh=CAdFDzey4cjPipXK3DQ8P3t8bGkhcF3+LUhaEy75l0A=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=qJot1x4PcRotgj2uMihWPMqTs8h4NMZcWNBPDDjuHMIC+WP8/BaEer0WttnOdgEseC3Ng43GicwH36HqkgBPeyiNXnF4H6Sy4cwQGFHTJ+c9PUaYP3VM+AJC7ZZ0lLdA6Qztfv7GsLI397zgkoeF6hgARkr9ksxJKPFiFOJSxko= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=Zs3cZJEL; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="Zs3cZJEL" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 66T7m0Aa3508489; Wed, 29 Jul 2026 10:57:06 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=Stv0WO aDE4IR45QDs8GMjRSElhfSi3PtQw8mUHO3jPs=; b=Zs3cZJELiMCQWx7hloGrJL 7D79aFL65X831iBUpn8sAuC3ODfaUIILsKPZfuHXqYgbhutC+57HVl7A1O39+OGJ HTdvyUZb5kRSc9IAxLQHlSIzdAVD9JNpEfzez999NA9GBU4QR31NfS3hW4rrhRQ2 tqEx8wiWV/8KL59ue1/37sQdskT2y0tuofcnrdgOCz5SSKSMhfZepIYhmtHWsl3S laNAyvmsiHxAV373BGr7tvf1fjsLz4Bd4vq8Qg7FFg5c5u77M29v6a0Kh06dz+8q hIrq8rujKjTz0UGk7/X/z5BINbyaOs8sJuf/GsQZFKKSzCZM6an7ysUqQ11ptVgw == Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fmv0nseyw-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 29 Jul 2026 10:57:05 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 66TAuKRJ030935; Wed, 29 Jul 2026 10:57:05 GMT Received: from smtprelay02.fra02v.mail.ibm.com ([9.218.2.226]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fn7uw6f92-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Wed, 29 Jul 2026 10:57:05 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay02.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 66TAv34d50397596 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Wed, 29 Jul 2026 10:57:03 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 1CE5B20043; Wed, 29 Jul 2026 10:57:03 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id E528420040; Wed, 29 Jul 2026 10:56:59 +0000 (GMT) Received: from [9.39.24.89] (unknown [9.39.24.89]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Wed, 29 Jul 2026 10:56:59 +0000 (GMT) Message-ID: Date: Wed, 29 Jul 2026 16:26:58 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] sched/fair: Prefer fully idle cores for NOHZ balancing To: K Prateek Nayak , Andrea Righi Cc: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Christian Loehle , Phil Auld , linux-kernel@vger.kernel.org References: <20260728214442.1648483-1-arighi@nvidia.com> Content-Language: en-US From: Shrikanth Hegde In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwNzI5MDA4OSBTYWx0ZWRfXzCFl0eflr8v/ KH+vezk+Gc8M5J23CYUfCL4ft7Ig0btJBiSTQaq9aMXWloC5KR9lgMiUckpB93QECppPfcZzq+t 8EE/pA/wxsjjMxwo5sSBrQHnf57mlC4= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNzI5MDA4OSBTYWx0ZWRfX9bQvkePY6sz4 r4QxnsB1ROAADkhLL2kfxPE7EAVyoKQ6n+Pl6ZEgLlWgJCs6kvK7UDGqjSFU4woNGfQ0WffTCaq xnFc28j2YwAMvGB4sduKH9YBa7g4gpl7cSTqjaU1tHqjG0DnV9XZhYiLaR69w0z/nIbSS8G7cKx MKiwfu7uNLqY/Pu5CQZXWFona+QmvoYLJMpt8YgsiCPQ/+hWjS5TgGLVzKADabx2YxakXx1ujoH ad4bor7nq2QP2iiLWlDlk6FtsM2EjCk+pHDt/RidcaRkphEEhlTz0Xnhec1uqa4YYDhwTtw8JD0 OKZ1Pos1E1x1/c4IGkVnFCS8H8Xx8iNozvkp9kYVzv5t7ImvJPkgfHSeOFfgUF6LXm0AULl5U8L UyjOfKTY+OegojIvcjLknPkmSgJSS61nK/pKzG+WmJkpNoEBcq99AAUAME1frDeOlMX89/Mn2rr XLy2H4gFxjhHpw4eqMQ== X-Authority-Analysis: v=2.4 cv=b5WCJNGx c=1 sm=1 tr=0 ts=6a69dc82 cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=IkcTkHD0fZMA:10 a=RAioF0-LDSMA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=VwQbUJbxAAAA:8 a=WsHKUha7AAAA:8 a=ektTJZ84M6uTkYcGK2sA:9 a=QEXdDO2ut3YA:10 a=H4LAKuo8djmI0KOkngUh:22 X-Proofpoint-GUID: ASwGXPaUFnCai5UE11nJRX0I7q0_LU6y X-Proofpoint-ORIG-GUID: BIe8DelQ0hwpNmU23zeyIRz0cw5a3hcW X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-07-29_03,2026-07-28_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 priorityscore=1501 spamscore=0 adultscore=0 malwarescore=0 impostorscore=0 bulkscore=0 phishscore=0 suspectscore=0 clxscore=1015 lowpriorityscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2607290089 On 7/29/26 4:09 PM, K Prateek Nayak wrote: > Hello Andrea, > > On 7/29/2026 3:05 PM, Andrea Righi wrote: >> Hi Prateek, >> >> On Wed, Jul 29, 2026 at 01:48:14PM +0530, K Prateek Nayak wrote: >>> Hello Andrea, >>> >>> On 7/29/2026 3:14 AM, Andrea Righi wrote: >>>> @@ -13974,19 +13974,32 @@ static inline int find_new_ilb(void) >>>> if (ilb_cpu == this_cpu) >>>> continue; >>>> >>>> - if (idle_cpu(ilb_cpu)) >>>> + if (!idle_cpu(ilb_cpu)) >>>> + continue; >>>> + >>>> + /* >>>> + * Running the idle load balancer on an idle sibling of a busy >>>> + * SMT core can reduce the capacity available to its sibling. Prefer >>>> + * a CPU whose entire core is idle, but retain the first idle CPU as >>>> + * a fallback so idle balancing can still make progress when no fully >>>> + * idle core exists. >>>> + */ >>>> + if (!sched_smt_active() || is_core_idle(ilb_cpu)) >>> >>> nit. >>> >>> is_core_idle() here would iterate all siblings and on systems with SMT-4 >>> and SMT-8, that overhead is apparently visible when one thread per core >>> is occupied based on past optimizations like f8858d96061f ("sched/fair: >>> Optimize should_we_balance() for large SMT systems"). >>> >>> Copying the same approach from that optimization, can we do: >> >> Good point. Without removing the remaining siblings, we may call is_core_idle() >> repeatedly for the same partially busy SMT core. I'll repeat my tests on my Vera >> system and incorporate your suggestion in v2 if I don't see any regression. > > Thank you! From my experience, it is not very visible on SMT-2 but > Shrikanth can vouch for the overheads on SMT-4, SMT-8 systems. > I had the same thought when I saw is_core_idle there. I will try to give the patch a try. >> >>> >>> (Only build tested) >>> >>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >>> index d78467ec6ee1..814bce21ccf1 100644 >>> --- a/kernel/sched/fair.c >>> +++ b/kernel/sched/fair.c >>> @@ -13849,21 +13849,35 @@ static inline int on_null_domain(struct rq *rq) >>> */ >>> static inline int find_new_ilb(void) >>> { >>> + struct cpumask *ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask); >>> int this_cpu = smp_processor_id(); >>> - const struct cpumask *hk_mask; >>> - int ilb_cpu; >>> + int ilb_cpu, fallback = -1; >>> >>> - hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); >>> + cpumask_and(ilb_cpus, nohz.idle_cpus_mask, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); >>> >>> - for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) { >>> + for_each_cpu(ilb_cpu, ilb_cpus) { >>> if (ilb_cpu == this_cpu) >>> continue; >>> >>> - if (idle_cpu(ilb_cpu)) >>> - return ilb_cpu; >>> + if (!idle_cpu(ilb_cpu)) >>> + continue; >>> + >>> + if (sched_smt_active() && !is_core_idle(ilb_cpu)) { >>> + if (fallback == -1) >>> + fallback = ilb_cpu; >>> + /* >>> + * If the core is not idle, and first SMT sibling which is >>> + * idle has been found, then its not needed to check other >>> + * SMT siblings for idleness: >>> + */ >>> + cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu)); >>> + continue; >>> + } >>> + >>> + return ilb_cpu; >>> } >>> >>> - return -1; >>> + return fallback; >>> } >>> >>> /* >>> --- >>> >>> It is safe to use "select_rq_mask" here since this is the tick handler >>> trying to find an ilb_cpu and "select_rq_mask" is only used in contexts >>> with IRQs disabled. It can probably be renamed to suggest that it is >>> safe to be used in any IRQ disabled context as a temporary mask. >>> >>> Thoughts? >> >> Agreed. We can also add lockdep_assert_irqs_disabled() to find_new_ilb() to >> better document and verify the condition that makes reusing select_rq_mask safe. > > That works too but it just looks a bot odd to have the selectrq_mask in > a load balancing function. > >> >> Speaking of that, instead of renaming it, would it be better to provide a helper >> to access select_rq_mask with lockdep_assert_irqs_disabled()? > > I'll defer to Peter on that :-) > > He had previously suggested renaming it when there were discussions to > reuse it here > https://lore.kernel.org/lkml/20260320114312.GB3558198@noisy.programming.kicks-ass.net/ >