From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B23F428470 for ; Fri, 31 Jul 2026 12:54:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785502471; cv=none; b=TtaRegzefzuUvRPWrG3kicAxqiuuzj1DIYlnLHkhkpJT9KHWEOKp64CV/zpXAIcZNgLMt7nhAHOvGqMoZfHpMwrPrvS3h05/UKhDDV2j+PgcQZIZKDs3XiV71s9TaH2s5N5JM4B7lRAyKEFnAPnBsAtqeT91ipEGsXuHVCIh/Yg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785502471; c=relaxed/simple; bh=VihO/NLyIJz8HIOKSvppdEIPjRiVzHZEm9ZV68XSFII=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=t67vqMWSJEVGYCu8txBFL6yk8puwSUShvJ153iNupWx7gn2lzrsI1JaohRF2YEtCmUqjUi3wuFXWUeIBGSLpkAm6gxzk7ggmIompQTQqzDToDzBbNrxnKOpXUT2GE+OHfYwI2P2u9KvLwlpbTWqR7W+9khbBynAukrZ/bv/hZck= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=Bz4pUXy9; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="Bz4pUXy9" Received: from pps.filterd (m0360083.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 66VBlY021730320; Fri, 31 Jul 2026 12:54:00 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=Ii0E1X EppV2kTyu/LVpe3CP305nv69RRKNWoVpJcHYQ=; b=Bz4pUXy9Zyxr1kr/lKmXqM aFLjdo6qU0iKIdDU6/tqhCEz+UJT12z5azFFgV0nTNFMK2OrNpS0r/+d107o7DqL qTraUjGDQhAG7fz8bzdTBLUlZ7h7AtgM6D2CFotaXTJz0hs5KlY2jAzSgw+b/4jk 4/eCUv8S+wZZp5l8DqbLNPAvrWWgOQCqxCigdADW7Uc6RVgbIsXM6Itf/DCyW5Iu WZDpHV6MLFNZehLq4cmflphVCW1V+Ntrvqw0jjbn5KKI+9n8fUSORnbszJhrS9ct RwbRW2xg6qwPc7/cHLoNZYQsdDXGursVRWeoCAeQPpF817dv3nvwP6WjYJaajqhg == Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fmuw7w2d0-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 31 Jul 2026 12:53:59 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 66VCQNog019446; Fri, 31 Jul 2026 12:53:58 GMT Received: from smtprelay07.fra02v.mail.ibm.com ([9.218.2.229]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fn7uwg1y2-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 31 Jul 2026 12:53:58 +0000 (GMT) Received: from smtpav04.fra02v.mail.ibm.com (smtpav04.fra02v.mail.ibm.com [10.20.54.103]) by smtprelay07.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 66VCrurf49873226 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 31 Jul 2026 12:53:56 GMT Received: from smtpav04.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 784C62004D; Fri, 31 Jul 2026 12:53:56 +0000 (GMT) Received: from smtpav04.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id F37CF20040; Fri, 31 Jul 2026 12:53:55 +0000 (GMT) Received: from [9.111.80.152] (unknown [9.111.80.152]) by smtpav04.fra02v.mail.ibm.com (Postfix) with ESMTP; Fri, 31 Jul 2026 12:53:55 +0000 (GMT) Message-ID: <45048fc9-119c-4770-a1d6-ab2b7e8f2a9e@linux.ibm.com> Date: Fri, 31 Jul 2026 14:53:55 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2] sched/fair: Prefer fully idle cores for NOHZ balancing To: Andrea Righi , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot Cc: Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Christian Loehle , Shrikanth Hegde , Phil Auld , linux-kernel@vger.kernel.org References: <20260729163225.1987068-1-arighi@nvidia.com> Content-Language: en-US From: Mete Durlu In-Reply-To: <20260729163225.1987068-1-arighi@nvidia.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: -lmtmlcPRkSTVABlJqY2KosE3tm8Kear X-Proofpoint-ORIG-GUID: nN05ppJNGbG5qTnz5JYXBQzbtNK8rL7h X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNzMxMDA5NiBTYWx0ZWRfXwIou3C+vBXiU XTvCSkJbzJHtTu9RUaJC44qdUm2D9B5kc/I6v6SVRJOvuWtuGTvP1swRdACjqG43ePSmHKLngHl iC6qvYdAeprUZtyxWq3z8kRu7JRMNQe6vvhL/HOfflK8uCOT/FEAewRBozshBuEK2Ls12YYYncq /glIclYQcH5edJo5W33+ybCG+I9waOjDOU9h/iK7VGncbSNbMz3FKPdUvOaIE2nG8hm/IORHY8I VcoJ7arVLx8AkvlX4h++dYZFPKN1v4dWkhL8U08eeio5jMZiq6VYiprkCB7MnWVuPhkrzk/Nn90 jiUyIlHFUgURIJGXQOm1c3FDF8tvViST/zQFuMsTx9Y4Z6AVjAIa6MQ7pTonNdaC2ZPNBHuPDRc tY4NrXD1je0sIA2Bpulzy2J/RA4clwK3ClbPacEfsOCKO8LO1byfV/cVUUr1+9mbl8cR6mcaxRV WPGQLARTCrq8jsFABQQ== X-Authority-Analysis: v=2.4 cv=SKFykuvH c=1 sm=1 tr=0 ts=6a6c9ae8 cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=IkcTkHD0fZMA:10 a=RAioF0-LDSMA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=iQ6ETzBq9ecOQQE5vZCe:22 a=VwQbUJbxAAAA:8 a=Ikd4Dj_1AAAA:8 a=zd2uoN0lAAAA:8 a=VnNF1IyMAAAA:8 a=oevgXaEZGow8fp0-k0IA:9 a=QEXdDO2ut3YA:10 X-Proofpoint-Spam-Info: AW1haW4tMjYwNzMxMDA5NiBTYWx0ZWRfX/iciCAL71gpl P9cEbO+q6O2lSqFF7LS22UrRiFTHbwcCy1q0lkqH/c/EAZOjlWh+r2qxByDQzJZGykQwq7QzUAC Ynlw2cM43OO9mmtZPXiy4MvoeB2A1D4= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-07-31_04,2026-07-30_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 priorityscore=1501 malwarescore=0 clxscore=1011 adultscore=0 lowpriorityscore=0 bulkscore=0 impostorscore=0 phishscore=0 spamscore=0 suspectscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2607310096 On 29/07/2026 18:32, Andrea Righi wrote: > find_new_ilb() selects the first idle housekeeping CPU without > considering whether another thread is running on the same physical core. > On an SMT system, the idle load balancer can therefore activate both > siblings even when another housekeeping CPU has an entirely idle core. > > On most SMT systems, this is not problematic because the idle load > balancer is a short-lived activity and the transient wakeup of a sibling > has negligible performance impact. > > However, this can be particularly costly on NVIDIA Olympus cores used in > Vera. Briefly activating an otherwise idle sibling can reduce the > performance available to the other sibling and this effect does not > necessarily end once the activated sibling becomes idle: after the ILB > finishes and its CPU enters WFI, full single-thread performance is > restored only after the sibling has remained idle for a qualification > interval (10 Ki cycles on the tested Vera system). Repeated short > sibling wakeups can therefore sustain the interference even with little > actual overlap. > > Prevent this by preferring an idle housekeeping CPU whose entire SMT > core is idle. Retain the first idle CPU as a fallback when no fully idle > core is available, so NOHZ balancing continues to make forward progress. > Once a partially busy core has been examined, skip its remaining SMT > siblings to avoid repeating the core-idle check on wide SMT systems. > > Tests performed using an ad hoc GEMM benchmark running one CPU-intensive > task per SMT core within its CPU affinity mask improved from > approximately 6.2 TFLOP/s to 9.4 TFLOP/s. > > Note that this preference may wake a fully idle physical core instead of > using an idle sibling of an active core, potentially increasing ILB > wakeup latency or energy consumption on some architectures. It may also > scan additional CPUs before selecting the one to run the ILB. The > selection falls back to the first idle CPU when no fully idle SMT core > is available. Non-SMT systems continue to select the first idle > housekeeping CPU. > > Cc: K Prateek Nayak > Cc: Shrikanth Hegde > Signed-off-by: Andrea Righi > --- > Changes in v2: > - Avoid repeated is_core_idle() checks on wide SMT systems by pruning > the remaining siblings of a partially busy core (Prateek Nayak) > - Link to v1: https://lore.kernel.org/r/20260728214442.1648483-1-arighi@nvidia.com/ > > kernel/sched/fair.c | 48 ++++++++++++++++++++++++++++++++++++--------- > 1 file changed, 39 insertions(+), 9 deletions(-) Hi, thank you for the interesting patch! I am testing this on s390 to see how it impacts our platform. After reading the discussion on v1 and reviewing the code I got minor nit. See below; > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index 37001c63452e5..b9e26938fb12a 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -13965,28 +13965,58 @@ static inline int on_null_domain(struct rq *rq) > static inline int find_new_ilb(void) > { > int this_cpu = smp_processor_id(); > - const struct cpumask *hk_mask; > - int ilb_cpu; > + struct cpumask *ilb_cpus; > + int ilb_cpu, fallback = -1; > + > + lockdep_assert_irqs_disabled(); > > - hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); > + /* > + * Reuse the per-CPU select_rq_mask, which is protected from concurrent > + * use on this CPU by having interrupts disabled. > + */ > + ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask); > + cpumask_and(ilb_cpus, nohz.idle_cpus_mask, > + housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); > > - for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) { > + for_each_cpu(ilb_cpu, ilb_cpus) { > if (ilb_cpu == this_cpu) > continue; > > - if (idle_cpu(ilb_cpu)) > - return ilb_cpu; > + if (!idle_cpu(ilb_cpu)) > + continue; > + > + /* > + * Running the idle load balancer on an idle sibling of a busy > + * SMT core can reduce the capacity available to its sibling. Prefer > + * a CPU whose entire core is idle, but retain the first idle CPU as > + * a fallback so idle balancing can still make progress when no fully > + * idle core exists. > + */ > + if (sched_smt_active() && !is_core_idle(ilb_cpu)) { > + if (fallback < 0) > + fallback = ilb_cpu; > + > + /* > + * The core is not idle, so there is no need to check > + * any of its other SMT siblings. > + */ > + cpumask_andnot(ilb_cpus, ilb_cpus, > + cpu_smt_mask(ilb_cpu)); Just a nit but; I think the purpose here is to move between cores if smt is active but current logic seems to be doing that only after an idle cpu is found. Consider that we first found an idle cpu but the siblings are busy, after we move to the next core we start traversing per cpu again. Wouldn't it be better if we always move per core after a fallback is found? How about sth like below? ... for_each_cpu(ilb_cpu, ilb_cpus) { if (ilb_cpu == this_cpu) continue; if (sched_smt_active()) { if (fallback < 0) { if (!idle_cpu(ilb_cpu)) continue; fallback = ilb_cpu; } /* * Running the idle load balancer on an idle sibling of * a busy SMT core can reduce the capacity available to * its sibling. Prefer a CPU whose entire core is idle, * but retain the first idle CPU as a fallback so idle * balancing can still make progress when no fully * idle core exists. */ if (is_core_idle(ilb_cpu)) return ilb_cpu; cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu)); continue; } if (idle_cpu(ilb_cpu)) return ilb_cpu; } return fallback;