From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from canpmsgout06.his.huawei.com (canpmsgout06.his.huawei.com [113.46.200.221]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6D4CD3793B0 for ; Thu, 16 Apr 2026 07:41:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=113.46.200.221 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776325314; cv=none; b=rxyKoFuJEWg3RvEUzpioOY7Lb/JApA6F5KA1p05py1I8Vmp3b1FqG44rq0FCuu1bFNCdW0RGFifSDqX8cigNIc8vj+g/q5z2Rya7ofphiKJgfcwp/Gp9uwP2kWmkRsDnPLowH2xIomwoQQZCEhWu2Mxu5UuW/zP7f9iZEGFfFxc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776325314; c=relaxed/simple; bh=U8aoTBrso/CGFc6yQivhj8nsbXCZMcyUL9/HkDdLAQg=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=EZ5Abg91X2an6BjNduy2dMtDdhc2cgRg3DzIKnMB1fR28bPnPurIRYwFfCSBBB9zxlG9zt6OWyGNb4w0ihih15QbhiZTbCMXcSjHSruGV/XvXRT53cXFkDlfg26tyCE0JKUbLadpa0a20xmbbHM8QkMGDqhMkmuUhAyAYKZsT1U= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=socFL0so; arc=none smtp.client-ip=113.46.200.221 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="socFL0so" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=6lHHJ2F3/GIFwnciQBBPgZ55vQuPYpU2vMfa7pl8kIk=; b=socFL0soiPZnlrAJs6RRUJcngRinTlnZSsEEZ9yx8dgVDdP5jkCP4Pjq0uI8ZL7tcFVNjgGBk TzxebFhkyz5ch/7EHWlNQmSs5weYj+qM+4+Q5JYOfL3u7uWbSMRKmtPkPjMiVlpT/cE/WbnufZS me9kzcmsmaxWg/stZuu9jIY= Received: from mail.maildlp.com (unknown [172.19.162.223]) by canpmsgout06.his.huawei.com (SkyGuard) with ESMTPS id 4fx8tn4lhwzRhR4; Thu, 16 Apr 2026 15:35:29 +0800 (CST) Received: from dggpemf100017.china.huawei.com (unknown [7.185.36.74]) by mail.maildlp.com (Postfix) with ESMTPS id 9309540561; Thu, 16 Apr 2026 15:41:48 +0800 (CST) Received: from [10.67.111.186] (10.67.111.186) by dggpemf100017.china.huawei.com (7.185.36.74) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Thu, 16 Apr 2026 15:41:47 +0800 Message-ID: <45712e75-f27e-804a-940b-79e932452e6e@huawei.com> Date: Thu, 16 Apr 2026 15:41:47 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.1.1 Subject: Re: [RFC PATCH] sched/fair: scale wake_wide() threshold by SMT width To: Shrikanth Hegde CC: , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , References: <20260407063915.2034198-1-zhangqiao22@huawei.com> <4276a8b7-141a-4e0b-bc35-47cb0a341801@linux.ibm.com> From: Zhang Qiao In-Reply-To: <4276a8b7-141a-4e0b-bc35-47cb0a341801@linux.ibm.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 8bit X-ClientProxiedBy: kwepems500001.china.huawei.com (7.221.188.70) To dggpemf100017.china.huawei.com (7.185.36.74) Hi Shrikanth, 在 2026/4/8 1:58, Shrikanth Hegde 写道: > Hi. > > On 4/7/26 12:09 PM, Zhang Qiao wrote: >> wake_wide() uses sd_llc_size as the spreading threshold to detect wide >> waker/wakee relationships and to disable wake_affine() for those cases. >> >> On SMT systems, sd_llc_size counts logical CPUs rather than physical >> cores. This inflates the wake_wide() threshold, allowing wake_affine() >> to pack more tasks into one LLC domain than the actual compute capacity >> of its physical cores can sustain. The resulting SMT interference may >> cost more than the cache-locality benefit wake_affine() intends to gain. >> > > Isn't load balance to move it out? What does the workload do? The workload is a producer-consumer model: one producer wakes up ~50 different consumers, with roughly 10+ consumers running concurrently. The total number of tasks is well below the CPU count. In this scenario, load balancing is largely ineffective. Each consumer spends most of its time sleeping, gets woken by the producer, runs briefly to process the message, then goes back to sleep. There is almost no window where a consumer sits on a CPU runqueue in the runnable state waiting to be pulled. Since load balancing can only migrate runnable tasks, it simply has no target to act on here. > >> Scale the factor by the SMT width of the current CPU so that it >> approximates the number of independent physical cores in the LLC domain, >> making wake_wide() more likely to kick in before SMT interference >> becomes significant. On non-SMT systems the SMT width is 1 and behaviour >> is unchanged. >> > > There are systems where LLC_SIZE == SMT_SIZE. i.e one core in the LLC. > This would effectively disable wake_affine feature in such systems. > > Power10 being a major example. > >> Signed-off-by: Zhang Qiao >> --- >>   kernel/sched/fair.c | 5 +++++ >>   1 file changed, 5 insertions(+) >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index f07df8987a5ef..4896582c6e904 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -7334,6 +7334,11 @@ static int wake_wide(struct task_struct *p) >>       unsigned int slave = p->wakee_flips; >>       int factor = __this_cpu_read(sd_llc_size); >>   +    /* Scale factor to physical-core count to account for SMT interference. */ >> +    if (sched_smt_active()) >> +        factor = DIV_ROUND_UP(factor, >> +                cpumask_weight(cpu_smt_mask(smp_processor_id()))); >> + >>       if (master < slave) >>           swap(master, slave); >>       if (slave < factor || master < slave * factor) > > .