From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 2B3083C872C for ; Wed, 22 Apr 2026 13:24:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776864250; cv=none; b=JbjJy+dLecMUYIy7PkyGA00qs8UGDAzUbLSAIK/ChIMb7Sq2wpzBexaBku8bXp1S28mq1NgTdCXE02cyLnP4gX8ocYN/ud/ISjXTys3C+nMZqoV3UQkYHoEJ9y0Tc2CE+SbddNuI5Khx9+LTGGfjemXYaqzlA3RtMn4NDVav+hM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776864250; c=relaxed/simple; bh=lPqfUW8lNwpzmoifUZdY7pT5WBp7n2Uv1UYnFKxll9o=; h=Message-ID:Date:MIME-Version:From:Subject:To:Cc:References: In-Reply-To:Content-Type; b=mVoc1D49x2vWXMUXTIVgISZHiYFhYTtVPE8CtiJtQeZmhVFnwzb4JnGf2pGm1iXM/McDWi989jrImQ6dUt94RTs5ovSC4sgCNIzTRvmqSHQNazh9SCxZhqIEjtKmv7xEVuScFv0rv7zDL9wND21GSok6P+CWT4ZGdZjRcDRVYxQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=SS55FhfT; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="SS55FhfT" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 87C151E2F; Wed, 22 Apr 2026 06:23:58 -0700 (PDT) Received: from [10.57.65.82] (unknown [10.57.65.82]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 69F703FD46; Wed, 22 Apr 2026 06:24:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1776864244; bh=lPqfUW8lNwpzmoifUZdY7pT5WBp7n2Uv1UYnFKxll9o=; h=Date:From:Subject:To:Cc:References:In-Reply-To:From; b=SS55FhfTrys6Oc1yLNCqJ4GcEpbQPKgkDw9XmLtEKns4zu0W+em3kEzxFQdD7gkC1 lPOBvX2GGCcFwiQAnESQxzjzvxJP1eQKfhQnviwZ13Sy8aj9a0hNd2QJqfSkGkpvqh mFT7ExQrye37h8cIO154nkzspoPD+zAfzY6XbsgU= Message-ID: Date: Wed, 22 Apr 2026 15:24:00 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird From: Dietmar Eggemann Subject: Re: [RFC PATCH] sched/fair: scale wake_wide() threshold by SMT width To: Shrikanth Hegde Cc: tanghui20@huawei.com, Zhang Qiao , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , linux-kernel@vger.kernel.org References: <20260407063915.2034198-1-zhangqiao22@huawei.com> <0aca0ee5-4102-426e-b290-6c50b6e819a7@arm.com> <45999d3a-4a83-4b58-bfda-9639a27eb2e4@linux.ibm.com> Content-Language: en-GB In-Reply-To: <45999d3a-4a83-4b58-bfda-9639a27eb2e4@linux.ibm.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On 07.04.26 20:16, Shrikanth Hegde wrote: > > > On 4/7/26 8:08 PM, Dietmar Eggemann wrote: >> On 07.04.26 08:39, Zhang Qiao wrote: >>> wake_wide() uses sd_llc_size as the spreading threshold to detect wide >>> waker/wakee relationships and to disable wake_affine() for those cases. >>> >>> On SMT systems, sd_llc_size counts logical CPUs rather than physical >>> cores. This inflates the wake_wide() threshold, allowing wake_affine() >>> to pack more tasks into one LLC domain than the actual compute capacity >>> of its physical cores can sustain. The resulting SMT interference may >>> cost more than the cache-locality benefit wake_affine() intends to gain. >>> >>> Scale the factor by the SMT width of the current CPU so that it >>> approximates the number of independent physical cores in the LLC domain, >>> making wake_wide() more likely to kick in before SMT interference >>> becomes significant. On non-SMT systems the SMT width is 1 and behaviour >>> is unchanged. >>> >>> Signed-off-by: Zhang Qiao >>> --- >>>   kernel/sched/fair.c | 5 +++++ >>>   1 file changed, 5 insertions(+) >>> >>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >>> index f07df8987a5ef..4896582c6e904 100644 >>> --- a/kernel/sched/fair.c >>> +++ b/kernel/sched/fair.c >>> @@ -7334,6 +7334,11 @@ static int wake_wide(struct task_struct *p) >>>       unsigned int slave = p->wakee_flips; >>>       int factor = __this_cpu_read(sd_llc_size); >>>   +    /* Scale factor to physical-core count to account for SMT >>> interference. */ >>> +    if (sched_smt_active()) >>> +        factor = DIV_ROUND_UP(factor, >>> +                cpumask_weight(cpu_smt_mask(smp_processor_id()))); >>> + >>>       if (master < slave) >>>           swap(master, slave); >>>       if (slave < factor || master < slave * factor) >> >> I assume not a lot of people care since this needs: > > wake_affine machinery needs SD_WAKE_AFFINE. No? Yes, the potential call to wake_affine() and forcing 'sd = NULL' but that's not forcing a wakeup (WF_TTWU) into the slow path (sched_balance_find_dst_cpu()), which IMHO is the actual wake wide. You need 'sd != NULL' which can only be set by (1)for a wakeup: for_each_domain(cpu, tmp) ... if (tmp->flags & sd_flag) <-- '(1) SD_BALANCE_WAKE == WF_TTWU' sd = tmp; and since SD_BALANCE_WAKE is never set per default in sd_init() [kernel/sched/topology.c] I wonder how they achieved this wide (i.e. not affine MC for this_cpu or prev_cpu) wakeup? By default, we only select wide for WF_FORK and WF_EXEC. Or do they just want to force 'wake_wide(p) == 1' into sis(..., new_cpu = prev_cpu) ? [...]