From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from canpmsgout03.his.huawei.com (canpmsgout03.his.huawei.com [113.46.200.218]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2F7C93AC00 for ; Wed, 29 Apr 2026 02:43:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=113.46.200.218 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777430597; cv=none; b=SVISvT/xemH05nyq/qixlwgssZoVcVFzZrsJLx+uMcg/vIHG2ydpF+AE/6xUysAWzxWSmh5rByfU5vNui1eW+mKK5WltJyIUplSECkpYhl8EPGPIrh+36DDEmUrOBIWzDjRvtVPwMADGuGGeu8CS4GhVAnX8jp0/qlPwSrcao9I= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777430597; c=relaxed/simple; bh=k7ykgpMpydAF76/zj1MoyWcGgZQGjMxrasCYy0QeCO8=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=QTFmGQuTBxA35RjkN1xSQ46/iU46kuUsHMAJjdMMdqlsYIo4HJPGE75UhpXTRcyIwZHD96xwgqvMsW8rB3zqIbAzjnNiFD5RrKg6GzWaAicig2xyYyAJR49afd9unVEgYBs4Re/Lt2zeyWEfWkUrZDEB9RdbY7ZfzFkvYNoJ1MM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=Pugf+AJL; arc=none smtp.client-ip=113.46.200.218 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="Pugf+AJL" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=wZ0rFY8on7dgQBFeqJt1XqXXUGeHVqXyLfSkTk+JQS4=; b=Pugf+AJLGJtIg7PGvyKuHTEm3+2fUi/ZubxENc/JpFym1LJlS+CZzP07ztwTcIJAbn+2WqzEa WVq/xYy+ZfK83PP/gYzru/xUmsaeJ7JrLaWFEwRrTsp3VMZQzL8CeZxciUT7A24oq0Owg7vnPG9 tuvkXzAonSL/Mqwcx8XYAyY= Received: from mail.maildlp.com (unknown [172.19.162.197]) by canpmsgout03.his.huawei.com (SkyGuard) with ESMTPS id 4g51dv0nC3zpSwY; Wed, 29 Apr 2026 10:36:35 +0800 (CST) Received: from dggpemf100017.china.huawei.com (unknown [7.185.36.74]) by mail.maildlp.com (Postfix) with ESMTPS id 8610640569; Wed, 29 Apr 2026 10:43:11 +0800 (CST) Received: from [10.67.111.186] (10.67.111.186) by dggpemf100017.china.huawei.com (7.185.36.74) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Wed, 29 Apr 2026 10:43:10 +0800 Message-ID: <309984ba-ec64-3186-2dfe-ffa2abfad00c@huawei.com> Date: Wed, 29 Apr 2026 10:43:10 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.1.1 Subject: Re: [RFC PATCH] sched/fair: scale wake_wide() threshold by SMT width To: Dietmar Eggemann , Shrikanth Hegde CC: , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , References: <20260407063915.2034198-1-zhangqiao22@huawei.com> <4276a8b7-141a-4e0b-bc35-47cb0a341801@linux.ibm.com> <45712e75-f27e-804a-940b-79e932452e6e@huawei.com> <8f71c69c-ddd6-4117-97ca-cefaa2a22180@arm.com> From: Zhang Qiao In-Reply-To: <8f71c69c-ddd6-4117-97ca-cefaa2a22180@arm.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 8bit X-ClientProxiedBy: kwepems200002.china.huawei.com (7.221.188.68) To dggpemf100017.china.huawei.com (7.185.36.74) Hi, 在 2026/4/22 21:26, Dietmar Eggemann 写道: > On 16.04.26 09:41, Zhang Qiao wrote: >> Hi Shrikanth, >> >> 在 2026/4/8 1:58, Shrikanth Hegde 写道: >>> Hi. >>> >>> On 4/7/26 12:09 PM, Zhang Qiao wrote: >>>> wake_wide() uses sd_llc_size as the spreading threshold to detect wide >>>> waker/wakee relationships and to disable wake_affine() for those cases. >>>> >>>> On SMT systems, sd_llc_size counts logical CPUs rather than physical >>>> cores. This inflates the wake_wide() threshold, allowing wake_affine() >>>> to pack more tasks into one LLC domain than the actual compute capacity >>>> of its physical cores can sustain. The resulting SMT interference may >>>> cost more than the cache-locality benefit wake_affine() intends to gain. >>>> >>> >>> Isn't load balance to move it out? What does the workload do? >> >> The workload is a producer-consumer model: one producer wakes up ~50 >> different consumers, with roughly 10+ consumers running concurrently. >> The total number of tasks is well below the CPU count. > > But higher than your MC core count I believe? Otherwise you wouldn't > care. I assume you have MC CPU count of 12-24. Do you have more than 2 > different MCs. My server has 10 different MCs (LLCs), with each MC containing 8 physical cores (16 threads with SMT-2). > >> In this scenario, load balancing is largely ineffective. Each consumer >> spends most of its time sleeping, gets woken by the producer, runs >> briefly to process the message, then goes back to sleep. There is >> almost no window where a consumer sits on a CPU runqueue in the runnable >> state waiting to be pulled. Since load balancing can only migrate >> runnable tasks, it simply has no target to act on here. > > OK, but SD_BALANCE_WAKE is not set by default, nobody would experience a SD_BALANCE_WAKE was not enabled in my tests. > difference in behaviour on an SMT machine in terms of waking tasks wide, > i.e. going through the slow path. Like I tried to explain in the > adjacent thread, your wakees would only end up in the slow path in case > your sched domains would have SD_BALANCE_WAKE set.> > Or do you just want to force wakeups which have wake_wide(p) return 1 > always into the fast path with 'new_cpu == prev_cpu'? But this wouldn't > be wake wide? The observed improvement comes from suppressing wake_affine() before it pulls wakees onto the waker's physical core. In the producer-consumer workload, without this patch, consumers are repeatedly affined into the waker's LLC and end up co-scheduled on the same physical core's SMT siblings. With the patch, wake_wide() fires earlier and wakees are left on prev_cpu, resulting in better spread across physical cores. Thanks Zhang Qiao > > [...] > > . >