From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 5210B41C2E7 for ; Mon, 11 May 2026 15:54:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778514870; cv=none; b=E+mQQo72wwV0C1DpETmtrWTbftYFSVqiFgzyrV0VxKNkdu8Ebqp4TJX4V9bqCeIzfDcbHFoRTZnoZhiNeFOir6gLdEhGTJOGh3okZf8nSGCNTD1rnGAm2a1JRmaZt1ZPwHXe7kDv3Y60O/BxjMN1WCEqcofumR98C2BWqxUAg/Y= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778514870; c=relaxed/simple; bh=BMBWjos0P4iE3h3Cz04xPSpPzHxJslZ1jh4qEd8hHqw=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=rM2uuxL2wQS9mnLdf95r7aejnCsReZZeFSv1pXx4zU/t/SxH6TYZn7/aEoT28kzRa9g1asjkZrlTdw5IoA20TdLerg7DRBPHNwLxpwLjhgQiTOEfbZoomI9tKy17lJulLiW1j1f003lm2XG8CwWH5QctDukkyj2g0U7Ng3iPUKE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=QO24xzDk; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="QO24xzDk" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 326091A2D; Mon, 11 May 2026 08:54:21 -0700 (PDT) Received: from [192.168.178.6] (usa-sjc-mx-foss1.foss.arm.com [172.31.20.19]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 8959C3F836; Mon, 11 May 2026 08:54:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1778514866; bh=BMBWjos0P4iE3h3Cz04xPSpPzHxJslZ1jh4qEd8hHqw=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=QO24xzDkYd3UQfxAEmDiFgiolqRT11Pdi9CcFpyzN4IRzXIV/GfobDMCpUFbsGpJT doDBwlHn8GXBAbbzl4DSqZNwI01gdXXu1UKabKeX8dZNmmfguSHoStA6WFuPtGycxT 1srBZ1uh9N9JshEJOUMgaSnx5sk8cczCjjF8Jujk= Message-ID: <1837536b-b77a-4312-b87d-ed04e067e234@arm.com> Date: Mon, 11 May 2026 17:54:16 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RFC PATCH] sched/fair: scale wake_wide() threshold by SMT width To: Zhang Qiao , Shrikanth Hegde Cc: tanghui20@huawei.com, Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , linux-kernel@vger.kernel.org References: <20260407063915.2034198-1-zhangqiao22@huawei.com> <4276a8b7-141a-4e0b-bc35-47cb0a341801@linux.ibm.com> <45712e75-f27e-804a-940b-79e932452e6e@huawei.com> <8f71c69c-ddd6-4117-97ca-cefaa2a22180@arm.com> <309984ba-ec64-3186-2dfe-ffa2abfad00c@huawei.com> From: Dietmar Eggemann Content-Language: en-GB In-Reply-To: <309984ba-ec64-3186-2dfe-ffa2abfad00c@huawei.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On 29.04.26 04:43, Zhang Qiao wrote: > > Hi, > > 在 2026/4/22 21:26, Dietmar Eggemann 写道: >> On 16.04.26 09:41, Zhang Qiao wrote: >>> Hi Shrikanth, >>> >>> 在 2026/4/8 1:58, Shrikanth Hegde 写道: >>>> Hi. >>>> >>>> On 4/7/26 12:09 PM, Zhang Qiao wrote: [...] >>> The workload is a producer-consumer model: one producer wakes up ~50 >>> different consumers, with roughly 10+ consumers running concurrently. >>> The total number of tasks is well below the CPU count. >> >> But higher than your MC core count I believe? Otherwise you wouldn't >> care. I assume you have MC CPU count of 12-24. Do you have more than 2 >> different MCs. > > My server has 10 different MCs (LLCs), with each MC containing 8 physical cores > (16 threads with SMT-2). Thanks. >>> In this scenario, load balancing is largely ineffective. Each consumer >>> spends most of its time sleeping, gets woken by the producer, runs >>> briefly to process the message, then goes back to sleep. There is >>> almost no window where a consumer sits on a CPU runqueue in the runnable >>> state waiting to be pulled. Since load balancing can only migrate >>> runnable tasks, it simply has no target to act on here. >> >> OK, but SD_BALANCE_WAKE is not set by default, nobody would experience a > > SD_BALANCE_WAKE was not enabled in my tests. Right, looks like I mixed up balance flags & fast/slow path with the wake affine vs. wake wide logic. >> difference in behaviour on an SMT machine in terms of waking tasks wide, >> i.e. going through the slow path. Like I tried to explain in the >> adjacent thread, your wakees would only end up in the slow path in case >> your sched domains would have SD_BALANCE_WAKE set.> >> Or do you just want to force wakeups which have wake_wide(p) return 1 >> always into the fast path with 'new_cpu == prev_cpu'? But this wouldn't >> be wake wide? > > The observed improvement comes from suppressing wake_affine() before it > pulls wakees onto the waker's physical core. In the producer-consumer > workload, without this patch, consumers are repeatedly affined into the > waker's LLC and end up co-scheduled on the same physical core's SMT > siblings. With the patch, wake_wide() fires earlier and wakees are left > on prev_cpu, resulting in better spread across physical cores. Makes sense. You mentioned having ~10+ consumers running concurrently. I’m curious why select_idle_sibling() isn’t doing a better job of distributing those tasks across idle cores, even though wakeups are affine to the waker and its LLC domain. Is this because you only have 8 cores per LLC, combined with general system noise?