From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.12]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E9786314D18 for ; Thu, 6 Aug 2026 03:25:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.12 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785986709; cv=none; b=j73PJzNoqdwDKUReoNhgdX3zdTpMflw9FlkcHq5bAJDKgvCH7FEBfslcjszxqYNh1NyT7UnO9LR53jua11x8rUupp8s4eoTwcFj6GJ3eKbLBkngJzjEf/M4SgHJVqAdO5x8hhmQ7Wc+/S9JcPrErH5lpN99VjwpYfn6DzP/WAME= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785986709; c=relaxed/simple; bh=R2kquQQEGhvQaEcAZk852HKLUusBWvLl2ac9ME0G/OQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=IwTjtTM7hg0qdyjnGJKTIoEZOA5JWB2ZA7pW0kzprDxAY3wsdmEyOwe6wn34RYzRk0ppr4wKCGNzX4Q4XxmMNt6+CmXoLevpsY/UlSjtSCXf9EuyvTdytph/uFaHOf8aHMoz5op978PxMnWrwKylFMSlioUWPeesaVtxSZxe/Kw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=VkVmpQPt; arc=none smtp.client-ip=198.175.65.12 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="VkVmpQPt" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785986706; x=1817522706; h=date:from:to:cc:subject:message-id:references: mime-version:in-reply-to; bh=R2kquQQEGhvQaEcAZk852HKLUusBWvLl2ac9ME0G/OQ=; b=VkVmpQPtjv/b5Mk9GxhgpydQ40nI/APyyWsoWwDT03AKr+GAsQhu98cC 2woVXgF9QQukb4fHF4vWZxDmJoHr4LtK+SWn19OpG5TH4YHGdlfRTC2Vj PYnH8r0NBew5cyKN9sML0HQ7HkMDAA+iuk1/lQnpaKTU7yTQYovVpkieE 0qOE8slGEYqWn9CcjEZQN7/TW0mZItljDARpe5BfHEeR0JvDf5ME/Vp2z rAva9NYFc6XRGOPwN+LYXW/a7BFB1tOgzp3aBjG32v2d0vh63+bTvhCmI zPoB+OG8NaU9av/UzNB/455vAyRCxNKQQqkieKA6epVaIofRwSWFI1td4 g==; X-CSE-ConnectionGUID: TzDRGW7yQb+28dZLkYuK6g== X-CSE-MsgGUID: 2LI8gg2aTRGptZcp7AH+IA== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="98083211" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="98083211" Received: from orviesa007.jf.intel.com ([10.64.159.147]) by orvoesa104.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 20:25:05 -0700 X-CSE-ConnectionGUID: sSfs/M1vQQiIORqluhUwfQ== X-CSE-MsgGUID: LSx0oF/1SbG81MEvw/N5tg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="262051946" Received: from ranerica-svr.sc.intel.com ([172.25.110.23]) by orviesa007.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 20:25:04 -0700 Date: Wed, 5 Aug 2026 20:34:07 -0700 From: Ricardo Neri To: Vincent Guittot Cc: Christian Loehle , Ingo Molnar , Peter Zijlstra , Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , K Prateek Nayak , Andrea Righi , Barry Song , "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org Subject: Re: [PATCH v6 5/6] sched/fair: Allow load balancing between CPUs of identical capacity Message-ID: <20260806033407.GA5786@ranerica-svr.sc.intel.com> References: <20260720-rneri-fix-cas-clusters-v6-0-bb500bf4afd4@linux.intel.com> <20260720-rneri-fix-cas-clusters-v6-5-bb500bf4afd4@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.9.4 (2018-02-28) On Tue, Aug 04, 2026 at 11:55:52AM +0200, Vincent Guittot wrote: > On Thu, 23 Jul 2026 at 09:11, Christian Loehle wrote: > > > > On 7/21/26 03:43, Ricardo Neri wrote: > > > sched_balance_find_src_rq() avoids selecting a runqueue with a single > > > running task as busiest if doing so results in migrating the task to a > > > CPU with less than ~5% of extra capacity. It also unintentionally > > > prevents migrations between CPUs of identical capacity. > > > > > > When CONFIG_SCHED_CLUSTER is enabled, load should be balanced across > > > clusters of CPUs with the same capacity. Allowing migration between CPUs > > > of identical capacity is necessary to meet this goal. > > > > > > Use get_actual_cpu_capacity() to reflect architectural capacity as well > > > as diminished capacity due to hardware or cpufreq pressure. Guard this > > > check with the sched_cluster_active static key so that systems without > > > cluster topology are unaffected. > > > > > > Tested-by: Christian Loehle > > > Tested-by: Andrea Righi > > > Signed-off-by: Ricardo Neri > > > --- > > > Changes in v6: > > > * Switched to use get_actual_cpu_capacity() instead of > > > arch_scale_cpu_capacity(). The former considers rq->avg_hw.load_avg and > > > cpufreq_pressure and their impact on CPU capacity. (Vincent) > > > * Renamed the variable same_arch_cluster as cluster_equal_cap for > > > clarity. (Andrea) > > > * Added Tested-by tag from Andrea. Thanks! > > > > > > Changes in v5: > > > * Optimized logic to identify same-arch clusters only when needed. > > > * Added Tested-by tag from Christian. Thanks! > > > > > > Changes in v4: > > > * Implemented the check for cluster with a local variable for improved > > > readability. > > > > > > Changes in v3: > > > * Reverted the inverted capacity check; the inverted form incorrectly > > > allows migrations to CPUs of slightly less capacity. > > > * Guarded the check for architectural capacity with the > > > sched_cluster_active static key. > > > > > > Changes in v2: > > > * Used arch_scale_cpu_capacity() instead of capacity_of() to ignore > > > runtime variability. > > > * Inverted the check for runtime capacity. (Christian) > > > * Reworded patch description for clarity. > > > --- > > > kernel/sched/fair.c | 9 ++++++++- > > > 1 file changed, 8 insertions(+), 1 deletion(-) > > > > > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > > > index feea47e6abea..de4189b562ac 100644 > > > --- a/kernel/sched/fair.c > > > +++ b/kernel/sched/fair.c > > > @@ -13104,13 +13104,20 @@ static struct rq *sched_balance_find_src_rq(struct lb_env *env, > > > */ > > > if (env->sd->flags & SD_ASYM_CPUCAPACITY && > > > nr_running == 1) { > > > + bool cluster_equal_cap = static_branch_unlikely(&sched_cluster_active) && > > > + (get_actual_cpu_capacity(env->dst_cpu) == > > > + get_actual_cpu_capacity(i)); > > > > > > I guess it's extremely unlikely, but it _feels_ wrong to have to clusters of different > > arch_scale_cpu_capacity() equal to true here because of system/thermal pressure (which is > > obviously considered more transient). > > Adding && arch_scale_cpu_capacity(env->dst_cpu) == arch_scale_cpu_capacity(i) might even > > make the check cheaper because it's better for the branch predictor than get_actual_cpu_capacity(). > > Vincent, would you be fine with requiring both: equal get_actual_cpu_capacity() and > > arch_scale_cpu_capacity()? > > TBH, I don't have a strong opinion on this. I have in mind that the > cpufreq pressure can last a long time (i.e., several hundreds of ms) > so it could make sense to spread tasks between clusters even if one > has lower max capacity. The task will migrate back to the cpu with > higher capacity once the cpufreq pressure is removed but this would > need some test results I am with Vincent in this: it would be better to spread tasks among clusters of equal _actual_ capacity even if they arch capacity is different. Besides diminished compute capacity, packing tasks in one cluser would cause L2 cache contention, which is what cluster scheduling wants to avoid. I am running some workloads and will report results.