From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F3152372ED0 for ; Mon, 22 Jun 2026 23:55:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782172545; cv=none; b=spr0mnuXZQhU+gETK0TmrgfeKWt4neWldAgqqNaQKnpCg+8wdGspDJjSHEep92uo63sg+UvgyEGUf3XgEGVKRyPsk0zKGpxoK7qtJDzkGX4rkcgifEmgF6AzY0H5SyRVF/L4QxaBoMgeoHgqjGRsdNvtFa3Bm568AjwR07O5m8Y= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782172545; c=relaxed/simple; bh=x2K39AFJmdRhNQGFrE7+3vKUL0yCoNjw5mmM5ALREMY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=m3VPCY7Ytm2KMjh3NuL3nfVXkfpfFqzjlEpVtn/9v1pFbSbNMfqt2othBngYcY1cTttOSYShINqWtmSA73o49p22D3mq7tzWybj0dHGqMYqH+aDrjBy46CLlTucG2/crxcUOKYcU+v6acrqT1FYY4/Fe3uV3XQisSl66/zjj4ZI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=Ig9fwkgG; arc=none smtp.client-ip=192.198.163.18 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="Ig9fwkgG" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1782172537; x=1813708537; h=from:date:subject:mime-version:content-transfer-encoding: message-id:references:in-reply-to:to:cc; bh=x2K39AFJmdRhNQGFrE7+3vKUL0yCoNjw5mmM5ALREMY=; b=Ig9fwkgGHK8KNe4xf7sKtpIt6AKzHsEdCNGCyH6uiApQ9cUc1YkvYh0y YuMlPnLqsClA+b5eugvoLs+RrJT3zNAAJEoVCw+fea/cLtTaDxTGVIfnh tT5s9wmENNuQg591tFtdnzjOwZqTMxvq1IxxyyV7Kd5j7WmwYhHO9Whdz bM9+zm7G1JjwARkKYIVxxfxDxNo66zbwgdlya5DkH71OnWPa1s6asqacI qsGiRA5168U+qL9Y8omNj2wcKKcmqyrVlI8YV9IZrVBfJfMnTFbxxPKtw TmxQtnk9RmPbM49F+gYT2e3Uk7ZTEDWL5Je2PYKttzh8VpN6Zs5T4nW9R A==; X-CSE-ConnectionGUID: mfSQX5nZQGWcr5OB8peFSA== X-CSE-MsgGUID: Syuhni2eQra0RZA7nkw8aw== X-IronPort-AV: E=McAfee;i="6800,10657,11825"; a="82014534" X-IronPort-AV: E=Sophos;i="6.24,219,1774335600"; d="scan'208";a="82014534" Received: from fmviesa003.fm.intel.com ([10.60.135.143]) by fmvoesa112.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 22 Jun 2026 16:55:32 -0700 X-CSE-ConnectionGUID: Rxl4dsmoS9KEg7CIXU1IDw== X-CSE-MsgGUID: 75raZ35DQHG9lLuKX6HreA== X-ExtLoop1: 1 Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by fmviesa003.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 22 Jun 2026 16:55:30 -0700 From: Ricardo Neri Date: Mon, 22 Jun 2026 17:05:56 -0700 Subject: [PATCH v5 6/6] sched/topology: Do not clear SD_PREFER_SIBLING in domains with clusters Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260622-rneri-fix-cas-clusters-v5-6-19968f2d1497@linux.intel.com> References: <20260622-rneri-fix-cas-clusters-v5-0-19968f2d1497@linux.intel.com> In-Reply-To: <20260622-rneri-fix-cas-clusters-v5-0-19968f2d1497@linux.intel.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , Christian Loehle , K Prateek Nayak , Barry Song Cc: "Rafael J. Wysocki" , Andrea Righi , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1782173186; l=3739; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=x2K39AFJmdRhNQGFrE7+3vKUL0yCoNjw5mmM5ALREMY=; b=0nqSXB7aCnrNcT1Eviaxx7hcs1N55rP8BHHvLVZSff1Eb8PILqwflhzBfVpn6WJmmsTC3n+C2 iga/LXSDNCLAcBUePimakCeE7c0kZAzdqWkfBbzrZqGUGYcRPFGEC5o X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= Some topologies have scheduling domains that contain CPUs of asymmetric capacity, grouped into two or more clusters of equal-capacity CPUs sharing an L2 cache. When CONFIG_SCHED_CLUSTER is enabled, load must be balanced across these clusters. Do not clear SD_PREFER_SIBLING in the child domains to indicate to the load balancer that it should spread load among cluster siblings. Checks for capacity in update_sd_pick_busiest(), sched_balance_find_src_group(), and sched_balance_find_src_rq() prevent migrations from high- to low-capacity CPUs if the busiest group is not overloaded. CPUs with spare capacity, big or small, have always helped overloaded groups. Once the overloading condition disappears, misfit load will still be used to move high-utilization tasks to bigger CPUs if they have spare capacity. Adding the SD_PREFER_SIBLING flag shifts load balancing in shared-LLC domains from equalizing the number of idle CPUs to equalizing the number of running tasks. This also enables migrations among clusters from newly- idle load balance, where the outgoing task is already dequeued but the CPU has not yet transitioned to idle. Reviewed-by: Tim Chen Tested-by: Christian Loehle Signed-off-by: Ricardo Neri --- Changes in v5: * Improved inline comments for accuracy. * Added Tested-by tag from Christian. Thanks! Changes in v4: * Added Reviewed-by tag from Tim. Thanks! Changes in v3: * Updated documentation of SD_PREFER_SIBLING. * Expanded the patch description to explain the behavior when overloaded groups are involved. Changes in v2: * Reworded the patch description for clarity. * Kept parentheses around bitwise operators for clarity. --- include/linux/sched/sd_flags.h | 3 ++- kernel/sched/topology.c | 14 ++++++++++++-- 2 files changed, 14 insertions(+), 3 deletions(-) diff --git a/include/linux/sched/sd_flags.h b/include/linux/sched/sd_flags.h index 42839cfa2778..f9a46fb8cacf 100644 --- a/include/linux/sched/sd_flags.h +++ b/include/linux/sched/sd_flags.h @@ -147,7 +147,8 @@ SD_FLAG(SD_ASYM_PACKING, SDF_NEEDS_GROUPS) * Prefer to place tasks in a sibling domain * * Set up until domains start spanning NUMA nodes. Close to being a SHARED_CHILD - * flag, but cleared below domains with SD_ASYM_CPUCAPACITY. + * flag, but cleared below domains with SD_ASYM_CPUCAPACITY unless those child + * domains have clusters of CPUs sharing cache. * * NEEDS_GROUPS: Load balancing flag. */ diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c index 622e2e01974c..261b407d0936 100644 --- a/kernel/sched/topology.c +++ b/kernel/sched/topology.c @@ -1995,8 +1995,18 @@ sd_init(struct sched_domain_topology_level *tl, /* * Convert topological properties into behaviour. */ - /* Don't attempt to spread across CPUs of different capacities. */ - if ((sd->flags & SD_ASYM_CPUCAPACITY) && sd->child) + /* + * Don't attempt to spread across CPUs of different capacities. + * + * If the child domain has clusters of CPUs sharing L2 cache, keep the + * flag to spread tasks across clusters of identical capacity. Checks in + * the load balancer prevent task migrations from high- to low-capacity + * CPUs unless the source group is overloaded. Migrations to a lower- + * capacity CPU can happen if a higher-capacity group is overloaded and + * a lower-capacity CPU has spare capacity. + */ + if ((sd->flags & SD_ASYM_CPUCAPACITY) && sd->child && + !(sd->child->flags & SD_CLUSTER)) sd->child->flags &= ~SD_PREFER_SIBLING; if (sd->flags & SD_SHARE_CPUCAPACITY) { -- 2.43.0