From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.13]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8553839935E for ; Thu, 14 May 2026 18:24:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.13 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778783051; cv=none; b=Er9s+xJkyqlFd+TSKzI1LaQJ119cdDUQAPF4RxT7tbvgv9BkL7ZnThwhLVMWMMouIaGdz4gPnzteH251+nxTiODQMBBzPg6RL7uTbOO8G8rNdgWbV4bXnKWh3SQzzKcwrmsu3bPoyTLnx7r8InoFREP7ImxVLXJNKyaPxlPa2YQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778783051; c=relaxed/simple; bh=cPPaIml3mystyoVfu07RSFsEBmIHYsEQUNd6cORguUw=; h=From:Subject:Date:Message-Id:MIME-Version:Content-Type:To:Cc; b=nWP8QRdAzuvWE2KIIaAXkFKaeVt9zobWS7oIncJLLzhbz/sGu/ahzX50sSLe5VmWoyYJvl7PXdmcUxsi05iQ6pJPzY/YEv1R4S8asTAj9luEsxyHyehGFRWSkWIs3AtiPp+TnWmYd6ZKG4FTH65bpuZ3+Tb465BvldQZ1YAYt3g= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=CMt2FQv9; arc=none smtp.client-ip=192.198.163.13 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="CMt2FQv9" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1778783049; x=1810319049; h=from:subject:date:message-id:mime-version: content-transfer-encoding:to:cc; bh=cPPaIml3mystyoVfu07RSFsEBmIHYsEQUNd6cORguUw=; b=CMt2FQv9jO7gBDg6bHdM8G4sZ7F/sbhZzS8VUf+04XtRQqMEX9F945D5 cJbT/OAmHL9sCrS+TycgtZrnXKMmaMgyNTlB2YnxagyYTOwvEPbeuNA62 wR1ARXvokPMW9uwUS1X07wJ7xlx1nFxA1uYaP8kVFHJYarx7zhhrYfUTj zBIGywGc+wMHGqIJp/8aP69cjT1xiILR2BGZWWAS6cNGafHsOEqwlVkhN 7pD2bnfUsw+PPdXVcg2dOGNWmdcBVNBuFN5em4lJwiD1cHAjR0XjBxtjc SIHhq+w3OicNa3B5LAn7p7Jl5Nma37Y+WeZs4hlqLE+ouwOmJUeROF7yH w==; X-CSE-ConnectionGUID: 44K11Xs5RdKoWOw1JO4Vpg== X-CSE-MsgGUID: cwX/fXO4RB2VTOp04iBx/g== X-IronPort-AV: E=McAfee;i="6800,10657,11786"; a="82303115" X-IronPort-AV: E=Sophos;i="6.23,235,1770624000"; d="scan'208";a="82303115" Received: from fmviesa010.fm.intel.com ([10.60.135.150]) by fmvoesa107.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 14 May 2026 11:24:08 -0700 X-CSE-ConnectionGUID: 0I9GDR3+R8KfBBBlYIGyzA== X-CSE-MsgGUID: YGz0wuRGQiW8GZy5qBK/Bg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.23,235,1770624000"; d="scan'208";a="234181032" Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by fmviesa010.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 14 May 2026 11:24:08 -0700 From: Ricardo Neri Subject: [PATCH v3 0/4] sched: Fix cluster scheduling in the presence of asymmetric capacity Date: Thu, 14 May 2026 11:34:36 -0700 Message-Id: <20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-B4-Tracking: v=1; b=H4sIALwVBmoC/3XOwQ6CMAwG4FcxOzuyFcbUk+9hPLBSZAkOs8GCI by7g4ST4fg3/b92ZoG8pcBup5l5ijbY3qWQn08M28q9iNs6ZQYClChBcO9SgTd24lgFjt0YBvK BG1PARdeSpAKWyh9PaWeDH8+UWxuG3n+3O1Gu053UR2SUXHAJsmmMEsagvnfWjVNm3UBdhv2br XCEHStFAddDDBKGtU4/Uq4Qy39sWZYfxskK1w8BAAA= To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , Christian Loehle , Barry Song Cc: "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1778783711; l=5354; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=cPPaIml3mystyoVfu07RSFsEBmIHYsEQUNd6cORguUw=; b=oiZ2dunacx5Ht98DL/sf+4LO4qjFtIr0qPAOhk3IWUvDQFfJ6jJ5nbJvS4PxCUMnyOXemTMuT hZLNO69ZxC6C2+mXa+yuSmK5JGPN+KiwbcuE3za6QZ5aaVQewLDtkfs X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= Hi, This is v3 of the series. It has a few but important changes. Please refer to the change log. Cluster scheduling aims to maximize performance by spreading load across clusters of CPUs that share mid-level resources [1]. It works well on uniform systems, but it breaks down on topologies with big and small cores arranged in clusters. As a result, it fails on several generations of Intel processors already shipped and upcoming. Consider the topology below of big (B) cores and clusters of small (s) cores. ------ ------ | B | | B | ----------------- ----------------- | | | | | s | s | s | s | | s | s | s | s | ------ ------ ----------------- ----------------- | L2 | | L2 | | L2 | | L2 | ------------------------------------------------------- | L3 | ------------------------------------------------------- On a partially busy system (one with idle CPUs; busy CPUs have one task each), scheduling for asymmetric capacity ensures that misfit tasks land on the big CPUs. The remaining tasks, misfit or not, run on the small CPUs. When CONFIG_SCHED_CLUSTER is enabled, these remaining tasks are supposed to be evenly spread among the small-CPU clusters. Today, this does not happen. Several issues in the load balancer prevent a small CPU in one cluster from pulling tasks from another: a) update_sd_pick_busiest() may select a fully_busy group with higher per-CPU capacity as the busiest, preventing a subsequent fully_busy group of equal capacity from being correctly selected. b) Misfit-load statistics are used to identify tasks that would benefit from migrating to bigger CPUs. Accounting misfit load is pointless if the destination CPU is equally small, and it also blocks balancing between clusters. c) Due to b), groups that are truly has_spare or fully_busy get misclassified as misfit_task. update_sd_pick_busiest() then skips them, since a small destination CPU cannot help with misfit tasks. d) Once a busiest group has been identified, sched_balance_find_src_rq() will refuse to migrate tasks to CPUs of equal capacity, even when doing so is precisely what is required to balance small-CPU clusters. e) The SD_PREFER_SIBLING flag is missing from scheduling domains with asymmetric capacity, preventing the balancer from equalizing load across sibling small-core clusters. Together, these issues prevent cluster-level balancing on systems with asymmetric CPU capacity. This series addresses each problem and restores the intended behavior. Details, rationale, and code changes are explained in each patch. I tested these patches on Alder Lake (with Hyper-Threading disabled), Lunar Lake and Panther Lake. I also tested configurations with only one CPU online per cluster to ensure that systems without cluster topology continue to behave as expected. Link: https://lore.kernel.org/r/20210924085104.44806-1-21cnbao@gmail.com/ [1] Changes in v3: - Patch 3: Reverted the inverted runtime capacity check. The inverted form resulted in migrations to CPUs of slightly lower capacity. Guarded the check for architectural capacity with the sched_cluster_active static key. - Patch 4: Expanded the patch description to explain the behavior of overloaded groups and low-capacity clusters with spare capacity. - Added Reviewed-by tags from Christian. Thanks! - Link to v2: https://lore.kernel.org/r/20260429-rneri-fix-cas-clusters-v2-0-cd787de35cc6@linux.intel.com Changes in v2: - Patch 1: Rewrote patch description for clarity. Added a note clarifying that SD_ASYM_CPUCAPACITY and SMT are mutually exclusive. (Tim) - Patch 2: Fixed a bug where the capacity check inadvertently broke the mutual exclusion of the sched_reduced_capacity() path. Keep marking the root domain as overloaded when misfit tasks are present to allow bigger CPUs to help via newly idle balance. (sashiko) Fixed the description to state that capacity_greater() looks for differences of ~5% or more, not 20%. (Christian) - Patch 3: Use arch_scale_cpu_capacity() instead of capacity_of() to ignore runtime capacity variability. Inverted the capacity check. (Christian) - Patch 4: Reworded the patch description for clarity. - Link to v1: https://lore.kernel.org/r/20260330-rneri-fix-cas-clusters-v1-0-1e465b6fecb2@linux.intel.com/ --- Ricardo Neri (4): sched/fair: Check CPU capacity before comparing group types during load balance sched/fair: Skip misfit load accounting when the destination CPU cannot help sched/fair: Allow load balancing between CPUs of identical capacity sched/topology: Do not clear SD_PREFER_SIBLING in domains with clusters include/linux/sched/sd_flags.h | 3 ++- kernel/sched/fair.c | 51 ++++++++++++++++++++++++++++++------------ kernel/sched/topology.c | 14 ++++++++++-- 3 files changed, 51 insertions(+), 17 deletions(-) --- base-commit: 4450349dc665603eb4fab0fb31b7df5b55d6af9b change-id: 20250620-rneri-fix-cas-clusters-bb4287d1e152 Best regards, -- Ricardo Neri