From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.11]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E391636DA1A for ; Tue, 21 Jul 2026 02:33:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.11 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784601201; cv=none; b=o0IJyxca0aDutfu7D9z2iNPWmhOPDPvxrnVVkk9QUoqWTUT9BSgiFuxyGjvuss6VeV4f8cYdXTP82EqE/sMtoiqkNESDROe9ptFX6TJJbs5E85Fh8YDF5hI3rKWUWPYpHYUYECTR0NcDSDO0D8+ie4VKJkFnsTx41Iuf7O8+xJE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784601201; c=relaxed/simple; bh=bgfWfID7trwrdRoHu3i0XmUP5TNm1g5e89XevH/5CmI=; h=From:Subject:Date:Message-Id:MIME-Version:Content-Type:To:Cc; b=luJWh05Gi2cTZ/s1eKnZLSrJzzZbC77KO2Txq6JBaxMlez1ZCIQE/hbcttO5JF9RhXaPgxUM+G2TgI0gEdxLc8hWsiQTzKMPoKXa6xhYNRYI7VXUkKhvbEhsMxfTzK4E5HIKg1VYmblXCM6K1cGtshoLL/WAhNzYXBUHQTlRDBY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=A4WAcNlq; arc=none smtp.client-ip=192.198.163.11 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="A4WAcNlq" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1784601199; x=1816137199; h=from:subject:date:message-id:mime-version: content-transfer-encoding:to:cc; bh=bgfWfID7trwrdRoHu3i0XmUP5TNm1g5e89XevH/5CmI=; b=A4WAcNlq/yulPs84cJLKctuxYNMqTVXVC5OF/HJ+3Svsr6DgrQOtZYF7 lkqkcPNUFzEYhToGp8OD4pEGLo77rRe3LJxlUwzneVmbQgMBeY1zBXPBJ BFtKqWRBB80SyPVV8M2sioPtUMYk+uC87yg0HMuw0DzwbSYnx/I2UoPi5 OZNM6KFgDP0BR+6rxH/c6HAbUCYtygqOVrq1+MtbtMQYD5jO711gH9s5g q5bxNfEzsGqg4ABKyux7fAcm5mkdleC4gWxEBLFBpv+L4TdLrR+X/FzQ2 em/P3jf/4gN/JTb0uDrNYpkU2ldaIK20GMW0UPA73eLwP9O4k9mUuUvH0 g==; X-CSE-ConnectionGUID: jp2x1pZRSqmGtawF8FVZIg== X-CSE-MsgGUID: fzAekMIST+6URSk7Ct5D4Q== X-IronPort-AV: E=McAfee;i="6800,10657,11852"; a="95785673" X-IronPort-AV: E=Sophos;i="6.25,175,1779174000"; d="scan'208";a="95785673" Received: from orviesa007.jf.intel.com ([10.64.159.147]) by fmvoesa105.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 20 Jul 2026 19:33:18 -0700 X-CSE-ConnectionGUID: 9e69uIFPRM2hb/9U5XBirQ== X-CSE-MsgGUID: /2BssE1WQ4yhNf4TjSpo7g== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,175,1779174000"; d="scan'208";a="257682253" Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by orviesa007.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 20 Jul 2026 19:33:17 -0700 From: Ricardo Neri Subject: [PATCH v6 0/6] sched: Fix cluster scheduling in the presence of asymmetric capacity Date: Mon, 20 Jul 2026 19:43:16 -0700 Message-Id: <20260720-rneri-fix-cas-clusters-v6-0-bb500bf4afd4@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-B4-Tracking: v=1; b=H4sIAMXcXmoC/3XQzW7DIAwH8FepOI8KOxjCTnuPaYfykRUpIxWkU acq7z5SqeoO4fi38M/Yd1ZCjqGw98Od5bDEEqdUg3o7MHc+pe/Ao6+ZoUASCgXPqTbwId64OxX uxmuZQy7cWom99hCAkNXmSw71zQP+/Kr5HMs85d/HnAW26pPULXIBLjggDIMlYa3TH2NM19sxp jmMRzf9sA1e8IkpIdE0MayY87r+MXTknNrHuhdGIJtYVzEhOt0rQySt38fkC1Oib2JyW5NQaQB npHT7GP3DEJsYbZgxqh/QgzQ7N1vX9Q8My5Si9gEAAA== To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , Christian Loehle , K Prateek Nayak , Andrea Righi , Barry Song Cc: "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784601848; l=8087; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=bgfWfID7trwrdRoHu3i0XmUP5TNm1g5e89XevH/5CmI=; b=vV9ZEnPLj5Ydqh3H/HmSTA0EMge07kjT8PJi5EQDp6/e49PKrgSR4IKFwJZKVPJCWjWiI7tXk GVfOR7UbjHiCOQj6UNPhRvgrCydqLZldwZGLH/B5Nh0S/zAK9TxHs2n X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= Hi, This is v6 of this series. The main change is restoring the SD_PREFER_SIBLING flag to scheduling domains with asymmetric capacity, not only those with child cluster domains. I also replaced arch_scale_cpu_capacity() with get_actual_cpu_capacity() to account for hardware and cpufreq pressure when identifying equal-capacity clusters, as Vincent suggested. Cluster scheduling aims to maximize performance by spreading load across clusters of CPUs that share mid-level resources [1]. It works well on uniform systems, but it breaks down on topologies with big and small cores arranged in clusters. As a result, it fails on several generations of Intel processors already shipped and upcoming. Consider the topology below of big (B) cores and clusters of small (s) cores. ------ ------ | B | | B | ----------------- ----------------- | | | | | s | s | s | s | | s | s | s | s | ------ ------ ----------------- ----------------- | L2 | | L2 | | L2 | | L2 | ------------------------------------------------------- | L3 | ------------------------------------------------------- On a partially busy system (one with idle CPUs; busy CPUs have one task each), scheduling for asymmetric capacity ensures that misfit tasks land on the big CPUs. The remaining tasks, misfit or not, run on the small CPUs. When CONFIG_SCHED_CLUSTER is enabled, these remaining tasks are supposed to be evenly spread among the small-CPU clusters. Today, this does not happen. Several issues in the load balancer prevent a small CPU in one cluster from pulling tasks from another: a) update_sd_pick_busiest() may select a fully_busy group with higher per-CPU capacity as the busiest, preventing a subsequent fully_busy group of equal capacity from being correctly selected. b) Misfit-load statistics are used to identify tasks that would benefit from migrating to bigger CPUs. Accounting misfit load is pointless if the destination CPU is equally small, and it also blocks balancing between clusters. c) Due to b), groups that are truly has_spare or fully_busy get misclassified as misfit_task. update_sd_pick_busiest() then skips them, since a small destination CPU cannot help with misfit tasks. d) Once a busiest group has been identified, sched_balance_find_src_rq() will refuse to migrate tasks to CPUs of equal capacity, even when doing so is precisely what is required to balance small-CPU clusters. e) The SD_PREFER_SIBLING flag is missing from scheduling domains with asymmetric capacity, preventing the balancer from equalizing load across sibling small-core clusters. Together, these issues prevent cluster-level balancing on systems with asymmetric CPU capacity. This series addresses each problem and restores the intended behavior. Details, rationale, and code changes are explained in each patch. I tested these patches on Alder Lake, which has both SMT Pcores and clusters of Ecores. I tested with SMT both disabled and enabled. I also tested on Lunar Lake and Panther Lake, which have an Ecore cluster not connected to the L3 cache. I repeated the same experiment with CONFIG_SCHED_CLUSTER disabled. The load balancer behaves as expected. I also tested the series using a patch from Rafael [2] to use policy->cpuinfo.max_freq as fallback for cpufreq_pressure when arch_scale_freq_ref() is not defined. Tasks duly spread and consolidate in the absence/presence of cpufreq pressure. Christian also tested this patchset on a synthetic arm64 qemu topology and observed the expected behavior [3]. Andrea tested this patchset on Vera Rubin (arm64) and found no regressions [4]. Link: https://lore.kernel.org/r/20210924085104.44806-1-21cnbao@gmail.com/ [1] Link: https://lore.kernel.org/all/5086499.GXAFRqVoOG@rafael.j.wysocki/ [2] Link: https://lore.kernel.org/all/e08492e0-d9f3-4574-8841-b633db008507@arm.com/ [3] Link: https://lore.kernel.org/all/akJu2S8SNgp1IaqH@gpd4/[4] Changes in v6: - Patch 6: Restored the SD_PREFER_SIBLING flag to all scheduling domains with asymmetric capacity, not only those with child domains with clusters. (Vincent) - Patch 5: Used get_actual_cpu_capacity() instead of arch_scale_cpu_capacity() to identify clusters of equal capacity. (Vincent) - Patch 5: Renamed a local variable in sched_balance_find_src_rq() for improved readability. (Andrea) - Added Reviewed-by tags from Vincent. Thanks! - Added Tested-by tags from Andrea. Thanks! - Link to v5: https://lore.kernel.org/r/20260622-rneri-fix-cas-clusters-v5-0-19968f2d1497@linux.intel.com Changes in v5: - Added Tested-by tags from Christian. Thanks! - Patch 1 (pre-work): Optimized logic to identify CPUs with busy SMT siblings only when needed. (Prateek, Chen Yu) - Patch 5: Optimized logic to check for architectural capacity only when needed. - Added Reviewed-by tag from Prateek. Thanks! - Link to v4: https://lore.kernel.org/r/20260608-rneri-fix-cas-clusters-v4-0-1526711c944c@linux.intel.com Changes in v4: - Patch 1 (pre-work): Fixed a bug that would block load balancing on SMT cores with more than one busy sibling. - Patch 2 (pre-work): Fixed a bug that would needlessly update sg_overloaded. - Patch 5: Reworked logic using a local variable for improved readability. - Added Reviewed-by tags from Chen Yu, Tim, and Vincent. Thanks! - Link to v3: https://lore.kernel.org/r/20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com Changes in v3: - Patch 3: Reverted the inverted runtime capacity check. The inverted form resulted in migrations to CPUs of slightly lower capacity. Guarded the check for architectural capacity with the sched_cluster_active static key. - Patch 4: Expanded the patch description to explain the behavior of overloaded groups and low-capacity clusters with spare capacity. - Added Reviewed-by tags from Christian. Thanks! - Link to v2: https://lore.kernel.org/r/20260429-rneri-fix-cas-clusters-v2-0-cd787de35cc6@linux.intel.com Changes in v2: - Patch 1: Rewrote patch description for clarity. Added a note clarifying that SD_ASYM_CPUCAPACITY and SMT are mutually exclusive. (Tim) - Patch 2: Fixed a bug where the capacity check inadvertently broke the mutual exclusion of the sched_reduced_capacity() path. Keep marking the root domain as overloaded when misfit tasks are present to allow bigger CPUs to help via newly idle balance. (sashiko) Fixed the description to state that capacity_greater() looks for differences of ~5% or more, not 20%. (Christian) - Patch 3: Use arch_scale_cpu_capacity() instead of capacity_of() to ignore runtime capacity variability. Inverted the capacity check. (Christian) - Patch 4: Reworded the patch description for clarity. - Link to v1: https://lore.kernel.org/r/20260330-rneri-fix-cas-clusters-v1-0-1e465b6fecb2@linux.intel.com/ --- Ricardo Neri (6): sched/fair: Do not skip CPUs of similar capacity with busy SMT siblings sched/fair: Also gate overloaded status update for SD_ASYM_CPUCAPACITY sched/fair: Check CPU capacity before comparing group types during load balance sched/fair: Skip misfit load accounting when the destination CPU cannot help sched/fair: Allow load balancing between CPUs of identical capacity sched/topology: Restore SD_PREFER_SIBLING in domains with asymmetric capacity include/linux/sched/sd_flags.h | 3 +- kernel/sched/fair.c | 66 ++++++++++++++++++++++++++++++------------ kernel/sched/topology.c | 4 --- 3 files changed, 49 insertions(+), 24 deletions(-) --- base-commit: 26b6066f005f1c3cb7c23b1800eb4c3c67ed85e0 change-id: 20250620-rneri-fix-cas-clusters-bb4287d1e152 Best regards, -- Ricardo Neri