From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 73C3F29CB24 for ; Mon, 22 Jun 2026 23:55:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782172534; cv=none; b=iZDW5zI+VPtwRqKQ6bJJSBnvJqVxAhWrKiyINUQQC1WcoY5UF+wjJhm/oQO2C97La4ldJIOUkhvkyJeUQTqn0dydhw1LAfpkRoVZlci+Vd8qaYH3Q4fQIPTIMZ69KOauiLEolns4HPU/8Fp1/bVu/oC4+gPzVWgzx78a+eAqlUc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782172534; c=relaxed/simple; bh=5RmXZfNnAa2wyUCXZlqG9P6+qM4FjPEf9s6bt7o6vZ8=; h=From:Subject:Date:Message-Id:MIME-Version:Content-Type:To:Cc; b=R7XEaf3RVTmOs7unNMxIWEJbK5fgMpffxCtFL9Sh0KPwkY3J1F/NP7Afj/AXHI1AMVq1fm1NHsB3y0KjDWzK/5aHBeA+KKKFV5G81PjniM+FflGZ1uoUT757Lk86KA6lk7qXH4Q978f8GkhArc7d3XwBujPAAlNKn9V3qFdVRbc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=n9t0gkBC; arc=none smtp.client-ip=192.198.163.18 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="n9t0gkBC" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1782172532; x=1813708532; h=from:subject:date:message-id:mime-version: content-transfer-encoding:to:cc; bh=5RmXZfNnAa2wyUCXZlqG9P6+qM4FjPEf9s6bt7o6vZ8=; b=n9t0gkBC+7Eh0nYemvKbvE3sJaP8nVQBBQiSqZMwx9eJE1x3N3pR0by5 mppFbZE/dz3Vu7IejbIP+GOXT9Ht3aZDSt1Z8vD4dEkLn+JFkksjyOCT0 8Fs8kuLGJ8o4vVu1Ye7Zkz5AtHOQc3VEHQyOVzZDtzfgPUrHoYTX8kA9O 2o9kQUekpF0qcayXz0BB1gTzObR1yAkVEdo3DobbXB1mWXBIVXcwg+Btq n54sSxlD1SL4SrkfHORkNmnNTuMZYEyDewfOKlSg8WkfSoLDugP8fGsp5 +GqJgjooeWT0S+ExtMOp9ayMTICxD8lzdN5gBLoWUrp//o3Zf++zIMye5 g==; X-CSE-ConnectionGUID: e5oLFGJvTtOOzTVjLhDusQ== X-CSE-MsgGUID: 6WBoEHOIQiWLB/uKW/nIZA== X-IronPort-AV: E=McAfee;i="6800,10657,11825"; a="82014469" X-IronPort-AV: E=Sophos;i="6.24,219,1774335600"; d="scan'208";a="82014469" Received: from fmviesa003.fm.intel.com ([10.60.135.143]) by fmvoesa112.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 22 Jun 2026 16:55:31 -0700 X-CSE-ConnectionGUID: romT3ozTR2WrQ7DAh15MTQ== X-CSE-MsgGUID: w3/i2Ko4ROWycI+JjEBqcg== X-ExtLoop1: 1 Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by fmviesa003.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 22 Jun 2026 16:55:30 -0700 From: Ricardo Neri Subject: [PATCH v5 0/6] sched: Fix cluster scheduling in the presence of asymmetric capacity Date: Mon, 22 Jun 2026 17:05:50 -0700 Message-Id: <20260622-rneri-fix-cas-clusters-v5-0-19968f2d1497@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-B4-Tracking: v=1; b=H4sIAN/NOWoC/3XPS47CMAwG4KugrAmK3TxaVtxjNAviukOkkqKkV CDUu5MiIRZDl79lf7YfInMKnMV+8xCJp5DDEEsw242g0zH+sQxtyQIVGmVRyRTLgOzCTdIxS+q veeSUpfcaa9cCg0FRhi+JS88L/vkt+RTyOKT7a88ES/VNujVyAqkkIHSdN8p7coc+xOttF+LI/ Y6Gs1jgCd+YVRqbVQwLRq0rN3JliOx3rPpgBvQqVhVMqcrVtjFG+/Y7pj+YVfUqppc3DVoHQI3 W9B+b5/kJrde6G6kBAAA= To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , Christian Loehle , K Prateek Nayak , Barry Song Cc: "Rafael J. Wysocki" , Andrea Righi , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri , Vincent Guittot X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1782173185; l=6867; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=5RmXZfNnAa2wyUCXZlqG9P6+qM4FjPEf9s6bt7o6vZ8=; b=nd3/aLomwBJz/CrI7stnNpQ2azPMMTAJQCCxeoOzDUJctR2e85NN1h4OnuFZaibpcxkeI+MqH QxoWfaqRNXAD0P+cEKFlg08loBG3vu5yv4uuKXo5lIiKSF1jz0aR9Z8 X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= Hi, This is v5 of the patch series. The only changes are optimizations to check same-capacity cluster and SMT siblings only when needed. Please read the changelog for details. Cluster scheduling aims to maximize performance by spreading load across clusters of CPUs that share mid-level resources [2]. It works well on uniform systems, but it breaks down on topologies with big and small cores arranged in clusters. As a result, it fails on several generations of Intel processors already shipped and upcoming. Consider the topology below of big (B) cores and clusters of small (s) cores. ------ ------ | B | | B | ----------------- ----------------- | | | | | s | s | s | s | | s | s | s | s | ------ ------ ----------------- ----------------- | L2 | | L2 | | L2 | | L2 | ------------------------------------------------------- | L3 | ------------------------------------------------------- On a partially busy system (one with idle CPUs; busy CPUs have one task each), scheduling for asymmetric capacity ensures that misfit tasks land on the big CPUs. The remaining tasks, misfit or not, run on the small CPUs. When CONFIG_SCHED_CLUSTER is enabled, these remaining tasks are supposed to be evenly spread among the small-CPU clusters. Today, this does not happen. Several issues in the load balancer prevent a small CPU in one cluster from pulling tasks from another: a) update_sd_pick_busiest() may select a fully_busy group with higher per-CPU capacity as the busiest, preventing a subsequent fully_busy group of equal capacity from being correctly selected. b) Misfit-load statistics are used to identify tasks that would benefit from migrating to bigger CPUs. Accounting misfit load is pointless if the destination CPU is equally small, and it also blocks balancing between clusters. c) Due to b), groups that are truly has_spare or fully_busy get misclassified as misfit_task. update_sd_pick_busiest() then skips them, since a small destination CPU cannot help with misfit tasks. d) Once a busiest group has been identified, sched_balance_find_src_rq() will refuse to migrate tasks to CPUs of equal capacity, even when doing so is precisely what is required to balance small-CPU clusters. e) The SD_PREFER_SIBLING flag is missing from scheduling domains with asymmetric capacity, preventing the balancer from equalizing load across sibling small-core clusters. Together, these issues prevent cluster-level balancing on systems with asymmetric CPU capacity. This series addresses each problem and restores the intended behavior. Details, rationale, and code changes are explained in each patch. I tested these patches on Alder Lake, which has both SMT Pcores and clusters of Ecores. I tested with SMT both disabled and enabled. I also tested on Lunar Lake and Panther Lake, which have an Ecore cluster not connected to the L3 cache. I repeated the same experiment with CONFIG_SCHED_CLUSTER disabled. The load balancer behaves as expected. Christian also tested this patchset on a synthetic arm64 qemu topology and the expected behavior [3]. Link: https://lore.kernel.org/all/20260509180955.1840064-1-arighi@nvidia.com/ [1] Link: https://lore.kernel.org/r/20210924085104.44806-1-21cnbao@gmail.com/ [2] Link: https://lore.kernel.org/all/e08492e0-d9f3-4574-8841-b633db008507@arm.com/ [3] Changes in v5: - Added Tested-by tags from Christian. Thanks! - Patch 1 (pre-work): Optimized logic to identify CPUs with busy SMT siblings only when needed. (Prateek, Chen Yu) - Patch 5: Optimized logic to check for architectural capacity only when needed. - Added Reviewed-by tag from Prateek. Thanks! - Link to v4: https://lore.kernel.org/r/20260608-rneri-fix-cas-clusters-v4-0-1526711c944c@linux.intel.com Changes in v4: - Patch 1 (pre-work): Fixed a bug that would block load balancing on SMT cores with more than one busy sibling. - Patch 2 (pre-work): Fixed a bug that would needlessly update sg_overloaded. - Patch 5: Reworked logic using a local variable for improved readability. - Added Reviewed-by tags from Chen Yu, Tim, and Vincent. Thanks! - Link to v3: https://lore.kernel.org/r/20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com Changes in v3: - Patch 3: Reverted the inverted runtime capacity check. The inverted form resulted in migrations to CPUs of slightly lower capacity. Guarded the check for architectural capacity with the sched_cluster_active static key. - Patch 4: Expanded the patch description to explain the behavior of overloaded groups and low-capacity clusters with spare capacity. - Added Reviewed-by tags from Christian. Thanks! - Link to v2: https://lore.kernel.org/r/20260429-rneri-fix-cas-clusters-v2-0-cd787de35cc6@linux.intel.com Changes in v2: - Patch 1: Rewrote patch description for clarity. Added a note clarifying that SD_ASYM_CPUCAPACITY and SMT are mutually exclusive. (Tim) - Patch 2: Fixed a bug where the capacity check inadvertently broke the mutual exclusion of the sched_reduced_capacity() path. Keep marking the root domain as overloaded when misfit tasks are present to allow bigger CPUs to help via newly idle balance. (sashiko) Fixed the description to state that capacity_greater() looks for differences of ~5% or more, not 20%. (Christian) - Patch 3: Use arch_scale_cpu_capacity() instead of capacity_of() to ignore runtime capacity variability. Inverted the capacity check. (Christian) - Patch 4: Reworded the patch description for clarity. - Link to v1: https://lore.kernel.org/r/20260330-rneri-fix-cas-clusters-v1-0-1e465b6fecb2@linux.intel.com/ --- Ricardo Neri (6): sched/fair: Do not skip CPUs of similar capacity with busy SMT siblings sched/fair: Also gate overloaded status update for SD_ASYM_CPUCAPACITY sched/fair: Check CPU capacity before comparing group types during load balance sched/fair: Skip misfit load accounting when the destination CPU cannot help sched/fair: Allow load balancing between CPUs of identical capacity sched/topology: Do not clear SD_PREFER_SIBLING in domains with clusters include/linux/sched/sd_flags.h | 3 +- kernel/sched/fair.c | 66 ++++++++++++++++++++++++++++++------------ kernel/sched/topology.c | 14 +++++++-- 3 files changed, 62 insertions(+), 21 deletions(-) --- base-commit: 50436392fe2359ea108fd27308f86c8283be1622 change-id: 20250620-rneri-fix-cas-clusters-bb4287d1e152 Best regards, -- Ricardo Neri