From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E59923E3C67 for ; Mon, 30 Mar 2026 22:22:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.9 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774909346; cv=none; b=AMZUCmSihq18Ne5fFXTGg5LRB6+ZV4/V8fnQ+WYQxYtff1EBSgA+ztrKxYUjtnqaidMKSHs56wTpctSTRITN+CZ3olS0/wz1xIv8F7hF0GuP3gqVtkXYck3DxZm+M4T3hxVGVmFsOUFzrOa3Z23XtbOr91LvVaSdZwKTwsSLKtc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774909346; c=relaxed/simple; bh=Ib/qYtkNBuD2ULEkcWydKJN8EeTwxREYetARM8T/PiU=; h=From:Subject:Date:Message-Id:MIME-Version:Content-Type:To:Cc; b=Olk+CZXaBS2K/o/c+6imV1rcBJJmzfJyVDs9XDd/hNXH8RMeSpHkAMKp3mHGmbiTtpeusWdq4mF+fRkFIp9pVppTu1hniDnm2m7PhF+ZuqB4wz5tyoWxCGB5a9LIS/Gqq4heUTQPzCHxxLATwhf/WGFgglmrpM02WSOMKBRRwCw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=lEyAofqo; arc=none smtp.client-ip=192.198.163.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="lEyAofqo" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1774909336; x=1806445336; h=from:subject:date:message-id:mime-version: content-transfer-encoding:to:cc; bh=Ib/qYtkNBuD2ULEkcWydKJN8EeTwxREYetARM8T/PiU=; b=lEyAofqo4tj3JWA8Gzu8mOacPvvmFIS6b1t7w827tYZdbBAfdyLo8nYm ADAQ2hlxxtwi3TNmyB/mIeEYAUZNO/u+PKpsUJbBYQl/S5HLMteoGfqSS xi9w5P8uTeakUEKs6psDZCet4/oqb/msy3uwBGJwZhWKWC+Erhv5VR76S V6WMs4Bo99R9wsab8DgvPs0ZHD5Oembjuhe6TdzSvETj1K1RJfKjP0Xlr HuKNp5qNfLof6v/FrVSXJjfMjaFzgRtQM8MhqTjXRruxryyHRoJHjuSAm NQYUntxOzCd3bUg+l306u9UCBha0dsQ4tuTtOgcmIFqWy7ONACRvJTD6b g==; X-CSE-ConnectionGUID: 0fRkjnTfT5iQLhRzu9bpzQ== X-CSE-MsgGUID: Arl29ve8SuiRjFHkL9Qz1Q== X-IronPort-AV: E=McAfee;i="6800,10657,11744"; a="86606134" X-IronPort-AV: E=Sophos;i="6.23,150,1770624000"; d="scan'208";a="86606134" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa103.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 30 Mar 2026 15:22:14 -0700 X-CSE-ConnectionGUID: yxGQ/1GmTkmGQgNY1qBhWw== X-CSE-MsgGUID: MLxwz2r4R2y3/fIbX4SFSQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.23,150,1770624000"; d="scan'208";a="230242278" Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by orviesa003.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 30 Mar 2026 15:22:13 -0700 From: Ricardo Neri Subject: [PATCH RESEND 0/4] sched: Fix cluster scheduling in the presence of asymmetric capacity Date: Mon, 30 Mar 2026 15:20:34 -0700 Message-Id: <20260330-rneri-fix-cas-clusters-v1-0-1e465b6fecb2@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Barry Song Cc: "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1774909272; l=3148; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=Ib/qYtkNBuD2ULEkcWydKJN8EeTwxREYetARM8T/PiU=; b=SU1OCAkL3nhdIR6TJTsw7vQE9dSlRyC9bDL8/ffuI0Un1BTvYa4DHLwaEFH3JdGQUS4MdakOG xEA3Yn7a5RoCGCVCWAjMd/DDmpuFykis856N/qC4cUQ7ck5YvcGw6Fo X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= Cluster scheduling balances load among clusters of CPUs sharing a resource [1]. It was broken on Intel hybrid processors using asymmetric packing of tasks. Tim fixed that [2]. It is broken again when combined with asymmetric CPU capacity. The diagram below shows a processor with big (B) and small (s) CPUs. Also, small CPUs are grouped in cluster sharing mid-level cache. This topology is common in Intel hybrid processors. ------ ------ | B | | B | ----------------- ----------------- | | | | | s | s | s | s | | s | s | s | s | ------ ------ ----------------- ----------------- | L2 | | L2 | | L2 | | L2 | ------------------------------------------------------- | L3 | ------------------------------------------------------- On a partially busy system (one with idle CPUs; busy CPUs have one task each), scheduling for asymmetric capacity ensures that misfit tasks are placed on the big CPUs. The remaining tasks, misfit or not, run on the small CPUs. If CONFIG_SCHED_CLUSTER is enabled, these remaining tasks should be evenly spread between the two small-CPU clusters. This does not happen today because various checks in the load balancer prevent a small CPU in one cluster from pulling tasks from another: * A bug in update_sd_pick_busiest() causes it to not check for capacity when preferring a fully_busy big CPU (which it cannot help) vs a has_ spare small-CPU cluster (which it can). * Accounting misfit load in a group is pointless if the destination CPU is equally a small CPU. Moreover, update_sd_pick_busiest() will not pick such group as busiest anyway. * Once a busiest group has been identified, sched_balance_find_src_rq() will refuse to migrate tasks to CPUs of equal capacity. * The SD_PREFER_SIBLING flag is removed from scheduling domains with asymmetric capacity. I address these issues in this series. Details are in the changelog of each patch. I tested these patches on an Alder Lake system with Hyper-Threading disabled. I also tested with CONFIG_SCHED_CLUSTER=n to ensure that processors without clusters continue to work. [1]. https://lore.kernel.org/r/20210924085104.44806-1-21cnbao@gmail.com/ [2]. https://lore.kernel.org/r/cover.1688770494.git.tim.c.chen@linux.intel.com/ --- Ricardo Neri (4): sched/fair: Always skip fully_busy higher-capacity groups for load balance sched/fair: Ignore misfit load if the destination CPU cannot help sched/fair: Allow load balancing between CPUs of equal capacity sched/topology: Keep SD_PREFER_SIBLING for domains with clusters kernel/sched/fair.c | 27 +++++++++++++++------------ kernel/sched/topology.c | 11 +++++++++-- 2 files changed, 24 insertions(+), 14 deletions(-) --- base-commit: e51a38e71974982abb3f2f16141763a1511f7a3f change-id: 20250620-rneri-fix-cas-clusters-bb4287d1e152 Best regards, -- Ricardo Neri