From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.13]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 39AE739FCB9 for ; Thu, 14 May 2026 18:24:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.13 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778783053; cv=none; b=Bvm/8Kjf1sg0uoeRBKPeOowndIbpUB4ja34w9A+8UIIklE2UcmY5lHQk5YTZ3aecQ30svAXJ5Dt67STwySVk/JBESmpwXtqnZOTNMkQFR8ual888SOz7I6jhRzMgIg/+UVAOaXXUku7i4tszduGWAneIdP0TTc+fxAolVIgRrQg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1778783053; c=relaxed/simple; bh=DJP1Odc4BgGcB8RgqGVoQ/OTqKr+5tC74UGMCU1Z/Wg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=J9as5z1RXA3o7BkLlLyuUu9zVpd1PElftLUmsRaJeMdkHT6dBAOt6i3K2BEBn1xQsgOCyOzJP4ESHhHDIfmLyh4/Ew6w3OZGwfqwF+exr+BcxwCG7thsbpPeR1d2sqqanyKDSpPmfIvHOWCOoGL3Tqqh+8fWgjREqLMvJH2waGc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=Lr5d09ir; arc=none smtp.client-ip=192.198.163.13 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="Lr5d09ir" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1778783052; x=1810319052; h=from:date:subject:mime-version:content-transfer-encoding: message-id:references:in-reply-to:to:cc; bh=DJP1Odc4BgGcB8RgqGVoQ/OTqKr+5tC74UGMCU1Z/Wg=; b=Lr5d09irvtL3AGkkgQOoViyB8R0DzFmA/ZQpJM7qiCE7mVvx98MkP1UI bxNDS1rcIivMGps4iCOD67S/09VingfPy3XXBreECmfYiKfuvi52WgQWh GpJvabVMMPFK8D5Gk2FbpuL7nfyeZYYaubfc9MphffjrUFArF3ZGcTE3o OOZChHxc4JK22cN1mheqPeVMn2RI3EMcfBPG3/Z4nJ8Ei/wkr/j6zg7w2 syAD9kDURECfMXvm+UCL4P3rttTdyLQMW8wofs9WZblmxTgORyp5fdRQ0 LMctqmVV/LPtU0uQc732Yqi/0k0/GRikjSxFN8T+regzsQeUhTOVG3KQA g==; X-CSE-ConnectionGUID: 97RTkM01TCSFgvwV9ZR0Hg== X-CSE-MsgGUID: ZR8kimzYSpa9RQjWkIlSSg== X-IronPort-AV: E=McAfee;i="6800,10657,11786"; a="82303124" X-IronPort-AV: E=Sophos;i="6.23,235,1770624000"; d="scan'208";a="82303124" Received: from fmviesa010.fm.intel.com ([10.60.135.150]) by fmvoesa107.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 14 May 2026 11:24:08 -0700 X-CSE-ConnectionGUID: 1zQbOlP8S8mjgI1MjgsEEg== X-CSE-MsgGUID: VIm9HtfUQimCaM3yqdUYQw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.23,235,1770624000"; d="scan'208";a="234181037" Received: from unknown (HELO [172.25.112.21]) ([172.25.112.21]) by fmviesa010.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 14 May 2026 11:24:08 -0700 From: Ricardo Neri Date: Thu, 14 May 2026 11:34:37 -0700 Subject: [PATCH v3 1/4] sched/fair: Check CPU capacity before comparing group types during load balance Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260514-rneri-fix-cas-clusters-v3-1-0037869554bd@linux.intel.com> References: <20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com> In-Reply-To: <20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Tim C Chen , Chen Yu , Christian Loehle , Barry Song Cc: "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org, Ricardo Neri X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1778783711; l=2820; i=ricardo.neri-calderon@linux.intel.com; s=20250602; h=from:subject:message-id; bh=DJP1Odc4BgGcB8RgqGVoQ/OTqKr+5tC74UGMCU1Z/Wg=; b=gFwcWbwoS097m/sRypzq6rC9ldtcrpkYplaOLjkgFlIfr7dNaRxgT8E3VL0vK5OiHQVOqadVT 9RGspL/B900A7p3KorZRD0DlRckpgH9tjWN/gFVweTJ+0dKvR/HkvoF X-Developer-Key: i=ricardo.neri-calderon@linux.intel.com; a=ed25519; pk=NfZw5SyQ2lxVfmNMaMR6KUj3+0OhcwDPyRzFDH9gY2w= update_sd_pick_busiest() may incorrectly select a fully_busy group as the busiest group when its per-CPU capacity exceeds that of the destination CPU. This happens because the type of busiest group is initialized to group_has_spare and allows the fully_busy group to win the type comparison. update_sd_pick_busiest() should not choose a candidate scheduling group with at most one runnable task if its per-CPU capacity is greater than that of the destination CPU. Such a check already exists, but it is done too late: after the type comparison, preventing a subsequent fully_busy group of equal per-CPU capacity from being correctly selected. Move this check to occur before comparing group types. Fixes: 0b0695f2b34a ("sched/fair: Rework load_balance()") Reviewed-by: Christian Loehle Signed-off-by: Ricardo Neri --- Changes in v3: * Added a Fixes tag. (Christian) * Added Reviewed-by tag from Christian. Thanks! Changes in v2: * Added a note clarifying that SMT and SD_ASYM_CPUCAPACITY are mutually exclusive. (Tim) * Kept parentheses around bitwise operators for clarity. * Rewrote patch description for clarity. --- kernel/sched/fair.c | 25 ++++++++++++++----------- 1 file changed, 14 insertions(+), 11 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 3ebec186f982..e06e74d9ce0e 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -10818,6 +10818,20 @@ static bool update_sd_pick_busiest(struct lb_env *env, sds->local_stat.group_type != group_has_spare)) return false; + /* + * Candidate sg has no more than one task per CPU and has higher + * per-CPU capacity. Migrating tasks to less capable CPUs may harm + * throughput. Maximize throughput, power/energy consequences are not + * considered. + * + * Systems with SMT are unaffected, as asymmetric capacity is not set + * in such cases. + */ + if ((env->sd->flags & SD_ASYM_CPUCAPACITY) && + (sgs->group_type <= group_fully_busy) && + (capacity_greater(sg->sgc->min_capacity, capacity_of(env->dst_cpu)))) + return false; + if (sgs->group_type > busiest->group_type) return true; @@ -10920,17 +10934,6 @@ static bool update_sd_pick_busiest(struct lb_env *env, break; } - /* - * Candidate sg has no more than one task per CPU and has higher - * per-CPU capacity. Migrating tasks to less capable CPUs may harm - * throughput. Maximize throughput, power/energy consequences are not - * considered. - */ - if ((env->sd->flags & SD_ASYM_CPUCAPACITY) && - (sgs->group_type <= group_fully_busy) && - (capacity_greater(sg->sgc->min_capacity, capacity_of(env->dst_cpu)))) - return false; - return true; } -- 2.43.0