From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.21]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 151E2625 for ; Mon, 18 May 2026 00:26:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.21 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779063991; cv=none; b=kr55HaGcUeQodUZplEcA72jg2IujxYlxWM9YR32eiIqqEr5TtdhGRyNbNxN2r2SRiv0weeBqAWr2dQKhUzkY+5uWr4F3ZD85CFNaYpqoYkMzP34l/NKmWNWeaEleTGHWhsnaf7jOoRHW5DVeUyyVwFAKcbHMmoDV8znKSeJ0u/0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779063991; c=relaxed/simple; bh=yDxhBN/FGRiPh+k+zc/S8UNufKEoGdpOF8+8NMWyr3I=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=uVZ17h7OQaf0vYL4OgpO/T9W4LVeWQ/Y/IeBtE1olRHGBXwBoXSOgcwGYUPE0rVFTfwoqcJuubvZL7CleZTLqMglHwzlrvqkPoLHl6N3Sg7YauMY4w7QY0qiKUD3SPcVBKqV8IbgtBx8G2+ctv2wDEe2KeV67pTqYVXowREld3A= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=eJHRkU4V; arc=none smtp.client-ip=198.175.65.21 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="eJHRkU4V" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1779063989; x=1810599989; h=date:from:to:cc:subject:message-id:references: mime-version:in-reply-to; bh=yDxhBN/FGRiPh+k+zc/S8UNufKEoGdpOF8+8NMWyr3I=; b=eJHRkU4Vlnu1nKVC2Ku8rdlGfp64AJjf2nlinpmfd5fRr03hh4uKDePq VEw+NKvfXRbrziPf8RXxnPnkWPHJlwLIiaXeBprEoS7paruvSK1Ok7LIt vMKxLzbyHwRl7BGmx2ytKKUefSxE96HKqHdxfYMGZAx/PWG26V6BKnSNF E8/OI7yzN3rk+gWRR6oViMILYznUcWCqlLpvoiVb2r+gxmpanyWtR+Sp8 6sqnSfEp1cV0nHjSbLqKMZwz7aXslB7G5TgJBbhUcFP2qP7pj5Ckwrnjv dymFDZUd7QjRj8kPZk8vRhMICdGkI/DOpD0irmvSeR/j5XFIiV9lW8V3z A==; X-CSE-ConnectionGUID: by8MNen7SrWSVjNI7eqwCQ== X-CSE-MsgGUID: VWlMs+eGTcGLGIZ2BEDoCQ== X-IronPort-AV: E=McAfee;i="6800,10657,11789"; a="79814687" X-IronPort-AV: E=Sophos;i="6.23,240,1770624000"; d="scan'208";a="79814687" Received: from fmviesa001.fm.intel.com ([10.60.135.141]) by orvoesa113.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 May 2026 17:26:29 -0700 X-CSE-ConnectionGUID: P+Z6P7nPTVGAnLYIYCbXLA== X-CSE-MsgGUID: q6LgbDokR1O87IayJ6GeWw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.23,240,1770624000"; d="scan'208";a="263033074" Received: from ranerica-svr.sc.intel.com ([172.25.110.23]) by fmviesa001.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 17 May 2026 17:26:28 -0700 Date: Sun, 17 May 2026 17:35:03 -0700 From: Ricardo Neri To: Tim Chen Cc: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Chen Yu , Christian Loehle , Barry Song , "Rafael J. Wysocki" , Len Brown , ricardo.neri@intel.com, linux-kernel@vger.kernel.org Subject: Re: [PATCH v3 2/4] sched/fair: Skip misfit load accounting when the destination CPU cannot help Message-ID: <20260518003503.GA3654@ranerica-svr.sc.intel.com> References: <20260514-rneri-fix-cas-clusters-v3-0-0037869554bd@linux.intel.com> <20260514-rneri-fix-cas-clusters-v3-2-0037869554bd@linux.intel.com> <6a2ae247d4a54d4b39c9295a0c1c1c680dba7ecc.camel@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <6a2ae247d4a54d4b39c9295a0c1c1c680dba7ecc.camel@linux.intel.com> User-Agent: Mutt/1.9.4 (2018-02-28) On Fri, May 15, 2026 at 01:12:28PM -0700, Tim Chen wrote: > On Thu, 2026-05-14 at 11:34 -0700, Ricardo Neri wrote: > > In domains with asymmetric capacity, identifying misfit load in a > > scheduling group is not useful when the destination CPU cannot help (i.e., > > its capacity exceeds the group's maximum CPU capacity by less than ~5%). In > > such cases, it also prevents load balance among clusters of equal capacity > > when CONFIG_SCHED_CLUSTER is enabled. This happens because > > update_sd_pick_busiest() skips candidate groups of type misfit_task if the > > destination CPU has similar capacity. > > > > Skipping misfit load accounting in this situation allows the group to be > > classified as has_spare or fully_busy and lets load balancing proceed. Keep > > marking scheduling groups as overloaded when misfit tasks are present. The > > sg_overloaded flag propagates to the root domain and allows bigger CPUs in > > it to help via newly idle balance. > > > > Reviewed-by: Christian Loehle > > Signed-off-by: Ricardo Neri > > --- > > Changes in v3: > > * Added Reviewed-by tag from Christian. Thanks! > > > > Changes in v2: > > * Moved the check of the destination CPU capacity inside the code block > > used for SD_ASYM_CPUCAPACITY. v1 inadvertently broke the mutual > > exclusion of the sched_reduced_capacity() path. > > * Keep marking the root domain as overloaded to allow bigger CPUs to > > help. (sashiko) > > * Fixed patch description to clarify that the capacity_greater() looks > > for differences of 5% or more. (Christian) > > * Reworded the patch description for clarity. > > * I did not include the Reviewed-by tag from Christian since the patch > > changed functionally. > > --- > > kernel/sched/fair.c | 20 +++++++++++++++++--- > > 1 file changed, 17 insertions(+), 3 deletions(-) > > > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > > index e06e74d9ce0e..dcc02ceb44b5 100644 > > --- a/kernel/sched/fair.c > > +++ b/kernel/sched/fair.c > > @@ -10749,10 +10749,24 @@ static inline void update_sg_lb_stats(struct lb_env *env, > > continue; > > > > if (sd_flags & SD_ASYM_CPUCAPACITY) { > > - /* Check for a misfit task on the cpu */ > > - if (sgs->group_misfit_task_load < rq->misfit_task_load) { > > - sgs->group_misfit_task_load = rq->misfit_task_load; > > + if (rq->misfit_task_load) { > > + /* > > + * Always mark the domain overloaded so big CPUs > > + * can pick up misfit tasks via newly idle > > + * balance. > > + */ > > *sg_overloaded = 1; > > + > > + /* > > + * Only account misfit load if @dst_cpu can > > + * help; otherwise, the group may be classified > > + * as misfit_task and update_sd_pick_busiest() > > + * will skip it. > > You mean "sd_pick_busiest() will pick it" instead of "skip it" for misfit task > load balancing in the above comment? Thank you for your review! I mean "skip it" because update_sd_pick_busiest() will skip a candidate group of type misfit if dst_cpu has less than 1.05 times the max capacity of such group. It is the first check in the function. Skipping misfit accounting allows the candidate group to be classified as fully_ busy or has_spare so that tasks can be balanced between clusters of equal capacity. I will rephrase this comment to make it more clear.