From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from casper.infradead.org (casper.infradead.org [90.155.50.34]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C9CF530DED1 for ; Wed, 10 Dec 2025 16:30:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=90.155.50.34 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765384252; cv=none; b=IzS3nDzDLsUGkY3/R4oCV3uO97yHr96YuhZz1CBBfIqXKlFdM88rEVVOdCjUmPxWlJb87wVzOQqaqdP0HQiEzfyoexbSuhVnkY8ZT27ZsyXqze6TNmGDS87q/48x1/z1Oh94yStiJ0gfIhfGO6J2gB2Z1bXLsnp6bIs23yDrMz0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765384252; c=relaxed/simple; bh=CLDXB9q1ax7C/hCwWLynnGmRr8CaHJPBGfLKwCvNSWg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=IOC8GBpzcpHoZc08oyP/SvMFdqiOD1fw01vQdXJiA1xseM8unxoGOw1zFbmvE7NjwKZRW7hRDACJXbqvwlc6q87E++1ese/CHc+ETEwMe/oKZqc9MnFhg4xkPfEl45NzEvS4gVzvt0zKTRCL70Cwc7KUQgpERQ6Hj2irzi/Weew= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org; spf=none smtp.mailfrom=infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=FDJ/hDO3; arc=none smtp.client-ip=90.155.50.34 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=infradead.org Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="FDJ/hDO3" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=casper.20170209; h=In-Reply-To:Content-Type:MIME-Version: References:Message-ID:Subject:Cc:To:From:Date:Sender:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description; bh=zt4lwlBQD90pOGMnooKi7X5cvHzGltCl3f/O1ZjDsVA=; b=FDJ/hDO39nIFIa1375qdJ9VpzD GP1iFwNoxxAmjP0HYl5j84rmXImtCgiNtBYSCBIn671phy52aacCKC/03pGRHQGhtuoE5KiVsXHDE mm1b3to6JxrUjpsR795k6KcIKM/eQJk1wtuwTXjjX8qTUNUSUcHOgUfKK3VWCGF4dDpl2VPl1elQ9 CrY4bXXclGLjb2b2uvvaWmAvwnhOABTSa2EiAq8P74XXRH+TqtAEAKm6G3f38/148DL3TmU2mnqb4 LMzIkAzOnVOedNIOpb0r8jCm6G17sJ6KbihwPN9tIIkeLq9zzdCaNMMUUbZbvxmWruebyxEBzIegi CRX6UI2A==; Received: from 77-249-17-252.cable.dynamic.v4.ziggo.nl ([77.249.17.252] helo=noisy.programming.kicks-ass.net) by casper.infradead.org with esmtpsa (Exim 4.98.2 #2 (Red Hat Linux)) id 1vTN4w-0000000D5Aj-46ri; Wed, 10 Dec 2025 16:30:15 +0000 Received: by noisy.programming.kicks-ass.net (Postfix, from userid 1000) id 81DF0300566; Wed, 10 Dec 2025 17:30:13 +0100 (CET) Date: Wed, 10 Dec 2025 17:30:13 +0100 From: Peter Zijlstra To: Tim Chen Cc: Ingo Molnar , K Prateek Nayak , "Gautham R . Shenoy" , Vincent Guittot , Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org Subject: Re: [PATCH v2 15/23] sched/cache: Respect LLC preference in task migration and detach Message-ID: <20251210163013.GW3707891@noisy.programming.kicks-ass.net> References: <1c75f54a2e259737eb9b15c98a5c1d1f142fdef6.1764801860.git.tim.c.chen@linux.intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <1c75f54a2e259737eb9b15c98a5c1d1f142fdef6.1764801860.git.tim.c.chen@linux.intel.com> On Wed, Dec 03, 2025 at 03:07:34PM -0800, Tim Chen wrote: > @@ -10025,6 +10025,13 @@ int can_migrate_task(struct task_struct *p, struct lb_env *env) > if (env->flags & LBF_ACTIVE_LB) > return 1; > > +#ifdef CONFIG_SCHED_CACHE > + if (sched_cache_enabled() && > + can_migrate_llc_task(env->src_cpu, env->dst_cpu, p) == mig_forbid && > + !task_has_sched_core(p)) > + return 0; > +#endif This seems wrong: - it does not let nr_balance_failed override things; - it takes precedence over migrate_degrade_locality(); you really want to migrate towards the preferred NUMA node over staying on your LLC. That is, this really wants to be done after migrate_degrades_locality() and only if degrades == 0 or something. > degrades = migrate_degrades_locality(p, env); > if (!degrades) > hot = task_hot(p, env); > @@ -10146,12 +10153,55 @@ static struct list_head > list_splice(&pref_old_llc, tasks); > return tasks; > } > + > +static bool stop_migrate_src_rq(struct task_struct *p, > + struct lb_env *env, > + int detached) > +{ > + if (!sched_cache_enabled() || p->preferred_llc == -1 || > + cpus_share_cache(env->src_cpu, env->dst_cpu) || > + env->sd->nr_balance_failed) > + return false; But you are allowing nr_balance_failed to override things here. > + /* > + * Stop migration for the src_rq and pull from a > + * different busy runqueue in the following cases: > + * > + * 1. Trying to migrate task to its preferred > + * LLC, but the chosen task does not prefer dest > + * LLC - case 3 in order_tasks_by_llc(). This violates > + * the goal of migrate_llc_task. However, we should > + * stop detaching only if some tasks have been detached > + * and the imbalance has been mitigated. > + * > + * 2. Don't detach more tasks if the remaining tasks want > + * to stay. We know the remaining tasks all prefer the > + * current LLC, because after order_tasks_by_llc(), the > + * tasks that prefer the current LLC are the least favored > + * candidates to be migrated out. > + */ > + if (env->migration_type == migrate_llc_task && > + detached && llc_id(env->dst_cpu) != p->preferred_llc) > + return true; > + > + if (llc_id(env->src_cpu) == p->preferred_llc) > + return true; > + > + return false; > +} Also, I think we have a problem with nr_balance_failed, cache_nice_tries is 1 for SHARE_LLC; this means for failed=0 we ignore: - ineligible tasks - llc fail - node-degrading / hot and then the very next round, we do all of them at once, without much grading. > @@ -10205,6 +10255,15 @@ static int detach_tasks(struct lb_env *env) > > p = list_last_entry(tasks, struct task_struct, se.group_node); > > + /* > + * Check if detaching current src_rq should be stopped, because > + * doing so would break cache aware load balance. If we stop > + * here, the env->flags has LBF_ALL_PINNED, which would cause > + * the load balance to pull from another busy runqueue. Uhh, can_migrate_task() will clear that ALL_PINNED thing if we've found at least one task before getting here. > + */ > + if (stop_migrate_src_rq(p, env, detached)) > + break; Perhaps split cfs_tasks into multiple lists from the get-go? That avoids this sorting.