From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.11]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 748251E8320 for ; Wed, 10 Dec 2025 18:36:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.11 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765391795; cv=none; b=qbCH0vmbzPz2dsBVOCoUNeyq8jQF9Kpk+okT4oGgo1JSz6Tg4xPhrGE/4nE1CjhIBK2O9VB5mpf1aYzkZ+8BCAdY8Bu9KGTflPOQtom5nUlS50HP7JRgx+ufhKldc/CLkrmZZGCqSAcstZ82sXoQtxqqiFDyhmeG/zL/UB4P3Ew= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765391795; c=relaxed/simple; bh=fi/tVpxQ1hF85QXkPgJH25RBK1W4oOm5VRWL7GnW6ec=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=rH+NtS2D1HoDGpNyIckHKb7GAKF+D21EnsD4LShsExRnSrzutZ3yHPb2JbDKG5D3qs9GDiWxc/+lA6jnnjgQMESA4qgY/Sk2xxqQwVYdkxJmbJ4JS01fdWjyBy/CuOltQRT4LjjdQTVrSXGaaReUsN6RyImCM5KfKYJxNAWELQM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=aa3fpvcH; arc=none smtp.client-ip=192.198.163.11 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="aa3fpvcH" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1765391793; x=1796927793; h=message-id:subject:from:to:cc:date:in-reply-to: references:content-transfer-encoding:mime-version; bh=fi/tVpxQ1hF85QXkPgJH25RBK1W4oOm5VRWL7GnW6ec=; b=aa3fpvcH/7Ci5692kYOakICkyw1/dVZTqCOeupSpF8GqYJi1aSOzKizu 1tAS1xO3aZ7jAqUVcXDKZc8Xj6xn/NZwW5v8ex527v7iiIcze3dIm99W7 K05qtvvdOl3gRP4dv3ljeXwccN42opbvgVE0IE6Kz2+tZNOlWoMN48dy4 X+LUR0b3JhC6omMP7p9MmzTOd6gubePGg05UXKa6HnlEM8ZViZIo6mpLx Vz7otJg8+loXCKLxk1uTcURYSGs48rCMG6b9NtwQSSv+5snJ1XLbt49GA 7v7J6wO2NgId8bUsFBUuGX5J7BnPsy6mUeWBD325WDyPIOjsv5pLEhU2i w==; X-CSE-ConnectionGUID: gx23FSDiTOSKnSlGza2cng== X-CSE-MsgGUID: 6REMULgXQk6XCeqqDhzW7g== X-IronPort-AV: E=McAfee;i="6800,10657,11638"; a="77988782" X-IronPort-AV: E=Sophos;i="6.20,264,1758610800"; d="scan'208";a="77988782" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa105.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 10 Dec 2025 10:36:32 -0800 X-CSE-ConnectionGUID: OaDlDk64QBC9izGKirEVUw== X-CSE-MsgGUID: ALdisrCcTD2RqECWBs2zgA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.20,264,1758610800"; d="scan'208";a="200755466" Received: from unknown (HELO [10.241.243.18]) ([10.241.243.18]) by ORVIESA003-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 10 Dec 2025 10:36:31 -0800 Message-ID: <3c3cc30f931a61eda1aed056abc03b0839291781.camel@linux.intel.com> Subject: Re: [PATCH v2 07/23] sched/cache: Introduce per runqueue task LLC preference counter From: Tim Chen To: Peter Zijlstra Cc: Ingo Molnar , K Prateek Nayak , "Gautham R . Shenoy" , Vincent Guittot , Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Madadi Vineeth Reddy , Hillf Danton , Shrikanth Hegde , Jianyong Wu , Yangyu Chen , Tingyin Duan , Vern Hao , Vern Hao , Len Brown , Aubrey Li , Zhao Liu , Chen Yu , Chen Yu , Adam Li , Aaron Lu , Tim Chen , linux-kernel@vger.kernel.org Date: Wed, 10 Dec 2025 10:36:30 -0800 In-Reply-To: <20251210124303.GR3707891@noisy.programming.kicks-ass.net> References: <63091f7ca7bb473fbc176af86a87d27a07a6e149.1764801860.git.tim.c.chen@linux.intel.com> <20251210124303.GR3707891@noisy.programming.kicks-ass.net> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.58.1 (3.58.1-1.fc43) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Wed, 2025-12-10 at 13:43 +0100, Peter Zijlstra wrote: > On Wed, Dec 03, 2025 at 03:07:26PM -0800, Tim Chen wrote: >=20 > > +static int resize_llc_pref(void) > > +{ > > + unsigned int *__percpu *tmp_llc_pref; > > + int i, ret =3D 0; > > + > > + if (new_max_llcs <=3D max_llcs) > > + return 0; > > + > > + /* > > + * Allocate temp percpu pointer for old llc_pref, > > + * which will be released after switching to the > > + * new buffer. > > + */ > > + tmp_llc_pref =3D alloc_percpu_noprof(unsigned int *); > > + if (!tmp_llc_pref) > > + return -ENOMEM; > > + > > + for_each_present_cpu(i) > > + *per_cpu_ptr(tmp_llc_pref, i) =3D NULL; > > + > > + /* > > + * Resize the per rq nr_pref_llc buffer and > > + * switch to this new buffer. > > + */ > > + for_each_present_cpu(i) { > > + struct rq_flags rf; > > + unsigned int *new; > > + struct rq *rq; > > + > > + rq =3D cpu_rq(i); > > + new =3D alloc_new_pref_llcs(rq->nr_pref_llc, per_cpu_ptr(tmp_llc_pre= f, i)); > > + if (!new) { > > + ret =3D -ENOMEM; > > + > > + goto release_old; > > + } > > + > > + /* > > + * Locking rq ensures that rq->nr_pref_llc values > > + * don't change with new task enqueue/dequeue > > + * when we repopulate the newly enlarged array. > > + */ > > + rq_lock_irqsave(rq, &rf); > > + populate_new_pref_llcs(rq->nr_pref_llc, new); > > + rq->nr_pref_llc =3D new; > > + rq_unlock_irqrestore(rq, &rf); > > + } > > + > > +release_old: > > + /* > > + * Load balance is done under rcu_lock. > > + * Wait for load balance before and during resizing to > > + * be done. They may refer to old nr_pref_llc[] > > + * that hasn't been resized. > > + */ > > + synchronize_rcu(); > > + for_each_present_cpu(i) > > + kfree(*per_cpu_ptr(tmp_llc_pref, i)); > > + > > + free_percpu(tmp_llc_pref); > > + > > + /* succeed and update */ > > + if (!ret) > > + max_llcs =3D new_max_llcs; > > + > > + return ret; > > +} >=20 > > @@ -2674,6 +2787,8 @@ build_sched_domains(const struct cpumask *cpu_map= , struct sched_domain_attr *att > > if (has_cluster) > > static_branch_inc_cpuslocked(&sched_cluster_active); > > =20 > > + resize_llc_pref(); > > + > > if (rq && sched_debug_verbose) > > pr_info("root domain span: %*pbl\n", cpumask_pr_args(cpu_map)); >=20 > I suspect people will hate on you for that synchronize_rcu() in there. >=20 > Specifically, we do build_sched_domain() for every CPU brought online, > this means booting 512 CPUs now includes 512 sync_rcu()s. > Worse, IIRC sync_rcu() is O(n) (or worse -- could be n*ln(n)) in number > of CPUs, so the total thing will be O(n^2) (or worse) for bringing CPUs > online. >=20 >=20 Though we only do sychronize_rcu in resize_llc_pref() when we encounter a n= ew LLC,=C2=A0 and need a larger array of LLCs, and not on every CPU. That said, I agree that free is better done in a RCU call back to avoid scynchronize_rcu overhead. Tim