From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.13]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A08FB33438F for ; Thu, 30 Jul 2026 02:22:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.13 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785378126; cv=none; b=RX5M9f9IpiWgmstAkFtzidUHtgMb3L4U+12ys3ST54HTrmphTRgD3MXFch1xpxbw1zGMqDo5kNmNMxIjPELqhh9odeq44rCoXKqL0TCt2zmcBdOq0cPCka0H7ao7pAuzbjmA4GUVxp402JkFjRLhTSTMmOPxEGqHQb0t2Q9lS2U= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785378126; c=relaxed/simple; bh=TjgNBwfCwr0iveDU+FS4QHpJR+GQGSB/RQ36XzrUH+c=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=QLbcnmdTDxQQ41nql8AbKck1Q7Q9nP4btHf/NvKKAx+4w4OQJFAe054w7rL9scKAFYeM6pFLFCqSxYOzLQ9+R1XW+JCfe5s2qPC5XigJ+IHaQB9ba4rTLV6dpaNNOu2iTEynt7adDB0JING5vmkmwCev/otPJbpohlKFuKoM9G0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=cSfn/Y52; arc=none smtp.client-ip=198.175.65.13 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="cSfn/Y52" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785378123; x=1816914123; h=message-id:subject:from:to:cc:date:in-reply-to: references:content-transfer-encoding:mime-version; bh=TjgNBwfCwr0iveDU+FS4QHpJR+GQGSB/RQ36XzrUH+c=; b=cSfn/Y5201QI/jUA4Lcthfs+Xh1M09gsNjfKtDd0NX82S3UewhSMrHRb JzSUaDUm7aRa5A8MhM6qi2ZlE3PpcgjbZTJsZG1xsS+WowQs2YAJLaP7S 0MsN3lu/HIkDNl/CVHcS9f5cJlIsldE5HZ18iTUGeZ1XyfKxEE4dXyo/H 4U/+6gZ68I69bHt2NPfm7mlVs1mTcZUsO2ELYHh2UP7rSo9q9XH2exm6N bPOCgFWooGphma6CKLdUVRDcqiu6n2oPRv4i3AQ9hjg4O/8y0WZRS0D/7 YGQkuhqBydQyabnSmE1bjW1ozKEhpsTQRJy5dVFYbS0kQbeY9mJZcgE5E w==; X-CSE-ConnectionGUID: N6M3NfBTTTu797Xqs1zd6w== X-CSE-MsgGUID: oXcOTyTNRwqsn+fcDVze9Q== X-IronPort-AV: E=McAfee;i="6800,10657,11859"; a="97148795" X-IronPort-AV: E=Sophos;i="6.25,193,1779174000"; d="scan'208";a="97148795" Received: from fmviesa006.fm.intel.com ([10.60.135.146]) by orvoesa105.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 29 Jul 2026 19:22:03 -0700 X-CSE-ConnectionGUID: ywcIMyviQUOOznZ+zBc3HQ== X-CSE-MsgGUID: GxwgqGY6RGK62FmrbvwhtQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,193,1779174000"; d="scan'208";a="255858858" Received: from schen9-mobl4.amr.corp.intel.com (HELO [10.125.111.216]) ([10.125.111.216]) by fmviesa006-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 29 Jul 2026 19:22:02 -0700 Message-ID: <2d191aead79140de43022f480c3b542c101613f1.camel@linux.intel.com> Subject: Re: : [PATCH v8 1/2] sched/cache: Reduce the overhead of task_cache_work by only scan the visisted cpus From: Tim Chen To: Luo Gengkun , Chen Yu Cc: "Chen, Yu C" , dietmar.eggemann@arm.com, rostedt@goodmis.org, peterz@infradead.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, linux-kernel@vger.kernel.org, mingo@redhat.com, juri.lelli@redhat.com, vincent.guittot@linaro.org Date: Wed, 29 Jul 2026 19:22:01 -0700 In-Reply-To: <7d1ae7709bbe052f1d5737b6f0ef6cddf2f5e541.camel@linux.intel.com> References: <20260723040429.630176-1-luogengkun2@huawei.com> <20260723040429.630176-2-luogengkun2@huawei.com> <1dc03c84-9bc2-4db1-bab4-3f603fba54cd@huawei.com> <8b7cd508-3b89-4fc0-85fd-5a8d35ed06ca@huawei.com> <7d1ae7709bbe052f1d5737b6f0ef6cddf2f5e541.camel@linux.intel.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.58.1 (3.58.1-1.fc43) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Wed, 2026-07-29 at 11:29 -0700, Tim Chen wrote: > On Wed, 2026-07-29 at 17:19 +0800, Luo Gengkun wrote: > >=20 > > On 2026/7/28 20:08, Chen Yu wrote: > > > On Tue, Jul 28, 2026 at 04:53:59PM +0800, Luo Gengkun wrote: > > > >=20 > > > > On 2026/7/27 9:09, Chen, Yu C wrote: > > > > > On 7/23/2026 12:04 PM, Luo Gengkun wrote: > > > > >=20 > > > > > [ ... ] > > > > >=20 > > > > > > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 guard(raw_spinlock_irqsave)(&rq= ->cpu_epoch_lock); > > > > > > =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 __update_mm_sched(rq, pcpu_sche= d); > > > > > > +=C2=A0=C2=A0=C2=A0 /* Skip the rq that has not been hit for a = long time */ > > > > > > +=C2=A0=C2=A0=C2=A0 if ((rq->cpu_epoch - pcpu_sched->epoch_last= _visit) > llc_epoch_affinity_timeout) { > > > > >=20 > > > > > In v2 there is a check if the cpu has been set before writing: > > > > > cpumask_test_cpu(cpu_of(rq), &mm->sc_stat.visited_cpus) > > > > > https://lore.kernel.org/all/20260414150745.225416-1-luogengkun2@h= uawei.com/ > > > > > do we need to bring that back? > > > > >=20 > > > > I don't think we need it back. Here is why: > > > >=20 > > > > In v2, for_each_cpu was used instead of for_each_cpu_and in the inn= er loop, > > > > meaning some CPUs being checked might not have been set. Therefore, > > > > cpumask_test_cpu was necessary to filter out those cases. > > > >=20 > > > > Now, with for_each_cpu_and(i, sched_domain_span(sd), &mm->sc_stat.v= isited_cpus), > > > > we can ensure each scanned CPU is set, so the issue no longer exist= s. > > > > Furthermore, the only place where the visited_cpus bits are cleared= is > > > > task_cache_work(), which is only called once per scan period, there= is no > > > > risk of the bit being cleared concurrently mid-loop. > > > >=20 > > >=20 > > > Make sense. > > > =20 > > > > However, is there a possibility that the current task_cache_work() = execution > > > > hasn't finished yet when the next scan window arrives? For instance= , if the > > > > current task work is heavily delayed or preempted by unexpected int= errupt, > > > > jiffies could advance past next_scan before the loop completes. > > > > If we move the `work->next =3D work;` to the very end of task_cache= _work(), > > > > would that resolve this issue? By doing so, the existing `work->nex= t =3D=3D work` > > > > check in task_tick_cache() should fail and no new task work will be= submitted. > > > >=20 > > > > Please let me know if I'm missing something. > > > >=20 > > >=20 > > > There are two layers of protection: first a cheap timeout gate (time_= before) > > > that skips scanning until the next period, and then a try_cmpxchg tha= t atomically > > > picks a single winner among the threads that pass the timeout =E2=80= =94 this actually > > > guarantees only one scanner per mm at a time, no? > >=20 > > What I am worried about is the following scenario: > >=20 > > Thread A (CPU 0) Thread B (CPU 1) > > =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D = =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D > > task_cache_work() > > | > > +-> try_cmpxchg() =3D=3D true > > | (Sets next_scan =3D now + 10) > > | > > +-> Enters Scanning Loop (jiffies =3D 100) > > | [ Delayed / Preempted ] > > | jiffies advances 100 -> 115. > > | Thread A STILL in the loop! > > | task_cache_work() (jiffie= s =3D 115) > >=20 > > | | > > | +-> next_scan =3D=3D 11= 0 (pass timeout check) > >=20 > > | | > > | +-> try_cmpxchg() =3D= =3D true > >=20 > > | | (Sets next_scan =3D= 115 + 10) > > | | > >=20 > > In other words, concurrent execution of task_cache_work() from two adja= cent periods can > > occur under extreme conditions; do we need to take this scenario into c= onsideration? > >=20 > > Perhaps we need an explicit state flag (e.g., using test_and_set_bit) t= o ensure > > that the previous task_cache_work() execution has fully completed befor= e allowing > > a new thread to proceed, regardless of whether the next scan window has= arrived. > > What do you think? >=20 > First, the EPOCH_PERIOD is fairly long (10 msec) so it is unlikely that T= hread A > hasn't completed its work. >=20 > Second, suppose the above scenario happened, the two thread above > are serialized in the work on updating occupancy > and visited cpus under the cpus_read_lock when looping through the cpus > in task_cache_work(). >=20 > So I think we should be okay without an explicit state flag. >=20 > That said, I think we should use mm->sc_stat.lock instead of cpus_read_lo= ck > in task_cache_work()'s cpu loop. That will improve scalability. It slipped my mind the cpus_read_lock is read lock, so the two threads could indeed go in parallel. But the update on each cpu is protected by rq->cpu_epoch_lock, and epochs don't go backwards. So we should still get consistent updates on epochs and occupancy stats. Tim