From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AB0363515DA for ; Tue, 25 Aug 2026 19:58:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787687921; cv=none; b=nurr92kEC7gc1j9TliIzrfSt1LYogbAs8fVMcx1/9CbggPZS9Ge56qLW2eHpHQ9zjaRYy65UA5h4xOtd4VLyrjUMktr6IU+8PeJfjdJCQe+81dxDQioVK391mb66V6Dt40ejMHDXzQ630Sw4P2AYafG6aZAQqy3y1Y6c0+SManM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787687921; c=relaxed/simple; bh=Rs1ly7UYBMlUYE7QVy/j2KMt5Co5GBrj7F01YNt8mGU=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=KR1EkT6jmuDQ4CXnQnlu1DqrLMZ9Qu3J0qEN0O/JrhOlueoZhgel5XXXlwejlfEVcZk/Ewjps0Ympy4DRMpsPDpDG22HSTm+qO4TtWjwFXhleU0HFvRxcC8SQDtTyB1QHgGgOZN4KgFK5TVGgCN9kjIfLoKiuI8S6QdCbFyqUgk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=hxLuTCt1; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="hxLuTCt1" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2d63bad3d09so2986755ad.3 for ; Tue, 25 Aug 2026 12:58:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787687919; x=1788292719; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=pw5HPga/49LaG41hAzsub+JW3+Qdxi0vmGANh8Af4Kg=; b=hxLuTCt1ZBCeDxidtUdCd0VdSlpnHOh6pS2h/sOyJgD9G4LA/9vLnC3kQe78GQC0Pr /V+aUsHqkTvbbHswCUtFsuBizyAibvoOrIzR16/TaHN+fKMwfYW2XzyV7nM+16dD+AI9 JhGi0NCu2+4QT8DOEMONuGuxuTk6SI37Wknj3o7mlzfzaOlRtEVyYCUo787Zzcc3/9v4 SdUVtZLTUgVT8iyHqufPD41vKrOCKgr8VxZMCBBrJJTvJctY72U6fc0bpt4bf68wp+vd g1CvNc+yzyShIeCAHYbYg+vrGod47nppGpP4+oPZLe7zG2S1So37w/l6Oo0iErOkOk9x SeeA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787687919; x=1788292719; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=pw5HPga/49LaG41hAzsub+JW3+Qdxi0vmGANh8Af4Kg=; b=klZX26gf48Fg3/7v//1GSixL61Oj/U/EOW7WupX9xd9knHGTPCefFFc+0+l+2/XjTC qLTZMFNcRTMUZtNgNPxiYnzP8F7n8YSusyHXHr/8MpAT4dI/Y2pJgU+ayriMDxrDX0E3 jxYMc3cte53csz+C8G+AbufQrxrEw5r30fyF1KNyzlZwndYgx/vucUQYcHL3EsOSxFGv 06b1bo+e+j/yOQlwnZBeypo9XgJEuS0zidjLvlBIBRRN7oWtYjFx1ggQHwJXXRePDkqA 1jCvDB7cZWuzEAP8vFaR7D4VTzbRFE9bOtlEeciEFyiAGYT1DP9jMZoEtKBYDwuZ6Vv2 KiNg== X-Forwarded-Encrypted: i=1; AHgh+RpoQTiiMPd+NSnD1Q1aNOeXIRqNoT/MqFuRN9rUTqSRYvP4w7obzZP4uUrNNA0txP4pI+5kGzaQW/lYgaE=@vger.kernel.org X-Gm-Message-State: AFuF++kJ6rUBzVK6UttWL190Rnw+AcW40YFFVqTRHqyBT/T4/5uRG9pj 2Emrc0pL8RdJGzCm7FaPGYJhRb/piWzm74aAbMUYCu/hClkkAnPdWN4IEfwixCoYNm2P+ck00/j OPFsScg== X-Received: from plyw15.prod.google.com ([2002:a17:902:d70f:b0:2cc:6206:58ad]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:19ce:b0:2d6:f6ba:263d with SMTP id d9443c01a7336-2d707aafd1fmr12835735ad.7.1787687918385; Tue, 25 Aug 2026 12:58:38 -0700 (PDT) Date: Tue, 25 Aug 2026 12:58:37 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <333c4cdb-8f34-4e3f-a47f-961f47e089a0@paulmck-laptop> <11dcf5a87125451897982f728c1f65a5151987e4.camel@infradead.org> <69567965-48fd-48c5-abb6-0699d7f5bd16@paulmck-laptop> <124af87fb49c267eeb26b8b4d952b9b1a5b3fb68.camel@infradead.org> <9427a8e0-3ba6-4f31-a35d-54429bd7e071@paulmck-laptop> <3d0463b6099d4fcde9a025a2f9e231d75303b61c.camel@infradead.org> Message-ID: Subject: Re: [PATCH] mm/mmu_notifier: Remove non_block_start/end() from notifier invocation From: Sean Christopherson To: "Paul E. McKenney" Cc: David Woodhouse , Jason Gunthorpe , Michal Hocko , Steven Rostedt , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Sebastian Andrzej Siewior , Clark Williams , Simona Vetter , Jerome Glisse , Christian Koenig , Paolo Bonzini , linux-mm@kvack.org, kvm@vger.kernel.org, linux-rt-devel@lists.linux.dev, linux-kernel@vger.kernel.org Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable On Tue, Aug 25, 2026, Paul E. McKenney wrote: > On Tue, Aug 25, 2026 at 06:48:08PM +0100, David Woodhouse wrote: > > On Tue, 2026-08-25 at 10:19 -0700, Paul E. McKenney wrote: > > > On Tue, Aug 25, 2026 at 06:05:54PM +0100, David Woodhouse wrote: > > > > On Tue, 2026-08-25 at 09:47 -0700, Paul E. McKenney wrote: > > > > >=20 > > > > > On the tail latencies... > > > > >=20 > > > > > The easiest way to reduce them is to require that preemption be d= isabled > > > > > across srcu_read_lock_atomic()/srcu_read_unlock_atomic() regions = and > > > > > across all calls to synchronize_srcu_atomic().=C2=A0 Without that= , the problem > > > > > is that the scheduler does not know that the spinning is pointles= s, > > > > > and we cannot use the blocking primitives that we could otherwise= use > > > > > to tell it what is going on. > > > > >=20 > > > > > So, is it feasible to simply require preemption be disabled as ca= lled > > > > > out above? > > > >=20 > > > > I'd experimented with disabling it around the GP driver loop in > > > > synchronize_srcu_atomic() as seen in > > > > https://git.infradead.org/?p=3Dusers/dwmw2/linux.git;a=3Dcommitdiff= ;h=3D07165e79340e > > > > and that didn't seem to change anything (which seems reasonable, as > > > > it's the *waiters* that were descheduled, not the threads driving t= he > > > > actual GP). So your suggestion that we do it around the whole funct= ion > > > > certainly makes sense too. I'll test it. > > > >=20 > > > > I do wonder if we're really doing the right thing here by selfishly > > > > blocking preemption because we want a specific tail latency to rema= in > > > > low in a contended system. Maybe we should allow preemption and tru= st > > > > that the right thing will happen? In my experience, preempting MMU operations, especially mmu_notifier invali= dations, is rarely a good idea. E.g. see commit d02c357e5bfa ("KVM: x86/mmu: Retry = fault before acquiring mmu_lock if mapping is changing"), which worked around an = issue where KVM would drop mmu_lock and yield in an mmu_notifier callback on pree= mptible kernels. We "fixed" the issue by avoiding mmu_lock contention, because it = was the easiest fix and benefited all setups, but the underlying problem that made = us take action was very specifically yielding mmu_lock on preemptible kernels. This isn't exactly the same, but it sounds quite similar: being greedy and = hogging the CPU to complete an operation can actually be beneficial for overall thr= oughput, not just for the immediate operation's latency, by avoiding trash and overh= ead that is incurred as a result of yielding or being preempted. > > > > Maybe the p100 isn't the right benchmark to be chasing... I'm looki= ng > > > > at it because Sean expressed concerns about it, but it's not the on= ly > > > > consideration. > > >=20 > > > My concern is algorithmic, not benchmark optimization. > > >=20 > > > Suppose that there is only one CPU, or, alternatively, that one of th= e > > > atomic SRCU readers is pinned to the same CPU occupied by the (higher > > > priority) task running synchronize_srcu_atomic().=C2=A0 In this case,= the > > > call to synchronize_srcu_atomic() uselessly burns CPU time until its > > > priority decays, real-time throttling kicks in, or in some configurat= ions, > > > maybe never. > >=20 > > I certainly have no problem with a blanket preempt_disable() around > > both sides for algorithmic reasons. As long as we aren't *just* doing > > it for the selfish reasons I described.=20 >=20 > Suppose I simply disable preemption in srcu_read_lock_atomic(), > enable it in srcu_read_unlock_atomic(), and disable it internally to > synchronize_srcu_atomic()? It might be against all RCU tradition, > but might also be easier to use. ;-) >=20 > > > Requiring preemption be disabled across both the atomic SRCU readers > > > and the synchronize_srcu_atomic() avoids this, at least when running = on > > > bare metal.=C2=A0 My (perhaps naive) hope is that guest OSes get some= use > > > out of those cpu_relax() calls. > >=20 > > Yeah, an overcommited guest vCPU should be able to get preempted there > > by the hypervisor, allowing other vCPUs to run. >=20 > Whew!!! ;-) Ya, and on KVM x86 at least, cpu_relax() =3D> PAUSE will conditionally trig= ger a VM-Exit after enough spins that causes KVM-the-host to try to yield the vCP= U to another vCPU in the same VM. The intended use case is to detect when a vCP= U is spinning waiting for a lock, to try and give cycles to the vCPU that is hol= ding said lock. IIUC, the same principle should apply here.