From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D627A18FDC8 for ; Mon, 20 Jan 2025 16:09:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737389388; cv=none; b=qDZJwkwhOFx3ABq4WXDYj+aeXrsLDDjlwwHAhwyxo1qZi82AnQVFeXx0CH/Jy2QyR0K1TLebpN+Nt037GIouONFITOQ92on3M9HAhLRQi3W022L6jbngzO/p2ioNDs+WD5u8y+993pi1O7q98yUd9YNDSFJJ6yuA5WDq+v3jrMc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737389388; c=relaxed/simple; bh=93WBb7KvYei1+bL8+Urz4YKDHhr7IwcnfNWXhTZuSug=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=r2SI1B8xasIDrRq2Fu9cgiSv3EE9VRXZWEAfAOOO9UFN/BugGp6Hu/8wve+1Oocp6Dmuo7Ha9BPAiwoRmeLJIwU73Ol0+CUAdrfTUrVPwUcWrPYFI/wvAFPG040CXe7fUBO6SryyoVz1Ib90GJc/Z3rKh8iiawYx1JpHIZ0RVOU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=shelob.surriel.com; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shelob.surriel.com Received: from fangorn.home.surriel.com ([10.0.13.7]) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1tZuL1-000000000yf-3zEo; Mon, 20 Jan 2025 11:09:19 -0500 Message-ID: Subject: Re: [PATCH v6 09/12] x86/mm: enable broadcast TLB invalidation for multi-threaded processes From: Rik van Riel To: Nadav Amit , x86@kernel.org Cc: linux-kernel@vger.kernel.org, bp@alien8.de, peterz@infradead.org, dave.hansen@linux.intel.com, zhengqi.arch@bytedance.com, thomas.lendacky@amd.com, kernel-team@meta.com, linux-mm@kvack.org, akpm@linux-foundation.org, jannh@google.com, mhklinux@outlook.com, andrew.cooper3@citrix.com Date: Mon, 20 Jan 2025 11:09:19 -0500 In-Reply-To: <84ba1c3e-d975-458f-89f5-a6f5d04a3d22@gmail.com> References: <20250120024104.1924753-1-riel@surriel.com> <20250120024104.1924753-10-riel@surriel.com> <84ba1c3e-d975-458f-89f5-a6f5d04a3d22@gmail.com> Autocrypt: addr=riel@surriel.com; prefer-encrypt=mutual; keydata=mQENBFIt3aUBCADCK0LicyCYyMa0E1lodCDUBf6G+6C5UXKG1jEYwQu49cc/gUBTTk33A eo2hjn4JinVaPF3zfZprnKMEGGv4dHvEOCPWiNhlz5RtqH3SKJllq2dpeMS9RqbMvDA36rlJIIo47 Z/nl6IA8MDhSqyqdnTY8z7LnQHqq16jAqwo7Ll9qALXz4yG1ZdSCmo80VPetBZZPw7WMjo+1hByv/ lvdFnLfiQ52tayuuC1r9x2qZ/SYWd2M4p/f5CLmvG9UcnkbYFsKWz8bwOBWKg1PQcaYHLx06sHGdY dIDaeVvkIfMFwAprSo5EFU+aes2VB2ZjugOTbkkW2aPSWTRsBhPHhV6dABEBAAG0HlJpayB2YW4gU mllbCA8cmllbEByZWRoYXQuY29tPokBHwQwAQIACQUCW5LcVgIdIAAKCRDOed6ShMTeg05SB/986o gEgdq4byrtaBQKFg5LWfd8e+h+QzLOg/T8mSS3dJzFXe5JBOfvYg7Bj47xXi9I5sM+I9Lu9+1XVb/ r2rGJrU1DwA09TnmyFtK76bgMF0sBEh1ECILYNQTEIemzNFwOWLZZlEhZFRJsZyX+mtEp/WQIygHV WjwuP69VJw+fPQvLOGn4j8W9QXuvhha7u1QJ7mYx4dLGHrZlHdwDsqpvWsW+3rsIqs1BBe5/Itz9o 6y9gLNtQzwmSDioV8KhF85VmYInslhv5tUtMEppfdTLyX4SUKh8ftNIVmH9mXyRCZclSoa6IMd635 Jq1Pj2/Lp64tOzSvN5Y9zaiCc5FucXtB9SaWsgdmFuIFJpZWwgPHJpZWxAc3VycmllbC5jb20+iQE +BBMBAgAoBQJSLd2lAhsjBQkSzAMABgsJCAcDAgYVCAIJCgsEFgIDAQIeAQIXgAAKCRDOed6ShMTe g4PpB/0ZivKYFt0LaB22ssWUrBoeNWCP1NY/lkq2QbPhR3agLB7ZXI97PF2z/5QD9Fuy/FD/jddPx KRTvFCtHcEzTOcFjBmf52uqgt3U40H9GM++0IM0yHusd9EzlaWsbp09vsAV2DwdqS69x9RPbvE/Ne fO5subhocH76okcF/aQiQ+oj2j6LJZGBJBVigOHg+4zyzdDgKM+jp0bvDI51KQ4XfxV593OhvkS3z 3FPx0CE7l62WhWrieHyBblqvkTYgJ6dq4bsYpqxxGJOkQ47WpEUx6onH+rImWmPJbSYGhwBzTo0Mm G1Nb1qGPG+mTrSmJjDRxrwf1zjmYqQreWVSFEt26tBpSaWsgdmFuIFJpZWwgPHJpZWxAZmIuY29tP okBPgQTAQIAKAUCW5LbiAIbIwUJEswDAAYLCQgHAwIGFQgCCQoLBBYCAwECHgECF4AACgkQznneko TE3oOUEQgAsrGxjTC1bGtZyuvyQPcXclap11Ogib6rQywGYu6/Mnkbd6hbyY3wpdyQii/cas2S44N cQj8HkGv91JLVE24/Wt0gITPCH3rLVJJDGQxprHTVDs1t1RAbsbp0XTksZPCNWDGYIBo2aHDwErhI omYQ0Xluo1WBtH/UmHgirHvclsou1Ks9jyTxiPyUKRfae7GNOFiX99+ZlB27P3t8CjtSO831Ij0Ip QrfooZ21YVlUKw0Wy6Ll8EyefyrEYSh8KTm8dQj4O7xxvdg865TLeLpho5PwDRF+/mR3qi8CdGbkE c4pYZQO8UDXUN4S+pe0aTeTqlYw8rRHWF9TnvtpcNzZw== Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.54.1 (3.54.1-1.fc41) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Sender: riel@surriel.com On Mon, 2025-01-20 at 16:02 +0200, Nadav Amit wrote: >=20 >=20 > On 20/01/2025 4:40, Rik van Riel wrote: > >=20 > > +static inline void broadcast_tlb_flush(struct flush_tlb_info > > *info) > > +{ > > + VM_WARN_ON_ONCE(1); >=20 > Not sure why not the use VM_WARN_ONCE() instead with some more=20 > informative message (anyhow, a string is allocated for it). >=20 VM_WARN_ON_ONCE only has a condition, not a message. > >=20 > > +static u16 get_global_asid(void) > > +{ > > + lockdep_assert_held(&global_asid_lock); > > + > > + do { > > + u16 start =3D last_global_asid; > > + u16 asid =3D find_next_zero_bit(global_asid_used, > > MAX_ASID_AVAILABLE, start); > > + > > + if (asid >=3D MAX_ASID_AVAILABLE) { > > + reset_global_asid_space(); > > + continue; > > + } >=20 > I think that unless something is awfully wrong, you are supposed to > at=20 > most call reset_global_asid_space() once. So if that's the case, why > not=20 > do it this way? >=20 > Instead, you can get rid of the loop and just do: >=20 > asid =3D find_next_zero_bit(global_asid_used, > MAX_ASID_AVAILABLE, start); >=20 > If you want, you can warn if asid >=3D MAX_ASID_AVAILABLE and have some > fallback. But the loop, is just confusing in my opinion for no > reason. I can get rid of the loop. You're right that the code can just call find_next_zero_bit after calling reset_global_asid_space. >=20 > > + /* Slower check to make sure. */ > > + for_each_cpu(cpu, mm_cpumask(mm)) { > > + /* Skip the CPUs that aren't really running this > > process. */ > > + if (per_cpu(cpu_tlbstate.loaded_mm, cpu) !=3D mm) > > + continue; >=20 > Then perhaps at least add a comment next to loaded_mm, that it's not=20 > private per-se, but rarely accessed by other cores? >=20 I don't see any comment in struct tlb_state that suggests it was ever private to begin with. Which comment are you referring to that should be edited? > >=20 > > + > > + /* > > + * The transition from IPI TLB flushing, with a dynamic > > ASID, > > + * and broadcast TLB flushing, using a global ASID, uses > > memory > > + * ordering for synchronization. > > + * > > + * While the process has threads still using a dynamic > > ASID, > > + * TLB invalidation IPIs continue to get sent. > > + * > > + * This code sets asid_transition first, before assigning > > the > > + * global ASID. > > + * > > + * The TLB flush code will only verify the ASID transition > > + * after it has seen the new global ASID for the process. > > + */ > > + WRITE_ONCE(mm->context.asid_transition, true); > > + WRITE_ONCE(mm->context.global_asid, get_global_asid()); >=20 > I know it is likely correct in practice (due to TSO memory model), > but=20 > it is not clear, at least for me, how those write order affects the > rest=20 > of the code. I managed to figure out how it relates to the reads in=20 > flush_tlb_mm_range() and native_flush_tlb_multi(), but I wouldn't say > it=20 > is trivial and doesn't worth a comment (or smp_wmb/smp_rmb). >=20 What kind of wording should we add here to make it easier to understand? "The TLB invalidation code reads these variables in the opposite order in which they are written" ? > > + /* > > + * If at least one CPU is not using the global > > ASID yet, > > + * send a TLB flush IPI. The IPI should cause > > stragglers > > + * to transition soon. > > + * > > + * This can race with the CPU switching to another > > task; > > + * that results in a (harmless) extra IPI. > > + */ > > + if (READ_ONCE(per_cpu(cpu_tlbstate.loaded_mm_asid, > > cpu)) !=3D bc_asid) { > > + flush_tlb_multi(mm_cpumask(info->mm), > > info); > > + return; >=20 > I am trying to figure out why we return here. The transition might > not=20 > be over? Why is it "soon"? Wouldn't flush_tlb_func() reload it=20 > unconditionally? The transition _should_ be over, but what if another CPU got an NMI while in the middle of switch_mm_irqs_off, and set its own bit in the mm_cpumask after we send this IPI? On the other hand, if it sets its mm_cpumask bit after this point, it will also load the mm->context.global_asid after this point, and should definitely get the new ASID. I think we are probably fine to set asid_transition to false here, but I've had to tweak this code so much over the past months that I don't feel super confident any more :) >=20 > > + /* > > + * TLB flushes with INVLPGB are kicked off asynchronously. > > + * The inc_mm_tlb_gen() guarantees page table updates are > > done > > + * before these TLB flushes happen. > > + */ > > + if (info->end =3D=3D TLB_FLUSH_ALL) { > > + invlpgb_flush_single_pcid_nosync(kern_pcid(asid)); > > + /* Do any CPUs supporting INVLPGB need PTI? */ > > + if (static_cpu_has(X86_FEATURE_PTI)) > > + invlpgb_flush_single_pcid_nosync(user_pcid > > (asid)); > > + } else for (; addr < info->end; addr +=3D nr << info- > > >stride_shift) { >=20 > I guess I was wrong, and do-while was cleaner here. >=20 > And I guess this is now a bug, if info->stride_shift > PMD_SHIFT... >=20 We set maxnr to 1 for larger stride shifts at the top of the function: /* Flushing multiple pages at once is not supported with 1GB pages. */ if (info->stride_shift > PMD_SHIFT) maxnr =3D 1; > [ I guess the cleanest way was to change get_flush_tlb_info to mask > the=20 > low bits of start and end based on ((1ull << stride_shift) - 1). But=20 > whatever... ] I'll change it back :) I'm just happy this code is getting lots of attention, and we're improving it with time. > > @@ -573,6 +874,23 @@ void switch_mm_irqs_off(struct mm_struct > > *unused, struct mm_struct *next, > > =C2=A0=C2=A0 !cpumask_test_cpu(cpu, > > mm_cpumask(next)))) > > =C2=A0=C2=A0 cpumask_set_cpu(cpu, mm_cpumask(next)); > > =C2=A0=20 > > + /* > > + * Check if the current mm is transitioning to a > > new ASID. > > + */ > > + if (needs_global_asid_reload(next, prev_asid)) { > > + next_tlb_gen =3D atomic64_read(&next- > > >context.tlb_gen); > > + > > + choose_new_asid(next, next_tlb_gen, > > &new_asid, &need_flush); > > + goto reload_tlb; >=20 > Not a fan of the goto's when they are not really needed, and I don't=20 > think it is really needed here. Especially that the name of the tag=20 > "reload_tlb" does not really convey that the page-tables are reloaded > at=20 > that point. In this particular case, the CPU continues running with the same page tables, but with a different PCID. >=20 > > + } > > + > > + /* > > + * Broadcast TLB invalidation keeps this PCID up > > to date > > + * all the time. > > + */ > > + if (is_global_asid(prev_asid)) > > + return; >=20 > Hard for me to convince myself When a process uses a global ASID, we always send out TLB invalidations using INVLPGB. The global ASID should always be up to date. >=20 > > @@ -769,6 +1092,16 @@ static void flush_tlb_func(void *info) > > =C2=A0=C2=A0 if (unlikely(loaded_mm =3D=3D &init_mm)) > > =C2=A0=C2=A0 return; > > =C2=A0=20 > > + /* Reload the ASID if transitioning into or out of a > > global ASID */ > > + if (needs_global_asid_reload(loaded_mm, loaded_mm_asid)) { > > + switch_mm_irqs_off(NULL, loaded_mm, NULL); >=20 > I understand you want to reuse that logic, but it doesn't seem=20 > reasonable to me. It both doesn't convey what you want to do, and can > lead to undesired operations - cpu_tlbstate_update_lam() for > instance.=20 > Probably the impact on performance is minor, but it is an opening for > future mistakes. My worry with having a separate code path here is that the separate code path could bit rot, and we could introduce bugs that way. I would rather have a tiny performance impact in what is a rare code path, than a rare (and hard to track down) memory corruption due to bit rot. >=20 > > + loaded_mm_asid =3D > > this_cpu_read(cpu_tlbstate.loaded_mm_asid); > > + } > > + > > + /* Broadcast ASIDs are always kept up to date with > > INVLPGB. */ > > + if (is_global_asid(loaded_mm_asid)) > > + return; >=20 > The comment does not clarify to me, and I don't manage to clearly=20 > explain to myself, why it is guaranteed that all the IPI TLB flushes, > which were potentially issued before the transition, are not needed. >=20 IPI TLB flushes that were issued before the transition went to the CPUs when they were using dynamic ASIDs (numbers 1-5). Reloading the TLB with a different PCID, even pointed at the same page tables, means that the TLB should load the translations fresh from the page tables, and not re-use any that it had previously loaded under a different PCID. --=20 All Rights Reversed.