From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7126236A34F for ; Thu, 25 Jun 2026 11:20:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782386437; cv=none; b=Nwz7iXLPZciRc/Yd3TTIjefPNJs3HLNdM9rGwhSbACmv4r2VAlYkRkGT73xRzAff0v4zyONZARdtYROvn9tz0Q4rIzzh4xB1sPYpKbziyZxhjuJd46TwkqsITtSbtVoeMEyX4eCvIuLDiPCf2bbd+Sc9X42m1mHH0SJr3NV7Jbw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782386437; c=relaxed/simple; bh=ftiN6fEJeTKvmg6K3zn1oC8wK5F7/yoWOleVoJarwzw=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=SA+d3KMkUC48bmG+tJ9rbL6Q35n5MsGUmSnjmvhINGexy3TBVaWcghsLwI2oH5dyINDiVZ/6CC7+IKwkF7A+73S1LpImp2NtxzZQqP9lZvuLlReJoNuIZhPYfysVNJxPTIQ7FeJluJSdU6wfflRJN/y4OhRMx3N/ire5KfIcnkU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=j+fUvyNt; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="j+fUvyNt" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=MIME-Version:Content-Transfer-Encoding:Content-Type:References: In-Reply-To:Date:Cc:To:From:Subject:Message-ID:Sender:Reply-To:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=JFk5c/1jNSMQLdFUSaHagIMvPm2hmWCfIn7xWSUzENg=; b=j+fUvyNtxzAZ/qnLIU0qO9PFEV BQxm5FHxb+olPWFNqAoGztRBjQuvVnOpQnKwGyVS7XFKPAlM1uLEnC5mOk1fQ0YZGayJ5Jpek1+Ll bAgYlL0PqqZeMDqKlbBmfVht5QHbJ3W2p/cRdK4TNWEiOcztdU8Dmm1M6vTwNLU6vou3stgzFz+C4 HmfUTPmunViN+bsbWtI2pEtCwM9Taabq4OSy+ujrbz3LXauDgofGfnS6+GA62rT06aumDUbfFaROF +8+4ALPbVvAIAwDIyLdlhfqNHTxUAQfXEOMJ+lQjyPqBBw9Y4fpMRBJuVtlFAhMwvjdm4ysoCbGq4 4UaWS68w==; Received: from fangorn.home.surriel.com ([10.0.13.7]) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wci86-000000001j0-1X3u; Thu, 25 Jun 2026 07:20:22 -0400 Message-ID: <610d1c95a2f9131d5c018304f9a90afb33e8a032.camel@surriel.com> Subject: Re: [PATCH 2/3] mm/pagewalk: let folio_walk_start() run under the per-VMA lock From: Rik van Riel To: Lorenzo Stoakes Cc: linux-kernel@vger.kernel.org, x86@kernel.org, linux-mm@kvack.org, Thomas Gleixner , Ingo Molnar , Dmitry Ilvokhin , Borislav Petkov , Dave Hansen , Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Suren Baghdasaryan , kernel-team@meta.com Date: Thu, 25 Jun 2026 07:20:22 -0400 In-Reply-To: References: <20260625015053.2445008-1-riel@surriel.com> <20260625015053.2445008-3-riel@surriel.com> Autocrypt: addr=riel@surriel.com; prefer-encrypt=mutual; keydata=mQENBFIt3aUBCADCK0LicyCYyMa0E1lodCDUBf6G+6C5UXKG1jEYwQu49cc/gUBTTk33A eo2hjn4JinVaPF3zfZprnKMEGGv4dHvEOCPWiNhlz5RtqH3SKJllq2dpeMS9RqbMvDA36rlJIIo47 Z/nl6IA8MDhSqyqdnTY8z7LnQHqq16jAqwo7Ll9qALXz4yG1ZdSCmo80VPetBZZPw7WMjo+1hByv/ lvdFnLfiQ52tayuuC1r9x2qZ/SYWd2M4p/f5CLmvG9UcnkbYFsKWz8bwOBWKg1PQcaYHLx06sHGdY dIDaeVvkIfMFwAprSo5EFU+aes2VB2ZjugOTbkkW2aPSWTRsBhPHhV6dABEBAAG0HlJpayB2YW4gU mllbCA8cmllbEByZWRoYXQuY29tPokBHwQwAQIACQUCW5LcVgIdIAAKCRDOed6ShMTeg05SB/986o gEgdq4byrtaBQKFg5LWfd8e+h+QzLOg/T8mSS3dJzFXe5JBOfvYg7Bj47xXi9I5sM+I9Lu9+1XVb/ r2rGJrU1DwA09TnmyFtK76bgMF0sBEh1ECILYNQTEIemzNFwOWLZZlEhZFRJsZyX+mtEp/WQIygHV WjwuP69VJw+fPQvLOGn4j8W9QXuvhha7u1QJ7mYx4dLGHrZlHdwDsqpvWsW+3rsIqs1BBe5/Itz9o 6y9gLNtQzwmSDioV8KhF85VmYInslhv5tUtMEppfdTLyX4SUKh8ftNIVmH9mXyRCZclSoa6IMd635 Jq1Pj2/Lp64tOzSvN5Y9zaiCc5FucXtB9SaWsgdmFuIFJpZWwgPHJpZWxAc3VycmllbC5jb20+iQE +BBMBAgAoBQJSLd2lAhsjBQkSzAMABgsJCAcDAgYVCAIJCgsEFgIDAQIeAQIXgAAKCRDOed6ShMTe g4PpB/0ZivKYFt0LaB22ssWUrBoeNWCP1NY/lkq2QbPhR3agLB7ZXI97PF2z/5QD9Fuy/FD/jddPx KRTvFCtHcEzTOcFjBmf52uqgt3U40H9GM++0IM0yHusd9EzlaWsbp09vsAV2DwdqS69x9RPbvE/Ne fO5subhocH76okcF/aQiQ+oj2j6LJZGBJBVigOHg+4zyzdDgKM+jp0bvDI51KQ4XfxV593OhvkS3z 3FPx0CE7l62WhWrieHyBblqvkTYgJ6dq4bsYpqxxGJOkQ47WpEUx6onH+rImWmPJbSYGhwBzTo0Mm G1Nb1qGPG+mTrSmJjDRxrwf1zjmYqQreWVSFEt26tBpSaWsgdmFuIFJpZWwgPHJpZWxAZmIuY29tP okBPgQTAQIAKAUCW5LbiAIbIwUJEswDAAYLCQgHAwIGFQgCCQoLBBYCAwECHgECF4AACgkQznneko TE3oOUEQgAsrGxjTC1bGtZyuvyQPcXclap11Ogib6rQywGYu6/Mnkbd6hbyY3wpdyQii/cas2S44N cQj8HkGv91JLVE24/Wt0gITPCH3rLVJJDGQxprHTVDs1t1RAbsbp0XTksZPCNWDGYIBo2aHDwErhI omYQ0Xluo1WBtH/UmHgirHvclsou1Ks9jyTxiPyUKRfae7GNOFiX99+ZlB27P3t8CjtSO831Ij0Ip QrfooZ21YVlUKw0Wy6Ll8EyefyrEYSh8KTm8dQj4O7xxvdg865TLeLpho5PwDRF+/mR3qi8CdGbkE c4pYZQO8UDXUN4S+pe0aTeTqlYw8rRHWF9TnvtpcNzZw== Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.56.2 (3.56.2-2.fc42) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Thu, 2026-06-25 at 08:34 +0100, Lorenzo Stoakes wrote: > Rik, it really would have helped if you'd replied to review :) >=20 > On Wed, Jun 24, 2026 at 09:50:52PM -0400, Rik van Riel wrote: > > folio_walk_start() asserts the mmap lock is held.=C2=A0 For callers tha= t > > only > > need to read a single, already-present page, the mmap lock is a > > heavy and > > often badly contended hammer.=C2=A0 Such a caller can instead hold the > > per-VMA > > lock, which keeps the VMA itself stable. >=20 > >=20 > > The per-VMA lock does not, however, keep the page tables walked > > below that > > VMA from being freed.=C2=A0 A concurrent munmap() or THP collapse of an > > adjacent region in the same mm can free a shared upper-level table, > > and >=20 > Yeah I need to update the documentation on this at > https://docs.kernel.org/mm/process_addrs.html=C2=A0it's more subtle than > written > there. >=20 > Firstly you're wrong about munmap() - it acquires the VMA lock of the > VMAs freed > in the range and will only remove an upper level table if the entire > range is > spanned. >=20 > And that's the only way higher level tables can be removed. >=20 > PTE page tables can be removed via MADV_DONTNEED, but that a. > acquires the VMA > lock and b. frees the PTE page table under RCU. >=20 > A THP collapse can happen concurrently, but PTEs are freed under RCU > so you > don't need to do this GUP fast imitating stuff. >=20 > > THP collapse (collapse_huge_page() -> retract_page_tables()) frees > > page > > tables of VMAs whose lock it does not hold.=C2=A0 Page table freeing >=20 > retract_page_tables() -> pte_free_defer() -> RCU > try_collapse_pte_mapped_thp() -> pte_free_defer() -> RCU One issue here is that while we can safely read the old page table under the RCU read lock, in the middle of a THP collapse there is no guarantee that the old page table points at the process's=C2=A0 current memory. Khugepaged could fix this in one of two ways: - zap all readers with an IPI, and use that as synchronization - make sure the old page table's PTEs point at the individual pages inside the new PMD Right now khugepaged does the first. Relying only on the RCU read lock to read the page table could result in us seeing old page table contents, that no longer point at the current process memory. Unless I'm missing something... --=20 All Rights Reversed.