From: Kiryl Shutsemau <kas@kernel.org>
To: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
"Liam R. Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>,
Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>, Guo Ren <guoren@kernel.org>,
Brian Cain <bcain@kernel.org>,
Geert Uytterhoeven <geert@linux-m68k.org>,
Dinh Nguyen <dinguyen@kernel.org>,
Simon Schuster <schuster.simon@siemens-energy.com>,
Jonas Bonn <jonas@southpole.se>,
Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>,
Stafford Horne <shorne@gmail.com>,
Yoshinori Sato <ysato@users.sourceforge.jp>,
Rich Felker <dalias@libc.org>,
John Paul Adrian Glaubitz <glaubitz@physik.fu-berlin.de>,
Paul Walmsley <pjw@kernel.org>,
Palmer Dabbelt <palmer@dabbelt.com>,
Albert Ou <aou@eecs.berkeley.edu>,
Alexandre Ghiti <alex@ghiti.fr>,
Russell King <linux@armlinux.org.uk>,
Vineet Gupta <vgupta@kernel.org>,
Michal Simek <monstr@monstr.eu>, Chris Zankel <chris@zankel.net>,
Max Filippov <jcmvbkbc@gmail.com>, Will Deacon <will@kernel.org>,
"Aneesh Kumar K.V" <aneesh.kumar@kernel.org>,
Nick Piggin <npiggin@gmail.com>,
Peter Zijlstra <peterz@infradead.org>,
"David S. Miller" <davem@davemloft.net>,
Andreas Larsson <andreas@gaisler.com>,
Richard Henderson <richard.henderson@linaro.org>,
Matt Turner <mattst88@gmail.com>,
Magnus Lindholm <linmag7@gmail.com>,
Catalin Marinas <catalin.marinas@arm.com>,
Mark Rutland <mark.rutland@arm.com>,
Huacai Chen <chenhuacai@kernel.org>,
WANG Xuerui <kernel@xen0n.name>,
Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
"James E.J. Bottomley" <James.Bottomley@hansenpartnership.com>,
Helge Deller <deller@gmx.de>,
Madhavan Srinivasan <maddy@linux.ibm.com>,
Michael Ellerman <mpe@ellerman.id.au>,
"Christophe Leroy (CS GROUP)" <chleroy@kernel.org>,
Heiko Carstens <hca@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Alexander Gordeev <agordeev@linux.ibm.com>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Sven Schnelle <svens@linux.ibm.com>,
Richard Weinberger <richard@nod.at>,
Anton Ivanov <anton.ivanov@cambridgegreys.com>,
Johannes Berg <johannes@sipsolutions.net>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Arnd Bergmann <arnd@arndb.de>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
linux-csky@vger.kernel.org, linux-hexagon@vger.kernel.org,
linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org,
linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org,
linux-arm-kernel@lists.infradead.org,
linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org,
sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org,
loongarch@lists.linux.dev, linux-mips@vger.kernel.org,
linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
linux-s390@vger.kernel.org, linux-um@lists.infradead.org,
Hugh Dickins <hughd@google.com>, Qi Zheng <qi.zheng@linux.dev>
Subject: Re: [PATCH v2 12/12] mm: change the contract for free_pgtables(), update docs
Date: Wed, 9 Sep 2026 10:24:21 +0100 [thread overview]
Message-ID: <aqEkfKww9XDq15ZF@thinkstation> (raw)
In-Reply-To: <20260908-rcu-pagetable-freeing-v2-12-1f60b64e878e@kernel.org>
On Tue, Sep 08, 2026 at 01:32:21PM +0100, Lorenzo Stoakes (ARM) wrote:
> Now that page tables are freed after an RCU grace period, it is safe for
> read-only page table walkers to walk page table ranges that are being
> concurrently torn down, provided the mm is kept alive via mmgrab().
>
> It is however unsafe for writers to do so, as they must obtain an
> appropriate lock to do so safely.
>
> Update the pte_offset_map_lock()'s comment block to reflect this.
>
> Similarly update the process addresses documentation.
>
> Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
> ---
> Documentation/mm/process_addrs.rst | 6 ++++++
> mm/pgtable-generic.c | 15 +++++++++++----
> 2 files changed, 17 insertions(+), 4 deletions(-)
>
> diff --git a/Documentation/mm/process_addrs.rst b/Documentation/mm/process_addrs.rst
> index a7296f251799..1e65b139f355 100644
> --- a/Documentation/mm/process_addrs.rst
> +++ b/Documentation/mm/process_addrs.rst
> @@ -537,6 +537,12 @@ We establish basic locking rules when interacting with page tables:
> * When changing a page table entry the page table lock for that page table
> **must** be held, except if you can safely assume nobody can access the page
> tables concurrently (such as on invocation of :c:func:`!free_pgtables`).
> +* Page tables may be *walked* under RCU alone, as page tables are freed only
> + after an RCU grace period has elapsed. However, any entry found must be
> + revalidated after the page table lock is taken (such as the
> + :c:func:`!pmd_same` recheck performed by :c:func:`!pte_offset_map_lock`)
> + before it is acted upon. Changing an entry always requires the page table
> + lock.
This says the PTL plus a recheck is enough to act on the entry. It is
not: free_pte_range() clears the PMD without the PTL, so pmd_same() can
pass and the table is torn down right after. That is the case we settled
on for the pte_offset_map_lock() comment, and the two now disagree.
Something like "Changing an entry requires the page table lock and one
of the locks that excludes teardown (mmap or VMA lock)" would match the
comment.
> * Reads from and writes to page table entries must be *appropriately*
> atomic. See the section on atomicity below for details.
> * Populating previously empty entries requires that the mmap or VMA locks are
> diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c
> index b91b1a98029c..a127e3e8f9b9 100644
> --- a/mm/pgtable-generic.c
> +++ b/mm/pgtable-generic.c
> @@ -385,10 +385,17 @@ pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
> * Note: "RO" / "RW" expresses the intended semantics, not that the *kmap* will
> * be read-only/read-write protected.
> *
> - * Note that free_pgtables(), used after unmapping detached vmas, or when
> - * exiting the whole mm, does not take page table lock before freeing a page
> - * table, and may not use RCU at all: "outsiders" like khugepaged should avoid
> - * pte_offset_map() and co once the vma is detached from mm or mm_users is zero.
> + * Note that free_pgtables(), used after unmapping detached vmas or when exiting
> + * the whole mm, does not take a page table lock before freeing a page table.
> + *
> + * As page table freeing itself is RCU-safe, page table readers can safely run
> + * concurrently with page table teardown.
> + *
> + * However, writers CANNOT as, without a lock being held, nothing prevents
> + * concurrent teardown.
> + *
> + * Also note that the PGD itself is freed at mmdrop() time, not under RCU - so
> + * the walker must keep the mm alive either by pinning the mm or the VMA.
> */
> pte_t *pte_offset_map_lock(struct mm_struct *mm, pmd_t *pmd,
> unsigned long addr, spinlock_t **ptlp)
>
> --
> 2.55.0
>
--
Kiryl Shutsemau / Kirill A. Shutemov
next prev parent reply other threads:[~2026-09-09 9:24 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-08 12:32 [PATCH v2 00/12] mm: make userland page table freeing RCU-safe Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 02/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for most 2-level architectures Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 03/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for MMU riscv Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 04/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for MMU arm Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 05/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for arc, microblaze, xtensa Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 06/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sparc64 Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 07/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for m68k-coldfire Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 08/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sh-X2 Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 09/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for m68k-motorola Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 10/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sparc32 Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 11/12] mm: make userland page table freeing RCU-safe Lorenzo Stoakes (ARM)
2026-09-09 9:15 ` Kiryl Shutsemau
2026-09-09 16:44 ` Lorenzo Stoakes (ARM)
2026-09-08 12:32 ` [PATCH v2 12/12] mm: change the contract for free_pgtables(), update docs Lorenzo Stoakes (ARM)
2026-09-09 9:24 ` Kiryl Shutsemau [this message]
2026-09-09 16:42 ` Lorenzo Stoakes (ARM)
2026-09-08 21:15 ` [PATCH v2 00/12] mm: make userland page table freeing RCU-safe Andrew Morton
2026-09-09 11:24 ` Lorenzo Stoakes (ARM)
2026-09-09 9:26 ` Kiryl Shutsemau
2026-09-09 11:07 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqEkfKww9XDq15ZF@thinkstation \
--to=kas@kernel.org \
--cc=James.Bottomley@hansenpartnership.com \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=andreas@gaisler.com \
--cc=aneesh.kumar@kernel.org \
--cc=anton.ivanov@cambridgegreys.com \
--cc=aou@eecs.berkeley.edu \
--cc=arnd@arndb.de \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bcain@kernel.org \
--cc=borntraeger@linux.ibm.com \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=chenhuacai@kernel.org \
--cc=chleroy@kernel.org \
--cc=chris@zankel.net \
--cc=dalias@libc.org \
--cc=dave.hansen@linux.intel.com \
--cc=davem@davemloft.net \
--cc=david@kernel.org \
--cc=deller@gmx.de \
--cc=dev.jain@arm.com \
--cc=dinguyen@kernel.org \
--cc=geert@linux-m68k.org \
--cc=glaubitz@physik.fu-berlin.de \
--cc=gor@linux.ibm.com \
--cc=guoren@kernel.org \
--cc=hca@linux.ibm.com \
--cc=hpa@zytor.com \
--cc=hughd@google.com \
--cc=jcmvbkbc@gmail.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=johannes@sipsolutions.net \
--cc=jonas@southpole.se \
--cc=kernel@xen0n.name \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linmag7@gmail.com \
--cc=linux-alpha@vger.kernel.org \
--cc=linux-arch@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-csky@vger.kernel.org \
--cc=linux-hexagon@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-m68k@lists.linux-m68k.org \
--cc=linux-mips@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-openrisc@vger.kernel.org \
--cc=linux-parisc@vger.kernel.org \
--cc=linux-riscv@lists.infradead.org \
--cc=linux-s390@vger.kernel.org \
--cc=linux-sh@vger.kernel.org \
--cc=linux-snps-arc@lists.infradead.org \
--cc=linux-um@lists.infradead.org \
--cc=linux@armlinux.org.uk \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=ljs@kernel.org \
--cc=loongarch@lists.linux.dev \
--cc=maddy@linux.ibm.com \
--cc=mark.rutland@arm.com \
--cc=mattst88@gmail.com \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=monstr@monstr.eu \
--cc=mpe@ellerman.id.au \
--cc=nico.pache@linux.dev \
--cc=npiggin@gmail.com \
--cc=palmer@dabbelt.com \
--cc=peterx@redhat.com \
--cc=peterz@infradead.org \
--cc=pjw@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=richard.henderson@linaro.org \
--cc=richard@nod.at \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=schuster.simon@siemens-energy.com \
--cc=shorne@gmail.com \
--cc=sparclinux@vger.kernel.org \
--cc=stefan.kristiansson@saunalahti.fi \
--cc=surenb@google.com \
--cc=svens@linux.ibm.com \
--cc=tglx@kernel.org \
--cc=tsbogend@alpha.franken.de \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vgupta@kernel.org \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=ysato@users.sourceforge.jp \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®