mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: James Houghton <jthoughton@google.com>
To: Will Deacon <will@kernel.org>,
	Catalin Marinas <catalin.marinas@arm.com>,
	 Muchun Song <muchun.song@linux.dev>,
	Oscar Salvador <osalvador@suse.de>,
	 Andrew Morton <akpm@linux-foundation.org>
Cc: Nikos Nikoleris <nikos.nikoleris@arm.com>,
	Linu Cherian <linu.cherian@arm.com>,
	 Mark Rutland <mark.rutland@arm.com>,
	David Hildenbrand <david@kernel.org>,
	 Ryan Roberts <ryan.roberts@arm.com>,
	Nanyong Sun <sunnanyong@huawei.com>,  Yu Zhao <yuzhao@google.com>,
	Frank van der Linden <fvdl@google.com>,
	 David Rientjes <rientjes@google.com>,
	James Houghton <jthoughton@google.com>,
	 linux-kernel@vger.kernel.org,
	linux-arm-kernel@lists.infradead.org,  linux-mm@kvack.org
Subject: [PATCH v2 12/20] arm64: Implement try_update_vmemmap_pte using the AF trick
Date: Sat,  3 Oct 2026 00:21:15 +0000	[thread overview]
Message-ID: <20261003002123.505555-13-jthoughton@google.com> (raw)
In-Reply-To: <20261003002123.505555-1-jthoughton@google.com>

try_update_vmemmap_pte() must modify vmemmap PTEs without introducing a
time window where other CPUs on the system might fault.

Normally a break-before-make sequence is required to avoid conflicts
with cached translations. However, if we can guarantee that the existing
translation cannot be cached, a BBM sequence is not needed.

Translations with the AF unset may not be cached (see Arm ARM Rule
DWZCQ); the implementation of try_update_vmemmap_pte() on arm64 takes
advantage of this fact to replace a PTE without BBM and therefore
without leaving a window open where PE might fault on this translation.

Of course, if some CPUs on the system do not support HW AF management,
clearing the AF will introduce potential faults.
system_supports_bbm_through_af() will return false if any CPUs on the
system do not support HW AF.

Signed-off-by: James Houghton <jthoughton@google.com>
---
 arch/arm64/include/asm/pgtable.h | 60 +++++++++++++++++++++++++++++---
 1 file changed, 56 insertions(+), 4 deletions(-)

diff --git a/arch/arm64/include/asm/pgtable.h b/arch/arm64/include/asm/pgtable.h
index e47c3d010715..f8e66bb22c0c 100644
--- a/arch/arm64/include/asm/pgtable.h
+++ b/arch/arm64/include/asm/pgtable.h
@@ -1249,14 +1249,24 @@ static inline void __pte_clear(struct mm_struct *mm,
 	__set_pte(ptep, __pte(0));
 }
 
-static inline bool __ptep_test_and_clear_young(struct vm_area_struct *vma,
-		unsigned long address, pte_t *ptep)
+/*
+ * Atomically clear the Accessed flag. Return the old value of the PTE.
+ */
+static inline pte_t __ptep_clear_young(pte_t *ptep)
 {
 	atomic64_t *pteval = (atomic64_t *)&pte_val(*ptep);
 	s64 af_mask = PTE_AF;
 
-	/* Atomically clear PTE_AF, checking that it was set before. */
-	return af_mask & atomic64_fetch_andnot_relaxed(af_mask, pteval);
+	/* Atomically clear PTE_AF. */
+	u64 oldval = atomic64_fetch_andnot_relaxed(af_mask, pteval);
+
+	return __pte(oldval);
+}
+
+static inline bool __ptep_test_and_clear_young(struct vm_area_struct *vma,
+		unsigned long address, pte_t *ptep)
+{
+	return pte_young(__ptep_clear_young(ptep));
 }
 
 static inline bool __ptep_clear_flush_young(struct vm_area_struct *vma,
@@ -1734,6 +1744,48 @@ static inline void pte_clear(struct mm_struct *mm,
 	__pte_clear(mm, addr, ptep);
 }
 
+#define __HAVE_ARCH_TRY_UPDATE_VMEMMAP_PTE
+static inline int try_update_vmemmap_pte(unsigned long addr, pte_t *ptep,
+					 const pte_t pte)
+{
+	const int max_attempts = 16;
+	int attempts = 0;
+	pte_t old_pte;
+
+	if (!system_supports_bbm_through_af())
+		return -EOPNOTSUPP;
+
+	/* This routine is only to be used for valid-to-valid transitions. */
+	if (WARN_ON_ONCE(!pte_valid(pte)))
+		return -EINVAL;
+
+	old_pte = __ptep_get(ptep);
+
+	do {
+		if (WARN_ON_ONCE(!pte_valid(old_pte)))
+			return -EINVAL;
+
+		/* We should never get a contiguous PTE here. */
+		if (WARN_ON_ONCE(pte_valid_cont(old_pte)))
+			return -EINVAL;
+
+		if (pte_young(old_pte)) {
+			/* __ptep_clear_young() returns the overwritten PTE */
+			old_pte = pte_mkold(__ptep_clear_young(ptep));
+
+			flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
+		}
+	/*
+	 * Translations without AF cannot be cached, so we can replace
+	 * them without BBM.
+	 */
+	} while (!try_cmpxchg_relaxed(&pte_val(*ptep), &pte_val(old_pte),
+				      pte_val(pte)) &&
+		 ++attempts < max_attempts);
+
+	return attempts == max_attempts ? -EAGAIN : 0;
+}
+
 #define clear_full_ptes clear_full_ptes
 static inline void clear_full_ptes(struct mm_struct *mm, unsigned long addr,
 				pte_t *ptep, unsigned int nr, int full)
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


  parent reply	other threads:[~2026-10-03  0:21 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03  0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
2026-10-03  0:21 ` [PATCH v2 01/20] hugetlb: Don't restore vmemmap of non-HVOed folios on bulk restore error James Houghton
2026-10-03  0:21 ` [PATCH v2 02/20] arm64/pgtable: Clear AF with LDCLR on supported systems James Houghton
2026-10-03  0:21 ` [PATCH v2 03/20] hugetlb_vmemmap: Always flush TLB if needed upon PTE remapping James Houghton
2026-10-03  0:21 ` [PATCH v2 04/20] hugetlb_vmemmap: Leave pages partially HVOed upon restore failure James Houghton
2026-10-03  0:21 ` [PATCH v2 05/20] hugetlb_vmemmap: Use try_update_vmemmap_pte to update in-use PTEs James Houghton
2026-10-03  0:21 ` [PATCH v2 06/20] hugetlb_vmemmap: Allow architectures to dynamically disallow HVO James Houghton
2026-10-03  0:21 ` [PATCH v2 07/20] hugetlb_vmemmap: Disable HVO sysctl if arch doesn't support HVO James Houghton
2026-10-03  0:21 ` [PATCH v2 08/20] hugetlb: Fully initialize tail struct pages of non-pre-HVOed bootmem folios James Houghton
2026-10-03  0:21 ` [PATCH v2 09/20] hugetlb_vmemmap: Allow architectures to make HVO enablement boot-time only James Houghton
2026-10-03  0:21 ` [PATCH v2 10/20] hugetlb_vmemmap: Expose whether HVO is enabled to architecture code James Houghton
2026-10-03  0:21 ` [PATCH v2 11/20] arm64: Add bbm_through_af capability James Houghton
2026-10-03  0:21 ` James Houghton [this message]
2026-10-03  0:21 ` [PATCH v2 13/20] arm64: Support hugetlb vmemmap optimization James Houghton
2026-10-03  0:21 ` [PATCH v2 14/20] hugetlb_vmemmap: Add fault injection for in-place vmemmap PTE updates James Houghton
2026-10-03  0:21 ` [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test James Houghton
2026-10-03  0:21 ` [PATCH v2 16/20 DO-NOT-MERGE] hugetlb_vmemmap: Use try_populate_vmemmap_pmd for replacing in-use PMDs James Houghton
2026-10-03  0:21 ` [PATCH v2 17/20 DO-NOT-MERGE] arm64: Implement try_populate_vmemmap_pmd using AF trick James Houghton
2026-10-03  0:21 ` [PATCH v2 18/20 DO-NOT-MERGE] arm64: Drop BBML3 requirement for HVO James Houghton
2026-10-03  0:21 ` [PATCH v2 19/20 DO-NOT-MERGE] hugetlb_vmemmap: Add fault injection for in-place vmemmap PMD splits James Houghton
2026-10-03  0:21 ` [PATCH v2 20/20 DO-NOT-MERGE] selftests/mm: Add HVO pmd-split fault injection tests James Houghton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261003002123.505555-13-jthoughton@google.com \
    --to=jthoughton@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=catalin.marinas@arm.com \
    --cc=david@kernel.org \
    --cc=fvdl@google.com \
    --cc=linu.cherian@arm.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mark.rutland@arm.com \
    --cc=muchun.song@linux.dev \
    --cc=nikos.nikoleris@arm.com \
    --cc=osalvador@suse.de \
    --cc=rientjes@google.com \
    --cc=ryan.roberts@arm.com \
    --cc=sunnanyong@huawei.com \
    --cc=will@kernel.org \
    --cc=yuzhao@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®