mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode
@ 2026-10-07 11:41 Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
                   ` (14 more replies)
  0 siblings, 15 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Hi All,

This is v8 of the batched PTE updates in lazy MMU mode rework.

Patches 1-9 are a prerequisite series [1] backported from mm-new to master.
It is only posted to facilitate the follow-up s390 lazy MMU rework review.

Patch 10 is a cleanup that could be picked independently and is
a prerequisite to patch 11

Patch 11 is the s390 variant of the ARM64 series [2] and provides
hw_pte_t type enablement

Patches 12,13 are the lazy MMU mode implementation and the main
objective of this rework.

Patches 14,15 are an additional KASAN-based lazy-PTE guard mechanism.

1. https://lore.kernel.org/linux-mm/20260922-pte0-v3-0-5670b8cb9059@arm.com/
2. https://lore.kernel.org/all/20260922-pte0_arm-v2-0-a3f1ddff0a8a@arm.com/

Changes since v7:
- CPUHP_BP_PREPARE_DYN hotplug event is used to (de-)allocate per-cpu data
- lazy_mmu_enabled static key is used to indicate the lazy mmu mode availability
- is_lazy_mmu_active() is implemented using CLIY alternative
- lazy_mmu_count lowcore variable is turned 1 byte to allow CLIY alternative

Changes since v6:
- bottom-halves are disabled on entering and leaving the lazy mmu mode

Changes since v5:
- IPTE optimization is not applied to secure guests [4]
- __kasan_(un)poison_pte() are marked as EXPORT_SYMBOL_GPL() [5]
- PTE table poisoning is not applied to architectures with PTE entry size=
s
  unaligned on KASAN_GRANULE_SIZE [5]
4. https://lore.kernel.org/linux-s390/cover.1783945507.git.agordeev@linux=
.ibm.com/T/#md98724bd3b66d0a0711deb6089fa541f420566d8
5. https://lore.kernel.org/linux-s390/cover.1783945507.git.agordeev@linux=
.ibm.com/T/#m9c8b31b4863416732d0cbbc5f5f63290db4522c1

Changes since v4:
- verified that a presumable sashiko performance regression finding [2]
  although appears valid does not really degrade performance
- applied sashiko suggestion [3] and added "direct-pte-access" kasan bug =
type
2. https://lore.kernel.org/linux-s390/20260623062703.269982B28-agordeev@l=
inux.ibm.com/#r
3. https://lore.kernel.org/linux-s390/20260623061321.269982A64-agordeev@l=
inux.ibm.com/

Changes since v3:
- all prerequisite patches are landed in -next and removed from the serie=
s

Changes since v2:
- lazy_mmu_mode_enable_for_pte_range() renamed to lazy_mmu_mode_enable_wi=
th_ptes()
  (David Hildenbrand)
- patch "mm/pgtable: Fix bogus comment to clear_not_present_full_ptes()"
  is dropped (David Hildenbrand)
- direct PTE dereferencing KASAN sanitizer added (Heiko Carstens)
- CONFIG_IPTE_BATCH option is dropped (Heiko Carstens)
- PTE_POISON changed from zero to 0x800 (Heiko Carstens)
- allocate per-cpu caches on CPU hot-plug (Heiko Carstens)
- introduced a lowcore field for fast lazy mode checking (Heiko Carstens)
- few minor code changes (Heiko Carstens)

Changes since v1:
- lazy_mmu_mode_enable_pte() renamed to lazy_mmu_mode_enable_for_pte_rang=
e()
- lazy_mmu_mode_enable_for_pte_range() semantics clarified
- some sashiko comments addressed [1] including one bug fix [1]
- patches 2-4 added
1. https://sashiko.dev/#/patchset/cover.1774420056.git.agordeev%40linux.i=
bm.com

This series addresses an s390-specific aspect of how page table entries
are modified. In many cases, changing a valid PTE (for example, setting
or clearing a hardware bit) requires issuing an Invalidate Page Table
Entry (IPTE) instruction beforehand.

A disadvantage of the IPTE instruction is that it may initiate a
machine-wide quiesce state. This state acts as an expensive global
hardware lock and should be avoided whenever possible.

Currently, IPTE is invoked for each individual PTE update in most code
paths. However, the instruction itself supports invalidating multiple
PTEs at once, covering up to 256 entries. Using this capability can
significantly reduce the number of quiesce events, with a positive
impact on overall system performance. At present, this feature is not
utilized.

An effort was therefore made to identify kernel code paths that update
large numbers of consecutive PTEs. Such updates can be batched and
handled by a single IPTE invocation, leveraging the hardware support
described above.

A natural candidate for this optimization is page-table walkers that
change attributes of memory ranges and thus modify contiguous ranges
of PTEs. Many memory-management system calls enter lazy MMU mode while
updating such ranges.

This lazy MMU mode can be leveraged to build on the already existing
infrastructure and implement a software-level lazy MMU mechanism,
allowing expensive PTE invalidations on s390 to be batched.

Thanks!

Alexander Gordeev (6):
  s390/mm: Cleanup pXXp_flush_lazy() routines
  s390: Distinguish hardware and software PTEs
  mm: Make lazy MMU mode context-aware
  s390/mm: Batch PTE updates in lazy MMU mode
  mm/kasan: Introduce helpers for lazy MMU mode sanitizer
  s390/mm: Lazy MMU mode sanitizer

Muhammad Usama Anjum (9):
  mm: introduce hw_pte_t for PTE table storage
  mm: rename pointers to software PTE values as ptentp
  mm: use hw_pte_t for generic PTE table storage
  mm: convert PTE table entries in ptep_get()
  mm: convert PTE table entry to pte
  mm: add hw_pte_val for HW PTE storage
  mm/kasan: use hw_pte_t for the early shadow PTE table
  drm/i915: use hw_pte_t for PTE range callbacks
  xen: use hw_pte_t for PTE range callbacks

 MAINTAINERS                                   |   1 +
 arch/s390/Kconfig                             |   2 +
 arch/s390/boot/startup.c                      |   2 +-
 arch/s390/boot/vmem.c                         |  17 +-
 arch/s390/include/asm/gmap_helpers.h          |   2 +-
 arch/s390/include/asm/hugetlb.h               |  18 +-
 arch/s390/include/asm/lowcore.h               |   3 +-
 arch/s390/include/asm/maccess.h               |   2 +-
 arch/s390/include/asm/page.h                  |   3 +-
 arch/s390/include/asm/pgalloc.h               |   6 +-
 arch/s390/include/asm/pgtable.h               | 215 +++++++--
 arch/s390/kernel/uv.c                         |   4 +-
 arch/s390/kvm/s390/pv.c                       |   2 +-
 arch/s390/mm/Makefile                         |   2 +-
 arch/s390/mm/gmap_helpers.c                   |  18 +-
 arch/s390/mm/hugetlbpage.c                    |  20 +-
 arch/s390/mm/lazy_mmu.c                       | 449 ++++++++++++++++++
 arch/s390/mm/maccess.c                        |   2 +-
 arch/s390/mm/pageattr.c                       |  10 +-
 arch/s390/mm/pgtable.c                        |  34 +-
 arch/s390/mm/vmem.c                           |  26 +-
 .../drm/i915/gem/selftests/i915_gem_mman.c    |   4 +-
 drivers/gpu/drm/i915/i915_mm.c                |   4 +-
 drivers/xen/gntdev.c                          |   2 +-
 drivers/xen/privcmd.c                         |   2 +-
 drivers/xen/xenbus/xenbus_client.c            |   2 +-
 drivers/xen/xlate_mmu.c                       |   4 +-
 fs/hugetlbfs/inode.c                          |   3 +-
 fs/proc/task_mmu.c                            |  35 +-
 include/asm-generic/hugetlb.h                 |  15 +-
 include/asm-generic/pgalloc.h                 |   6 +-
 include/asm-generic/tlb.h                     |   5 +-
 include/linux/hugetlb.h                       |  53 ++-
 include/linux/kasan.h                         |  21 +-
 include/linux/mm.h                            |  26 +-
 include/linux/page_table_check.h              |  10 +-
 include/linux/pagewalk.h                      |  10 +-
 include/linux/pgtable.h                       | 127 +++--
 include/linux/pgtable_types.h                 |  23 +
 include/linux/rmap.h                          |   2 +-
 include/linux/swapops.h                       |   6 +-
 include/linux/vmalloc.h                       |   4 +-
 include/trace/events/xen.h                    |  10 +-
 kernel/bpf/arena.c                            |   9 +-
 kernel/events/core.c                          |   3 +-
 mm/Kconfig                                    |   3 +
 mm/damon/ops-common.c                         |   2 +-
 mm/damon/ops-common.h                         |   2 +-
 mm/damon/vaddr.c                              |  20 +-
 mm/debug_vm_pgtable.c                         |   2 +-
 mm/filemap.c                                  |   4 +-
 mm/gup.c                                      |   9 +-
 mm/highmem.c                                  |  15 +-
 mm/hmm.c                                      |   6 +-
 mm/huge_memory.c                              |   4 +-
 mm/hugetlb.c                                  |  60 +--
 mm/hugetlb_vmemmap.c                          |  13 +-
 mm/internal.h                                 |  16 +-
 mm/kasan/common.c                             |  14 +
 mm/kasan/init.c                               |  14 +-
 mm/kasan/kasan.h                              |   2 +
 mm/kasan/report_generic.c                     |   3 +
 mm/kasan/shadow.c                             |   6 +-
 mm/khugepaged.c                               |  50 +-
 mm/ksm.c                                      |  11 +-
 mm/madvise.c                                  |  26 +-
 mm/mapping_dirty_helpers.c                    |   4 +-
 mm/memory-failure.c                           |   6 +-
 mm/memory.c                                   |  86 ++--
 mm/mempolicy.c                                |   4 +-
 mm/migrate.c                                  |   4 +-
 mm/migrate_device.c                           |   4 +-
 mm/mincore.c                                  |   4 +-
 mm/mlock.c                                    |   4 +-
 mm/mprotect.c                                 |  21 +-
 mm/mremap.c                                   |   6 +-
 mm/page_table_check.c                         |   4 +-
 mm/pagewalk.c                                 |   9 +-
 mm/percpu.c                                   |   2 +-
 mm/pgtable-generic.c                          |  20 +-
 mm/ptdump.c                                   |   4 +-
 mm/rmap.c                                     |   6 +-
 mm/sparse-vmemmap.c                           |  22 +-
 mm/swap_state.c                               |   3 +-
 mm/swapfile.c                                 |   5 +-
 mm/userfaultfd.c                              |  32 +-
 mm/util.c                                     |   2 +-
 mm/vmalloc.c                                  |  17 +-
 mm/vmscan.c                                   |   6 +-
 89 files changed, 1269 insertions(+), 512 deletions(-)
 create mode 100644 arch/s390/mm/lazy_mmu.c
 create mode 100644 include/linux/pgtable_types.h

-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp Alexander Gordeev
                   ` (13 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

pte_t is used both for software PTE values and for entries stored in a PTE
table, so pte_t * does not distinguish a pointer to a software PTE
value from a pointer to table storage.

Introduce hw_pte_t as the generic name for a PTE table element. Define it
as a macro alias of pte_t by default. When an architecture selects
ARCH_HAS_HW_PTE_T, define it as a structure containing a pte_t instead.
This preserves the representation while allowing converted architectures
to enforce the distinction at compile time.

Name the generic wrapper structure __hw_pte_t so architectures can
forward-declare it when pgtable_t must be defined before the generic
hw_pte_t typedef is visible. This avoids header-order dependencies.

Keep the C type definitions behind an __ASSEMBLY__ check because
architecture assembly sources can include this header indirectly. Include
asm/page.h so consumers such as linux/vmalloc.h retain the page definitions
they previously obtained from that header.

Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 MAINTAINERS                   |  1 +
 include/linux/pgtable_types.h | 17 +++++++++++++++++
 mm/Kconfig                    |  3 +++
 3 files changed, 21 insertions(+)
 create mode 100644 include/linux/pgtable_types.h

diff --git a/MAINTAINERS b/MAINTAINERS
index 65e8a4b5c90b..76f809c607ae 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -17142,6 +17142,7 @@ F:	include/linux/mmu_notifier.h
 F:	include/linux/pagewalk.h
 F:	include/linux/pgalloc.h
 F:	include/linux/pgtable.h
+F:	include/linux/pgtable_types.h
 F:	include/linux/ptdump.h
 F:	include/linux/vmpressure.h
 F:	include/linux/vmstat.h
diff --git a/include/linux/pgtable_types.h b/include/linux/pgtable_types.h
new file mode 100644
index 000000000000..07da05d375c2
--- /dev/null
+++ b/include/linux/pgtable_types.h
@@ -0,0 +1,17 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_PGTABLE_TYPES_H
+#define _LINUX_PGTABLE_TYPES_H
+
+#include <asm/page.h>
+
+#ifndef __ASSEMBLY__
+
+#ifdef CONFIG_ARCH_HAS_HW_PTE_T
+typedef struct __hw_pte_t { pte_t __pte; } hw_pte_t;
+#else
+#define hw_pte_t pte_t
+#endif
+
+#endif /* !__ASSEMBLY__ */
+
+#endif /* _LINUX_PGTABLE_TYPES_H */
diff --git a/mm/Kconfig b/mm/Kconfig
index 604c58199acb..2e65b93b133d 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -1317,6 +1317,9 @@ comment "GUP_TEST needs to have DEBUG_FS enabled"
 config GUP_GET_PXX_LOW_HIGH
 	bool
 
+config ARCH_HAS_HW_PTE_T
+	bool
+
 config DMAPOOL_TEST
 	tristate "Enable a module to run time tests on dma_pool"
 	depends on HAS_DMA
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage Alexander Gordeev
                   ` (12 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

Some interfaces use pte_t * for a software PTE value rather than for an
entry stored in a PTE table. These pointers must remain pte_t * when
pointers to PTE table storage are converted to hw_pte_t *.

Rename these parameters in the install_pte callback,
write_protect_page(), and guard_install_set_pte() to ptentp. The later
Coccinelle conversion skips pointers named ptentp, allowing it to convert
the remaining PTE table pointers without changing these interfaces.

Some functions already use ptentp for such pointers, including:
- madvise_folio_pte_batch()
- folio_pte_batch_flags()
No need to rename them.

This patch only renames parameters and makes no functional change.

Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 include/linux/pagewalk.h | 2 +-
 mm/ksm.c                 | 4 ++--
 mm/madvise.c             | 4 ++--
 3 files changed, 5 insertions(+), 5 deletions(-)

diff --git a/include/linux/pagewalk.h b/include/linux/pagewalk.h
index b41d7265c01b..c34d826c5e4a 100644
--- a/include/linux/pagewalk.h
+++ b/include/linux/pagewalk.h
@@ -89,7 +89,7 @@ struct mm_walk_ops {
 		       struct mm_walk *walk);
 	void (*post_vma)(struct mm_walk *walk);
 	int (*install_pte)(unsigned long addr, unsigned long next,
-			   pte_t *ptep, struct mm_walk *walk);
+			   pte_t *ptentp, struct mm_walk *walk);
 	enum page_walk_lock walk_lock;
 };
 
diff --git a/mm/ksm.c b/mm/ksm.c
index 49d48d1e0998..892bf4a0d6e4 100644
--- a/mm/ksm.c
+++ b/mm/ksm.c
@@ -1291,7 +1291,7 @@ static u32 calc_checksum(struct page *page)
 }
 
 static int write_protect_page(struct vm_area_struct *vma, struct folio *folio,
-			      pte_t *orig_pte)
+			      pte_t *ptentp)
 {
 	struct mm_struct *mm = vma->vm_mm;
 	DEFINE_FOLIO_VMA_WALK(pvmw, folio, vma, 0, 0);
@@ -1370,7 +1370,7 @@ static int write_protect_page(struct vm_area_struct *vma, struct folio *folio,
 
 		set_pte_at(mm, pvmw.address, pvmw.pte, entry);
 	}
-	*orig_pte = entry;
+	*ptentp = entry;
 	err = 0;
 
 out_unlock:
diff --git a/mm/madvise.c b/mm/madvise.c
index eeee82cf2b3f..c2133b36b24a 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -1110,12 +1110,12 @@ static int guard_install_pte_entry(pte_t *pte, unsigned long addr,
 }
 
 static int guard_install_set_pte(unsigned long addr, unsigned long next,
-				 pte_t *ptep, struct mm_walk *walk)
+				 pte_t *ptentp, struct mm_walk *walk)
 {
 	unsigned long *nr_pages = (unsigned long *)walk->private;
 
 	/* Simply install a PTE marker, this causes segfault on access. */
-	*ptep = make_pte_marker(PTE_MARKER_GUARD);
+	*ptentp = make_pte_marker(PTE_MARKER_GUARD);
 	(*nr_pages)++;
 
 	return 0;
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 04/15] mm: convert PTE table entries in ptep_get() Alexander Gordeev
                   ` (11 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

Generic page-table interfaces use pte_t * for both pointers to PTE table
storage and pointers to software PTE values. Convert the generic
declarations and their MM, fs, and kernel users together so parameters,
return types, callbacks, and local pointers that designate table storage
use hw_pte_t *.

Include linux/pgtable_types.h from headers that expose the converted
interfaces. It continues to provide pgprot_t to vmalloc.h through
asm/page.h.

Keep software PTE values as pte_t and retain pte_t * for value interfaces.

No architecture selects ARCH_HAS_HW_PTE_T at this point, so hw_pte_t
remains an alias of pte_t and this changes the interface vocabulary without
changing representation or behavior.

Most of this mechanical conversion was generated with the Coccinelle script
included in the cover letter. The script deliberately ignores pte_t *
pointers named ptentp because they designate software PTE values. The
result was then audited, and sites the script could not convert were
updated by hand.

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 fs/hugetlbfs/inode.c             |  3 +-
 fs/proc/task_mmu.c               | 33 ++++++-------
 include/asm-generic/hugetlb.h    | 15 +++---
 include/asm-generic/pgalloc.h    |  6 +--
 include/asm-generic/tlb.h        |  5 +-
 include/linux/hugetlb.h          | 50 +++++++++++---------
 include/linux/mm.h               | 26 +++++------
 include/linux/page_table_check.h | 10 ++--
 include/linux/pagewalk.h         |  8 ++--
 include/linux/pgtable.h          | 79 +++++++++++++++++---------------
 include/linux/rmap.h             |  2 +-
 include/linux/swapops.h          |  6 ++-
 include/linux/vmalloc.h          |  4 +-
 include/trace/events/xen.h       | 10 ++--
 kernel/bpf/arena.c               |  9 ++--
 kernel/events/core.c             |  3 +-
 mm/damon/ops-common.c            |  2 +-
 mm/damon/ops-common.h            |  2 +-
 mm/damon/vaddr.c                 | 20 ++++----
 mm/debug_vm_pgtable.c            |  2 +-
 mm/filemap.c                     |  4 +-
 mm/gup.c                         |  9 ++--
 mm/highmem.c                     | 15 +++---
 mm/hmm.c                         |  6 +--
 mm/huge_memory.c                 |  4 +-
 mm/hugetlb.c                     | 60 ++++++++++++------------
 mm/hugetlb_vmemmap.c             | 13 +++---
 mm/internal.h                    | 16 +++----
 mm/kasan/init.c                  | 12 ++---
 mm/kasan/shadow.c                |  6 +--
 mm/khugepaged.c                  | 50 ++++++++++++--------
 mm/ksm.c                         |  7 +--
 mm/madvise.c                     | 14 +++---
 mm/mapping_dirty_helpers.c       |  4 +-
 mm/memory-failure.c              |  6 +--
 mm/memory.c                      | 78 ++++++++++++++++---------------
 mm/mempolicy.c                   |  4 +-
 mm/migrate.c                     |  4 +-
 mm/migrate_device.c              |  4 +-
 mm/mincore.c                     |  4 +-
 mm/mlock.c                       |  4 +-
 mm/mprotect.c                    | 19 ++++----
 mm/mremap.c                      |  4 +-
 mm/page_table_check.c            |  4 +-
 mm/pagewalk.c                    |  9 ++--
 mm/percpu.c                      |  2 +-
 mm/pgtable-generic.c             | 20 ++++----
 mm/ptdump.c                      |  2 +-
 mm/rmap.c                        |  6 +--
 mm/sparse-vmemmap.c              | 22 ++++-----
 mm/swap_state.c                  |  3 +-
 mm/swapfile.c                    |  5 +-
 mm/userfaultfd.c                 | 32 +++++++------
 mm/util.c                        |  2 +-
 mm/vmalloc.c                     | 11 +++--
 mm/vmscan.c                      |  6 +--
 56 files changed, 410 insertions(+), 356 deletions(-)

diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c
index 7611a8470ea2..ddaab714c3e0 100644
--- a/fs/hugetlbfs/inode.c
+++ b/fs/hugetlbfs/inode.c
@@ -328,7 +328,8 @@ static void hugetlb_delete_from_page_cache(struct folio *folio)
 static bool hugetlb_vma_maps_pfn(struct vm_area_struct *vma,
 				unsigned long addr, unsigned long pfn)
 {
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	ptep = hugetlb_walk(vma, addr, huge_page_size(hstate_vma(vma)));
 	if (!ptep)
diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index 5c54aebe2118..2b5af72bbe45 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -1046,7 +1046,7 @@ static void smaps_pte_hole_lookup(unsigned long addr, struct mm_walk *walk)
 #endif
 }
 
-static void smaps_pte_entry(pte_t *pte, unsigned long addr,
+static void smaps_pte_entry(hw_pte_t *pte, unsigned long addr,
 		struct mm_walk *walk)
 {
 	struct mem_size_stats *mss = walk->private;
@@ -1140,7 +1140,7 @@ static int smaps_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 			   struct mm_walk *walk)
 {
 	struct vm_area_struct *vma = walk->vma;
-	pte_t *pte;
+	hw_pte_t *pte;
 	spinlock_t *ptl;
 
 	ptl = pmd_trans_huge_lock(pmd, vma);
@@ -1263,7 +1263,7 @@ static void show_smap_vma_flags(struct seq_file *m, struct vm_area_struct *vma)
 }
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int smaps_hugetlb_range(pte_t *pte, unsigned long hmask,
+static int smaps_hugetlb_range(hw_pte_t *pte, unsigned long hmask,
 				 unsigned long addr, unsigned long end,
 				 struct mm_walk *walk)
 {
@@ -1704,7 +1704,7 @@ static inline bool pte_is_pinned(struct vm_area_struct *vma, unsigned long addr,
 }
 
 static inline void clear_soft_dirty(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *pte)
+		unsigned long addr, hw_pte_t *pte)
 {
 	if (!pgtable_supports_soft_dirty())
 		return;
@@ -1772,7 +1772,8 @@ static int clear_refs_pte_range(pmd_t *pmd, unsigned long addr,
 {
 	struct clear_refs_private *cp = walk->private;
 	struct vm_area_struct *vma = walk->vma;
-	pte_t *pte, ptent;
+	hw_pte_t *pte;
+	pte_t ptent;
 	spinlock_t *ptl;
 	struct folio *folio;
 
@@ -2172,7 +2173,7 @@ static int pagemap_pmd_range(pmd_t *pmdp, unsigned long addr, unsigned long end,
 	struct vm_area_struct *vma = walk->vma;
 	struct pagemapread *pm = walk->private;
 	spinlock_t *ptl;
-	pte_t *pte, *orig_pte;
+	hw_pte_t *pte, *orig_pte;
 	int err = 0;
 
 #ifdef CONFIG_TRANSPARENT_HUGEPAGE
@@ -2210,7 +2211,7 @@ static int pagemap_pmd_range(pmd_t *pmdp, unsigned long addr, unsigned long end,
 
 #ifdef CONFIG_HUGETLB_PAGE
 /* This function walks within one hugetlb entry in the single call */
-static int pagemap_hugetlb_range(pte_t *ptep, unsigned long hmask,
+static int pagemap_hugetlb_range(hw_pte_t *ptep, unsigned long hmask,
 				 unsigned long addr, unsigned long end,
 				 struct mm_walk *walk)
 {
@@ -2501,7 +2502,7 @@ static unsigned long pagemap_page_category(struct pagemap_scan_private *p,
 }
 
 static void make_uffd_wp_pte(struct vm_area_struct *vma,
-			     unsigned long addr, pte_t *pte, pte_t ptent)
+			     unsigned long addr, hw_pte_t *pte, pte_t ptent)
 {
 	if (pte_present(ptent)) {
 		pte_t old_pte;
@@ -2634,7 +2635,7 @@ static unsigned long pagemap_hugetlb_category(struct vm_area_struct *vma,
 }
 
 static void make_uffd_wp_huge_pte(struct vm_area_struct *vma,
-				  unsigned long addr, pte_t *ptep,
+				  unsigned long addr, hw_pte_t *ptep,
 				  pte_t ptent)
 {
 	const unsigned long psize = huge_page_size(hstate_vma(vma));
@@ -2868,7 +2869,7 @@ static int pagemap_scan_pmd_entry(pmd_t *pmd, unsigned long start,
 	struct pagemap_scan_private *p = walk->private;
 	struct vm_area_struct *vma = walk->vma;
 	unsigned long addr, flush_end = 0;
-	pte_t *pte, *start_pte;
+	hw_pte_t *pte, *start_pte;
 	spinlock_t *ptl;
 	int ret;
 
@@ -2962,7 +2963,7 @@ static int pagemap_scan_pmd_entry(pmd_t *pmd, unsigned long start,
 }
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int pagemap_scan_hugetlb_entry(pte_t *ptep, unsigned long hmask,
+static int pagemap_scan_hugetlb_entry(hw_pte_t *ptep, unsigned long hmask,
 				      unsigned long start, unsigned long end,
 				      struct mm_walk *walk)
 {
@@ -3043,7 +3044,7 @@ static int pagemap_scan_hugetlb_hole_wp(struct vm_area_struct *vma,
 	unsigned long psize = huge_page_size(h);
 	struct mm_struct *mm = vma->vm_mm;
 	spinlock_t *ptl;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t pte;
 
 	for (addr = ALIGN_DOWN(addr, psize); addr < end; addr += psize) {
@@ -3427,8 +3428,8 @@ static int gather_pte_stats(pmd_t *pmd, unsigned long addr,
 	struct numa_maps *md = walk->private;
 	struct vm_area_struct *vma = walk->vma;
 	spinlock_t *ptl;
-	pte_t *orig_pte;
-	pte_t *pte;
+	hw_pte_t *orig_pte;
+	hw_pte_t *pte;
 
 #ifdef CONFIG_TRANSPARENT_HUGEPAGE
 	ptl = pmd_trans_huge_lock(pmd, vma);
@@ -3461,7 +3462,7 @@ static int gather_pte_stats(pmd_t *pmd, unsigned long addr,
 	return 0;
 }
 #ifdef CONFIG_HUGETLB_PAGE
-static int gather_hugetlb_stats(pte_t *pte, unsigned long hmask,
+static int gather_hugetlb_stats(hw_pte_t *pte, unsigned long hmask,
 		unsigned long addr, unsigned long end, struct mm_walk *walk)
 {
 	pte_t huge_pte;
@@ -3484,7 +3485,7 @@ static int gather_hugetlb_stats(pte_t *pte, unsigned long hmask,
 }
 
 #else
-static int gather_hugetlb_stats(pte_t *pte, unsigned long hmask,
+static int gather_hugetlb_stats(hw_pte_t *pte, unsigned long hmask,
 		unsigned long addr, unsigned long end, struct mm_walk *walk)
 {
 	return 0;
diff --git a/include/asm-generic/hugetlb.h b/include/asm-generic/hugetlb.h
index 635c41cc3479..3bfad4a99b64 100644
--- a/include/asm-generic/hugetlb.h
+++ b/include/asm-generic/hugetlb.h
@@ -60,7 +60,7 @@ static inline int huge_pte_uffd(pte_t pte)
 
 #ifndef __HAVE_ARCH_HUGE_PTE_CLEAR
 static inline void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
-		    pte_t *ptep, unsigned long sz)
+		    hw_pte_t *ptep, unsigned long sz)
 {
 	pte_clear(mm, addr, ptep);
 }
@@ -68,7 +68,7 @@ static inline void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
 
 #ifndef __HAVE_ARCH_HUGE_SET_HUGE_PTE_AT
 static inline void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, pte_t pte, unsigned long sz)
+		hw_pte_t *ptep, pte_t pte, unsigned long sz)
 {
 	set_pte_at(mm, addr, ptep, pte);
 }
@@ -76,7 +76,7 @@ static inline void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
 
 #ifndef __HAVE_ARCH_HUGE_PTEP_GET_AND_CLEAR
 static inline pte_t huge_ptep_get_and_clear(struct mm_struct *mm,
-		unsigned long addr, pte_t *ptep, unsigned long sz)
+		unsigned long addr, hw_pte_t *ptep, unsigned long sz)
 {
 	return ptep_get_and_clear(mm, addr, ptep);
 }
@@ -84,7 +84,7 @@ static inline pte_t huge_ptep_get_and_clear(struct mm_struct *mm,
 
 #ifndef __HAVE_ARCH_HUGE_PTEP_CLEAR_FLUSH
 static inline pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep)
+		unsigned long addr, hw_pte_t *ptep)
 {
 	return ptep_clear_flush(vma, addr, ptep);
 }
@@ -99,7 +99,7 @@ static inline int huge_pte_none(pte_t pte)
 
 #ifndef __HAVE_ARCH_HUGE_PTEP_SET_WRPROTECT
 static inline void huge_ptep_set_wrprotect(struct mm_struct *mm,
-		unsigned long addr, pte_t *ptep)
+		unsigned long addr, hw_pte_t *ptep)
 {
 	ptep_set_wrprotect(mm, addr, ptep);
 }
@@ -107,7 +107,7 @@ static inline void huge_ptep_set_wrprotect(struct mm_struct *mm,
 
 #ifndef __HAVE_ARCH_HUGE_PTEP_SET_ACCESS_FLAGS
 static inline int huge_ptep_set_access_flags(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep,
+		unsigned long addr, hw_pte_t *ptep,
 		pte_t pte, int dirty)
 {
 	return ptep_set_access_flags(vma, addr, ptep, pte, dirty);
@@ -115,7 +115,8 @@ static inline int huge_ptep_set_access_flags(struct vm_area_struct *vma,
 #endif
 
 #ifndef __HAVE_ARCH_HUGE_PTEP_GET
-static inline pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+static inline pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr,
+		hw_pte_t *ptep)
 {
 	return ptep_get(ptep);
 }
diff --git a/include/asm-generic/pgalloc.h b/include/asm-generic/pgalloc.h
index 051aa1331051..b35e5a2158ad 100644
--- a/include/asm-generic/pgalloc.h
+++ b/include/asm-generic/pgalloc.h
@@ -16,7 +16,7 @@
  *
  * Return: pointer to the allocated memory or %NULL on error
  */
-static inline pte_t *__pte_alloc_one_kernel_noprof(struct mm_struct *mm)
+static inline hw_pte_t *__pte_alloc_one_kernel_noprof(struct mm_struct *mm)
 {
 	struct ptdesc *ptdesc = pagetable_alloc_noprof(GFP_PGTABLE_KERNEL, 0);
 
@@ -40,7 +40,7 @@ static inline pte_t *__pte_alloc_one_kernel_noprof(struct mm_struct *mm)
  *
  * Return: pointer to the allocated memory or %NULL on error
  */
-static inline pte_t *pte_alloc_one_kernel_noprof(struct mm_struct *mm)
+static inline hw_pte_t *pte_alloc_one_kernel_noprof(struct mm_struct *mm)
 {
 	return __pte_alloc_one_kernel_noprof(mm);
 }
@@ -52,7 +52,7 @@ static inline pte_t *pte_alloc_one_kernel_noprof(struct mm_struct *mm)
  * @mm: the mm_struct of the current context
  * @pte: pointer to the memory containing the page table
  */
-static inline void pte_free_kernel(struct mm_struct *mm, pte_t *pte)
+static inline void pte_free_kernel(struct mm_struct *mm, hw_pte_t *pte)
 {
 	pagetable_dtor_free(virt_to_ptdesc(pte));
 }
diff --git a/include/asm-generic/tlb.h b/include/asm-generic/tlb.h
index bdcc2778ac64..2c3517800a9a 100644
--- a/include/asm-generic/tlb.h
+++ b/include/asm-generic/tlb.h
@@ -644,7 +644,8 @@ static inline void tlb_flush_p4d_range(struct mmu_gather *tlb,
 }
 
 #ifndef __tlb_remove_tlb_entry
-static inline void __tlb_remove_tlb_entry(struct mmu_gather *tlb, pte_t *ptep, unsigned long address)
+static inline void __tlb_remove_tlb_entry(struct mmu_gather *tlb,
+		hw_pte_t *ptep, unsigned long address)
 {
 }
 #endif
@@ -670,7 +671,7 @@ static inline void __tlb_remove_tlb_entry(struct mmu_gather *tlb, pte_t *ptep, u
  * consecutive ptes instead of only a single one.
  */
 static inline void tlb_remove_tlb_entries(struct mmu_gather *tlb,
-		pte_t *ptep, unsigned int nr, unsigned long address)
+		hw_pte_t *ptep, unsigned int nr, unsigned long address)
 {
 	tlb_flush_pte_range(tlb, address, PAGE_SIZE * nr);
 	for (;;) {
diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index 16c4c4caa126..bc0b9c65aa1d 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -141,7 +141,7 @@ unsigned long hugetlb_total_pages(void);
 vm_fault_t hugetlb_fault(struct mm_struct *mm, struct vm_area_struct *vma,
 			unsigned long address, unsigned int flags);
 #ifdef CONFIG_USERFAULTFD
-int hugetlb_mfill_atomic_pte(pte_t *dst_pte,
+int hugetlb_mfill_atomic_pte(hw_pte_t *dst_pte,
 			     struct vm_area_struct *dst_vma,
 			     unsigned long dst_addr,
 			     unsigned long src_addr,
@@ -161,7 +161,7 @@ void hugetlb_fix_reserve_counts(struct inode *inode);
 extern struct mutex *hugetlb_fault_mutex_table;
 u32 hugetlb_fault_mutex_hash(struct address_space *mapping, pgoff_t idx);
 
-pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
 		      unsigned long addr, pud_t *pud);
 bool hugetlbfs_pagecache_present(struct hstate *h,
 				 struct vm_area_struct *vma,
@@ -186,22 +186,22 @@ void hugetlb_bootmem_set_nodes(void);
  * which may go down to the lowest PTE level in their huge_pte_offset() and
  * huge_pte_alloc(): to avoid reliance on pte_offset_map() without pte_unmap().
  */
-static inline pte_t *pte_offset_huge(pmd_t *pmd, unsigned long address)
+static inline hw_pte_t *pte_offset_huge(pmd_t *pmd, unsigned long address)
 {
 	return pte_offset_kernel(pmd, address);
 }
-static inline pte_t *pte_alloc_huge(struct mm_struct *mm, pmd_t *pmd,
+static inline hw_pte_t *pte_alloc_huge(struct mm_struct *mm, pmd_t *pmd,
 				    unsigned long address)
 {
 	return pte_alloc(mm, pmd) ? NULL : pte_offset_huge(pmd, address);
 }
 #endif
 
-pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
 			unsigned long addr, unsigned long sz);
 /*
  * huge_pte_offset(): Walk the hugetlb pgtable until the last level PTE.
- * Returns the pte_t* if found, or NULL if the address is not mapped.
+ * Returns the hw_pte_t* if found, or NULL if the address is not mapped.
  *
  * IMPORTANT: we should normally not directly call this function, instead
  * this is only a common interface to implement arch-specific
@@ -236,11 +236,11 @@ pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
  * a concurrent pmd unshare, but it makes sure the pgtable page is safe to
  * access.
  */
-pte_t *huge_pte_offset(struct mm_struct *mm,
+hw_pte_t *huge_pte_offset(struct mm_struct *mm,
 		       unsigned long addr, unsigned long sz);
 unsigned long hugetlb_mask_last_page(struct hstate *h);
 int huge_pmd_unshare(struct mmu_gather *tlb, struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep);
+		unsigned long addr, hw_pte_t *ptep);
 void huge_pmd_unshare_flush(struct mmu_gather *tlb, struct vm_area_struct *vma);
 void adjust_range_if_pmd_sharing_possible(struct vm_area_struct *vma,
 				unsigned long *start, unsigned long *end);
@@ -302,7 +302,8 @@ static inline struct address_space *hugetlb_folio_mapping_lock_write(
 }
 
 static inline int huge_pmd_unshare(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, unsigned long addr, pte_t *ptep)
+		struct vm_area_struct *vma, unsigned long addr,
+		hw_pte_t *ptep)
 {
 	return 0;
 }
@@ -394,7 +395,7 @@ static inline int is_hugepage_only_range(struct mm_struct *mm,
 }
 
 #ifdef CONFIG_USERFAULTFD
-static inline int hugetlb_mfill_atomic_pte(pte_t *dst_pte,
+static inline int hugetlb_mfill_atomic_pte(hw_pte_t *dst_pte,
 					   struct vm_area_struct *dst_vma,
 					   unsigned long dst_addr,
 					   unsigned long src_addr,
@@ -406,7 +407,7 @@ static inline int hugetlb_mfill_atomic_pte(pte_t *dst_pte,
 }
 #endif /* CONFIG_USERFAULTFD */
 
-static inline pte_t *huge_pte_offset(struct mm_struct *mm, unsigned long addr,
+static inline hw_pte_t *huge_pte_offset(struct mm_struct *mm, unsigned long addr,
 					unsigned long sz)
 {
 	return NULL;
@@ -989,7 +990,8 @@ static inline bool htlb_allow_alloc_fallback(enum migrate_reason reason)
 }
 
 static inline spinlock_t *huge_pte_lockptr(struct hstate *h,
-					   struct mm_struct *mm, pte_t *pte)
+					   struct mm_struct *mm,
+					   hw_pte_t *pte)
 {
 	const unsigned long size = huge_page_size(h);
 
@@ -1053,7 +1055,8 @@ static inline void hugetlb_count_sub(long l, struct mm_struct *mm)
 #ifndef huge_ptep_modify_prot_start
 #define huge_ptep_modify_prot_start huge_ptep_modify_prot_start
 static inline pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma,
-						unsigned long addr, pte_t *ptep)
+						unsigned long addr,
+						hw_pte_t *ptep)
 {
 	unsigned long psize = huge_page_size(hstate_vma(vma));
 
@@ -1064,7 +1067,8 @@ static inline pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma,
 #ifndef huge_ptep_modify_prot_commit
 #define huge_ptep_modify_prot_commit huge_ptep_modify_prot_commit
 static inline void huge_ptep_modify_prot_commit(struct vm_area_struct *vma,
-						unsigned long addr, pte_t *ptep,
+						unsigned long addr,
+						hw_pte_t *ptep,
 						pte_t old_pte, pte_t pte)
 {
 	unsigned long psize = huge_page_size(hstate_vma(vma));
@@ -1252,7 +1256,8 @@ static inline bool htlb_allow_alloc_fallback(enum migrate_reason reason)
 }
 
 static inline spinlock_t *huge_pte_lockptr(struct hstate *h,
-					   struct mm_struct *mm, pte_t *pte)
+					   struct mm_struct *mm,
+					   hw_pte_t *pte)
 {
 	return &mm->page_table_lock;
 }
@@ -1269,11 +1274,11 @@ static inline void hugetlb_count_sub(long l, struct mm_struct *mm)
 {
 }
 
-pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep);
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep);
 unsigned long huge_pte_dirty(pte_t pte);
 
 static inline pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
-					  unsigned long addr, pte_t *ptep)
+					  unsigned long addr, hw_pte_t *ptep)
 {
 #ifdef CONFIG_MMU
 	return ptep_get(ptep);
@@ -1283,7 +1288,8 @@ static inline pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
 }
 
 static inline void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-				   pte_t *ptep, pte_t pte, unsigned long sz)
+				   hw_pte_t *ptep, pte_t pte,
+				   unsigned long sz)
 {
 }
 
@@ -1311,7 +1317,7 @@ static inline void hugetlb_bootmem_struct_page_init(void)
 #endif	/* CONFIG_HUGETLB_PAGE */
 
 static inline spinlock_t *huge_pte_lock(struct hstate *h,
-					struct mm_struct *mm, pte_t *pte)
+					struct mm_struct *mm, hw_pte_t *pte)
 {
 	spinlock_t *ptl;
 
@@ -1329,12 +1335,12 @@ static inline __init void hugetlb_cma_reserve(void)
 #endif
 
 #ifdef CONFIG_HUGETLB_PMD_PAGE_TABLE_SHARING
-static inline bool hugetlb_pmd_shared(pte_t *pte)
+static inline bool hugetlb_pmd_shared(hw_pte_t *pte)
 {
 	return ptdesc_pmd_is_shared(virt_to_ptdesc(pte));
 }
 #else
-static inline bool hugetlb_pmd_shared(pte_t *pte)
+static inline bool hugetlb_pmd_shared(hw_pte_t *pte)
 {
 	return false;
 }
@@ -1361,7 +1367,7 @@ bool __vma_private_lock(struct vm_area_struct *vma);
  * Safe version of huge_pte_offset() to check the locks.  See comments
  * above huge_pte_offset().
  */
-static inline pte_t *
+static inline hw_pte_t *
 hugetlb_walk(struct vm_area_struct *vma, unsigned long addr, unsigned long sz)
 {
 #if defined(CONFIG_HUGETLB_PMD_PAGE_TABLE_SHARING) && defined(CONFIG_LOCKDEP)
diff --git a/include/linux/mm.h b/include/linux/mm.h
index dd09c438fa23..b7b2848fdf41 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -779,7 +779,7 @@ struct vm_fault {
 					 * VM_FAULT_ERROR).
 					 */
 	/* These three entries are valid only while holding ptl lock */
-	pte_t *pte;			/* Pointer to pte entry matching
+	hw_pte_t *pte;			/* Pointer to pte entry matching
 					 * the 'address'. NULL if the page
 					 * table hasn't been allocated.
 					 */
@@ -3249,7 +3249,7 @@ struct follow_pfnmap_args {
 	 * The caller shouldn't touch any of these.
 	 */
 	spinlock_t *lock;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	/**
 	 * Outputs:
 	 *
@@ -3586,7 +3586,7 @@ static inline pud_t pud_mkspecial(pud_t pud)
 }
 #endif	/* CONFIG_ARCH_SUPPORTS_PUD_PFNMAP */
 
-extern pte_t *get_locked_pte(struct mm_struct *mm, unsigned long addr,
+extern hw_pte_t *get_locked_pte(struct mm_struct *mm, unsigned long addr,
 			     spinlock_t **ptl);
 
 #ifdef __PAGETABLE_P4D_FOLDED
@@ -3864,7 +3864,7 @@ static inline spinlock_t *pte_lockptr(struct mm_struct *mm, pmd_t *pmd)
 	return ptlock_ptr(page_ptdesc(pmd_page(*pmd)));
 }
 
-static inline spinlock_t *ptep_lockptr(struct mm_struct *mm, pte_t *pte)
+static inline spinlock_t *ptep_lockptr(struct mm_struct *mm, hw_pte_t *pte)
 {
 	BUILD_BUG_ON(IS_ENABLED(CONFIG_HIGHPTE));
 	BUILD_BUG_ON(MAX_PTRS_PER_PTE * sizeof(pte_t) > PAGE_SIZE);
@@ -3895,7 +3895,7 @@ static inline spinlock_t *pte_lockptr(struct mm_struct *mm, pmd_t *pmd)
 {
 	return &mm->page_table_lock;
 }
-static inline spinlock_t *ptep_lockptr(struct mm_struct *mm, pte_t *pte)
+static inline spinlock_t *ptep_lockptr(struct mm_struct *mm, hw_pte_t *pte)
 {
 	return &mm->page_table_lock;
 }
@@ -3936,19 +3936,19 @@ static inline bool pagetable_pte_ctor(struct mm_struct *mm,
 	return true;
 }
 
-pte_t *__pte_offset_map(pmd_t *pmd, unsigned long addr, pmd_t *pmdvalp);
+hw_pte_t *__pte_offset_map(pmd_t *pmd, unsigned long addr, pmd_t *pmdvalp);
 
-static inline pte_t *pte_offset_map(pmd_t *pmd, unsigned long addr)
+static inline hw_pte_t *pte_offset_map(pmd_t *pmd, unsigned long addr)
 {
 	return __pte_offset_map(pmd, addr, NULL);
 }
 
-pte_t *pte_offset_map_lock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_lock(struct mm_struct *mm, pmd_t *pmd,
 			   unsigned long addr, spinlock_t **ptlp);
 
-pte_t *pte_offset_map_ro_nolock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_ro_nolock(struct mm_struct *mm, pmd_t *pmd,
 				unsigned long addr, spinlock_t **ptlp);
-pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
 				unsigned long addr, pmd_t *pmdvalp,
 				spinlock_t **ptlp);
 
@@ -4898,7 +4898,7 @@ static inline bool gup_can_follow_protnone(const struct vm_area_struct *vma,
 	return !vma_is_accessible(vma);
 }
 
-typedef int (*pte_fn_t)(pte_t *pte, unsigned long addr, void *data);
+typedef int (*pte_fn_t)(hw_pte_t *pte, unsigned long addr, void *data);
 extern int apply_to_page_range(struct mm_struct *mm, unsigned long address,
 			       unsigned long size, pte_fn_t fn, void *data);
 extern int apply_to_existing_page_range(struct mm_struct *mm,
@@ -5148,7 +5148,7 @@ void *vmemmap_alloc_block(unsigned long size, int node);
 struct vmem_altmap;
 void *vmemmap_alloc_block_buf(unsigned long size, int node,
 			      struct vmem_altmap *altmap);
-void vmemmap_verify(pte_t *, int, unsigned long, unsigned long);
+void vmemmap_verify(hw_pte_t *, int, unsigned long, unsigned long);
 void vmemmap_set_pmd(pmd_t *pmd, void *p, int node,
 		     unsigned long addr, unsigned long next);
 int vmemmap_check_pmd(pmd_t *pmd, int node,
@@ -5516,7 +5516,7 @@ static inline bool snapshot_page_is_faithful(const struct page_snapshot *ps)
 
 void snapshot_page(struct page_snapshot *ps, const struct page *page);
 
-void map_anon_folio_pte_nopf(struct folio *folio, pte_t *pte,
+void map_anon_folio_pte_nopf(struct folio *folio, hw_pte_t *pte,
 		struct vm_area_struct *vma, unsigned long addr,
 		bool uffd_wp);
 
diff --git a/include/linux/page_table_check.h b/include/linux/page_table_check.h
index 12268a32e8be..933c1028c223 100644
--- a/include/linux/page_table_check.h
+++ b/include/linux/page_table_check.h
@@ -7,6 +7,8 @@
 #ifndef __LINUX_PAGE_TABLE_CHECK_H
 #define __LINUX_PAGE_TABLE_CHECK_H
 
+#include <linux/pgtable_types.h>
+
 #ifdef CONFIG_PAGE_TABLE_CHECK
 #include <linux/jump_label.h>
 
@@ -21,7 +23,7 @@ void __page_table_check_pmd_clear(struct mm_struct *mm, unsigned long addr,
 void __page_table_check_pud_clear(struct mm_struct *mm, unsigned long addr,
 				  pud_t pud);
 void __page_table_check_ptes_set(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, pte_t pte, unsigned int nr);
+		hw_pte_t *ptep, pte_t pte, unsigned int nr);
 void __page_table_check_pmds_set(struct mm_struct *mm, unsigned long addr,
 		pmd_t *pmdp, pmd_t pmd, unsigned int nr);
 void __page_table_check_puds_set(struct mm_struct *mm, unsigned long addr,
@@ -74,7 +76,8 @@ static inline void page_table_check_pud_clear(struct mm_struct *mm,
 }
 
 static inline void page_table_check_ptes_set(struct mm_struct *mm,
-					     unsigned long addr, pte_t *ptep,
+					     unsigned long addr,
+					     hw_pte_t *ptep,
 					     pte_t pte, unsigned int nr)
 {
 	if (static_branch_likely(&page_table_check_disabled))
@@ -137,7 +140,8 @@ static inline void page_table_check_pud_clear(struct mm_struct *mm,
 }
 
 static inline void page_table_check_ptes_set(struct mm_struct *mm,
-					     unsigned long addr, pte_t *ptep,
+					     unsigned long addr,
+					     hw_pte_t *ptep,
 					     pte_t pte, unsigned int nr)
 {
 }
diff --git a/include/linux/pagewalk.h b/include/linux/pagewalk.h
index c34d826c5e4a..7ab2eb39e02c 100644
--- a/include/linux/pagewalk.h
+++ b/include/linux/pagewalk.h
@@ -38,7 +38,7 @@ enum page_walk_lock {
  *			not trigger for any populated ranges.
  * @hugetlb_entry:	if set, called for each hugetlb entry. This hook
  *			function is called with the vma lock held, in order to
- *			protect against a concurrent freeing of the pte_t* or
+ *			protect against a concurrent freeing of the hw_pte_t* or
  *			the ptl. In some cases, the hook function needs to drop
  *			and retake the vma lock in order to avoid deadlocks
  *			while calling other functions. In such cases the hook
@@ -76,11 +76,11 @@ struct mm_walk_ops {
 			 unsigned long next, struct mm_walk *walk);
 	int (*pmd_entry)(pmd_t *pmd, unsigned long addr,
 			 unsigned long next, struct mm_walk *walk);
-	int (*pte_entry)(pte_t *pte, unsigned long addr,
+	int (*pte_entry)(hw_pte_t *pte, unsigned long addr,
 			 unsigned long next, struct mm_walk *walk);
 	int (*pte_hole)(unsigned long addr, unsigned long next,
 			int depth, struct mm_walk *walk);
-	int (*hugetlb_entry)(pte_t *pte, unsigned long hmask,
+	int (*hugetlb_entry)(hw_pte_t *pte, unsigned long hmask,
 			     unsigned long addr, unsigned long next,
 			     struct mm_walk *walk);
 	int (*test_walk)(unsigned long addr, unsigned long next,
@@ -173,7 +173,7 @@ struct folio_walk {
 	struct page *page;
 	enum folio_walk_level level;
 	union {
-		pte_t *ptep;
+		hw_pte_t *ptep;
 		pud_t *pudp;
 		pmd_t *pmdp;
 	};
diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h
index 8c093c119e5a..dad80d264aac 100644
--- a/include/linux/pgtable.h
+++ b/include/linux/pgtable.h
@@ -4,6 +4,7 @@
 
 #include <linux/pfn.h>
 #include <asm/pgtable.h>
+#include <linux/pgtable_types.h>
 
 #define PMD_ORDER	(PMD_SHIFT - PAGE_SHIFT)
 #define PUD_ORDER	(PUD_SHIFT - PAGE_SHIFT)
@@ -93,26 +94,26 @@ static inline void pud_init(void *addr)
 #endif
 
 #ifndef pte_offset_kernel
-static inline pte_t *pte_offset_kernel(pmd_t *pmd, unsigned long address)
+static inline hw_pte_t *pte_offset_kernel(pmd_t *pmd, unsigned long address)
 {
-	return (pte_t *)pmd_page_vaddr(*pmd) + pte_index(address);
+	return (hw_pte_t *)pmd_page_vaddr(*pmd) + pte_index(address);
 }
 #define pte_offset_kernel pte_offset_kernel
 #endif
 
 #ifdef CONFIG_HIGHPTE
 #define __pte_map(pmd, address) \
-	((pte_t *)kmap_local_page(pmd_page(*(pmd))) + pte_index((address)))
+	((hw_pte_t *)kmap_local_page(pmd_page(*(pmd))) + pte_index((address)))
 #define pte_unmap(pte)	do {	\
 	kunmap_local((pte));	\
 	rcu_read_unlock();	\
 } while (0)
 #else
-static inline pte_t *__pte_map(pmd_t *pmd, unsigned long address)
+static inline hw_pte_t *__pte_map(pmd_t *pmd, unsigned long address)
 {
 	return pte_offset_kernel(pmd, address);
 }
-static inline void pte_unmap(pte_t *pte)
+static inline void pte_unmap(hw_pte_t *pte)
 {
 	rcu_read_unlock();
 }
@@ -172,7 +173,7 @@ static inline pmd_t *pmd_off_k(unsigned long va)
 	return pmd_offset(pud_offset(p4d_offset(pgd_offset_k(va), va), va), va);
 }
 
-static inline pte_t *virt_to_kpte(unsigned long vaddr)
+static inline hw_pte_t *virt_to_kpte(unsigned long vaddr)
 {
 	pmd_t *pmd = pmd_off_k(vaddr);
 
@@ -407,7 +408,7 @@ static inline void lazy_mmu_mode_resume(void) {}
  *
  * May be overridden by the architecture, else pte_batch_hint is always 1.
  */
-static inline unsigned int pte_batch_hint(pte_t *ptep, pte_t pte)
+static inline unsigned int pte_batch_hint(hw_pte_t *ptep, pte_t pte)
 {
 	return 1;
 }
@@ -442,7 +443,7 @@ static inline pte_t pte_advance_pfn(pte_t pte, unsigned long nr)
  * to the same folio.  The PTEs are all in the same PMD.
  */
 static inline void set_ptes(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, pte_t pte, unsigned int nr)
+		hw_pte_t *ptep, pte_t pte, unsigned int nr)
 {
 	page_table_check_ptes_set(mm, addr, ptep, pte, nr);
 
@@ -459,7 +460,7 @@ static inline void set_ptes(struct mm_struct *mm, unsigned long addr,
 
 #ifndef __HAVE_ARCH_PTEP_SET_ACCESS_FLAGS
 extern int ptep_set_access_flags(struct vm_area_struct *vma,
-				 unsigned long address, pte_t *ptep,
+				 unsigned long address, hw_pte_t *ptep,
 				 pte_t entry, int dirty);
 #endif
 
@@ -490,7 +491,7 @@ static inline int pudp_set_access_flags(struct vm_area_struct *vma,
 #endif
 
 #ifndef ptep_get
-static inline pte_t ptep_get(pte_t *ptep)
+static inline pte_t ptep_get(hw_pte_t *ptep)
 {
 	return READ_ONCE(*ptep);
 }
@@ -526,7 +527,7 @@ static inline pgd_t pgdp_get(pgd_t *pgdp)
 
 #ifndef __HAVE_ARCH_PTEP_TEST_AND_CLEAR_YOUNG
 static inline bool ptep_test_and_clear_young(struct vm_area_struct *vma,
-		unsigned long address, pte_t *ptep)
+		unsigned long address, hw_pte_t *ptep)
 {
 	pte_t pte = ptep_get(ptep);
 	bool young = true;
@@ -565,7 +566,7 @@ static inline bool pmdp_test_and_clear_young(struct vm_area_struct *vma,
 
 #ifndef __HAVE_ARCH_PTEP_CLEAR_YOUNG_FLUSH
 bool ptep_clear_flush_young(struct vm_area_struct *vma,
-		unsigned long address, pte_t *ptep);
+		unsigned long address, hw_pte_t *ptep);
 #endif
 
 #ifndef __HAVE_ARCH_PMDP_CLEAR_YOUNG_FLUSH
@@ -644,7 +645,7 @@ static inline void arch_check_zapped_pud(struct vm_area_struct *vma, pud_t pud)
 #ifndef __HAVE_ARCH_PTEP_GET_AND_CLEAR
 static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
 				       unsigned long address,
-				       pte_t *ptep)
+				       hw_pte_t *ptep)
 {
 	pte_t pte = ptep_get(ptep);
 	pte_clear(mm, address, ptep);
@@ -673,7 +674,7 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline void clear_young_dirty_ptes(struct vm_area_struct *vma,
-					  unsigned long addr, pte_t *ptep,
+					  unsigned long addr, hw_pte_t *ptep,
 					  unsigned int nr, cydp_t flags)
 {
 	pte_t pte;
@@ -698,7 +699,7 @@ static inline void clear_young_dirty_ptes(struct vm_area_struct *vma,
 #endif
 
 static inline void ptep_clear(struct mm_struct *mm, unsigned long addr,
-			      pte_t *ptep)
+			      hw_pte_t *ptep)
 {
 	pte_t pte = ptep_get(ptep);
 
@@ -739,7 +740,7 @@ static inline void ptep_clear(struct mm_struct *mm, unsigned long addr,
  * present bit set *unless* it is 'l'. Because get_user_pages_fast() only
  * operates on present ptes we're safe.
  */
-static inline pte_t ptep_get_lockless(pte_t *ptep)
+static inline pte_t ptep_get_lockless(hw_pte_t *ptep)
 {
 	pte_t pte;
 
@@ -777,7 +778,7 @@ static inline pmd_t pmdp_get_lockless(pmd_t *pmdp)
  * We require that the PTE can be read atomically.
  */
 #ifndef ptep_get_lockless
-static inline pte_t ptep_get_lockless(pte_t *ptep)
+static inline pte_t ptep_get_lockless(hw_pte_t *ptep)
 {
 	return ptep_get(ptep);
 }
@@ -844,7 +845,8 @@ static inline pud_t pudp_huge_get_and_clear_full(struct vm_area_struct *vma,
 
 #ifndef __HAVE_ARCH_PTEP_GET_AND_CLEAR_FULL
 static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
-					    unsigned long address, pte_t *ptep,
+					    unsigned long address,
+					    hw_pte_t *ptep,
 					    int full)
 {
 	return ptep_get_and_clear(mm, address, ptep);
@@ -872,7 +874,7 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline pte_t get_and_clear_full_ptes(struct mm_struct *mm,
-		unsigned long addr, pte_t *ptep, unsigned int nr, int full)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr, int full)
 {
 	pte_t pte, tmp_pte;
 
@@ -908,7 +910,7 @@ static inline pte_t get_and_clear_full_ptes(struct mm_struct *mm,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline pte_t get_and_clear_ptes(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, unsigned int nr)
+		hw_pte_t *ptep, unsigned int nr)
 {
 	return get_and_clear_full_ptes(mm, addr, ptep, nr, 0);
 }
@@ -933,7 +935,7 @@ static inline pte_t get_and_clear_ptes(struct mm_struct *mm, unsigned long addr,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline void clear_full_ptes(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, unsigned int nr, int full)
+		hw_pte_t *ptep, unsigned int nr, int full)
 {
 	for (;;) {
 		ptep_get_and_clear_full(mm, addr, ptep, full);
@@ -962,7 +964,7 @@ static inline void clear_full_ptes(struct mm_struct *mm, unsigned long addr,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline void clear_ptes(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, unsigned int nr)
+		hw_pte_t *ptep, unsigned int nr)
 {
 	clear_full_ptes(mm, addr, ptep, nr, 0);
 }
@@ -977,13 +979,14 @@ static inline void clear_ptes(struct mm_struct *mm, unsigned long addr,
  */
 #ifndef update_mmu_tlb_range
 static inline void update_mmu_tlb_range(struct vm_area_struct *vma,
-				unsigned long address, pte_t *ptep, unsigned int nr)
+				unsigned long address, hw_pte_t *ptep,
+				unsigned int nr)
 {
 }
 #endif
 
 static inline void update_mmu_tlb(struct vm_area_struct *vma,
-				unsigned long address, pte_t *ptep)
+				unsigned long address, hw_pte_t *ptep)
 {
 	update_mmu_tlb_range(vma, address, ptep, 1);
 }
@@ -1000,7 +1003,7 @@ static inline void update_mmu_tlb(struct vm_area_struct *vma,
  * The PTEs are all in the same PMD.
  */
 static inline void clear_nonpresent_ptes(struct mm_struct *mm,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	(void)addr;
 
@@ -1016,7 +1019,7 @@ static inline void clear_nonpresent_ptes(struct mm_struct *mm,
 #ifndef __HAVE_ARCH_PTEP_CLEAR_FLUSH
 extern pte_t ptep_clear_flush(struct vm_area_struct *vma,
 			      unsigned long address,
-			      pte_t *ptep);
+			      hw_pte_t *ptep);
 #endif
 
 #ifndef __HAVE_ARCH_PMDP_HUGE_CLEAR_FLUSH
@@ -1044,7 +1047,8 @@ static inline pmd_t pmd_mkwrite(pmd_t pmd, struct vm_area_struct *vma)
 
 #ifndef __HAVE_ARCH_PTEP_SET_WRPROTECT
 struct mm_struct;
-static inline void ptep_set_wrprotect(struct mm_struct *mm, unsigned long address, pte_t *ptep)
+static inline void ptep_set_wrprotect(struct mm_struct *mm, unsigned long address,
+				      hw_pte_t *ptep)
 {
 	pte_t old_pte = ptep_get(ptep);
 	set_pte_at(mm, address, ptep, pte_wrprotect(old_pte));
@@ -1070,7 +1074,7 @@ static inline void ptep_set_wrprotect(struct mm_struct *mm, unsigned long addres
  * ptep_try_set as an identity macro. The generic stub returns false, which is
  * correct for callers that fall through to oops on failure.
  */
-static inline bool ptep_try_set(pte_t *ptep, pte_t new_pte)
+static inline bool ptep_try_set(hw_pte_t *ptep, pte_t new_pte)
 {
 	return false;
 }
@@ -1113,7 +1117,7 @@ static inline void flush_tlb_before_set(unsigned long addr)
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline void wrprotect_ptes(struct mm_struct *mm, unsigned long addr,
-		pte_t *ptep, unsigned int nr)
+		hw_pte_t *ptep, unsigned int nr)
 {
 	for (;;) {
 		ptep_set_wrprotect(mm, addr, ptep);
@@ -1144,7 +1148,7 @@ static inline void wrprotect_ptes(struct mm_struct *mm, unsigned long addr,
  * pages that belong to the same folio.  The PTEs are all in the same PMD.
  */
 static inline bool clear_flush_young_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	bool young = false;
 
@@ -1181,7 +1185,7 @@ static inline bool clear_flush_young_ptes(struct vm_area_struct *vma,
  * Returns: whether any PTE was young.
  */
 static inline bool test_and_clear_young_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	bool young = false;
 
@@ -1585,7 +1589,7 @@ static inline int pmd_none_or_clear_bad(pmd_t *pmd)
 
 static inline pte_t __ptep_modify_prot_start(struct vm_area_struct *vma,
 					     unsigned long addr,
-					     pte_t *ptep)
+					     hw_pte_t *ptep)
 {
 	/*
 	 * Get the current pte state, but zero it out to make it
@@ -1597,7 +1601,7 @@ static inline pte_t __ptep_modify_prot_start(struct vm_area_struct *vma,
 
 static inline void __ptep_modify_prot_commit(struct vm_area_struct *vma,
 					     unsigned long addr,
-					     pte_t *ptep, pte_t pte)
+					     hw_pte_t *ptep, pte_t pte)
 {
 	/*
 	 * The pte is non-present, so there's no hardware state to
@@ -1623,7 +1627,7 @@ static inline void __ptep_modify_prot_commit(struct vm_area_struct *vma,
  */
 static inline pte_t ptep_modify_prot_start(struct vm_area_struct *vma,
 					   unsigned long addr,
-					   pte_t *ptep)
+					   hw_pte_t *ptep)
 {
 	return __ptep_modify_prot_start(vma, addr, ptep);
 }
@@ -1636,7 +1640,8 @@ static inline pte_t ptep_modify_prot_start(struct vm_area_struct *vma,
  */
 static inline void ptep_modify_prot_commit(struct vm_area_struct *vma,
 					   unsigned long addr,
-					   pte_t *ptep, pte_t old_pte, pte_t pte)
+					   hw_pte_t *ptep, pte_t old_pte,
+					   pte_t pte)
 {
 	__ptep_modify_prot_commit(vma, addr, ptep, pte);
 }
@@ -1667,7 +1672,7 @@ static inline void ptep_modify_prot_commit(struct vm_area_struct *vma,
  */
 #ifndef modify_prot_start_ptes
 static inline pte_t modify_prot_start_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	pte_t pte, tmp_pte;
 
@@ -1707,7 +1712,7 @@ static inline pte_t modify_prot_start_ptes(struct vm_area_struct *vma,
  */
 #ifndef modify_prot_commit_ptes
 static inline void modify_prot_commit_ptes(struct vm_area_struct *vma, unsigned long addr,
-		pte_t *ptep, pte_t old_pte, pte_t pte, unsigned int nr)
+		hw_pte_t *ptep, pte_t old_pte, pte_t pte, unsigned int nr)
 {
 	int i;
 
diff --git a/include/linux/rmap.h b/include/linux/rmap.h
index 0b332770abee..d28a7fb6f2c0 100644
--- a/include/linux/rmap.h
+++ b/include/linux/rmap.h
@@ -868,7 +868,7 @@ struct page_vma_mapped_walk {
 	struct vm_area_struct *vma;
 	unsigned long address;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 	spinlock_t *ptl;
 	unsigned int flags;
 	bool pgoff_is_anon : 1;
diff --git a/include/linux/swapops.h b/include/linux/swapops.h
index e7d0d529f3e0..40d2fd920558 100644
--- a/include/linux/swapops.h
+++ b/include/linux/swapops.h
@@ -216,7 +216,8 @@ static inline swp_entry_t make_migration_entry_dirty(swp_entry_t entry)
 
 extern void migration_entry_wait(struct mm_struct *mm, pmd_t *pmd,
 					unsigned long address);
-extern void migration_entry_wait_huge(struct vm_area_struct *vma, unsigned long addr, pte_t *pte);
+extern void migration_entry_wait_huge(struct vm_area_struct *vma, unsigned long addr,
+				      hw_pte_t *pte);
 #else  /* CONFIG_MIGRATION */
 static inline swp_entry_t make_readable_migration_entry(pgoff_t offset)
 {
@@ -236,7 +237,8 @@ static inline swp_entry_t make_writable_migration_entry(pgoff_t offset)
 static inline void migration_entry_wait(struct mm_struct *mm, pmd_t *pmd,
 					unsigned long address) { }
 static inline void migration_entry_wait_huge(struct vm_area_struct *vma,
-					     unsigned long addr, pte_t *pte) { }
+					     unsigned long addr,
+					     hw_pte_t *pte) { }
 
 static inline swp_entry_t make_migration_entry_young(swp_entry_t entry)
 {
diff --git a/include/linux/vmalloc.h b/include/linux/vmalloc.h
index aed121d729b0..666ff2b3741d 100644
--- a/include/linux/vmalloc.h
+++ b/include/linux/vmalloc.h
@@ -8,7 +8,7 @@
 #include <linux/init.h>
 #include <linux/list.h>
 #include <linux/llist.h>
-#include <asm/page.h>		/* pgprot_t */
+#include <linux/pgtable_types.h>	/* pgprot_t, hw_pte_t */
 #include <linux/rbtree.h>
 #include <linux/overflow.h>
 
@@ -120,7 +120,7 @@ static inline unsigned long arch_vmap_pte_range_map_size(unsigned long addr, uns
 
 #ifndef arch_vmap_pte_range_unmap_size
 static inline unsigned long arch_vmap_pte_range_unmap_size(unsigned long addr,
-							   pte_t *ptep)
+							   hw_pte_t *ptep)
 {
 	return PAGE_SIZE;
 }
diff --git a/include/trace/events/xen.h b/include/trace/events/xen.h
index ad384969e2cb..1972d50b043a 100644
--- a/include/trace/events/xen.h
+++ b/include/trace/events/xen.h
@@ -138,10 +138,10 @@ TRACE_EVENT(xen_mc_extend_args,
 TRACE_DEFINE_SIZEOF(pteval_t);
 
 TRACE_EVENT(xen_mmu_set_pte,
-	    TP_PROTO(pte_t *ptep, pte_t pteval),
+	    TP_PROTO(hw_pte_t *ptep, pte_t pteval),
 	    TP_ARGS(ptep, pteval),
 	    TP_STRUCT__entry(
-		    __field(pte_t *, ptep)
+		    __field(hw_pte_t *, ptep)
 		    __field(pteval_t, pteval)
 		    ),
 	    TP_fast_assign(__entry->ptep = ptep;
@@ -207,12 +207,12 @@ TRACE_EVENT(xen_mmu_set_p4d,
 
 DECLARE_EVENT_CLASS(xen_mmu_ptep_modify_prot,
 	    TP_PROTO(struct mm_struct *mm, unsigned long addr,
-		     pte_t *ptep, pte_t pteval),
+		     hw_pte_t *ptep, pte_t pteval),
 	    TP_ARGS(mm, addr, ptep, pteval),
 	    TP_STRUCT__entry(
 		    __field(struct mm_struct *, mm)
 		    __field(unsigned long, addr)
-		    __field(pte_t *, ptep)
+		    __field(hw_pte_t *, ptep)
 		    __field(pteval_t, pteval)
 		    ),
 	    TP_fast_assign(__entry->mm = mm;
@@ -227,7 +227,7 @@ DECLARE_EVENT_CLASS(xen_mmu_ptep_modify_prot,
 #define DEFINE_XEN_MMU_PTEP_MODIFY_PROT(name)				\
 	DEFINE_EVENT(xen_mmu_ptep_modify_prot, name,			\
 		     TP_PROTO(struct mm_struct *mm, unsigned long addr,	\
-			      pte_t *ptep, pte_t pteval),		\
+			      hw_pte_t *ptep, pte_t pteval),		\
 		     TP_ARGS(mm, addr, ptep, pteval))
 
 DEFINE_XEN_MMU_PTEP_MODIFY_PROT(xen_mmu_ptep_modify_prot_start);
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 7b6847200b43..62005229b2e9 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -155,7 +155,7 @@ struct clear_range_data {
 	struct llist_head *free_pages;
 };
 
-static int apply_range_set_cb(pte_t *pte, unsigned long addr, void *data)
+static int apply_range_set_cb(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct apply_range_data *d = data;
 	struct page *page;
@@ -207,7 +207,7 @@ static void flush_vmap_cache(unsigned long start, unsigned long size)
 	flush_cache_vmap(start, start + size);
 }
 
-static int apply_range_clear_cb(pte_t *pte, unsigned long addr, void *data)
+static int apply_range_clear_cb(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct clear_range_data *d = data;
 	pte_t old_pte;
@@ -238,7 +238,8 @@ static int apply_range_clear_cb(pte_t *pte, unsigned long addr, void *data)
 	return 0;
 }
 
-static int apply_range_set_scratch_cb(pte_t *pte, unsigned long addr, void *data)
+static int apply_range_set_scratch_cb(hw_pte_t *pte, unsigned long addr,
+				      void *data)
 {
 	struct page *scratch_page = data;
 
@@ -340,7 +341,7 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 	return ERR_PTR(err);
 }
 
-static int existing_page_cb(pte_t *ptep, unsigned long addr, void *data)
+static int existing_page_cb(hw_pte_t *ptep, unsigned long addr, void *data)
 {
 	struct bpf_arena *arena = data;
 	struct page *page;
diff --git a/kernel/events/core.c b/kernel/events/core.c
index 7846d70be57f..0eaa0e745423 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -8534,7 +8534,8 @@ static u64 perf_get_pgtable_size(struct mm_struct *mm, unsigned long addr)
 	p4d_t *p4dp, p4d;
 	pud_t *pudp, pud;
 	pmd_t *pmdp, pmd;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	pgdp = pgd_offset(mm, addr);
 	pgd = pgdp_get(pgdp);
diff --git a/mm/damon/ops-common.c b/mm/damon/ops-common.c
index 8fc61d06d358..2c73842dd2aa 100644
--- a/mm/damon/ops-common.c
+++ b/mm/damon/ops-common.c
@@ -39,7 +39,7 @@ struct folio *damon_get_folio(unsigned long pfn)
 	return folio;
 }
 
-void damon_ptep_mkold(pte_t *pte, struct vm_area_struct *vma, unsigned long addr)
+void damon_ptep_mkold(hw_pte_t *pte, struct vm_area_struct *vma, unsigned long addr)
 {
 	pte_t pteval = ptep_get(pte);
 	struct folio *folio;
diff --git a/mm/damon/ops-common.h b/mm/damon/ops-common.h
index 38d295488fa1..34b7e5715fcf 100644
--- a/mm/damon/ops-common.h
+++ b/mm/damon/ops-common.h
@@ -7,7 +7,7 @@
 
 struct folio *damon_get_folio(unsigned long pfn);
 
-void damon_ptep_mkold(pte_t *pte, struct vm_area_struct *vma, unsigned long addr);
+void damon_ptep_mkold(hw_pte_t *pte, struct vm_area_struct *vma, unsigned long addr);
 void damon_pmdp_mkold(pmd_t *pmd, struct vm_area_struct *vma, unsigned long addr);
 void damon_folio_mkold(struct folio *folio);
 bool damon_folio_young(struct folio *folio);
diff --git a/mm/damon/vaddr.c b/mm/damon/vaddr.c
index 04ee2a2c6a4d..3ce4be80d80d 100644
--- a/mm/damon/vaddr.c
+++ b/mm/damon/vaddr.c
@@ -268,7 +268,7 @@ static void damon_va_walk_page_range(struct mm_struct *mm, unsigned long start,
 static int damon_mkold_pmd_entry(pmd_t *pmd, unsigned long addr,
 		unsigned long next, struct mm_walk *walk)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	spinlock_t *ptl;
 
 	ptl = pmd_trans_huge_lock(pmd, walk->vma);
@@ -306,7 +306,7 @@ static bool damon_hugetlb_ptep_mkold(pte_t *pte, struct mm_struct *mm,
 	return true;
 }
 
-static void damon_hugetlb_mkold(pte_t *pte, struct mm_struct *mm,
+static void damon_hugetlb_mkold(hw_pte_t *pte, struct mm_struct *mm,
 				struct vm_area_struct *vma, unsigned long addr)
 {
 	bool referenced = false;
@@ -327,7 +327,7 @@ static void damon_hugetlb_mkold(pte_t *pte, struct mm_struct *mm,
 	folio_put(folio);
 }
 
-static int damon_mkold_hugetlb_entry(pte_t *pte, unsigned long hmask,
+static int damon_mkold_hugetlb_entry(hw_pte_t *pte, unsigned long hmask,
 				     unsigned long addr, unsigned long end,
 				     struct mm_walk *walk)
 {
@@ -396,7 +396,7 @@ struct damon_young_walk_private {
 static int damon_young_pmd_entry(pmd_t *pmd, unsigned long addr,
 		unsigned long next, struct mm_walk *walk)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	pte_t ptent;
 	spinlock_t *ptl;
 	struct folio *folio;
@@ -440,7 +440,7 @@ static int damon_young_pmd_entry(pmd_t *pmd, unsigned long addr,
 }
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int damon_young_hugetlb_entry(pte_t *pte, unsigned long hmask,
+static int damon_young_hugetlb_entry(hw_pte_t *pte, unsigned long hmask,
 				     unsigned long addr, unsigned long end,
 				     struct mm_walk *walk)
 {
@@ -529,7 +529,7 @@ static unsigned int damon_va_check_accesses(struct damon_ctx *ctx)
 
 static bool damos_va_filter_young_match(struct damos_filter *filter,
 		struct folio *folio, struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, pmd_t *pmdp)
+		unsigned long addr, hw_pte_t *ptep, pmd_t *pmdp)
 {
 	bool young = false;
 
@@ -551,7 +551,7 @@ static bool damos_va_filter_young_match(struct damos_filter *filter,
 
 static bool damos_va_filter_out(struct damos *scheme, struct folio *folio,
 		struct vm_area_struct *vma, unsigned long addr,
-		pte_t *ptep, pmd_t *pmdp)
+		hw_pte_t *ptep, pmd_t *pmdp)
 {
 	struct damos_filter *filter;
 	bool matched;
@@ -648,7 +648,8 @@ static int damos_va_migrate_pmd_entry(pmd_t *pmd, unsigned long addr,
 	struct damos_migrate_dests *dests = &s->migrate_dests;
 	struct folio *folio;
 	spinlock_t *ptl;
-	pte_t *start_pte, *pte, ptent;
+	hw_pte_t *start_pte, *pte;
+	pte_t ptent;
 	int nr;
 
 #ifdef CONFIG_TRANSPARENT_HUGEPAGE
@@ -808,7 +809,8 @@ static int damos_va_stat_pmd_entry(pmd_t *pmd, unsigned long addr,
 	struct vm_area_struct *vma = walk->vma;
 	struct folio *folio;
 	spinlock_t *ptl;
-	pte_t *start_pte, *pte, ptent;
+	hw_pte_t *start_pte, *pte;
+	pte_t ptent;
 	int nr;
 
 #ifdef CONFIG_TRANSPARENT_HUGEPAGE
diff --git a/mm/debug_vm_pgtable.c b/mm/debug_vm_pgtable.c
index 2875fd22d7bb..54e194d2a585 100644
--- a/mm/debug_vm_pgtable.c
+++ b/mm/debug_vm_pgtable.c
@@ -50,7 +50,7 @@ struct pgtable_debug_args {
 	p4d_t			*p4dp;
 	pud_t			*pudp;
 	pmd_t			*pmdp;
-	pte_t			*ptep;
+	hw_pte_t		*ptep;
 
 	p4d_t			*start_p4dp;
 	pud_t			*start_pudp;
diff --git a/mm/filemap.c b/mm/filemap.c
index 00fd89cf6f55..98f89dcc5560 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -3490,7 +3490,7 @@ static vm_fault_t filemap_fault_recheck_pte_none(struct vm_fault *vmf)
 {
 	struct vm_area_struct *vma = vmf->vma;
 	vm_fault_t ret = 0;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	/*
 	 * We might have COW'ed a pagecache folio and might now have an mlocked
@@ -3796,7 +3796,7 @@ static vm_fault_t filemap_map_folio_range(struct vm_fault *vmf,
 	vm_fault_t ret = 0;
 	struct page *page = folio_page(folio, start);
 	unsigned int count = 0;
-	pte_t *old_ptep = vmf->pte;
+	hw_pte_t *old_ptep = vmf->pte;
 	unsigned long addr0;
 
 	/*
diff --git a/mm/gup.c b/mm/gup.c
index eb898ea1ee22..8108e4015d4e 100644
--- a/mm/gup.c
+++ b/mm/gup.c
@@ -761,7 +761,7 @@ static struct page *follow_huge_pmd(struct vm_area_struct *vma,
 #endif	/* CONFIG_PGTABLE_HAS_HUGE_LEAVES */
 
 static int follow_pfn_pte(struct vm_area_struct *vma, unsigned long address,
-		pte_t *pte, unsigned int flags)
+		hw_pte_t *pte, unsigned int flags)
 {
 	if (flags & FOLL_TOUCH) {
 		pte_t orig_entry = ptep_get(pte);
@@ -806,7 +806,8 @@ static struct page *follow_page_pte(struct vm_area_struct *vma,
 	struct folio *folio;
 	struct page *page;
 	spinlock_t *ptl;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 	int ret;
 
 	ptep = pte_offset_map_lock(mm, pmd, address, &ptl);
@@ -1035,7 +1036,7 @@ static int get_gate_page(struct mm_struct *mm, unsigned long address,
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 	pte_t entry;
 	int ret = -EFAULT;
 
@@ -2831,7 +2832,7 @@ static int gup_fast_pte_range(pmd_t pmd, pmd_t *pmdp, unsigned long addr,
 		int *nr)
 {
 	int ret = 0;
-	pte_t *ptep, *ptem;
+	hw_pte_t *ptep, *ptem;
 
 	ptem = ptep = pte_offset_map(&pmd, addr);
 	if (!ptep)
diff --git a/mm/highmem.c b/mm/highmem.c
index a33e41183951..87f94cac6106 100644
--- a/mm/highmem.c
+++ b/mm/highmem.c
@@ -141,7 +141,7 @@ EXPORT_SYMBOL(__totalhigh_pages);
 static int pkmap_count[LAST_PKMAP];
 static  __cacheline_aligned_in_smp DEFINE_SPINLOCK(kmap_lock);
 
-pte_t *pkmap_page_table;
+hw_pte_t *pkmap_page_table;
 
 /*
  * Most architectures have no use for kmap_high_get(), so let's abstract
@@ -532,9 +532,9 @@ static inline bool kmap_high_unmap_local(unsigned long vaddr)
 	return false;
 }
 
-static pte_t *__kmap_pte;
+static hw_pte_t *__kmap_pte;
 
-static pte_t *kmap_get_pte(unsigned long vaddr, int idx)
+static hw_pte_t *kmap_get_pte(unsigned long vaddr, int idx)
 {
 	if (IS_ENABLED(CONFIG_KMAP_LOCAL_NON_LINEAR_PTE_ARRAY))
 		/*
@@ -549,8 +549,9 @@ static pte_t *kmap_get_pte(unsigned long vaddr, int idx)
 
 void *__kmap_local_pfn_prot(unsigned long pfn, pgprot_t prot)
 {
-	pte_t pteval, *kmap_pte;
 	unsigned long vaddr;
+	hw_pte_t *kmap_pte;
+	pte_t pteval;
 	int idx;
 
 	/*
@@ -597,7 +598,7 @@ EXPORT_SYMBOL(__kmap_local_page_prot);
 void kunmap_local_indexed(const void *vaddr)
 {
 	unsigned long addr = (unsigned long) vaddr & PAGE_MASK;
-	pte_t *kmap_pte;
+	hw_pte_t *kmap_pte;
 	int idx;
 
 	if (addr < __fix_to_virt(FIX_KMAP_END) ||
@@ -646,7 +647,7 @@ EXPORT_SYMBOL(kunmap_local_indexed);
 void __kmap_local_sched_out(void)
 {
 	struct task_struct *tsk = current;
-	pte_t *kmap_pte;
+	hw_pte_t *kmap_pte;
 	int i;
 
 	/* Clear kmaps */
@@ -683,7 +684,7 @@ void __kmap_local_sched_out(void)
 void __kmap_local_sched_in(void)
 {
 	struct task_struct *tsk = current;
-	pte_t *kmap_pte;
+	hw_pte_t *kmap_pte;
 	int i;
 
 	/* Restore kmaps */
diff --git a/mm/hmm.c b/mm/hmm.c
index 2f1e98c6b644..0c959611eda0 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -240,7 +240,7 @@ static inline unsigned long pte_to_hmm_pfn_flags(struct hmm_range *range,
 }
 
 static int hmm_vma_handle_pte(struct mm_walk *walk, unsigned long addr,
-			      unsigned long end, pmd_t *pmdp, pte_t *ptep,
+			      unsigned long end, pmd_t *pmdp, hw_pte_t *ptep,
 			      unsigned long *hmm_pfn)
 {
 	struct hmm_vma_walk *hmm_vma_walk = walk->private;
@@ -411,7 +411,7 @@ static int hmm_vma_walk_pmd(pmd_t *pmdp,
 		&range->hmm_pfns[(start - range->start) >> PAGE_SHIFT];
 	unsigned long npages = (end - start) >> PAGE_SHIFT;
 	unsigned long addr = start;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pmd_t pmd;
 
 again:
@@ -547,7 +547,7 @@ static int hmm_vma_walk_pud(pud_t *pudp, unsigned long start, unsigned long end,
 #endif
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int hmm_vma_walk_hugetlb_entry(pte_t *pte, unsigned long hmask,
+static int hmm_vma_walk_hugetlb_entry(hw_pte_t *pte, unsigned long hmask,
 				      unsigned long start, unsigned long end,
 				      struct mm_walk *walk)
 {
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 1e5d68acf62a..b70ea24f73fa 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3142,7 +3142,7 @@ static void __split_huge_zero_page_pmd(struct vm_area_struct *vma,
 	pgtable_t pgtable;
 	pmd_t _pmd, old_pmd;
 	unsigned long addr;
-	pte_t *pte;
+	hw_pte_t *pte;
 	int i;
 
 	/*
@@ -3192,7 +3192,7 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
 	bool soft_dirty, uffd_wp = false, young = false, write = false;
 	bool anon_exclusive = false, dirty = false;
 	unsigned long addr;
-	pte_t *pte;
+	hw_pte_t *pte;
 	int i;
 
 	VM_BUG_ON(haddr & ~HPAGE_PMD_MASK);
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index cea25773a6c9..1e5048861b07 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -120,7 +120,7 @@ static void hugetlb_vma_lock_free(struct vm_area_struct *vma);
 static void hugetlb_vma_lock_alloc(struct vm_area_struct *vma);
 static void __hugetlb_vma_unlock_write_free(struct vm_area_struct *vma);
 static int __huge_pmd_unshare(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, unsigned long addr, pte_t *ptep,
+		struct vm_area_struct *vma, unsigned long addr, hw_pte_t *ptep,
 		bool check_locks);
 static void hugetlb_unshare_pmds(struct vm_area_struct *vma,
 		unsigned long start, unsigned long end, bool take_locks);
@@ -4853,7 +4853,7 @@ static pte_t make_huge_pte(struct vm_area_struct *vma, struct folio *folio,
 }
 
 static void set_huge_ptep_writable(struct vm_area_struct *vma,
-				   unsigned long address, pte_t *ptep)
+				   unsigned long address, hw_pte_t *ptep)
 {
 	pte_t entry;
 
@@ -4863,14 +4863,14 @@ static void set_huge_ptep_writable(struct vm_area_struct *vma,
 }
 
 static void set_huge_ptep_maybe_writable(struct vm_area_struct *vma,
-					 unsigned long address, pte_t *ptep)
+					 unsigned long address, hw_pte_t *ptep)
 {
 	if (vma->vm_flags & VM_WRITE)
 		set_huge_ptep_writable(vma, address, ptep);
 }
 
 static void
-hugetlb_install_folio(struct vm_area_struct *vma, pte_t *ptep, unsigned long addr,
+hugetlb_install_folio(struct vm_area_struct *vma, hw_pte_t *ptep, unsigned long addr,
 		      struct folio *new_folio, pte_t old, unsigned long sz)
 {
 	pte_t newpte = make_huge_pte(vma, new_folio, true);
@@ -4896,7 +4896,8 @@ int copy_hugetlb_page_range(struct mm_struct *dst, struct mm_struct *src,
 			    struct vm_area_struct *dst_vma,
 			    struct vm_area_struct *src_vma)
 {
-	pte_t *src_pte, *dst_pte, entry;
+	hw_pte_t *src_pte, *dst_pte;
+	pte_t entry;
 	struct folio *pte_folio;
 	unsigned long addr;
 	bool cow = vma_is_cow_mapping(src_vma);
@@ -5088,7 +5089,8 @@ int copy_hugetlb_page_range(struct mm_struct *dst, struct mm_struct *src,
 }
 
 static void move_huge_pte(struct vm_area_struct *vma, unsigned long old_addr,
-			  unsigned long new_addr, pte_t *src_pte, pte_t *dst_pte,
+			  unsigned long new_addr, hw_pte_t *src_pte,
+			  hw_pte_t *dst_pte,
 			  unsigned long sz)
 {
 	bool need_clear_uffd_wp = vma_has_uffd_without_event_remap(vma);
@@ -5150,7 +5152,7 @@ int move_hugetlb_page_tables(struct vm_area_struct *vma,
 	struct mm_struct *mm = vma->vm_mm;
 	unsigned long old_end = old_addr + len;
 	unsigned long last_addr_mask;
-	pte_t *src_pte, *dst_pte;
+	hw_pte_t *src_pte, *dst_pte;
 	struct mmu_notifier_range range;
 	struct mmu_gather tlb;
 
@@ -5214,7 +5216,7 @@ void __unmap_hugepage_range(struct mmu_gather *tlb, struct vm_area_struct *vma,
 	struct mm_struct *mm = vma->vm_mm;
 	const bool folio_provided = !!folio;
 	unsigned long address;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t pte;
 	spinlock_t *ptl;
 	struct hstate *h = hstate_vma(vma);
@@ -5748,7 +5750,7 @@ static inline vm_fault_t hugetlb_handle_userfault(struct vm_fault *vmf,
  * false if pte changed or is changing.
  */
 static bool hugetlb_pte_stable(struct hstate *h, struct mm_struct *mm, unsigned long addr,
-			       pte_t *ptep, pte_t old_pte)
+			       hw_pte_t *ptep, pte_t old_pte)
 {
 	spinlock_t *ptl;
 	bool same;
@@ -6273,7 +6275,7 @@ static struct folio *alloc_hugetlb_folio_vma(struct hstate *h,
  * Used by userfaultfd UFFDIO_* ioctls. Based on userfaultfd's mfill_atomic_pte
  * with modifications for hugetlb pages.
  */
-int hugetlb_mfill_atomic_pte(pte_t *dst_pte,
+int hugetlb_mfill_atomic_pte(hw_pte_t *dst_pte,
 			     struct vm_area_struct *dst_vma,
 			     unsigned long dst_addr,
 			     unsigned long src_addr,
@@ -6332,7 +6334,7 @@ int hugetlb_mfill_atomic_pte(pte_t *dst_pte,
 
 		folio = alloc_hugetlb_folio(dst_vma, dst_addr, false);
 		if (IS_ERR(folio)) {
-			pte_t *actual_pte = hugetlb_walk(dst_vma, dst_addr, PMD_SIZE);
+			hw_pte_t *actual_pte = hugetlb_walk(dst_vma, dst_addr, PMD_SIZE);
 			if (actual_pte) {
 				ret = -EEXIST;
 				goto out;
@@ -6500,7 +6502,7 @@ long hugetlb_change_protection(struct vm_area_struct *vma,
 {
 	struct mm_struct *mm = vma->vm_mm;
 	unsigned long start = address;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t pte;
 	struct hstate *h = hstate_vma(vma);
 	long pages = 0, psize = huge_page_size(h);
@@ -6985,15 +6987,15 @@ void adjust_range_if_pmd_sharing_possible(struct vm_area_struct *vma,
  * racing tasks could either miss the sharing (see huge_pte_offset) or select a
  * bad pmd for sharing.
  */
-pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
 		      unsigned long addr, pud_t *pud)
 {
 	struct address_space *mapping = vma->vm_file->f_mapping;
 	const pgoff_t idx = linear_page_index(vma, addr);
 	struct vm_area_struct *svma;
 	unsigned long saddr;
-	pte_t *spte = NULL;
-	pte_t *pte;
+	hw_pte_t *spte = NULL;
+	hw_pte_t *pte;
 
 	i_mmap_lock_read(mapping);
 	mapping_rmap_tree_foreach(svma, mapping, idx, idx) {
@@ -7024,13 +7026,13 @@ pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
 	}
 	spin_unlock(&mm->page_table_lock);
 out:
-	pte = (pte_t *)pmd_alloc(mm, pud, addr);
+	pte = (hw_pte_t *)pmd_alloc(mm, pud, addr);
 	i_mmap_unlock_read(mapping);
 	return pte;
 }
 
 static int __huge_pmd_unshare(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, unsigned long addr, pte_t *ptep,
+		struct vm_area_struct *vma, unsigned long addr, hw_pte_t *ptep,
 		bool check_locks)
 {
 	unsigned long sz = huge_page_size(hstate_vma(vma));
@@ -7071,7 +7073,7 @@ static int __huge_pmd_unshare(struct mmu_gather *tlb,
  *	    was not a shared PMD table.
  */
 int huge_pmd_unshare(struct mmu_gather *tlb, struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep)
+		unsigned long addr, hw_pte_t *ptep)
 {
 	return __huge_pmd_unshare(tlb, vma, addr, ptep, /*check_locks=*/true);
 }
@@ -7101,21 +7103,21 @@ void huge_pmd_unshare_flush(struct mmu_gather *tlb, struct vm_area_struct *vma)
 
 #else /* !CONFIG_HUGETLB_PMD_PAGE_TABLE_SHARING */
 
-pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pmd_share(struct mm_struct *mm, struct vm_area_struct *vma,
 		      unsigned long addr, pud_t *pud)
 {
 	return NULL;
 }
 
 static int __huge_pmd_unshare(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, unsigned long addr, pte_t *ptep,
+		struct vm_area_struct *vma, unsigned long addr, hw_pte_t *ptep,
 		bool check_locks)
 {
 	return 0;
 }
 
 int huge_pmd_unshare(struct mmu_gather *tlb, struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep)
+		unsigned long addr, hw_pte_t *ptep)
 {
 	return 0;
 }
@@ -7136,13 +7138,13 @@ bool want_pmd_share(struct vm_area_struct *vma, unsigned long addr)
 #endif /* CONFIG_HUGETLB_PMD_PAGE_TABLE_SHARING */
 
 #ifdef CONFIG_ARCH_WANT_GENERAL_HUGETLB
-pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
 			unsigned long addr, unsigned long sz)
 {
 	pgd_t *pgd;
 	p4d_t *p4d;
 	pud_t *pud;
-	pte_t *pte = NULL;
+	hw_pte_t *pte = NULL;
 
 	pgd = pgd_offset(mm, addr);
 	p4d = p4d_alloc(mm, pgd, addr);
@@ -7151,13 +7153,13 @@ pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
 	pud = pud_alloc(mm, p4d, addr);
 	if (pud) {
 		if (sz == PUD_SIZE) {
-			pte = (pte_t *)pud;
+			pte = (hw_pte_t *)pud;
 		} else {
 			BUG_ON(sz != PMD_SIZE);
 			if (want_pmd_share(vma, addr) && pud_none(*pud))
 				pte = huge_pmd_share(mm, vma, addr, pud);
 			else
-				pte = (pte_t *)pmd_alloc(mm, pud, addr);
+				pte = (hw_pte_t *)pmd_alloc(mm, pud, addr);
 		}
 	}
 
@@ -7179,7 +7181,7 @@ pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
  * size @sz doesn't match the hugepage size at this level of the page
  * table.
  */
-pte_t *huge_pte_offset(struct mm_struct *mm,
+hw_pte_t *huge_pte_offset(struct mm_struct *mm,
 		       unsigned long addr, unsigned long sz)
 {
 	pgd_t *pgd;
@@ -7197,14 +7199,14 @@ pte_t *huge_pte_offset(struct mm_struct *mm,
 	pud = pud_offset(p4d, addr);
 	if (sz == PUD_SIZE)
 		/* must be pud huge, non-present or none */
-		return (pte_t *)pud;
+		return (hw_pte_t *)pud;
 	if (!pud_present(*pud))
 		return NULL;
 	/* must have a valid entry and size to go further */
 
 	pmd = pmd_offset(pud, addr);
 	/* must be pmd huge, non-present or none */
-	return (pte_t *)pmd;
+	return (hw_pte_t *)pmd;
 }
 
 /*
@@ -7383,7 +7385,7 @@ static void hugetlb_unshare_pmds(struct vm_area_struct *vma,
 	struct mmu_gather tlb;
 	unsigned long address;
 	spinlock_t *ptl;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	if (!(vma->vm_flags & VM_MAYSHARE))
 		return;
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 917db0984143..2850c19c4528 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -33,7 +33,7 @@
  *			operations.
  */
 struct vmemmap_remap_walk {
-	void			(*remap_pte)(pte_t *pte, unsigned long addr,
+	void			(*remap_pte)(hw_pte_t *pte, unsigned long addr,
 					     struct vmemmap_remap_walk *walk);
 
 	unsigned long		nr_walked;
@@ -55,7 +55,7 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
 	pmd_t __pmd;
 	int i;
 	unsigned long addr = start;
-	pte_t *pgtable;
+	hw_pte_t *pgtable;
 
 	pgtable = pte_alloc_one_kernel(&init_mm);
 	if (!pgtable)
@@ -64,7 +64,8 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
 	pmd_populate_kernel(&init_mm, &__pmd, pgtable);
 
 	for (i = 0; i < PTRS_PER_PTE; i++, addr += PAGE_SIZE) {
-		pte_t entry, *pte;
+		pte_t entry;
+		hw_pte_t *pte;
 		pgprot_t pgprot = PAGE_KERNEL;
 
 		entry = mk_pte(head + i, pgprot);
@@ -136,7 +137,7 @@ static int vmemmap_pmd_entry(pmd_t *pmd, unsigned long addr,
 	return vmemmap_split_pmd(pmd, head, addr & PMD_MASK, vmemmap_walk);
 }
 
-static int vmemmap_pte_entry(pte_t *pte, unsigned long addr,
+static int vmemmap_pte_entry(hw_pte_t *pte, unsigned long addr,
 			     unsigned long next, struct mm_walk *walk)
 {
 	struct vmemmap_remap_walk *vmemmap_walk = walk->private;
@@ -198,7 +199,7 @@ static void free_vmemmap_page_list(struct list_head *list)
 		free_vmemmap_page(page);
 }
 
-static void vmemmap_remap_pte(pte_t *pte, unsigned long addr,
+static void vmemmap_remap_pte(hw_pte_t *pte, unsigned long addr,
 			      struct vmemmap_remap_walk *walk)
 {
 	struct page *page = pte_page(ptep_get(pte));
@@ -232,7 +233,7 @@ static void vmemmap_remap_pte(pte_t *pte, unsigned long addr,
 	set_pte_at(&init_mm, addr, pte, entry);
 }
 
-static void vmemmap_restore_pte(pte_t *pte, unsigned long addr,
+static void vmemmap_restore_pte(hw_pte_t *pte, unsigned long addr,
 				struct vmemmap_remap_walk *walk)
 {
 	struct page *src = pte_page(ptep_get(pte)), *dst;
diff --git a/mm/internal.h b/mm/internal.h
index 38b1165212c9..c4a3adc6e586 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -274,7 +274,7 @@ void unmap_vmas(struct mmu_gather *tlb, struct unmap_desc *unmap);
 #ifdef CONFIG_MMU
 
 bool cond_install_uffd_wp_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, pte_t pte,
+		unsigned long addr, hw_pte_t *ptep, pte_t pte,
 		unsigned long nr_ptes);
 
 static inline void get_anon_vma(struct anon_vma *anon_vma)
@@ -410,7 +410,7 @@ static inline pte_t __pte_batch_clear_ignored(pte_t pte, fpb_t flags)
  * Return: the number of table entries in the batch.
  */
 static inline unsigned int folio_pte_batch_flags(struct folio *folio,
-		struct vm_area_struct *vma, pte_t *ptep, pte_t *ptentp,
+		struct vm_area_struct *vma, hw_pte_t *ptep, pte_t *ptentp,
 		unsigned int max_nr, fpb_t flags)
 {
 	bool any_writable = false, any_young = false, any_dirty = false;
@@ -464,7 +464,7 @@ static inline unsigned int folio_pte_batch_flags(struct folio *folio,
 	return min(nr, max_nr);
 }
 
-unsigned int folio_pte_batch(struct folio *folio, pte_t *ptep, pte_t pte,
+unsigned int folio_pte_batch(struct folio *folio, hw_pte_t *ptep, pte_t pte,
 		unsigned int max_nr);
 
 /**
@@ -521,11 +521,11 @@ static inline pte_t pte_next_swp_offset(pte_t pte)
  *
  * Return: the number of table entries in the batch.
  */
-static inline int swap_pte_batch(pte_t *start_ptep, int max_nr, pte_t pte)
+static inline int swap_pte_batch(hw_pte_t *start_ptep, int max_nr, pte_t pte)
 {
 	pte_t expected_pte = pte_next_swp_offset(pte);
-	const pte_t *end_ptep = start_ptep + max_nr;
-	pte_t *ptep = start_ptep + 1;
+	const hw_pte_t *end_ptep = start_ptep + max_nr;
+	hw_pte_t *ptep = start_ptep + 1;
 
 	VM_WARN_ON(max_nr < 1);
 	VM_WARN_ON(!softleaf_is_swap(softleaf_from_pte(pte)));
@@ -1563,7 +1563,7 @@ static inline void maybe_rmap_unlock_action(struct vm_area_struct *vma,
 
 #ifdef CONFIG_MMU_NOTIFIER
 static inline bool clear_flush_young_ptes_notify(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	bool young;
 
@@ -1584,7 +1584,7 @@ static inline bool pmdp_clear_flush_young_notify(struct vm_area_struct *vma,
 }
 
 static inline bool test_and_clear_young_ptes_notify(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, unsigned int nr)
+		unsigned long addr, hw_pte_t *ptep, unsigned int nr)
 {
 	bool young;
 
diff --git a/mm/kasan/init.c b/mm/kasan/init.c
index 66a883887987..30c5266eaf37 100644
--- a/mm/kasan/init.c
+++ b/mm/kasan/init.c
@@ -92,7 +92,7 @@ static __init void *early_alloc(size_t size, int node)
 static void __ref zero_pte_populate(pmd_t *pmd, unsigned long addr,
 				unsigned long end)
 {
-	pte_t *pte = pte_offset_kernel(pmd, addr);
+	hw_pte_t *pte = pte_offset_kernel(pmd, addr);
 	pte_t zero_pte;
 
 	zero_pte = pfn_pte(PFN_DOWN(__pa_symbol(kasan_early_shadow_page)),
@@ -122,7 +122,7 @@ static int __ref zero_pmd_populate(pud_t *pud, unsigned long addr,
 		}
 
 		if (pmd_none(*pmd)) {
-			pte_t *p;
+			hw_pte_t *p;
 
 			if (slab_is_available())
 				p = pte_alloc_one_kernel(&init_mm);
@@ -281,9 +281,9 @@ int __ref kasan_populate_early_shadow(const void *shadow_start,
 	return 0;
 }
 
-static void kasan_free_pte(pte_t *pte_start, pmd_t *pmd)
+static void kasan_free_pte(hw_pte_t *pte_start, pmd_t *pmd)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	int i;
 
 	for (i = 0; i < PTRS_PER_PTE; i++) {
@@ -341,7 +341,7 @@ static void kasan_free_p4d(p4d_t *p4d_start, pgd_t *pgd)
 	pgd_clear(pgd);
 }
 
-static void kasan_remove_pte_table(pte_t *pte, unsigned long addr,
+static void kasan_remove_pte_table(hw_pte_t *pte, unsigned long addr,
 				unsigned long end)
 {
 	unsigned long next;
@@ -369,7 +369,7 @@ static void kasan_remove_pmd_table(pmd_t *pmd, unsigned long addr,
 	unsigned long next;
 
 	for (; addr < end; addr = next, pmd++) {
-		pte_t *pte;
+		hw_pte_t *pte;
 
 		next = pmd_addr_end(addr, end);
 
diff --git a/mm/kasan/shadow.c b/mm/kasan/shadow.c
index d286e0a04543..86fc7ed45dcd 100644
--- a/mm/kasan/shadow.c
+++ b/mm/kasan/shadow.c
@@ -189,7 +189,7 @@ static bool shadow_mapped(unsigned long addr)
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	if (pgd_none(*pgd))
 		return false;
@@ -297,7 +297,7 @@ struct vmalloc_populate_data {
 	struct page **pages;
 };
 
-static int kasan_populate_vmalloc_pte(pte_t *ptep, unsigned long addr,
+static int kasan_populate_vmalloc_pte(hw_pte_t *ptep, unsigned long addr,
 				      void *_data)
 {
 	struct vmalloc_populate_data *data = _data;
@@ -465,7 +465,7 @@ int __kasan_populate_vmalloc(unsigned long addr, unsigned long size, gfp_t gfp_m
 	return 0;
 }
 
-static int kasan_depopulate_vmalloc_pte(pte_t *ptep, unsigned long addr,
+static int kasan_depopulate_vmalloc_pte(hw_pte_t *ptep, unsigned long addr,
 					void *unused)
 {
 	pte_t pte;
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 75639298efc2..b1b042547970 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -639,8 +639,8 @@ static void release_pte_folio(struct folio *folio)
 	folio_putback_lru(folio);
 }
 
-static void release_pte_pages(pte_t *pte, pte_t *_pte,
-		struct list_head *compound_pagelist)
+static void release_pte_pages(hw_pte_t *pte, hw_pte_t *_pte,
+			      struct list_head *compound_pagelist)
 {
 	struct folio *folio, *tmp;
 
@@ -685,7 +685,8 @@ static void count_collapse_event(unsigned int order, enum vm_event_item vm_event
 }
 
 static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct *vma,
-		unsigned long start_addr, pte_t *pte, struct collapse_control *cc,
+		unsigned long start_addr, hw_pte_t *pte,
+		struct collapse_control *cc,
 		unsigned int order, struct list_head *compound_pagelist)
 {
 	const unsigned int max_ptes_none = collapse_max_ptes_none(cc, vma, order);
@@ -694,7 +695,7 @@ static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct *vma,
 	struct page *page = NULL;
 	struct folio *folio = NULL;
 	unsigned long addr = start_addr;
-	pte_t *_pte;
+	hw_pte_t *_pte;
 	int none_or_zero = 0, shared = 0, referenced = 0;
 	enum scan_result result = SCAN_FAIL;
 
@@ -840,16 +841,18 @@ static enum scan_result __collapse_huge_page_isolate(struct vm_area_struct *vma,
 	return result;
 }
 
-static void __collapse_huge_page_copy_succeeded(pte_t *pte,
-		struct vm_area_struct *vma, unsigned long address,
-		spinlock_t *ptl, unsigned int order,
-		struct list_head *compound_pagelist)
+static void __collapse_huge_page_copy_succeeded(hw_pte_t *pte,
+						struct vm_area_struct *vma,
+						unsigned long address,
+						spinlock_t *ptl,
+						unsigned int order,
+						struct list_head *compound_pagelist)
 {
 	const unsigned long nr_pages = 1UL << order;
 	unsigned long end = address + (PAGE_SIZE * nr_pages);
 	struct folio *src, *tmp;
 	pte_t pteval;
-	pte_t *_pte;
+	hw_pte_t *_pte;
 	unsigned int nr_ptes;
 
 	for (_pte = pte; _pte < pte + nr_pages; _pte += nr_ptes,
@@ -904,9 +907,11 @@ static void __collapse_huge_page_copy_succeeded(pte_t *pte,
 	}
 }
 
-static void __collapse_huge_page_copy_failed(pte_t *pte,
-		pmd_t *pmd, pmd_t orig_pmd, struct vm_area_struct *vma,
-		unsigned int order, struct list_head *compound_pagelist)
+static void __collapse_huge_page_copy_failed(hw_pte_t *pte,
+					     pmd_t *pmd, pmd_t orig_pmd,
+					     struct vm_area_struct *vma,
+					     unsigned int order,
+					     struct list_head *compound_pagelist)
 {
 	const unsigned long nr_pages = 1UL << order;
 	spinlock_t *pmd_ptl;
@@ -942,10 +947,14 @@ static void __collapse_huge_page_copy_failed(pte_t *pte,
  * @ptl: lock on raw pages' PTEs
  * @compound_pagelist: list that stores compound pages
  */
-static enum scan_result __collapse_huge_page_copy(pte_t *pte, struct folio *folio,
-		pmd_t *pmd, pmd_t orig_pmd, struct vm_area_struct *vma,
-		unsigned long address, spinlock_t *ptl, unsigned int order,
-		struct list_head *compound_pagelist)
+static enum scan_result __collapse_huge_page_copy(hw_pte_t *pte,
+						  struct folio *folio,
+						  pmd_t *pmd, pmd_t orig_pmd,
+						  struct vm_area_struct *vma,
+						  unsigned long address,
+						  spinlock_t *ptl,
+						  unsigned int order,
+						  struct list_head *compound_pagelist)
 {
 	const unsigned long nr_pages = 1UL << order;
 	unsigned int i;
@@ -1164,7 +1173,7 @@ static enum scan_result __collapse_huge_page_swapin(struct mm_struct *mm,
 	vm_fault_t ret = 0;
 	unsigned long addr, end = start_addr + (PAGE_SIZE << order);
 	enum scan_result result;
-	pte_t *pte = NULL;
+	hw_pte_t *pte = NULL;
 	spinlock_t *ptl;
 
 	for (addr = start_addr; addr < end; addr += PAGE_SIZE) {
@@ -1292,7 +1301,7 @@ static enum scan_result collapse_huge_page(struct mm_struct *mm, unsigned long s
 	const unsigned long end_addr = start_addr + (PAGE_SIZE << order);
 	LIST_HEAD(compound_pagelist);
 	pmd_t *pmd, _pmd;
-	pte_t *pte = NULL;
+	hw_pte_t *pte = NULL;
 	pgtable_t pgtable;
 	struct folio *folio;
 	spinlock_t *pmd_ptl, *pte_ptl;
@@ -1606,7 +1615,8 @@ static enum scan_result collapse_scan_pmd(struct mm_struct *mm,
 	unsigned int max_ptes_none = collapse_max_ptes_none(cc, vma, HPAGE_PMD_ORDER);
 	enum tva_type tva_flags = cc->is_khugepaged ? TVA_KHUGEPAGED : TVA_FORCED_COLLAPSE;
 	pmd_t *pmd;
-	pte_t *pte, *_pte, pteval;
+	hw_pte_t *pte, *_pte;
+	pte_t pteval;
 	int i;
 	int none_or_zero = 0, shared = 0, referenced = 0;
 	enum scan_result result = SCAN_FAIL;
@@ -1860,7 +1870,7 @@ static enum scan_result try_collapse_pte_mapped_thp(struct mm_struct *mm, unsign
 	unsigned long end = haddr + HPAGE_PMD_SIZE;
 	struct vm_area_struct *vma = vma_lookup(mm, haddr);
 	struct folio *folio;
-	pte_t *start_pte, *pte;
+	hw_pte_t *start_pte, *pte;
 	pmd_t *pmd, pgt_pmd;
 	spinlock_t *pml = NULL, *ptl;
 	int i;
diff --git a/mm/ksm.c b/mm/ksm.c
index 892bf4a0d6e4..e78708651498 100644
--- a/mm/ksm.c
+++ b/mm/ksm.c
@@ -618,7 +618,7 @@ static int break_ksm_pmd_entry(pmd_t *pmdp, unsigned long addr, unsigned long en
 {
 	unsigned long *found_addr = (unsigned long *) walk->private;
 	struct mm_struct *mm = walk->mm;
-	pte_t *start_ptep, *ptep;
+	hw_pte_t *start_ptep, *ptep;
 	spinlock_t *ptl;
 	int found = 0;
 
@@ -1398,7 +1398,7 @@ static int replace_page(struct vm_area_struct *vma, struct page *page,
 	struct folio *folio = page_folio(page);
 	pmd_t *pmd;
 	pmd_t pmde;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t newpte;
 	spinlock_t *ptl;
 	unsigned long addr;
@@ -2529,7 +2529,8 @@ static int ksm_next_page_pmd_entry(pmd_t *pmdp, unsigned long addr, unsigned lon
 {
 	struct ksm_next_page_arg *private = walk->private;
 	struct vm_area_struct *vma = walk->vma;
-	pte_t *start_ptep = NULL, *ptep, pte;
+	hw_pte_t *start_ptep = NULL, *ptep;
+	pte_t pte;
 	struct mm_struct *mm = walk->mm;
 	struct folio *folio;
 	struct page *page;
diff --git a/mm/madvise.c b/mm/madvise.c
index c2133b36b24a..40ce0d5940ef 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -198,7 +198,7 @@ static int swapin_walk_pmd_entry(pmd_t *pmd, unsigned long start,
 {
 	struct vm_area_struct *vma = walk->private;
 	struct swap_io_ctx ctx = {};
-	pte_t *ptep = NULL;
+	hw_pte_t *ptep = NULL;
 	spinlock_t *ptl;
 	unsigned long addr;
 
@@ -350,7 +350,7 @@ static inline bool can_do_file_pageout(struct vm_area_struct *vma)
 }
 
 static inline int madvise_folio_pte_batch(unsigned long addr, unsigned long end,
-					  struct folio *folio, pte_t *ptep,
+					  struct folio *folio, hw_pte_t *ptep,
 					  pte_t *ptentp)
 {
 	int max_nr = (end - addr) / PAGE_SIZE;
@@ -368,7 +368,8 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
 	bool pageout = private->pageout;
 	struct mm_struct *mm = tlb->mm;
 	struct vm_area_struct *vma = walk->vma;
-	pte_t *start_pte, *pte, ptent;
+	hw_pte_t *start_pte, *pte;
+	pte_t ptent;
 	spinlock_t *ptl;
 	struct folio *folio = NULL;
 	LIST_HEAD(folio_list);
@@ -667,7 +668,8 @@ static int madvise_free_pte_range(pmd_t *pmd, unsigned long addr,
 	struct mm_struct *mm = tlb->mm;
 	struct vm_area_struct *vma = walk->vma;
 	spinlock_t *ptl;
-	pte_t *start_pte, *pte, ptent;
+	hw_pte_t *start_pte, *pte;
+	pte_t ptent;
 	struct folio *folio;
 	int nr_swap = 0;
 	unsigned long next;
@@ -1092,7 +1094,7 @@ static int guard_install_pmd_entry(pmd_t *pmd, unsigned long addr,
 	return pmd_trans_huge(pmdval);
 }
 
-static int guard_install_pte_entry(pte_t *pte, unsigned long addr,
+static int guard_install_pte_entry(hw_pte_t *pte, unsigned long addr,
 				   unsigned long next, struct mm_walk *walk)
 {
 	pte_t pteval = ptep_get(pte);
@@ -1235,7 +1237,7 @@ static int guard_remove_pmd_entry(pmd_t *pmd, unsigned long addr,
 	return 0;
 }
 
-static int guard_remove_pte_entry(pte_t *pte, unsigned long addr,
+static int guard_remove_pte_entry(hw_pte_t *pte, unsigned long addr,
 				  unsigned long next, struct mm_walk *walk)
 {
 	pte_t ptent = ptep_get(pte);
diff --git a/mm/mapping_dirty_helpers.c b/mm/mapping_dirty_helpers.c
index e0efa36e0a07..dcbd39912f81 100644
--- a/mm/mapping_dirty_helpers.c
+++ b/mm/mapping_dirty_helpers.c
@@ -31,7 +31,7 @@ struct wp_walk {
  * The function write-protects a pte and records the range in
  * virtual address space of touched ptes for efficient range TLB flushes.
  */
-static int wp_pte(pte_t *pte, unsigned long addr, unsigned long end,
+static int wp_pte(hw_pte_t *pte, unsigned long addr, unsigned long end,
 		  struct mm_walk *walk)
 {
 	struct wp_walk *wpwalk = walk->private;
@@ -86,7 +86,7 @@ struct clean_walk {
  * in the address_space, as well as the first and last of the bits
  * touched.
  */
-static int clean_record_pte(pte_t *pte, unsigned long addr,
+static int clean_record_pte(hw_pte_t *pte, unsigned long addr,
 			    unsigned long end, struct mm_walk *walk)
 {
 	struct wp_walk *wpwalk = walk->private;
diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index a8b03e2920ba..a89f3fde47a5 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -342,7 +342,7 @@ static unsigned long dev_pagemap_mapping_shift(struct vm_area_struct *vma,
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 	pte_t ptent;
 
 	VM_BUG_ON_VMA(address == -EFAULT, vma);
@@ -742,7 +742,7 @@ static int hwpoison_pte_range(pmd_t *pmdp, unsigned long addr,
 {
 	struct hwpoison_walk *hwp = walk->private;
 	int ret = 0;
-	pte_t *ptep, *mapped_pte;
+	hw_pte_t *ptep, *mapped_pte;
 	spinlock_t *ptl;
 
 	ptl = pmd_trans_huge_lock(pmdp, walk->vma);
@@ -770,7 +770,7 @@ static int hwpoison_pte_range(pmd_t *pmdp, unsigned long addr,
 }
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int hwpoison_hugetlb_range(pte_t *ptep, unsigned long hmask,
+static int hwpoison_hugetlb_range(hw_pte_t *ptep, unsigned long hmask,
 			    unsigned long addr, unsigned long end,
 			    struct mm_walk *walk)
 {
diff --git a/mm/memory.c b/mm/memory.c
index 8b0c2c735d3d..48d6ca89b0fe 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -462,7 +462,7 @@ int __pte_alloc(struct mm_struct *mm, pmd_t *pmd)
 
 int __pte_alloc_kernel(pmd_t *pmd)
 {
-	pte_t *new = pte_alloc_one_kernel(&init_mm);
+	hw_pte_t *new = pte_alloc_one_kernel(&init_mm);
 	if (!new)
 		return -ENOMEM;
 
@@ -946,7 +946,7 @@ struct page *vm_normal_page_pud(struct vm_area_struct *vma,
  */
 static void restore_exclusive_pte(struct vm_area_struct *vma,
 		struct folio *folio, struct page *page, unsigned long address,
-		pte_t *ptep, pte_t orig_pte)
+		hw_pte_t *ptep, pte_t orig_pte)
 {
 	pte_t pte;
 
@@ -983,7 +983,7 @@ static void restore_exclusive_pte(struct vm_area_struct *vma,
  * sleeping.
  */
 static int try_restore_exclusive_pte(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, pte_t orig_pte)
+		unsigned long addr, hw_pte_t *ptep, pte_t orig_pte)
 {
 	const softleaf_t entry = softleaf_from_pte(orig_pte);
 	struct page *page = softleaf_to_page(entry);
@@ -1006,7 +1006,7 @@ static int try_restore_exclusive_pte(struct vm_area_struct *vma,
 
 static unsigned long
 copy_nonpresent_pte(struct mm_struct *dst_mm, struct mm_struct *src_mm,
-		pte_t *dst_pte, pte_t *src_pte, struct vm_area_struct *dst_vma,
+		hw_pte_t *dst_pte, hw_pte_t *src_pte, struct vm_area_struct *dst_vma,
 		struct vm_area_struct *src_vma, unsigned long addr, int *rss)
 {
 	pte_t orig_pte = ptep_get(src_pte);
@@ -1120,7 +1120,7 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct mm_struct *src_mm,
  */
 static inline int
 copy_present_page(struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
-		  pte_t *dst_pte, pte_t *src_pte, unsigned long addr, int *rss,
+		  hw_pte_t *dst_pte, hw_pte_t *src_pte, unsigned long addr, int *rss,
 		  struct folio **prealloc, struct page *page)
 {
 	struct folio *new_folio;
@@ -1159,7 +1159,7 @@ copy_present_page(struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma
 }
 
 static __always_inline void __copy_present_ptes(struct vm_area_struct *dst_vma,
-		struct vm_area_struct *src_vma, pte_t *dst_pte, pte_t *src_pte,
+		struct vm_area_struct *src_vma, hw_pte_t *dst_pte, hw_pte_t *src_pte,
 		pte_t pte, unsigned long addr, int nr)
 {
 	struct mm_struct *src_mm = src_vma->vm_mm;
@@ -1209,7 +1209,7 @@ static __always_inline void __copy_present_ptes(struct vm_area_struct *dst_vma,
  */
 static inline int
 copy_present_ptes(struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
-		 pte_t *dst_pte, pte_t *src_pte, pte_t pte, unsigned long addr,
+		 hw_pte_t *dst_pte, hw_pte_t *src_pte, pte_t pte, unsigned long addr,
 		 int max_nr, int *rss, struct folio **prealloc)
 {
 	fpb_t flags = FPB_MERGE_WRITE;
@@ -1309,8 +1309,8 @@ copy_pte_range(struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
 {
 	struct mm_struct *dst_mm = dst_vma->vm_mm;
 	struct mm_struct *src_mm = src_vma->vm_mm;
-	pte_t *orig_src_pte, *orig_dst_pte;
-	pte_t *src_pte, *dst_pte;
+	hw_pte_t *orig_src_pte, *orig_dst_pte;
+	hw_pte_t *src_pte, *dst_pte;
 	pmd_t dummy_pmdval;
 	pte_t ptent;
 	spinlock_t *src_ptl, *dst_ptl;
@@ -1703,7 +1703,7 @@ static inline bool zap_drop_markers(struct zap_details *details)
  * Returns true if uffd-wp PTEs were installed, false otherwise.
  */
 bool cond_install_uffd_wp_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, pte_t pte,
+		unsigned long addr, hw_pte_t *ptep, pte_t pte,
 		unsigned long nr_ptes)
 {
 	bool arm_uffd_pte = false;
@@ -1757,7 +1757,7 @@ bool cond_install_uffd_wp_ptes(struct vm_area_struct *vma,
  */
 static inline bool
 zap_install_uffd_wp_if_needed(struct vm_area_struct *vma,
-			      unsigned long addr, pte_t *pte, int nr,
+			      unsigned long addr, hw_pte_t *pte, int nr,
 			      struct zap_details *details, pte_t pteval)
 {
 	if (zap_drop_markers(details))
@@ -1768,7 +1768,7 @@ zap_install_uffd_wp_if_needed(struct vm_area_struct *vma,
 
 static __always_inline void zap_present_folio_ptes(struct mmu_gather *tlb,
 		struct vm_area_struct *vma, struct folio *folio,
-		struct page *page, pte_t *pte, pte_t ptent, unsigned int nr,
+		struct page *page, hw_pte_t *pte, pte_t ptent, unsigned int nr,
 		unsigned long addr, struct zap_details *details, int *rss,
 		bool *force_flush, bool *force_break, bool *any_skipped)
 {
@@ -1818,7 +1818,7 @@ static __always_inline void zap_present_folio_ptes(struct mmu_gather *tlb,
  * Returns the number of processed (skipped or zapped) PTEs (at least 1).
  */
 static inline int zap_present_ptes(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, pte_t *pte, pte_t ptent,
+		struct vm_area_struct *vma, hw_pte_t *pte, pte_t ptent,
 		unsigned int max_nr, unsigned long addr,
 		struct zap_details *details, int *rss, bool *force_flush,
 		bool *force_break, bool *any_skipped)
@@ -1864,7 +1864,7 @@ static inline int zap_present_ptes(struct mmu_gather *tlb,
 }
 
 static inline int zap_nonpresent_ptes(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, pte_t *pte, pte_t ptent,
+		struct vm_area_struct *vma, hw_pte_t *pte, pte_t ptent,
 		unsigned int max_nr, unsigned long addr,
 		struct zap_details *details, int *rss, bool *any_skipped)
 {
@@ -1935,7 +1935,7 @@ static inline int zap_nonpresent_ptes(struct mmu_gather *tlb,
 }
 
 static inline int do_zap_pte_range(struct mmu_gather *tlb,
-				   struct vm_area_struct *vma, pte_t *pte,
+				   struct vm_area_struct *vma, hw_pte_t *pte,
 				   unsigned long addr, unsigned long end,
 				   struct zap_details *details, int *rss,
 				   bool *force_flush, bool *force_break,
@@ -1998,7 +1998,7 @@ static bool zap_pte_table_if_empty(struct mm_struct *mm, pmd_t *pmd,
 		unsigned long addr, pmd_t *pmdval)
 {
 	spinlock_t *pml, *ptl = NULL;
-	pte_t *start_pte, *pte;
+	hw_pte_t *start_pte, *pte;
 	int i;
 
 	pml = pmd_lock(mm, pmd);
@@ -2038,8 +2038,8 @@ static unsigned long zap_pte_range(struct mmu_gather *tlb,
 	struct mm_struct *mm = tlb->mm;
 	int rss[NR_MM_COUNTERS];
 	spinlock_t *ptl;
-	pte_t *start_pte;
-	pte_t *pte;
+	hw_pte_t *start_pte;
+	hw_pte_t *pte;
 	pmd_t pmdval;
 	unsigned long start = addr;
 	bool direct_reclaim = true;
@@ -2421,7 +2421,7 @@ static pmd_t *walk_to_pmd(struct mm_struct *mm, unsigned long addr)
 	return pmd;
 }
 
-pte_t *get_locked_pte(struct mm_struct *mm, unsigned long addr,
+hw_pte_t *get_locked_pte(struct mm_struct *mm, unsigned long addr,
 		      spinlock_t **ptl)
 {
 	pmd_t *pmd = walk_to_pmd(mm, addr);
@@ -2479,7 +2479,7 @@ static int validate_page_before_insert(struct vm_area_struct *vma,
 	return 0;
 }
 
-static int insert_page_into_pte_locked(struct vm_area_struct *vma, pte_t *pte,
+static int insert_page_into_pte_locked(struct vm_area_struct *vma, hw_pte_t *pte,
 				unsigned long addr, struct page *page,
 				pgprot_t prot, bool mkwrite)
 {
@@ -2524,7 +2524,7 @@ static int insert_page(struct vm_area_struct *vma, unsigned long addr,
 			struct page *page, pgprot_t prot, bool mkwrite)
 {
 	int retval;
-	pte_t *pte;
+	hw_pte_t *pte;
 	spinlock_t *ptl;
 
 	retval = validate_page_before_insert(vma, page);
@@ -2541,7 +2541,7 @@ static int insert_page(struct vm_area_struct *vma, unsigned long addr,
 	return retval;
 }
 
-static int insert_page_in_batch_locked(struct vm_area_struct *vma, pte_t *pte,
+static int insert_page_in_batch_locked(struct vm_area_struct *vma, hw_pte_t *pte,
 			unsigned long addr, struct page *page, pgprot_t prot)
 {
 	int err;
@@ -2559,7 +2559,7 @@ static int insert_pages(struct vm_area_struct *vma, unsigned long addr,
 			struct page **pages, unsigned long *num, pgprot_t prot)
 {
 	pmd_t *pmd = NULL;
-	pte_t *start_pte, *pte;
+	hw_pte_t *start_pte, *pte;
 	spinlock_t *pte_lock;
 	struct mm_struct *const mm = vma->vm_mm;
 	unsigned long curr_page_idx = 0;
@@ -2802,7 +2802,8 @@ static vm_fault_t insert_pfn(struct vm_area_struct *vma, unsigned long addr,
 			unsigned long pfn, pgprot_t prot, bool mkwrite)
 {
 	struct mm_struct *mm = vma->vm_mm;
-	pte_t *pte, entry;
+	hw_pte_t *pte;
+	pte_t entry;
 	spinlock_t *ptl;
 
 	pte = get_locked_pte(mm, addr, &ptl);
@@ -3043,7 +3044,7 @@ static int remap_pte_range(struct mm_struct *mm, pmd_t *pmd,
 			unsigned long addr, unsigned long end,
 			unsigned long pfn, pgprot_t prot)
 {
-	pte_t *pte, *mapped_pte;
+	hw_pte_t *pte, *mapped_pte;
 	spinlock_t *ptl;
 	int err = 0;
 
@@ -3444,7 +3445,7 @@ static int apply_to_pte_range(struct mm_struct *mm, pmd_t *pmd,
 				     pte_fn_t fn, void *data, bool create,
 				     pgtbl_mod_mask *mask)
 {
-	pte_t *pte, *mapped_pte;
+	hw_pte_t *pte, *mapped_pte;
 	int err = 0;
 	spinlock_t *ptl;
 
@@ -4747,7 +4748,7 @@ static vm_fault_t handle_pte_marker(struct vm_fault *vmf)
 /*
  * Check if the PTEs within a range are contiguous swap entries.
  */
-static bool can_swapin_thp(struct vm_fault *vmf, pte_t *ptep, int nr_pages)
+static bool can_swapin_thp(struct vm_fault *vmf, hw_pte_t *ptep, int nr_pages)
 {
 	unsigned long addr;
 	int idx;
@@ -4799,7 +4800,7 @@ static unsigned long thp_swapin_suitable_orders(struct vm_fault *vmf)
 	unsigned long addr;
 	softleaf_t entry;
 	spinlock_t *ptl;
-	pte_t *pte;
+	hw_pte_t *pte;
 	int order;
 
 	/*
@@ -4893,7 +4894,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
 	int nr_pages;
 	unsigned long page_idx;
 	unsigned long address;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	if (!pte_unmap_same(vmf))
 		goto out;
@@ -5056,7 +5057,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
 		unsigned long idx = folio_page_idx(folio, page);
 		unsigned long folio_start = address - idx * PAGE_SIZE;
 		unsigned long folio_end = folio_start + nr * PAGE_SIZE;
-		pte_t *folio_ptep;
+		hw_pte_t *folio_ptep;
 		pte_t folio_pte;
 
 		if (unlikely(folio_start < max(address & PMD_MASK, vma->vm_start)))
@@ -5286,7 +5287,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
 	return ret;
 }
 
-static bool pte_range_none(pte_t *pte, int nr_pages)
+static bool pte_range_none(hw_pte_t *pte, int nr_pages)
 {
 	int i;
 
@@ -5305,7 +5306,7 @@ static struct folio *alloc_anon_folio(struct vm_fault *vmf)
 	unsigned long orders;
 	struct folio *folio;
 	unsigned long addr;
-	pte_t *pte;
+	hw_pte_t *pte;
 	gfp_t gfp;
 	int order;
 
@@ -5387,7 +5388,7 @@ static struct folio *alloc_anon_folio(struct vm_fault *vmf)
 	return folio_prealloc(vma->vm_mm, vma, vmf->address, true);
 }
 
-void map_anon_folio_pte_nopf(struct folio *folio, pte_t *pte,
+void map_anon_folio_pte_nopf(struct folio *folio, hw_pte_t *pte,
 		struct vm_area_struct *vma, unsigned long addr,
 		bool uffd_wp)
 {
@@ -5408,7 +5409,7 @@ void map_anon_folio_pte_nopf(struct folio *folio, pte_t *pte,
 	update_mmu_cache_range(NULL, vma, addr, pte, nr_pages);
 }
 
-static void map_anon_folio_pte_pf(struct folio *folio, pte_t *pte,
+static void map_anon_folio_pte_pf(struct folio *folio, hw_pte_t *pte,
 		struct vm_area_struct *vma, unsigned long addr, bool uffd_wp)
 {
 	const unsigned int order = folio_order(folio);
@@ -6193,7 +6194,7 @@ int numa_migrate_check(struct folio *folio, struct vm_fault *vmf,
 }
 
 static void numa_rebuild_single_mapping(struct vm_fault *vmf, struct vm_area_struct *vma,
-					unsigned long fault_addr, pte_t *fault_pte,
+					unsigned long fault_addr, hw_pte_t *fault_pte,
 					bool writable)
 {
 	pte_t pte, old_pte;
@@ -6215,7 +6216,7 @@ static void numa_rebuild_large_mapping(struct vm_fault *vmf, struct vm_area_stru
 	unsigned long start, end, addr = vmf->address;
 	unsigned long addr_start = addr - (nr << PAGE_SHIFT);
 	unsigned long pt_start = ALIGN_DOWN(addr, PMD_SIZE);
-	pte_t *start_ptep;
+	hw_pte_t *start_ptep;
 
 	/* Stay within the VMA and within the page table. */
 	start = max3(addr_start, pt_start, vma->vm_start);
@@ -6977,7 +6978,7 @@ int __pmd_alloc(struct mm_struct *mm, pud_t *pud, unsigned long address)
 #endif /* __PAGETABLE_PMD_FOLDED */
 
 static inline void pfnmap_args_setup(struct follow_pfnmap_args *args,
-				     spinlock_t *lock, pte_t *ptep,
+				     spinlock_t *lock, hw_pte_t *ptep,
 				     pgprot_t pgprot, unsigned long pfn_base,
 				     unsigned long addr_mask, bool writable,
 				     bool special)
@@ -7046,7 +7047,8 @@ int follow_pfnmap_start(struct follow_pfnmap_args *args)
 	p4d_t *p4dp, p4d;
 	pud_t *pudp, pud;
 	pmd_t *pmdp, pmd;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	pfnmap_lockdep_assert(vma);
 
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 79053ece02cd..21f3b3b3507c 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -691,7 +691,7 @@ static int queue_folios_pte_range(pmd_t *pmd, unsigned long addr,
 	struct folio *folio;
 	struct queue_pages *qp = walk->private;
 	unsigned long flags = qp->flags;
-	pte_t *pte, *mapped_pte;
+	hw_pte_t *pte, *mapped_pte;
 	pte_t ptent;
 	spinlock_t *ptl;
 	int max_nr, nr;
@@ -771,7 +771,7 @@ static int queue_folios_pte_range(pmd_t *pmd, unsigned long addr,
 	return 0;
 }
 
-static int queue_folios_hugetlb(pte_t *pte, unsigned long hmask,
+static int queue_folios_hugetlb(hw_pte_t *pte, unsigned long hmask,
 			       unsigned long addr, unsigned long end,
 			       struct mm_walk *walk)
 {
diff --git a/mm/migrate.c b/mm/migrate.c
index 15b45832bcfa..1786ab5e941c 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -494,7 +494,7 @@ void migration_entry_wait(struct mm_struct *mm, pmd_t *pmd,
 			  unsigned long address)
 {
 	spinlock_t *ptl;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t pte;
 	softleaf_t entry;
 
@@ -525,7 +525,7 @@ void migration_entry_wait(struct mm_struct *mm, pmd_t *pmd,
  *
  * This function will release the vma lock before returning.
  */
-void migration_entry_wait_huge(struct vm_area_struct *vma, unsigned long addr, pte_t *ptep)
+void migration_entry_wait_huge(struct vm_area_struct *vma, unsigned long addr, hw_pte_t *ptep)
 {
 	spinlock_t *ptl = huge_pte_lockptr(hstate_vma(vma), vma->vm_mm, ptep);
 	softleaf_t entry;
diff --git a/mm/migrate_device.c b/mm/migrate_device.c
index 009bfa8b212d..6138833fa902 100644
--- a/mm/migrate_device.c
+++ b/mm/migrate_device.c
@@ -254,7 +254,7 @@ static int migrate_vma_collect_pmd(pmd_t *pmdp,
 	spinlock_t *ptl;
 	struct folio *fault_folio = migrate->fault_page ?
 		page_folio(migrate->fault_page) : NULL;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 again:
 	if (pmd_trans_huge(*pmdp) || !pmd_present(*pmdp)) {
@@ -989,7 +989,7 @@ static void migrate_vma_insert_page(struct migrate_vma *migrate,
 	p4d_t *p4dp;
 	pud_t *pudp;
 	pmd_t *pmdp;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	pte_t orig_pte;
 
 	/* Only allow populating anonymous memory */
diff --git a/mm/mincore.c b/mm/mincore.c
index ff4ac8281768..15454ced345a 100644
--- a/mm/mincore.c
+++ b/mm/mincore.c
@@ -24,7 +24,7 @@
 #include "swap.h"
 #include "internal.h"
 
-static int mincore_hugetlb(pte_t *pte, unsigned long hmask, unsigned long addr,
+static int mincore_hugetlb(hw_pte_t *pte, unsigned long hmask, unsigned long addr,
 			unsigned long end, struct mm_walk *walk)
 {
 #ifdef CONFIG_HUGETLB_PAGE
@@ -164,7 +164,7 @@ static int mincore_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 {
 	spinlock_t *ptl;
 	struct vm_area_struct *vma = walk->vma;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 	unsigned char *vec = walk->private;
 	int nr = (end - addr) >> PAGE_SHIFT;
 	int step, i;
diff --git a/mm/mlock.c b/mm/mlock.c
index 39215a3eab1f..cc6c24e46348 100644
--- a/mm/mlock.c
+++ b/mm/mlock.c
@@ -305,7 +305,7 @@ void munlock_folio(struct folio *folio)
 }
 
 static inline unsigned int folio_mlock_step(struct folio *folio,
-		pte_t *pte, unsigned long addr, unsigned long end)
+		hw_pte_t *pte, unsigned long addr, unsigned long end)
 {
 	unsigned int count = (end - addr) >> PAGE_SHIFT;
 	pte_t ptent = ptep_get(pte);
@@ -353,7 +353,7 @@ static int mlock_pte_range(pmd_t *pmd, unsigned long addr,
 {
 	struct vm_area_struct *vma = walk->vma;
 	spinlock_t *ptl;
-	pte_t *start_pte, *pte;
+	hw_pte_t *start_pte, *pte;
 	pte_t ptent;
 	struct folio *folio;
 	unsigned int step = 1;
diff --git a/mm/mprotect.c b/mm/mprotect.c
index 2888ee638d87..1d23475e0bb7 100644
--- a/mm/mprotect.c
+++ b/mm/mprotect.c
@@ -103,7 +103,7 @@ bool can_change_pte_writable(struct vm_area_struct *vma, unsigned long addr,
 	return can_change_shared_pte_writable(vma, pte);
 }
 
-static int mprotect_folio_pte_batch(struct folio *folio, pte_t *ptep,
+static int mprotect_folio_pte_batch(struct folio *folio, hw_pte_t *ptep,
 				    pte_t pte, int max_nr_ptes, fpb_t flags)
 {
 	/* No underlying folio, so cannot batch */
@@ -118,7 +118,7 @@ static int mprotect_folio_pte_batch(struct folio *folio, pte_t *ptep,
 
 /* Set nr_ptes number of ptes, starting from idx */
 static __always_inline void prot_commit_flush_ptes(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep, pte_t oldpte, pte_t ptent,
+		unsigned long addr, hw_pte_t *ptep, pte_t oldpte, pte_t ptent,
 		int nr_ptes, int idx, bool set_write, struct mmu_gather *tlb)
 {
 	/*
@@ -170,7 +170,7 @@ static __always_inline int page_anon_exclusive_batch(int start_idx, int max_len,
  * retrieve sub-batches.
  */
 static __always_inline void commit_anon_folio_batch(struct vm_area_struct *vma,
-		struct folio *folio, struct page *first_page, unsigned long addr, pte_t *ptep,
+		struct folio *folio, struct page *first_page, unsigned long addr, hw_pte_t *ptep,
 		pte_t oldpte, pte_t ptent, int nr_ptes, struct mmu_gather *tlb)
 {
 	bool expected_anon_exclusive;
@@ -189,7 +189,7 @@ static __always_inline void commit_anon_folio_batch(struct vm_area_struct *vma,
 }
 
 static __always_inline void set_write_prot_commit_flush_ptes(struct vm_area_struct *vma,
-		struct folio *folio, struct page *page, unsigned long addr, pte_t *ptep,
+		struct folio *folio, struct page *page, unsigned long addr, hw_pte_t *ptep,
 		pte_t oldpte, pte_t ptent, int nr_ptes, struct mmu_gather *tlb)
 {
 	bool set_write;
@@ -212,7 +212,7 @@ static __always_inline void set_write_prot_commit_flush_ptes(struct vm_area_stru
 }
 
 static long change_softleaf_pte(struct vm_area_struct *vma,
-	unsigned long addr, pte_t *pte, pte_t oldpte, unsigned long cp_flags)
+	unsigned long addr, hw_pte_t *pte, pte_t oldpte, unsigned long cp_flags)
 {
 	const bool uffd_prot = cp_flags & (MM_CP_UFFD_WP | MM_CP_UFFD_RWP);
 	const bool uffd_prot_resolve = cp_flags &
@@ -279,7 +279,7 @@ static long change_softleaf_pte(struct vm_area_struct *vma,
 }
 
 static __always_inline void change_present_ptes(struct mmu_gather *tlb,
-		struct vm_area_struct *vma, unsigned long addr, pte_t *ptep,
+		struct vm_area_struct *vma, unsigned long addr, hw_pte_t *ptep,
 		int nr_ptes, unsigned long end, pgprot_t newprot,
 		struct folio *folio, struct page *page, unsigned long cp_flags)
 {
@@ -332,7 +332,8 @@ static long change_pte_range(struct mmu_gather *tlb,
 		struct vm_area_struct *vma, pmd_t *pmd, unsigned long addr,
 		unsigned long end, pgprot_t newprot, unsigned long cp_flags)
 {
-	pte_t *pte, oldpte;
+	hw_pte_t *pte;
+	pte_t oldpte;
 	spinlock_t *ptl;
 	long pages = 0;
 	bool is_private_single_threaded;
@@ -727,7 +728,7 @@ long change_protection(struct mmu_gather *tlb,
 	return pages;
 }
 
-static int prot_none_pte_entry(pte_t *pte, unsigned long addr,
+static int prot_none_pte_entry(hw_pte_t *pte, unsigned long addr,
 			       unsigned long next, struct mm_walk *walk)
 {
 	return pfn_modify_allowed(pte_pfn(ptep_get(pte)),
@@ -736,7 +737,7 @@ static int prot_none_pte_entry(pte_t *pte, unsigned long addr,
 }
 
 #ifdef CONFIG_HUGETLB_PAGE
-static int prot_none_hugetlb_entry(pte_t *pte, unsigned long hmask,
+static int prot_none_hugetlb_entry(hw_pte_t *pte, unsigned long hmask,
 				   unsigned long addr, unsigned long next,
 				   struct mm_walk *walk)
 {
diff --git a/mm/mremap.c b/mm/mremap.c
index 7c368440fafe..5272cea12f80 100644
--- a/mm/mremap.c
+++ b/mm/mremap.c
@@ -176,7 +176,7 @@ static pte_t move_soft_dirty_pte(pte_t pte)
 }
 
 static int mremap_folio_pte_batch(struct vm_area_struct *vma, unsigned long addr,
-		pte_t *ptep, pte_t pte, int max_nr)
+		hw_pte_t *ptep, pte_t pte, int max_nr)
 {
 	struct folio *folio;
 
@@ -200,7 +200,7 @@ static int move_ptes(struct pagetable_move_control *pmc,
 	struct vm_area_struct *vma = pmc->old;
 	bool need_clear_uffd_wp = vma_has_uffd_without_event_remap(vma);
 	struct mm_struct *mm = vma->vm_mm;
-	pte_t *old_ptep, *new_ptep;
+	hw_pte_t *old_ptep, *new_ptep;
 	pte_t old_pte, pte;
 	pmd_t dummy_pmdval;
 	spinlock_t *old_ptl, *new_ptl;
diff --git a/mm/page_table_check.c b/mm/page_table_check.c
index 6ffc536359cd..70984b4e3cde 100644
--- a/mm/page_table_check.c
+++ b/mm/page_table_check.c
@@ -208,7 +208,7 @@ static void page_table_check_pte_flags(pte_t pte)
 }
 
 void __page_table_check_ptes_set(struct mm_struct *mm, unsigned long addr,
-				 pte_t *ptep, pte_t pte, unsigned int nr)
+				 hw_pte_t *ptep, pte_t pte, unsigned int nr)
 {
 	unsigned int i;
 
@@ -279,7 +279,7 @@ void __page_table_check_pte_clear_range(struct mm_struct *mm,
 		return;
 
 	if (!pmd_bad(pmd) && !pmd_leaf(pmd)) {
-		pte_t *ptep = pte_offset_map(&pmd, addr);
+		hw_pte_t *ptep = pte_offset_map(&pmd, addr);
 		unsigned long i;
 
 		if (WARN_ON(!ptep))
diff --git a/mm/pagewalk.c b/mm/pagewalk.c
index cc07fcf50e87..71157ec0b305 100644
--- a/mm/pagewalk.c
+++ b/mm/pagewalk.c
@@ -26,7 +26,7 @@ static int real_depth(int depth)
 	return depth;
 }
 
-static int walk_pte_range_inner(pte_t *pte, unsigned long addr,
+static int walk_pte_range_inner(hw_pte_t *pte, unsigned long addr,
 				unsigned long end, struct mm_walk *walk)
 {
 	const struct mm_walk_ops *ops = walk->ops;
@@ -61,7 +61,7 @@ static int walk_pte_range_inner(pte_t *pte, unsigned long addr,
 static int walk_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 			  struct mm_walk *walk)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	int err = 0;
 	spinlock_t *ptl;
 
@@ -341,7 +341,7 @@ static int walk_hugetlb_range(unsigned long addr, unsigned long end,
 	unsigned long next;
 	unsigned long hmask = huge_page_mask(h);
 	unsigned long sz = huge_page_size(h);
-	pte_t *pte;
+	hw_pte_t *pte;
 	const struct mm_walk_ops *ops = walk->ops;
 	int err = 0;
 
@@ -907,7 +907,8 @@ struct folio *folio_walk_start(struct folio_walk *fw,
 	struct page *page;
 	pud_t *pudp, pud;
 	pmd_t *pmdp, pmd;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 	spinlock_t *ptl;
 	pgd_t *pgdp;
 	p4d_t *p4dp;
diff --git a/mm/percpu.c b/mm/percpu.c
index a802d72c116f..fa79a1295c5b 100644
--- a/mm/percpu.c
+++ b/mm/percpu.c
@@ -3176,7 +3176,7 @@ void __init __weak pcpu_populate_pte(unsigned long addr)
 
 	pmd = pmd_offset(pud, addr);
 	if (!pmd_present(*pmd)) {
-		pte_t *new;
+		hw_pte_t *new;
 
 		new = memblock_alloc_or_panic(PTE_TABLE_SIZE, PTE_TABLE_SIZE);
 		pmd_populate_kernel(&init_mm, pmd, new);
diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c
index b91b1a98029c..52f286da3eec 100644
--- a/mm/pgtable-generic.c
+++ b/mm/pgtable-generic.c
@@ -68,7 +68,7 @@ void pmd_clear_bad(pmd_t *pmd)
  * force that call on sun4c so we changed this macro slightly
  */
 int ptep_set_access_flags(struct vm_area_struct *vma,
-			  unsigned long address, pte_t *ptep,
+			  unsigned long address, hw_pte_t *ptep,
 			  pte_t entry, int dirty)
 {
 	int changed = !pte_same(ptep_get(ptep), entry);
@@ -82,7 +82,7 @@ int ptep_set_access_flags(struct vm_area_struct *vma,
 
 #ifndef __HAVE_ARCH_PTEP_CLEAR_YOUNG_FLUSH
 bool ptep_clear_flush_young(struct vm_area_struct *vma,
-		unsigned long address, pte_t *ptep)
+		unsigned long address, hw_pte_t *ptep)
 {
 	bool young;
 
@@ -95,7 +95,7 @@ bool ptep_clear_flush_young(struct vm_area_struct *vma,
 
 #ifndef __HAVE_ARCH_PTEP_CLEAR_FLUSH
 pte_t ptep_clear_flush(struct vm_area_struct *vma, unsigned long address,
-		       pte_t *ptep)
+		       hw_pte_t *ptep)
 {
 	struct mm_struct *mm = (vma)->vm_mm;
 	pte_t pte;
@@ -282,7 +282,7 @@ static unsigned long pmdp_get_lockless_start(void) { return 0; }
 static void pmdp_get_lockless_end(unsigned long irqflags) { }
 #endif
 
-pte_t *__pte_offset_map(pmd_t *pmd, unsigned long addr, pmd_t *pmdvalp)
+hw_pte_t *__pte_offset_map(pmd_t *pmd, unsigned long addr, pmd_t *pmdvalp)
 {
 	unsigned long irqflags;
 	pmd_t pmdval;
@@ -308,11 +308,11 @@ pte_t *__pte_offset_map(pmd_t *pmd, unsigned long addr, pmd_t *pmdvalp)
 	return NULL;
 }
 
-pte_t *pte_offset_map_ro_nolock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_ro_nolock(struct mm_struct *mm, pmd_t *pmd,
 				unsigned long addr, spinlock_t **ptlp)
 {
 	pmd_t pmdval;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	pte = __pte_offset_map(pmd, addr, &pmdval);
 	if (likely(pte))
@@ -320,11 +320,11 @@ pte_t *pte_offset_map_ro_nolock(struct mm_struct *mm, pmd_t *pmd,
 	return pte;
 }
 
-pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
 				unsigned long addr, pmd_t *pmdvalp,
 				spinlock_t **ptlp)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	VM_WARN_ON_ONCE(!pmdvalp);
 	pte = __pte_offset_map(pmd, addr, pmdvalp);
@@ -390,12 +390,12 @@ pte_t *pte_offset_map_rw_nolock(struct mm_struct *mm, pmd_t *pmd,
  * table, and may not use RCU at all: "outsiders" like khugepaged should avoid
  * pte_offset_map() and co once the vma is detached from mm or mm_users is zero.
  */
-pte_t *pte_offset_map_lock(struct mm_struct *mm, pmd_t *pmd,
+hw_pte_t *pte_offset_map_lock(struct mm_struct *mm, pmd_t *pmd,
 			   unsigned long addr, spinlock_t **ptlp)
 {
 	spinlock_t *ptl;
 	pmd_t pmdval;
-	pte_t *pte;
+	hw_pte_t *pte;
 again:
 	pte = __pte_offset_map(pmd, addr, &pmdval);
 	if (unlikely(!pte))
diff --git a/mm/ptdump.c b/mm/ptdump.c
index 5851096e6f65..376880071ca2 100644
--- a/mm/ptdump.c
+++ b/mm/ptdump.c
@@ -117,7 +117,7 @@ static int ptdump_pmd_entry(pmd_t *pmd, unsigned long addr,
 	return 0;
 }
 
-static int ptdump_pte_entry(pte_t *pte, unsigned long addr,
+static int ptdump_pte_entry(hw_pte_t *pte, unsigned long addr,
 			    unsigned long next, struct mm_walk *walk)
 {
 	struct ptdump_state *st = walk->private;
diff --git a/mm/rmap.c b/mm/rmap.c
index f3b21aaa34ee..aafcd6b52893 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -1128,7 +1128,7 @@ static int page_vma_mkclean_one(struct page_vma_mapped_walk *pvmw)
 
 		address = pvmw->address;
 		if (pvmw->pte) {
-			pte_t *pte = pvmw->pte;
+			hw_pte_t *pte = pvmw->pte;
 			pte_t entry = ptep_get(pte);
 
 			/*
@@ -2145,7 +2145,7 @@ static pte_t swp_pte_prepare(swp_entry_t entry, pte_t old_pte,
 
 static bool ttu_anon_swapbacked_folio(struct vm_area_struct *vma,
 		struct folio *folio, struct page *page, unsigned long address,
-		pte_t *ptep, pte_t pteval)
+		hw_pte_t *ptep, pte_t pteval)
 {
 	const bool anon_exclusive = folio_test_anon(folio) &&
 				    PageAnonExclusive(page);
@@ -2180,7 +2180,7 @@ static bool ttu_anon_swapbacked_folio(struct vm_area_struct *vma,
 }
 
 static bool ttu_anon_folio(struct vm_area_struct *vma, struct folio *folio,
-		struct page *page, unsigned long address, pte_t *ptep,
+		struct page *page, unsigned long address, hw_pte_t *ptep,
 		pte_t pteval, unsigned long nr_pages)
 {
 	/*
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 5a2469fb1838..1d7a56884905 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -137,7 +137,7 @@ static void * __meminit altmap_alloc_block_buf(unsigned long size,
 	return __va(__pfn_to_phys(pfn));
 }
 
-void __meminit vmemmap_verify(pte_t *pte, int node,
+void __meminit vmemmap_verify(hw_pte_t *pte, int node,
 				unsigned long start, unsigned long end)
 {
 	unsigned long pfn = pte_pfn(ptep_get(pte));
@@ -148,11 +148,11 @@ void __meminit vmemmap_verify(pte_t *pte, int node,
 			start, end - 1);
 }
 
-static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
+static hw_pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
 				       struct vmem_altmap *altmap,
 				       unsigned long ptpfn, unsigned long flags)
 {
-	pte_t *pte = pte_offset_kernel(pmd, addr);
+	hw_pte_t *pte = pte_offset_kernel(pmd, addr);
 	if (pte_none(ptep_get(pte))) {
 		pte_t entry;
 		void *p;
@@ -243,7 +243,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
 	return pgd;
 }
 
-static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
+static hw_pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 					      struct vmem_altmap *altmap,
 					      unsigned long ptpfn,
 					      unsigned long flags)
@@ -252,7 +252,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	pgd = vmemmap_pgd_populate(addr, node);
 	if (!pgd)
@@ -281,7 +281,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
 					    unsigned long flags)
 {
 	unsigned long addr = start;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	for (; addr < end; addr += PAGE_SIZE) {
 		pte = vmemmap_populate_address(addr, node, altmap,
@@ -314,7 +314,7 @@ void vmemmap_wrprotect_hvo(unsigned long addr, unsigned long end,
 				    int node, unsigned long headsize)
 {
 	unsigned long maddr;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	for (maddr = addr + headsize; maddr < end; maddr += PAGE_SIZE) {
 		pte = virt_to_kpte(maddr);
@@ -396,7 +396,7 @@ int __weak __meminit vmemmap_check_pmd(pmd_t *pmd, int node,
 {
 	if (!pmd_leaf(pmdp_get(pmd)))
 		return 0;
-	vmemmap_verify((pte_t *)pmd, node, addr, next);
+	vmemmap_verify((hw_pte_t *)pmd, node, addr, next);
 
 	return 1;
 }
@@ -474,9 +474,9 @@ static bool __meminit reuse_compound_section(unsigned long start_pfn,
 	return !IS_ALIGNED(offset, nr_pages) && nr_pages > PAGES_PER_SUBSECTION;
 }
 
-static pte_t * __meminit compound_section_tail_page(unsigned long addr)
+static hw_pte_t * __meminit compound_section_tail_page(unsigned long addr)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	addr -= PAGE_SIZE;
 
@@ -497,7 +497,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
 						     struct dev_pagemap *pgmap)
 {
 	unsigned long size, addr;
-	pte_t *pte;
+	hw_pte_t *pte;
 	int rc;
 
 	if (reuse_compound_section(start_pfn, pgmap)) {
diff --git a/mm/swap_state.c b/mm/swap_state.c
index f3961fdd857d..865eb5422c2e 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -918,7 +918,8 @@ static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
 	struct swap_io_ctx ctx = {};
 	struct blk_plug plug;
 	struct folio *folio;
-	pte_t *pte = NULL, pentry;
+	hw_pte_t *pte = NULL;
+	pte_t pentry;
 	int win;
 	unsigned long start, end, addr;
 	pgoff_t ilx = targ_ilx;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 601979b97f95..6adeee9c95ec 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2422,7 +2422,8 @@ static int unuse_pte(struct vm_area_struct *vma, pmd_t *pmd,
 	struct page *page;
 	struct folio *swapcache;
 	spinlock_t *ptl;
-	pte_t *pte, new_pte, old_pte;
+	hw_pte_t *pte;
+	pte_t new_pte, old_pte;
 	bool hwpoisoned = false;
 	int ret = 1;
 
@@ -2533,7 +2534,7 @@ static int unuse_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
 			unsigned long addr, unsigned long end,
 			unsigned int type)
 {
-	pte_t *pte = NULL;
+	hw_pte_t *pte = NULL;
 
 	do {
 		struct folio *folio;
diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index 74f04c323c50..8fa3c960fb1a 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -361,13 +361,13 @@ static int mfill_atomic_install_pte(pmd_t *dst_pmd,
 {
 	int ret;
 	struct mm_struct *dst_mm = dst_vma->vm_mm;
-	pte_t _dst_pte, *dst_pte;
+	hw_pte_t *dst_pte;
 	bool writable = dst_vma->vm_flags & VM_WRITE;
 	bool vm_shared = dst_vma->vm_flags & VM_SHARED;
 	spinlock_t *ptl;
 	struct folio *folio = page_folio(page);
 	bool page_in_cache = folio_mapping(folio);
-	pte_t dst_ptep;
+	pte_t _dst_pte, dst_ptep;
 
 	_dst_pte = mk_pte(page, dst_vma->vm_page_prot);
 	_dst_pte = pte_mkdirty(_dst_pte);
@@ -656,7 +656,8 @@ static int mfill_atomic_pte_zeropage(struct mfill_state *state)
 	struct vm_area_struct *dst_vma = state->vma;
 	unsigned long dst_addr = state->dst_addr;
 	pmd_t *dst_pmd = state->pmd;
-	pte_t _dst_pte, *dst_pte;
+	hw_pte_t *dst_pte;
+	pte_t _dst_pte;
 	spinlock_t *ptl;
 	int ret;
 
@@ -737,7 +738,8 @@ static int mfill_atomic_pte_poison(struct mfill_state *state)
 	struct mm_struct *dst_mm = dst_vma->vm_mm;
 	unsigned long dst_addr = state->dst_addr;
 	pmd_t *dst_pmd = state->pmd;
-	pte_t _dst_pte, *dst_pte;
+	hw_pte_t *dst_pte;
+	pte_t _dst_pte;
 	spinlock_t *ptl;
 	int ret;
 
@@ -784,7 +786,7 @@ static __always_inline ssize_t mfill_atomic_hugetlb(
 {
 	struct mm_struct *dst_mm = dst_vma->vm_mm;
 	ssize_t err;
-	pte_t *dst_pte;
+	hw_pte_t *dst_pte;
 	unsigned long src_addr, dst_addr;
 	long copied;
 	struct folio *folio;
@@ -1261,7 +1263,7 @@ void double_pt_unlock(spinlock_t *ptl1,
 		__release(ptl2);
 }
 
-static inline bool is_pte_pages_stable(pte_t *dst_pte, pte_t *src_pte,
+static inline bool is_pte_pages_stable(hw_pte_t *dst_pte, hw_pte_t *src_pte,
 				       pte_t orig_dst_pte, pte_t orig_src_pte,
 				       pmd_t *dst_pmd, pmd_t dst_pmdval)
 {
@@ -1279,7 +1281,8 @@ static inline bool is_pte_pages_stable(pte_t *dst_pte, pte_t *src_pte,
  */
 static struct folio *check_ptes_for_batched_move(struct vm_area_struct *src_vma,
 						 unsigned long src_addr,
-						 pte_t *src_pte, pte_t *dst_pte)
+						 hw_pte_t *src_pte,
+						 hw_pte_t *dst_pte)
 {
 	pte_t orig_dst_pte, orig_src_pte;
 	struct folio *folio;
@@ -1310,7 +1313,7 @@ static long move_present_ptes(struct mm_struct *mm,
 			      struct vm_area_struct *dst_vma,
 			      struct vm_area_struct *src_vma,
 			      unsigned long dst_addr, unsigned long src_addr,
-			      pte_t *dst_pte, pte_t *src_pte,
+			      hw_pte_t *dst_pte, hw_pte_t *src_pte,
 			      pte_t orig_dst_pte, pte_t orig_src_pte,
 			      pmd_t *dst_pmd, pmd_t dst_pmdval,
 			      spinlock_t *dst_ptl, spinlock_t *src_ptl,
@@ -1397,7 +1400,7 @@ static long move_present_ptes(struct mm_struct *mm,
 
 static int move_swap_pte(struct mm_struct *mm, struct vm_area_struct *dst_vma,
 			 unsigned long dst_addr, unsigned long src_addr,
-			 pte_t *dst_pte, pte_t *src_pte,
+			 hw_pte_t *dst_pte, hw_pte_t *src_pte,
 			 pte_t orig_dst_pte, pte_t orig_src_pte,
 			 pmd_t *dst_pmd, pmd_t dst_pmdval,
 			 spinlock_t *dst_ptl, spinlock_t *src_ptl,
@@ -1462,7 +1465,7 @@ static int move_zeropage_pte(struct mm_struct *mm,
 			     struct vm_area_struct *dst_vma,
 			     struct vm_area_struct *src_vma,
 			     unsigned long dst_addr, unsigned long src_addr,
-			     pte_t *dst_pte, pte_t *src_pte,
+			     hw_pte_t *dst_pte, hw_pte_t *src_pte,
 			     pte_t orig_dst_pte, pte_t orig_src_pte,
 			     pmd_t *dst_pmd, pmd_t dst_pmdval,
 			     spinlock_t *dst_ptl, spinlock_t *src_ptl)
@@ -1508,8 +1511,8 @@ static long move_pages_ptes(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd
 	pte_t orig_src_pte, orig_dst_pte;
 	pte_t src_folio_pte;
 	spinlock_t *src_ptl, *dst_ptl;
-	pte_t *src_pte = NULL;
-	pte_t *dst_pte = NULL;
+	hw_pte_t *src_pte = NULL;
+	hw_pte_t *dst_pte = NULL;
 	pmd_t dummy_pmdval;
 	pmd_t dst_pmdval;
 	struct folio *src_folio = NULL;
@@ -2650,7 +2653,8 @@ static inline bool userfaultfd_huge_must_wait(struct userfaultfd_ctx *ctx,
 					      unsigned long reason)
 {
 	struct vm_area_struct *vma = vmf->vma;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	assert_fault_locked(vmf);
 
@@ -2723,7 +2727,7 @@ static inline bool userfaultfd_must_wait(struct userfaultfd_ctx *ctx,
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd, _pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 	pte_t ptent;
 	bool ret;
 
diff --git a/mm/util.c b/mm/util.c
index bf0513d1d3d0..cc253050f689 100644
--- a/mm/util.c
+++ b/mm/util.c
@@ -1564,7 +1564,7 @@ EXPORT_SYMBOL(mmap_action_complete);
  *
  * Return: the number of table entries in the batch.
  */
-unsigned int folio_pte_batch(struct folio *folio, pte_t *ptep, pte_t pte,
+unsigned int folio_pte_batch(struct folio *folio, hw_pte_t *ptep, pte_t pte,
 		unsigned int max_nr)
 {
 	return folio_pte_batch_flags(folio, NULL, ptep, &pte, max_nr, 0);
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index bea9f76ed7e7..4f97f99c340e 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -97,7 +97,7 @@ static int vmap_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 			phys_addr_t phys_addr, pgprot_t prot,
 			unsigned int max_page_shift, pgtbl_mod_mask *mask)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	u64 pfn;
 	struct page *page;
 	unsigned long size = PAGE_SIZE;
@@ -389,7 +389,7 @@ int ioremap_page_range(unsigned long addr, unsigned long end,
 static void vunmap_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 			     pgtbl_mod_mask *mask)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	pte_t ptent;
 	unsigned long size = PAGE_SIZE;
 
@@ -550,7 +550,7 @@ static int vmap_pages_pte_range(pmd_t *pmd, unsigned long addr,
 		pgtbl_mod_mask *mask)
 {
 	int err = 0;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	/*
 	 * nr is a running index into the array which helps higher level
@@ -827,7 +827,8 @@ struct page *vmalloc_to_page(const void *vmalloc_addr)
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	/*
 	 * XXX we might need to change this if we add VIRTUAL_BUG_ON for
@@ -3605,7 +3606,7 @@ struct vmap_pfn_data {
 	unsigned int	idx;
 };
 
-static int vmap_pfn_apply(pte_t *pte, unsigned long addr, void *private)
+static int vmap_pfn_apply(hw_pte_t *pte, unsigned long addr, void *private)
 {
 	struct vmap_pfn_data *data = private;
 	unsigned long pfn = data->pfns[data->idx];
diff --git a/mm/vmscan.c b/mm/vmscan.c
index f11491ee9ed5..acc77c673a95 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3540,7 +3540,7 @@ static bool walk_pte_range(pmd_t *pmd, unsigned long start, unsigned long end,
 {
 	int i;
 	bool dirty;
-	pte_t *pte;
+	hw_pte_t *pte;
 	spinlock_t *ptl;
 	unsigned long addr;
 	int total = 0;
@@ -3573,7 +3573,7 @@ static bool walk_pte_range(pmd_t *pmd, unsigned long start, unsigned long end,
 	for (i = pte_index(start), addr = start; addr != end; i += nr, addr += nr * PAGE_SIZE) {
 		unsigned long pfn;
 		struct folio *folio;
-		pte_t *cur_pte = pte + i;
+		hw_pte_t *cur_pte = pte + i;
 		pte_t ptent = ptep_get(cur_pte);
 
 		nr = 1;
@@ -4262,7 +4262,7 @@ bool lru_gen_look_around(struct page_vma_mapped_walk *pvmw, unsigned int nr)
 	struct lru_gen_mm_walk *walk;
 	struct folio *last = NULL;
 	int young = nr;
-	pte_t *pte = pvmw->pte;
+	hw_pte_t *pte = pvmw->pte;
 	unsigned long addr = pvmw->address;
 	struct vm_area_struct *vma = pvmw->vma;
 	struct folio *folio = pfn_folio(pvmw->pfn);
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 04/15] mm: convert PTE table entries in ptep_get()
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (2 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 05/15] mm: convert PTE table entry to pte Alexander Gordeev
                   ` (10 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

ptep_get() now accepts a pointer to hw_pte_t storage but must continue to
return a software PTE value. Add __pte_from_hw for both generic hw_pte_t
definitions. Read the hw_pte_t table element atomically before converting
it to pte_t.

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 include/linux/pgtable.h       | 2 +-
 include/linux/pgtable_types.h | 2 ++
 2 files changed, 3 insertions(+), 1 deletion(-)

diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h
index dad80d264aac..1768421755a9 100644
--- a/include/linux/pgtable.h
+++ b/include/linux/pgtable.h
@@ -493,7 +493,7 @@ static inline int pudp_set_access_flags(struct vm_area_struct *vma,
 #ifndef ptep_get
 static inline pte_t ptep_get(hw_pte_t *ptep)
 {
-	return READ_ONCE(*ptep);
+	return __pte_from_hw(READ_ONCE(*ptep));
 }
 #endif
 
diff --git a/include/linux/pgtable_types.h b/include/linux/pgtable_types.h
index 07da05d375c2..d6c5a7548550 100644
--- a/include/linux/pgtable_types.h
+++ b/include/linux/pgtable_types.h
@@ -8,8 +8,10 @@
 
 #ifdef CONFIG_ARCH_HAS_HW_PTE_T
 typedef struct __hw_pte_t { pte_t __pte; } hw_pte_t;
+#define __pte_from_hw(pte)	((pte).__pte)
 #else
 #define hw_pte_t pte_t
+#define __pte_from_hw(pte)	(pte)
 #endif
 
 #endif /* !__ASSEMBLY__ */
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 05/15] mm: convert PTE table entry to pte
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (3 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 04/15] mm: convert PTE table entries in ptep_get() Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage Alexander Gordeev
                   ` (9 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

The non-MMU stub receives hw_pte_t but returns a software PTE value. It
has no attached PTE that requires ptep_get(), so convert only the stored
entry through __pte_from_hw().

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 include/linux/hugetlb.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index bc0b9c65aa1d..be31fcf93654 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -1283,7 +1283,8 @@ static inline pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
 #ifdef CONFIG_MMU
 	return ptep_get(ptep);
 #else
-	return *ptep;
+	/* No attached PTE that requires ptep_get(). */
+	return __pte_from_hw(*ptep);
 #endif
 }
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (4 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 05/15] mm: convert PTE table entry to pte Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table Alexander Gordeev
                   ` (8 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

Atomic PTE updates need an lvalue for the bits stored in an HW PTE.
pte_val() only accepts a SW PTE value, so it cannot operate directly on
a distinct hw_pte_t.

Add hw_pte_val() to expose the underlying pte_val() lvalue. Access the
wrapper's __pte member when hw_pte_t is distinct, and use pte_val()
directly when it remains an alias of pte_t.

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 include/linux/pgtable_types.h | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/include/linux/pgtable_types.h b/include/linux/pgtable_types.h
index d6c5a7548550..ee4eace5c3e1 100644
--- a/include/linux/pgtable_types.h
+++ b/include/linux/pgtable_types.h
@@ -9,9 +9,13 @@
 #ifdef CONFIG_ARCH_HAS_HW_PTE_T
 typedef struct __hw_pte_t { pte_t __pte; } hw_pte_t;
 #define __pte_from_hw(pte)	((pte).__pte)
+
+#define hw_pte_val(x)  pte_val((x).__pte)
 #else
 #define hw_pte_t pte_t
 #define __pte_from_hw(pte)	(pte)
+
+#define hw_pte_val(x)  pte_val(x)
 #endif
 
 #endif /* !__ASSEMBLY__ */
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (5 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks Alexander Gordeev
                   ` (7 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

kasan_early_shadow_pte is a complete PTE table rather than a software
PTE value. Declare and define its elements as hw_pte_t so the object uses
the PTE table storage type.

Read the first element through ptep_get(), which returns the software PTE
value expected by note_page_pte(), instead of accessing hw_pte_t storage
directly. This also uses the accessor selected by the architecture.

No architecture selects ARCH_HAS_HW_PTE_T at this point, so the storage
representation remains unchanged.

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 include/linux/kasan.h | 2 +-
 mm/kasan/init.c       | 2 +-
 mm/ptdump.c           | 2 +-
 3 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/include/linux/kasan.h b/include/linux/kasan.h
index bf233bde68c7..ff41949c0ac0 100644
--- a/include/linux/kasan.h
+++ b/include/linux/kasan.h
@@ -51,7 +51,7 @@ typedef unsigned int __bitwise kasan_vmalloc_flags_t;
 #endif
 
 extern unsigned char kasan_early_shadow_page[PAGE_SIZE];
-extern pte_t kasan_early_shadow_pte[MAX_PTRS_PER_PTE + PTE_HWTABLE_PTRS];
+extern hw_pte_t kasan_early_shadow_pte[MAX_PTRS_PER_PTE + PTE_HWTABLE_PTRS];
 extern pmd_t kasan_early_shadow_pmd[MAX_PTRS_PER_PMD];
 extern pud_t kasan_early_shadow_pud[MAX_PTRS_PER_PUD];
 extern p4d_t kasan_early_shadow_p4d[MAX_PTRS_PER_P4D];
diff --git a/mm/kasan/init.c b/mm/kasan/init.c
index 30c5266eaf37..9bbc2a41d23f 100644
--- a/mm/kasan/init.c
+++ b/mm/kasan/init.c
@@ -64,7 +64,7 @@ static inline bool kasan_pmd_table(pud_t pud)
 	return false;
 }
 #endif
-pte_t kasan_early_shadow_pte[MAX_PTRS_PER_PTE + PTE_HWTABLE_PTRS]
+hw_pte_t kasan_early_shadow_pte[MAX_PTRS_PER_PTE + PTE_HWTABLE_PTRS]
 	__bss_pgtbl;
 
 static inline bool kasan_pte_table(pmd_t pmd)
diff --git a/mm/ptdump.c b/mm/ptdump.c
index 376880071ca2..8f19f20be3c4 100644
--- a/mm/ptdump.c
+++ b/mm/ptdump.c
@@ -19,7 +19,7 @@ static inline int note_kasan_page_table(struct mm_walk *walk,
 {
 	struct ptdump_state *st = walk->private;
 
-	st->note_page_pte(st, addr, kasan_early_shadow_pte[0]);
+	st->note_page_pte(st, addr, ptep_get(kasan_early_shadow_pte));
 
 	walk->action = ACTION_CONTINUE;
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (6 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 09/15] xen: " Alexander Gordeev
                   ` (6 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

apply_to_page_range() now passes PTE table storage to its callback as
hw_pte_t *. Update the i915 remap and selftest callbacks to match the new
type.

Continue to use ptep_get() for software PTE values and set_pte_at() for
updates.
The type remains an alias of pte_t on x86 until that architecture opts in.

Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 drivers/gpu/drm/i915/gem/selftests/i915_gem_mman.c | 4 ++--
 drivers/gpu/drm/i915/i915_mm.c                     | 4 ++--
 2 files changed, 4 insertions(+), 4 deletions(-)

diff --git a/drivers/gpu/drm/i915/gem/selftests/i915_gem_mman.c b/drivers/gpu/drm/i915/gem/selftests/i915_gem_mman.c
index d01acfb7d93d..056faf4a3618 100644
--- a/drivers/gpu/drm/i915/gem/selftests/i915_gem_mman.c
+++ b/drivers/gpu/drm/i915/gem/selftests/i915_gem_mman.c
@@ -1690,7 +1690,7 @@ static int igt_mmap_gpu(void *arg)
 	return 0;
 }
 
-static int check_present_pte(pte_t *pte, unsigned long addr, void *data)
+static int check_present_pte(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	pte_t ptent = ptep_get(pte);
 
@@ -1703,7 +1703,7 @@ static int check_present_pte(pte_t *pte, unsigned long addr, void *data)
 	return 0;
 }
 
-static int check_absent_pte(pte_t *pte, unsigned long addr, void *data)
+static int check_absent_pte(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	pte_t ptent = ptep_get(pte);
 
diff --git a/drivers/gpu/drm/i915/i915_mm.c b/drivers/gpu/drm/i915/i915_mm.c
index fd89e7c7d8d6..aab88e8edf94 100644
--- a/drivers/gpu/drm/i915/i915_mm.c
+++ b/drivers/gpu/drm/i915/i915_mm.c
@@ -48,7 +48,7 @@ static inline unsigned long sgt_pfn(const struct remap_pfn *r)
 		return r->sgt.pfn + (r->sgt.curr >> PAGE_SHIFT);
 }
 
-static int remap_sg(pte_t *pte, unsigned long addr, void *data)
+static int remap_sg(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct remap_pfn *r = data;
 
@@ -70,7 +70,7 @@ static int remap_sg(pte_t *pte, unsigned long addr, void *data)
 #define EXPECTED_FLAGS (VM_PFNMAP | VM_DONTEXPAND | VM_DONTDUMP)
 
 #if IS_ENABLED(CONFIG_X86)
-static int remap_pfn(pte_t *pte, unsigned long addr, void *data)
+static int remap_pfn(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct remap_pfn *r = data;
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 09/15] xen: use hw_pte_t for PTE range callbacks
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (7 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines Alexander Gordeev
                   ` (5 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

From: Muhammad Usama Anjum <usama.anjum@arm.com>

Generic PTE range and remapping helpers now pass pointers to PTE table
storage as hw_pte_t *. Update the Xen callbacks to match those interfaces.

Keep software PTE values as pte_t and continue to access them through the
existing PTE helpers. This is required when Xen is built for an
architecture that selects the distinct hw_pte_t wrapper.

Reviewed-by: Juergen Gross <jgross@suse.com>
Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com>
---
 drivers/xen/gntdev.c               | 2 +-
 drivers/xen/privcmd.c              | 2 +-
 drivers/xen/xenbus/xenbus_client.c | 2 +-
 drivers/xen/xlate_mmu.c            | 4 ++--
 4 files changed, 5 insertions(+), 5 deletions(-)

diff --git a/drivers/xen/gntdev.c b/drivers/xen/gntdev.c
index 1dcc4675580e..b013bcad99b5 100644
--- a/drivers/xen/gntdev.c
+++ b/drivers/xen/gntdev.c
@@ -301,7 +301,7 @@ void gntdev_put_map(struct gntdev_priv *priv, struct gntdev_grant_map *map)
 
 /* ------------------------------------------------------------------ */
 
-static int find_grant_ptes(pte_t *pte, unsigned long addr, void *data)
+static int find_grant_ptes(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct gntdev_grant_map *map = data;
 	unsigned int pgnr = (addr - map->pages_vm_start) >> PAGE_SHIFT;
diff --git a/drivers/xen/privcmd.c b/drivers/xen/privcmd.c
index 7cfc28f1bb86..e724cc053e26 100644
--- a/drivers/xen/privcmd.c
+++ b/drivers/xen/privcmd.c
@@ -1658,7 +1658,7 @@ static int privcmd_mmap(struct file *file, struct vm_area_struct *vma)
  * on a per pfn/pte basis. Mapping calls that fail with ENOENT
  * can be then retried until success.
  */
-static int is_mapped_fn(pte_t *pte, unsigned long addr, void *data)
+static int is_mapped_fn(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	return pte_none(ptep_get(pte)) ? 0 : -EBUSY;
 }
diff --git a/drivers/xen/xenbus/xenbus_client.c b/drivers/xen/xenbus/xenbus_client.c
index 27682cb5e58a..387f97a55559 100644
--- a/drivers/xen/xenbus/xenbus_client.c
+++ b/drivers/xen/xenbus/xenbus_client.c
@@ -749,7 +749,7 @@ int xenbus_unmap_ring_vfree(struct xenbus_device *dev, void *vaddr)
 EXPORT_SYMBOL_GPL(xenbus_unmap_ring_vfree);
 
 #ifdef CONFIG_XEN_PV
-static int map_ring_apply(pte_t *pte, unsigned long addr, void *data)
+static int map_ring_apply(hw_pte_t *pte, unsigned long addr, void *data)
 {
 	struct map_ring_valloc *info = data;
 
diff --git a/drivers/xen/xlate_mmu.c b/drivers/xen/xlate_mmu.c
index 8efd7b55223f..d7244f9f9a54 100644
--- a/drivers/xen/xlate_mmu.c
+++ b/drivers/xen/xlate_mmu.c
@@ -93,7 +93,7 @@ static void setup_hparams(unsigned long gfn, void *data)
 	info->fgfn++;
 }
 
-static int remap_pte_fn(pte_t *ptep, unsigned long addr, void *data)
+static int remap_pte_fn(hw_pte_t *ptep, unsigned long addr, void *data)
 {
 	struct remap_data *info = data;
 	struct page *page = info->pages[info->index++];
@@ -269,7 +269,7 @@ struct remap_pfn {
 	unsigned long i;
 };
 
-static int remap_pfn_fn(pte_t *ptep, unsigned long addr, void *data)
+static int remap_pfn_fn(hw_pte_t *ptep, unsigned long addr, void *data)
 {
 	struct remap_pfn *r = data;
 	struct page *page = r->pages[r->i];
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (8 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 09/15] xen: " Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 11/15] s390: Distinguish hardware and software PTEs Alexander Gordeev
                   ` (4 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Avoid unnecessary pgtable entry read in ptep_flush_lazy() and
pmdp_flush_lazy() routines. Use the already obtained value of
a pgtable entry instead of dereferencing it again.

It might appear that the optimized out second dereferencing happens
under mm_context_t::flush_count lock and therefore a race condition
is possible. However, both routines write invalid entry bits to the
pgtable entries and therefore no reader is allowed to interpret the
rest of the value afterwards.

Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 arch/s390/mm/pgtable.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/arch/s390/mm/pgtable.c b/arch/s390/mm/pgtable.c
index 4acd8b140c4b..da076928b63b 100644
--- a/arch/s390/mm/pgtable.c
+++ b/arch/s390/mm/pgtable.c
@@ -107,7 +107,7 @@ static inline pte_t ptep_flush_lazy(struct mm_struct *mm,
 	atomic_inc(&mm->context.flush_count);
 	if (cpumask_equal(&mm->context.cpu_attach_mask,
 			  cpumask_of(smp_processor_id()))) {
-		set_pte(ptep, set_pte_bit(*ptep, __pgprot(_PAGE_INVALID)));
+		set_pte(ptep, set_pte_bit(old, __pgprot(_PAGE_INVALID)));
 		mm->context.flush_mm = 1;
 	} else
 		ptep_ipte_global(mm, addr, ptep, nodat);
@@ -227,7 +227,7 @@ static inline pmd_t pmdp_flush_lazy(struct mm_struct *mm,
 	atomic_inc(&mm->context.flush_count);
 	if (cpumask_equal(&mm->context.cpu_attach_mask,
 			  cpumask_of(smp_processor_id()))) {
-		set_pmd(pmdp, set_pmd_bit(*pmdp, __pgprot(_SEGMENT_ENTRY_INVALID)));
+		set_pmd(pmdp, set_pmd_bit(old, __pgprot(_SEGMENT_ENTRY_INVALID)));
 		mm->context.flush_mm = 1;
 	} else {
 		pmdp_idte_global(mm, addr, pmdp);
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 11/15] s390: Distinguish hardware and software PTEs
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (9 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 12/15] mm: Make lazy MMU mode context-aware Alexander Gordeev
                   ` (3 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

This is s390 version of ARM64 commit XXXXXXXXXXXX ("arm64:
distinguish HW PTE pointers from SW PTE value pointers").

It distinguishes hardware and software PTEs, introduces hw_pte_t
type for hardware PTEs, while software PTEs still use pte_t type
and enables ARCH_HAS_HW_PTE_T configuration for s390.

Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 arch/s390/Kconfig                    |  1 +
 arch/s390/boot/startup.c             |  2 +-
 arch/s390/boot/vmem.c                | 17 ++++-----
 arch/s390/include/asm/gmap_helpers.h |  2 +-
 arch/s390/include/asm/hugetlb.h      | 18 +++++-----
 arch/s390/include/asm/maccess.h      |  2 +-
 arch/s390/include/asm/page.h         |  3 +-
 arch/s390/include/asm/pgalloc.h      |  6 ++--
 arch/s390/include/asm/pgtable.h      | 54 +++++++++++++++-------------
 arch/s390/kernel/uv.c                |  4 +--
 arch/s390/kvm/s390/pv.c              |  2 +-
 arch/s390/mm/gmap_helpers.c          | 18 +++++-----
 arch/s390/mm/hugetlbpage.c           | 20 +++++------
 arch/s390/mm/maccess.c               |  2 +-
 arch/s390/mm/pageattr.c              | 10 +++---
 arch/s390/mm/pgtable.c               | 26 +++++++-------
 arch/s390/mm/vmem.c                  | 26 +++++++-------
 17 files changed, 112 insertions(+), 101 deletions(-)

diff --git a/arch/s390/Kconfig b/arch/s390/Kconfig
index 4b51bc6e8948..b9bc0e5e7d7d 100644
--- a/arch/s390/Kconfig
+++ b/arch/s390/Kconfig
@@ -100,6 +100,7 @@ config S390
 	select ARCH_HAS_FORTIFY_SOURCE
 	select ARCH_HAS_GCOV_PROFILE_ALL
 	select ARCH_HAS_GIGANTIC_PAGE
+	select ARCH_HAS_HW_PTE_T
 	select ARCH_HAS_KCOV
 	select ARCH_HAS_MEMBARRIER_SYNC_CORE
 	select ARCH_HAS_MEM_ENCRYPT
diff --git a/arch/s390/boot/startup.c b/arch/s390/boot/startup.c
index c1dcf8b1b579..d8732b416161 100644
--- a/arch/s390/boot/startup.c
+++ b/arch/s390/boot/startup.c
@@ -30,7 +30,7 @@
 struct vm_layout __bootdata_preserved(vm_layout);
 unsigned long __bootdata_preserved(__abs_lowcore);
 unsigned long __bootdata_preserved(__memcpy_real_area);
-pte_t *__bootdata_preserved(memcpy_real_ptep);
+hw_pte_t *__bootdata_preserved(memcpy_real_ptep);
 unsigned long __bootdata_preserved(VMALLOC_START);
 unsigned long __bootdata_preserved(VMALLOC_END);
 struct page *__bootdata_preserved(vmemmap);
diff --git a/arch/s390/boot/vmem.c b/arch/s390/boot/vmem.c
index ff6d58a476ba..cc619cc85d37 100644
--- a/arch/s390/boot/vmem.c
+++ b/arch/s390/boot/vmem.c
@@ -75,7 +75,7 @@ static void pgtable_populate(unsigned long addr, unsigned long end, enum populat
 #ifdef CONFIG_KASAN
 
 #define kasan_early_shadow_page	vmlinux.kasan_early_shadow_page_off
-#define kasan_early_shadow_pte	((pte_t *)vmlinux.kasan_early_shadow_pte_off)
+#define kasan_early_shadow_pte	((hw_pte_t *)vmlinux.kasan_early_shadow_pte_off)
 #define kasan_early_shadow_pmd	((pmd_t *)vmlinux.kasan_early_shadow_pmd_off)
 #define kasan_early_shadow_pud	((pud_t *)vmlinux.kasan_early_shadow_pud_off)
 #define kasan_early_shadow_p4d	((p4d_t *)vmlinux.kasan_early_shadow_p4d_off)
@@ -178,7 +178,7 @@ static bool kasan_pmd_populate_zero_shadow(pmd_t *pmd, unsigned long addr,
 	return false;
 }
 
-static bool kasan_pte_populate_zero_shadow(pte_t *pte, enum populate_mode mode)
+static bool kasan_pte_populate_zero_shadow(hw_pte_t *pte, enum populate_mode mode)
 {
 	if (mode == POPULATE_KASAN_ZERO_SHADOW) {
 		set_pte(pte, pte_z);
@@ -216,7 +216,7 @@ static inline bool kasan_pmd_populate_zero_shadow(pmd_t *pmd, unsigned long addr
 	return false;
 }
 
-static bool kasan_pte_populate_zero_shadow(pte_t *pte, enum populate_mode mode)
+static bool kasan_pte_populate_zero_shadow(hw_pte_t *pte, enum populate_mode mode)
 {
 	return false;
 }
@@ -226,7 +226,7 @@ static bool kasan_pte_populate_zero_shadow(pte_t *pte, enum populate_mode mode)
 /*
  * Mimic virt_to_kpte() in lack of init_mm symbol. Skip pmd NULL check though.
  */
-static inline pte_t *__virt_to_kpte(unsigned long va)
+static inline hw_pte_t *__virt_to_kpte(unsigned long va)
 {
 	return pte_offset_kernel(pmd_offset(pud_offset(p4d_offset(pgd_offset_k(va), va), va), va), va);
 }
@@ -242,9 +242,9 @@ static void *boot_crst_alloc(unsigned long val)
 	return table;
 }
 
-static pte_t *boot_pte_alloc(void)
+static hw_pte_t *boot_pte_alloc(void)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	pte = (void *)physmem_alloc_or_die(RR_VMEM, PAGE_SIZE, PAGE_SIZE);
 	__arch_set_page_dat(pte, 1);
@@ -334,7 +334,8 @@ static void pgtable_pte_populate(pmd_t *pmd, unsigned long addr, unsigned long e
 				 enum populate_mode mode)
 {
 	unsigned long pages = 0;
-	pte_t *pte, entry;
+	hw_pte_t *pte;
+	pte_t entry;
 
 	pte = pte_offset_kernel(pmd, addr);
 	for (; addr < end; addr += PAGE_SIZE, pte++) {
@@ -356,7 +357,7 @@ static void pgtable_pmd_populate(pud_t *pud, unsigned long addr, unsigned long e
 {
 	unsigned long pa, next, pages = 0;
 	pmd_t *pmd, entry, large_entry;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	pmd = pmd_offset(pud, addr);
 	for (; addr < end; addr = next, pmd++) {
diff --git a/arch/s390/include/asm/gmap_helpers.h b/arch/s390/include/asm/gmap_helpers.h
index d2b616604a46..d23907ed78f3 100644
--- a/arch/s390/include/asm/gmap_helpers.h
+++ b/arch/s390/include/asm/gmap_helpers.h
@@ -12,6 +12,6 @@ void gmap_helper_zap_one_page(struct mm_struct *mm, unsigned long vmaddr);
 void gmap_helper_discard(struct mm_struct *mm, unsigned long vmaddr, unsigned long end);
 int gmap_helper_disable_cow_sharing(void);
 void gmap_helper_try_set_pte_unused(struct mm_struct *mm, unsigned long vmaddr);
-pte_t *try_get_locked_pte(struct mm_struct *mm, unsigned long addr, spinlock_t **ptl);
+hw_pte_t *try_get_locked_pte(struct mm_struct *mm, unsigned long addr, spinlock_t **ptl);
 
 #endif /* _ASM_S390_GMAP_HELPERS_H */
diff --git a/arch/s390/include/asm/hugetlb.h b/arch/s390/include/asm/hugetlb.h
index aea754b67c89..7c1cf08725c5 100644
--- a/arch/s390/include/asm/hugetlb.h
+++ b/arch/s390/include/asm/hugetlb.h
@@ -19,19 +19,19 @@
 
 #define __HAVE_ARCH_HUGE_SET_HUGE_PTE_AT
 void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-		     pte_t *ptep, pte_t pte, unsigned long sz);
+		     hw_pte_t *ptep, pte_t pte, unsigned long sz);
 void __set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-		       pte_t *ptep, pte_t pte);
+		       hw_pte_t *ptep, pte_t pte);
 
 #define __HAVE_ARCH_HUGE_PTEP_GET
-pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep);
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep);
 
 pte_t __huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
-				pte_t *ptep);
+				hw_pte_t *ptep);
 
 #define __HAVE_ARCH_HUGE_PTEP_GET_AND_CLEAR
 static inline pte_t huge_ptep_get_and_clear(struct mm_struct *mm,
-					    unsigned long addr, pte_t *ptep,
+					    unsigned long addr, hw_pte_t *ptep,
 					    unsigned long sz)
 {
 	return __huge_ptep_get_and_clear(mm, addr, ptep);
@@ -39,7 +39,7 @@ static inline pte_t huge_ptep_get_and_clear(struct mm_struct *mm,
 
 #define __HAVE_ARCH_HUGE_PTE_CLEAR
 static inline void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
-				  pte_t *ptep, unsigned long sz)
+				  hw_pte_t *ptep, unsigned long sz)
 {
 	if ((pte_val(ptep_get(ptep)) & _REGION_ENTRY_TYPE_MASK) == _REGION_ENTRY_TYPE_R3)
 		set_pud((pud_t *)ptep, __pud(_REGION3_ENTRY_EMPTY));
@@ -49,14 +49,14 @@ static inline void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
 
 #define __HAVE_ARCH_HUGE_PTEP_CLEAR_FLUSH
 static inline pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
-					  unsigned long address, pte_t *ptep)
+					  unsigned long address, hw_pte_t *ptep)
 {
 	return __huge_ptep_get_and_clear(vma->vm_mm, address, ptep);
 }
 
 #define  __HAVE_ARCH_HUGE_PTEP_SET_ACCESS_FLAGS
 static inline int huge_ptep_set_access_flags(struct vm_area_struct *vma,
-					     unsigned long addr, pte_t *ptep,
+					     unsigned long addr, hw_pte_t *ptep,
 					     pte_t pte, int dirty)
 {
 	int changed = !pte_same(huge_ptep_get(vma->vm_mm, addr, ptep), pte);
@@ -70,7 +70,7 @@ static inline int huge_ptep_set_access_flags(struct vm_area_struct *vma,
 
 #define __HAVE_ARCH_HUGE_PTEP_SET_WRPROTECT
 static inline void huge_ptep_set_wrprotect(struct mm_struct *mm,
-					   unsigned long addr, pte_t *ptep)
+					   unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t pte = __huge_ptep_get_and_clear(mm, addr, ptep);
 
diff --git a/arch/s390/include/asm/maccess.h b/arch/s390/include/asm/maccess.h
index 50225940d971..dafd9fe290ce 100644
--- a/arch/s390/include/asm/maccess.h
+++ b/arch/s390/include/asm/maccess.h
@@ -10,7 +10,7 @@
 struct iov_iter;
 
 extern unsigned long __memcpy_real_area;
-extern pte_t *memcpy_real_ptep;
+extern hw_pte_t *memcpy_real_ptep;
 size_t memcpy_real_iter(struct iov_iter *iter, unsigned long src, size_t count);
 int memcpy_real(void *dest, unsigned long src, size_t count);
 #ifdef CONFIG_CRASH_DUMP
diff --git a/arch/s390/include/asm/page.h b/arch/s390/include/asm/page.h
index 56da819a79e6..42b4fecd2de3 100644
--- a/arch/s390/include/asm/page.h
+++ b/arch/s390/include/asm/page.h
@@ -113,7 +113,8 @@ DEFINE_PGVAL_FUNC(pud)
 DEFINE_PGVAL_FUNC(p4d)
 DEFINE_PGVAL_FUNC(pgd)
 
-typedef pte_t *pgtable_t;
+struct __hw_pte_t;
+typedef struct __hw_pte_t *pgtable_t;
 
 #define __pgprot(x)	((pgprot_t) { (x) } )
 #define __pte(x)        ((pte_t) { (x) } )
diff --git a/arch/s390/include/asm/pgalloc.h b/arch/s390/include/asm/pgalloc.h
index a5de9e61ea9e..83a732898819 100644
--- a/arch/s390/include/asm/pgalloc.h
+++ b/arch/s390/include/asm/pgalloc.h
@@ -159,8 +159,8 @@ static inline void pmd_populate(struct mm_struct *mm,
 /*
  * page table entry allocation/free routines.
  */
-#define pte_alloc_one_kernel(mm) ((pte_t *)page_table_alloc(mm))
-#define pte_alloc_one(mm) ((pte_t *)page_table_alloc(mm))
+#define pte_alloc_one_kernel(mm) ((hw_pte_t *)page_table_alloc(mm))
+#define pte_alloc_one(mm) ((hw_pte_t *)page_table_alloc(mm))
 
 #define pte_free_kernel(mm, pte) page_table_free(mm, (unsigned long *) pte)
 #define pte_free(mm, pte) page_table_free(mm, (unsigned long *) pte)
@@ -171,7 +171,7 @@ void pte_free_defer(struct mm_struct *mm, pgtable_t pgtable);
 
 void vmem_map_init(void);
 void *vmem_crst_alloc(unsigned long val);
-pte_t *vmem_pte_alloc(void);
+hw_pte_t *vmem_pte_alloc(void);
 
 unsigned long base_asce_alloc(unsigned long addr, unsigned long num_pages);
 void base_asce_free(unsigned long asce);
diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
index e882663a58e7..c47264f3abf2 100644
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -978,17 +978,19 @@ static inline void set_pmd(pmd_t *pmdp, pmd_t pmd)
 	WRITE_ONCE(*pmdp, pmd);
 }
 
-static inline void set_pte(pte_t *ptep, pte_t pte)
+static inline void set_pte(hw_pte_t *ptep, pte_t pte)
 {
+	hw_pte_t hwpte = (hw_pte_t) { (pte) };
+
 	if (pte_present(pte))
 		pte = clear_pte_bit(pte, __pgprot(_PAGE_UNUSED));
-	WRITE_ONCE(*ptep, pte);
+	WRITE_ONCE(*ptep, hwpte);
 }
 
 #define ptep_get ptep_get
-static inline pte_t ptep_get(pte_t *ptep)
+static inline pte_t ptep_get(hw_pte_t *ptep)
 {
-	return READ_ONCE(*ptep);
+	return __pte_from_hw(READ_ONCE(*ptep));
 }
 
 #define pmdp_get pmdp_get
@@ -1015,7 +1017,7 @@ static inline pgd_t pgdp_get(pgd_t *pgdp)
 	return READ_ONCE(*pgdp);
 }
 
-static inline void pte_clear(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+static inline void pte_clear(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep)
 {
 	set_pte(ptep, __pte(_PAGE_INVALID));
 }
@@ -1133,7 +1135,7 @@ static inline unsigned long sske_frame(unsigned long addr, unsigned char skey)
 #define IPTE_NODAT	0x400
 #define IPTE_GUEST_ASCE	0x800
 
-static __always_inline void __ptep_rdp(unsigned long addr, pte_t *ptep, int local)
+static __always_inline void __ptep_rdp(unsigned long addr, hw_pte_t *ptep, int local)
 {
 	unsigned long pto;
 
@@ -1144,7 +1146,7 @@ static __always_inline void __ptep_rdp(unsigned long addr, pte_t *ptep, int loca
 		       [m4] "i" (local));
 }
 
-static __always_inline void __ptep_ipte(unsigned long address, pte_t *ptep,
+static __always_inline void __ptep_ipte(unsigned long address, hw_pte_t *ptep,
 					unsigned long opt, unsigned long asce,
 					int local)
 {
@@ -1168,7 +1170,7 @@ static __always_inline void __ptep_ipte(unsigned long address, pte_t *ptep,
 }
 
 static __always_inline void __ptep_ipte_range(unsigned long address, int nr,
-					      pte_t *ptep, int local)
+					      hw_pte_t *ptep, int local)
 {
 	unsigned long pto = __pa(ptep);
 
@@ -1194,12 +1196,12 @@ static __always_inline void __ptep_ipte_range(unsigned long address, int nr,
  * have ptep_get_and_clear do the tlb flush. In exchange flush_tlb_range
  * is a nop.
  */
-pte_t ptep_xchg_direct(struct mm_struct *, unsigned long, pte_t *, pte_t);
-pte_t ptep_xchg_lazy(struct mm_struct *, unsigned long, pte_t *, pte_t);
+pte_t ptep_xchg_direct(struct mm_struct *, unsigned long, hw_pte_t *, pte_t);
+pte_t ptep_xchg_lazy(struct mm_struct *, unsigned long, hw_pte_t *, pte_t);
 
 #define __HAVE_ARCH_PTEP_TEST_AND_CLEAR_YOUNG
 static inline bool ptep_test_and_clear_young(struct vm_area_struct *vma,
-		unsigned long addr, pte_t *ptep)
+		unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t pte = ptep_get(ptep);
 
@@ -1209,14 +1211,14 @@ static inline bool ptep_test_and_clear_young(struct vm_area_struct *vma,
 
 #define __HAVE_ARCH_PTEP_CLEAR_YOUNG_FLUSH
 static inline bool ptep_clear_flush_young(struct vm_area_struct *vma,
-		unsigned long address, pte_t *ptep)
+		unsigned long address, hw_pte_t *ptep)
 {
 	return ptep_test_and_clear_young(vma, address, ptep);
 }
 
 #define __HAVE_ARCH_PTEP_GET_AND_CLEAR
 static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
-				       unsigned long addr, pte_t *ptep)
+				       unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t res;
 
@@ -1229,13 +1231,13 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
 }
 
 #define __HAVE_ARCH_PTEP_MODIFY_PROT_TRANSACTION
-pte_t ptep_modify_prot_start(struct vm_area_struct *, unsigned long, pte_t *);
+pte_t ptep_modify_prot_start(struct vm_area_struct *, unsigned long, hw_pte_t *);
 void ptep_modify_prot_commit(struct vm_area_struct *, unsigned long,
-			     pte_t *, pte_t, pte_t);
+			     hw_pte_t *, pte_t, pte_t);
 
 #define __HAVE_ARCH_PTEP_CLEAR_FLUSH
 static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
-				     unsigned long addr, pte_t *ptep)
+				     unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t res;
 
@@ -1257,7 +1259,7 @@ static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
 #define __HAVE_ARCH_PTEP_GET_AND_CLEAR_FULL
 static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
 					    unsigned long addr,
-					    pte_t *ptep, int full)
+					    hw_pte_t *ptep, int full)
 {
 	pte_t res;
 
@@ -1289,7 +1291,7 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
 
 #define __HAVE_ARCH_PTEP_SET_WRPROTECT
 static inline void ptep_set_wrprotect(struct mm_struct *mm,
-				      unsigned long addr, pte_t *ptep)
+				      unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t pte = ptep_get(ptep);
 
@@ -1315,7 +1317,7 @@ static inline int pte_allow_rdp(pte_t old, pte_t new)
 
 static inline void flush_tlb_fix_spurious_fault(struct vm_area_struct *vma,
 						unsigned long address,
-						pte_t *ptep)
+						hw_pte_t *ptep)
 {
 	/*
 	 * RDP might not have propagated the PTE protection reset to all CPUs,
@@ -1332,17 +1334,19 @@ static inline void flush_tlb_fix_spurious_fault(struct vm_area_struct *vma,
 }
 #define flush_tlb_fix_spurious_fault flush_tlb_fix_spurious_fault
 
-void ptep_reset_dat_prot(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+void ptep_reset_dat_prot(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep,
 			 pte_t new);
 
 #define __HAVE_ARCH_PTEP_SET_ACCESS_FLAGS
 static inline int ptep_set_access_flags(struct vm_area_struct *vma,
-					unsigned long addr, pte_t *ptep,
+					unsigned long addr, hw_pte_t *ptep,
 					pte_t entry, int dirty)
 {
-	if (pte_same(*ptep, entry))
+	pte_t pte = ptep_get(ptep);
+
+	if (pte_same(pte, entry))
 		return 0;
-	if (cpu_has_rdp() && pte_allow_rdp(*ptep, entry))
+	if (cpu_has_rdp() && pte_allow_rdp(pte, entry))
 		ptep_reset_dat_prot(vma->vm_mm, addr, ptep, entry);
 	else
 		ptep_xchg_direct(vma->vm_mm, addr, ptep, entry);
@@ -1359,7 +1363,7 @@ pgprot_t pgprot_writecombine(pgprot_t prot);
  * are within the same folio, PMD and VMA.
  */
 static inline void set_ptes(struct mm_struct *mm, unsigned long addr,
-			      pte_t *ptep, pte_t entry, unsigned int nr)
+		hw_pte_t *ptep, pte_t entry, unsigned int nr)
 {
 	page_table_check_ptes_set(mm, addr, ptep, entry, nr);
 	for (;;) {
@@ -1997,7 +2001,7 @@ extern void vmem_remove_mapping(unsigned long start, unsigned long size);
 extern int __vmem_map_4k_page(unsigned long addr, unsigned long phys, pgprot_t prot, bool alloc);
 extern int vmem_map_4k_page(unsigned long addr, unsigned long phys, pgprot_t prot);
 extern void vmem_unmap_4k_page(unsigned long addr);
-extern pte_t *vmem_get_alloc_pte(unsigned long addr, bool alloc);
+extern hw_pte_t *vmem_get_alloc_pte(unsigned long addr, bool alloc);
 
 /* s390 has a private copy of get unmapped area to deal with cache synonyms */
 #define HAVE_ARCH_UNMAPPED_AREA
diff --git a/arch/s390/kernel/uv.c b/arch/s390/kernel/uv.c
index 8ea9dd7704ff..f81c47db421e 100644
--- a/arch/s390/kernel/uv.c
+++ b/arch/s390/kernel/uv.c
@@ -210,7 +210,7 @@ int uv_convert_from_secure_pte(pte_t pte)
 	return uv_convert_from_secure_folio(pfn_folio(pte_pfn(pte)));
 }
 
-static int uv_free_range_cb(pte_t *ptep, unsigned long addr, void *data)
+static int uv_free_range_cb(hw_pte_t *ptep, unsigned long addr, void *data)
 {
 	pte_t pte = ptep_get(ptep);
 
@@ -242,7 +242,7 @@ void uv_free_stor_var(void *stor_var)
 }
 EXPORT_SYMBOL_FOR_MODULES(uv_free_stor_var, "kvm");
 
-static int uv_alloc_range_cb(pte_t *ptep, unsigned long addr, void *data)
+static int uv_alloc_range_cb(hw_pte_t *ptep, unsigned long addr, void *data)
 {
 	struct page *page;
 	pte_t pte;
diff --git a/arch/s390/kvm/s390/pv.c b/arch/s390/kvm/s390/pv.c
index b18abd0e29ef..bec85e434514 100644
--- a/arch/s390/kvm/s390/pv.c
+++ b/arch/s390/kvm/s390/pv.c
@@ -106,7 +106,7 @@ static void _kvm_s390_pv_make_secure(struct guest_fault *f)
 	struct pv_make_secure *priv = f->priv;
 	struct folio *folio;
 	spinlock_t *ptl;	/* pte lock from try_get_locked_pte() */
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	folio = pfn_folio(f->pfn);
 	priv->rc = -EAGAIN;
diff --git a/arch/s390/mm/gmap_helpers.c b/arch/s390/mm/gmap_helpers.c
index ff63ffb1dbd2..b39f939b308e 100644
--- a/arch/s390/mm/gmap_helpers.c
+++ b/arch/s390/mm/gmap_helpers.c
@@ -39,14 +39,14 @@
  * * the pointer to the pte corresponding to @addr in @mm, if it can be reached
  *   and locked.
  */
-pte_t *try_get_locked_pte(struct mm_struct *mm, unsigned long vmaddr, spinlock_t **ptl)
+hw_pte_t *try_get_locked_pte(struct mm_struct *mm, unsigned long vmaddr, spinlock_t **ptl)
 __context_unsafe(/* Returns nonnull if lock taken or not taken */)
 {
 	pmd_t *pmdp, pmd, pmdval;
 	pud_t *pudp, pud;
 	p4d_t *p4dp, p4d;
 	pgd_t *pgdp, pgd;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	pgdp = pgd_offset(mm, vmaddr);
 	pgd = pgdp_get(pgdp);
@@ -96,7 +96,7 @@ __context_unsafe(/* pte_unmap_unlock() not instrumented */)
 	struct vm_area_struct *vma;
 	spinlock_t *ptl;	/* Lock for the host (userspace) page table */
 	softleaf_t sl;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	mmap_assert_locked(mm);
 
@@ -109,8 +109,8 @@ __context_unsafe(/* pte_unmap_unlock() not instrumented */)
 	ptep = try_get_locked_pte(mm, vmaddr, &ptl);
 	if (IS_ERR_OR_NULL(ptep))
 		return;
-	sl = softleaf_from_pte(*ptep);
-	if (pte_swap(*ptep) && softleaf_is_swap(sl)) {
+	sl = softleaf_from_pte(ptep_get(ptep));
+	if (pte_swap(ptep_get(ptep)) && softleaf_is_swap(sl)) {
 		dec_mm_counter(mm, MM_SWAPENTS);
 		swap_put_entries_direct(sl, 1);
 		pte_clear(mm, vmaddr, ptep);
@@ -166,7 +166,7 @@ void gmap_helper_try_set_pte_unused(struct mm_struct *mm, unsigned long vmaddr)
 __context_unsafe(/* pte_unmap_unlock() not instrumented */)
 {
 	spinlock_t *ptl;	/* Lock for the host (userspace) page table */
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	/*
 	 * Several paths exists that takes the ptl lock and then call the
@@ -184,19 +184,19 @@ __context_unsafe(/* pte_unmap_unlock() not instrumented */)
 	if (IS_ERR_OR_NULL(ptep))
 		return;
 
-	if (pte_present(*ptep))
+	if (pte_present(ptep_get(ptep)))
 		__atomic64_or(_PAGE_UNUSED, (long *)ptep);
 	pte_unmap_unlock(ptep, ptl);
 }
 EXPORT_SYMBOL_GPL(gmap_helper_try_set_pte_unused);
 
-static int find_zeropage_pte_entry(pte_t *pte, unsigned long addr,
+static int find_zeropage_pte_entry(hw_pte_t *pte, unsigned long addr,
 				   unsigned long end, struct mm_walk *walk)
 {
 	unsigned long *found_addr = walk->private;
 
 	/* Return 1 of the page is a zeropage. */
-	if (is_zero_pfn(pte_pfn(*pte))) {
+	if (is_zero_pfn(pte_pfn(ptep_get(pte)))) {
 		/*
 		 * Shared zeropage in e.g., a FS DAX mapping? We cannot do the
 		 * right thing and likely don't care: FAULT_FLAG_UNSHARE
diff --git a/arch/s390/mm/hugetlbpage.c b/arch/s390/mm/hugetlbpage.c
index f84aa9265430..35b7ae166fe6 100644
--- a/arch/s390/mm/hugetlbpage.c
+++ b/arch/s390/mm/hugetlbpage.c
@@ -136,7 +136,7 @@ static inline pte_t __rste_to_pte(unsigned long rste)
 }
 
 void __set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-		     pte_t *ptep, pte_t pte)
+		     hw_pte_t *ptep, pte_t pte)
 {
 	unsigned long rste;
 
@@ -156,18 +156,18 @@ void __set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
 }
 
 void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
-		     pte_t *ptep, pte_t pte, unsigned long sz)
+		     hw_pte_t *ptep, pte_t pte, unsigned long sz)
 {
 	__set_huge_pte_at(mm, addr, ptep, pte);
 }
 
-pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep)
 {
 	return __rste_to_pte(pte_val(ptep_get(ptep)));
 }
 
 pte_t __huge_ptep_get_and_clear(struct mm_struct *mm,
-				unsigned long addr, pte_t *ptep)
+				unsigned long addr, hw_pte_t *ptep)
 {
 	pte_t pte = huge_ptep_get(mm, addr, ptep);
 	pmd_t *pmdp = (pmd_t *) ptep;
@@ -180,7 +180,7 @@ pte_t __huge_ptep_get_and_clear(struct mm_struct *mm,
 	return pte;
 }
 
-pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+hw_pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
 			unsigned long addr, unsigned long sz)
 {
 	pgd_t *pgdp;
@@ -194,15 +194,15 @@ pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
 		pudp = pud_alloc(mm, p4dp, addr);
 		if (pudp) {
 			if (sz == PUD_SIZE)
-				return (pte_t *) pudp;
+				return (hw_pte_t *) pudp;
 			else if (sz == PMD_SIZE)
 				pmdp = pmd_alloc(mm, pudp, addr);
 		}
 	}
-	return (pte_t *) pmdp;
+	return (hw_pte_t *) pmdp;
 }
 
-pte_t *huge_pte_offset(struct mm_struct *mm,
+hw_pte_t *huge_pte_offset(struct mm_struct *mm,
 		       unsigned long addr, unsigned long sz)
 {
 	pgd_t *pgdp;
@@ -216,12 +216,12 @@ pte_t *huge_pte_offset(struct mm_struct *mm,
 		if (p4d_present(p4dp_get(p4dp))) {
 			pudp = pud_offset(p4dp, addr);
 			if (sz == PUD_SIZE)
-				return (pte_t *)pudp;
+				return (hw_pte_t *)pudp;
 			if (pud_present(pudp_get(pudp)))
 				pmdp = pmd_offset(pudp, addr);
 		}
 	}
-	return (pte_t *) pmdp;
+	return (hw_pte_t *) pmdp;
 }
 
 bool __init arch_hugetlb_valid_size(unsigned long size)
diff --git a/arch/s390/mm/maccess.c b/arch/s390/mm/maccess.c
index f39968dd8063..9026d6a3e657 100644
--- a/arch/s390/mm/maccess.c
+++ b/arch/s390/mm/maccess.c
@@ -22,7 +22,7 @@
 #include <asm/ctlreg.h>
 
 unsigned long __bootdata_preserved(__memcpy_real_area);
-pte_t *__bootdata_preserved(memcpy_real_ptep);
+hw_pte_t *__bootdata_preserved(memcpy_real_ptep);
 static DEFINE_MUTEX(memcpy_real_mutex);
 
 static notrace long s390_kernel_write_odd(void *dst, const void *src, size_t size)
diff --git a/arch/s390/mm/pageattr.c b/arch/s390/mm/pageattr.c
index 1e202e3d08e7..b5a055791611 100644
--- a/arch/s390/mm/pageattr.c
+++ b/arch/s390/mm/pageattr.c
@@ -79,7 +79,8 @@ static void pgt_set(unsigned long *old, unsigned long new, unsigned long addr,
 static int walk_pte_level(pmd_t *pmdp, unsigned long addr, unsigned long end,
 			  unsigned long flags)
 {
-	pte_t *ptep, new;
+	hw_pte_t *ptep;
+	pte_t new;
 
 	if (flags == SET_MEMORY_4K)
 		return 0;
@@ -112,7 +113,7 @@ static int walk_pte_level(pmd_t *pmdp, unsigned long addr, unsigned long end,
 static int split_pmd_page(pmd_t *pmdp, unsigned long addr)
 {
 	unsigned long pte_addr, prot;
-	pte_t *pt_dir, *ptep;
+	hw_pte_t *pt_dir, *ptep;
 	pmd_t new, pmd;
 	int i, ro, nx;
 
@@ -421,7 +422,7 @@ bool kernel_page_present(struct page *page)
 
 #if defined(CONFIG_DEBUG_PAGEALLOC) || defined(CONFIG_KFENCE)
 
-static void ipte_range(pte_t *pte, unsigned long address, int nr)
+static void ipte_range(hw_pte_t *pte, unsigned long address, int nr)
 {
 	int i;
 
@@ -439,7 +440,8 @@ static void ipte_range(pte_t *pte, unsigned long address, int nr)
 void __kernel_map_pages(struct page *page, int numpages, int enable)
 {
 	unsigned long address;
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 	int nr, i, j;
 
 	for (i = 0; i < numpages;) {
diff --git a/arch/s390/mm/pgtable.c b/arch/s390/mm/pgtable.c
index da076928b63b..b5eca8bfb539 100644
--- a/arch/s390/mm/pgtable.c
+++ b/arch/s390/mm/pgtable.c
@@ -37,7 +37,7 @@ pgprot_t pgprot_writecombine(pgprot_t prot)
 EXPORT_SYMBOL_GPL(pgprot_writecombine);
 
 static inline void ptep_ipte_local(struct mm_struct *mm, unsigned long addr,
-				   pte_t *ptep, int nodat)
+				   hw_pte_t *ptep, int nodat)
 {
 	unsigned long opt, asce;
 
@@ -57,7 +57,7 @@ static inline void ptep_ipte_local(struct mm_struct *mm, unsigned long addr,
 }
 
 static inline void ptep_ipte_global(struct mm_struct *mm, unsigned long addr,
-				    pte_t *ptep, int nodat)
+				    hw_pte_t *ptep, int nodat)
 {
 	unsigned long opt, asce;
 
@@ -77,12 +77,12 @@ static inline void ptep_ipte_global(struct mm_struct *mm, unsigned long addr,
 }
 
 static inline pte_t ptep_flush_direct(struct mm_struct *mm,
-				      unsigned long addr, pte_t *ptep,
+				      unsigned long addr, hw_pte_t *ptep,
 				      int nodat)
 {
 	pte_t old;
 
-	old = *ptep;
+	old = ptep_get(ptep);
 	if (unlikely(pte_val(old) & _PAGE_INVALID))
 		return old;
 	atomic_inc(&mm->context.flush_count);
@@ -96,12 +96,12 @@ static inline pte_t ptep_flush_direct(struct mm_struct *mm,
 }
 
 static inline pte_t ptep_flush_lazy(struct mm_struct *mm,
-				    unsigned long addr, pte_t *ptep,
+				    unsigned long addr, hw_pte_t *ptep,
 				    int nodat)
 {
 	pte_t old;
 
-	old = *ptep;
+	old = ptep_get(ptep);
 	if (unlikely(pte_val(old) & _PAGE_INVALID))
 		return old;
 	atomic_inc(&mm->context.flush_count);
@@ -116,7 +116,7 @@ static inline pte_t ptep_flush_lazy(struct mm_struct *mm,
 }
 
 pte_t ptep_xchg_direct(struct mm_struct *mm, unsigned long addr,
-		       pte_t *ptep, pte_t new)
+		       hw_pte_t *ptep, pte_t new)
 {
 	pte_t old;
 
@@ -132,7 +132,7 @@ EXPORT_SYMBOL(ptep_xchg_direct);
  * Caller must check that new PTE only differs in _PAGE_PROTECT HW bit, so that
  * RDP can be used instead of IPTE. See also comments at pte_allow_rdp().
  */
-void ptep_reset_dat_prot(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+void ptep_reset_dat_prot(struct mm_struct *mm, unsigned long addr, hw_pte_t *ptep,
 			 pte_t new)
 {
 	preempt_disable();
@@ -154,7 +154,7 @@ void ptep_reset_dat_prot(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
 EXPORT_SYMBOL(ptep_reset_dat_prot);
 
 pte_t ptep_xchg_lazy(struct mm_struct *mm, unsigned long addr,
-		     pte_t *ptep, pte_t new)
+		     hw_pte_t *ptep, pte_t new)
 {
 	pte_t old;
 
@@ -167,13 +167,13 @@ pte_t ptep_xchg_lazy(struct mm_struct *mm, unsigned long addr,
 EXPORT_SYMBOL(ptep_xchg_lazy);
 
 pte_t ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
-			     pte_t *ptep)
+			     hw_pte_t *ptep)
 {
 	return ptep_flush_lazy(vma->vm_mm, addr, ptep, 1);
 }
 
 void ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
-			     pte_t *ptep, pte_t old_pte, pte_t pte)
+			     hw_pte_t *ptep, pte_t old_pte, pte_t pte)
 {
 	set_pte(ptep, pte);
 }
@@ -333,7 +333,7 @@ pgtable_t pgtable_trans_huge_withdraw(struct mm_struct *mm, pmd_t *pmdp)
 {
 	struct list_head *lh;
 	pgtable_t pgtable;
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	assert_spin_locked(pmd_lockptr(mm, pmdp));
 
@@ -346,7 +346,7 @@ pgtable_t pgtable_trans_huge_withdraw(struct mm_struct *mm, pmd_t *pmdp)
 		pmd_huge_pte(mm, pmdp) = (pgtable_t) lh->next;
 		list_del(lh);
 	}
-	ptep = (pte_t *) pgtable;
+	ptep = (hw_pte_t *) pgtable;
 	set_pte(ptep, __pte(_PAGE_INVALID));
 	ptep++;
 	set_pte(ptep, __pte(_PAGE_INVALID));
diff --git a/arch/s390/mm/vmem.c b/arch/s390/mm/vmem.c
index d2879ce860a1..22737319a4c8 100644
--- a/arch/s390/mm/vmem.c
+++ b/arch/s390/mm/vmem.c
@@ -66,14 +66,14 @@ void *vmem_crst_alloc(unsigned long val)
 	return table;
 }
 
-pte_t __ref *vmem_pte_alloc(void)
+hw_pte_t __ref *vmem_pte_alloc(void)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	if (slab_is_available())
-		pte = (pte_t *)page_table_alloc(&init_mm);
+		pte = (hw_pte_t *)page_table_alloc(&init_mm);
 	else
-		pte = (pte_t *)memblock_alloc(PAGE_SIZE, PAGE_SIZE);
+		pte = (hw_pte_t *)memblock_alloc(PAGE_SIZE, PAGE_SIZE);
 	if (!pte)
 		return NULL;
 	memset64((u64 *)pte, _PAGE_INVALID, PTRS_PER_PTE);
@@ -168,7 +168,8 @@ static int __ref modify_pte_table(pmd_t *pmd, unsigned long addr,
 {
 	unsigned long prot, pages = 0;
 	int ret = -ENOMEM;
-	pte_t *pte, entry;
+	hw_pte_t *pte;
+	pte_t entry;
 
 	prot = pgprot_val(PAGE_KERNEL);
 	pte = pte_offset_kernel(pmd, addr);
@@ -204,7 +205,7 @@ static int __ref modify_pte_table(pmd_t *pmd, unsigned long addr,
 
 static void try_free_pte_table(pmd_t *pmd, unsigned long start)
 {
-	pte_t *pte;
+	hw_pte_t *pte;
 	int i;
 
 	/* We can safely assume this is fully in 1:1 mapping & vmemmap area */
@@ -226,7 +227,7 @@ static int __ref modify_pmd_table(pud_t *pud, unsigned long addr,
 	int ret = -ENOMEM;
 	pmd_t entry;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	prot = pgprot_val(SEGMENT_KERNEL);
 	pmd = pmd_offset(pud, addr);
@@ -575,16 +576,16 @@ int vmem_add_mapping(unsigned long start, unsigned long size)
  * while traversing is an error, since the function is expected to be
  * called against virtual regions reserved for 4KB mappings only.
  */
-pte_t *vmem_get_alloc_pte(unsigned long addr, bool alloc)
+hw_pte_t *vmem_get_alloc_pte(unsigned long addr, bool alloc)
 {
-	pte_t *ptep = NULL;
+	hw_pte_t *ptep = NULL;
 	pud_t pud_entry;
 	pmd_t pmd_entry;
 	pgd_t *pgd;
 	p4d_t *p4d;
 	pud_t *pud;
 	pmd_t *pmd;
-	pte_t *pte;
+	hw_pte_t *pte;
 
 	pgd = pgd_offset_k(addr);
 	if (pgd_none(pgdp_get(pgd))) {
@@ -635,7 +636,8 @@ pte_t *vmem_get_alloc_pte(unsigned long addr, bool alloc)
 
 int __vmem_map_4k_page(unsigned long addr, unsigned long phys, pgprot_t prot, bool alloc)
 {
-	pte_t *ptep, pte;
+	hw_pte_t *ptep;
+	pte_t pte;
 
 	if (!IS_ALIGNED(addr, PAGE_SIZE))
 		return -EINVAL;
@@ -660,7 +662,7 @@ int vmem_map_4k_page(unsigned long addr, unsigned long phys, pgprot_t prot)
 
 void vmem_unmap_4k_page(unsigned long addr)
 {
-	pte_t *ptep;
+	hw_pte_t *ptep;
 
 	mutex_lock(&vmem_mutex);
 	ptep = virt_to_kpte(addr);
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 12/15] mm: Make lazy MMU mode context-aware
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (10 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 11/15] s390: Distinguish hardware and software PTEs Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (2 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Lazy MMU mode is assumed to be context-independent, in the sense
that it does not need any additional information while operating.
However, the s390 architecture benefits from knowing the exact
page table entries being modified.

Introduce lazy_mmu_mode_enable_with_ptes(), which is provided with
the process address space and the page table being operated on.
This information is required to enable s390-specific optimizations.

The function takes parameters that are typically passed to page-
table level walkers, which implies that the span of PTE entries
never crosses a page table boundary.

Architectures that do not require such information simply do not
need to define the lazy_mmu_mode_enable_with_ptes() callback.

Reviewed-by: Kevin Brodsky <kevin.brodsky@arm.com>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 fs/proc/task_mmu.c      |  2 +-
 include/linux/pgtable.h | 46 +++++++++++++++++++++++++++++++++++++++++
 mm/madvise.c            |  8 +++----
 mm/memory.c             |  8 +++----
 mm/mprotect.c           |  2 +-
 mm/mremap.c             |  2 +-
 mm/vmalloc.c            |  6 +++---
 7 files changed, 60 insertions(+), 14 deletions(-)

diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index 2b5af72bbe45..3ae43f4e78b2 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -2884,7 +2884,7 @@ static int pagemap_scan_pmd_entry(pmd_t *pmd, unsigned long start,
 		return 0;
 	}
 
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(vma->vm_mm, start, end, start_pte);
 
 	if ((p->arg.flags & PM_SCAN_WP_MATCHING) && !p->vec_out) {
 		/* Fast path for performing exclusive WP */
diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h
index 1768421755a9..1e6240269d4f 100644
--- a/include/linux/pgtable.h
+++ b/include/linux/pgtable.h
@@ -272,6 +272,50 @@ static inline void lazy_mmu_mode_enable(void)
 		arch_enter_lazy_mmu_mode();
 }
 
+#ifndef arch_enter_lazy_mmu_mode_with_ptes
+static inline void arch_enter_lazy_mmu_mode_with_ptes(struct mm_struct *mm,
+		unsigned long addr, unsigned long end, hw_pte_t *ptep)
+{
+	arch_enter_lazy_mmu_mode();
+}
+#endif
+
+/**
+ * lazy_mmu_mode_enable_with_ptes() - Enable the lazy MMU mode with a speedup hint.
+ * @mm: Address space the pages are mapped into.
+ * @addr: Start address of the range.
+ * @end: End address of the range.
+ * @ptep: Page table pointer for the first entry.
+ *
+ * Enters a new lazy MMU mode section; if the mode was not already enabled,
+ * enables it and calls arch_enter_lazy_mmu_mode_with_ptes().
+ *
+ * PTEs that fall within the specified range might observe update speedups.
+ * The PTEs must belong to the specified address space and be in the same PMD.
+ *
+ * There are no requirements on the order or range completeness of PTE
+ * updates for the specified range.
+ *
+ * Must be paired with a call to lazy_mmu_mode_disable().
+ *
+ * Has no effect if called:
+ * - While paused - see lazy_mmu_mode_pause()
+ * - In interrupt context
+ */
+static inline void lazy_mmu_mode_enable_with_ptes(struct mm_struct *mm,
+		unsigned long addr, unsigned long end, hw_pte_t *ptep)
+{
+	struct lazy_mmu_state *state = &current->lazy_mmu_state;
+
+	if (in_interrupt() || state->pause_count > 0)
+		return;
+
+	VM_WARN_ON_ONCE(state->enable_count == U8_MAX);
+
+	if (state->enable_count++ == 0)
+		arch_enter_lazy_mmu_mode_with_ptes(mm, addr, end, ptep);
+}
+
 /**
  * lazy_mmu_mode_disable() - Disable the lazy MMU mode.
  *
@@ -388,6 +432,8 @@ static inline void lazy_mmu_mode_resume(void)
 }
 #else
 static inline void lazy_mmu_mode_enable(void) {}
+static inline void lazy_mmu_mode_enable_with_ptes(struct mm_struct *mm,
+		unsigned long addr, unsigned long end, hw_pte_t *ptep) {}
 static inline void lazy_mmu_mode_disable(void) {}
 static inline void lazy_mmu_mode_pause(void) {}
 static inline void lazy_mmu_mode_resume(void) {}
diff --git a/mm/madvise.c b/mm/madvise.c
index 40ce0d5940ef..12830d445499 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -462,7 +462,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
 	if (!start_pte)
 		return 0;
 	flush_tlb_batched_pending(mm);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, addr, end, start_pte);
 	for (; addr < end; pte += nr, addr += nr * PAGE_SIZE) {
 		nr = 1;
 		ptent = ptep_get(pte);
@@ -517,7 +517,7 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd,
 				if (!start_pte)
 					break;
 				flush_tlb_batched_pending(mm);
-				lazy_mmu_mode_enable();
+				lazy_mmu_mode_enable_with_ptes(mm, addr, end, start_pte);
 				if (!err)
 					nr = 0;
 				continue;
@@ -685,7 +685,7 @@ static int madvise_free_pte_range(pmd_t *pmd, unsigned long addr,
 	if (!start_pte)
 		return 0;
 	flush_tlb_batched_pending(mm);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, addr, end, start_pte);
 	for (; addr != end; pte += nr, addr += PAGE_SIZE * nr) {
 		nr = 1;
 		ptent = ptep_get(pte);
@@ -745,7 +745,7 @@ static int madvise_free_pte_range(pmd_t *pmd, unsigned long addr,
 				if (!start_pte)
 					break;
 				flush_tlb_batched_pending(mm);
-				lazy_mmu_mode_enable();
+				lazy_mmu_mode_enable_with_ptes(mm, addr, end, start_pte);
 				if (!err)
 					nr = 0;
 				continue;
diff --git a/mm/memory.c b/mm/memory.c
index 48d6ca89b0fe..69a81e867083 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -1354,7 +1354,7 @@ copy_pte_range(struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
 	spin_lock_nested(src_ptl, SINGLE_DEPTH_NESTING);
 	orig_src_pte = src_pte;
 	orig_dst_pte = dst_pte;
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(src_mm, addr, end, src_pte);
 
 	do {
 		nr = 1;
@@ -2053,7 +2053,7 @@ static unsigned long zap_pte_range(struct mmu_gather *tlb,
 		return addr;
 
 	flush_tlb_batched_pending(mm);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, addr, end, start_pte);
 	do {
 		bool any_skipped = false;
 
@@ -3051,7 +3051,7 @@ static int remap_pte_range(struct mm_struct *mm, pmd_t *pmd,
 	mapped_pte = pte = pte_alloc_map_lock(mm, pmd, addr, &ptl);
 	if (!pte)
 		return -ENOMEM;
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, addr, end, mapped_pte);
 	do {
 		BUG_ON(!pte_none(ptep_get(pte)));
 		if (!pfn_modify_allowed(pfn, prot)) {
@@ -3463,7 +3463,7 @@ static int apply_to_pte_range(struct mm_struct *mm, pmd_t *pmd,
 			return -EINVAL;
 	}
 
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, addr, end, mapped_pte);
 
 	if (fn) {
 		do {
diff --git a/mm/mprotect.c b/mm/mprotect.c
index 1d23475e0bb7..6bb112ba1b91 100644
--- a/mm/mprotect.c
+++ b/mm/mprotect.c
@@ -351,7 +351,7 @@ static long change_pte_range(struct mmu_gather *tlb,
 		is_private_single_threaded = vma_is_single_threaded_private(vma);
 
 	flush_tlb_batched_pending(vma->vm_mm);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(vma->vm_mm, addr, end, pte);
 	do {
 		nr_ptes = 1;
 		oldpte = ptep_get(pte);
diff --git a/mm/mremap.c b/mm/mremap.c
index 5272cea12f80..6a19393759de 100644
--- a/mm/mremap.c
+++ b/mm/mremap.c
@@ -260,7 +260,7 @@ static int move_ptes(struct pagetable_move_control *pmc,
 	if (new_ptl != old_ptl)
 		spin_lock_nested(new_ptl, SINGLE_DEPTH_NESTING);
 	flush_tlb_batched_pending(vma->vm_mm);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(mm, old_addr, old_end, old_ptep);
 
 	for (; old_addr < old_end; old_ptep += nr_ptes, old_addr += nr_ptes * PAGE_SIZE,
 		new_ptep += nr_ptes, new_addr += nr_ptes * PAGE_SIZE) {
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index 4f97f99c340e..978029669cf4 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -110,7 +110,7 @@ static int vmap_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 	if (!pte)
 		return -ENOMEM;
 
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(&init_mm, addr, end, pte);
 
 	do {
 		if (unlikely(!pte_none(ptep_get(pte)))) {
@@ -394,7 +394,7 @@ static void vunmap_pte_range(pmd_t *pmd, unsigned long addr, unsigned long end,
 	unsigned long size = PAGE_SIZE;
 
 	pte = pte_offset_kernel(pmd, addr);
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(&init_mm, addr, end, pte);
 
 	do {
 #ifdef CONFIG_HUGETLB_PAGE
@@ -561,7 +561,7 @@ static int vmap_pages_pte_range(pmd_t *pmd, unsigned long addr,
 	if (!pte)
 		return -ENOMEM;
 
-	lazy_mmu_mode_enable();
+	lazy_mmu_mode_enable_with_ptes(&init_mm, addr, end, pte);
 
 	do {
 		struct page *page = pages[*nr];
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (11 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 12/15] mm: Make lazy MMU mode context-aware Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 15/15] s390/mm: Lazy " Alexander Gordeev
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Make use of the IPTE instruction's "Additional Entries" field to
invalidate multiple PTEs in one go while in lazy MMU mode. This
is the mode in which many memory-management system calls (like
mremap(), mprotect(), etc.) update memory attributes.

To achieve that, the set_pte() and ptep_get() primitives use a
per-CPU cache to store and retrieve PTE values and apply the
cached values to the real page table once lazy MMU mode is left.

The same is done for memory-management platform callbacks that
would otherwise cause intense per-PTE IPTE traffic, reducing the
number of IPTE instructions from up to PTRS_PER_PTE to a single
instruction in the best case. The average reduction is of course
smaller.

Since all existing page table iterators called in lazy MMU mode
handle one table at a time, the per-CPU cache does not need to be
larger than PTRS_PER_PTE entries. That also naturally aligns with
the IPTE instruction, which must not cross a page table boundary.

Before this change, the system calls did:

    lazy_mmu_mode_enable_with_ptes()
    ...
    <update PTEs>		// up to PTRS_PER_PTE single-IPTEs
    ...
    lazy_mmu_mode_disable()

With this change, the system calls do:

    lazy_mmu_mode_enable_with_ptes()
    ...
    <store new PTE values in the per-CPU cache>
    ...
    lazy_mmu_mode_disable()	// apply cache with one multi-IPTE

When applied to large memory ranges, some system calls show
significant speedups:

    mprotect()    ~15x
    munmap()      ~3x
    mremap()      ~28x

The overall results depend on memory size and access patterns,
but the change generally does not degrade performance.

In addition to a process-wide impact, the rework affects the
whole Central Electronics Complex (CEC). Each (global) IPTE
instruction initiates a quiesce state in a CEC, so reducing
the number of IPTE calls relieves CEC-wide quiesce traffic.

In an extreme case of mprotect() contiguously triggering the
quiesce state on four LPARs in parallel, measurements show
~25x fewer quiesce events.

If an interrupt arrives in the middle of enter_ipte_range() or
leave_ipte_range() the state of the per-cpu struct ipte_range
may be inconsistent. That could lead to crashes if the interrupt
handler tries to access a PTE. e.g:

handle_softirqs()->...->free_percpu()->pcpu_chunk_addr_search()->
pcpu_addr_to_page()->vmalloc_to_page()->ptep_get()

To avoid that disable bottom halves on the lazy mmu mode entering
or leaving.

Though unexpected nothing prevents a top half from accessing a PTE
that belongs to an active lazy mmu mode range. In that case do
VM_BUG_ON() and consider disabling the interrupts altogether if it
ever hits.

Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 arch/s390/Kconfig               |   1 +
 arch/s390/include/asm/lowcore.h |   3 +-
 arch/s390/include/asm/pgtable.h | 163 ++++++++++--
 arch/s390/mm/Makefile           |   2 +-
 arch/s390/mm/lazy_mmu.c         | 430 ++++++++++++++++++++++++++++++++
 arch/s390/mm/pgtable.c          |   8 +-
 6 files changed, 583 insertions(+), 24 deletions(-)
 create mode 100644 arch/s390/mm/lazy_mmu.c

diff --git a/arch/s390/Kconfig b/arch/s390/Kconfig
index b9bc0e5e7d7d..65482c3c8588 100644
--- a/arch/s390/Kconfig
+++ b/arch/s390/Kconfig
@@ -102,6 +102,7 @@ config S390
 	select ARCH_HAS_GIGANTIC_PAGE
 	select ARCH_HAS_HW_PTE_T
 	select ARCH_HAS_KCOV
+	select ARCH_HAS_LAZY_MMU_MODE
 	select ARCH_HAS_MEMBARRIER_SYNC_CORE
 	select ARCH_HAS_MEM_ENCRYPT
 	select ARCH_HAS_NMI_SAFE_THIS_CPU_OPS
diff --git a/arch/s390/include/asm/lowcore.h b/arch/s390/include/asm/lowcore.h
index 5cef215d30e7..6ec815e7feed 100644
--- a/arch/s390/include/asm/lowcore.h
+++ b/arch/s390/include/asm/lowcore.h
@@ -175,7 +175,8 @@ struct lowcore {
 
 	__u32	return_lpswe;			/* 0x0400 */
 	__u32	return_mcck_lpswe;		/* 0x0404 */
-	__u8	pad_0x040a[0x0e00-0x0408];	/* 0x0408 */
+	__u8	lazy_mmu_count;			/* 0x0408 */
+	__u8	pad_0x0409[0x0e00-0x0409];	/* 0x0409 */
 
 	/*
 	 * 0xe00 contains the address of the IPL Parameter Information
diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
index c47264f3abf2..decce76fac74 100644
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -39,6 +39,76 @@ enum {
 
 extern atomic_long_t direct_pages_count[PG_DIRECT_MAP_MAX];
 
+bool __lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res);
+bool __lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res);
+bool __lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res);
+bool __lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep, pte_t old_pte, pte_t pte);
+bool __lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep);
+bool __lazy_mmu_set_pte(hw_pte_t *ptep, pte_t pte);
+bool __lazy_mmu_ptep_get(hw_pte_t *ptep, pte_t *res);
+
+static __always_inline bool is_lazy_mmu_active(void)
+{
+	unsigned long lc_lazy_mmu_count;
+	int cc;
+
+	if (__is_defined(__DECOMPRESSOR))
+		return false;
+	lc_lazy_mmu_count = offsetof(struct lowcore, lazy_mmu_count);
+	asm_inline(
+		ALTERNATIVE("	cliy	%[offzero](%%r0),0\n",
+			    "	cliy	%[offalt](%%r0),0\n",
+			    ALT_FEATURE(MFEATURE_LOWCORE))
+		CC_IPM(cc)
+		: CC_OUT(cc, cc)
+		: [offzero] "i" (lc_lazy_mmu_count),
+		  [offalt] "i" (lc_lazy_mmu_count + LOWCORE_ALT_ADDRESS),
+		  "m" (((struct lowcore *)0)->lazy_mmu_count)
+		: CC_CLOBBER);
+	return CC_TRANSFORM(cc);
+}
+
+static inline
+bool lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res)
+{
+	if (!is_lazy_mmu_active())
+		return false;
+	return __lazy_mmu_ptep_test_and_clear_young(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+	if (!is_lazy_mmu_active())
+		return false;
+	return __lazy_mmu_ptep_get_and_clear(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+	if (!is_lazy_mmu_active())
+		return false;
+	return __lazy_mmu_ptep_modify_prot_start(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep,
+				      pte_t old_pte, pte_t pte)
+{
+	if (!is_lazy_mmu_active())
+		return false;
+	return __lazy_mmu_ptep_modify_prot_commit(addr, ptep, old_pte, pte);
+}
+
+static inline
+bool lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep)
+{
+	if (!is_lazy_mmu_active())
+		return false;
+	return __lazy_mmu_ptep_set_wrprotect(addr, ptep);
+}
+
 static inline void update_page_count(int level, long count)
 {
 	if (IS_ENABLED(CONFIG_PROC_FS))
@@ -978,7 +1048,7 @@ static inline void set_pmd(pmd_t *pmdp, pmd_t pmd)
 	WRITE_ONCE(*pmdp, pmd);
 }
 
-static inline void set_pte(hw_pte_t *ptep, pte_t pte)
+static inline void __set_pte(hw_pte_t *ptep, pte_t pte)
 {
 	hw_pte_t hwpte = (hw_pte_t) { (pte) };
 
@@ -987,10 +1057,25 @@ static inline void set_pte(hw_pte_t *ptep, pte_t pte)
 	WRITE_ONCE(*ptep, hwpte);
 }
 
+static inline void set_pte(hw_pte_t *ptep, pte_t pte)
+{
+	if (!is_lazy_mmu_active() || !__lazy_mmu_set_pte(ptep, pte))
+		__set_pte(ptep, pte);
+}
+
+static inline pte_t __ptep_get(hw_pte_t *ptep)
+{
+	return __pte_from_hw(READ_ONCE(*ptep));
+}
+
 #define ptep_get ptep_get
 static inline pte_t ptep_get(hw_pte_t *ptep)
 {
-	return __pte_from_hw(READ_ONCE(*ptep));
+	pte_t res;
+
+	if (!is_lazy_mmu_active() || !__lazy_mmu_ptep_get(ptep, &res))
+		res = __ptep_get(ptep);
+	return res;
 }
 
 #define pmdp_get pmdp_get
@@ -1183,6 +1268,15 @@ static __always_inline void __ptep_ipte_range(unsigned long address, int nr,
 	} while (nr != 255);
 }
 
+void arch_enter_lazy_mmu_mode_with_ptes(struct mm_struct *mm,
+					unsigned long addr, unsigned long end,
+					hw_pte_t *pte);
+#define arch_enter_lazy_mmu_mode_with_ptes arch_enter_lazy_mmu_mode_with_ptes
+
+void arch_enter_lazy_mmu_mode(void);
+void arch_leave_lazy_mmu_mode(void);
+void arch_flush_lazy_mmu_mode(void);
+
 /*
  * This is hard to understand. ptep_get_and_clear and ptep_clear_flush
  * both clear the TLB for the unmapped pte. The reason is that
@@ -1203,10 +1297,16 @@ pte_t ptep_xchg_lazy(struct mm_struct *, unsigned long, hw_pte_t *, pte_t);
 static inline bool ptep_test_and_clear_young(struct vm_area_struct *vma,
 		unsigned long addr, hw_pte_t *ptep)
 {
-	pte_t pte = ptep_get(ptep);
+	pte_t pte;
+	int res;
 
-	pte = ptep_xchg_direct(vma->vm_mm, addr, ptep, pte_mkold(pte));
-	return pte_young(pte);
+	if (!lazy_mmu_ptep_test_and_clear_young(addr, ptep, &res)) {
+		pte = __ptep_get(ptep);
+		pte = pte_mkold(pte);
+		pte = ptep_xchg_direct(vma->vm_mm, addr, ptep, pte);
+		res = pte_young(pte);
+	}
+	return res;
 }
 
 #define __HAVE_ARCH_PTEP_CLEAR_YOUNG_FLUSH
@@ -1222,7 +1322,8 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
 {
 	pte_t res;
 
-	res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+	if (!lazy_mmu_ptep_get_and_clear(addr, ptep, &res))
+		res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
 	page_table_check_pte_clear(mm, addr, res);
 	/* At this point the reference through the mapping is still present */
 	if (mm_is_protected(mm) && pte_present(res))
@@ -1231,9 +1332,28 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
 }
 
 #define __HAVE_ARCH_PTEP_MODIFY_PROT_TRANSACTION
-pte_t ptep_modify_prot_start(struct vm_area_struct *, unsigned long, hw_pte_t *);
-void ptep_modify_prot_commit(struct vm_area_struct *, unsigned long,
-			     hw_pte_t *, pte_t, pte_t);
+pte_t ___ptep_modify_prot_start(struct vm_area_struct *, unsigned long, hw_pte_t *);
+void ___ptep_modify_prot_commit(struct vm_area_struct *, unsigned long,
+				hw_pte_t *, pte_t, pte_t);
+
+static inline
+pte_t ptep_modify_prot_start(struct vm_area_struct *vma,
+			     unsigned long addr, hw_pte_t *ptep)
+{
+	pte_t res;
+
+	if (!lazy_mmu_ptep_modify_prot_start(addr, ptep, &res))
+		res = ___ptep_modify_prot_start(vma, addr, ptep);
+	return res;
+}
+
+static inline
+void ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
+			     hw_pte_t *ptep, pte_t old_pte, pte_t pte)
+{
+	if (!lazy_mmu_ptep_modify_prot_commit(addr, ptep, old_pte, pte))
+		___ptep_modify_prot_commit(vma, addr, ptep, old_pte, pte);
+}
 
 #define __HAVE_ARCH_PTEP_CLEAR_FLUSH
 static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
@@ -1263,11 +1383,13 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
 {
 	pte_t res;
 
-	if (full) {
-		res = ptep_get(ptep);
-		set_pte(ptep, __pte(_PAGE_INVALID));
-	} else {
-		res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+	if (!lazy_mmu_ptep_get_and_clear(addr, ptep, &res)) {
+		if (full) {
+			res = __ptep_get(ptep);
+			__set_pte(ptep, __pte(_PAGE_INVALID));
+		} else {
+			res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+		}
 	}
 	page_table_check_pte_clear(mm, addr, res);
 	/* At this point the reference through the mapping is still present */
@@ -1293,10 +1415,15 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
 static inline void ptep_set_wrprotect(struct mm_struct *mm,
 				      unsigned long addr, hw_pte_t *ptep)
 {
-	pte_t pte = ptep_get(ptep);
+	pte_t pte;
 
-	if (pte_write(pte))
-		ptep_xchg_lazy(mm, addr, ptep, pte_wrprotect(pte));
+	if (!lazy_mmu_ptep_set_wrprotect(addr, ptep)) {
+		pte = __ptep_get(ptep);
+		if (pte_write(pte)) {
+			pte = pte_wrprotect(pte);
+			ptep_xchg_lazy(mm, addr, ptep, pte);
+		}
+	}
 }
 
 /*
@@ -1329,7 +1456,7 @@ static inline void flush_tlb_fix_spurious_fault(struct vm_area_struct *vma,
 	 * PTE does not have _PAGE_PROTECT set, to avoid unnecessary overhead.
 	 * A local RDP can be used to do the flush.
 	 */
-	if (cpu_has_rdp() && !(pte_val(ptep_get(ptep)) & _PAGE_PROTECT))
+	if (cpu_has_rdp() && !(pte_val(__ptep_get(ptep)) & _PAGE_PROTECT))
 		__ptep_rdp(address, ptep, 1);
 }
 #define flush_tlb_fix_spurious_fault flush_tlb_fix_spurious_fault
diff --git a/arch/s390/mm/Makefile b/arch/s390/mm/Makefile
index 7dea37a5ad3b..6a9f856a94c1 100644
--- a/arch/s390/mm/Makefile
+++ b/arch/s390/mm/Makefile
@@ -5,7 +5,7 @@
 
 CONTEXT_ANALYSIS := y
 
-obj-y		:= init.o fault.o extmem.o mmap.o vmem.o maccess.o
+obj-y		:= init.o fault.o extmem.o mmap.o vmem.o maccess.o lazy_mmu.o
 obj-y		+= page-states.o pageattr.o pgtable.o pgalloc.o extable.o
 
 obj-$(CONFIG_CMM)		+= cmm.o
diff --git a/arch/s390/mm/lazy_mmu.c b/arch/s390/mm/lazy_mmu.c
new file mode 100644
index 000000000000..8c1a62f23736
--- /dev/null
+++ b/arch/s390/mm/lazy_mmu.c
@@ -0,0 +1,430 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <linux/pgtable.h>
+#include <linux/kasan.h>
+#include <linux/slab.h>
+#include <linux/cpuhotplug.h>
+#include <asm/facility.h>
+#include <kunit/visibility.h>
+
+#define PTE_POISON	_PAGE_LARGE
+
+struct ipte_range {
+	struct mm_struct *mm;
+	unsigned long base_addr;
+	unsigned long base_end;
+	hw_pte_t *base_pte;
+	hw_pte_t *start_pte;
+	hw_pte_t *end_pte;
+	pte_t cache[PTRS_PER_PTE];
+};
+
+static DEFINE_PER_CPU(struct ipte_range *, ipte_range);
+static DEFINE_STATIC_KEY_FALSE(lazy_mmu_enabled);
+
+static int count_contiguous(hw_pte_t *start, hw_pte_t *end, bool *valid)
+{
+	unsigned long page_invalid_bit;
+	hw_pte_t *ptep;
+	pte_t pte;
+
+	pte = __ptep_get(start);
+	page_invalid_bit = pte_val(pte) & _PAGE_INVALID;
+
+	for (ptep = start + 1; ptep < end; ptep++) {
+		pte = __ptep_get(ptep);
+		if ((pte_val(pte) & _PAGE_INVALID) != page_invalid_bit)
+			break;
+	}
+
+	*valid = !(page_invalid_bit);
+	return ptep - start;
+}
+
+static void __invalidate_pte_range(struct mm_struct *mm, unsigned long addr,
+				   int nr_ptes, hw_pte_t *ptep)
+{
+	atomic_inc(&mm->context.flush_count);
+	if (cpu_has_tlb_lc() && cpumask_equal(mm_cpumask(mm), cpumask_of(smp_processor_id())))
+		__ptep_ipte_range(addr, nr_ptes - 1, ptep, IPTE_LOCAL);
+	else
+		__ptep_ipte_range(addr, nr_ptes - 1, ptep, IPTE_GLOBAL);
+	atomic_dec(&mm->context.flush_count);
+}
+
+static int invalidate_pte_range(struct mm_struct *mm, unsigned long addr,
+				hw_pte_t *start, hw_pte_t *end)
+{
+	int nr_ptes;
+	bool valid;
+
+	nr_ptes = count_contiguous(start, end, &valid);
+	if (valid)
+		__invalidate_pte_range(mm, addr, nr_ptes, start);
+
+	return nr_ptes;
+}
+
+static void set_pte_range(struct mm_struct *mm, unsigned long addr,
+			  hw_pte_t *ptep, hw_pte_t *end, pte_t *cache)
+{
+	int i, nr_ptes;
+
+	while (ptep < end) {
+		nr_ptes = invalidate_pte_range(mm, addr, ptep, end);
+
+		for (i = 0; i < nr_ptes; i++, ptep++, cache++) {
+			__set_pte(ptep, *cache);
+			*cache = __pte(PTE_POISON);
+		}
+
+		addr += nr_ptes * PAGE_SIZE;
+	}
+}
+
+static void enter_ipte_norange(void)
+{
+	struct ipte_range __maybe_unused *range;
+
+	if (!static_branch_likely(&lazy_mmu_enabled))
+		return;
+
+	range = get_cpu_var(ipte_range);
+	local_bh_disable();
+	get_lowcore()->lazy_mmu_count++;
+	local_bh_enable();
+}
+
+static void enter_ipte_range(struct mm_struct *mm,
+			     unsigned long addr, unsigned long end, hw_pte_t *pte)
+{
+	struct ipte_range *range;
+
+	if (!static_branch_likely(&lazy_mmu_enabled))
+		return;
+
+	range = get_cpu_var(ipte_range);
+	local_bh_disable();
+	get_lowcore()->lazy_mmu_count++;
+
+	if (mm_is_protected(mm)) {
+		local_bh_enable();
+		return;
+	}
+
+	range->mm = mm;
+	range->base_addr = addr;
+	range->base_end = end;
+	range->base_pte = pte;
+
+	local_bh_enable();
+}
+
+static void leave_ipte_range(void)
+{
+	unsigned long start_addr, addr;
+	pte_t *start_cache, *cache;
+	hw_pte_t *ptep, *start;
+	struct ipte_range *range;
+	int start_idx;
+
+	if (!static_branch_likely(&lazy_mmu_enabled))
+		return;
+
+	lockdep_assert_preemption_disabled();
+	range = this_cpu_read(ipte_range);
+	local_bh_disable();
+
+	if (!range->mm)
+		goto norange;
+	if (!range->start_pte)
+		goto done;
+
+	start = range->start_pte;
+	start_idx = range->start_pte - range->base_pte;
+	start_addr = range->base_addr + start_idx * PAGE_SIZE;
+	addr = start_addr;
+	start_cache = &range->cache[start_idx];
+	cache = start_cache;
+	for (ptep = start; ptep < range->end_pte; ptep++, cache++, addr += PAGE_SIZE) {
+		if (pte_val(*cache) == PTE_POISON) {
+			if (start) {
+				set_pte_range(range->mm, start_addr, start, ptep, start_cache);
+				start = NULL;
+			}
+		} else if (!start) {
+			start = ptep;
+			start_addr = addr;
+			start_cache = cache;
+		}
+	}
+	set_pte_range(range->mm, start_addr, start, ptep, start_cache);
+
+	range->start_pte = NULL;
+	range->end_pte = NULL;
+
+done:
+	range->mm = NULL;
+	range->base_addr = 0;
+	range->base_end = 0;
+	range->base_pte = NULL;
+
+norange:
+	get_lowcore()->lazy_mmu_count--;
+	local_bh_enable();
+
+	put_cpu_var(ipte_range);
+}
+
+static void flush_lazy_mmu_mode(void)
+{
+	unsigned long addr, end;
+	struct ipte_range *range;
+	struct mm_struct *mm;
+	hw_pte_t *pte;
+
+	if (!static_branch_likely(&lazy_mmu_enabled))
+		return;
+
+	range = get_cpu_var(ipte_range);
+	if (range->mm) {
+		mm = range->mm;
+		addr = range->base_addr;
+		end = range->base_end;
+		pte = range->base_pte;
+
+		leave_ipte_range();
+		enter_ipte_range(mm, addr, end, pte);
+	}
+	put_cpu_var(ipte_range);
+}
+
+void arch_enter_lazy_mmu_mode(void)
+{
+	enter_ipte_norange();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_enter_lazy_mmu_mode);
+
+void arch_enter_lazy_mmu_mode_with_ptes(struct mm_struct *mm,
+					unsigned long addr, unsigned long end,
+					hw_pte_t *pte)
+{
+	enter_ipte_range(mm, addr, end, pte);
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_enter_lazy_mmu_mode_with_ptes);
+
+void arch_leave_lazy_mmu_mode(void)
+{
+	leave_ipte_range();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_leave_lazy_mmu_mode);
+
+void arch_flush_lazy_mmu_mode(void)
+{
+	flush_lazy_mmu_mode();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_flush_lazy_mmu_mode);
+
+static void __ipte_range_set_pte(struct ipte_range *range, hw_pte_t *ptep, pte_t pte)
+{
+	unsigned int idx = ptep - range->base_pte;
+
+	lockdep_assert_preemption_disabled();
+	range->cache[idx] = pte;
+
+	if (!range->start_pte) {
+		range->start_pte = ptep;
+		range->end_pte = ptep + 1;
+	} else if (ptep < range->start_pte) {
+		range->start_pte = ptep;
+	} else if (ptep + 1 > range->end_pte) {
+		range->end_pte = ptep + 1;
+	}
+}
+
+static pte_t __ipte_range_ptep_get(struct ipte_range *range, hw_pte_t *ptep)
+{
+	unsigned int idx = ptep - range->base_pte;
+
+	lockdep_assert_preemption_disabled();
+	if (pte_val(range->cache[idx]) == PTE_POISON)
+		return __ptep_get(ptep);
+	return range->cache[idx];
+}
+
+static struct ipte_range *this_ipte_range(hw_pte_t *ptep)
+{
+	struct ipte_range *range;
+	unsigned int nr_ptes;
+
+	range = this_cpu_read(ipte_range);
+	if (ptep < range->base_pte)
+		return NULL;
+	nr_ptes = (range->base_end - range->base_addr) / PAGE_SIZE;
+	if (ptep >= range->base_pte + nr_ptes)
+		return NULL;
+
+	/*
+	 * The user pages are not expected to get accessed from a
+	 * hardware interrupt handler. Should such a code exist,
+	 * not only bottom, but also top halves must be disabled.
+	 */
+	VM_BUG_ON(in_hardirq());
+
+	return range;
+}
+
+bool __lazy_mmu_set_pte(hw_pte_t *ptep, pte_t pte)
+{
+	struct ipte_range *range;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	__ipte_range_set_pte(range, ptep, pte);
+
+	return true;
+}
+
+bool __lazy_mmu_ptep_get(hw_pte_t *ptep, pte_t *res)
+{
+	struct ipte_range *range;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	*res = __ipte_range_ptep_get(range, ptep);
+
+	return true;
+}
+
+bool __lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res)
+{
+	struct ipte_range *range;
+	pte_t pte, old;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	old = __ipte_range_ptep_get(range, ptep);
+	pte = pte_mkold(old);
+	__ipte_range_set_pte(range, ptep, pte);
+	*res = pte_young(old);
+
+	return true;
+}
+
+bool __lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+	struct ipte_range *range;
+	pte_t pte, old;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	old = __ipte_range_ptep_get(range, ptep);
+	pte = __pte(_PAGE_INVALID);
+	__ipte_range_set_pte(range, ptep, pte);
+	*res = old;
+
+	return true;
+}
+
+bool __lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+	return __lazy_mmu_ptep_get_and_clear(addr, ptep, res);
+}
+
+bool __lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep,
+					pte_t old_pte, pte_t pte)
+{
+	struct ipte_range *range;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	__ipte_range_set_pte(range, ptep, pte);
+
+	return true;
+}
+
+bool __lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep)
+{
+	struct ipte_range *range;
+	pte_t pte;
+
+	range = this_ipte_range(ptep);
+	if (!range)
+		return false;
+
+	pte = __ipte_range_ptep_get(range, ptep);
+	if (pte_write(pte)) {
+		pte = pte_wrprotect(pte);
+		__ipte_range_set_pte(range, ptep, pte);
+	}
+
+	return true;
+}
+
+static int lazy_mmu_alloc(unsigned int cpu)
+{
+	struct ipte_range *range;
+	int i;
+
+	range = kzalloc_obj(*range, GFP_KERNEL);
+	if (!range)
+		return -ENOMEM;
+
+	for (i = 0; i < ARRAY_SIZE(range->cache); i++)
+		range->cache[i] = __pte(PTE_POISON);
+	per_cpu(ipte_range, cpu) = range;
+
+	return 0;
+}
+
+static void lazy_mmu_free(unsigned int cpu)
+{
+	struct ipte_range *range;
+
+	range = per_cpu(ipte_range, cpu);
+	per_cpu(ipte_range, cpu) = NULL;
+	kfree(range);
+}
+
+static int lazy_mmu_cpu_online(unsigned int cpu)
+{
+	int rc;
+
+	if (!cpu) {
+		rc = lazy_mmu_alloc(0);
+		if (rc) {
+			pr_warn("Not enough memory to enable the lazy MMU mode\n");
+			return rc;
+		}
+
+		static_branch_enable_cpuslocked(&lazy_mmu_enabled);
+		return 0;
+	}
+
+	return lazy_mmu_alloc(cpu);
+}
+
+static int lazy_mmu_cpu_offline(unsigned int cpu)
+{
+	lazy_mmu_free(cpu);
+	return 0;
+}
+
+static int __init lazy_mmu_init(void)
+{
+	if (!test_facility(13))
+		return 0;
+
+	return cpuhp_setup_state(CPUHP_BP_PREPARE_DYN, "s390/lazy_mmu:online",
+				 lazy_mmu_cpu_online, lazy_mmu_cpu_offline);
+}
+early_initcall(lazy_mmu_init);
diff --git a/arch/s390/mm/pgtable.c b/arch/s390/mm/pgtable.c
index b5eca8bfb539..32aec9cc9a74 100644
--- a/arch/s390/mm/pgtable.c
+++ b/arch/s390/mm/pgtable.c
@@ -166,14 +166,14 @@ pte_t ptep_xchg_lazy(struct mm_struct *mm, unsigned long addr,
 }
 EXPORT_SYMBOL(ptep_xchg_lazy);
 
-pte_t ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
-			     hw_pte_t *ptep)
+pte_t ___ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
+				hw_pte_t *ptep)
 {
 	return ptep_flush_lazy(vma->vm_mm, addr, ptep, 1);
 }
 
-void ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
-			     hw_pte_t *ptep, pte_t old_pte, pte_t pte)
+void ___ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
+				hw_pte_t *ptep, pte_t old_pte, pte_t pte)
 {
 	set_pte(ptep, pte);
 }
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (12 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  2026-10-07 11:41 ` [PATCH v8 15/15] s390/mm: Lazy " Alexander Gordeev
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Provide helpers that allow architectures implement
illegitimate PTE direct accesses while the lazy MMU
mode is enabled, such as:

	pte_t pte = *ptep;
	*ptep = pte;

By contrast, these would have to be:

	pte_t pte = ptep_get(ptep);
	set_pte(ptep, pte);

The direct PTE accesses pose a real issue on s390.

Suggested-by: Ilya Leoshkevich <iii@linux.ibm.com>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 include/linux/kasan.h     | 19 +++++++++++++++++--
 mm/kasan/common.c         | 14 ++++++++++++++
 mm/kasan/kasan.h          |  2 ++
 mm/kasan/report_generic.c |  3 +++
 4 files changed, 36 insertions(+), 2 deletions(-)

diff --git a/include/linux/kasan.h b/include/linux/kasan.h
index ff41949c0ac0..01f710f03114 100644
--- a/include/linux/kasan.h
+++ b/include/linux/kasan.h
@@ -6,6 +6,7 @@
 #include <linux/kasan-enabled.h>
 #include <linux/kasan-tags.h>
 #include <linux/kernel.h>
+#include <linux/pgtable.h>
 #include <linux/static_key.h>
 #include <linux/types.h>
 
@@ -35,8 +36,6 @@ typedef unsigned int __bitwise kasan_vmalloc_flags_t;
 
 #if defined(CONFIG_KASAN_GENERIC) || defined(CONFIG_KASAN_SW_TAGS)
 
-#include <linux/pgtable.h>
-
 /* Software KASAN implementations use shadow memory. */
 
 #ifdef CONFIG_KASAN_SW_TAGS
@@ -134,6 +133,20 @@ static __always_inline void kasan_poison_slab(struct slab *slab)
 		__kasan_poison_slab(slab);
 }
 
+void __kasan_poison_pte(hw_pte_t *pte, int nr);
+static __always_inline void kasan_poison_pte(hw_pte_t *pte, int nr)
+{
+	if (kasan_enabled())
+		__kasan_poison_pte(pte, nr);
+}
+
+void __kasan_unpoison_pte(hw_pte_t *pte, int nr);
+static __always_inline void kasan_unpoison_pte(hw_pte_t *pte, int nr)
+{
+	if (kasan_enabled())
+		__kasan_unpoison_pte(pte, nr);
+}
+
 void __kasan_unpoison_new_object(struct kmem_cache *cache, void *object);
 /**
  * kasan_unpoison_new_object - Temporarily unpoison a new slab object.
@@ -414,6 +427,8 @@ static inline bool kasan_unpoison_pages(struct page *page, unsigned int order,
 	return false;
 }
 static inline void kasan_poison_slab(struct slab *slab) {}
+static inline void kasan_poison_pte(hw_pte_t *pte, int nr) {}
+static inline void kasan_unpoison_pte(hw_pte_t *pte, int nr) {}
 static inline void kasan_unpoison_new_object(struct kmem_cache *cache,
 					void *object) {}
 static inline void kasan_poison_new_object(struct kmem_cache *cache,
diff --git a/mm/kasan/common.c b/mm/kasan/common.c
index 7d21c4db68d2..f3cbbeb8ad0e 100644
--- a/mm/kasan/common.c
+++ b/mm/kasan/common.c
@@ -174,6 +174,20 @@ void __kasan_poison_slab(struct slab *slab)
 		     KASAN_SLAB_REDZONE, false);
 }
 
+void __kasan_poison_pte(hw_pte_t *pte, int nr)
+{
+	if (IS_ALIGNED(sizeof(*pte), KASAN_GRANULE_SIZE))
+		kasan_poison(pte, sizeof(*pte) * nr, KASAN_LAZY_MMU_PTE, false);
+}
+EXPORT_SYMBOL_GPL(__kasan_poison_pte);
+
+void __kasan_unpoison_pte(hw_pte_t *pte, int nr)
+{
+	if (IS_ALIGNED(sizeof(*pte), KASAN_GRANULE_SIZE))
+		kasan_unpoison(pte, sizeof(*pte) * nr, false);
+}
+EXPORT_SYMBOL_GPL(__kasan_unpoison_pte);
+
 void __kasan_unpoison_new_object(struct kmem_cache *cache, void *object)
 {
 	kasan_unpoison(object, cache->object_size, false);
diff --git a/mm/kasan/kasan.h b/mm/kasan/kasan.h
index fc9169a54766..1a2d18cdb21d 100644
--- a/mm/kasan/kasan.h
+++ b/mm/kasan/kasan.h
@@ -144,12 +144,14 @@ static inline bool kasan_requires_meta(void)
 #define KASAN_PAGE_REDZONE	0xFE  /* redzone for kmalloc_large allocation */
 #define KASAN_SLAB_REDZONE	0xFC  /* redzone for slab object */
 #define KASAN_SLAB_FREE		0xFB  /* freed slab object */
+#define KASAN_LAZY_MMU_PTE	0xFD  /* direct pte access in lazy mmu mode */
 #define KASAN_VMALLOC_INVALID	0xF8  /* inaccessible space in vmap area */
 #else
 #define KASAN_PAGE_FREE		KASAN_TAG_INVALID
 #define KASAN_PAGE_REDZONE	KASAN_TAG_INVALID
 #define KASAN_SLAB_REDZONE	KASAN_TAG_INVALID
 #define KASAN_SLAB_FREE		KASAN_TAG_INVALID
+#define KASAN_LAZY_MMU_PTE	KASAN_TAG_INVALID
 #define KASAN_VMALLOC_INVALID	KASAN_TAG_INVALID /* only used for SW_TAGS */
 #endif
 
diff --git a/mm/kasan/report_generic.c b/mm/kasan/report_generic.c
index f5b8e37b3805..489d4a8d6902 100644
--- a/mm/kasan/report_generic.c
+++ b/mm/kasan/report_generic.c
@@ -113,6 +113,9 @@ static const char *get_shadow_bug_type(struct kasan_report_info *info)
 	case KASAN_SLAB_FREE_META:
 		bug_type = "slab-use-after-free";
 		break;
+	case KASAN_LAZY_MMU_PTE:
+		bug_type = "lazy-mmu-pte-access";
+		break;
 	case KASAN_ALLOCA_LEFT:
 	case KASAN_ALLOCA_RIGHT:
 		bug_type = "alloca-out-of-bounds";
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v8 15/15] s390/mm: Lazy MMU mode sanitizer
  2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
                   ` (13 preceding siblings ...)
  2026-10-07 11:41 ` [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer Alexander Gordeev
@ 2026-10-07 11:41 ` Alexander Gordeev
  14 siblings, 0 replies; 16+ messages in thread
From: Alexander Gordeev @ 2026-10-07 11:41 UTC (permalink / raw)
  To: Gerald Schaefer, Heiko Carstens, Christian Borntraeger,
	Vasily Gorbik, Claudio Imbrenda, Andrey Ryabinin
  Cc: Muhammad Usama Anjum, linux-s390, linux-mm, linux-kernel, kasan-dev

Detect PTE entries access in lazy MMU mode by means other
than set_pte() and ptep_get() primitives, which would be
a read hazard.

The access to kasan shadow memory from ptep_get_lockless()
mistakenly hits invalid access in case a concurrent lazy
MMU access to the same PTE is happening. To avoid that
disable instrumentation for ptep_get_lockless() altogether.

Suggested-by: Ilya Leoshkevich <iii@linux.ibm.com>
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
 arch/s390/include/asm/pgtable.h |  6 ++++++
 arch/s390/mm/lazy_mmu.c         | 27 +++++++++++++++++++++++----
 2 files changed, 29 insertions(+), 4 deletions(-)

diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
index decce76fac74..47f0adfc1029 100644
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -1063,6 +1063,12 @@ static inline void set_pte(hw_pte_t *ptep, pte_t pte)
 		__set_pte(ptep, pte);
 }
 
+#define ptep_get_lockless ptep_get_lockless
+static inline __no_sanitize_address pte_t ptep_get_lockless(hw_pte_t *ptep)
+{
+	return __pte_from_hw(READ_ONCE(*ptep));
+}
+
 static inline pte_t __ptep_get(hw_pte_t *ptep)
 {
 	return __pte_from_hw(READ_ONCE(*ptep));
diff --git a/arch/s390/mm/lazy_mmu.c b/arch/s390/mm/lazy_mmu.c
index 8c1a62f23736..e29afa081061 100644
--- a/arch/s390/mm/lazy_mmu.c
+++ b/arch/s390/mm/lazy_mmu.c
@@ -65,10 +65,13 @@ static int invalidate_pte_range(struct mm_struct *mm, unsigned long addr,
 }
 
 static void set_pte_range(struct mm_struct *mm, unsigned long addr,
-			  hw_pte_t *ptep, hw_pte_t *end, pte_t *cache)
+			  hw_pte_t *start, hw_pte_t *end, pte_t *cache)
 {
-	int i, nr_ptes;
+	int nr_ptes, nr_total = end - start;
+	hw_pte_t *ptep = start;
+	int i;
 
+	kasan_unpoison_pte(start, nr_total);
 	while (ptep < end) {
 		nr_ptes = invalidate_pte_range(mm, addr, ptep, end);
 
@@ -79,6 +82,7 @@ static void set_pte_range(struct mm_struct *mm, unsigned long addr,
 
 		addr += nr_ptes * PAGE_SIZE;
 	}
+	kasan_poison_pte(start, nr_total);
 }
 
 static void enter_ipte_norange(void)
@@ -98,6 +102,7 @@ static void enter_ipte_range(struct mm_struct *mm,
 			     unsigned long addr, unsigned long end, hw_pte_t *pte)
 {
 	struct ipte_range *range;
+	unsigned int nr_ptes;
 
 	if (!static_branch_likely(&lazy_mmu_enabled))
 		return;
@@ -116,6 +121,9 @@ static void enter_ipte_range(struct mm_struct *mm,
 	range->base_end = end;
 	range->base_pte = pte;
 
+	nr_ptes = (range->base_end - range->base_addr) / PAGE_SIZE;
+	kasan_poison_pte(range->base_pte, nr_ptes);
+
 	local_bh_enable();
 }
 
@@ -125,6 +133,7 @@ static void leave_ipte_range(void)
 	pte_t *start_cache, *cache;
 	hw_pte_t *ptep, *start;
 	struct ipte_range *range;
+	unsigned int nr_ptes;
 	int start_idx;
 
 	if (!static_branch_likely(&lazy_mmu_enabled))
@@ -163,6 +172,9 @@ static void leave_ipte_range(void)
 	range->end_pte = NULL;
 
 done:
+	nr_ptes = (range->base_end - range->base_addr) / PAGE_SIZE;
+	kasan_unpoison_pte(range->base_pte, nr_ptes);
+
 	range->mm = NULL;
 	range->base_addr = 0;
 	range->base_end = 0;
@@ -244,10 +256,17 @@ static void __ipte_range_set_pte(struct ipte_range *range, hw_pte_t *ptep, pte_t
 static pte_t __ipte_range_ptep_get(struct ipte_range *range, hw_pte_t *ptep)
 {
 	unsigned int idx = ptep - range->base_pte;
+	pte_t pte;
 
 	lockdep_assert_preemption_disabled();
-	if (pte_val(range->cache[idx]) == PTE_POISON)
-		return __ptep_get(ptep);
+	if (pte_val(range->cache[idx]) == PTE_POISON) {
+		kasan_unpoison_pte(ptep, 1);
+		pte = __ptep_get(ptep);
+		kasan_poison_pte(ptep, 1);
+
+		return pte;
+	}
+
 	return range->cache[idx];
 }
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2026-10-07 11:42 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 04/15] mm: convert PTE table entries in ptep_get() Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 05/15] mm: convert PTE table entry to pte Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 09/15] xen: " Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 11/15] s390: Distinguish hardware and software PTEs Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 12/15] mm: Make lazy MMU mode context-aware Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 15/15] s390/mm: Lazy " Alexander Gordeev

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®