From: Alexander Gordeev <agordeev@linux.ibm.com>
To: Gerald Schaefer <gerald.schaefer@linux.ibm.com>,
Heiko Carstens <hca@linux.ibm.com>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Claudio Imbrenda <imbrenda@linux.ibm.com>,
Andrey Ryabinin <ryabinin.a.a@gmail.com>
Cc: Muhammad Usama Anjum <usama.anjum@arm.com>,
linux-s390@vger.kernel.org, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, kasan-dev@googlegroups.com
Subject: [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode
Date: Wed, 7 Oct 2026 13:41:56 +0200 [thread overview]
Message-ID: <277e24d315e55592c367ac0c347d44cac5e429f4.1791365932.git.agordeev@linux.ibm.com> (raw)
In-Reply-To: <cover.1791365932.git.agordeev@linux.ibm.com>
Make use of the IPTE instruction's "Additional Entries" field to
invalidate multiple PTEs in one go while in lazy MMU mode. This
is the mode in which many memory-management system calls (like
mremap(), mprotect(), etc.) update memory attributes.
To achieve that, the set_pte() and ptep_get() primitives use a
per-CPU cache to store and retrieve PTE values and apply the
cached values to the real page table once lazy MMU mode is left.
The same is done for memory-management platform callbacks that
would otherwise cause intense per-PTE IPTE traffic, reducing the
number of IPTE instructions from up to PTRS_PER_PTE to a single
instruction in the best case. The average reduction is of course
smaller.
Since all existing page table iterators called in lazy MMU mode
handle one table at a time, the per-CPU cache does not need to be
larger than PTRS_PER_PTE entries. That also naturally aligns with
the IPTE instruction, which must not cross a page table boundary.
Before this change, the system calls did:
lazy_mmu_mode_enable_with_ptes()
...
<update PTEs> // up to PTRS_PER_PTE single-IPTEs
...
lazy_mmu_mode_disable()
With this change, the system calls do:
lazy_mmu_mode_enable_with_ptes()
...
<store new PTE values in the per-CPU cache>
...
lazy_mmu_mode_disable() // apply cache with one multi-IPTE
When applied to large memory ranges, some system calls show
significant speedups:
mprotect() ~15x
munmap() ~3x
mremap() ~28x
The overall results depend on memory size and access patterns,
but the change generally does not degrade performance.
In addition to a process-wide impact, the rework affects the
whole Central Electronics Complex (CEC). Each (global) IPTE
instruction initiates a quiesce state in a CEC, so reducing
the number of IPTE calls relieves CEC-wide quiesce traffic.
In an extreme case of mprotect() contiguously triggering the
quiesce state on four LPARs in parallel, measurements show
~25x fewer quiesce events.
If an interrupt arrives in the middle of enter_ipte_range() or
leave_ipte_range() the state of the per-cpu struct ipte_range
may be inconsistent. That could lead to crashes if the interrupt
handler tries to access a PTE. e.g:
handle_softirqs()->...->free_percpu()->pcpu_chunk_addr_search()->
pcpu_addr_to_page()->vmalloc_to_page()->ptep_get()
To avoid that disable bottom halves on the lazy mmu mode entering
or leaving.
Though unexpected nothing prevents a top half from accessing a PTE
that belongs to an active lazy mmu mode range. In that case do
VM_BUG_ON() and consider disabling the interrupts altogether if it
ever hits.
Signed-off-by: Alexander Gordeev <agordeev@linux.ibm.com>
---
arch/s390/Kconfig | 1 +
arch/s390/include/asm/lowcore.h | 3 +-
arch/s390/include/asm/pgtable.h | 163 ++++++++++--
arch/s390/mm/Makefile | 2 +-
arch/s390/mm/lazy_mmu.c | 430 ++++++++++++++++++++++++++++++++
arch/s390/mm/pgtable.c | 8 +-
6 files changed, 583 insertions(+), 24 deletions(-)
create mode 100644 arch/s390/mm/lazy_mmu.c
diff --git a/arch/s390/Kconfig b/arch/s390/Kconfig
index b9bc0e5e7d7d..65482c3c8588 100644
--- a/arch/s390/Kconfig
+++ b/arch/s390/Kconfig
@@ -102,6 +102,7 @@ config S390
select ARCH_HAS_GIGANTIC_PAGE
select ARCH_HAS_HW_PTE_T
select ARCH_HAS_KCOV
+ select ARCH_HAS_LAZY_MMU_MODE
select ARCH_HAS_MEMBARRIER_SYNC_CORE
select ARCH_HAS_MEM_ENCRYPT
select ARCH_HAS_NMI_SAFE_THIS_CPU_OPS
diff --git a/arch/s390/include/asm/lowcore.h b/arch/s390/include/asm/lowcore.h
index 5cef215d30e7..6ec815e7feed 100644
--- a/arch/s390/include/asm/lowcore.h
+++ b/arch/s390/include/asm/lowcore.h
@@ -175,7 +175,8 @@ struct lowcore {
__u32 return_lpswe; /* 0x0400 */
__u32 return_mcck_lpswe; /* 0x0404 */
- __u8 pad_0x040a[0x0e00-0x0408]; /* 0x0408 */
+ __u8 lazy_mmu_count; /* 0x0408 */
+ __u8 pad_0x0409[0x0e00-0x0409]; /* 0x0409 */
/*
* 0xe00 contains the address of the IPL Parameter Information
diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
index c47264f3abf2..decce76fac74 100644
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -39,6 +39,76 @@ enum {
extern atomic_long_t direct_pages_count[PG_DIRECT_MAP_MAX];
+bool __lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res);
+bool __lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res);
+bool __lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res);
+bool __lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep, pte_t old_pte, pte_t pte);
+bool __lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep);
+bool __lazy_mmu_set_pte(hw_pte_t *ptep, pte_t pte);
+bool __lazy_mmu_ptep_get(hw_pte_t *ptep, pte_t *res);
+
+static __always_inline bool is_lazy_mmu_active(void)
+{
+ unsigned long lc_lazy_mmu_count;
+ int cc;
+
+ if (__is_defined(__DECOMPRESSOR))
+ return false;
+ lc_lazy_mmu_count = offsetof(struct lowcore, lazy_mmu_count);
+ asm_inline(
+ ALTERNATIVE(" cliy %[offzero](%%r0),0\n",
+ " cliy %[offalt](%%r0),0\n",
+ ALT_FEATURE(MFEATURE_LOWCORE))
+ CC_IPM(cc)
+ : CC_OUT(cc, cc)
+ : [offzero] "i" (lc_lazy_mmu_count),
+ [offalt] "i" (lc_lazy_mmu_count + LOWCORE_ALT_ADDRESS),
+ "m" (((struct lowcore *)0)->lazy_mmu_count)
+ : CC_CLOBBER);
+ return CC_TRANSFORM(cc);
+}
+
+static inline
+bool lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res)
+{
+ if (!is_lazy_mmu_active())
+ return false;
+ return __lazy_mmu_ptep_test_and_clear_young(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+ if (!is_lazy_mmu_active())
+ return false;
+ return __lazy_mmu_ptep_get_and_clear(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+ if (!is_lazy_mmu_active())
+ return false;
+ return __lazy_mmu_ptep_modify_prot_start(addr, ptep, res);
+}
+
+static inline
+bool lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep,
+ pte_t old_pte, pte_t pte)
+{
+ if (!is_lazy_mmu_active())
+ return false;
+ return __lazy_mmu_ptep_modify_prot_commit(addr, ptep, old_pte, pte);
+}
+
+static inline
+bool lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep)
+{
+ if (!is_lazy_mmu_active())
+ return false;
+ return __lazy_mmu_ptep_set_wrprotect(addr, ptep);
+}
+
static inline void update_page_count(int level, long count)
{
if (IS_ENABLED(CONFIG_PROC_FS))
@@ -978,7 +1048,7 @@ static inline void set_pmd(pmd_t *pmdp, pmd_t pmd)
WRITE_ONCE(*pmdp, pmd);
}
-static inline void set_pte(hw_pte_t *ptep, pte_t pte)
+static inline void __set_pte(hw_pte_t *ptep, pte_t pte)
{
hw_pte_t hwpte = (hw_pte_t) { (pte) };
@@ -987,10 +1057,25 @@ static inline void set_pte(hw_pte_t *ptep, pte_t pte)
WRITE_ONCE(*ptep, hwpte);
}
+static inline void set_pte(hw_pte_t *ptep, pte_t pte)
+{
+ if (!is_lazy_mmu_active() || !__lazy_mmu_set_pte(ptep, pte))
+ __set_pte(ptep, pte);
+}
+
+static inline pte_t __ptep_get(hw_pte_t *ptep)
+{
+ return __pte_from_hw(READ_ONCE(*ptep));
+}
+
#define ptep_get ptep_get
static inline pte_t ptep_get(hw_pte_t *ptep)
{
- return __pte_from_hw(READ_ONCE(*ptep));
+ pte_t res;
+
+ if (!is_lazy_mmu_active() || !__lazy_mmu_ptep_get(ptep, &res))
+ res = __ptep_get(ptep);
+ return res;
}
#define pmdp_get pmdp_get
@@ -1183,6 +1268,15 @@ static __always_inline void __ptep_ipte_range(unsigned long address, int nr,
} while (nr != 255);
}
+void arch_enter_lazy_mmu_mode_with_ptes(struct mm_struct *mm,
+ unsigned long addr, unsigned long end,
+ hw_pte_t *pte);
+#define arch_enter_lazy_mmu_mode_with_ptes arch_enter_lazy_mmu_mode_with_ptes
+
+void arch_enter_lazy_mmu_mode(void);
+void arch_leave_lazy_mmu_mode(void);
+void arch_flush_lazy_mmu_mode(void);
+
/*
* This is hard to understand. ptep_get_and_clear and ptep_clear_flush
* both clear the TLB for the unmapped pte. The reason is that
@@ -1203,10 +1297,16 @@ pte_t ptep_xchg_lazy(struct mm_struct *, unsigned long, hw_pte_t *, pte_t);
static inline bool ptep_test_and_clear_young(struct vm_area_struct *vma,
unsigned long addr, hw_pte_t *ptep)
{
- pte_t pte = ptep_get(ptep);
+ pte_t pte;
+ int res;
- pte = ptep_xchg_direct(vma->vm_mm, addr, ptep, pte_mkold(pte));
- return pte_young(pte);
+ if (!lazy_mmu_ptep_test_and_clear_young(addr, ptep, &res)) {
+ pte = __ptep_get(ptep);
+ pte = pte_mkold(pte);
+ pte = ptep_xchg_direct(vma->vm_mm, addr, ptep, pte);
+ res = pte_young(pte);
+ }
+ return res;
}
#define __HAVE_ARCH_PTEP_CLEAR_YOUNG_FLUSH
@@ -1222,7 +1322,8 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
{
pte_t res;
- res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+ if (!lazy_mmu_ptep_get_and_clear(addr, ptep, &res))
+ res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
page_table_check_pte_clear(mm, addr, res);
/* At this point the reference through the mapping is still present */
if (mm_is_protected(mm) && pte_present(res))
@@ -1231,9 +1332,28 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
}
#define __HAVE_ARCH_PTEP_MODIFY_PROT_TRANSACTION
-pte_t ptep_modify_prot_start(struct vm_area_struct *, unsigned long, hw_pte_t *);
-void ptep_modify_prot_commit(struct vm_area_struct *, unsigned long,
- hw_pte_t *, pte_t, pte_t);
+pte_t ___ptep_modify_prot_start(struct vm_area_struct *, unsigned long, hw_pte_t *);
+void ___ptep_modify_prot_commit(struct vm_area_struct *, unsigned long,
+ hw_pte_t *, pte_t, pte_t);
+
+static inline
+pte_t ptep_modify_prot_start(struct vm_area_struct *vma,
+ unsigned long addr, hw_pte_t *ptep)
+{
+ pte_t res;
+
+ if (!lazy_mmu_ptep_modify_prot_start(addr, ptep, &res))
+ res = ___ptep_modify_prot_start(vma, addr, ptep);
+ return res;
+}
+
+static inline
+void ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
+ hw_pte_t *ptep, pte_t old_pte, pte_t pte)
+{
+ if (!lazy_mmu_ptep_modify_prot_commit(addr, ptep, old_pte, pte))
+ ___ptep_modify_prot_commit(vma, addr, ptep, old_pte, pte);
+}
#define __HAVE_ARCH_PTEP_CLEAR_FLUSH
static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
@@ -1263,11 +1383,13 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
{
pte_t res;
- if (full) {
- res = ptep_get(ptep);
- set_pte(ptep, __pte(_PAGE_INVALID));
- } else {
- res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+ if (!lazy_mmu_ptep_get_and_clear(addr, ptep, &res)) {
+ if (full) {
+ res = __ptep_get(ptep);
+ __set_pte(ptep, __pte(_PAGE_INVALID));
+ } else {
+ res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
+ }
}
page_table_check_pte_clear(mm, addr, res);
/* At this point the reference through the mapping is still present */
@@ -1293,10 +1415,15 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
static inline void ptep_set_wrprotect(struct mm_struct *mm,
unsigned long addr, hw_pte_t *ptep)
{
- pte_t pte = ptep_get(ptep);
+ pte_t pte;
- if (pte_write(pte))
- ptep_xchg_lazy(mm, addr, ptep, pte_wrprotect(pte));
+ if (!lazy_mmu_ptep_set_wrprotect(addr, ptep)) {
+ pte = __ptep_get(ptep);
+ if (pte_write(pte)) {
+ pte = pte_wrprotect(pte);
+ ptep_xchg_lazy(mm, addr, ptep, pte);
+ }
+ }
}
/*
@@ -1329,7 +1456,7 @@ static inline void flush_tlb_fix_spurious_fault(struct vm_area_struct *vma,
* PTE does not have _PAGE_PROTECT set, to avoid unnecessary overhead.
* A local RDP can be used to do the flush.
*/
- if (cpu_has_rdp() && !(pte_val(ptep_get(ptep)) & _PAGE_PROTECT))
+ if (cpu_has_rdp() && !(pte_val(__ptep_get(ptep)) & _PAGE_PROTECT))
__ptep_rdp(address, ptep, 1);
}
#define flush_tlb_fix_spurious_fault flush_tlb_fix_spurious_fault
diff --git a/arch/s390/mm/Makefile b/arch/s390/mm/Makefile
index 7dea37a5ad3b..6a9f856a94c1 100644
--- a/arch/s390/mm/Makefile
+++ b/arch/s390/mm/Makefile
@@ -5,7 +5,7 @@
CONTEXT_ANALYSIS := y
-obj-y := init.o fault.o extmem.o mmap.o vmem.o maccess.o
+obj-y := init.o fault.o extmem.o mmap.o vmem.o maccess.o lazy_mmu.o
obj-y += page-states.o pageattr.o pgtable.o pgalloc.o extable.o
obj-$(CONFIG_CMM) += cmm.o
diff --git a/arch/s390/mm/lazy_mmu.c b/arch/s390/mm/lazy_mmu.c
new file mode 100644
index 000000000000..8c1a62f23736
--- /dev/null
+++ b/arch/s390/mm/lazy_mmu.c
@@ -0,0 +1,430 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <linux/pgtable.h>
+#include <linux/kasan.h>
+#include <linux/slab.h>
+#include <linux/cpuhotplug.h>
+#include <asm/facility.h>
+#include <kunit/visibility.h>
+
+#define PTE_POISON _PAGE_LARGE
+
+struct ipte_range {
+ struct mm_struct *mm;
+ unsigned long base_addr;
+ unsigned long base_end;
+ hw_pte_t *base_pte;
+ hw_pte_t *start_pte;
+ hw_pte_t *end_pte;
+ pte_t cache[PTRS_PER_PTE];
+};
+
+static DEFINE_PER_CPU(struct ipte_range *, ipte_range);
+static DEFINE_STATIC_KEY_FALSE(lazy_mmu_enabled);
+
+static int count_contiguous(hw_pte_t *start, hw_pte_t *end, bool *valid)
+{
+ unsigned long page_invalid_bit;
+ hw_pte_t *ptep;
+ pte_t pte;
+
+ pte = __ptep_get(start);
+ page_invalid_bit = pte_val(pte) & _PAGE_INVALID;
+
+ for (ptep = start + 1; ptep < end; ptep++) {
+ pte = __ptep_get(ptep);
+ if ((pte_val(pte) & _PAGE_INVALID) != page_invalid_bit)
+ break;
+ }
+
+ *valid = !(page_invalid_bit);
+ return ptep - start;
+}
+
+static void __invalidate_pte_range(struct mm_struct *mm, unsigned long addr,
+ int nr_ptes, hw_pte_t *ptep)
+{
+ atomic_inc(&mm->context.flush_count);
+ if (cpu_has_tlb_lc() && cpumask_equal(mm_cpumask(mm), cpumask_of(smp_processor_id())))
+ __ptep_ipte_range(addr, nr_ptes - 1, ptep, IPTE_LOCAL);
+ else
+ __ptep_ipte_range(addr, nr_ptes - 1, ptep, IPTE_GLOBAL);
+ atomic_dec(&mm->context.flush_count);
+}
+
+static int invalidate_pte_range(struct mm_struct *mm, unsigned long addr,
+ hw_pte_t *start, hw_pte_t *end)
+{
+ int nr_ptes;
+ bool valid;
+
+ nr_ptes = count_contiguous(start, end, &valid);
+ if (valid)
+ __invalidate_pte_range(mm, addr, nr_ptes, start);
+
+ return nr_ptes;
+}
+
+static void set_pte_range(struct mm_struct *mm, unsigned long addr,
+ hw_pte_t *ptep, hw_pte_t *end, pte_t *cache)
+{
+ int i, nr_ptes;
+
+ while (ptep < end) {
+ nr_ptes = invalidate_pte_range(mm, addr, ptep, end);
+
+ for (i = 0; i < nr_ptes; i++, ptep++, cache++) {
+ __set_pte(ptep, *cache);
+ *cache = __pte(PTE_POISON);
+ }
+
+ addr += nr_ptes * PAGE_SIZE;
+ }
+}
+
+static void enter_ipte_norange(void)
+{
+ struct ipte_range __maybe_unused *range;
+
+ if (!static_branch_likely(&lazy_mmu_enabled))
+ return;
+
+ range = get_cpu_var(ipte_range);
+ local_bh_disable();
+ get_lowcore()->lazy_mmu_count++;
+ local_bh_enable();
+}
+
+static void enter_ipte_range(struct mm_struct *mm,
+ unsigned long addr, unsigned long end, hw_pte_t *pte)
+{
+ struct ipte_range *range;
+
+ if (!static_branch_likely(&lazy_mmu_enabled))
+ return;
+
+ range = get_cpu_var(ipte_range);
+ local_bh_disable();
+ get_lowcore()->lazy_mmu_count++;
+
+ if (mm_is_protected(mm)) {
+ local_bh_enable();
+ return;
+ }
+
+ range->mm = mm;
+ range->base_addr = addr;
+ range->base_end = end;
+ range->base_pte = pte;
+
+ local_bh_enable();
+}
+
+static void leave_ipte_range(void)
+{
+ unsigned long start_addr, addr;
+ pte_t *start_cache, *cache;
+ hw_pte_t *ptep, *start;
+ struct ipte_range *range;
+ int start_idx;
+
+ if (!static_branch_likely(&lazy_mmu_enabled))
+ return;
+
+ lockdep_assert_preemption_disabled();
+ range = this_cpu_read(ipte_range);
+ local_bh_disable();
+
+ if (!range->mm)
+ goto norange;
+ if (!range->start_pte)
+ goto done;
+
+ start = range->start_pte;
+ start_idx = range->start_pte - range->base_pte;
+ start_addr = range->base_addr + start_idx * PAGE_SIZE;
+ addr = start_addr;
+ start_cache = &range->cache[start_idx];
+ cache = start_cache;
+ for (ptep = start; ptep < range->end_pte; ptep++, cache++, addr += PAGE_SIZE) {
+ if (pte_val(*cache) == PTE_POISON) {
+ if (start) {
+ set_pte_range(range->mm, start_addr, start, ptep, start_cache);
+ start = NULL;
+ }
+ } else if (!start) {
+ start = ptep;
+ start_addr = addr;
+ start_cache = cache;
+ }
+ }
+ set_pte_range(range->mm, start_addr, start, ptep, start_cache);
+
+ range->start_pte = NULL;
+ range->end_pte = NULL;
+
+done:
+ range->mm = NULL;
+ range->base_addr = 0;
+ range->base_end = 0;
+ range->base_pte = NULL;
+
+norange:
+ get_lowcore()->lazy_mmu_count--;
+ local_bh_enable();
+
+ put_cpu_var(ipte_range);
+}
+
+static void flush_lazy_mmu_mode(void)
+{
+ unsigned long addr, end;
+ struct ipte_range *range;
+ struct mm_struct *mm;
+ hw_pte_t *pte;
+
+ if (!static_branch_likely(&lazy_mmu_enabled))
+ return;
+
+ range = get_cpu_var(ipte_range);
+ if (range->mm) {
+ mm = range->mm;
+ addr = range->base_addr;
+ end = range->base_end;
+ pte = range->base_pte;
+
+ leave_ipte_range();
+ enter_ipte_range(mm, addr, end, pte);
+ }
+ put_cpu_var(ipte_range);
+}
+
+void arch_enter_lazy_mmu_mode(void)
+{
+ enter_ipte_norange();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_enter_lazy_mmu_mode);
+
+void arch_enter_lazy_mmu_mode_with_ptes(struct mm_struct *mm,
+ unsigned long addr, unsigned long end,
+ hw_pte_t *pte)
+{
+ enter_ipte_range(mm, addr, end, pte);
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_enter_lazy_mmu_mode_with_ptes);
+
+void arch_leave_lazy_mmu_mode(void)
+{
+ leave_ipte_range();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_leave_lazy_mmu_mode);
+
+void arch_flush_lazy_mmu_mode(void)
+{
+ flush_lazy_mmu_mode();
+}
+EXPORT_SYMBOL_IF_KUNIT(arch_flush_lazy_mmu_mode);
+
+static void __ipte_range_set_pte(struct ipte_range *range, hw_pte_t *ptep, pte_t pte)
+{
+ unsigned int idx = ptep - range->base_pte;
+
+ lockdep_assert_preemption_disabled();
+ range->cache[idx] = pte;
+
+ if (!range->start_pte) {
+ range->start_pte = ptep;
+ range->end_pte = ptep + 1;
+ } else if (ptep < range->start_pte) {
+ range->start_pte = ptep;
+ } else if (ptep + 1 > range->end_pte) {
+ range->end_pte = ptep + 1;
+ }
+}
+
+static pte_t __ipte_range_ptep_get(struct ipte_range *range, hw_pte_t *ptep)
+{
+ unsigned int idx = ptep - range->base_pte;
+
+ lockdep_assert_preemption_disabled();
+ if (pte_val(range->cache[idx]) == PTE_POISON)
+ return __ptep_get(ptep);
+ return range->cache[idx];
+}
+
+static struct ipte_range *this_ipte_range(hw_pte_t *ptep)
+{
+ struct ipte_range *range;
+ unsigned int nr_ptes;
+
+ range = this_cpu_read(ipte_range);
+ if (ptep < range->base_pte)
+ return NULL;
+ nr_ptes = (range->base_end - range->base_addr) / PAGE_SIZE;
+ if (ptep >= range->base_pte + nr_ptes)
+ return NULL;
+
+ /*
+ * The user pages are not expected to get accessed from a
+ * hardware interrupt handler. Should such a code exist,
+ * not only bottom, but also top halves must be disabled.
+ */
+ VM_BUG_ON(in_hardirq());
+
+ return range;
+}
+
+bool __lazy_mmu_set_pte(hw_pte_t *ptep, pte_t pte)
+{
+ struct ipte_range *range;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ __ipte_range_set_pte(range, ptep, pte);
+
+ return true;
+}
+
+bool __lazy_mmu_ptep_get(hw_pte_t *ptep, pte_t *res)
+{
+ struct ipte_range *range;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ *res = __ipte_range_ptep_get(range, ptep);
+
+ return true;
+}
+
+bool __lazy_mmu_ptep_test_and_clear_young(unsigned long addr, hw_pte_t *ptep, int *res)
+{
+ struct ipte_range *range;
+ pte_t pte, old;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ old = __ipte_range_ptep_get(range, ptep);
+ pte = pte_mkold(old);
+ __ipte_range_set_pte(range, ptep, pte);
+ *res = pte_young(old);
+
+ return true;
+}
+
+bool __lazy_mmu_ptep_get_and_clear(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+ struct ipte_range *range;
+ pte_t pte, old;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ old = __ipte_range_ptep_get(range, ptep);
+ pte = __pte(_PAGE_INVALID);
+ __ipte_range_set_pte(range, ptep, pte);
+ *res = old;
+
+ return true;
+}
+
+bool __lazy_mmu_ptep_modify_prot_start(unsigned long addr, hw_pte_t *ptep, pte_t *res)
+{
+ return __lazy_mmu_ptep_get_and_clear(addr, ptep, res);
+}
+
+bool __lazy_mmu_ptep_modify_prot_commit(unsigned long addr, hw_pte_t *ptep,
+ pte_t old_pte, pte_t pte)
+{
+ struct ipte_range *range;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ __ipte_range_set_pte(range, ptep, pte);
+
+ return true;
+}
+
+bool __lazy_mmu_ptep_set_wrprotect(unsigned long addr, hw_pte_t *ptep)
+{
+ struct ipte_range *range;
+ pte_t pte;
+
+ range = this_ipte_range(ptep);
+ if (!range)
+ return false;
+
+ pte = __ipte_range_ptep_get(range, ptep);
+ if (pte_write(pte)) {
+ pte = pte_wrprotect(pte);
+ __ipte_range_set_pte(range, ptep, pte);
+ }
+
+ return true;
+}
+
+static int lazy_mmu_alloc(unsigned int cpu)
+{
+ struct ipte_range *range;
+ int i;
+
+ range = kzalloc_obj(*range, GFP_KERNEL);
+ if (!range)
+ return -ENOMEM;
+
+ for (i = 0; i < ARRAY_SIZE(range->cache); i++)
+ range->cache[i] = __pte(PTE_POISON);
+ per_cpu(ipte_range, cpu) = range;
+
+ return 0;
+}
+
+static void lazy_mmu_free(unsigned int cpu)
+{
+ struct ipte_range *range;
+
+ range = per_cpu(ipte_range, cpu);
+ per_cpu(ipte_range, cpu) = NULL;
+ kfree(range);
+}
+
+static int lazy_mmu_cpu_online(unsigned int cpu)
+{
+ int rc;
+
+ if (!cpu) {
+ rc = lazy_mmu_alloc(0);
+ if (rc) {
+ pr_warn("Not enough memory to enable the lazy MMU mode\n");
+ return rc;
+ }
+
+ static_branch_enable_cpuslocked(&lazy_mmu_enabled);
+ return 0;
+ }
+
+ return lazy_mmu_alloc(cpu);
+}
+
+static int lazy_mmu_cpu_offline(unsigned int cpu)
+{
+ lazy_mmu_free(cpu);
+ return 0;
+}
+
+static int __init lazy_mmu_init(void)
+{
+ if (!test_facility(13))
+ return 0;
+
+ return cpuhp_setup_state(CPUHP_BP_PREPARE_DYN, "s390/lazy_mmu:online",
+ lazy_mmu_cpu_online, lazy_mmu_cpu_offline);
+}
+early_initcall(lazy_mmu_init);
diff --git a/arch/s390/mm/pgtable.c b/arch/s390/mm/pgtable.c
index b5eca8bfb539..32aec9cc9a74 100644
--- a/arch/s390/mm/pgtable.c
+++ b/arch/s390/mm/pgtable.c
@@ -166,14 +166,14 @@ pte_t ptep_xchg_lazy(struct mm_struct *mm, unsigned long addr,
}
EXPORT_SYMBOL(ptep_xchg_lazy);
-pte_t ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
- hw_pte_t *ptep)
+pte_t ___ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
+ hw_pte_t *ptep)
{
return ptep_flush_lazy(vma->vm_mm, addr, ptep, 1);
}
-void ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
- hw_pte_t *ptep, pte_t old_pte, pte_t pte)
+void ___ptep_modify_prot_commit(struct vm_area_struct *vma, unsigned long addr,
+ hw_pte_t *ptep, pte_t old_pte, pte_t pte)
{
set_pte(ptep, pte);
}
--
2.53.0
next prev parent reply other threads:[~2026-10-07 11:42 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 11:41 [PATCH v8 00/15] " Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 04/15] mm: convert PTE table entries in ptep_get() Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 05/15] mm: convert PTE table entry to pte Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 09/15] xen: " Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 11/15] s390: Distinguish hardware and software PTEs Alexander Gordeev
2026-10-08 10:32 ` kernel test robot
2026-10-07 11:41 ` [PATCH v8 12/15] mm: Make lazy MMU mode context-aware Alexander Gordeev
2026-10-07 11:41 ` Alexander Gordeev [this message]
2026-10-07 11:41 ` [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer Alexander Gordeev
2026-10-07 20:24 ` kernel test robot
2026-10-07 20:35 ` kernel test robot
2026-10-08 1:11 ` kernel test robot
2026-10-07 11:41 ` [PATCH v8 15/15] s390/mm: Lazy " Alexander Gordeev
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=277e24d315e55592c367ac0c347d44cac5e429f4.1791365932.git.agordeev@linux.ibm.com \
--to=agordeev@linux.ibm.com \
--cc=borntraeger@linux.ibm.com \
--cc=gerald.schaefer@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=imbrenda@linux.ibm.com \
--cc=kasan-dev@googlegroups.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-s390@vger.kernel.org \
--cc=ryabinin.a.a@gmail.com \
--cc=usama.anjum@arm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®