* [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
2026-10-08 3:24 ` [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner
Alpha's CONFIG_COMPACTION-gated ptep_get_and_clear() overrides the
generic version in include/linux/pgtable.h but omits its
page_table_check_pte_clear() call. With CONFIG_PAGE_TABLE_CHECK=y the
map count taken by set_ptes() is therefore never dropped when a PTE is
cleared through this path. zap_pte_range() clears PTEs this way, so the
first page freed after it hits the BUG_ON() in
__page_table_check_zero().
The call was instead placed in ptep_clear_flush() by commit dd5712f3379c
("alpha: fix user-space corruption during memory compaction"), where it
was harmless until alpha enabled page table check.
Add the missing call, and drop the now-duplicate one in
ptep_clear_flush(), which already goes through ptep_get_and_clear().
Fixes: e761b6fe4085 ("alpha: add ARCH_SUPPORTS_PAGE_TABLE_CHECK support")
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 6ebdd90f3035..8a175e0c2b42 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -305,6 +305,7 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
pte_t pte = READ_ONCE(*ptep);
pte_clear(mm, address, ptep);
+ page_table_check_pte_clear(mm, address, pte);
return pte;
}
@@ -316,7 +317,6 @@ static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
struct mm_struct *mm = vma->vm_mm;
pte_t pte = ptep_get_and_clear(mm, addr, ptep);
- page_table_check_pte_clear(mm, addr, pte);
migrate_flush_tlb_page(vma, addr);
return pte;
}
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault()
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-08 3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
2026-10-08 3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner, stable
do_page_fault() handles VM_FAULT_OOM, VM_FAULT_SIGSEGV and
VM_FAULT_SIGBUS and then falls into BUG() for anything else in
VM_FAULT_ERROR. VM_FAULT_HWPOISON and VM_FAULT_HWPOISON_LARGE are in
that set, so a fault on a poisoned page takes down the kernel instead of
delivering a signal.
Alpha does not select ARCH_SUPPORTS_MEMORY_FAILURE, so the machine check
paths cannot produce these. UFFDIO_POISON can: it installs a poison PTE
marker without any memory failure support, and a subsequent access
returns VM_FAULT_HWPOISON. Registering an anonymous range with
userfaultfd, poisoning it and reading it back reliably hits the BUG():
Kernel bug at arch/alpha/mm/fault.c:188
poison(50809): Kernel Bug 1
pc is at do_page_fault+0x4e8/0x550
ra is at do_page_fault+0xfc/0x550
The BUG() fires with mmap_read_lock() still held, so the faulting task
is left unkillable in D state holding the lock, and shutdown stalls
behind it. A UFFD_USER_MODE_ONLY userfaultfd is always allowed, so any
local user can do this.
Deliver SIGBUS with BUS_MCEERR_AR instead, reporting the size of the
poisoned area, as the other architectures do. That is the huge page size
for VM_FAULT_HWPOISON_LARGE, which becomes reachable once alpha
implements huge pages.
Tested on an UP1500 (EV68AL): the reproducer above now takes a SIGBUS
and the kernel logs
poison[373]: hardware memory error at 0000020000030000 pc 00000200010007ec
with no oops, no wedged task and no taint.
A backport needs the show_signal_msg() call dropped, since alpha only
has it since commit 9ca6af1aac40 ("alpha: select
SYSCTL_EXCEPTION_TRACE").
Fixes: fc71884a5f59 ("mm: userfaultfd: add new UFFDIO_POISON ioctl")
Cc: stable@vger.kernel.org # 6.6+
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/mm/fault.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..24408b53197c 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -8,6 +8,7 @@
#include <linux/sched/signal.h>
#include <linux/kernel.h>
#include <linux/mm.h>
+#include <linux/hugetlb.h>
#include <asm/io.h>
#define __EXTERN_INLINE inline
@@ -113,6 +114,7 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
struct mm_struct *mm = current->mm;
const struct exception_table_entry *fixup;
int si_code = SEGV_MAPERR;
+ unsigned int lsb;
vm_fault_t fault;
unsigned int flags = FAULT_FLAG_DEFAULT;
@@ -185,6 +187,8 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto bad_area;
else if (fault & VM_FAULT_SIGBUS)
goto do_sigbus;
+ else if (fault & (VM_FAULT_HWPOISON | VM_FAULT_HWPOISON_LARGE))
+ goto do_sigbus_mceerr;
BUG();
}
@@ -248,6 +252,21 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto no_context;
return;
+ do_sigbus_mceerr:
+ mmap_read_unlock(mm);
+ if (!user_mode(regs))
+ goto no_context;
+ /*
+ * Report the size of the poisoned area, which for a hugetlb fault
+ * is the size of the huge page that could not be mapped.
+ */
+ lsb = PAGE_SHIFT;
+ if (fault & VM_FAULT_HWPOISON_LARGE)
+ lsb = hstate_index_to_shift(VM_FAULT_GET_HINDEX(fault));
+ show_signal_msg(regs, address, SIGBUS, "hardware memory error");
+ force_sig_mceerr(BUS_MCEERR_AR, (void __user *) address, lsb);
+ return;
+
do_sigsegv:
show_signal_msg(regs, address, SIGSEGV,
si_code == SEGV_MAPERR ? "unmapped access"
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 3/6] alpha: describe the PTE read and write enable bits
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-08 3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
2026-10-08 3:24 ` [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
2026-10-08 3:24 ` [PATCH v2 4/6] alpha: define granularity hint PTE bits Matt Turner
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner
The comments on _PAGE_KRE and _PAGE_URE said only "xxx". Describe the
four enable bits as the Alpha Linux PTE defines them, in Table 22-3 of
the Alpha Architecture Reference Manual: kernel and user read enable in
bits 8 and 9, kernel and user write enable in bits 12 and 13, with bits
<11:10> and <15:14> reserved. There are only two processor modes, user
and kernel (section 22.5.1). Linux uses the read enables as its accessed
bit and the write enables as its dirty bit, so say which of
__ACCESS_BITS and __DIRTY_BITS each one belongs to.
The names differ from what the hardware calls those bits on the 21264.
Its PALcode loads the PTE unchanged into DTB_PTE (21264/EV67 Hardware
Reference Manual, section 6.9), where bits 9 and 13 are the Executive
read and write enables (Figure 5-27), Executive being mode 1 of the four
the processor implements (Table 5-5). That is a detail below the PALcode
interface, and it is also the layout of the OpenVMS PTE in Table 11-2,
which is not the one Linux uses.
No functional change.
Suggested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 8a175e0c2b42..d22e28b956fb 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,10 +65,10 @@ struct vm_area_struct;
#define _PAGE_FOW 0x0004 /* used for page protection (fault on write) */
#define _PAGE_FOE 0x0008 /* used for page protection (fault on exec) */
#define _PAGE_ASM 0x0010
-#define _PAGE_KRE 0x0100 /* xxx - see below on the "accessed" bit */
-#define _PAGE_URE 0x0200 /* xxx */
-#define _PAGE_KWE 0x1000 /* used to do the dirty bit in software */
-#define _PAGE_UWE 0x2000 /* used to do the dirty bit in software */
+#define _PAGE_KRE 0x0100 /* kernel read enable, in __ACCESS_BITS */
+#define _PAGE_URE 0x0200 /* user read enable, in __ACCESS_BITS */
+#define _PAGE_KWE 0x1000 /* kernel write enable, in __DIRTY_BITS */
+#define _PAGE_UWE 0x2000 /* user write enable, in __DIRTY_BITS */
/* .. and these are ours ... */
#define _PAGE_DIRTY 0x20000
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 4/6] alpha: define granularity hint PTE bits
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (2 preceding siblings ...)
2026-10-08 3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
2026-10-08 3:24 ` [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
2026-10-08 3:24 ` [PATCH v2 6/6] alpha: implement hugetlb support Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner
Bits <6:5> of the Alpha PTE are the granularity hint, described in Table
22-3 of the Alpha Architecture Reference Manual. A hint of N marks the
PTE as one of a block of 8^N physically contiguous, naturally aligned
pages that the translation buffer may map with a single entry. With 8KB
pages that gives 64KB, 512KB and 4MB blocks.
Define the field and the helpers to encode and decode it. Nothing sets a
non-zero hint yet.
Add the hint to _PAGE_CHG_MASK so that pte_modify() preserves it. Swap
PTEs leave bits <31:0> clear, so pte_huge() is false on swap, migration
and marker entries, and page table entries above the last level keep a
zero hint where pmd_bad() would reject anything else.
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 51 +++++++++++++++++++++++++++++++++++++++-
1 file changed, 50 insertions(+), 1 deletion(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index d22e28b956fb..7cf685a694b6 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,6 +65,8 @@ struct vm_area_struct;
#define _PAGE_FOW 0x0004 /* used for page protection (fault on write) */
#define _PAGE_FOE 0x0008 /* used for page protection (fault on exec) */
#define _PAGE_ASM 0x0010
+#define _PAGE_GH_MASK 0x0060 /* granularity hint, PTE<6:5> */
+#define _PAGE_GH_SHIFT 5
#define _PAGE_KRE 0x0100 /* kernel read enable, in __ACCESS_BITS */
#define _PAGE_URE 0x0200 /* user read enable, in __ACCESS_BITS */
#define _PAGE_KWE 0x1000 /* kernel write enable, in __DIRTY_BITS */
@@ -95,7 +97,36 @@ struct vm_area_struct;
#define _PFN_MASK 0xFFFFFFFF00000000UL
#define _PAGE_TABLE (_PAGE_VALID | __DIRTY_BITS | __ACCESS_BITS)
-#define _PAGE_CHG_MASK (_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | _PAGE_SPECIAL)
+/*
+ * pte_modify() must preserve the granularity hint, or
+ * hugetlb_change_protection() would turn a block into single-page PTEs.
+ */
+#define _PAGE_CHG_MASK (_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | \
+ _PAGE_GH_MASK | _PAGE_SPECIAL)
+
+/*
+ * Granularity hints. A PTE with hint order N is one of a block of 8^N
+ * physically contiguous pages, naturally aligned both virtually and
+ * physically, that the TB may map with a single entry. With 8KB pages the
+ * four encodings give 8KB, 64KB, 512KB and 4MB.
+ *
+ * Every PTE of a block must still carry the PFN of its own page, and all of
+ * them must agree in bits <15:0>: the protection, fault, hint and valid bits.
+ * __ACCESS_BITS and __DIRTY_BITS reach into that range, so young and dirty
+ * updates have to be applied to the whole block too.
+ *
+ * The hint is advisory: an implementation that ignores it still translates
+ * correctly through the individual PTEs, so it is safe to set on any Alpha.
+ */
+#define GH_ORDER_MAX 4 /* hint orders 0..3 */
+
+#define gh_cont_shift(order) (PAGE_SHIFT + 3 * (order))
+#define gh_cont_size(order) (1UL << gh_cont_shift(order))
+#define gh_cont_mask(order) (~(gh_cont_size(order) - 1))
+#define gh_pte_num(order) (1UL << (3 * (order)))
+
+#define for_each_gh_order(order) \
+ for ((order) = 1; (order) < GH_ORDER_MAX; (order)++)
/*
* All the normal masks have the "page accessed" bits on, as any time they are used,
@@ -261,6 +292,24 @@ extern inline pte_t pte_mkyoung(pte_t pte) { pte_val(pte) |= __ACCESS_BITS; retu
extern inline int pte_special(pte_t pte) { return pte_val(pte) & _PAGE_SPECIAL; }
extern inline pte_t pte_mkspecial(pte_t pte) { pte_val(pte) |= _PAGE_SPECIAL; return pte; }
+extern inline unsigned int pte_gh_order(pte_t pte)
+{
+ return (pte_val(pte) & _PAGE_GH_MASK) >> _PAGE_GH_SHIFT;
+}
+
+extern inline pte_t pte_mkgh(pte_t pte, unsigned int order)
+{
+ pte_val(pte) &= ~_PAGE_GH_MASK;
+ pte_val(pte) |= (unsigned long)order << _PAGE_GH_SHIFT;
+ return pte;
+}
+
+/*
+ * Swap PTEs leave bits <31:0> clear, so this is false for every swap,
+ * migration and marker entry, not just for ordinary small pages.
+ */
+extern inline int pte_huge(pte_t pte) { return pte_val(pte) & _PAGE_GH_MASK; }
+
/*
* The smp_rmb() in the following functions are required to order the load of
* *dir (the pointer in the top level page table) with any subsequent load of
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area()
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (3 preceding siblings ...)
2026-10-08 3:24 ` [PATCH v2 4/6] alpha: define granularity hint PTE bits Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
2026-10-08 3:24 ` [PATCH v2 6/6] alpha: implement hugetlb support Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner
hugetlb_get_unmapped_area() does not align the address itself; it
delegates to mm_get_unmapped_area_vmaflags(), and the huge page
alignment is applied by generic_get_unmapped_area() and
generic_get_unmapped_area_topdown() via info.align_mask.
Alpha defines HAVE_ARCH_UNMAPPED_AREA and supplies its own
arch_get_unmapped_area(), which never sets align_mask. An
mmap(MAP_HUGETLB) with no address hint would therefore return a merely
page-aligned address.
That has been the architecture's job since commit 7bd3f1e1a9ae ("mm:
make hugetlb mappings go through mm_get_unmapped_area_vmflags"), whose
series converted the generic code along with x86, s390, sparc and
powerpc. LoongArch, which also has its own arch_get_unmapped_area(), was
missed and hit a BUG in LTP's hugefork02 until commit 3109d5ff484b
("LoongArch: Set hugetlb mmap base address aligned with pmd size").
Pass the file down to arch_get_unmapped_area_1() and set align_mask for
hugetlbfs mappings, as x86, s390, sparc and loongarch already do. This
is a no-op until alpha gains hugetlb support.
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/kernel/osf_sys.c | 14 +++++++++-----
1 file changed, 9 insertions(+), 5 deletions(-)
diff --git a/arch/alpha/kernel/osf_sys.c b/arch/alpha/kernel/osf_sys.c
index 7b6543d2cca3..61d806176213 100644
--- a/arch/alpha/kernel/osf_sys.c
+++ b/arch/alpha/kernel/osf_sys.c
@@ -38,6 +38,7 @@
#include <linux/namei.h>
#include <linux/mount.h>
#include <linux/uio.h>
+#include <linux/hugetlb.h>
#include <linux/vfs.h>
#include <linux/rcupdate.h>
#include <linux/slab.h>
@@ -1201,14 +1202,16 @@ SYSCALL_DEFINE1(old_adjtimex, struct timex32 __user *, txc_p)
/* Get an address range which is currently unmapped. */
static unsigned long
-arch_get_unmapped_area_1(unsigned long addr, unsigned long len,
- unsigned long limit)
+arch_get_unmapped_area_1(struct file *filp, unsigned long addr,
+ unsigned long len, unsigned long limit)
{
struct vm_unmapped_area_info info = {};
info.length = len;
info.low_limit = addr;
info.high_limit = limit;
+ if (filp && is_file_hugepages(filp))
+ info.align_mask = huge_page_mask_align(filp);
return vm_unmapped_area(&info);
}
@@ -1236,19 +1239,20 @@ arch_get_unmapped_area(struct file *filp, unsigned long addr,
this feature should be incorporated into all ports? */
if (addr) {
- addr = arch_get_unmapped_area_1 (PAGE_ALIGN(addr), len, limit);
+ addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(addr), len,
+ limit);
if (addr != (unsigned long) -ENOMEM)
return addr;
}
/* Next, try allocating at TASK_UNMAPPED_BASE. */
- addr = arch_get_unmapped_area_1 (PAGE_ALIGN(TASK_UNMAPPED_BASE),
+ addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(TASK_UNMAPPED_BASE),
len, limit);
if (addr != (unsigned long) -ENOMEM)
return addr;
/* Finally, try allocating in low memory. */
- addr = arch_get_unmapped_area_1 (PAGE_SIZE, len, limit);
+ addr = arch_get_unmapped_area_1 (filp, PAGE_SIZE, len, limit);
return addr;
}
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH v2 6/6] alpha: implement hugetlb support
2026-10-08 3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (4 preceding siblings ...)
2026-10-08 3:24 ` [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
@ 2026-10-08 3:24 ` Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-08 3:24 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
Andrew Morton, Peter Xu
Cc: linux-alpha, linux-kernel, Matt Turner
Alpha has no leaf entry above the last page table level, so the only
huge pages it can offer are granularity hint blocks: 64KB, 512KB and 4MB
with 8KB pages. Register all three as hstates, with 4MB as the default.
There is no PMD sized huge page, hence no transparent huge pages, no PMD
page table sharing and no gigantic pages.
As with arm64's contiguous PTEs, each PTE of a block keeps the frame
number of its own page, so a block is written with set_ptes(). Each
hstate's buddy order is its hint order, so a hugetlb folio is always
naturally aligned physically, as the hint requires, including after a
demote.
All PTEs of a block must agree in bits <15:0> (note 2 of Table 22-3 of
the Alpha Architecture Reference Manual), and both __ACCESS_BITS and
__DIRTY_BITS reach into that range, so every update rewrites the whole
block and huge_ptep_get() merges the young and dirty state back
together. Rewriting a valid block in place would leave its PTEs
disagreeing part way through, so changes to a valid block go through
break before make, following the procedure in section 11.6.1.
The hint is advisory, so this is safe on implementations that ignore it,
and gup_fast needs no changes since it walks the individual PTEs.
Enabling CONFIG_HUGETLB_PAGE derives pageblock_order from
HUGETLB_PAGE_ORDER, which lowers it from 10 to 9 and so halves the
anti-fragmentation and compaction granularity from 8MB to 4MB.
Hugepage migration is left disabled, since it has not been tested.
Co-developed-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/Kconfig | 1 +
arch/alpha/include/asm/hugetlb.h | 43 ++++++
arch/alpha/include/asm/page.h | 13 ++
arch/alpha/mm/Makefile | 2 +
arch/alpha/mm/hugetlbpage.c | 308 +++++++++++++++++++++++++++++++++++++++
5 files changed, 367 insertions(+)
diff --git a/arch/alpha/Kconfig b/arch/alpha/Kconfig
index b5b02bb42aca..b2c359b947f6 100644
--- a/arch/alpha/Kconfig
+++ b/arch/alpha/Kconfig
@@ -19,6 +19,7 @@ config ALPHA
select ARCH_NO_PREEMPT
select ARCH_NO_SG_CHAIN
select ARCH_SUPPORTS_ATOMIC_RMW
+ select ARCH_SUPPORTS_HUGETLBFS
select ARCH_SUPPORTS_INT128 if CC_HAS_INT128
select ARCH_SUPPORTS_PAGE_TABLE_CHECK
select ARCH_USE_CMPXCHG_LOCKREF
diff --git a/arch/alpha/include/asm/hugetlb.h b/arch/alpha/include/asm/hugetlb.h
new file mode 100644
index 000000000000..ff69e24c775d
--- /dev/null
+++ b/arch/alpha/include/asm/hugetlb.h
@@ -0,0 +1,43 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _ALPHA_HUGETLB_H
+#define _ALPHA_HUGETLB_H
+
+#include <asm/page.h>
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags);
+#define arch_make_huge_pte arch_make_huge_pte
+
+#define __HAVE_ARCH_HUGE_PTEP_GET
+extern pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep);
+#define __HAVE_ARCH_HUGE_SET_HUGE_PTE_AT
+extern void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, pte_t pte, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_GET_AND_CLEAR
+extern pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_CLEAR_FLUSH
+extern pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_WRPROTECT
+extern void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_ACCESS_FLAGS
+extern int huge_ptep_set_access_flags(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep,
+ pte_t pte, int dirty);
+#define __HAVE_ARCH_HUGE_PTE_CLEAR
+extern void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz);
+
+/*
+ * The generic version does not flush, which would leave the old translation
+ * live across a restrictive change to PTE bits <15:0>.
+ */
+#define huge_ptep_modify_prot_start huge_ptep_modify_prot_start
+extern pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep);
+
+#include <asm-generic/hugetlb.h>
+
+#endif /* _ALPHA_HUGETLB_H */
diff --git a/arch/alpha/include/asm/page.h b/arch/alpha/include/asm/page.h
index 59d01f9b77f6..67053c52953f 100644
--- a/arch/alpha/include/asm/page.h
+++ b/arch/alpha/include/asm/page.h
@@ -6,6 +6,19 @@
#include <asm/pal.h>
#include <vdso/page.h>
+#ifdef CONFIG_HUGETLB_PAGE
+/*
+ * The default huge page is the largest granularity hint block, 4MB. All
+ * three hint sizes are registered as hstates, hence HUGE_MAX_HSTATE, which
+ * linux/hugetlb.h needs before it includes asm/hugetlb.h.
+ */
+#define HPAGE_SHIFT 22
+#define HPAGE_SIZE (_AC(1, UL) << HPAGE_SHIFT)
+#define HPAGE_MASK (~(HPAGE_SIZE - 1))
+#define HUGETLB_PAGE_ORDER (HPAGE_SHIFT - PAGE_SHIFT)
+#define HUGE_MAX_HSTATE 3
+#endif
+
#ifndef __ASSEMBLER__
#define STRICT_MM_TYPECHECKS
diff --git a/arch/alpha/mm/Makefile b/arch/alpha/mm/Makefile
index 2d05664058f6..c022e55f231a 100644
--- a/arch/alpha/mm/Makefile
+++ b/arch/alpha/mm/Makefile
@@ -4,3 +4,5 @@
#
obj-y := init.o fault.o tlbflush.o
+
+obj-$(CONFIG_HUGETLB_PAGE) += hugetlbpage.o
diff --git a/arch/alpha/mm/hugetlbpage.c b/arch/alpha/mm/hugetlbpage.c
new file mode 100644
index 000000000000..ca36af4667dc
--- /dev/null
+++ b/arch/alpha/mm/hugetlbpage.c
@@ -0,0 +1,308 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Alpha huge page support, built on the page table entry granularity hint
+ * described in asm/pgtable.h. All PTEs of a block have to agree in bits
+ * <15:0>, so every update below rewrites the whole block.
+ */
+
+#include <linux/hugetlb.h>
+#include <linux/mm.h>
+#include <linux/pgtable.h>
+
+#include <asm/tlbflush.h>
+
+/*
+ * Every hint size is smaller than PMD_SIZE, so a block always fits inside a
+ * single last level page table and its entry count is just its page count.
+ */
+static inline unsigned long num_contig_ptes(unsigned long sz)
+{
+ return sz >> PAGE_SHIFT;
+}
+
+/* Hint order that maps @sz, or 0 if @sz is not a hint size. */
+static unsigned int gh_order_from_size(unsigned long sz)
+{
+ unsigned int order;
+
+ for_each_gh_order(order)
+ if (gh_cont_size(order) == sz)
+ return order;
+
+ return 0;
+}
+
+pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+ unsigned long addr, unsigned long sz)
+{
+ unsigned int order = gh_order_from_size(sz);
+ pgd_t *pgd;
+ p4d_t *p4d;
+ pud_t *pud;
+ pmd_t *pmd;
+
+ if (!order)
+ return NULL;
+
+ pgd = pgd_offset(mm, addr);
+ p4d = p4d_alloc(mm, pgd, addr);
+ if (!p4d)
+ return NULL;
+
+ pud = pud_alloc(mm, p4d, addr);
+ if (!pud)
+ return NULL;
+
+ pmd = pmd_alloc(mm, pud, addr);
+ if (!pmd)
+ return NULL;
+
+ return pte_alloc_huge(mm, pmd, addr & gh_cont_mask(order));
+}
+
+pte_t *huge_pte_offset(struct mm_struct *mm, unsigned long addr,
+ unsigned long sz)
+{
+ unsigned int order = gh_order_from_size(sz);
+ pgd_t *pgd;
+ p4d_t *p4d;
+ pud_t *pud;
+ pmd_t *pmd;
+
+ if (!order)
+ return NULL;
+
+ pgd = pgd_offset(mm, addr);
+ if (!pgd_present(pgdp_get(pgd)))
+ return NULL;
+
+ p4d = p4d_offset(pgd, addr);
+ if (!p4d_present(p4dp_get(p4d)))
+ return NULL;
+
+ pud = pud_offset(p4d, addr);
+ if (!pud_present(pudp_get(pud)))
+ return NULL;
+
+ pmd = pmd_offset(pud, addr);
+ if (!pmd_present(pmdp_get(pmd)))
+ return NULL;
+
+ return pte_offset_huge(pmd, addr & gh_cont_mask(order));
+}
+
+/*
+ * A block is only as large as its own hint, so the walk can skip to the end of
+ * the containing last level page table but no further.
+ */
+unsigned long hugetlb_mask_last_page(struct hstate *h)
+{
+ return PMD_SIZE - huge_page_size(h);
+}
+
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+{
+ pte_t orig_pte = ptep_get(ptep);
+ unsigned long i, ncontig;
+
+ /* Swap, migration and marker entries carry no hint. */
+ if (!pte_huge(orig_pte))
+ return orig_pte;
+
+ ncontig = gh_pte_num(pte_gh_order(orig_pte));
+
+ for (i = 0; i < ncontig; i++, ptep++) {
+ pte_t pte = ptep_get(ptep);
+
+ if (pte_dirty(pte))
+ orig_pte = pte_mkdirty(orig_pte);
+ if (pte_young(pte))
+ orig_pte = pte_mkyoung(orig_pte);
+ }
+
+ return orig_pte;
+}
+
+static pte_t get_clear_contig(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long ncontig)
+{
+ pte_t pte = ptep_get_and_clear(mm, addr, ptep);
+ bool present = pte_present(pte);
+
+ while (--ncontig) {
+ pte_t tmp_pte;
+
+ ptep++;
+ addr += PAGE_SIZE;
+ tmp_pte = ptep_get_and_clear(mm, addr, ptep);
+ if (present) {
+ if (pte_dirty(tmp_pte))
+ pte = pte_mkdirty(pte);
+ if (pte_young(tmp_pte))
+ pte = pte_mkyoung(pte);
+ }
+ }
+
+ return pte;
+}
+
+static pte_t get_clear_contig_flush(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long ncontig)
+{
+ pte_t orig_pte = get_clear_contig(mm, addr, ptep, ncontig);
+ struct vm_area_struct vma = TLB_FLUSH_VMA(mm, 0);
+
+ if (!pte_none(orig_pte))
+ flush_tlb_range(&vma, addr, addr + ncontig * PAGE_SIZE);
+
+ return orig_pte;
+}
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags)
+{
+ unsigned int order = gh_order_from_size(1UL << shift);
+
+ /*
+ * Not every caller asks for a size Alpha can hint; mm/debug_vm_pgtable.c
+ * probes with PMD_SHIFT. Leave the entry alone rather than complaining.
+ */
+ if (order)
+ entry = pte_mkgh(entry, order);
+
+ return entry;
+}
+
+void set_huge_pte_at(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+ pte_t pte, unsigned long sz)
+{
+ unsigned long i, ncontig = num_contig_ptes(sz);
+
+ /* The hint requires the block to be naturally aligned in the VA too. */
+ VM_WARN_ON_ONCE(addr & (sz - 1));
+
+ if (!pte_present(pte)) {
+ /*
+ * Swap, migration and marker entries hold no frame number, so
+ * every slot gets the same value.
+ */
+ for (i = 0; i < ncontig; i++, ptep++, addr += PAGE_SIZE)
+ set_ptes(mm, addr, ptep, pte, 1);
+ return;
+ }
+
+ /*
+ * All PTEs of a block must agree in bits <15:0> (ARM Table 22-3, note
+ * 2), which rewriting a valid block in place would violate part way
+ * through. Invalidate it everywhere first, following ARM 11.6.1: the
+ * invalidate has to assume the hint is zero, so it must cover every
+ * page of the block.
+ */
+ if (pte_present(ptep_get(ptep)))
+ get_clear_contig_flush(mm, addr, ptep, ncontig);
+
+ /* set_ptes() advances the frame number for each entry. */
+ set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz)
+{
+ return get_clear_contig(mm, addr, ptep, num_contig_ptes(sz));
+}
+
+/*
+ * The block size comes from the hstate rather than from the entry, because
+ * this is also reached for non-present entries, which carry no hint.
+ */
+pte_t huge_ptep_clear_flush(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep)
+{
+ unsigned long sz = huge_page_size(hstate_vma(vma));
+
+ return get_clear_contig_flush(vma->vm_mm, addr, ptep,
+ num_contig_ptes(sz));
+}
+
+pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep)
+{
+ return huge_ptep_clear_flush(vma, addr, ptep);
+}
+
+/* True if any entry of the block differs from @pte outside the frame number. */
+static bool gh_access_flags_changed(pte_t *ptep, pte_t pte,
+ unsigned long ncontig)
+{
+ unsigned long i;
+
+ for (i = 0; i < ncontig; i++)
+ if ((pte_val(ptep_get(ptep + i)) ^ pte_val(pte)) & ~_PFN_MASK)
+ return true;
+
+ return false;
+}
+
+int huge_ptep_set_access_flags(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep, pte_t pte, int dirty)
+{
+ struct mm_struct *mm = vma->vm_mm;
+ unsigned long sz = huge_page_size(hstate_vma(vma));
+ unsigned long ncontig = num_contig_ptes(sz);
+ pte_t orig_pte = huge_ptep_get(mm, addr, ptep);
+
+ /* Keep the young and dirty state the block already has. */
+ if (pte_dirty(orig_pte))
+ pte = pte_mkdirty(pte);
+ if (pte_young(orig_pte))
+ pte = pte_mkyoung(pte);
+
+ /* Breaking an unchanged block only makes concurrent touchers refault. */
+ if (!gh_access_flags_changed(ptep, pte, ncontig))
+ return 0;
+
+ get_clear_contig_flush(mm, addr, ptep, ncontig);
+ set_ptes(mm, addr, ptep, pte, ncontig);
+
+ return 1;
+}
+
+void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep)
+{
+ pte_t pte = ptep_get(ptep);
+ unsigned long ncontig = gh_pte_num(pte_gh_order(pte));
+
+ VM_WARN_ON_ONCE(!pte_huge(pte));
+
+ /*
+ * pte_wrprotect() sets _PAGE_FOW, which lives in the bits the whole
+ * block must agree on, so this takes the break before make path.
+ */
+ pte = pte_wrprotect(get_clear_contig_flush(mm, addr, ptep, ncontig));
+ set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+void huge_pte_clear(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+ unsigned long sz)
+{
+ unsigned long i, ncontig = num_contig_ptes(sz);
+
+ for (i = 0; i < ncontig; i++, addr += PAGE_SIZE, ptep++)
+ pte_clear(mm, addr, ptep);
+}
+
+bool __init arch_hugetlb_valid_size(unsigned long size)
+{
+ return gh_order_from_size(size) != 0;
+}
+
+static __init int alpha_hugetlbpage_init(void)
+{
+ unsigned int order;
+
+ for_each_gh_order(order)
+ hugetlb_add_hstate(gh_cont_shift(order) - PAGE_SHIFT);
+
+ return 0;
+}
+arch_initcall(alpha_hugetlbpage_init);
--
2.55.0
^ permalink raw reply [flat|nested] 7+ messages in thread