* [PATCH 0/6] alpha: hugetlb support using granularity hints
@ 2026-10-06 14:04 Matt Turner
2026-10-06 14:04 ` [PATCH 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
` (5 more replies)
0 siblings, 6 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
Alpha has no leaf entry above the last page table level, so the usual
PMD sized huge page is not available to it. What it does have is the
granularity hint, bits <6:5> of the PTE, described in Table 17-3 of the
Alpha Architecture Reference Manual: a hint of order N marks a PTE as
one of 8^N physically contiguous, naturally aligned pages that the
translation buffer is permitted to map with a single entry. With 8KB
pages that gives 64KB, 512KB and 4MB blocks, and this series registers
all three as hstates.
This works like arm64's contiguous PTE support rather than RISC-V's
NAPOT: every PTE of a block keeps the frame number of its own page, so a
block is written with set_ptes() and the frame number advances across
it. There is no PMD sized huge page, and hence no transparent huge
pages, no PMD page table sharing and no gigantic pages.
The architecture requires all PTEs of a block to agree in bits <15:0>,
and both __ACCESS_BITS and __DIRTY_BITS reach into that range, so every
update rewrites the whole block and huge_ptep_get() merges the young and
dirty state back together. Changes to a valid block go through break
before make, as section 11.6.1 requires; the invalidate there has to
assume the hint is zero and so must cover every page of the block, which
flush_tlb_range() on alpha already exceeds, since it rolls the address
space number.
The hint is advisory. An implementation that ignores it still translates
correctly through the individual PTEs, so this is safe on every Alpha,
and gup_fast needs no changes. The console block that would report which
hint sizes the translation buffer implements was never filled in by any
console through EV7, so which sizes an implementation honors can only be
established by measuring, as the last section does for EV7.
Patches 1 and 2 are independent fixes. Patch 1 adds a
page_table_check_pte_clear() call missing from alpha's
ptep_get_and_clear(). Patch 2 fixes a BUG() that testing this series
turned up: do_page_fault() fell through to BUG() on VM_FAULT_HWPOISON,
reachable through UFFDIO_POISON without any memory failure support.
Patch 3 documents that the bits Linux calls _PAGE_URE and _PAGE_UWE are
the architecture's ERE and EWE, which matters when reading PTE layout
tables alongside this code. Patches 4 and 5 are groundwork and patch 6
is the implementation.
Testing
=======
Tested on megalith (EV7 Marvel, 8GB), with two CPUs and again with one,
cross-built with alpha-unknown-linux-gnu-gcc, with CONFIG_DEBUG_VM=y,
CONFIG_DEBUG_VM_PGTABLE=y, CONFIG_PAGE_TABLE_CHECK_ENFORCED=y and
CONFIG_CGROUP_HUGETLB=y. All three hint sizes register as hstates.
mm selftests, the same in both configurations: hugetlb 10/10,
userfaultfd and cow 7 pass 1 skip (uffd-wp-mremap), 0 fail, including
uffd-stress hugetlb and hugetlb-private at 128MB/32 threads.
debug_vm_pgtable validates clean at boot. No page_table_check reports
and no DEBUG_VM splats in any run.
Ad hoc tests written for this series also pass at all three sizes: dense
per-base-page write and read back across a block, which catches a wrong
frame number inside a block that a strided pattern would alias over;
mprotect down to PROT_READ and back, including a middle-block-only case,
exercising break before make; fork COW; hugetlbfs shared mappings
checked through pread and through a second independent mapping; hole
punch of a middle block with both neighbors and the refaulted hole
checked; ftruncate down and back up; and the 4MB -> 512KB -> 64KB demote
chain with exact count checks. All of it again with four concurrent
copies, and the whole suite ten times in a row.
An earlier revision of the series was also tested on up1500 (EV68AL
Nautilus, UP, 4GB) with CONFIG_DEBUG_VM=y and
CONFIG_PAGE_TABLE_CHECK_ENFORCED=y. That machine has not been retested
with this revision.
Magnus Lindholm also tested the series on SMP, on a UP2000+ (2x EV68AL
833 MHz), including a multithreaded stress test. It found that
huge_ptep_set_access_flags() broke and rewrote a block on every fault,
so two threads on different CPUs could keep each other faulting: 17,316
faults in 6 seconds after one permission change on a 4MB page. With the
check now in patch 6 the same change causes one fault. No writes were
lost with or without it.
Hugepage migration is not enabled. It has no test coverage: the
migration, rmap and ksm selftests do not cross-build for want of libnuma
in the alpha sysroot, and neither machine is multi-node.
Does the hint do anything?
==========================
On EV7, yes. Since the hint is advisory there is no way to ask the
hardware whether it implements one, and the console block that would
report it was never filled in, so the only way to find out is to
measure.
Comparing the three hint sizes against each other rather than against
normal pages avoids the hugetlb-versus-anonymous confound. A random
pointer chase touches one cache line per 8KB base page, with the
permutation seeded only from the page count, so every backing walks an
identical sequence over an identical footprint and the only variable is
how many DTB entries the working set needs. EV6 and EV7 have a 128 entry
fully associative DTB, so coverage is 128 times the page size: 1MB with
base pages, 8MB at 64KB, 64MB at 512KB, 512MB at 4MB.
Nanoseconds per access on megalith with one CPU, 8,000,000 accesses per
pass. Each figure is the best of ten runs, every run on freshly
allocated memory: single runs differ by 40% or more from one allocation
to the next, at every page size including base pages, so the best run is
the one to compare.
size base(8KB) 64KB 512KB 4096KB
1 MB 11.24 11.11 15.02 11.72
4 MB 155.47 104.54 105.03 104.17
16 MB 156.69 127.50 109.55 105.52
64 MB 154.61 146.79 110.16 105.76
256 MB 172.56 171.34 175.74 107.47
1024 MB 232.40 230.08 222.15 190.03
Each size keeps its advantage until the working set passes its own
coverage and then converges on the base page column: 64KB is useful to
about 8MB, 512KB to about 64MB, 4MB to about 512MB.
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
Matt Turner (6):
alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
alpha: handle VM_FAULT_HWPOISON in do_page_fault()
alpha: clarify that _PAGE_URE and _PAGE_UWE are the Executive bits
alpha: define granularity hint PTE bits
alpha: align hugetlb mappings in arch_get_unmapped_area()
alpha: implement hugetlb support
arch/alpha/Kconfig | 1 +
arch/alpha/include/asm/hugetlb.h | 43 ++++++
arch/alpha/include/asm/page.h | 13 ++
arch/alpha/include/asm/pgtable.h | 69 ++++++++-
arch/alpha/kernel/osf_sys.c | 14 +-
arch/alpha/mm/Makefile | 2 +
arch/alpha/mm/fault.c | 19 +++
arch/alpha/mm/hugetlbpage.c | 306 +++++++++++++++++++++++++++++++++++++++
8 files changed, 456 insertions(+), 11 deletions(-)
---
base-commit: e946efcc89066c5d80acbae42d015a4da33a11de
change-id: 20260912-alpha-hugepages-0c6cb6c0995d
Best regards,
--
Matt Turner <mattst88@gmail.com>
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
2026-10-06 14:04 ` [PATCH 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
Alpha's CONFIG_COMPACTION-gated ptep_get_and_clear() overrides the
generic version in include/linux/pgtable.h but omits its
page_table_check_pte_clear() call. With CONFIG_PAGE_TABLE_CHECK=y the
map count taken by set_ptes() is therefore never dropped when a PTE is
cleared through this path.
Add the missing call, and drop the now-duplicate one in
ptep_clear_flush(), which already goes through ptep_get_and_clear().
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 6ebdd90f3035..8a175e0c2b42 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -305,6 +305,7 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
pte_t pte = READ_ONCE(*ptep);
pte_clear(mm, address, ptep);
+ page_table_check_pte_clear(mm, address, pte);
return pte;
}
@@ -316,7 +317,6 @@ static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
struct mm_struct *mm = vma->vm_mm;
pte_t pte = ptep_get_and_clear(mm, addr, ptep);
- page_table_check_pte_clear(mm, addr, pte);
migrate_flush_tlb_page(vma, addr);
return pte;
}
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault()
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-06 14:04 ` [PATCH 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
2026-10-06 14:04 ` [PATCH 3/6] alpha: clarify that _PAGE_URE and _PAGE_UWE are the Executive bits Matt Turner
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
do_page_fault() handles VM_FAULT_OOM, VM_FAULT_SIGSEGV and
VM_FAULT_SIGBUS and then falls into BUG() for anything else in
VM_FAULT_ERROR. VM_FAULT_HWPOISON and VM_FAULT_HWPOISON_LARGE are in
that set, so a fault on a poisoned page takes down the kernel instead of
delivering a signal.
Alpha does not select ARCH_SUPPORTS_MEMORY_FAILURE, so the machine check
paths cannot produce these. UFFDIO_POISON can: it installs a poison PTE
marker without any memory failure support, and a subsequent access
returns VM_FAULT_HWPOISON. Registering an anonymous range with
userfaultfd, poisoning it and reading it back reliably hits the BUG():
Kernel bug at arch/alpha/mm/fault.c:188
poison(50809): Kernel Bug 1
pc is at do_page_fault+0x4e8/0x550
ra is at do_page_fault+0xfc/0x550
The BUG() fires with mmap_read_lock() still held, so the faulting task is left
unkillable in D state holding the lock, and shutdown stalls behind it.
Deliver SIGBUS with BUS_MCEERR_AR instead, reporting the size of the
poisoned area, as the other architectures do. That is the huge page size
for VM_FAULT_HWPOISON_LARGE, which becomes reachable once alpha
implements huge pages.
Tested on an UP1500 (EV68AL): the reproducer above now takes a SIGBUS
and the kernel logs
poison[373]: hardware memory error at 0000020000030000 pc 00000200010007ec
with no oops, no wedged task and no taint.
Fixes: fc71884a5f59 ("mm: userfaultfd: add new UFFDIO_POISON ioctl")
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/mm/fault.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..24408b53197c 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -8,6 +8,7 @@
#include <linux/sched/signal.h>
#include <linux/kernel.h>
#include <linux/mm.h>
+#include <linux/hugetlb.h>
#include <asm/io.h>
#define __EXTERN_INLINE inline
@@ -113,6 +114,7 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
struct mm_struct *mm = current->mm;
const struct exception_table_entry *fixup;
int si_code = SEGV_MAPERR;
+ unsigned int lsb;
vm_fault_t fault;
unsigned int flags = FAULT_FLAG_DEFAULT;
@@ -185,6 +187,8 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto bad_area;
else if (fault & VM_FAULT_SIGBUS)
goto do_sigbus;
+ else if (fault & (VM_FAULT_HWPOISON | VM_FAULT_HWPOISON_LARGE))
+ goto do_sigbus_mceerr;
BUG();
}
@@ -248,6 +252,21 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto no_context;
return;
+ do_sigbus_mceerr:
+ mmap_read_unlock(mm);
+ if (!user_mode(regs))
+ goto no_context;
+ /*
+ * Report the size of the poisoned area, which for a hugetlb fault
+ * is the size of the huge page that could not be mapped.
+ */
+ lsb = PAGE_SHIFT;
+ if (fault & VM_FAULT_HWPOISON_LARGE)
+ lsb = hstate_index_to_shift(VM_FAULT_GET_HINDEX(fault));
+ show_signal_msg(regs, address, SIGBUS, "hardware memory error");
+ force_sig_mceerr(BUS_MCEERR_AR, (void __user *) address, lsb);
+ return;
+
do_sigsegv:
show_signal_msg(regs, address, SIGSEGV,
si_code == SEGV_MAPERR ? "unmapped access"
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 3/6] alpha: clarify that _PAGE_URE and _PAGE_UWE are the Executive bits
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-06 14:04 ` [PATCH 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
2026-10-06 14:04 ` [PATCH 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
2026-10-06 14:04 ` [PATCH 4/6] alpha: define granularity hint PTE bits Matt Turner
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
The Alpha Architecture Reference Manual defines eight protection bits in
PTE<15:8>: KRE, ERE, SRE, URE, KWE, EWE, SWE and UWE, for Kernel,
Executive, Supervisor and User Read and Write Enable. Linux uses only
two of the four privilege modes, and its user processes run in Executive
mode, so the bits it calls _PAGE_URE and _PAGE_UWE are in fact the
architecture's ERE (bit 9) and EWE (bit 13). The true URE and UWE bits,
11 and 15, are unused.
The names have been misleading since the beginning, and the comments on
those defines said only "xxx". Spell out what the bits actually are. No
functional change.
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 16 ++++++++++++----
1 file changed, 12 insertions(+), 4 deletions(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 8a175e0c2b42..d123fd0d4091 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,10 +65,18 @@ struct vm_area_struct;
#define _PAGE_FOW 0x0004 /* used for page protection (fault on write) */
#define _PAGE_FOE 0x0008 /* used for page protection (fault on exec) */
#define _PAGE_ASM 0x0010
-#define _PAGE_KRE 0x0100 /* xxx - see below on the "accessed" bit */
-#define _PAGE_URE 0x0200 /* xxx */
-#define _PAGE_KWE 0x1000 /* used to do the dirty bit in software */
-#define _PAGE_UWE 0x2000 /* used to do the dirty bit in software */
+/*
+ * The Alpha AXP ARM defines 8 protection bits in PTE[15:8]: KRE, ERE, SRE,
+ * URE, KWE, EWE, SWE, UWE (Kernel/Executive/Supervisor/User Read/Write
+ * Enable). Linux only uses two privilege modes: kernel (mode 0) and user.
+ * User processes run in Executive mode (mode 1), so _PAGE_URE and _PAGE_UWE
+ * correspond to the architecture's ERE (bit 9) and EWE (bit 13) bits. The
+ * true User Read/Write Enable bits (bits 11 and 15) are unused.
+ */
+#define _PAGE_KRE 0x0100 /* Kernel Read Enable (bit 8) */
+#define _PAGE_URE 0x0200 /* Executive Read Enable (bit 9); "user" in Linux */
+#define _PAGE_KWE 0x1000 /* Kernel Write Enable (bit 12) */
+#define _PAGE_UWE 0x2000 /* Executive Write Enable (bit 13); "user" in Linux */
/* .. and these are ours ... */
#define _PAGE_DIRTY 0x20000
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 4/6] alpha: define granularity hint PTE bits
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (2 preceding siblings ...)
2026-10-06 14:04 ` [PATCH 3/6] alpha: clarify that _PAGE_URE and _PAGE_UWE are the Executive bits Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
2026-10-06 14:04 ` [PATCH 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
2026-10-06 14:04 ` [PATCH 6/6] alpha: implement hugetlb support Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
Bits <6:5> of the Alpha PTE are the granularity hint, described in Table
17-3 of the Alpha Architecture Reference Manual. A hint of N marks the
PTE as one of a block of 8^N physically contiguous, naturally aligned
pages that the translation buffer may map with a single entry. With 8KB
pages that gives 64KB, 512KB and 4MB blocks.
Define the field and the helpers to encode and decode it. Nothing sets a
non-zero hint yet.
Add the hint to _PAGE_CHG_MASK so that pte_modify() preserves it. Swap
PTEs leave bits <31:0> clear, so pte_huge() is false on swap, migration
and marker entries, and page table entries above the last level keep a
zero hint where pmd_bad() would reject anything else.
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/include/asm/pgtable.h | 51 +++++++++++++++++++++++++++++++++++++++-
1 file changed, 50 insertions(+), 1 deletion(-)
diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index d123fd0d4091..605b82f73a98 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,6 +65,8 @@ struct vm_area_struct;
#define _PAGE_FOW 0x0004 /* used for page protection (fault on write) */
#define _PAGE_FOE 0x0008 /* used for page protection (fault on exec) */
#define _PAGE_ASM 0x0010
+#define _PAGE_GH_MASK 0x0060 /* granularity hint, PTE<6:5> */
+#define _PAGE_GH_SHIFT 5
/*
* The Alpha AXP ARM defines 8 protection bits in PTE[15:8]: KRE, ERE, SRE,
* URE, KWE, EWE, SWE, UWE (Kernel/Executive/Supervisor/User Read/Write
@@ -103,7 +105,36 @@ struct vm_area_struct;
#define _PFN_MASK 0xFFFFFFFF00000000UL
#define _PAGE_TABLE (_PAGE_VALID | __DIRTY_BITS | __ACCESS_BITS)
-#define _PAGE_CHG_MASK (_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | _PAGE_SPECIAL)
+/*
+ * pte_modify() must preserve the granularity hint, or
+ * hugetlb_change_protection() would turn a block into single-page PTEs.
+ */
+#define _PAGE_CHG_MASK (_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | \
+ _PAGE_GH_MASK | _PAGE_SPECIAL)
+
+/*
+ * Granularity hints. A PTE with hint order N is one of a block of 8^N
+ * physically contiguous pages, naturally aligned both virtually and
+ * physically, that the TB may map with a single entry. With 8KB pages the
+ * four encodings give 8KB, 64KB, 512KB and 4MB.
+ *
+ * Every PTE of a block must still carry the PFN of its own page, and all of
+ * them must agree in bits <15:0>: the protection, fault, hint and valid bits.
+ * __ACCESS_BITS and __DIRTY_BITS reach into that range, so young and dirty
+ * updates have to be applied to the whole block too.
+ *
+ * The hint is advisory: an implementation that ignores it still translates
+ * correctly through the individual PTEs, so it is safe to set on any Alpha.
+ */
+#define GH_ORDER_MAX 4 /* hint orders 0..3 */
+
+#define gh_cont_shift(order) (PAGE_SHIFT + 3 * (order))
+#define gh_cont_size(order) (1UL << gh_cont_shift(order))
+#define gh_cont_mask(order) (~(gh_cont_size(order) - 1))
+#define gh_pte_num(order) (1UL << (3 * (order)))
+
+#define for_each_gh_order(order) \
+ for ((order) = 1; (order) < GH_ORDER_MAX; (order)++)
/*
* All the normal masks have the "page accessed" bits on, as any time they are used,
@@ -269,6 +300,24 @@ extern inline pte_t pte_mkyoung(pte_t pte) { pte_val(pte) |= __ACCESS_BITS; retu
extern inline int pte_special(pte_t pte) { return pte_val(pte) & _PAGE_SPECIAL; }
extern inline pte_t pte_mkspecial(pte_t pte) { pte_val(pte) |= _PAGE_SPECIAL; return pte; }
+extern inline unsigned int pte_gh_order(pte_t pte)
+{
+ return (pte_val(pte) & _PAGE_GH_MASK) >> _PAGE_GH_SHIFT;
+}
+
+extern inline pte_t pte_mkgh(pte_t pte, unsigned int order)
+{
+ pte_val(pte) &= ~_PAGE_GH_MASK;
+ pte_val(pte) |= (unsigned long)order << _PAGE_GH_SHIFT;
+ return pte;
+}
+
+/*
+ * Swap PTEs leave bits <31:0> clear, so this is false for every swap,
+ * migration and marker entry, not just for ordinary small pages.
+ */
+extern inline int pte_huge(pte_t pte) { return pte_val(pte) & _PAGE_GH_MASK; }
+
/*
* The smp_rmb() in the following functions are required to order the load of
* *dir (the pointer in the top level page table) with any subsequent load of
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area()
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (3 preceding siblings ...)
2026-10-06 14:04 ` [PATCH 4/6] alpha: define granularity hint PTE bits Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
2026-10-06 14:04 ` [PATCH 6/6] alpha: implement hugetlb support Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
hugetlb_get_unmapped_area() does not align the address itself; it
delegates to mm_get_unmapped_area_vmaflags(), and the huge page
alignment is applied by generic_get_unmapped_area() and
generic_get_unmapped_area_topdown() via info.align_mask.
Alpha defines HAVE_ARCH_UNMAPPED_AREA and supplies its own
arch_get_unmapped_area(), which never sets align_mask. An
mmap(MAP_HUGETLB) with no address hint would therefore return a merely
page-aligned address.
Pass the file down to arch_get_unmapped_area_1() and set align_mask for
hugetlbfs mappings, as x86, s390, sparc and loongarch already do. This
is a no-op until alpha gains hugetlb support.
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/kernel/osf_sys.c | 14 +++++++++-----
1 file changed, 9 insertions(+), 5 deletions(-)
diff --git a/arch/alpha/kernel/osf_sys.c b/arch/alpha/kernel/osf_sys.c
index 7b6543d2cca3..61d806176213 100644
--- a/arch/alpha/kernel/osf_sys.c
+++ b/arch/alpha/kernel/osf_sys.c
@@ -38,6 +38,7 @@
#include <linux/namei.h>
#include <linux/mount.h>
#include <linux/uio.h>
+#include <linux/hugetlb.h>
#include <linux/vfs.h>
#include <linux/rcupdate.h>
#include <linux/slab.h>
@@ -1201,14 +1202,16 @@ SYSCALL_DEFINE1(old_adjtimex, struct timex32 __user *, txc_p)
/* Get an address range which is currently unmapped. */
static unsigned long
-arch_get_unmapped_area_1(unsigned long addr, unsigned long len,
- unsigned long limit)
+arch_get_unmapped_area_1(struct file *filp, unsigned long addr,
+ unsigned long len, unsigned long limit)
{
struct vm_unmapped_area_info info = {};
info.length = len;
info.low_limit = addr;
info.high_limit = limit;
+ if (filp && is_file_hugepages(filp))
+ info.align_mask = huge_page_mask_align(filp);
return vm_unmapped_area(&info);
}
@@ -1236,19 +1239,20 @@ arch_get_unmapped_area(struct file *filp, unsigned long addr,
this feature should be incorporated into all ports? */
if (addr) {
- addr = arch_get_unmapped_area_1 (PAGE_ALIGN(addr), len, limit);
+ addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(addr), len,
+ limit);
if (addr != (unsigned long) -ENOMEM)
return addr;
}
/* Next, try allocating at TASK_UNMAPPED_BASE. */
- addr = arch_get_unmapped_area_1 (PAGE_ALIGN(TASK_UNMAPPED_BASE),
+ addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(TASK_UNMAPPED_BASE),
len, limit);
if (addr != (unsigned long) -ENOMEM)
return addr;
/* Finally, try allocating in low memory. */
- addr = arch_get_unmapped_area_1 (PAGE_SIZE, len, limit);
+ addr = arch_get_unmapped_area_1 (filp, PAGE_SIZE, len, limit);
return addr;
}
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 6/6] alpha: implement hugetlb support
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
` (4 preceding siblings ...)
2026-10-06 14:04 ` [PATCH 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
@ 2026-10-06 14:04 ` Matt Turner
5 siblings, 0 replies; 7+ messages in thread
From: Matt Turner @ 2026-10-06 14:04 UTC (permalink / raw)
To: Richard Henderson, Magnus Lindholm; +Cc: linux-alpha, linux-kernel, Matt Turner
Alpha has no leaf entry above the last page table level, so the only
huge pages it can offer are granularity hint blocks: 64KB, 512KB and 4MB
with 8KB pages. Register all three as hstates, with 4MB as the default.
There is no PMD sized huge page, hence no transparent huge pages, no PMD
page table sharing and no gigantic pages.
As with arm64's contiguous PTEs, each PTE of a block keeps the frame
number of its own page, so a block is written with set_ptes(). Each
hstate's buddy order is its hint order, so a hugetlb folio is always
naturally aligned physically, as the hint requires, including after a
demote.
All PTEs of a block must agree in bits <15:0>, and both __ACCESS_BITS
and __DIRTY_BITS reach into that range, so every update rewrites the
whole block and huge_ptep_get() merges the young and dirty state back
together. Changes to a valid block go through break before make, as
section 11.6.1 of the architecture manual requires.
huge_ptep_set_access_flags() skips the break when no entry would change:
hugetlb_fault() calls it on every fault on a present entry, and breaking
an unchanged block makes threads on other CPUs refault and break it
again.
The hint is advisory, so this is safe on implementations that ignore it,
and gup_fast needs no changes since it walks the individual PTEs.
Enabling CONFIG_HUGETLB_PAGE derives pageblock_order from
HUGETLB_PAGE_ORDER, which lowers it from 10 to 9 and so halves the
anti-fragmentation and compaction granularity from 8MB to 4MB.
Hugepage migration is left disabled, since it has not been tested.
Co-developed-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
arch/alpha/Kconfig | 1 +
arch/alpha/include/asm/hugetlb.h | 43 ++++++
arch/alpha/include/asm/page.h | 13 ++
arch/alpha/mm/Makefile | 2 +
arch/alpha/mm/hugetlbpage.c | 306 +++++++++++++++++++++++++++++++++++++++
5 files changed, 365 insertions(+)
diff --git a/arch/alpha/Kconfig b/arch/alpha/Kconfig
index b5b02bb42aca..b2c359b947f6 100644
--- a/arch/alpha/Kconfig
+++ b/arch/alpha/Kconfig
@@ -19,6 +19,7 @@ config ALPHA
select ARCH_NO_PREEMPT
select ARCH_NO_SG_CHAIN
select ARCH_SUPPORTS_ATOMIC_RMW
+ select ARCH_SUPPORTS_HUGETLBFS
select ARCH_SUPPORTS_INT128 if CC_HAS_INT128
select ARCH_SUPPORTS_PAGE_TABLE_CHECK
select ARCH_USE_CMPXCHG_LOCKREF
diff --git a/arch/alpha/include/asm/hugetlb.h b/arch/alpha/include/asm/hugetlb.h
new file mode 100644
index 000000000000..ff69e24c775d
--- /dev/null
+++ b/arch/alpha/include/asm/hugetlb.h
@@ -0,0 +1,43 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _ALPHA_HUGETLB_H
+#define _ALPHA_HUGETLB_H
+
+#include <asm/page.h>
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags);
+#define arch_make_huge_pte arch_make_huge_pte
+
+#define __HAVE_ARCH_HUGE_PTEP_GET
+extern pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep);
+#define __HAVE_ARCH_HUGE_SET_HUGE_PTE_AT
+extern void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, pte_t pte, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_GET_AND_CLEAR
+extern pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_CLEAR_FLUSH
+extern pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_WRPROTECT
+extern void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_ACCESS_FLAGS
+extern int huge_ptep_set_access_flags(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep,
+ pte_t pte, int dirty);
+#define __HAVE_ARCH_HUGE_PTE_CLEAR
+extern void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz);
+
+/*
+ * The generic version does not flush, which would leave the old translation
+ * live across a restrictive change to PTE bits <15:0>.
+ */
+#define huge_ptep_modify_prot_start huge_ptep_modify_prot_start
+extern pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma,
+ unsigned long addr, pte_t *ptep);
+
+#include <asm-generic/hugetlb.h>
+
+#endif /* _ALPHA_HUGETLB_H */
diff --git a/arch/alpha/include/asm/page.h b/arch/alpha/include/asm/page.h
index 59d01f9b77f6..67053c52953f 100644
--- a/arch/alpha/include/asm/page.h
+++ b/arch/alpha/include/asm/page.h
@@ -6,6 +6,19 @@
#include <asm/pal.h>
#include <vdso/page.h>
+#ifdef CONFIG_HUGETLB_PAGE
+/*
+ * The default huge page is the largest granularity hint block, 4MB. All
+ * three hint sizes are registered as hstates, hence HUGE_MAX_HSTATE, which
+ * linux/hugetlb.h needs before it includes asm/hugetlb.h.
+ */
+#define HPAGE_SHIFT 22
+#define HPAGE_SIZE (_AC(1, UL) << HPAGE_SHIFT)
+#define HPAGE_MASK (~(HPAGE_SIZE - 1))
+#define HUGETLB_PAGE_ORDER (HPAGE_SHIFT - PAGE_SHIFT)
+#define HUGE_MAX_HSTATE 3
+#endif
+
#ifndef __ASSEMBLER__
#define STRICT_MM_TYPECHECKS
diff --git a/arch/alpha/mm/Makefile b/arch/alpha/mm/Makefile
index 2d05664058f6..c022e55f231a 100644
--- a/arch/alpha/mm/Makefile
+++ b/arch/alpha/mm/Makefile
@@ -4,3 +4,5 @@
#
obj-y := init.o fault.o tlbflush.o
+
+obj-$(CONFIG_HUGETLB_PAGE) += hugetlbpage.o
diff --git a/arch/alpha/mm/hugetlbpage.c b/arch/alpha/mm/hugetlbpage.c
new file mode 100644
index 000000000000..94b02de4b370
--- /dev/null
+++ b/arch/alpha/mm/hugetlbpage.c
@@ -0,0 +1,306 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Alpha huge page support, built on the page table entry granularity hint
+ * described in asm/pgtable.h. All PTEs of a block have to agree in bits
+ * <15:0>, so every update below rewrites the whole block.
+ */
+
+#include <linux/hugetlb.h>
+#include <linux/mm.h>
+#include <linux/pgtable.h>
+
+#include <asm/tlbflush.h>
+
+/*
+ * Every hint size is smaller than PMD_SIZE, so a block always fits inside a
+ * single last level page table and its entry count is just its page count.
+ */
+static inline unsigned long num_contig_ptes(unsigned long sz)
+{
+ return sz >> PAGE_SHIFT;
+}
+
+/* Hint order that maps @sz, or 0 if @sz is not a hint size. */
+static unsigned int gh_order_from_size(unsigned long sz)
+{
+ unsigned int order;
+
+ for_each_gh_order(order)
+ if (gh_cont_size(order) == sz)
+ return order;
+
+ return 0;
+}
+
+pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+ unsigned long addr, unsigned long sz)
+{
+ unsigned int order = gh_order_from_size(sz);
+ pgd_t *pgd;
+ p4d_t *p4d;
+ pud_t *pud;
+ pmd_t *pmd;
+
+ if (!order)
+ return NULL;
+
+ pgd = pgd_offset(mm, addr);
+ p4d = p4d_alloc(mm, pgd, addr);
+ if (!p4d)
+ return NULL;
+
+ pud = pud_alloc(mm, p4d, addr);
+ if (!pud)
+ return NULL;
+
+ pmd = pmd_alloc(mm, pud, addr);
+ if (!pmd)
+ return NULL;
+
+ return pte_alloc_huge(mm, pmd, addr & gh_cont_mask(order));
+}
+
+pte_t *huge_pte_offset(struct mm_struct *mm, unsigned long addr,
+ unsigned long sz)
+{
+ unsigned int order = gh_order_from_size(sz);
+ pgd_t *pgd;
+ p4d_t *p4d;
+ pud_t *pud;
+ pmd_t *pmd;
+
+ if (!order)
+ return NULL;
+
+ pgd = pgd_offset(mm, addr);
+ if (!pgd_present(pgdp_get(pgd)))
+ return NULL;
+
+ p4d = p4d_offset(pgd, addr);
+ if (!p4d_present(p4dp_get(p4d)))
+ return NULL;
+
+ pud = pud_offset(p4d, addr);
+ if (!pud_present(pudp_get(pud)))
+ return NULL;
+
+ pmd = pmd_offset(pud, addr);
+ if (!pmd_present(pmdp_get(pmd)))
+ return NULL;
+
+ return pte_offset_huge(pmd, addr & gh_cont_mask(order));
+}
+
+/*
+ * A block is only as large as its own hint, so the walk can skip to the end of
+ * the containing last level page table but no further.
+ */
+unsigned long hugetlb_mask_last_page(struct hstate *h)
+{
+ return PMD_SIZE - huge_page_size(h);
+}
+
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+{
+ pte_t orig_pte = ptep_get(ptep);
+ unsigned long i, ncontig;
+
+ /* Swap, migration and marker entries carry no hint. */
+ if (!pte_huge(orig_pte))
+ return orig_pte;
+
+ ncontig = gh_pte_num(pte_gh_order(orig_pte));
+
+ for (i = 0; i < ncontig; i++, ptep++) {
+ pte_t pte = ptep_get(ptep);
+
+ if (pte_dirty(pte))
+ orig_pte = pte_mkdirty(orig_pte);
+ if (pte_young(pte))
+ orig_pte = pte_mkyoung(orig_pte);
+ }
+
+ return orig_pte;
+}
+
+static pte_t get_clear_contig(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long ncontig)
+{
+ pte_t pte = ptep_get_and_clear(mm, addr, ptep);
+ bool present = pte_present(pte);
+
+ while (--ncontig) {
+ pte_t tmp_pte;
+
+ ptep++;
+ addr += PAGE_SIZE;
+ tmp_pte = ptep_get_and_clear(mm, addr, ptep);
+ if (present) {
+ if (pte_dirty(tmp_pte))
+ pte = pte_mkdirty(pte);
+ if (pte_young(tmp_pte))
+ pte = pte_mkyoung(pte);
+ }
+ }
+
+ return pte;
+}
+
+static pte_t get_clear_contig_flush(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long ncontig)
+{
+ pte_t orig_pte = get_clear_contig(mm, addr, ptep, ncontig);
+ struct vm_area_struct vma = TLB_FLUSH_VMA(mm, 0);
+
+ if (!pte_none(orig_pte))
+ flush_tlb_range(&vma, addr, addr + ncontig * PAGE_SIZE);
+
+ return orig_pte;
+}
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags)
+{
+ unsigned int order = gh_order_from_size(1UL << shift);
+
+ /*
+ * Not every caller asks for a size Alpha can hint; mm/debug_vm_pgtable.c
+ * probes with PMD_SHIFT. Leave the entry alone rather than complaining.
+ */
+ if (order)
+ entry = pte_mkgh(entry, order);
+
+ return entry;
+}
+
+void set_huge_pte_at(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+ pte_t pte, unsigned long sz)
+{
+ unsigned long i, ncontig = num_contig_ptes(sz);
+
+ /* The hint requires the block to be naturally aligned in the VA too. */
+ VM_WARN_ON_ONCE(addr & (sz - 1));
+
+ if (!pte_present(pte)) {
+ /*
+ * Swap, migration and marker entries hold no frame number, so
+ * every slot gets the same value.
+ */
+ for (i = 0; i < ncontig; i++, ptep++, addr += PAGE_SIZE)
+ set_ptes(mm, addr, ptep, pte, 1);
+ return;
+ }
+
+ /*
+ * ARM 11.6.1 requires an already valid entry to be invalidated
+ * everywhere before it is rewritten, and the invalidate has to assume
+ * the hint is zero, so it must cover every page of the block.
+ */
+ if (pte_present(ptep_get(ptep)))
+ get_clear_contig_flush(mm, addr, ptep, ncontig);
+
+ /* set_ptes() advances the frame number for each entry. */
+ set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep, unsigned long sz)
+{
+ return get_clear_contig(mm, addr, ptep, num_contig_ptes(sz));
+}
+
+/*
+ * The block size comes from the hstate rather than from the entry, because
+ * this is also reached for non-present entries, which carry no hint.
+ */
+pte_t huge_ptep_clear_flush(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep)
+{
+ unsigned long sz = huge_page_size(hstate_vma(vma));
+
+ return get_clear_contig_flush(vma->vm_mm, addr, ptep,
+ num_contig_ptes(sz));
+}
+
+pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep)
+{
+ return huge_ptep_clear_flush(vma, addr, ptep);
+}
+
+/* True if any entry of the block differs from @pte outside the frame number. */
+static bool gh_access_flags_changed(pte_t *ptep, pte_t pte,
+ unsigned long ncontig)
+{
+ unsigned long i;
+
+ for (i = 0; i < ncontig; i++)
+ if ((pte_val(ptep_get(ptep + i)) ^ pte_val(pte)) & ~_PFN_MASK)
+ return true;
+
+ return false;
+}
+
+int huge_ptep_set_access_flags(struct vm_area_struct *vma, unsigned long addr,
+ pte_t *ptep, pte_t pte, int dirty)
+{
+ struct mm_struct *mm = vma->vm_mm;
+ unsigned long sz = huge_page_size(hstate_vma(vma));
+ unsigned long ncontig = num_contig_ptes(sz);
+ pte_t orig_pte = huge_ptep_get(mm, addr, ptep);
+
+ /* Keep the young and dirty state the block already has. */
+ if (pte_dirty(orig_pte))
+ pte = pte_mkdirty(pte);
+ if (pte_young(orig_pte))
+ pte = pte_mkyoung(pte);
+
+ /* Breaking an unchanged block only makes concurrent touchers refault. */
+ if (!gh_access_flags_changed(ptep, pte, ncontig))
+ return 0;
+
+ get_clear_contig_flush(mm, addr, ptep, ncontig);
+ set_ptes(mm, addr, ptep, pte, ncontig);
+
+ return 1;
+}
+
+void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+ pte_t *ptep)
+{
+ pte_t pte = ptep_get(ptep);
+ unsigned long ncontig = gh_pte_num(pte_gh_order(pte));
+
+ VM_WARN_ON_ONCE(!pte_huge(pte));
+
+ /*
+ * pte_wrprotect() sets _PAGE_FOW, which lives in the bits the whole
+ * block must agree on, so this takes the break before make path.
+ */
+ pte = pte_wrprotect(get_clear_contig_flush(mm, addr, ptep, ncontig));
+ set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+void huge_pte_clear(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+ unsigned long sz)
+{
+ unsigned long i, ncontig = num_contig_ptes(sz);
+
+ for (i = 0; i < ncontig; i++, addr += PAGE_SIZE, ptep++)
+ pte_clear(mm, addr, ptep);
+}
+
+bool __init arch_hugetlb_valid_size(unsigned long size)
+{
+ return gh_order_from_size(size) != 0;
+}
+
+static __init int alpha_hugetlbpage_init(void)
+{
+ unsigned int order;
+
+ for_each_gh_order(order)
+ hugetlb_add_hstate(gh_cont_shift(order) - PAGE_SHIFT);
+
+ return 0;
+}
+arch_initcall(alpha_hugetlbpage_init);
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-06 14:04 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 14:04 [PATCH 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-06 14:04 ` [PATCH 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
2026-10-06 14:04 ` [PATCH 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
2026-10-06 14:04 ` [PATCH 3/6] alpha: clarify that _PAGE_URE and _PAGE_UWE are the Executive bits Matt Turner
2026-10-06 14:04 ` [PATCH 4/6] alpha: define granularity hint PTE bits Matt Turner
2026-10-06 14:04 ` [PATCH 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
2026-10-06 14:04 ` [PATCH 6/6] alpha: implement hugetlb support Matt Turner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®