mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/6] alpha: hugetlb support using granularity hints
@ 2026-10-08  3:24 Matt Turner
  2026-10-08  3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
                   ` (5 more replies)
  0 siblings, 6 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner, stable

Alpha has no leaf entry above the last page table level, so the usual
PMD sized huge page is not available to it. What it does have is the
granularity hint, bits <6:5> of the PTE, described in Table 22-3 of the
Alpha Architecture Reference Manual: a hint of order N marks a PTE as
one of 8^N physically contiguous, naturally aligned pages that the
translation buffer is permitted to map with a single entry. With 8KB
pages that gives 64KB, 512KB and 4MB blocks, and this series registers
all three as hstates.

This works like arm64's contiguous PTE support rather than RISC-V's
NAPOT: every PTE of a block keeps the frame number of its own page, so a
block is written with set_ptes() and the frame number advances across
it. There is no PMD sized huge page, and hence no transparent huge
pages, no PMD page table sharing and no gigantic pages.

The architecture requires all PTEs of a block to agree in bits <15:0>
(note 2 of Table 22-3), and both __ACCESS_BITS and __DIRTY_BITS reach
into that range, so every update rewrites the whole block and
huge_ptep_get() merges the young and dirty state back together.
Rewriting a valid block in place would leave its PTEs disagreeing part
way through, so changes to a valid block go through break before make,
following the procedure in section 11.6.1. The invalidate there has to
assume the hint is zero and so must cover every page of the block, which
flush_tlb_range() on alpha already exceeds, since it rolls the address
space number.

On SMP, break before make is only as good as the flush: flush_tlb_mm()
has to reach every CPU that may hold a translation. Magnus Lindholm's
TLB shootdown series [1] fixes flushes that miss a CPU after fork, under
lazy TLB, or on the calling CPU, and this series depends on it there.

[1] https://lore.kernel.org/linux-alpha/20260923074903.862898-1-linmag7@gmail.com/

The hint is advisory. An implementation that ignores it still translates
correctly through the individual PTEs, so this is safe on every Alpha,
and gup_fast needs no changes. The console block that would report which
hint sizes the translation buffer implements was never filled in by any
console through EV7, so which sizes an implementation honors can only be
established by measuring, as the last section does for EV7.

Patches 1 and 2 are independent fixes. Patch 1 adds a
page_table_check_pte_clear() call missing from alpha's
ptep_get_and_clear(). Patch 2 fixes a BUG() that testing this series
turned up: do_page_fault() fell through to BUG() on VM_FAULT_HWPOISON,
reachable through UFFDIO_POISON without any memory failure support.
Patch 3 replaces the "xxx" comments on the PTE read and write enable
bits with their descriptions from Table 22-3. Patches 4 and 5 are
groundwork and patch 6 is the implementation.

Testing
=======

Tested on megalith (EV7 Marvel, 8GB), with two CPUs and again with one,
cross-built with alpha-unknown-linux-gnu-gcc, with CONFIG_DEBUG_VM=y,
CONFIG_DEBUG_VM_PGTABLE=y, CONFIG_PAGE_TABLE_CHECK_ENFORCED=y and
CONFIG_CGROUP_HUGETLB=y. All three hint sizes register as hstates.

mm selftests, the same in both configurations: hugetlb 10/10,
userfaultfd and cow 7 pass 1 skip (uffd-wp-mremap), 0 fail, including
uffd-stress hugetlb and hugetlb-private at 128MB/32 threads.
debug_vm_pgtable validates clean at boot. No page_table_check reports
and no DEBUG_VM splats in any run.

Ad hoc tests written for this series also pass at all three sizes: dense
per-base-page write and read back across a block, which catches a wrong
frame number inside a block that a strided pattern would alias over;
mprotect down to PROT_READ and back, including a middle-block-only case,
exercising break before make; fork COW; hugetlbfs shared mappings
checked through pread and through a second independent mapping; hole
punch of a middle block with both neighbors and the refaulted hole
checked; ftruncate down and back up; and the 4MB -> 512KB -> 64KB demote
chain with exact count checks. All of it again with four concurrent
copies, and the whole suite ten times in a row.

An earlier revision of the series was also tested on up1500 (EV68AL
Nautilus, UP, 4GB) with CONFIG_DEBUG_VM=y and
CONFIG_PAGE_TABLE_CHECK_ENFORCED=y. That machine has not been retested
with this revision.

Magnus Lindholm also tested the series on SMP, on a UP2000+ (2x EV68AL
833 MHz), including a multithreaded stress test. It found that
huge_ptep_set_access_flags() broke and rewrote a block on every fault,
so two threads on different CPUs could keep each other faulting: 17,316
faults in 6 seconds after one permission change on a 4MB page. With the
check now in patch 6 the same change causes one fault. No writes were
lost with or without it.

Magnus has since run v1 on the UP2000+ again, both as posted and on top
of [1]: his test script 200 times, the hugetlb, userfaultfd and cow
selftests (17 pass, 1 skip) and the stress test at all three sizes, with
no lost writes and at most one fault after a permission change on either
kernel. v2 changes only comments and changelogs; the code is unchanged.

Hugepage migration is not enabled. It has no test coverage: the
migration, rmap and ksm selftests do not cross-build for want of libnuma
in the alpha sysroot, and neither machine is multi-node.

Does the hint do anything?
==========================

On EV7, yes. Since the hint is advisory there is no way to ask the
hardware whether it implements one, and the console block that would
report it was never filled in, so the only way to find out is to
measure.

Comparing the three hint sizes against each other rather than against
normal pages avoids the hugetlb-versus-anonymous confound. A random
pointer chase touches one cache line per 8KB base page, with the
permutation seeded only from the page count, so every backing walks an
identical sequence over an identical footprint and the only variable is
how many DTB entries the working set needs. EV6 and EV7 have a 128 entry
fully associative DTB, so coverage is 128 times the page size: 1MB with
base pages, 8MB at 64KB, 64MB at 512KB, 512MB at 4MB.

Nanoseconds per access on megalith with one CPU, 8,000,000 accesses per
pass. Each figure is the best of ten runs, every run on freshly
allocated memory: single runs differ by 40% or more from one allocation
to the next, at every page size including base pages, so the best run is
the one to compare.

     size      base(8KB)       64KB      512KB     4096KB
      1 MB          11.24      11.11      15.02      11.72
      4 MB         155.47     104.54     105.03     104.17
     16 MB         156.69     127.50     109.55     105.52
     64 MB         154.61     146.79     110.16     105.76
    256 MB         172.56     171.34     175.74     107.47
   1024 MB         232.40     230.08     222.15     190.03

Each size keeps its advantage until the working set passes its own
coverage and then converges on the base page column: 64KB is useful to
about 8MB, 512KB to about 64MB, 4MB to about 512MB.

Magnus Lindholm reports that EV68AL honors all three hint sizes too. On
the UP2000+, retired instructions per access, which include the PALcode
DTB fill, drop to the no-miss level exactly where each size fits the 128
entry DTB, whose entries can each map 1, 8, 64 or 512 pages (21264/EV67
Hardware Reference Manual, section 2.1.6.4).

Signed-off-by: Matt Turner <mattst88@gmail.com>
---
Changes in v2:
- Collect Reviewed-by and Tested-by tags from Magnus Lindholm.
- Patch 1: say which BUG_ON() the missing call leads to, add Fixes.
- Patch 2: Cc stable, note that any local user can reach the BUG() and
  what a backport has to drop.
- Patch 3: describe the enable bits as the Alpha Linux PTE (Table 22-3)
  defines them rather than the OpenVMS one, move the 21264 Executive
  mode detail to the changelog, and retitle to match (Magnus).
- Patch 4: cite Table 22-3 rather than its Tru64 twin, Table 17-3.
- Patch 5: name the commits that made the alignment the architecture's
  job (Magnus).
- Patch 6: cite note 2 of Table 22-3 as the reason for break before
  make and section 11.6.1 as the procedure followed, in the changelog
  and in set_huge_pte_at() (Magnus).
- Cover letter: state the SMP dependency on the TLB shootdown series,
  add the UP2000+ results for v1 and the EV68AL hint measurement.
- Link to v1: https://lore.kernel.org/r/20261006-alpha-hugepages-v1-0-a673a18aaa70@gmail.com

---
Matt Turner (6):
      alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
      alpha: handle VM_FAULT_HWPOISON in do_page_fault()
      alpha: describe the PTE read and write enable bits
      alpha: define granularity hint PTE bits
      alpha: align hugetlb mappings in arch_get_unmapped_area()
      alpha: implement hugetlb support

 arch/alpha/Kconfig               |   1 +
 arch/alpha/include/asm/hugetlb.h |  43 ++++++
 arch/alpha/include/asm/page.h    |  13 ++
 arch/alpha/include/asm/pgtable.h |  61 +++++++-
 arch/alpha/kernel/osf_sys.c      |  14 +-
 arch/alpha/mm/Makefile           |   2 +
 arch/alpha/mm/fault.c            |  19 +++
 arch/alpha/mm/hugetlbpage.c      | 308 +++++++++++++++++++++++++++++++++++++++
 8 files changed, 450 insertions(+), 11 deletions(-)
---
base-commit: e946efcc89066c5d80acbae42d015a4da33a11de
change-id: 20260912-alpha-hugepages-0c6cb6c0995d

Best regards,
-- 
Matt Turner <mattst88@gmail.com>


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear()
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  2026-10-08  3:24 ` [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
                   ` (4 subsequent siblings)
  5 siblings, 0 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner

Alpha's CONFIG_COMPACTION-gated ptep_get_and_clear() overrides the
generic version in include/linux/pgtable.h but omits its
page_table_check_pte_clear() call. With CONFIG_PAGE_TABLE_CHECK=y the
map count taken by set_ptes() is therefore never dropped when a PTE is
cleared through this path. zap_pte_range() clears PTEs this way, so the
first page freed after it hits the BUG_ON() in
__page_table_check_zero().

The call was instead placed in ptep_clear_flush() by commit dd5712f3379c
("alpha: fix user-space corruption during memory compaction"), where it
was harmless until alpha enabled page table check.

Add the missing call, and drop the now-duplicate one in
ptep_clear_flush(), which already goes through ptep_get_and_clear().

Fixes: e761b6fe4085 ("alpha: add ARCH_SUPPORTS_PAGE_TABLE_CHECK support")
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/include/asm/pgtable.h | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 6ebdd90f3035..8a175e0c2b42 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -305,6 +305,7 @@ static inline pte_t ptep_get_and_clear(struct mm_struct *mm,
 	pte_t pte = READ_ONCE(*ptep);
 
 	pte_clear(mm, address, ptep);
+	page_table_check_pte_clear(mm, address, pte);
 	return pte;
 }
 
@@ -316,7 +317,6 @@ static inline pte_t ptep_clear_flush(struct vm_area_struct *vma,
 	struct mm_struct *mm = vma->vm_mm;
 	pte_t pte = ptep_get_and_clear(mm, addr, ptep);
 
-	page_table_check_pte_clear(mm, addr, pte);
 	migrate_flush_tlb_page(vma, addr);
 	return pte;
 }

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault()
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
  2026-10-08  3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  2026-10-08  3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner, stable

do_page_fault() handles VM_FAULT_OOM, VM_FAULT_SIGSEGV and
VM_FAULT_SIGBUS and then falls into BUG() for anything else in
VM_FAULT_ERROR. VM_FAULT_HWPOISON and VM_FAULT_HWPOISON_LARGE are in
that set, so a fault on a poisoned page takes down the kernel instead of
delivering a signal.

Alpha does not select ARCH_SUPPORTS_MEMORY_FAILURE, so the machine check
paths cannot produce these. UFFDIO_POISON can: it installs a poison PTE
marker without any memory failure support, and a subsequent access
returns VM_FAULT_HWPOISON. Registering an anonymous range with
userfaultfd, poisoning it and reading it back reliably hits the BUG():

  Kernel bug at arch/alpha/mm/fault.c:188
  poison(50809): Kernel Bug 1
  pc is at do_page_fault+0x4e8/0x550
  ra is at do_page_fault+0xfc/0x550

The BUG() fires with mmap_read_lock() still held, so the faulting task
is left unkillable in D state holding the lock, and shutdown stalls
behind it. A UFFD_USER_MODE_ONLY userfaultfd is always allowed, so any
local user can do this.

Deliver SIGBUS with BUS_MCEERR_AR instead, reporting the size of the
poisoned area, as the other architectures do. That is the huge page size
for VM_FAULT_HWPOISON_LARGE, which becomes reachable once alpha
implements huge pages.

Tested on an UP1500 (EV68AL): the reproducer above now takes a SIGBUS
and the kernel logs

  poison[373]: hardware memory error at 0000020000030000 pc 00000200010007ec

with no oops, no wedged task and no taint.

A backport needs the show_signal_msg() call dropped, since alpha only
has it since commit 9ca6af1aac40 ("alpha: select
SYSCTL_EXCEPTION_TRACE").

Fixes: fc71884a5f59 ("mm: userfaultfd: add new UFFDIO_POISON ioctl")
Cc: stable@vger.kernel.org # 6.6+
Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/mm/fault.c | 19 +++++++++++++++++++
 1 file changed, 19 insertions(+)

diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..24408b53197c 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -8,6 +8,7 @@
 #include <linux/sched/signal.h>
 #include <linux/kernel.h>
 #include <linux/mm.h>
+#include <linux/hugetlb.h>
 #include <asm/io.h>
 
 #define __EXTERN_INLINE inline
@@ -113,6 +114,7 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
 	struct mm_struct *mm = current->mm;
 	const struct exception_table_entry *fixup;
 	int si_code = SEGV_MAPERR;
+	unsigned int lsb;
 	vm_fault_t fault;
 	unsigned int flags = FAULT_FLAG_DEFAULT;
 
@@ -185,6 +187,8 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
 			goto bad_area;
 		else if (fault & VM_FAULT_SIGBUS)
 			goto do_sigbus;
+		else if (fault & (VM_FAULT_HWPOISON | VM_FAULT_HWPOISON_LARGE))
+			goto do_sigbus_mceerr;
 		BUG();
 	}
 
@@ -248,6 +252,21 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
 		goto no_context;
 	return;
 
+ do_sigbus_mceerr:
+	mmap_read_unlock(mm);
+	if (!user_mode(regs))
+		goto no_context;
+	/*
+	 * Report the size of the poisoned area, which for a hugetlb fault
+	 * is the size of the huge page that could not be mapped.
+	 */
+	lsb = PAGE_SHIFT;
+	if (fault & VM_FAULT_HWPOISON_LARGE)
+		lsb = hstate_index_to_shift(VM_FAULT_GET_HINDEX(fault));
+	show_signal_msg(regs, address, SIGBUS, "hardware memory error");
+	force_sig_mceerr(BUS_MCEERR_AR, (void __user *) address, lsb);
+	return;
+
  do_sigsegv:
 	show_signal_msg(regs, address, SIGSEGV,
 			si_code == SEGV_MAPERR ? "unmapped access"

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 3/6] alpha: describe the PTE read and write enable bits
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
  2026-10-08  3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
  2026-10-08  3:24 ` [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  2026-10-08 14:27   ` Magnus Lindholm
  2026-10-08  3:24 ` [PATCH v2 4/6] alpha: define granularity hint PTE bits Matt Turner
                   ` (2 subsequent siblings)
  5 siblings, 1 reply; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner

The comments on _PAGE_KRE and _PAGE_URE said only "xxx". Describe the
four enable bits as the Alpha Linux PTE defines them, in Table 22-3 of
the Alpha Architecture Reference Manual: kernel and user read enable in
bits 8 and 9, kernel and user write enable in bits 12 and 13, with bits
<11:10> and <15:14> reserved. There are only two processor modes, user
and kernel (section 22.5.1). Linux uses the read enables as its accessed
bit and the write enables as its dirty bit, so say which of
__ACCESS_BITS and __DIRTY_BITS each one belongs to.

The names differ from what the hardware calls those bits on the 21264.
Its PALcode loads the PTE unchanged into DTB_PTE (21264/EV67 Hardware
Reference Manual, section 6.9), where bits 9 and 13 are the Executive
read and write enables (Figure 5-27), Executive being mode 1 of the four
the processor implements (Table 5-5). That is a detail below the PALcode
interface, and it is also the layout of the OpenVMS PTE in Table 11-2,
which is not the one Linux uses.

No functional change.

Suggested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/include/asm/pgtable.h | 8 ++++----
 1 file changed, 4 insertions(+), 4 deletions(-)

diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index 8a175e0c2b42..d22e28b956fb 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,10 +65,10 @@ struct vm_area_struct;
 #define _PAGE_FOW	0x0004	/* used for page protection (fault on write) */
 #define _PAGE_FOE	0x0008	/* used for page protection (fault on exec) */
 #define _PAGE_ASM	0x0010
-#define _PAGE_KRE	0x0100	/* xxx - see below on the "accessed" bit */
-#define _PAGE_URE	0x0200	/* xxx */
-#define _PAGE_KWE	0x1000	/* used to do the dirty bit in software */
-#define _PAGE_UWE	0x2000	/* used to do the dirty bit in software */
+#define _PAGE_KRE	0x0100	/* kernel read enable, in __ACCESS_BITS */
+#define _PAGE_URE	0x0200	/* user read enable, in __ACCESS_BITS */
+#define _PAGE_KWE	0x1000	/* kernel write enable, in __DIRTY_BITS */
+#define _PAGE_UWE	0x2000	/* user write enable, in __DIRTY_BITS */
 
 /* .. and these are ours ... */
 #define _PAGE_DIRTY	0x20000

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 4/6] alpha: define granularity hint PTE bits
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
                   ` (2 preceding siblings ...)
  2026-10-08  3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  2026-10-08  3:24 ` [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
  2026-10-08  3:24 ` [PATCH v2 6/6] alpha: implement hugetlb support Matt Turner
  5 siblings, 0 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner

Bits <6:5> of the Alpha PTE are the granularity hint, described in Table
22-3 of the Alpha Architecture Reference Manual. A hint of N marks the
PTE as one of a block of 8^N physically contiguous, naturally aligned
pages that the translation buffer may map with a single entry. With 8KB
pages that gives 64KB, 512KB and 4MB blocks.

Define the field and the helpers to encode and decode it. Nothing sets a
non-zero hint yet.

Add the hint to _PAGE_CHG_MASK so that pte_modify() preserves it. Swap
PTEs leave bits <31:0> clear, so pte_huge() is false on swap, migration
and marker entries, and page table entries above the last level keep a
zero hint where pmd_bad() would reject anything else.

Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/include/asm/pgtable.h | 51 +++++++++++++++++++++++++++++++++++++++-
 1 file changed, 50 insertions(+), 1 deletion(-)

diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
index d22e28b956fb..7cf685a694b6 100644
--- a/arch/alpha/include/asm/pgtable.h
+++ b/arch/alpha/include/asm/pgtable.h
@@ -65,6 +65,8 @@ struct vm_area_struct;
 #define _PAGE_FOW	0x0004	/* used for page protection (fault on write) */
 #define _PAGE_FOE	0x0008	/* used for page protection (fault on exec) */
 #define _PAGE_ASM	0x0010
+#define _PAGE_GH_MASK	0x0060	/* granularity hint, PTE<6:5> */
+#define _PAGE_GH_SHIFT	5
 #define _PAGE_KRE	0x0100	/* kernel read enable, in __ACCESS_BITS */
 #define _PAGE_URE	0x0200	/* user read enable, in __ACCESS_BITS */
 #define _PAGE_KWE	0x1000	/* kernel write enable, in __DIRTY_BITS */
@@ -95,7 +97,36 @@ struct vm_area_struct;
 #define _PFN_MASK	0xFFFFFFFF00000000UL
 
 #define _PAGE_TABLE	(_PAGE_VALID | __DIRTY_BITS | __ACCESS_BITS)
-#define _PAGE_CHG_MASK	(_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | _PAGE_SPECIAL)
+/*
+ * pte_modify() must preserve the granularity hint, or
+ * hugetlb_change_protection() would turn a block into single-page PTEs.
+ */
+#define _PAGE_CHG_MASK	(_PFN_MASK | __DIRTY_BITS | __ACCESS_BITS | \
+			 _PAGE_GH_MASK | _PAGE_SPECIAL)
+
+/*
+ * Granularity hints. A PTE with hint order N is one of a block of 8^N
+ * physically contiguous pages, naturally aligned both virtually and
+ * physically, that the TB may map with a single entry. With 8KB pages the
+ * four encodings give 8KB, 64KB, 512KB and 4MB.
+ *
+ * Every PTE of a block must still carry the PFN of its own page, and all of
+ * them must agree in bits <15:0>: the protection, fault, hint and valid bits.
+ * __ACCESS_BITS and __DIRTY_BITS reach into that range, so young and dirty
+ * updates have to be applied to the whole block too.
+ *
+ * The hint is advisory: an implementation that ignores it still translates
+ * correctly through the individual PTEs, so it is safe to set on any Alpha.
+ */
+#define GH_ORDER_MAX	4			/* hint orders 0..3 */
+
+#define gh_cont_shift(order)	(PAGE_SHIFT + 3 * (order))
+#define gh_cont_size(order)	(1UL << gh_cont_shift(order))
+#define gh_cont_mask(order)	(~(gh_cont_size(order) - 1))
+#define gh_pte_num(order)	(1UL << (3 * (order)))
+
+#define for_each_gh_order(order) \
+	for ((order) = 1; (order) < GH_ORDER_MAX; (order)++)
 
 /*
  * All the normal masks have the "page accessed" bits on, as any time they are used,
@@ -261,6 +292,24 @@ extern inline pte_t pte_mkyoung(pte_t pte)	{ pte_val(pte) |= __ACCESS_BITS; retu
 extern inline int pte_special(pte_t pte)	{ return pte_val(pte) & _PAGE_SPECIAL; }
 extern inline pte_t pte_mkspecial(pte_t pte)	{ pte_val(pte) |= _PAGE_SPECIAL; return pte; }
 
+extern inline unsigned int pte_gh_order(pte_t pte)
+{
+	return (pte_val(pte) & _PAGE_GH_MASK) >> _PAGE_GH_SHIFT;
+}
+
+extern inline pte_t pte_mkgh(pte_t pte, unsigned int order)
+{
+	pte_val(pte) &= ~_PAGE_GH_MASK;
+	pte_val(pte) |= (unsigned long)order << _PAGE_GH_SHIFT;
+	return pte;
+}
+
+/*
+ * Swap PTEs leave bits <31:0> clear, so this is false for every swap,
+ * migration and marker entry, not just for ordinary small pages.
+ */
+extern inline int pte_huge(pte_t pte)		{ return pte_val(pte) & _PAGE_GH_MASK; }
+
 /*
  * The smp_rmb() in the following functions are required to order the load of
  * *dir (the pointer in the top level page table) with any subsequent load of

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area()
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
                   ` (3 preceding siblings ...)
  2026-10-08  3:24 ` [PATCH v2 4/6] alpha: define granularity hint PTE bits Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  2026-10-08  3:24 ` [PATCH v2 6/6] alpha: implement hugetlb support Matt Turner
  5 siblings, 0 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner

hugetlb_get_unmapped_area() does not align the address itself; it
delegates to mm_get_unmapped_area_vmaflags(), and the huge page
alignment is applied by generic_get_unmapped_area() and
generic_get_unmapped_area_topdown() via info.align_mask.

Alpha defines HAVE_ARCH_UNMAPPED_AREA and supplies its own
arch_get_unmapped_area(), which never sets align_mask. An
mmap(MAP_HUGETLB) with no address hint would therefore return a merely
page-aligned address.

That has been the architecture's job since commit 7bd3f1e1a9ae ("mm:
make hugetlb mappings go through mm_get_unmapped_area_vmflags"), whose
series converted the generic code along with x86, s390, sparc and
powerpc. LoongArch, which also has its own arch_get_unmapped_area(), was
missed and hit a BUG in LTP's hugefork02 until commit 3109d5ff484b
("LoongArch: Set hugetlb mmap base address aligned with pmd size").

Pass the file down to arch_get_unmapped_area_1() and set align_mask for
hugetlbfs mappings, as x86, s390, sparc and loongarch already do. This
is a no-op until alpha gains hugetlb support.

Reviewed-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/kernel/osf_sys.c | 14 +++++++++-----
 1 file changed, 9 insertions(+), 5 deletions(-)

diff --git a/arch/alpha/kernel/osf_sys.c b/arch/alpha/kernel/osf_sys.c
index 7b6543d2cca3..61d806176213 100644
--- a/arch/alpha/kernel/osf_sys.c
+++ b/arch/alpha/kernel/osf_sys.c
@@ -38,6 +38,7 @@
 #include <linux/namei.h>
 #include <linux/mount.h>
 #include <linux/uio.h>
+#include <linux/hugetlb.h>
 #include <linux/vfs.h>
 #include <linux/rcupdate.h>
 #include <linux/slab.h>
@@ -1201,14 +1202,16 @@ SYSCALL_DEFINE1(old_adjtimex, struct timex32 __user *, txc_p)
 /* Get an address range which is currently unmapped. */
 
 static unsigned long
-arch_get_unmapped_area_1(unsigned long addr, unsigned long len,
-		         unsigned long limit)
+arch_get_unmapped_area_1(struct file *filp, unsigned long addr,
+			 unsigned long len, unsigned long limit)
 {
 	struct vm_unmapped_area_info info = {};
 
 	info.length = len;
 	info.low_limit = addr;
 	info.high_limit = limit;
+	if (filp && is_file_hugepages(filp))
+		info.align_mask = huge_page_mask_align(filp);
 	return vm_unmapped_area(&info);
 }
 
@@ -1236,19 +1239,20 @@ arch_get_unmapped_area(struct file *filp, unsigned long addr,
 	   this feature should be incorporated into all ports?  */
 
 	if (addr) {
-		addr = arch_get_unmapped_area_1 (PAGE_ALIGN(addr), len, limit);
+		addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(addr), len,
+						 limit);
 		if (addr != (unsigned long) -ENOMEM)
 			return addr;
 	}
 
 	/* Next, try allocating at TASK_UNMAPPED_BASE.  */
-	addr = arch_get_unmapped_area_1 (PAGE_ALIGN(TASK_UNMAPPED_BASE),
+	addr = arch_get_unmapped_area_1 (filp, PAGE_ALIGN(TASK_UNMAPPED_BASE),
 					 len, limit);
 	if (addr != (unsigned long) -ENOMEM)
 		return addr;
 
 	/* Finally, try allocating in low memory.  */
-	addr = arch_get_unmapped_area_1 (PAGE_SIZE, len, limit);
+	addr = arch_get_unmapped_area_1 (filp, PAGE_SIZE, len, limit);
 
 	return addr;
 }

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH v2 6/6] alpha: implement hugetlb support
  2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
                   ` (4 preceding siblings ...)
  2026-10-08  3:24 ` [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
@ 2026-10-08  3:24 ` Matt Turner
  5 siblings, 0 replies; 8+ messages in thread
From: Matt Turner @ 2026-10-08  3:24 UTC (permalink / raw)
  To: Richard Henderson, Magnus Lindholm, Axel Rasmussen,
	Andrew Morton, Peter Xu
  Cc: linux-alpha, linux-kernel, Matt Turner

Alpha has no leaf entry above the last page table level, so the only
huge pages it can offer are granularity hint blocks: 64KB, 512KB and 4MB
with 8KB pages. Register all three as hstates, with 4MB as the default.
There is no PMD sized huge page, hence no transparent huge pages, no PMD
page table sharing and no gigantic pages.

As with arm64's contiguous PTEs, each PTE of a block keeps the frame
number of its own page, so a block is written with set_ptes(). Each
hstate's buddy order is its hint order, so a hugetlb folio is always
naturally aligned physically, as the hint requires, including after a
demote.

All PTEs of a block must agree in bits <15:0> (note 2 of Table 22-3 of
the Alpha Architecture Reference Manual), and both __ACCESS_BITS and
__DIRTY_BITS reach into that range, so every update rewrites the whole
block and huge_ptep_get() merges the young and dirty state back
together. Rewriting a valid block in place would leave its PTEs
disagreeing part way through, so changes to a valid block go through
break before make, following the procedure in section 11.6.1.

The hint is advisory, so this is safe on implementations that ignore it,
and gup_fast needs no changes since it walks the individual PTEs.

Enabling CONFIG_HUGETLB_PAGE derives pageblock_order from
HUGETLB_PAGE_ORDER, which lowers it from 10 to 9 and so halves the
anti-fragmentation and compaction granularity from 8MB to 4MB.

Hugepage migration is left disabled, since it has not been tested.

Co-developed-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Magnus Lindholm <linmag7@gmail.com>
Tested-by: Magnus Lindholm <linmag7@gmail.com>
Signed-off-by: Matt Turner <mattst88@gmail.com>
---
 arch/alpha/Kconfig               |   1 +
 arch/alpha/include/asm/hugetlb.h |  43 ++++++
 arch/alpha/include/asm/page.h    |  13 ++
 arch/alpha/mm/Makefile           |   2 +
 arch/alpha/mm/hugetlbpage.c      | 308 +++++++++++++++++++++++++++++++++++++++
 5 files changed, 367 insertions(+)

diff --git a/arch/alpha/Kconfig b/arch/alpha/Kconfig
index b5b02bb42aca..b2c359b947f6 100644
--- a/arch/alpha/Kconfig
+++ b/arch/alpha/Kconfig
@@ -19,6 +19,7 @@ config ALPHA
 	select ARCH_NO_PREEMPT
 	select ARCH_NO_SG_CHAIN
 	select ARCH_SUPPORTS_ATOMIC_RMW
+	select ARCH_SUPPORTS_HUGETLBFS
 	select ARCH_SUPPORTS_INT128 if CC_HAS_INT128
 	select ARCH_SUPPORTS_PAGE_TABLE_CHECK
 	select ARCH_USE_CMPXCHG_LOCKREF
diff --git a/arch/alpha/include/asm/hugetlb.h b/arch/alpha/include/asm/hugetlb.h
new file mode 100644
index 000000000000..ff69e24c775d
--- /dev/null
+++ b/arch/alpha/include/asm/hugetlb.h
@@ -0,0 +1,43 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _ALPHA_HUGETLB_H
+#define _ALPHA_HUGETLB_H
+
+#include <asm/page.h>
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags);
+#define arch_make_huge_pte arch_make_huge_pte
+
+#define __HAVE_ARCH_HUGE_PTEP_GET
+extern pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr,
+			   pte_t *ptep);
+#define __HAVE_ARCH_HUGE_SET_HUGE_PTE_AT
+extern void set_huge_pte_at(struct mm_struct *mm, unsigned long addr,
+			    pte_t *ptep, pte_t pte, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_GET_AND_CLEAR
+extern pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+				     pte_t *ptep, unsigned long sz);
+#define __HAVE_ARCH_HUGE_PTEP_CLEAR_FLUSH
+extern pte_t huge_ptep_clear_flush(struct vm_area_struct *vma,
+				   unsigned long addr, pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_WRPROTECT
+extern void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+				    pte_t *ptep);
+#define __HAVE_ARCH_HUGE_PTEP_SET_ACCESS_FLAGS
+extern int huge_ptep_set_access_flags(struct vm_area_struct *vma,
+				      unsigned long addr, pte_t *ptep,
+				      pte_t pte, int dirty);
+#define __HAVE_ARCH_HUGE_PTE_CLEAR
+extern void huge_pte_clear(struct mm_struct *mm, unsigned long addr,
+			   pte_t *ptep, unsigned long sz);
+
+/*
+ * The generic version does not flush, which would leave the old translation
+ * live across a restrictive change to PTE bits <15:0>.
+ */
+#define huge_ptep_modify_prot_start huge_ptep_modify_prot_start
+extern pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma,
+					 unsigned long addr, pte_t *ptep);
+
+#include <asm-generic/hugetlb.h>
+
+#endif /* _ALPHA_HUGETLB_H */
diff --git a/arch/alpha/include/asm/page.h b/arch/alpha/include/asm/page.h
index 59d01f9b77f6..67053c52953f 100644
--- a/arch/alpha/include/asm/page.h
+++ b/arch/alpha/include/asm/page.h
@@ -6,6 +6,19 @@
 #include <asm/pal.h>
 #include <vdso/page.h>
 
+#ifdef CONFIG_HUGETLB_PAGE
+/*
+ * The default huge page is the largest granularity hint block, 4MB. All
+ * three hint sizes are registered as hstates, hence HUGE_MAX_HSTATE, which
+ * linux/hugetlb.h needs before it includes asm/hugetlb.h.
+ */
+#define HPAGE_SHIFT		22
+#define HPAGE_SIZE		(_AC(1, UL) << HPAGE_SHIFT)
+#define HPAGE_MASK		(~(HPAGE_SIZE - 1))
+#define HUGETLB_PAGE_ORDER	(HPAGE_SHIFT - PAGE_SHIFT)
+#define HUGE_MAX_HSTATE		3
+#endif
+
 #ifndef __ASSEMBLER__
 
 #define STRICT_MM_TYPECHECKS
diff --git a/arch/alpha/mm/Makefile b/arch/alpha/mm/Makefile
index 2d05664058f6..c022e55f231a 100644
--- a/arch/alpha/mm/Makefile
+++ b/arch/alpha/mm/Makefile
@@ -4,3 +4,5 @@
 #
 
 obj-y	:= init.o fault.o tlbflush.o
+
+obj-$(CONFIG_HUGETLB_PAGE) += hugetlbpage.o
diff --git a/arch/alpha/mm/hugetlbpage.c b/arch/alpha/mm/hugetlbpage.c
new file mode 100644
index 000000000000..ca36af4667dc
--- /dev/null
+++ b/arch/alpha/mm/hugetlbpage.c
@@ -0,0 +1,308 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Alpha huge page support, built on the page table entry granularity hint
+ * described in asm/pgtable.h. All PTEs of a block have to agree in bits
+ * <15:0>, so every update below rewrites the whole block.
+ */
+
+#include <linux/hugetlb.h>
+#include <linux/mm.h>
+#include <linux/pgtable.h>
+
+#include <asm/tlbflush.h>
+
+/*
+ * Every hint size is smaller than PMD_SIZE, so a block always fits inside a
+ * single last level page table and its entry count is just its page count.
+ */
+static inline unsigned long num_contig_ptes(unsigned long sz)
+{
+	return sz >> PAGE_SHIFT;
+}
+
+/* Hint order that maps @sz, or 0 if @sz is not a hint size. */
+static unsigned int gh_order_from_size(unsigned long sz)
+{
+	unsigned int order;
+
+	for_each_gh_order(order)
+		if (gh_cont_size(order) == sz)
+			return order;
+
+	return 0;
+}
+
+pte_t *huge_pte_alloc(struct mm_struct *mm, struct vm_area_struct *vma,
+		      unsigned long addr, unsigned long sz)
+{
+	unsigned int order = gh_order_from_size(sz);
+	pgd_t *pgd;
+	p4d_t *p4d;
+	pud_t *pud;
+	pmd_t *pmd;
+
+	if (!order)
+		return NULL;
+
+	pgd = pgd_offset(mm, addr);
+	p4d = p4d_alloc(mm, pgd, addr);
+	if (!p4d)
+		return NULL;
+
+	pud = pud_alloc(mm, p4d, addr);
+	if (!pud)
+		return NULL;
+
+	pmd = pmd_alloc(mm, pud, addr);
+	if (!pmd)
+		return NULL;
+
+	return pte_alloc_huge(mm, pmd, addr & gh_cont_mask(order));
+}
+
+pte_t *huge_pte_offset(struct mm_struct *mm, unsigned long addr,
+		       unsigned long sz)
+{
+	unsigned int order = gh_order_from_size(sz);
+	pgd_t *pgd;
+	p4d_t *p4d;
+	pud_t *pud;
+	pmd_t *pmd;
+
+	if (!order)
+		return NULL;
+
+	pgd = pgd_offset(mm, addr);
+	if (!pgd_present(pgdp_get(pgd)))
+		return NULL;
+
+	p4d = p4d_offset(pgd, addr);
+	if (!p4d_present(p4dp_get(p4d)))
+		return NULL;
+
+	pud = pud_offset(p4d, addr);
+	if (!pud_present(pudp_get(pud)))
+		return NULL;
+
+	pmd = pmd_offset(pud, addr);
+	if (!pmd_present(pmdp_get(pmd)))
+		return NULL;
+
+	return pte_offset_huge(pmd, addr & gh_cont_mask(order));
+}
+
+/*
+ * A block is only as large as its own hint, so the walk can skip to the end of
+ * the containing last level page table but no further.
+ */
+unsigned long hugetlb_mask_last_page(struct hstate *h)
+{
+	return PMD_SIZE - huge_page_size(h);
+}
+
+pte_t huge_ptep_get(struct mm_struct *mm, unsigned long addr, pte_t *ptep)
+{
+	pte_t orig_pte = ptep_get(ptep);
+	unsigned long i, ncontig;
+
+	/* Swap, migration and marker entries carry no hint. */
+	if (!pte_huge(orig_pte))
+		return orig_pte;
+
+	ncontig = gh_pte_num(pte_gh_order(orig_pte));
+
+	for (i = 0; i < ncontig; i++, ptep++) {
+		pte_t pte = ptep_get(ptep);
+
+		if (pte_dirty(pte))
+			orig_pte = pte_mkdirty(orig_pte);
+		if (pte_young(pte))
+			orig_pte = pte_mkyoung(orig_pte);
+	}
+
+	return orig_pte;
+}
+
+static pte_t get_clear_contig(struct mm_struct *mm, unsigned long addr,
+			      pte_t *ptep, unsigned long ncontig)
+{
+	pte_t pte = ptep_get_and_clear(mm, addr, ptep);
+	bool present = pte_present(pte);
+
+	while (--ncontig) {
+		pte_t tmp_pte;
+
+		ptep++;
+		addr += PAGE_SIZE;
+		tmp_pte = ptep_get_and_clear(mm, addr, ptep);
+		if (present) {
+			if (pte_dirty(tmp_pte))
+				pte = pte_mkdirty(pte);
+			if (pte_young(tmp_pte))
+				pte = pte_mkyoung(pte);
+		}
+	}
+
+	return pte;
+}
+
+static pte_t get_clear_contig_flush(struct mm_struct *mm, unsigned long addr,
+				    pte_t *ptep, unsigned long ncontig)
+{
+	pte_t orig_pte = get_clear_contig(mm, addr, ptep, ncontig);
+	struct vm_area_struct vma = TLB_FLUSH_VMA(mm, 0);
+
+	if (!pte_none(orig_pte))
+		flush_tlb_range(&vma, addr, addr + ncontig * PAGE_SIZE);
+
+	return orig_pte;
+}
+
+pte_t arch_make_huge_pte(pte_t entry, unsigned int shift, vm_flags_t flags)
+{
+	unsigned int order = gh_order_from_size(1UL << shift);
+
+	/*
+	 * Not every caller asks for a size Alpha can hint; mm/debug_vm_pgtable.c
+	 * probes with PMD_SHIFT. Leave the entry alone rather than complaining.
+	 */
+	if (order)
+		entry = pte_mkgh(entry, order);
+
+	return entry;
+}
+
+void set_huge_pte_at(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+		     pte_t pte, unsigned long sz)
+{
+	unsigned long i, ncontig = num_contig_ptes(sz);
+
+	/* The hint requires the block to be naturally aligned in the VA too. */
+	VM_WARN_ON_ONCE(addr & (sz - 1));
+
+	if (!pte_present(pte)) {
+		/*
+		 * Swap, migration and marker entries hold no frame number, so
+		 * every slot gets the same value.
+		 */
+		for (i = 0; i < ncontig; i++, ptep++, addr += PAGE_SIZE)
+			set_ptes(mm, addr, ptep, pte, 1);
+		return;
+	}
+
+	/*
+	 * All PTEs of a block must agree in bits <15:0> (ARM Table 22-3, note
+	 * 2), which rewriting a valid block in place would violate part way
+	 * through. Invalidate it everywhere first, following ARM 11.6.1: the
+	 * invalidate has to assume the hint is zero, so it must cover every
+	 * page of the block.
+	 */
+	if (pte_present(ptep_get(ptep)))
+		get_clear_contig_flush(mm, addr, ptep, ncontig);
+
+	/* set_ptes() advances the frame number for each entry. */
+	set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+pte_t huge_ptep_get_and_clear(struct mm_struct *mm, unsigned long addr,
+			      pte_t *ptep, unsigned long sz)
+{
+	return get_clear_contig(mm, addr, ptep, num_contig_ptes(sz));
+}
+
+/*
+ * The block size comes from the hstate rather than from the entry, because
+ * this is also reached for non-present entries, which carry no hint.
+ */
+pte_t huge_ptep_clear_flush(struct vm_area_struct *vma, unsigned long addr,
+			    pte_t *ptep)
+{
+	unsigned long sz = huge_page_size(hstate_vma(vma));
+
+	return get_clear_contig_flush(vma->vm_mm, addr, ptep,
+				      num_contig_ptes(sz));
+}
+
+pte_t huge_ptep_modify_prot_start(struct vm_area_struct *vma, unsigned long addr,
+				  pte_t *ptep)
+{
+	return huge_ptep_clear_flush(vma, addr, ptep);
+}
+
+/* True if any entry of the block differs from @pte outside the frame number. */
+static bool gh_access_flags_changed(pte_t *ptep, pte_t pte,
+				    unsigned long ncontig)
+{
+	unsigned long i;
+
+	for (i = 0; i < ncontig; i++)
+		if ((pte_val(ptep_get(ptep + i)) ^ pte_val(pte)) & ~_PFN_MASK)
+			return true;
+
+	return false;
+}
+
+int huge_ptep_set_access_flags(struct vm_area_struct *vma, unsigned long addr,
+			       pte_t *ptep, pte_t pte, int dirty)
+{
+	struct mm_struct *mm = vma->vm_mm;
+	unsigned long sz = huge_page_size(hstate_vma(vma));
+	unsigned long ncontig = num_contig_ptes(sz);
+	pte_t orig_pte = huge_ptep_get(mm, addr, ptep);
+
+	/* Keep the young and dirty state the block already has. */
+	if (pte_dirty(orig_pte))
+		pte = pte_mkdirty(pte);
+	if (pte_young(orig_pte))
+		pte = pte_mkyoung(pte);
+
+	/* Breaking an unchanged block only makes concurrent touchers refault. */
+	if (!gh_access_flags_changed(ptep, pte, ncontig))
+		return 0;
+
+	get_clear_contig_flush(mm, addr, ptep, ncontig);
+	set_ptes(mm, addr, ptep, pte, ncontig);
+
+	return 1;
+}
+
+void huge_ptep_set_wrprotect(struct mm_struct *mm, unsigned long addr,
+			     pte_t *ptep)
+{
+	pte_t pte = ptep_get(ptep);
+	unsigned long ncontig = gh_pte_num(pte_gh_order(pte));
+
+	VM_WARN_ON_ONCE(!pte_huge(pte));
+
+	/*
+	 * pte_wrprotect() sets _PAGE_FOW, which lives in the bits the whole
+	 * block must agree on, so this takes the break before make path.
+	 */
+	pte = pte_wrprotect(get_clear_contig_flush(mm, addr, ptep, ncontig));
+	set_ptes(mm, addr, ptep, pte, ncontig);
+}
+
+void huge_pte_clear(struct mm_struct *mm, unsigned long addr, pte_t *ptep,
+		    unsigned long sz)
+{
+	unsigned long i, ncontig = num_contig_ptes(sz);
+
+	for (i = 0; i < ncontig; i++, addr += PAGE_SIZE, ptep++)
+		pte_clear(mm, addr, ptep);
+}
+
+bool __init arch_hugetlb_valid_size(unsigned long size)
+{
+	return gh_order_from_size(size) != 0;
+}
+
+static __init int alpha_hugetlbpage_init(void)
+{
+	unsigned int order;
+
+	for_each_gh_order(order)
+		hugetlb_add_hstate(gh_cont_shift(order) - PAGE_SHIFT);
+
+	return 0;
+}
+arch_initcall(alpha_hugetlbpage_init);

-- 
2.55.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH v2 3/6] alpha: describe the PTE read and write enable bits
  2026-10-08  3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
@ 2026-10-08 14:27   ` Magnus Lindholm
  0 siblings, 0 replies; 8+ messages in thread
From: Magnus Lindholm @ 2026-10-08 14:27 UTC (permalink / raw)
  To: Matt Turner
  Cc: Richard Henderson, Axel Rasmussen, Andrew Morton, Peter Xu,
	linux-alpha, linux-kernel

Hi Matt,

On Thu, Oct 8, 2026 at 5:24 AM Matt Turner <mattst88@gmail.com> wrote:
>
> The comments on _PAGE_KRE and _PAGE_URE said only "xxx". Describe the
> four enable bits as the Alpha Linux PTE defines them, in Table 22-3 of
> the Alpha Architecture Reference Manual: kernel and user read enable in
> bits 8 and 9, kernel and user write enable in bits 12 and 13, with bits
> <11:10> and <15:14> reserved. There are only two processor modes, user
> and kernel (section 22.5.1). Linux uses the read enables as its accessed
> bit and the write enables as its dirty bit, so say which of
> __ACCESS_BITS and __DIRTY_BITS each one belongs to.
>
> The names differ from what the hardware calls those bits on the 21264.
> Its PALcode loads the PTE unchanged into DTB_PTE (21264/EV67 Hardware
> Reference Manual, section 6.9), where bits 9 and 13 are the Executive
> read and write enables (Figure 5-27), Executive being mode 1 of the four
> the processor implements (Table 5-5). That is a detail below the PALcode
> interface, and it is also the layout of the OpenVMS PTE in Table 11-2,
> which is not the one Linux uses.
>
> No functional change.
>
> Suggested-by: Magnus Lindholm <linmag7@gmail.com>
> Signed-off-by: Matt Turner <mattst88@gmail.com>
> ---
>  arch/alpha/include/asm/pgtable.h | 8 ++++----
>  1 file changed, 4 insertions(+), 4 deletions(-)
>
> diff --git a/arch/alpha/include/asm/pgtable.h b/arch/alpha/include/asm/pgtable.h
> index 8a175e0c2b42..d22e28b956fb 100644
> --- a/arch/alpha/include/asm/pgtable.h
> +++ b/arch/alpha/include/asm/pgtable.h
> @@ -65,10 +65,10 @@ struct vm_area_struct;
>  #define _PAGE_FOW      0x0004  /* used for page protection (fault on write) */
>  #define _PAGE_FOE      0x0008  /* used for page protection (fault on exec) */
>  #define _PAGE_ASM      0x0010
> -#define _PAGE_KRE      0x0100  /* xxx - see below on the "accessed" bit */
> -#define _PAGE_URE      0x0200  /* xxx */
> -#define _PAGE_KWE      0x1000  /* used to do the dirty bit in software */
> -#define _PAGE_UWE      0x2000  /* used to do the dirty bit in software */
> +#define _PAGE_KRE      0x0100  /* kernel read enable, in __ACCESS_BITS */
> +#define _PAGE_URE      0x0200  /* user read enable, in __ACCESS_BITS */
> +#define _PAGE_KWE      0x1000  /* kernel write enable, in __DIRTY_BITS */
> +#define _PAGE_UWE      0x2000  /* user write enable, in __DIRTY_BITS */
>
>  /* .. and these are ours ... */
>  #define _PAGE_DIRTY    0x20000
>
> --
> 2.55.0
>

Looks good to me.

Reviewed-by: Magnus Lindholm <linmag7@gmail.com>

^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-10-08 14:28 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08  3:24 [PATCH v2 0/6] alpha: hugetlb support using granularity hints Matt Turner
2026-10-08  3:24 ` [PATCH v2 1/6] alpha: add missing page_table_check_pte_clear() to ptep_get_and_clear() Matt Turner
2026-10-08  3:24 ` [PATCH v2 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault() Matt Turner
2026-10-08  3:24 ` [PATCH v2 3/6] alpha: describe the PTE read and write enable bits Matt Turner
2026-10-08 14:27   ` Magnus Lindholm
2026-10-08  3:24 ` [PATCH v2 4/6] alpha: define granularity hint PTE bits Matt Turner
2026-10-08  3:24 ` [PATCH v2 5/6] alpha: align hugetlb mappings in arch_get_unmapped_area() Matt Turner
2026-10-08  3:24 ` [PATCH v2 6/6] alpha: implement hugetlb support Matt Turner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®