* [PATCH 0/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range()
@ 2026-02-02 4:26 Anshuman Khandual
2026-02-02 4:26 ` [PATCH 1/2] " Anshuman Khandual
2026-02-02 4:26 ` [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping Anshuman Khandual
0 siblings, 2 replies; 8+ messages in thread
From: Anshuman Khandual @ 2026-02-02 4:26 UTC (permalink / raw)
To: linux-arm-kernel
Cc: Anshuman Khandual, Catalin Marinas, Will Deacon, Ryan Roberts,
Yang Shi, Christoph Lameter, linux-kernel
This series enables batched TLB flush in unmap_hotplug_range() which avoids
individual page TLB flush for potential CONT blocks in linear mapping while
also improving performance due to range based TLB operation along with less
synchronization barrier instructions.
It also now rejects memory removal that might split a leaf entry in kernel
mapping, which would have otherwise required re-structuring using the break
before make (BBM) sematics.
This series applies on 6.19-rc8 and tested on KVM guest.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Yang Shi <yang@os.amperecomputing.com>
Cc: Christoph Lameter <cl@gentwo.org>
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-kernel@vger.kernel.org
Anshuman Khandual (2):
arm64/mm: Enable batched TLB flush in unmap_hotplug_range()
arm64/mm: Reject memory removal that splits a kernel leaf mapping
arch/arm64/mm/mmu.c | 207 +++++++++++++++++++++++++++++++++++++++++---
1 file changed, 193 insertions(+), 14 deletions(-)
--
2.30.2
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 1/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range()
2026-02-02 4:26 [PATCH 0/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range() Anshuman Khandual
@ 2026-02-02 4:26 ` Anshuman Khandual
2026-02-02 9:18 ` Ryan Roberts
2026-02-02 4:26 ` [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping Anshuman Khandual
1 sibling, 1 reply; 8+ messages in thread
From: Anshuman Khandual @ 2026-02-02 4:26 UTC (permalink / raw)
To: linux-arm-kernel
Cc: Anshuman Khandual, Catalin Marinas, Will Deacon, Ryan Roberts,
Yang Shi, Christoph Lameter, linux-kernel, stable
During a memory hot remove operartion both linear and vmemmap mappings for
the memory range being removed, get unmapped via unmap_hotplug_range() but
mapped pages get freed only for vmemmap mapping. This is just a sequential
operation where each table entry gets cleared, followed by a leaf specific
TLB flush, and then followed by memory free operation when applicable.
This approach was simple and uniform both for vmemmap and linear mappings.
But linear mapping might contain CONT marked block memory where it becomes
necessary to first clear out all entire in the range before a TLB flush.
This is as per the architecture requirement. Hence batch all TLB flushes
during the table tear down walk and finally do it in unmap_hotplug_range().
Besides it is helps in improving the performance via TLBI range operation
along with reduced synchronization instructions. The time spent executing
unmap_hotplug_range() improved 97% measured over a 2GB memory hot removal
in KVM guest.
This scheme is not applicable during vmemmap mapping tear down where memory
needs to be freed and hence a TLB flush is required after clearing out page
table entry.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-kernel@vger.kernel.org
Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
Cc: stable@vger.kernel.org
Signed-off-by: Ryan Roberts <ryan.roberts@arm.com>
Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
---
arch/arm64/mm/mmu.c | 81 +++++++++++++++++++++++++++++++++++++--------
1 file changed, 67 insertions(+), 14 deletions(-)
diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
index 8e1d80a7033e..8ec8a287aaa1 100644
--- a/arch/arm64/mm/mmu.c
+++ b/arch/arm64/mm/mmu.c
@@ -1458,10 +1458,32 @@ static void unmap_hotplug_pte_range(pmd_t *pmdp, unsigned long addr,
WARN_ON(!pte_present(pte));
__pte_clear(&init_mm, addr, ptep);
- flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
- if (free_mapped)
+ if (free_mapped) {
+ /*
+ * If page is part of an existing contiguous
+ * memory block, individual TLB invalidation
+ * here would not be appropriate. Instead it
+ * will require clearing all entries for the
+ * memory block and subsequently a TLB flush
+ * for the entire range.
+ */
+ WARN_ON(pte_cont(pte));
+
+ /*
+ * TLB flush is essential for freeing memory.
+ */
+ flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
free_hotplug_page_range(pte_page(pte),
PAGE_SIZE, altmap);
+ }
+
+ /*
+ * TLB flush is batched in unmap_hotplug_range()
+ * for the entire range, when memory need not be
+ * freed. Besides linear mapping might have CONT
+ * blocks where TLB flush needs to be done after
+ * clearing all relevant entries.
+ */
} while (addr += PAGE_SIZE, addr < end);
}
@@ -1482,15 +1504,32 @@ static void unmap_hotplug_pmd_range(pud_t *pudp, unsigned long addr,
WARN_ON(!pmd_present(pmd));
if (pmd_sect(pmd)) {
pmd_clear(pmdp);
+ if (free_mapped) {
+ /*
+ * If page is part of an existing contiguous
+ * memory block, individual TLB invalidation
+ * here would not be appropriate. Instead it
+ * will require clearing all entries for the
+ * memory block and subsequently a TLB flush
+ * for the entire range.
+ */
+ WARN_ON(pmd_cont(pmd));
+
+ /*
+ * TLB flush is essential for freeing memory.
+ */
+ flush_tlb_kernel_range(addr, addr + PMD_SIZE);
+ free_hotplug_page_range(pmd_page(pmd),
+ PMD_SIZE, altmap);
+ }
/*
- * One TLBI should be sufficient here as the PMD_SIZE
- * range is mapped with a single block entry.
+ * TLB flush is batched in unmap_hotplug_range()
+ * for the entire range, when memory need not be
+ * freed. Besides linear mapping might have CONT
+ * blocks where TLB flush needs to be done after
+ * clearing all relevant entries.
*/
- flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
- if (free_mapped)
- free_hotplug_page_range(pmd_page(pmd),
- PMD_SIZE, altmap);
continue;
}
WARN_ON(!pmd_table(pmd));
@@ -1515,15 +1554,20 @@ static void unmap_hotplug_pud_range(p4d_t *p4dp, unsigned long addr,
WARN_ON(!pud_present(pud));
if (pud_sect(pud)) {
pud_clear(pudp);
+ if (free_mapped) {
+ /*
+ * TLB flush is essential for freeing memory.
+ */
+ flush_tlb_kernel_range(addr, addr + PUD_SIZE);
+ free_hotplug_page_range(pud_page(pud),
+ PUD_SIZE, altmap);
+ }
/*
- * One TLBI should be sufficient here as the PUD_SIZE
- * range is mapped with a single block entry.
+ * TLB flush is batched in unmap_hotplug_range()
+ * for the entire range, when memory need not be
+ * freed.
*/
- flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
- if (free_mapped)
- free_hotplug_page_range(pud_page(pud),
- PUD_SIZE, altmap);
continue;
}
WARN_ON(!pud_table(pud));
@@ -1553,6 +1597,7 @@ static void unmap_hotplug_p4d_range(pgd_t *pgdp, unsigned long addr,
static void unmap_hotplug_range(unsigned long addr, unsigned long end,
bool free_mapped, struct vmem_altmap *altmap)
{
+ unsigned long start = addr;
unsigned long next;
pgd_t *pgdp, pgd;
@@ -1574,6 +1619,14 @@ static void unmap_hotplug_range(unsigned long addr, unsigned long end,
WARN_ON(!pgd_present(pgd));
unmap_hotplug_p4d_range(pgdp, addr, next, free_mapped, altmap);
} while (addr = next, addr < end);
+
+ /*
+ * Batched TLB flush only for linear mapping which
+ * might contain CONT blocks, and does not require
+ * freeing up memory as well.
+ */
+ if (!free_mapped)
+ flush_tlb_kernel_range(start, end);
}
static void free_empty_pte_table(pmd_t *pmdp, unsigned long addr,
--
2.30.2
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping
2026-02-02 4:26 [PATCH 0/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range() Anshuman Khandual
2026-02-02 4:26 ` [PATCH 1/2] " Anshuman Khandual
@ 2026-02-02 4:26 ` Anshuman Khandual
2026-02-02 9:42 ` Ryan Roberts
1 sibling, 1 reply; 8+ messages in thread
From: Anshuman Khandual @ 2026-02-02 4:26 UTC (permalink / raw)
To: linux-arm-kernel
Cc: Anshuman Khandual, Catalin Marinas, Will Deacon, Ryan Roberts,
Yang Shi, Christoph Lameter, linux-kernel, stable
Linear and vmemmap mapings that get teared down during a memory hot remove
operation might contain leaf level entries on any page table level. If the
requested memory range's linear or vmemmap mappings falls within such leaf
entries, new mappings need to be created for the remaning memory mapped on
the leaf entry earlier, following standard break before make aka BBM rules.
Currently memory hot remove operation does not perform such restructuring,
and so removing memory ranges that could split a kernel leaf level mapping
need to be rejected.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-kernel@vger.kernel.org
Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
Cc: stable@vger.kernel.org
Suggested-by: Ryan Roberts <ryan.roberts@arm.com>
Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
---
arch/arm64/mm/mmu.c | 126 ++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 126 insertions(+)
diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
index 8ec8a287aaa1..9d59e10fb3de 100644
--- a/arch/arm64/mm/mmu.c
+++ b/arch/arm64/mm/mmu.c
@@ -2063,6 +2063,129 @@ void arch_remove_memory(u64 start, u64 size, struct vmem_altmap *altmap)
__remove_pgd_mapping(swapper_pg_dir, __phys_to_virt(start), size);
}
+
+static bool split_kernel_leaf_boundary(unsigned long addr)
+{
+ pgd_t *pgdp, pgd;
+ p4d_t *p4dp, p4d;
+ pud_t *pudp, pud;
+ pmd_t *pmdp, pmd;
+ pte_t *ptep, pte;
+
+ /*
+ * PGD: If addr is PGD aligned then addr already
+ * describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, PGDIR_SIZE) == addr)
+ return false;
+
+ pgdp = pgd_offset_k(addr);
+ pgd = pgdp_get(pgdp);
+ if (!pgd_present(pgd))
+ return false;
+
+ /*
+ * P4D: If addr is P4D aligned then addr already
+ * describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, P4D_SIZE) == addr)
+ return false;
+
+ p4dp = p4d_offset(pgdp, addr);
+ p4d = p4dp_get(p4dp);
+ if (!p4d_present(p4d))
+ return false;
+
+ /*
+ * PUD: If addr is PUD aligned then addr already
+ * describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, PUD_SIZE) == addr)
+ return false;
+
+ pudp = pud_offset(p4dp, addr);
+ pud = pudp_get(pudp);
+ if (!pud_present(pud))
+ return false;
+
+ if (pud_leaf(pud))
+ return true;
+
+ /*
+ * CONT_PMD: If addr is CONT_PMD aligned then
+ * addr already describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, CONT_PMD_SIZE) == addr)
+ return false;
+
+ pmdp = pmd_offset(pudp, addr);
+ pmd = pmdp_get(pmdp);
+ if (!pmd_present(pmd))
+ return false;
+
+ if (pmd_leaf(pmd) && pmd_cont(pmd))
+ return true;
+
+ /*
+ * PMD: If addr is PMD aligned then addr already
+ * describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, PMD_SIZE) == addr)
+ return false;
+
+ if (pmd_leaf(pmd))
+ return true;
+
+ /*
+ * CONT_PTE: If addr is CONT_PTE aligned then addr
+ * already describes a leaf boundary.
+ */
+ if (ALIGN_DOWN(addr, CONT_PTE_SIZE) == addr)
+ return false;
+
+ ptep = pte_offset_kernel(pmdp, addr);
+ pte = __ptep_get(ptep);
+ if (!pte_present(pte))
+ return false;
+
+ if (pte_valid(pte) && pte_cont(pte))
+ return true;
+
+ if (ALIGN_DOWN(addr, PAGE_SIZE) == addr)
+ return false;
+ return true;
+}
+
+static bool can_unmap_without_split(unsigned long pfn, unsigned long nr_pages)
+{
+ unsigned long linear_start, linear_end, phys_start, phys_end;
+ unsigned long vmemmap_size, vmemmap_start, vmemmap_end;
+
+ /* Assert linear map edges do not split a leaf entry */
+ phys_start = PFN_PHYS(pfn);
+ phys_end = phys_start + nr_pages * PAGE_SIZE;
+ linear_start = __phys_to_virt(phys_start);
+ linear_end = __phys_to_virt(phys_end);
+ if (split_kernel_leaf_boundary(linear_start) ||
+ split_kernel_leaf_boundary(linear_end)) {
+ pr_warn("[%lx %lx] splits a leaf entry in linear map\n",
+ phys_start, phys_end);
+ return false;
+ }
+
+ /* Assert vmemmap edges do not split a leaf entry */
+ vmemmap_size = nr_pages * sizeof(struct page);
+ vmemmap_start = (unsigned long) pfn_to_page(pfn);
+ vmemmap_end = vmemmap_start + vmemmap_size;
+ if (split_kernel_leaf_boundary(vmemmap_start) ||
+ split_kernel_leaf_boundary(vmemmap_end)) {
+ pr_warn("[%lx %lx] splits a leaf entry in vmemmap\n",
+ phys_start, phys_end);
+ return false;
+ }
+ return true;
+}
+
/*
* This memory hotplug notifier helps prevent boot memory from being
* inadvertently removed as it blocks pfn range offlining process in
@@ -2083,6 +2206,9 @@ static int prevent_bootmem_remove_notifier(struct notifier_block *nb,
if ((action != MEM_GOING_OFFLINE) && (action != MEM_OFFLINE))
return NOTIFY_OK;
+ if (!can_unmap_without_split(pfn, arg->nr_pages))
+ return NOTIFY_BAD;
+
for (; pfn < end_pfn; pfn += PAGES_PER_SECTION) {
unsigned long start = PFN_PHYS(pfn);
unsigned long end = start + (1UL << PA_SECTION_SHIFT);
--
2.30.2
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range()
2026-02-02 4:26 ` [PATCH 1/2] " Anshuman Khandual
@ 2026-02-02 9:18 ` Ryan Roberts
2026-02-02 10:48 ` Anshuman Khandual
0 siblings, 1 reply; 8+ messages in thread
From: Ryan Roberts @ 2026-02-02 9:18 UTC (permalink / raw)
To: Anshuman Khandual, linux-arm-kernel
Cc: Catalin Marinas, Will Deacon, Yang Shi, Christoph Lameter,
linux-kernel, stable
On 02/02/2026 04:26, Anshuman Khandual wrote:
> During a memory hot remove operartion both linear and vmemmap mappings for
> the memory range being removed, get unmapped via unmap_hotplug_range() but
> mapped pages get freed only for vmemmap mapping. This is just a sequential
> operation where each table entry gets cleared, followed by a leaf specific
> TLB flush, and then followed by memory free operation when applicable.
>
> This approach was simple and uniform both for vmemmap and linear mappings.
> But linear mapping might contain CONT marked block memory where it becomes
> necessary to first clear out all entire in the range before a TLB flush.
> This is as per the architecture requirement. Hence batch all TLB flushes
> during the table tear down walk and finally do it in unmap_hotplug_range().
I might be worth mentioning the impact of not bein architecture compliant here?
Something like:
Prior to this fix, it was hypothetically possible for a speculative access to
a higher address in the contiguous block to fill the TLB with shattered
entries for the entire contiguous range after a lower address had already been
cleared and invalidated. Due to the entries being shattered, the subsequent
tlbi for the higher address would not then clear the TLB entries for the lower
address, meaning stale TLB entries could persist.
>
> Besides it is helps in improving the performance via TLBI range operation
nit: ^^ (remove)
> along with reduced synchronization instructions. The time spent executing
> unmap_hotplug_range() improved 97% measured over a 2GB memory hot removal
> in KVM guest.
That's a great improvement :)
>
> This scheme is not applicable during vmemmap mapping tear down where memory
> needs to be freed and hence a TLB flush is required after clearing out page
> table entry.
>
> Cc: Catalin Marinas <catalin.marinas@arm.com>
> Cc: Will Deacon <will@kernel.org>
> Cc: linux-arm-kernel@lists.infradead.org
> Cc: linux-kernel@vger.kernel.org
> Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
> Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
> Cc: stable@vger.kernel.org
> Signed-off-by: Ryan Roberts <ryan.roberts@arm.com>
> Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
I suggested the original shape of this and I see you have added my SOB. Final
patch looks good to me - I'm not sure if it's correct for me to add Rb, but here
it is regardless:
Reviewed-by: Ryan Roberts <ryan.roberts@arm.com>
> ---
> arch/arm64/mm/mmu.c | 81 +++++++++++++++++++++++++++++++++++++--------
> 1 file changed, 67 insertions(+), 14 deletions(-)
>
> diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
> index 8e1d80a7033e..8ec8a287aaa1 100644
> --- a/arch/arm64/mm/mmu.c
> +++ b/arch/arm64/mm/mmu.c
> @@ -1458,10 +1458,32 @@ static void unmap_hotplug_pte_range(pmd_t *pmdp, unsigned long addr,
>
> WARN_ON(!pte_present(pte));
> __pte_clear(&init_mm, addr, ptep);
> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
> - if (free_mapped)
> + if (free_mapped) {
> + /*
> + * If page is part of an existing contiguous
> + * memory block, individual TLB invalidation
> + * here would not be appropriate. Instead it
> + * will require clearing all entries for the
> + * memory block and subsequently a TLB flush
> + * for the entire range.
> + */
> + WARN_ON(pte_cont(pte));
> +
> + /*
> + * TLB flush is essential for freeing memory.
> + */
> + flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
> free_hotplug_page_range(pte_page(pte),
> PAGE_SIZE, altmap);
> + }
> +
> + /*
> + * TLB flush is batched in unmap_hotplug_range()
> + * for the entire range, when memory need not be
> + * freed. Besides linear mapping might have CONT
> + * blocks where TLB flush needs to be done after
> + * clearing all relevant entries.
> + */
> } while (addr += PAGE_SIZE, addr < end);
> }
>
> @@ -1482,15 +1504,32 @@ static void unmap_hotplug_pmd_range(pud_t *pudp, unsigned long addr,
> WARN_ON(!pmd_present(pmd));
> if (pmd_sect(pmd)) {
> pmd_clear(pmdp);
> + if (free_mapped) {
> + /*
> + * If page is part of an existing contiguous
> + * memory block, individual TLB invalidation
> + * here would not be appropriate. Instead it
> + * will require clearing all entries for the
> + * memory block and subsequently a TLB flush
> + * for the entire range.
> + */
> + WARN_ON(pmd_cont(pmd));
> +
> + /*
> + * TLB flush is essential for freeing memory.
> + */
> + flush_tlb_kernel_range(addr, addr + PMD_SIZE);
> + free_hotplug_page_range(pmd_page(pmd),
> + PMD_SIZE, altmap);
> + }
>
> /*
> - * One TLBI should be sufficient here as the PMD_SIZE
> - * range is mapped with a single block entry.
> + * TLB flush is batched in unmap_hotplug_range()
> + * for the entire range, when memory need not be
> + * freed. Besides linear mapping might have CONT
> + * blocks where TLB flush needs to be done after
> + * clearing all relevant entries.
> */
> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
> - if (free_mapped)
> - free_hotplug_page_range(pmd_page(pmd),
> - PMD_SIZE, altmap);
> continue;
> }
> WARN_ON(!pmd_table(pmd));
> @@ -1515,15 +1554,20 @@ static void unmap_hotplug_pud_range(p4d_t *p4dp, unsigned long addr,
> WARN_ON(!pud_present(pud));
> if (pud_sect(pud)) {
> pud_clear(pudp);
> + if (free_mapped) {
> + /*
> + * TLB flush is essential for freeing memory.
> + */
> + flush_tlb_kernel_range(addr, addr + PUD_SIZE);
> + free_hotplug_page_range(pud_page(pud),
> + PUD_SIZE, altmap);
> + }
>
> /*
> - * One TLBI should be sufficient here as the PUD_SIZE
> - * range is mapped with a single block entry.
> + * TLB flush is batched in unmap_hotplug_range()
> + * for the entire range, when memory need not be
> + * freed.
> */
> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
> - if (free_mapped)
> - free_hotplug_page_range(pud_page(pud),
> - PUD_SIZE, altmap);
> continue;
> }
> WARN_ON(!pud_table(pud));
> @@ -1553,6 +1597,7 @@ static void unmap_hotplug_p4d_range(pgd_t *pgdp, unsigned long addr,
> static void unmap_hotplug_range(unsigned long addr, unsigned long end,
> bool free_mapped, struct vmem_altmap *altmap)
> {
> + unsigned long start = addr;
> unsigned long next;
> pgd_t *pgdp, pgd;
>
> @@ -1574,6 +1619,14 @@ static void unmap_hotplug_range(unsigned long addr, unsigned long end,
> WARN_ON(!pgd_present(pgd));
> unmap_hotplug_p4d_range(pgdp, addr, next, free_mapped, altmap);
> } while (addr = next, addr < end);
> +
> + /*
> + * Batched TLB flush only for linear mapping which
> + * might contain CONT blocks, and does not require
> + * freeing up memory as well.
> + */
> + if (!free_mapped)
> + flush_tlb_kernel_range(start, end);
> }
>
> static void free_empty_pte_table(pmd_t *pmdp, unsigned long addr,
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping
2026-02-02 4:26 ` [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping Anshuman Khandual
@ 2026-02-02 9:42 ` Ryan Roberts
2026-02-02 11:06 ` Anshuman Khandual
0 siblings, 1 reply; 8+ messages in thread
From: Ryan Roberts @ 2026-02-02 9:42 UTC (permalink / raw)
To: Anshuman Khandual, linux-arm-kernel
Cc: Catalin Marinas, Will Deacon, Yang Shi, Christoph Lameter,
linux-kernel, stable
On 02/02/2026 04:26, Anshuman Khandual wrote:
> Linear and vmemmap mapings that get teared down during a memory hot remove
> operation might contain leaf level entries on any page table level. If the
> requested memory range's linear or vmemmap mappings falls within such leaf
> entries, new mappings need to be created for the remaning memory mapped on
> the leaf entry earlier, following standard break before make aka BBM rules.
I think it would be good to mention that the kernel cannot tolerate BBM so
remapping to fine grained leaves would not be possible on systems without
BBML2_NOABORT.
>
> Currently memory hot remove operation does not perform such restructuring,
> and so removing memory ranges that could split a kernel leaf level mapping
> need to be rejected.
Perhaps it is useful to mention that while memory_hotplug.c does appear to
permit hot-unplugging arbitrary ranges of memory, the higher layers that drive
memory_hotplug (e.g. ACPI, virtio, ...) all appear to treat memory as fixed size
devices so it is impossible to hotunplug a different amount than was previously
hotplugged, and so we should never see a rejection in practice, but adding the
check makes us robust against a future change.
>
> Cc: Catalin Marinas <catalin.marinas@arm.com>
> Cc: Will Deacon <will@kernel.org>
> Cc: linux-arm-kernel@lists.infradead.org
> Cc: linux-kernel@vger.kernel.org
> Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
> Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
> Cc: stable@vger.kernel.org
> Suggested-by: Ryan Roberts <ryan.roberts@arm.com>
> Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
> ---
> arch/arm64/mm/mmu.c | 126 ++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 126 insertions(+)
>
> diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
> index 8ec8a287aaa1..9d59e10fb3de 100644
> --- a/arch/arm64/mm/mmu.c
> +++ b/arch/arm64/mm/mmu.c
> @@ -2063,6 +2063,129 @@ void arch_remove_memory(u64 start, u64 size, struct vmem_altmap *altmap)
> __remove_pgd_mapping(swapper_pg_dir, __phys_to_virt(start), size);
> }
>
> +
> +static bool split_kernel_leaf_boundary(unsigned long addr)
The name currently makes it sound like we are asking for the mapping to be split
(we have existing functions to do this that are named similarly). Perhaps a
better name would be addr_splits_leaf()?
> +{
> + pgd_t *pgdp, pgd;
> + p4d_t *p4dp, p4d;
> + pud_t *pudp, pud;
> + pmd_t *pmdp, pmd;
> + pte_t *ptep, pte;
> +
> + /*
> + * PGD: If addr is PGD aligned then addr already
> + * describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, PGDIR_SIZE) == addr)
> + return false;
> +
> + pgdp = pgd_offset_k(addr);
> + pgd = pgdp_get(pgdp);
> + if (!pgd_present(pgd))
> + return false;
> +
> + /*
> + * P4D: If addr is P4D aligned then addr already
> + * describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, P4D_SIZE) == addr)
> + return false;
> +
> + p4dp = p4d_offset(pgdp, addr);
> + p4d = p4dp_get(p4dp);
> + if (!p4d_present(p4d))
> + return false;
> +
> + /*
> + * PUD: If addr is PUD aligned then addr already
> + * describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, PUD_SIZE) == addr)
> + return false;
> +
> + pudp = pud_offset(p4dp, addr);
> + pud = pudp_get(pudp);
> + if (!pud_present(pud))
> + return false;
> +
> + if (pud_leaf(pud))
> + return true;
> +
> + /*
> + * CONT_PMD: If addr is CONT_PMD aligned then
> + * addr already describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, CONT_PMD_SIZE) == addr)
> + return false;
> +
> + pmdp = pmd_offset(pudp, addr);
> + pmd = pmdp_get(pmdp);
> + if (!pmd_present(pmd))
> + return false;
> +
> + if (pmd_leaf(pmd) && pmd_cont(pmd))
> + return true;
> +
> + /*
> + * PMD: If addr is PMD aligned then addr already
> + * describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, PMD_SIZE) == addr)
> + return false;
> +
> + if (pmd_leaf(pmd))
> + return true;
> +
> + /*
> + * CONT_PTE: If addr is CONT_PTE aligned then addr
> + * already describes a leaf boundary.
> + */
> + if (ALIGN_DOWN(addr, CONT_PTE_SIZE) == addr)
> + return false;
> +
> + ptep = pte_offset_kernel(pmdp, addr);
> + pte = __ptep_get(ptep);
> + if (!pte_present(pte))
> + return false;
> +
> + if (pte_valid(pte) && pte_cont(pte))
Why do you need pte_valid() here? You have already checked !pte_present(). Are
you expecting a case of present but not valid (PTE_PRESENT_INVALID)? If so, do
you need to consider that for the other levels too? (pmd_leaf() only checks
pmd_present()).
Personally I think you can just drop the pte_valid() check here.
> + return true;
> +
> + if (ALIGN_DOWN(addr, PAGE_SIZE) == addr)
> + return false;
> + return true;
> +}
> +
> +static bool can_unmap_without_split(unsigned long pfn, unsigned long nr_pages)
> +{
> + unsigned long linear_start, linear_end, phys_start, phys_end;
> + unsigned long vmemmap_size, vmemmap_start, vmemmap_end;
nit: do we need all these variables. Perhaps just:
unsigned long sz, start, end, phys_start, phys_end;
are sufficient?
> +
> + /* Assert linear map edges do not split a leaf entry */
> + phys_start = PFN_PHYS(pfn);
> + phys_end = phys_start + nr_pages * PAGE_SIZE;
> + linear_start = __phys_to_virt(phys_start);
> + linear_end = __phys_to_virt(phys_end);
> + if (split_kernel_leaf_boundary(linear_start) ||
> + split_kernel_leaf_boundary(linear_end)) {
> + pr_warn("[%lx %lx] splits a leaf entry in linear map\n",
> + phys_start, phys_end);
> + return false;
> + }
> +
> + /* Assert vmemmap edges do not split a leaf entry */
> + vmemmap_size = nr_pages * sizeof(struct page);
> + vmemmap_start = (unsigned long) pfn_to_page(pfn);
nit: ^
I don't think we would normally have that space?
> + vmemmap_end = vmemmap_start + vmemmap_size;
> + if (split_kernel_leaf_boundary(vmemmap_start) ||
> + split_kernel_leaf_boundary(vmemmap_end)) {
> + pr_warn("[%lx %lx] splits a leaf entry in vmemmap\n",
> + phys_start, phys_end);
> + return false;
> + }
> + return true;
> +}
> +
> /*
> * This memory hotplug notifier helps prevent boot memory from being
> * inadvertently removed as it blocks pfn range offlining process in
> @@ -2083,6 +2206,9 @@ static int prevent_bootmem_remove_notifier(struct notifier_block *nb,
> if ((action != MEM_GOING_OFFLINE) && (action != MEM_OFFLINE))
> return NOTIFY_OK;
>
> + if (!can_unmap_without_split(pfn, arg->nr_pages))
> + return NOTIFY_BAD;
> +
Personally, I'd keep the bootmem check first and do this check after. That means
an existing warning will not change.
Thanks,
Ryan
> for (; pfn < end_pfn; pfn += PAGES_PER_SECTION) {
> unsigned long start = PFN_PHYS(pfn);
> unsigned long end = start + (1UL << PA_SECTION_SHIFT);
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range()
2026-02-02 9:18 ` Ryan Roberts
@ 2026-02-02 10:48 ` Anshuman Khandual
0 siblings, 0 replies; 8+ messages in thread
From: Anshuman Khandual @ 2026-02-02 10:48 UTC (permalink / raw)
To: Ryan Roberts, linux-arm-kernel
Cc: Catalin Marinas, Will Deacon, Yang Shi, Christoph Lameter,
linux-kernel, stable
On 02/02/26 2:48 PM, Ryan Roberts wrote:
> On 02/02/2026 04:26, Anshuman Khandual wrote:
>> During a memory hot remove operartion both linear and vmemmap mappings for
>> the memory range being removed, get unmapped via unmap_hotplug_range() but
>> mapped pages get freed only for vmemmap mapping. This is just a sequential
>> operation where each table entry gets cleared, followed by a leaf specific
>> TLB flush, and then followed by memory free operation when applicable.
>>
>> This approach was simple and uniform both for vmemmap and linear mappings.
>> But linear mapping might contain CONT marked block memory where it becomes
>> necessary to first clear out all entire in the range before a TLB flush.
>> This is as per the architecture requirement. Hence batch all TLB flushes
>> during the table tear down walk and finally do it in unmap_hotplug_range().
>
> I might be worth mentioning the impact of not bein architecture compliant here?
>
> Something like:
>
> Prior to this fix, it was hypothetically possible for a speculative access to
> a higher address in the contiguous block to fill the TLB with shattered
> entries for the entire contiguous range after a lower address had already been
> cleared and invalidated. Due to the entries being shattered, the subsequent
> tlbi for the higher address would not then clear the TLB entries for the lower
> address, meaning stale TLB entries could persist.
Sounds good - will add in the commit message.
>
>>
>> Besides it is helps in improving the performance via TLBI range operation
>
> nit: ^^ (remove)
Will fix that.
>
>> along with reduced synchronization instructions. The time spent executing
>> unmap_hotplug_range() improved 97% measured over a 2GB memory hot removal
>> in KVM guest.
>
> That's a great improvement :)
>
>>
>> This scheme is not applicable during vmemmap mapping tear down where memory
>> needs to be freed and hence a TLB flush is required after clearing out page
>> table entry.
>>
>> Cc: Catalin Marinas <catalin.marinas@arm.com>
>> Cc: Will Deacon <will@kernel.org>
>> Cc: linux-arm-kernel@lists.infradead.org
>> Cc: linux-kernel@vger.kernel.org
>> Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
>> Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
>> Cc: stable@vger.kernel.org
>> Signed-off-by: Ryan Roberts <ryan.roberts@arm.com>
>> Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
>
> I suggested the original shape of this and I see you have added my SOB. Final
> patch looks good to me - I'm not sure if it's correct for me to add Rb, but here
> it is regardless:
>
> Reviewed-by: Ryan Roberts <ryan.roberts@arm.com>
Thanks !
>
>
>> ---
>> arch/arm64/mm/mmu.c | 81 +++++++++++++++++++++++++++++++++++++--------
>> 1 file changed, 67 insertions(+), 14 deletions(-)
>>
>> diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
>> index 8e1d80a7033e..8ec8a287aaa1 100644
>> --- a/arch/arm64/mm/mmu.c
>> +++ b/arch/arm64/mm/mmu.c
>> @@ -1458,10 +1458,32 @@ static void unmap_hotplug_pte_range(pmd_t *pmdp, unsigned long addr,
>>
>> WARN_ON(!pte_present(pte));
>> __pte_clear(&init_mm, addr, ptep);
>> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
>> - if (free_mapped)
>> + if (free_mapped) {
>> + /*
>> + * If page is part of an existing contiguous
>> + * memory block, individual TLB invalidation
>> + * here would not be appropriate. Instead it
>> + * will require clearing all entries for the
>> + * memory block and subsequently a TLB flush
>> + * for the entire range.
>> + */
>> + WARN_ON(pte_cont(pte));
>> +
>> + /*
>> + * TLB flush is essential for freeing memory.
>> + */
>> + flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
>> free_hotplug_page_range(pte_page(pte),
>> PAGE_SIZE, altmap);
>> + }
>> +
>> + /*
>> + * TLB flush is batched in unmap_hotplug_range()
>> + * for the entire range, when memory need not be
>> + * freed. Besides linear mapping might have CONT
>> + * blocks where TLB flush needs to be done after
>> + * clearing all relevant entries.
>> + */
>> } while (addr += PAGE_SIZE, addr < end);
>> }
>>
>> @@ -1482,15 +1504,32 @@ static void unmap_hotplug_pmd_range(pud_t *pudp, unsigned long addr,
>> WARN_ON(!pmd_present(pmd));
>> if (pmd_sect(pmd)) {
>> pmd_clear(pmdp);
>> + if (free_mapped) {
>> + /*
>> + * If page is part of an existing contiguous
>> + * memory block, individual TLB invalidation
>> + * here would not be appropriate. Instead it
>> + * will require clearing all entries for the
>> + * memory block and subsequently a TLB flush
>> + * for the entire range.
>> + */
>> + WARN_ON(pmd_cont(pmd));
>> +
>> + /*
>> + * TLB flush is essential for freeing memory.
>> + */
>> + flush_tlb_kernel_range(addr, addr + PMD_SIZE);
>> + free_hotplug_page_range(pmd_page(pmd),
>> + PMD_SIZE, altmap);
>> + }
>>
>> /*
>> - * One TLBI should be sufficient here as the PMD_SIZE
>> - * range is mapped with a single block entry.
>> + * TLB flush is batched in unmap_hotplug_range()
>> + * for the entire range, when memory need not be
>> + * freed. Besides linear mapping might have CONT
>> + * blocks where TLB flush needs to be done after
>> + * clearing all relevant entries.
>> */
>> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
>> - if (free_mapped)
>> - free_hotplug_page_range(pmd_page(pmd),
>> - PMD_SIZE, altmap);
>> continue;
>> }
>> WARN_ON(!pmd_table(pmd));
>> @@ -1515,15 +1554,20 @@ static void unmap_hotplug_pud_range(p4d_t *p4dp, unsigned long addr,
>> WARN_ON(!pud_present(pud));
>> if (pud_sect(pud)) {
>> pud_clear(pudp);
>> + if (free_mapped) {
>> + /*
>> + * TLB flush is essential for freeing memory.
>> + */
>> + flush_tlb_kernel_range(addr, addr + PUD_SIZE);
>> + free_hotplug_page_range(pud_page(pud),
>> + PUD_SIZE, altmap);
>> + }
>>
>> /*
>> - * One TLBI should be sufficient here as the PUD_SIZE
>> - * range is mapped with a single block entry.
>> + * TLB flush is batched in unmap_hotplug_range()
>> + * for the entire range, when memory need not be
>> + * freed.
>> */
>> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
>> - if (free_mapped)
>> - free_hotplug_page_range(pud_page(pud),
>> - PUD_SIZE, altmap);
>> continue;
>> }
>> WARN_ON(!pud_table(pud));
>> @@ -1553,6 +1597,7 @@ static void unmap_hotplug_p4d_range(pgd_t *pgdp, unsigned long addr,
>> static void unmap_hotplug_range(unsigned long addr, unsigned long end,
>> bool free_mapped, struct vmem_altmap *altmap)
>> {
>> + unsigned long start = addr;
>> unsigned long next;
>> pgd_t *pgdp, pgd;
>>
>> @@ -1574,6 +1619,14 @@ static void unmap_hotplug_range(unsigned long addr, unsigned long end,
>> WARN_ON(!pgd_present(pgd));
>> unmap_hotplug_p4d_range(pgdp, addr, next, free_mapped, altmap);
>> } while (addr = next, addr < end);
>> +
>> + /*
>> + * Batched TLB flush only for linear mapping which
>> + * might contain CONT blocks, and does not require
>> + * freeing up memory as well.
>> + */
>> + if (!free_mapped)
>> + flush_tlb_kernel_range(start, end);
>> }
>>
>> static void free_empty_pte_table(pmd_t *pmdp, unsigned long addr,
>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping
2026-02-02 9:42 ` Ryan Roberts
@ 2026-02-02 11:06 ` Anshuman Khandual
2026-02-02 11:35 ` Ryan Roberts
0 siblings, 1 reply; 8+ messages in thread
From: Anshuman Khandual @ 2026-02-02 11:06 UTC (permalink / raw)
To: Ryan Roberts, linux-arm-kernel
Cc: Catalin Marinas, Will Deacon, Yang Shi, Christoph Lameter,
linux-kernel, stable
On 02/02/26 3:12 PM, Ryan Roberts wrote:
> On 02/02/2026 04:26, Anshuman Khandual wrote:
>> Linear and vmemmap mapings that get teared down during a memory hot remove
>> operation might contain leaf level entries on any page table level. If the
>> requested memory range's linear or vmemmap mappings falls within such leaf
>> entries, new mappings need to be created for the remaning memory mapped on
>> the leaf entry earlier, following standard break before make aka BBM rules.
>
> I think it would be good to mention that the kernel cannot tolerate BBM so
> remapping to fine grained leaves would not be possible on systems without
> BBML2_NOABORT.
Sure will add that.
>
>>
>> Currently memory hot remove operation does not perform such restructuring,
>> and so removing memory ranges that could split a kernel leaf level mapping
>> need to be rejected.
>
> Perhaps it is useful to mention that while memory_hotplug.c does appear to
> permit hot-unplugging arbitrary ranges of memory, the higher layers that drive
> memory_hotplug (e.g. ACPI, virtio, ...) all appear to treat memory as fixed size
> devices so it is impossible to hotunplug a different amount than was previously
> hotplugged, and so we should never see a rejection in practice, but adding the
> check makes us robust against a future change.
Agreed, will update the commit message.
>
>>
>> Cc: Catalin Marinas <catalin.marinas@arm.com>
>> Cc: Will Deacon <will@kernel.org>
>> Cc: linux-arm-kernel@lists.infradead.org
>> Cc: linux-kernel@vger.kernel.org
>> Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
>> Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
>> Cc: stable@vger.kernel.org
>> Suggested-by: Ryan Roberts <ryan.roberts@arm.com>
>> Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
>> ---
>> arch/arm64/mm/mmu.c | 126 ++++++++++++++++++++++++++++++++++++++++++++
>> 1 file changed, 126 insertions(+)
>>
>> diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
>> index 8ec8a287aaa1..9d59e10fb3de 100644
>> --- a/arch/arm64/mm/mmu.c
>> +++ b/arch/arm64/mm/mmu.c
>> @@ -2063,6 +2063,129 @@ void arch_remove_memory(u64 start, u64 size, struct vmem_altmap *altmap)
>> __remove_pgd_mapping(swapper_pg_dir, __phys_to_virt(start), size);
>> }
>>
>> +
>> +static bool split_kernel_leaf_boundary(unsigned long addr)
>
> The name currently makes it sound like we are asking for the mapping to be split
> (we have existing functions to do this that are named similarly). Perhaps a
> better name would be addr_splits_leaf()?
Agreed that name sounds bit confusing and ambiguous as there is already
a similarly named function. Will rename it as addr_splits_kernel_leaf()
instead.
>
>> +{
>> + pgd_t *pgdp, pgd;
>> + p4d_t *p4dp, p4d;
>> + pud_t *pudp, pud;
>> + pmd_t *pmdp, pmd;
>> + pte_t *ptep, pte;
>> +
>> + /*
>> + * PGD: If addr is PGD aligned then addr already
>> + * describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, PGDIR_SIZE) == addr)
>> + return false;
>> +
>> + pgdp = pgd_offset_k(addr);
>> + pgd = pgdp_get(pgdp);
>> + if (!pgd_present(pgd))
>> + return false;
>> +
>> + /*
>> + * P4D: If addr is P4D aligned then addr already
>> + * describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, P4D_SIZE) == addr)
>> + return false;
>> +
>> + p4dp = p4d_offset(pgdp, addr);
>> + p4d = p4dp_get(p4dp);
>> + if (!p4d_present(p4d))
>> + return false;
>> +
>> + /*
>> + * PUD: If addr is PUD aligned then addr already
>> + * describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, PUD_SIZE) == addr)
>> + return false;
>> +
>> + pudp = pud_offset(p4dp, addr);
>> + pud = pudp_get(pudp);
>> + if (!pud_present(pud))
>> + return false;
>> +
>> + if (pud_leaf(pud))
>> + return true;
>> +
>> + /*
>> + * CONT_PMD: If addr is CONT_PMD aligned then
>> + * addr already describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, CONT_PMD_SIZE) == addr)
>> + return false;
>> +
>> + pmdp = pmd_offset(pudp, addr);
>> + pmd = pmdp_get(pmdp);
>> + if (!pmd_present(pmd))
>> + return false;
>> +
>> + if (pmd_leaf(pmd) && pmd_cont(pmd))
>> + return true;
>> +
>> + /*
>> + * PMD: If addr is PMD aligned then addr already
>> + * describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, PMD_SIZE) == addr)
>> + return false;
>> +
>> + if (pmd_leaf(pmd))
>> + return true;
>> +
>> + /*
>> + * CONT_PTE: If addr is CONT_PTE aligned then addr
>> + * already describes a leaf boundary.
>> + */
>> + if (ALIGN_DOWN(addr, CONT_PTE_SIZE) == addr)
>> + return false;
>> +
>> + ptep = pte_offset_kernel(pmdp, addr);
>> + pte = __ptep_get(ptep);
>> + if (!pte_present(pte))
>> + return false;
>> +
>> + if (pte_valid(pte) && pte_cont(pte))
>
> Why do you need pte_valid() here? You have already checked !pte_present(). Are
> you expecting a case of present but not valid (PTE_PRESENT_INVALID)? If so, do
> you need to consider that for the other levels too? (pmd_leaf() only checks
> pmd_present()).
> > Personally I think you can just drop the pte_valid() check here.
Added pte_valid() for abundance of caution but it is not really
necessary though. Sure will drop it off.
>
>> + return true;
>> +
>> + if (ALIGN_DOWN(addr, PAGE_SIZE) == addr)
>> + return false;
>> + return true;
>> +}
>> +
>> +static bool can_unmap_without_split(unsigned long pfn, unsigned long nr_pages)
>> +{
>> + unsigned long linear_start, linear_end, phys_start, phys_end;
>> + unsigned long vmemmap_size, vmemmap_start, vmemmap_end;
>
> nit: do we need all these variables. Perhaps just:
>
> unsigned long sz, start, end, phys_start, phys_end;
>
> are sufficient?
Alright. I guess start and end can be re-used both for linear and
vmemmap mapping.
>
>> +
>> + /* Assert linear map edges do not split a leaf entry */
>> + phys_start = PFN_PHYS(pfn);
>> + phys_end = phys_start + nr_pages * PAGE_SIZE;
>> + linear_start = __phys_to_virt(phys_start);
>> + linear_end = __phys_to_virt(phys_end);
>> + if (split_kernel_leaf_boundary(linear_start) ||
>> + split_kernel_leaf_boundary(linear_end)) {
>> + pr_warn("[%lx %lx] splits a leaf entry in linear map\n",
>> + phys_start, phys_end);
>> + return false;
>> + }
>> +
>> + /* Assert vmemmap edges do not split a leaf entry */
>> + vmemmap_size = nr_pages * sizeof(struct page);
>> + vmemmap_start = (unsigned long) pfn_to_page(pfn);
>
> nit: ^
>
> I don't think we would normally have that space?
Sure will drop that.
>
>> + vmemmap_end = vmemmap_start + vmemmap_size;
>> + if (split_kernel_leaf_boundary(vmemmap_start) ||
>> + split_kernel_leaf_boundary(vmemmap_end)) {
>> + pr_warn("[%lx %lx] splits a leaf entry in vmemmap\n",
>> + phys_start, phys_end);
>> + return false;
>> + }
>> + return true;
>> +}
>> +
>> /*
>> * This memory hotplug notifier helps prevent boot memory from being
>> * inadvertently removed as it blocks pfn range offlining process in
>> @@ -2083,6 +2206,9 @@ static int prevent_bootmem_remove_notifier(struct notifier_block *nb,
>> if ((action != MEM_GOING_OFFLINE) && (action != MEM_OFFLINE))
>> return NOTIFY_OK;
>>
>> + if (!can_unmap_without_split(pfn, arg->nr_pages))
>> + return NOTIFY_BAD;
>> +
>
> Personally, I'd keep the bootmem check first and do this check after. That means
> an existing warning will not change.
Makes sense, will move it after existing bootmem check. BTW the function
still named as prevent_bootmem_remove_notifier() although now it's going
to check leaf boundaries as well. Should the function be renamed as well
to something more generic e.g prevent_memory_remove_notifier() ?
>
> Thanks,
> Ryan
>
>> for (; pfn < end_pfn; pfn += PAGES_PER_SECTION) {
>> unsigned long start = PFN_PHYS(pfn);
>> unsigned long end = start + (1UL << PA_SECTION_SHIFT);
>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping
2026-02-02 11:06 ` Anshuman Khandual
@ 2026-02-02 11:35 ` Ryan Roberts
0 siblings, 0 replies; 8+ messages in thread
From: Ryan Roberts @ 2026-02-02 11:35 UTC (permalink / raw)
To: Anshuman Khandual, linux-arm-kernel
Cc: Catalin Marinas, Will Deacon, Yang Shi, Christoph Lameter,
linux-kernel, stable
On 02/02/2026 11:06, Anshuman Khandual wrote:
> On 02/02/26 3:12 PM, Ryan Roberts wrote:
>> On 02/02/2026 04:26, Anshuman Khandual wrote:
>>> Linear and vmemmap mapings that get teared down during a memory hot remove
>>> operation might contain leaf level entries on any page table level. If the
>>> requested memory range's linear or vmemmap mappings falls within such leaf
>>> entries, new mappings need to be created for the remaning memory mapped on
>>> the leaf entry earlier, following standard break before make aka BBM rules.
>>
>> I think it would be good to mention that the kernel cannot tolerate BBM so
>> remapping to fine grained leaves would not be possible on systems without
>> BBML2_NOABORT.
>
> Sure will add that.
>
>>
>>>
>>> Currently memory hot remove operation does not perform such restructuring,
>>> and so removing memory ranges that could split a kernel leaf level mapping
>>> need to be rejected.
>>
>> Perhaps it is useful to mention that while memory_hotplug.c does appear to
>> permit hot-unplugging arbitrary ranges of memory, the higher layers that drive
>> memory_hotplug (e.g. ACPI, virtio, ...) all appear to treat memory as fixed size
>> devices so it is impossible to hotunplug a different amount than was previously
>> hotplugged, and so we should never see a rejection in practice, but adding the
>> check makes us robust against a future change.
>
> Agreed, will update the commit message.
>
>>
>>>
>>> Cc: Catalin Marinas <catalin.marinas@arm.com>
>>> Cc: Will Deacon <will@kernel.org>
>>> Cc: linux-arm-kernel@lists.infradead.org
>>> Cc: linux-kernel@vger.kernel.org
>>> Closes: https://lore.kernel.org/all/aWZYXhrT6D2M-7-N@willie-the-truck/
>>> Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove")
>>> Cc: stable@vger.kernel.org
>>> Suggested-by: Ryan Roberts <ryan.roberts@arm.com>
>>> Signed-off-by: Anshuman Khandual <anshuman.khandual@arm.com>
>>> ---
>>> arch/arm64/mm/mmu.c | 126 ++++++++++++++++++++++++++++++++++++++++++++
>>> 1 file changed, 126 insertions(+)
>>>
>>> diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
>>> index 8ec8a287aaa1..9d59e10fb3de 100644
>>> --- a/arch/arm64/mm/mmu.c
>>> +++ b/arch/arm64/mm/mmu.c
>>> @@ -2063,6 +2063,129 @@ void arch_remove_memory(u64 start, u64 size, struct vmem_altmap *altmap)
>>> __remove_pgd_mapping(swapper_pg_dir, __phys_to_virt(start), size);
>>> }
>>>
>>> +
>>> +static bool split_kernel_leaf_boundary(unsigned long addr)
>>
>> The name currently makes it sound like we are asking for the mapping to be split
>> (we have existing functions to do this that are named similarly). Perhaps a
>> better name would be addr_splits_leaf()?
>
> Agreed that name sounds bit confusing and ambiguous as there is already
> a similarly named function. Will rename it as addr_splits_kernel_leaf()
> instead.
>
>>
>>> +{
>>> + pgd_t *pgdp, pgd;
>>> + p4d_t *p4dp, p4d;
>>> + pud_t *pudp, pud;
>>> + pmd_t *pmdp, pmd;
>>> + pte_t *ptep, pte;
>>> +
>>> + /*
>>> + * PGD: If addr is PGD aligned then addr already
>>> + * describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, PGDIR_SIZE) == addr)
>>> + return false;
>>> +
>>> + pgdp = pgd_offset_k(addr);
>>> + pgd = pgdp_get(pgdp);
>>> + if (!pgd_present(pgd))
>>> + return false;
>>> +
>>> + /*
>>> + * P4D: If addr is P4D aligned then addr already
>>> + * describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, P4D_SIZE) == addr)
>>> + return false;
>>> +
>>> + p4dp = p4d_offset(pgdp, addr);
>>> + p4d = p4dp_get(p4dp);
>>> + if (!p4d_present(p4d))
>>> + return false;
>>> +
>>> + /*
>>> + * PUD: If addr is PUD aligned then addr already
>>> + * describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, PUD_SIZE) == addr)
>>> + return false;
>>> +
>>> + pudp = pud_offset(p4dp, addr);
>>> + pud = pudp_get(pudp);
>>> + if (!pud_present(pud))
>>> + return false;
>>> +
>>> + if (pud_leaf(pud))
>>> + return true;
>>> +
>>> + /*
>>> + * CONT_PMD: If addr is CONT_PMD aligned then
>>> + * addr already describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, CONT_PMD_SIZE) == addr)
>>> + return false;
>>> +
>>> + pmdp = pmd_offset(pudp, addr);
>>> + pmd = pmdp_get(pmdp);
>>> + if (!pmd_present(pmd))
>>> + return false;
>>> +
>>> + if (pmd_leaf(pmd) && pmd_cont(pmd))
>>> + return true;
>>> +
>>> + /*
>>> + * PMD: If addr is PMD aligned then addr already
>>> + * describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, PMD_SIZE) == addr)
>>> + return false;
>>> +
>>> + if (pmd_leaf(pmd))
>>> + return true;
>>> +
>>> + /*
>>> + * CONT_PTE: If addr is CONT_PTE aligned then addr
>>> + * already describes a leaf boundary.
>>> + */
>>> + if (ALIGN_DOWN(addr, CONT_PTE_SIZE) == addr)
>>> + return false;
>>> +
>>> + ptep = pte_offset_kernel(pmdp, addr);
>>> + pte = __ptep_get(ptep);
>>> + if (!pte_present(pte))
>>> + return false;
>>> +
>>> + if (pte_valid(pte) && pte_cont(pte))
>>
>> Why do you need pte_valid() here? You have already checked !pte_present(). Are
>> you expecting a case of present but not valid (PTE_PRESENT_INVALID)? If so, do
>> you need to consider that for the other levels too? (pmd_leaf() only checks
>> pmd_present()).
>>> Personally I think you can just drop the pte_valid() check here.
>
> Added pte_valid() for abundance of caution but it is not really
> necessary though. Sure will drop it off.
>
>>
>>> + return true;
>>> +
>>> + if (ALIGN_DOWN(addr, PAGE_SIZE) == addr)
>>> + return false;
>>> + return true;
>>> +}
>>> +
>>> +static bool can_unmap_without_split(unsigned long pfn, unsigned long nr_pages)
>>> +{
>>> + unsigned long linear_start, linear_end, phys_start, phys_end;
>>> + unsigned long vmemmap_size, vmemmap_start, vmemmap_end;
>>
>> nit: do we need all these variables. Perhaps just:
>>
>> unsigned long sz, start, end, phys_start, phys_end;
>>
>> are sufficient?
>
> Alright. I guess start and end can be re-used both for linear and
> vmemmap mapping.
>
>>
>>> +
>>> + /* Assert linear map edges do not split a leaf entry */
>>> + phys_start = PFN_PHYS(pfn);
>>> + phys_end = phys_start + nr_pages * PAGE_SIZE;
>>> + linear_start = __phys_to_virt(phys_start);
>>> + linear_end = __phys_to_virt(phys_end);
>>> + if (split_kernel_leaf_boundary(linear_start) ||
>>> + split_kernel_leaf_boundary(linear_end)) {
>>> + pr_warn("[%lx %lx] splits a leaf entry in linear map\n",
>>> + phys_start, phys_end);
>>> + return false;
>>> + }
>>> +
>>> + /* Assert vmemmap edges do not split a leaf entry */
>>> + vmemmap_size = nr_pages * sizeof(struct page);
>>> + vmemmap_start = (unsigned long) pfn_to_page(pfn);
>>
>> nit: ^
>>
>> I don't think we would normally have that space?
>
> Sure will drop that.
>
>>
>>> + vmemmap_end = vmemmap_start + vmemmap_size;
>>> + if (split_kernel_leaf_boundary(vmemmap_start) ||
>>> + split_kernel_leaf_boundary(vmemmap_end)) {
>>> + pr_warn("[%lx %lx] splits a leaf entry in vmemmap\n",
>>> + phys_start, phys_end);
>>> + return false;
>>> + }
>>> + return true;
>>> +}
>>> +
>>> /*
>>> * This memory hotplug notifier helps prevent boot memory from being
>>> * inadvertently removed as it blocks pfn range offlining process in
>>> @@ -2083,6 +2206,9 @@ static int prevent_bootmem_remove_notifier(struct notifier_block *nb,
>>> if ((action != MEM_GOING_OFFLINE) && (action != MEM_OFFLINE))
>>> return NOTIFY_OK;
>>>
>>> + if (!can_unmap_without_split(pfn, arg->nr_pages))
>>> + return NOTIFY_BAD;
>>> +
>>
>> Personally, I'd keep the bootmem check first and do this check after. That means
>> an existing warning will not change.
>
> Makes sense, will move it after existing bootmem check. BTW the function
> still named as prevent_bootmem_remove_notifier() although now it's going
> to check leaf boundaries as well. Should the function be renamed as well
> to something more generic e.g prevent_memory_remove_notifier() ?
Works for me.
>
>>
>> Thanks,
>> Ryan
>>
>>> for (; pfn < end_pfn; pfn += PAGES_PER_SECTION) {
>>> unsigned long start = PFN_PHYS(pfn);
>>> unsigned long end = start + (1UL << PA_SECTION_SHIFT);
>>
>
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-02-02 11:35 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-02-02 4:26 [PATCH 0/2] arm64/mm: Enable batched TLB flush in unmap_hotplug_range() Anshuman Khandual
2026-02-02 4:26 ` [PATCH 1/2] " Anshuman Khandual
2026-02-02 9:18 ` Ryan Roberts
2026-02-02 10:48 ` Anshuman Khandual
2026-02-02 4:26 ` [PATCH 2/2] arm64/mm: Reject memory removal that splits a kernel leaf mapping Anshuman Khandual
2026-02-02 9:42 ` Ryan Roberts
2026-02-02 11:06 ` Anshuman Khandual
2026-02-02 11:35 ` Ryan Roberts
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®