* [PATCH v3 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-30 4:24 ` Lance Yang
2026-09-29 5:32 ` [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path Muchun Song
` (5 subsequent siblings)
6 siblings, 1 reply; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in
two places: it keeps allocations on the normal path and takes a reference
when a backing page is supplied for reuse.
After Device DAX switched to the common per-zone shared tail page, both
conditions can be determined locally. DAX supplies ptpfn for every shared
tail mapping and requests an allocation only for compound head mappings,
whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
selects the correct allocation path.
When ptpfn is supplied, the caller is reusing an existing backing page.
Once the slab allocator is available, take a reference for each reused
mapping to balance the release performed by vmemmap_free(). Although the
buddy allocator is available before slab, no vmemmap population occurs in
that interval. Earlier mappings are backed by memblock/reserved memory and
do not need page reference accounting.
Remove VMEMMAP_POPULATE_DAX and the flags argument from the vmemmap
population helpers.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v2:
- Explain why the buddy-to-slab initialization gap is safe (suggested by
Qi Zheng)
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 36 +++++++++++++-----------------------
1 file changed, 13 insertions(+), 23 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 111efc802465..8219abc6c3e5 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -32,12 +32,6 @@
#include <asm/dma.h>
#include <asm/tlbflush.h>
-/*
- * Flags for vmemmap_populate_range and friends.
- */
-/* Vmemmap population for ZONE_DEVICE compound pages */
-#define VMEMMAP_POPULATE_DAX 0x0001
-
#include "internal.h"
#include "mm_init.h"
#include "sparse.h"
@@ -237,7 +231,7 @@ struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zon
#endif
static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
- struct vmem_altmap *altmap, unsigned long flags)
+ struct vmem_altmap *altmap)
{
struct zone *zone;
struct page *page;
@@ -247,7 +241,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
* Device DAX still relies on vmemmap_populate_compound_pages() for
* head/first-tail allocation and tail-page reuse.
*/
- if (!vmemmap_optimizable_pfn(pfn) || flags & VMEMMAP_POPULATE_DAX)
+ if (!vmemmap_optimizable_pfn(pfn))
return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
zone = pfn_to_zone(pfn, node);
@@ -259,8 +253,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
}
static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
- struct vmem_altmap *altmap,
- unsigned long ptpfn, unsigned long flags)
+ struct vmem_altmap *altmap, unsigned long ptpfn)
{
pte_t *pte = pte_offset_kernel(pmd, addr);
unsigned long pfn = page_to_pfn((struct page *)addr);
@@ -269,7 +262,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
pte_t entry;
if (ptpfn == (unsigned long)-1) {
- void *p = vmemmap_alloc_pte(pfn, node, altmap, flags);
+ void *p = vmemmap_alloc_pte(pfn, node, altmap);
if (!p)
return NULL;
@@ -287,7 +280,7 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
* Use try_get_page() to prevent the shared page refcount
* from overflowing.
*/
- if ((flags & VMEMMAP_POPULATE_DAX) &&
+ if (slab_is_available() &&
!try_get_page(pfn_to_page(ptpfn)))
return NULL;
}
@@ -351,8 +344,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
struct vmem_altmap *altmap,
- unsigned long ptpfn,
- unsigned long flags)
+ unsigned long ptpfn)
{
pgd_t *pgd;
p4d_t *p4d;
@@ -372,7 +364,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
pmd = vmemmap_pmd_populate(pud, addr, node);
if (!pmd)
return NULL;
- pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn, flags);
+ pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn);
if (!pte)
return NULL;
vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
@@ -383,15 +375,14 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
static int __meminit vmemmap_populate_range(unsigned long start,
unsigned long end, int node,
struct vmem_altmap *altmap,
- unsigned long ptpfn,
- unsigned long flags)
+ unsigned long ptpfn)
{
unsigned long addr = start;
pte_t *pte;
for (; addr < end; addr += PAGE_SIZE) {
pte = vmemmap_populate_address(addr, node, altmap,
- ptpfn, flags);
+ ptpfn);
if (!pte)
return -ENOMEM;
}
@@ -402,7 +393,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
int node, struct vmem_altmap *altmap)
{
- return vmemmap_populate_range(start, end, node, altmap, -1, 0);
+ return vmemmap_populate_range(start, end, node, altmap, -1);
}
/*
@@ -531,7 +522,6 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
unsigned long size, addr;
pte_t *pte;
int rc;
- unsigned long flags = VMEMMAP_POPULATE_DAX;
struct page *page;
unsigned int order = pfn_to_section_compound_order(start_pfn);
@@ -541,14 +531,14 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
if (reuse_compound_section(start_pfn, pgmap))
return vmemmap_populate_range(start, end, node, NULL,
- page_to_pfn(page), flags);
+ page_to_pfn(page));
size = min(end - start, (1UL << order) * sizeof(struct page));
for (addr = start; addr < end; addr += size) {
unsigned long next, last = addr + size;
/* Populate the head page vmemmap page */
- pte = vmemmap_populate_address(addr, node, NULL, -1, flags);
+ pte = vmemmap_populate_address(addr, node, NULL, -1);
if (!pte)
return -ENOMEM;
@@ -558,7 +548,7 @@ static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
*/
next = addr + PAGE_SIZE;
rc = vmemmap_populate_range(next, last, node, NULL,
- page_to_pfn(page), flags);
+ page_to_pfn(page));
if (rc)
return -ENOMEM;
}
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX
2026-09-29 5:32 ` [PATCH v3 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX Muchun Song
@ 2026-09-30 4:24 ` Lance Yang
0 siblings, 0 replies; 15+ messages in thread
From: Lance Yang @ 2026-09-30 4:24 UTC (permalink / raw)
To: songmuchun
Cc: maddy, rppt, akpm, david, mpe, npiggin, chleroy, ritesh.list,
sshegde, ljs, liam, vbabka, surenb, mhocko, qi.zheng,
linuxppc-dev, linux-kernel, linux-mm, muchun.song, Lance Yang
On Tue, Sep 29, 2026 at 01:32:26PM +0800, Muchun Song wrote:
>VMEMMAP_POPULATE_DAX currently distinguishes DAX vmemmap population in
>two places: it keeps allocations on the normal path and takes a reference
>when a backing page is supplied for reuse.
>
>After Device DAX switched to the common per-zone shared tail page, both
>conditions can be determined locally. DAX supplies ptpfn for every shared
>tail mapping and requests an allocation only for compound head mappings,
>whose PFNs are not optimizable. Therefore, vmemmap_optimizable_pfn() alone
>selects the correct allocation path.
>
>When ptpfn is supplied, the caller is reusing an existing backing page.
>Once the slab allocator is available, take a reference for each reused
>mapping to balance the release performed by vmemmap_free(). Although the
>buddy allocator is available before slab, no vmemmap population occurs in
>that interval. Earlier mappings are backed by memblock/reserved memory and
>do not need page reference accounting.
>
>Remove VMEMMAP_POPULATE_DAX and the flags argument from the vmemmap
>population helpers.
>
>Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>Acked-by: Qi Zheng <qi.zheng@linux.dev>
>---
LGTM! Feel free to add:
Reviewed-by: Lance Yang <lance.yang@linux.dev>
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
2026-09-29 5:32 ` [PATCH v3 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-30 8:53 ` Lance Yang
2026-09-30 10:02 ` Lance Yang
2026-09-29 5:32 ` [PATCH v3 3/6] mm/sparse-vmemmap: drop Device DAX-specific population path Muchun Song
` (4 subsequent siblings)
6 siblings, 2 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
The common vmemmap population path cannot yet handle optimized Device DAX
mappings on its own. It uses pfn_to_zone() to find the shared tail page,
but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
is initialized.
Teach the common path to use device_zone() for runtime optimized vmemmap
population while retaining pfn_to_zone() for early boot. This allows the
same path to support both early boot mappings and Device DAX.
The backing PFN supplied by the Device DAX-specific population path is no
longer used, allowing the redundant lookup and population code to be
removed later.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
v2:
- Expand comments around slab initialization to explain zone lookup and
page refcounting (suggested by Qi Zheng)
---
mm/sparse-vmemmap.c | 64 +++++++++++++++++++++++++--------------------
1 file changed, 35 insertions(+), 29 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 8219abc6c3e5..ee4c113ca938 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -237,18 +237,43 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
struct page *page;
const unsigned int order = pfn_to_section_compound_order(pfn);
- /*
- * Device DAX still relies on vmemmap_populate_compound_pages() for
- * head/first-tail allocation and tail-page reuse.
- */
if (!vmemmap_optimizable_pfn(pfn))
return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
- zone = pfn_to_zone(pfn, node);
+ /*
+ * Before slab is available, vmemmap optimization is used for early
+ * system RAM, whose zone can be determined from the PFN.
+ *
+ * Once slab is available, only ZONE_DEVICE memory reaches this
+ * optimized population path. Its zone span has not been initialized
+ * while its vmemmap is being populated, so pfn_to_zone() cannot be
+ * used. Obtain ZONE_DEVICE directly from the node instead.
+ */
+ zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
page = vmemmap_shared_tail_page(order, zone);
if (!page)
return NULL;
+ /*
+ * During early vmemmap population, the shared tail vmemmap backing
+ * page is allocated from memblock before its struct page can safely
+ * participate in page refcounting. Therefore, no reference can be
+ * held for each shared PTE mapping, and the mappings must be unshared
+ * before the vmemmap is depopulated.
+ *
+ * Once slab is available, the shared backing page is allocated from
+ * the buddy allocator and can be refcounted. Hold one reference for
+ * each shared PTE mapping. The architecture vmemmap teardown drops
+ * the reference through __free_pages() when removing the mapping,
+ * preventing the backing page from being freed while it is shared.
+ *
+ * The backing page may be shared by enough PTE mappings to exhaust
+ * the positive range of its reference count. Stop populating the
+ * vmemmap if another reference cannot be acquired.
+ */
+ if (slab_is_available() && !try_get_page(page))
+ return NULL;
+
return page_address(page);
}
@@ -260,31 +285,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
if (pte_none(ptep_get(pte))) {
pte_t entry;
+ void *p = vmemmap_alloc_pte(pfn, node, altmap);
- if (ptpfn == (unsigned long)-1) {
- void *p = vmemmap_alloc_pte(pfn, node, altmap);
-
- if (!p)
- return NULL;
- ptpfn = PHYS_PFN(__pa(p));
- } else {
- /*
- * When a PTE/PMD entry is freed from the init_mm
- * there's a free_pages() call to this page allocated
- * above. Thus this try_get_page() is paired with the
- * put_page_testzero() on the freeing path.
- * This can only called by certain ZONE_DEVICE path,
- * and through vmemmap_populate_compound_pages() when
- * slab is available.
- *
- * Use try_get_page() to prevent the shared page refcount
- * from overflowing.
- */
- if (slab_is_available() &&
- !try_get_page(pfn_to_page(ptpfn)))
- return NULL;
- }
- entry = pfn_pte(ptpfn, PAGE_KERNEL);
+ if (!p)
+ return NULL;
+
+ entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
set_pte_at(&init_mm, addr, pte, entry);
} else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
return NULL;
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-29 5:32 ` [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path Muchun Song
@ 2026-09-30 8:53 ` Lance Yang
2026-09-30 10:09 ` Muchun Song
2026-09-30 11:11 ` Muchun Song
2026-09-30 10:02 ` Lance Yang
1 sibling, 2 replies; 15+ messages in thread
From: Lance Yang @ 2026-09-30 8:53 UTC (permalink / raw)
To: songmuchun
Cc: maddy, rppt, akpm, david, mpe, npiggin, chleroy, ritesh.list,
sshegde, ljs, liam, vbabka, surenb, mhocko, qi.zheng,
linuxppc-dev, linux-kernel, linux-mm, muchun.song, Lance Yang
On Tue, Sep 29, 2026 at 01:32:27PM +0800, Muchun Song wrote:
>The common vmemmap population path cannot yet handle optimized Device DAX
>mappings on its own. It uses pfn_to_zone() to find the shared tail page,
>but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
>is initialized.
>
>Teach the common path to use device_zone() for runtime optimized vmemmap
>population while retaining pfn_to_zone() for early boot. This allows the
>same path to support both early boot mappings and Device DAX.
>
>The backing PFN supplied by the Device DAX-specific population path is no
>longer used, allowing the redundant lookup and population code to be
>removed later.
>
>Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>Acked-by: Qi Zheng <qi.zheng@linux.dev>
>---
>v3:
>- Collect Acked-by from Qi Zheng
>
>v2:
>- Expand comments around slab initialization to explain zone lookup and
> page refcounting (suggested by Qi Zheng)
>---
> mm/sparse-vmemmap.c | 64 +++++++++++++++++++++++++--------------------
> 1 file changed, 35 insertions(+), 29 deletions(-)
>
>diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
>index 8219abc6c3e5..ee4c113ca938 100644
>--- a/mm/sparse-vmemmap.c
>+++ b/mm/sparse-vmemmap.c
>@@ -237,18 +237,43 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
> struct page *page;
> const unsigned int order = pfn_to_section_compound_order(pfn);
>
>- /*
>- * Device DAX still relies on vmemmap_populate_compound_pages() for
>- * head/first-tail allocation and tail-page reuse.
>- */
> if (!vmemmap_optimizable_pfn(pfn))
> return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
>
>- zone = pfn_to_zone(pfn, node);
>+ /*
>+ * Before slab is available, vmemmap optimization is used for early
>+ * system RAM, whose zone can be determined from the PFN.
>+ *
>+ * Once slab is available, only ZONE_DEVICE memory reaches this
>+ * optimized population path. Its zone span has not been initialized
>+ * while its vmemmap is being populated, so pfn_to_zone() cannot be
>+ * used. Obtain ZONE_DEVICE directly from the node instead.
>+ */
>+ zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
> page = vmemmap_shared_tail_page(order, zone);
> if (!page)
> return NULL;
>
>+ /*
>+ * During early vmemmap population, the shared tail vmemmap backing
>+ * page is allocated from memblock before its struct page can safely
>+ * participate in page refcounting. Therefore, no reference can be
>+ * held for each shared PTE mapping, and the mappings must be unshared
>+ * before the vmemmap is depopulated.
>+ *
>+ * Once slab is available, the shared backing page is allocated from
>+ * the buddy allocator and can be refcounted. Hold one reference for
>+ * each shared PTE mapping. The architecture vmemmap teardown drops
>+ * the reference through __free_pages() when removing the mapping,
>+ * preventing the backing page from being freed while it is shared.
>+ *
>+ * The backing page may be shared by enough PTE mappings to exhaust
>+ * the positive range of its reference count. Stop populating the
>+ * vmemmap if another reference cannot be acquired.
>+ */
>+ if (slab_is_available() && !try_get_page(page))
>+ return NULL;
BTW, shouldn't __add_pages() undo the earlier sections on a population
failure? Say the first section is added successfully, but populating the
next one fails, e.g. due to an allocation failure:
void *memremap_pages(struct dev_pagemap *pgmap, int nid)
{
...
const int nr_range = pgmap->nr_range;
int error, i;
...
pgmap->nr_range = 0;
error = 0;
for (i = 0; i < nr_range; i++) {
error = pagemap_range(pgmap, ¶ms, i, nid);
if (error)
break;
pgmap->nr_range++;
}
if (i < nr_range) {
memunmap_pages(pgmap);
pgmap->nr_range = nr_range;
return ERR_PTR(error);
}
...
}
We still need to undo the sections already added in the failed range,
though ... memunmap_pages() won't touch those, since it only removes
completed ranges.
If the range starts at a section boundary, we'd hit -EEXIST in
fill_subsection_map() on retry while those subsection bits are still set.
The old DAX path had this issue too. Could we roll back [start_pfn, pfn)
in __add_pages() as a separate fix? Something like this:
---8<---
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index 796af1028ee2..16a0a2c885bc 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -380,6 +380,7 @@ EXPORT_SYMBOL_GPL(pfn_to_online_page);
int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
struct mhp_params *params)
{
+ const unsigned long start_pfn = pfn;
const unsigned long end_pfn = pfn + nr_pages;
unsigned long cur_nr_pages;
int err;
@@ -417,6 +418,10 @@ int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
break;
cond_resched();
}
+
+ /* Roll back the sections added before the failure. */
+ if (err && pfn != start_pfn)
+ __remove_pages(start_pfn, pfn - start_pfn, altmap, params->pgmap);
vmemmap_populate_print_last();
return err;
}
--
Hope I haven't missed anything :)
Cheers, Lance
>+
> return page_address(page);
> }
>
>@@ -260,31 +285,12 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
>
> if (pte_none(ptep_get(pte))) {
> pte_t entry;
>+ void *p = vmemmap_alloc_pte(pfn, node, altmap);
>
>- if (ptpfn == (unsigned long)-1) {
>- void *p = vmemmap_alloc_pte(pfn, node, altmap);
>-
>- if (!p)
>- return NULL;
>- ptpfn = PHYS_PFN(__pa(p));
>- } else {
>- /*
>- * When a PTE/PMD entry is freed from the init_mm
>- * there's a free_pages() call to this page allocated
>- * above. Thus this try_get_page() is paired with the
>- * put_page_testzero() on the freeing path.
>- * This can only called by certain ZONE_DEVICE path,
>- * and through vmemmap_populate_compound_pages() when
>- * slab is available.
>- *
>- * Use try_get_page() to prevent the shared page refcount
>- * from overflowing.
>- */
>- if (slab_is_available() &&
>- !try_get_page(pfn_to_page(ptpfn)))
>- return NULL;
>- }
>- entry = pfn_pte(ptpfn, PAGE_KERNEL);
>+ if (!p)
>+ return NULL;
>+
>+ entry = pfn_pte(PHYS_PFN(__pa(p)), PAGE_KERNEL);
> set_pte_at(&init_mm, addr, pte, entry);
> } else if (WARN_ON_ONCE(vmemmap_optimizable_pfn(pfn)))
> return NULL;
>--
>2.54.0
>
>
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-30 8:53 ` Lance Yang
@ 2026-09-30 10:09 ` Muchun Song
2026-09-30 11:11 ` Muchun Song
1 sibling, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-30 10:09 UTC (permalink / raw)
To: Lance Yang
Cc: Muchun Song, maddy, rppt, akpm, david, mpe, npiggin, chleroy,
ritesh.list, sshegde, ljs, liam, vbabka, surenb, mhocko,
qi.zheng, linuxppc-dev, linux-kernel, linux-mm
> On Sep 30, 2026, at 16:53, Lance Yang <lance.yang@linux.dev> wrote:
>
>
> On Tue, Sep 29, 2026 at 01:32:27PM +0800, Muchun Song wrote:
>> The common vmemmap population path cannot yet handle optimized Device DAX
>> mappings on its own. It uses pfn_to_zone() to find the shared tail page,
>> but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
>> is initialized.
>>
>> Teach the common path to use device_zone() for runtime optimized vmemmap
>> population while retaining pfn_to_zone() for early boot. This allows the
>> same path to support both early boot mappings and Device DAX.
>>
>> The backing PFN supplied by the Device DAX-specific population path is no
>> longer used, allowing the redundant lookup and population code to be
>> removed later.
>>
>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>> Acked-by: Qi Zheng <qi.zheng@linux.dev>
>> ---
>> v3:
>> - Collect Acked-by from Qi Zheng
>>
>> v2:
>> - Expand comments around slab initialization to explain zone lookup and
>> page refcounting (suggested by Qi Zheng)
>> ---
>> mm/sparse-vmemmap.c | 64 +++++++++++++++++++++++++--------------------
>> 1 file changed, 35 insertions(+), 29 deletions(-)
>>
>> diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
>> index 8219abc6c3e5..ee4c113ca938 100644
>> --- a/mm/sparse-vmemmap.c
>> +++ b/mm/sparse-vmemmap.c
>> @@ -237,18 +237,43 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
>> struct page *page;
>> const unsigned int order = pfn_to_section_compound_order(pfn);
>>
>> - /*
>> - * Device DAX still relies on vmemmap_populate_compound_pages() for
>> - * head/first-tail allocation and tail-page reuse.
>> - */
>> if (!vmemmap_optimizable_pfn(pfn))
>> return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
>>
>> - zone = pfn_to_zone(pfn, node);
>> + /*
>> + * Before slab is available, vmemmap optimization is used for early
>> + * system RAM, whose zone can be determined from the PFN.
>> + *
>> + * Once slab is available, only ZONE_DEVICE memory reaches this
>> + * optimized population path. Its zone span has not been initialized
>> + * while its vmemmap is being populated, so pfn_to_zone() cannot be
>> + * used. Obtain ZONE_DEVICE directly from the node instead.
>> + */
>> + zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
>> page = vmemmap_shared_tail_page(order, zone);
>> if (!page)
>> return NULL;
>>
>> + /*
>> + * During early vmemmap population, the shared tail vmemmap backing
>> + * page is allocated from memblock before its struct page can safely
>> + * participate in page refcounting. Therefore, no reference can be
>> + * held for each shared PTE mapping, and the mappings must be unshared
>> + * before the vmemmap is depopulated.
>> + *
>> + * Once slab is available, the shared backing page is allocated from
>> + * the buddy allocator and can be refcounted. Hold one reference for
>> + * each shared PTE mapping. The architecture vmemmap teardown drops
>> + * the reference through __free_pages() when removing the mapping,
>> + * preventing the backing page from being freed while it is shared.
>> + *
>> + * The backing page may be shared by enough PTE mappings to exhaust
>> + * the positive range of its reference count. Stop populating the
>> + * vmemmap if another reference cannot be acquired.
>> + */
>> + if (slab_is_available() && !try_get_page(page))
>> + return NULL;
>
> BTW, shouldn't __add_pages() undo the earlier sections on a population
> failure? Say the first section is added successfully, but populating the
> next one fails, e.g. due to an allocation failure:
>
> void *memremap_pages(struct dev_pagemap *pgmap, int nid)
> {
> ...
> const int nr_range = pgmap->nr_range;
> int error, i;
> ...
> pgmap->nr_range = 0;
> error = 0;
> for (i = 0; i < nr_range; i++) {
> error = pagemap_range(pgmap, ¶ms, i, nid);
> if (error)
> break;
> pgmap->nr_range++;
> }
>
> if (i < nr_range) {
> memunmap_pages(pgmap);
> pgmap->nr_range = nr_range;
> return ERR_PTR(error);
> }
> ...
> }
>
> We still need to undo the sections already added in the failed range,
> though ... memunmap_pages() won't touch those, since it only removes
> completed ranges.
>
> If the range starts at a section boundary, we'd hit -EEXIST in
> fill_subsection_map() on retry while those subsection bits are still set.
>
> The old DAX path had this issue too. Could we roll back [start_pfn, pfn)
> in __add_pages() as a separate fix? Something like this:
>
> ---8<---
> diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
> index 796af1028ee2..16a0a2c885bc 100644
> --- a/mm/memory_hotplug.c
> +++ b/mm/memory_hotplug.c
> @@ -380,6 +380,7 @@ EXPORT_SYMBOL_GPL(pfn_to_online_page);
> int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
> struct mhp_params *params)
> {
> + const unsigned long start_pfn = pfn;
> const unsigned long end_pfn = pfn + nr_pages;
> unsigned long cur_nr_pages;
> int err;
> @@ -417,6 +418,10 @@ int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
> break;
> cond_resched();
> }
> +
> + /* Roll back the sections added before the failure. */
> + if (err && pfn != start_pfn)
> + __remove_pages(start_pfn, pfn - start_pfn, altmap, params->pgmap);
> vmemmap_populate_print_last();
> return err;
> }
> --
>
> Hope I haven't missed anything :)
Good catch. This is indeed a pre-existing issue, and the old DAX path was
affected as well.
Your proposed fix looks correct to me. I would slightly prefer keeping the
rollback close to the failure:
err = sparse_add_section(nid, pfn, cur_nr_pages, altmap,
params->pgmap);
if (err) {
__remove_pages(start_pfn, pfn - start_pfn, altmap,
params->pgmap);
break;
}
If the first section fails, this simply calls __remove_pages() with an
empty range, which is a harmless no-op.
Would you mind sending this as a separate bug fix? I will ACK it.
Thanks,
Muchun
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-30 8:53 ` Lance Yang
2026-09-30 10:09 ` Muchun Song
@ 2026-09-30 11:11 ` Muchun Song
2026-09-30 13:38 ` Lance Yang
1 sibling, 1 reply; 15+ messages in thread
From: Muchun Song @ 2026-09-30 11:11 UTC (permalink / raw)
To: Lance Yang
Cc: Muchun Song, maddy, rppt, akpm, david, mpe, npiggin, chleroy,
ritesh.list, sshegde, ljs, liam, vbabka, surenb, mhocko,
qi.zheng, linuxppc-dev, linux-kernel, linux-mm
> On Sep 30, 2026, at 16:53, Lance Yang <lance.yang@linux.dev> wrote:
>
>
> On Tue, Sep 29, 2026 at 01:32:27PM +0800, Muchun Song wrote:
>> The common vmemmap population path cannot yet handle optimized Device DAX
>> mappings on its own. It uses pfn_to_zone() to find the shared tail page,
>> but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
>> is initialized.
>>
>> Teach the common path to use device_zone() for runtime optimized vmemmap
>> population while retaining pfn_to_zone() for early boot. This allows the
>> same path to support both early boot mappings and Device DAX.
>>
>> The backing PFN supplied by the Device DAX-specific population path is no
>> longer used, allowing the redundant lookup and population code to be
>> removed later.
>>
>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>> Acked-by: Qi Zheng <qi.zheng@linux.dev>
>> ---
>> v3:
>> - Collect Acked-by from Qi Zheng
>>
>> v2:
>> - Expand comments around slab initialization to explain zone lookup and
>> page refcounting (suggested by Qi Zheng)
>> ---
>> mm/sparse-vmemmap.c | 64 +++++++++++++++++++++++++--------------------
>> 1 file changed, 35 insertions(+), 29 deletions(-)
>>
>> diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
>> index 8219abc6c3e5..ee4c113ca938 100644
>> --- a/mm/sparse-vmemmap.c
>> +++ b/mm/sparse-vmemmap.c
>> @@ -237,18 +237,43 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
>> struct page *page;
>> const unsigned int order = pfn_to_section_compound_order(pfn);
>>
>> - /*
>> - * Device DAX still relies on vmemmap_populate_compound_pages() for
>> - * head/first-tail allocation and tail-page reuse.
>> - */
>> if (!vmemmap_optimizable_pfn(pfn))
>> return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
>>
>> - zone = pfn_to_zone(pfn, node);
>> + /*
>> + * Before slab is available, vmemmap optimization is used for early
>> + * system RAM, whose zone can be determined from the PFN.
>> + *
>> + * Once slab is available, only ZONE_DEVICE memory reaches this
>> + * optimized population path. Its zone span has not been initialized
>> + * while its vmemmap is being populated, so pfn_to_zone() cannot be
>> + * used. Obtain ZONE_DEVICE directly from the node instead.
>> + */
>> + zone = slab_is_available() ? device_zone(node) : pfn_to_zone(pfn, node);
>> page = vmemmap_shared_tail_page(order, zone);
>> if (!page)
>> return NULL;
>>
>> + /*
>> + * During early vmemmap population, the shared tail vmemmap backing
>> + * page is allocated from memblock before its struct page can safely
>> + * participate in page refcounting. Therefore, no reference can be
>> + * held for each shared PTE mapping, and the mappings must be unshared
>> + * before the vmemmap is depopulated.
>> + *
>> + * Once slab is available, the shared backing page is allocated from
>> + * the buddy allocator and can be refcounted. Hold one reference for
>> + * each shared PTE mapping. The architecture vmemmap teardown drops
>> + * the reference through __free_pages() when removing the mapping,
>> + * preventing the backing page from being freed while it is shared.
>> + *
>> + * The backing page may be shared by enough PTE mappings to exhaust
>> + * the positive range of its reference count. Stop populating the
>> + * vmemmap if another reference cannot be acquired.
>> + */
>> + if (slab_is_available() && !try_get_page(page))
>> + return NULL;
>
> BTW, shouldn't __add_pages() undo the earlier sections on a population
> failure? Say the first section is added successfully, but populating the
> next one fails, e.g. due to an allocation failure:
>
> void *memremap_pages(struct dev_pagemap *pgmap, int nid)
> {
> ...
> const int nr_range = pgmap->nr_range;
> int error, i;
> ...
> pgmap->nr_range = 0;
> error = 0;
> for (i = 0; i < nr_range; i++) {
> error = pagemap_range(pgmap, ¶ms, i, nid);
> if (error)
> break;
> pgmap->nr_range++;
> }
>
> if (i < nr_range) {
> memunmap_pages(pgmap);
> pgmap->nr_range = nr_range;
> return ERR_PTR(error);
> }
> ...
> }
>
> We still need to undo the sections already added in the failed range,
> though ... memunmap_pages() won't touch those, since it only removes
> completed ranges.
>
> If the range starts at a section boundary, we'd hit -EEXIST in
> fill_subsection_map() on retry while those subsection bits are still set.
>
> The old DAX path had this issue too. Could we roll back [start_pfn, pfn)
> in __add_pages() as a separate fix? Something like this:
>
> ---8<---
> diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
> index 796af1028ee2..16a0a2c885bc 100644
> --- a/mm/memory_hotplug.c
> +++ b/mm/memory_hotplug.c
> @@ -380,6 +380,7 @@ EXPORT_SYMBOL_GPL(pfn_to_online_page);
> int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
> struct mhp_params *params)
> {
> + const unsigned long start_pfn = pfn;
> const unsigned long end_pfn = pfn + nr_pages;
> unsigned long cur_nr_pages;
> int err;
> @@ -417,6 +418,10 @@ int __add_pages(int nid, unsigned long pfn, unsigned long nr_pages,
> break;
> cond_resched();
> }
> +
> + /* Roll back the sections added before the failure. */
> + if (err && pfn != start_pfn)
> + __remove_pages(start_pfn, pfn - start_pfn, altmap, params->pgmap);
> vmemmap_populate_print_last();
> return err;
> }
> --
>
> Hope I haven't missed anything :)
Good catch. This is indeed a pre-existing issue, and the old DAX path was
affected as well.
Your proposed fix looks correct to me. I would slightly prefer keeping the
rollback close to the failure:
err = sparse_add_section(nid, pfn, cur_nr_pages, altmap,
params->pgmap);
if (err) {
__remove_pages(start_pfn, pfn - start_pfn, altmap,
params->pgmap);
break;
}
If the first section fails, this simply calls __remove_pages() with an
empty range, which is a harmless no-op.
Would you mind sending this as a separate bug fix? I will ACK it.
Thanks,
Muchun
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-30 11:11 ` Muchun Song
@ 2026-09-30 13:38 ` Lance Yang
0 siblings, 0 replies; 15+ messages in thread
From: Lance Yang @ 2026-09-30 13:38 UTC (permalink / raw)
To: Muchun Song
Cc: Muchun Song, maddy, rppt, akpm, david, mpe, npiggin, chleroy,
ritesh.list, sshegde, ljs, liam, vbabka, surenb, mhocko,
qi.zheng, linuxppc-dev, linux-kernel, linux-mm
On 2026/9/30 19:11, Muchun Song wrote:
[...]
>
> Your proposed fix looks correct to me. I would slightly prefer keeping the
> rollback close to the failure:
>
> err = sparse_add_section(nid, pfn, cur_nr_pages, altmap,
> params->pgmap);
> if (err) {
> __remove_pages(start_pfn, pfn - start_pfn, altmap,
> params->pgmap);
> break;
> }
>
Cool. Will shamelessly steal this approach :D
> If the first section fails, this simply calls __remove_pages() with an
> empty range, which is a harmless no-op.
>
> Would you mind sending this as a separate bug fix? I will ACK it.
Certainly! Lance
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path
2026-09-29 5:32 ` [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path Muchun Song
2026-09-30 8:53 ` Lance Yang
@ 2026-09-30 10:02 ` Lance Yang
1 sibling, 0 replies; 15+ messages in thread
From: Lance Yang @ 2026-09-30 10:02 UTC (permalink / raw)
To: songmuchun
Cc: maddy, rppt, akpm, david, mpe, npiggin, chleroy, ritesh.list,
sshegde, ljs, liam, vbabka, surenb, mhocko, qi.zheng,
linuxppc-dev, linux-kernel, linux-mm, muchun.song, Lance Yang
On Tue, Sep 29, 2026 at 01:32:27PM +0800, Muchun Song wrote:
>The common vmemmap population path cannot yet handle optimized Device DAX
>mappings on its own. It uses pfn_to_zone() to find the shared tail page,
>but Device DAX populates its vmemmap at runtime before the ZONE_DEVICE span
>is initialized.
>
>Teach the common path to use device_zone() for runtime optimized vmemmap
>population while retaining pfn_to_zone() for early boot. This allows the
>same path to support both early boot mappings and Device DAX.
>
>The backing PFN supplied by the Device DAX-specific population path is no
>longer used, allowing the redundant lookup and population code to be
>removed later.
>
>Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>Acked-by: Qi Zheng <qi.zheng@linux.dev>
>---
Nothing jumped out at me, feel free to add:
Reviewed-by: Lance Yang <lance.yang@linux.dev>
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v3 3/6] mm/sparse-vmemmap: drop Device DAX-specific population path
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
2026-09-29 5:32 ` [PATCH v3 1/6] mm/sparse-vmemmap: drop VMEMMAP_POPULATE_DAX Muchun Song
2026-09-29 5:32 ` [PATCH v3 2/6] mm/sparse-vmemmap: support device DAX in common vmemmap path Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-29 5:32 ` [PATCH v3 4/6] mm/sparse-vmemmap: remove the unused ptpfn argument Muchun Song
` (3 subsequent siblings)
6 siblings, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
The common vmemmap path selects the shared page for optimized mappings
itself, so Device DAX no longer needs vmemmap_populate_compound_pages()
to find a shared tail page and pass its backing PFN through the generic
population helpers.
Remove the Device DAX-specific population path and let section memmap
population always use vmemmap_populate(). The powerpc retains an
architecture-specific compound-page implementation, so select it directly
from radix__vmemmap_populate() for optimizable sections.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Use the public order-based optimization predicate in the powerpc path
v2:
- Collect Acked-by from Qi Zheng
---
arch/powerpc/mm/book3s64/radix_pgtable.c | 3 +
mm/mm_init.c | 2 +-
mm/sparse-vmemmap.c | 71 +-----------------------
3 files changed, 5 insertions(+), 71 deletions(-)
diff --git a/arch/powerpc/mm/book3s64/radix_pgtable.c b/arch/powerpc/mm/book3s64/radix_pgtable.c
index 9ca28e4a610a..fd13fa91e5c0 100644
--- a/arch/powerpc/mm/book3s64/radix_pgtable.c
+++ b/arch/powerpc/mm/book3s64/radix_pgtable.c
@@ -1122,7 +1122,10 @@ int __meminit radix__vmemmap_populate(unsigned long start, unsigned long end, in
pud_t *pud;
pmd_t *pmd;
pte_t *pte;
+ unsigned long pfn = page_to_pfn((struct page *)start);
+ if (vmemmap_optimizable_order(pfn_to_section_compound_order(pfn)))
+ return vmemmap_populate_compound_pages(pfn, start, end, node, NULL);
/*
* If altmap is present, Make sure we align the start vmemmap addr
* to PAGE_SIZE so that we calculate the correct start_pfn in
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 56bb4567a494..1650d6bc1211 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1046,7 +1046,7 @@ static void zone_device_page_init_from_template(struct page *page,
* initialize is a lot smaller that the total amount of struct pages being
* mapped. This is a paired / mild layering violation with explicit knowledge
* of how the sparse_vmemmap internals handle compound pages in the lack
- * of an altmap. See vmemmap_populate_compound_pages().
+ * of an altmap.
*/
static inline unsigned long compound_nr_pages(unsigned long pfn,
struct dev_pagemap *pgmap)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index ee4c113ca938..dee95e520376 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -499,71 +499,6 @@ int __meminit vmemmap_populate_hugepages(unsigned long start, unsigned long end,
return 0;
}
-#ifndef vmemmap_populate_compound_pages
-/*
- * For compound pages bigger than section size (e.g. x86 1G compound
- * pages with 2M subsection size) fill the rest of sections as tail
- * pages.
- *
- * Note that memremap_pages() resets @nr_range value and will increment
- * it after each range successful onlining. Thus the value or @nr_range
- * at section memmap populate corresponds to the in-progress range
- * being onlined here.
- */
-static bool __meminit reuse_compound_section(unsigned long start_pfn,
- struct dev_pagemap *pgmap)
-{
- unsigned long nr_pages = pgmap_vmemmap_nr(pgmap);
- unsigned long offset = start_pfn -
- PHYS_PFN(pgmap->ranges[pgmap->nr_range].start);
-
- return !IS_ALIGNED(offset, nr_pages) && nr_pages > PAGES_PER_SUBSECTION;
-}
-
-static int __meminit vmemmap_populate_compound_pages(unsigned long start_pfn,
- unsigned long start,
- unsigned long end, int node,
- struct dev_pagemap *pgmap)
-{
- unsigned long size, addr;
- pte_t *pte;
- int rc;
- struct page *page;
- unsigned int order = pfn_to_section_compound_order(start_pfn);
-
- page = vmemmap_shared_tail_page(order, device_zone(node));
- if (!page)
- return -ENOMEM;
-
- if (reuse_compound_section(start_pfn, pgmap))
- return vmemmap_populate_range(start, end, node, NULL,
- page_to_pfn(page));
-
- size = min(end - start, (1UL << order) * sizeof(struct page));
- for (addr = start; addr < end; addr += size) {
- unsigned long next, last = addr + size;
-
- /* Populate the head page vmemmap page */
- pte = vmemmap_populate_address(addr, node, NULL, -1);
- if (!pte)
- return -ENOMEM;
-
- /*
- * Reuse the shared page for the rest of tail pages
- * See layout diagram in Documentation/mm/vmemmap_dedup.rst
- */
- next = addr + PAGE_SIZE;
- rc = vmemmap_populate_range(next, last, node, NULL,
- page_to_pfn(page));
- if (rc)
- return -ENOMEM;
- }
-
- return 0;
-}
-
-#endif
-
struct page * __meminit __populate_section_memmap(unsigned long pfn,
unsigned long nr_pages, int nid, struct vmem_altmap *altmap,
struct dev_pagemap *pgmap)
@@ -576,11 +511,7 @@ struct page * __meminit __populate_section_memmap(unsigned long pfn,
!IS_ALIGNED(nr_pages, PAGES_PER_SUBSECTION)))
return NULL;
- if (pgmap && section_vmemmap_optimizable(__pfn_to_section(pfn)))
- r = vmemmap_populate_compound_pages(pfn, start, end, nid, pgmap);
- else
- r = vmemmap_populate(start, end, nid, altmap);
-
+ r = vmemmap_populate(start, end, nid, altmap);
if (r < 0)
return NULL;
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* [PATCH v3 4/6] mm/sparse-vmemmap: remove the unused ptpfn argument
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
` (2 preceding siblings ...)
2026-09-29 5:32 ` [PATCH v3 3/6] mm/sparse-vmemmap: drop Device DAX-specific population path Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-29 5:32 ` [PATCH v3 5/6] mm/sparse-vmemmap: open-code vmemmap_populate_address() Muchun Song
` (2 subsequent siblings)
6 siblings, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
vmemmap_pte_populate() no longer uses ptpfn as an input. Drop the
argument to simplify the code.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v2:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 15 ++++++---------
1 file changed, 6 insertions(+), 9 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index dee95e520376..5fd2df2b516f 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -278,7 +278,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node,
}
static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
- struct vmem_altmap *altmap, unsigned long ptpfn)
+ struct vmem_altmap *altmap)
{
pte_t *pte = pte_offset_kernel(pmd, addr);
unsigned long pfn = page_to_pfn((struct page *)addr);
@@ -349,8 +349,7 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
}
static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
- struct vmem_altmap *altmap,
- unsigned long ptpfn)
+ struct vmem_altmap *altmap)
{
pgd_t *pgd;
p4d_t *p4d;
@@ -370,7 +369,7 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
pmd = vmemmap_pmd_populate(pud, addr, node);
if (!pmd)
return NULL;
- pte = vmemmap_pte_populate(pmd, addr, node, altmap, ptpfn);
+ pte = vmemmap_pte_populate(pmd, addr, node, altmap);
if (!pte)
return NULL;
vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
@@ -380,15 +379,13 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
static int __meminit vmemmap_populate_range(unsigned long start,
unsigned long end, int node,
- struct vmem_altmap *altmap,
- unsigned long ptpfn)
+ struct vmem_altmap *altmap)
{
unsigned long addr = start;
pte_t *pte;
for (; addr < end; addr += PAGE_SIZE) {
- pte = vmemmap_populate_address(addr, node, altmap,
- ptpfn);
+ pte = vmemmap_populate_address(addr, node, altmap);
if (!pte)
return -ENOMEM;
}
@@ -399,7 +396,7 @@ static int __meminit vmemmap_populate_range(unsigned long start,
int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
int node, struct vmem_altmap *altmap)
{
- return vmemmap_populate_range(start, end, node, altmap, -1);
+ return vmemmap_populate_range(start, end, node, altmap);
}
/*
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* [PATCH v3 5/6] mm/sparse-vmemmap: open-code vmemmap_populate_address()
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
` (3 preceding siblings ...)
2026-09-29 5:32 ` [PATCH v3 4/6] mm/sparse-vmemmap: remove the unused ptpfn argument Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-29 5:32 ` [PATCH v3 6/6] mm/mm_init: add zone mismatch warning during page init Muchun Song
2026-09-29 21:32 ` [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Andrew Morton
6 siblings, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
vmemmap_populate_address() no longer has any callers that need the
returned PTE. Its only remaining user, vmemmap_populate_range(), only
checks whether population succeeded.
Open-code vmemmap_populate_address() directly in
vmemmap_populate_basepages(), remove the now-redundant range helper,
and return -ENOMEM directly on failure.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v2:
- Collect Acked-by from Qi Zheng
---
mm/sparse-vmemmap.c | 54 ++++++++++++++-------------------------------
1 file changed, 17 insertions(+), 37 deletions(-)
diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
index 5fd2df2b516f..9a58323dd628 100644
--- a/mm/sparse-vmemmap.c
+++ b/mm/sparse-vmemmap.c
@@ -348,8 +348,8 @@ static pgd_t * __meminit vmemmap_pgd_populate(unsigned long addr, int node)
return pgd;
}
-static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
- struct vmem_altmap *altmap)
+int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
+ int node, struct vmem_altmap *altmap)
{
pgd_t *pgd;
p4d_t *p4d;
@@ -357,48 +357,28 @@ static pte_t * __meminit vmemmap_populate_address(unsigned long addr, int node,
pmd_t *pmd;
pte_t *pte;
- pgd = vmemmap_pgd_populate(addr, node);
- if (!pgd)
- return NULL;
- p4d = vmemmap_p4d_populate(pgd, addr, node);
- if (!p4d)
- return NULL;
- pud = vmemmap_pud_populate(p4d, addr, node);
- if (!pud)
- return NULL;
- pmd = vmemmap_pmd_populate(pud, addr, node);
- if (!pmd)
- return NULL;
- pte = vmemmap_pte_populate(pmd, addr, node, altmap);
- if (!pte)
- return NULL;
- vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
-
- return pte;
-}
-
-static int __meminit vmemmap_populate_range(unsigned long start,
- unsigned long end, int node,
- struct vmem_altmap *altmap)
-{
- unsigned long addr = start;
- pte_t *pte;
-
- for (; addr < end; addr += PAGE_SIZE) {
- pte = vmemmap_populate_address(addr, node, altmap);
+ for (unsigned long addr = start; addr < end; addr += PAGE_SIZE) {
+ pgd = vmemmap_pgd_populate(addr, node);
+ if (!pgd)
+ return -ENOMEM;
+ p4d = vmemmap_p4d_populate(pgd, addr, node);
+ if (!p4d)
+ return -ENOMEM;
+ pud = vmemmap_pud_populate(p4d, addr, node);
+ if (!pud)
+ return -ENOMEM;
+ pmd = vmemmap_pmd_populate(pud, addr, node);
+ if (!pmd)
+ return -ENOMEM;
+ pte = vmemmap_pte_populate(pmd, addr, node, altmap);
if (!pte)
return -ENOMEM;
+ vmemmap_verify(pte, node, addr, addr + PAGE_SIZE);
}
return 0;
}
-int __meminit vmemmap_populate_basepages(unsigned long start, unsigned long end,
- int node, struct vmem_altmap *altmap)
-{
- return vmemmap_populate_range(start, end, node, altmap);
-}
-
/*
* Write protect the mirrored tail page structs for HVO. This will be
* called from the hugetlb code when gathering and initializing the
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* [PATCH v3 6/6] mm/mm_init: add zone mismatch warning during page init
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
` (4 preceding siblings ...)
2026-09-29 5:32 ` [PATCH v3 5/6] mm/sparse-vmemmap: open-code vmemmap_populate_address() Muchun Song
@ 2026-09-29 5:32 ` Muchun Song
2026-09-29 21:32 ` [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Andrew Morton
6 siblings, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-29 5:32 UTC (permalink / raw)
To: Madhavan Srinivasan, Mike Rapoport, Andrew Morton, David Hildenbrand
Cc: Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
Muchun Song, muchun.song
For vmemmap-optimized sections, tail struct pages may be backed by
shared vmemmap pages. Those shared pages must carry the same page zone
ID as the struct pages initialized for the section.
Warn in __init_single_page() if the shared tail page has a different
page_zone_id(), which would indicate inconsistent initialization.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
---
v3:
- Collect Acked-by from Qi Zheng
v2:
- New patch.
---
mm/mm_init.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 1650d6bc1211..bd02e8d06965 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -609,6 +609,9 @@ void __meminit __init_single_page(struct page *page, unsigned long pfn,
if (!is_highmem_idx(zone))
set_page_address(page, __va(pfn << PAGE_SHIFT));
#endif
+ VM_WARN_ON_ONCE(vmemmap_optimizable_order(pfn_to_section_compound_order(pfn)) &&
+ page_zone_id(page + VMEMMAP_OPTIMIZATION_NR_STRUCT_PAGES) !=
+ page_zone_id(page));
}
#ifdef CONFIG_NUMA
--
2.54.0
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths
2026-09-29 5:32 [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Muchun Song
` (5 preceding siblings ...)
2026-09-29 5:32 ` [PATCH v3 6/6] mm/mm_init: add zone mismatch warning during page init Muchun Song
@ 2026-09-29 21:32 ` Andrew Morton
2026-09-30 1:26 ` Muchun Song
6 siblings, 1 reply; 15+ messages in thread
From: Andrew Morton @ 2026-09-29 21:32 UTC (permalink / raw)
To: Muchun Song
Cc: Madhavan Srinivasan, Mike Rapoport, David Hildenbrand,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Ritesh Harjani, Shrikanth Hegde, Lorenzo Stoakes,
Liam R . Howlett, Vlastimil Babka, Suren Baghdasaryan,
Michal Hocko, Qi Zheng, linuxppc-dev, linux-kernel, linux-mm,
muchun.song
On Tue, 29 Sep 2026 13:32:25 +0800 Muchun Song <songmuchun@bytedance.com> wrote:
> This v3 is based on mm-new commit 2ddb90ee544a, which contains v5 of
> "mm: Switch device DAX to section-based vmemmap optimization" [1].
>
> This series is split out from the earlier, larger series "mm: Generalize
> HVO for HugeTLB and device DAX" [2]. While the parent series generalizes
> vmemmap optimization across HugeTLB and device DAX, this subset addresses
> a single, self-contained step: unifying their vmemmap population paths.
>
> After the preceding Device DAX conversion, both HugeTLB and Device DAX
> describe optimized vmemmap mappings through memory-section metadata and
> use per-zone shared tail vmemmap pages. The generic code, however, still
> carries a Device DAX-specific population flag and compound-page population
> path, along with arguments and helpers needed only by that path.
>
> This series first removes VMEMMAP_POPULATE_DAX and moves selection and
> reference handling for the shared tail page into the common vmemmap
> population path. It then removes the generic Device DAX-specific
> compound-page population path and routes section vmemmap population
> through vmemmap_populate(). The powerpc radix path continues to use its
> architecture-specific compound-page population implementation for
> optimizable sections.
>
> The remaining patches remove the unused ptpfn argument, open-code
> vmemmap_populate_address() now that no caller needs its returned PTE, and
> add a warning for inconsistent zone initialization of shared tail vmemmap
> pages.
>
> This is the fourth smaller step toward the broader HVO generalization.
> After this series, HugeTLB and Device DAX use the same population model
> instead of parallel generic paths, while powerpc keeps its
> architecture-specific implementation.
afaict the whole series is "no functional changes intended". Apart
from [6/6]'s new WARN_ON.
And [3/6] is the one which could cause unintended fuctional changes!
Thanks. I'll queue it. Reluctantly. We're up to 688 MM patches this
cycle and it's time to stop.
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths
2026-09-29 21:32 ` [PATCH v3 0/6] mm: Unify device DAX and HugeTLB vmemmap population paths Andrew Morton
@ 2026-09-30 1:26 ` Muchun Song
0 siblings, 0 replies; 15+ messages in thread
From: Muchun Song @ 2026-09-30 1:26 UTC (permalink / raw)
To: Andrew Morton
Cc: Muchun Song, Madhavan Srinivasan, Mike Rapoport,
David Hildenbrand, Michael Ellerman, Nicholas Piggin,
Christophe Leroy, Ritesh Harjani, Shrikanth Hegde,
Lorenzo Stoakes, Liam R . Howlett, Vlastimil Babka,
Suren Baghdasaryan, Michal Hocko, Qi Zheng, linuxppc-dev,
linux-kernel, linux-mm
> On Sep 30, 2026, at 05:32, Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Tue, 29 Sep 2026 13:32:25 +0800 Muchun Song <songmuchun@bytedance.com> wrote:
>
>> This v3 is based on mm-new commit 2ddb90ee544a, which contains v5 of
>> "mm: Switch device DAX to section-based vmemmap optimization" [1].
>>
>> This series is split out from the earlier, larger series "mm: Generalize
>> HVO for HugeTLB and device DAX" [2]. While the parent series generalizes
>> vmemmap optimization across HugeTLB and device DAX, this subset addresses
>> a single, self-contained step: unifying their vmemmap population paths.
>>
>> After the preceding Device DAX conversion, both HugeTLB and Device DAX
>> describe optimized vmemmap mappings through memory-section metadata and
>> use per-zone shared tail vmemmap pages. The generic code, however, still
>> carries a Device DAX-specific population flag and compound-page population
>> path, along with arguments and helpers needed only by that path.
>>
>> This series first removes VMEMMAP_POPULATE_DAX and moves selection and
>> reference handling for the shared tail page into the common vmemmap
>> population path. It then removes the generic Device DAX-specific
>> compound-page population path and routes section vmemmap population
>> through vmemmap_populate(). The powerpc radix path continues to use its
>> architecture-specific compound-page population implementation for
>> optimizable sections.
>>
>> The remaining patches remove the unused ptpfn argument, open-code
>> vmemmap_populate_address() now that no caller needs its returned PTE, and
>> add a warning for inconsistent zone initialization of shared tail vmemmap
>> pages.
>>
>> This is the fourth smaller step toward the broader HVO generalization.
>> After this series, HugeTLB and Device DAX use the same population model
>> instead of parallel generic paths, while powerpc keeps its
>> architecture-specific implementation.
>
> afaict the whole series is "no functional changes intended". Apart
> from [6/6]'s new WARN_ON.
Yes.
>
> And [3/6] is the one which could cause unintended fuctional changes!
Understood. However, the resulting vmemmap layout is intended to remain
unchanged.
>
> Thanks. I'll queue it. Reluctantly. We're up to 688 MM patches this
> cycle and it's time to stop.
Thanks for taking it despite the load.
Muchun,
Thanks.
^ permalink raw reply [flat|nested] 15+ messages in thread