* [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization
@ 2026-05-27 3:36 Li Zhe
2026-05-27 3:36 ` [PATCH v3 1/8] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
` (8 more replies)
0 siblings, 9 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
This series reduces latency in paths which create or recreate
ZONE_DEVICE memmaps. Pmem hotplug is a concrete example.
memmap_init_zone_device() can spend a substantial amount of time
initializing large ZONE_DEVICE ranges because it repeats nearly
identical struct page setup for every PFN.
The series addresses that overhead in eight patches.
The first patch updates a stale comment in __init_zone_device_page() so
it matches the current ZONE_DEVICE refcount rules.
The second patch factors the reusable pieces out of
__init_zone_device_page() so later patches can share the same logic
without changing the existing slow path.
The third patch adds set_page_section_from_pfn(), so generic callers
can update section bits from a PFN without open-coding
SECTION_IN_PAGE_FLAGS handling.
The fourth patch adds a template-based fast path for ZONE_DEVICE head
pages. Instead of rebuilding the same struct page state for every PFN,
it prepares one reusable head-page template through the existing slow
path, refreshes the PFN-dependent fields before each copy, and then
copies that template into the destination page.
The fifth patch extends the same template-based approach to compound
tails, so pfns_per_compound > 1 can also benefit from the fast path.
The sixth patch introduces memcpy_streaming() and
memcpy_streaming_drain() as a generic interface for write-once copies,
with a memcpy() fallback for architectures that do not provide a
specialized backend, or for transfers which an architecture-specific
backend cannot safely handle with non-temporal stores.
The seventh patch extends x86 memcpy_flushcache() small fixed-size
fastpaths for naturally aligned copies, so common struct-page-sized
streaming copies can stay on the inline movnti path while preserving
forward store order.
The last patch switches the zone-device template-copy path over to
memcpy_streaming(). It keeps pageblock-aligned PFNs on regular memcpy()
so pageblock setup can immediately read page metadata back through the
usual helpers, and it drains streaming stores before later normal
stores update overlapping or dependent compound-page metadata.
The optimized path is disabled when the page_ref_set tracepoint is
enabled, and sanitized builds remain on the slow path so their
instrumented stores are preserved.
Testing
=======
Tests were run in a VM on an Intel Ice Lake server.
Two PMEM configurations were used:
- a 100 GB fsdax namespace configured with map=dev, which exercises
the nd_pmem rebind path (pfns_per_compound == 1)
- a 100 GB devdax namespace configured with align=2097152, which
exercises the dax_pmem rebind path (pfns_per_compound > 1)
For each configuration, the corresponding driver was unbound and
rebound 30 times. Memmap initialization latency was collected from the
pr_debug() output of memmap_init_zone_device().
The first bind is reported separately, and the average of subsequent
rebinds is used as the repeat-run result.
Performance
===========
nd_pmem rebind, 100 GB fsdax namespace, map=dev
Base(v7.1-rc3):
First binding: 1486 ms
Average of subsequent rebinds: 273.52 ms
With patches 1-4 applied:
First binding: 1422 ms
Average of subsequent rebinds: 246.65 ms
Full series:
First binding: 1285 ms
Average of subsequent rebinds: 114.31 ms
dax_pmem rebind, 100 GB devdax namespace, align=2097152
Base(v7.1-rc3):
First binding: 1515 ms
Average of subsequent rebinds: 313.45 ms
With patches 1-5 applied:
First binding: 1422 ms
Average of subsequent rebinds: 240.42 ms
Full series:
First binding: 1331 ms
Average of subsequent rebinds: 99.37 ms
Li Zhe (8):
mm: fix stale ZONE_DEVICE refcount comment
mm: factor zone-device page init helpers out of
__init_zone_device_page
mm: add a set_page_section_from_pfn() helper
mm: add a template-based fast path for zone-device page init
mm: extend the template fast path to zone-device compound tails
string: introduce memcpy_streaming() helpers
x86/string: extend memcpy_flushcache() fixed-size fastpaths
mm: use memcpy_streaming() in zone-device template copies
arch/x86/include/asm/string_64.h | 157 ++++++++++++++++++++++---
include/linux/mm.h | 19 ++-
include/linux/string.h | 20 ++++
mm/mm_init.c | 195 +++++++++++++++++++++++++++----
4 files changed, 352 insertions(+), 39 deletions(-)
---
v2: https://lore.kernel.org/all/20260521040124.10608-1-lizhe.67@bytedance.com/
v1: https://lore.kernel.org/all/20260515082045.63029-1-lizhe.67@bytedance.com/
Changelogs:
v2->v3:
- Add a leading comment-only cleanup patch to fix the stale ZONE_DEVICE
refcount comment, so later refactoring patches no longer carry that
stale wording forward.
- Rename the refcount-policy helper to pagemap_resets_refcount() and
make it a bool predicate. Suggested by Mike Rapoport.
- Update the reusable template in place from the first template
fast-path patch onward, so later patches only change the copy
primitive instead of introducing an intermediate post-copy fixup
step. Suggested by Mike Rapoport.
- Clean up helper indentation and use early continue in the
compound-tail loop. Suggested by Mike Rapoport.
- Narrow the x86 memcpy_streaming() backend to transfers that can stay
entirely on the non-temporal store path, keeping zero-length and
unaligned cases on memcpy(). Suggested by Andrew Morton.
- Restrict the x86 memcpy_flushcache() small-copy fastpath to naturally
aligned cases, add volatile/memory clobbers, and preserve forward
movnti store order. Suggested by Andrew Morton.
- Drop the obsolete struct page size/alignment eligibility check from
the template fast path now that the intermediate patches use
memcpy() and the streaming backend has a safe fallback.
- Refresh the benchmark results.
v1->v2:
- Move the pageblock-helper split into patch 1, and add a dedicated
set_page_section_from_pfn() helper so generic callers no longer
open-code SECTION_IN_PAGE_FLAGS handling. Suggested by Mike Rapoport.
- Drop the v1 32-bit gating and document instead that CONFIG_ZONE_DEVICE
is 64-bit only because it depends on MEMORY_HOTPLUG. Suggested by
Mike Rapoport.
- Replace the v1 BUILD_BUG_ON() struct page layout checks with a runtime
fast-path eligibility check that falls back to the slow path when the
layout is unsuitable. Suggested by Mike Rapoport.
- Rename the template fast-path helpers to zone_device_* names for
clarity. Suggested by Mike Rapoport.
- Replace the v1 open-coded arch_optimize_store_u64()/drain() approach
with a generic memcpy_streaming()/memcpy_streaming_drain() interface,
and move the x86 optimization under memcpy_flushcache(). Suggested by
Alistair Popple.
- Split the old v1 streaming-copy patch into three patches: introduce
the generic helper, extend the x86 backend, and then switch mm over
to the new interface. This is part of the memcpy-based rework
suggested by Alistair Popple.
- Refresh the performance section and report whole-series results as
series-level numbers.
- Fix a missing memcpy_streaming_drain() in the compound-page
initialization path, following review feedback from Andrew Morton.
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 1/8] mm: fix stale ZONE_DEVICE refcount comment
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 2/8] mm: factor zone-device page init helpers out of __init_zone_device_page Li Zhe
` (7 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
The comment in __init_zone_device_page() still uses the old
MEMORY_TYPE_* names and implies that FS_DAX pages regain a
refcount of 1 in the free path. That no longer matches the code.
Update the comment to describe the current policy correctly:
MEMORY_DEVICE_GENERIC pages regain a refcount of 1 in the free path,
while the remaining ZONE_DEVICE types start from 0 here and raise the
count again when the allocator or driver hands the page out.
No functional change intended.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
mm/mm_init.c | 10 +++-------
1 file changed, 3 insertions(+), 7 deletions(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index f9f8e1af921c..35de3b6a186d 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1027,13 +1027,9 @@ static void __ref __init_zone_device_page(struct page *page, unsigned long pfn,
}
/*
- * ZONE_DEVICE pages other than MEMORY_TYPE_GENERIC are released
- * directly to the driver page allocator which will set the page count
- * to 1 when allocating the page.
- *
- * MEMORY_TYPE_GENERIC and MEMORY_TYPE_FS_DAX pages automatically have
- * their refcount reset to one whenever they are freed (ie. after
- * their refcount drops to 0).
+ * MEMORY_DEVICE_GENERIC pages regain a refcount of 1 in the free
+ * path. The remaining ZONE_DEVICE types start from 0 here and raise
+ * the count again when the allocator or driver hands the page out.
*/
switch (pgmap->type) {
case MEMORY_DEVICE_FS_DAX:
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 2/8] mm: factor zone-device page init helpers out of __init_zone_device_page
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
2026-05-27 3:36 ` [PATCH v3 1/8] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 3/8] mm: add a set_page_section_from_pfn() helper Li Zhe
` (6 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
memmap_init_zone_device() currently mixes refcount policy, core
ZONE_DEVICE page setup, and pageblock metadata handling in a single
helper.
Factor the refcount-reset predicate into pagemap_resets_refcount(), move
the common page initialization into __zone_device_page_init(), split
pageblock handling into zone_device_page_init_pageblock(), and wrap the
existing slow path in zone_device_page_init_slow().
This keeps the slow-path behaviour unchanged and gives later patches
reusable helper boundaries.
No functional change intended.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
mm/mm_init.c | 62 ++++++++++++++++++++++++++++++++++++----------------
1 file changed, 43 insertions(+), 19 deletions(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 35de3b6a186d..2e5899c5cf35 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -987,11 +987,38 @@ static void __init memmap_init(void)
}
#ifdef CONFIG_ZONE_DEVICE
-static void __ref __init_zone_device_page(struct page *page, unsigned long pfn,
+/*
+ * Return true when the free path for this pagemap type restores the page
+ * refcount to 1, so memmap_init_zone_device() can keep the count set by
+ * __init_single_page(). Otherwise initialize the refcount to 0 and leave
+ * it to the allocator or pgmap callbacks to raise it when the page is
+ * handed out again.
+ */
+static inline bool pagemap_resets_refcount(const struct dev_pagemap *pgmap)
+{
+ /*
+ * MEMORY_DEVICE_GENERIC pages regain a refcount of 1 in the free
+ * path. The remaining ZONE_DEVICE types start from 0 here and raise
+ * the count again when the allocator or driver hands the page out.
+ */
+ switch (pgmap->type) {
+ case MEMORY_DEVICE_FS_DAX:
+ case MEMORY_DEVICE_PRIVATE:
+ case MEMORY_DEVICE_COHERENT:
+ case MEMORY_DEVICE_PCI_P2PDMA:
+ return false;
+ case MEMORY_DEVICE_GENERIC:
+ return true;
+ default:
+ WARN_ONCE(1, "Unknown memory type!");
+ return true;
+ }
+}
+
+static void __ref __zone_device_page_init(struct page *page, unsigned long pfn,
unsigned long zone_idx, int nid,
struct dev_pagemap *pgmap)
{
-
__init_single_page(page, pfn, zone_idx, nid);
/*
@@ -1010,7 +1037,11 @@ static void __ref __init_zone_device_page(struct page *page, unsigned long pfn,
*/
page_folio(page)->pgmap = pgmap;
page->zone_device_data = NULL;
+}
+static void __ref zone_device_page_init_pageblock(struct page *page,
+ unsigned long pfn)
+{
/*
* Mark the block movable so that blocks are reserved for
* movable at startup. This will force kernel allocations
@@ -1025,23 +1056,16 @@ static void __ref __init_zone_device_page(struct page *page, unsigned long pfn,
init_pageblock_migratetype(page, MIGRATE_MOVABLE, false);
cond_resched();
}
+}
- /*
- * MEMORY_DEVICE_GENERIC pages regain a refcount of 1 in the free
- * path. The remaining ZONE_DEVICE types start from 0 here and raise
- * the count again when the allocator or driver hands the page out.
- */
- switch (pgmap->type) {
- case MEMORY_DEVICE_FS_DAX:
- case MEMORY_DEVICE_PRIVATE:
- case MEMORY_DEVICE_COHERENT:
- case MEMORY_DEVICE_PCI_P2PDMA:
+static void __ref zone_device_page_init_slow(struct page *page,
+ unsigned long pfn, unsigned long zone_idx, int nid,
+ struct dev_pagemap *pgmap)
+{
+ __zone_device_page_init(page, pfn, zone_idx, nid, pgmap);
+ if (!pagemap_resets_refcount(pgmap))
set_page_count(page, 0);
- break;
-
- case MEMORY_DEVICE_GENERIC:
- break;
- }
+ zone_device_page_init_pageblock(page, pfn);
}
/*
@@ -1080,7 +1104,7 @@ static void __ref memmap_init_compound(struct page *head,
for (pfn = head_pfn + 1; pfn < end_pfn; pfn++) {
struct page *page = pfn_to_page(pfn);
- __init_zone_device_page(page, pfn, zone_idx, nid, pgmap);
+ zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
prep_compound_tail(page, head, order);
set_page_count(page, 0);
}
@@ -1116,7 +1140,7 @@ void __ref memmap_init_zone_device(struct zone *zone,
for (pfn = start_pfn; pfn < end_pfn; pfn += pfns_per_compound) {
struct page *page = pfn_to_page(pfn);
- __init_zone_device_page(page, pfn, zone_idx, nid, pgmap);
+ zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
if (pfns_per_compound == 1)
continue;
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 3/8] mm: add a set_page_section_from_pfn() helper
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
2026-05-27 3:36 ` [PATCH v3 1/8] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
2026-05-27 3:36 ` [PATCH v3 2/8] mm: factor zone-device page init helpers out of __init_zone_device_page Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 4/8] mm: add a template-based fast path for zone-device page init Li Zhe
` (5 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
Callers that want to update section bits from a PFN currently need to
open-code:
set_page_section(page, pfn_to_section_nr(pfn));
and guard that sequence with #ifdef SECTION_IN_PAGE_FLAGS.
Add set_page_section_from_pfn() to wrap that update in one place. When
section bits are stored in page flags, the helper derives the section
number from the PFN and updates the page flags. Otherwise it degrades to
a no-op.
Convert set_page_links() to use the new helper so later ZONE_DEVICE
fast-path patches can also update section bits without open-coding
SECTION_IN_PAGE_FLAGS at each callsite.
This keeps the PFN-to-section translation local to the configurations
that actually store section bits in struct page flags, and avoids
exposing that detail to generic callers.
No functional change intended.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
include/linux/mm.h | 19 ++++++++++++++++---
1 file changed, 16 insertions(+), 3 deletions(-)
diff --git a/include/linux/mm.h b/include/linux/mm.h
index af23453e9dbd..bf84e698385c 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -2507,11 +2507,26 @@ static inline void set_page_section(struct page *page, unsigned long section)
page->flags.f |= (section & SECTIONS_MASK) << SECTIONS_PGSHIFT;
}
+static inline void set_page_section_from_pfn(struct page *page,
+ unsigned long pfn)
+{
+ set_page_section(page, pfn_to_section_nr(pfn));
+}
+
static inline unsigned long memdesc_section(memdesc_flags_t mdf)
{
return (mdf.f >> SECTIONS_PGSHIFT) & SECTIONS_MASK;
}
#else /* !SECTION_IN_PAGE_FLAGS */
+static inline void set_page_section(struct page *page, unsigned long section)
+{
+}
+
+static inline void set_page_section_from_pfn(struct page *page,
+ unsigned long pfn)
+{
+}
+
static inline unsigned long memdesc_section(memdesc_flags_t mdf)
{
return 0;
@@ -2734,9 +2749,7 @@ static inline void set_page_links(struct page *page, enum zone_type zone,
{
set_page_zone(page, zone);
set_page_node(page, node);
-#ifdef SECTION_IN_PAGE_FLAGS
- set_page_section(page, pfn_to_section_nr(pfn));
-#endif
+ set_page_section_from_pfn(page, pfn);
}
/**
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 4/8] mm: add a template-based fast path for zone-device page init
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (2 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 3/8] mm: add a set_page_section_from_pfn() helper Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 5/8] mm: extend the template fast path to zone-device compound tails Li Zhe
` (4 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
memmap_init_zone_device() repeats nearly identical head-page
initialization for each PFN. Prepare one reusable ZONE_DEVICE head-page
template through the existing slow path, refresh the PFN-dependent
fields in that template before each copy, and memcpy it into each
destination page.
The optimized path assigns _refcount through the copied template, so
keep it disabled when the page_ref_set tracepoint is enabled.
This patch accelerates the pfns_per_compound == 1 case. Compound tails
are handled in the next patch.
Tested in a VM with a 100 GB fsdax namespace device configured with
map=dev on Intel Ice Lake server. This test exercises the nd_pmem rebind
path (pfns_per_compound == 1).
Test procedure:
Rebind the nd_pmem driver 30 times and collect the memmap initialization
time from the pr_debug() output of memmap_init_zone_device().
Base(v7.1-rc3):
First binding: 1486 ms
Average of subsequent rebinds: 273.52 ms
With this patch and its prerequisites applied:
First binding: 1422 ms
Average of subsequent rebinds: 246.65 ms
This reduces the average rebind time from 273.52 ms to 246.65 ms, or
about 10%.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
mm/mm_init.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 62 insertions(+), 1 deletion(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 2e5899c5cf35..53c0241c66b7 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1068,6 +1068,56 @@ static void __ref zone_device_page_init_slow(struct page *page,
zone_device_page_init_pageblock(page, pfn);
}
+static inline bool zone_device_page_init_optimization_enabled(void)
+{
+ /*
+ * The template fast path copies a preinitialized struct page image.
+ * Skip it when the page_ref_set tracepoint is enabled.
+ */
+ return !page_ref_tracepoint_active(page_ref_set);
+}
+
+static inline void zone_device_template_page_init(struct page *template,
+ unsigned long pfn,
+ unsigned long zone_idx,
+ int nid,
+ struct dev_pagemap *pgmap)
+{
+ __zone_device_page_init(template, pfn, zone_idx, nid, pgmap);
+ if (!pagemap_resets_refcount(pgmap))
+ set_page_count(template, 0);
+}
+
+/*
+ * 'template' is a reusable page prototype rather than a strictly immutable
+ * object. Most ZONE_DEVICE fields stay constant across the pages covered by
+ * the current template, but section bits and page->virtual may still depend
+ * on the PFN. Refresh those PFN-dependent fields in the template before
+ * copying it into @page.
+ */
+static inline void zone_device_page_update_template(struct page *template,
+ unsigned long pfn)
+{
+ set_page_section_from_pfn(template, pfn);
+#ifdef WANT_PAGE_VIRTUAL
+ if (!is_highmem_idx(ZONE_DEVICE))
+ set_page_address(template, __va(pfn << PAGE_SHIFT));
+#endif
+}
+
+static void zone_device_page_init_from_template(struct page *page,
+ unsigned long pfn, struct page *template)
+{
+ /*
+ * 'template' carries the invariant portion of a ZONE_DEVICE struct
+ * page. Update the PFN-dependent fields in place before copying it
+ * to the destination page.
+ */
+ zone_device_page_update_template(template, pfn);
+ memcpy(page, template, sizeof(*page));
+ zone_device_page_init_pageblock(page, pfn);
+}
+
/*
* With compound page geometry and when struct pages are stored in ram most
* tail pages are reused. Consequently, the amount of unique struct pages to
@@ -1116,6 +1166,7 @@ void __ref memmap_init_zone_device(struct zone *zone,
unsigned long nr_pages,
struct dev_pagemap *pgmap)
{
+ bool use_template = zone_device_page_init_optimization_enabled();
unsigned long pfn, end_pfn = start_pfn + nr_pages;
struct pglist_data *pgdat = zone->zone_pgdat;
struct vmem_altmap *altmap = pgmap_altmap(pgmap);
@@ -1123,6 +1174,7 @@ void __ref memmap_init_zone_device(struct zone *zone,
unsigned long zone_idx = zone_idx(zone);
unsigned long start = jiffies;
int nid = pgdat->node_id;
+ struct page template;
if (WARN_ON_ONCE(!pgmap || zone_idx != ZONE_DEVICE))
return;
@@ -1137,10 +1189,19 @@ void __ref memmap_init_zone_device(struct zone *zone,
nr_pages = end_pfn - start_pfn;
}
+ if (use_template)
+ zone_device_template_page_init(&template, start_pfn, zone_idx,
+ nid, pgmap);
+
for (pfn = start_pfn; pfn < end_pfn; pfn += pfns_per_compound) {
struct page *page = pfn_to_page(pfn);
- zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
+ if (use_template)
+ zone_device_page_init_from_template(page, pfn,
+ &template);
+ else
+ zone_device_page_init_slow(page, pfn, zone_idx,
+ nid, pgmap);
if (pfns_per_compound == 1)
continue;
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 5/8] mm: extend the template fast path to zone-device compound tails
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (3 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 4/8] mm: add a template-based fast path for zone-device page init Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 6/8] string: introduce memcpy_streaming() helpers Li Zhe
` (3 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
The template fast path from the previous patch only accelerates head
pages. Compound tails in memmap_init_compound() still go through the
slow path one by one.
Build separate head and tail templates and reuse one prepared tail
template across the tail pages in a compound range. Head pages preserve
the existing refcount policy, while compound tails always start with a
refcount of 0 after prep_compound_tail().
This extends the template-copy fast path to pfns_per_compound > 1
without changing the existing slow path. Tail-page PFN-dependent fields
are refreshed in the reusable tail template before each copy.
Tested in a VM with a 100 GB devdax namespace (align=2097152) on Intel
Ice Lake server. This test exercises the dax_pmem rebind path and
measures memmap initialization latency.
Test procedure:
Unbind and rebind the dax_pmem driver 30 times, collect memmap
initialization time from the pr_debug() output of memmap_init_zone_device().
Base(v7.1-rc3):
First binding: 1515 ms
Average of subsequent rebinds: 313.45 ms
With this patch and its prerequisites applied:
First binding: 1422 ms
Average of subsequent rebinds: 240.42 ms
This reduces the average rebind time from 313.45 ms to 240.42 ms, or
about 23.3%.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Reviewed-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
---
mm/mm_init.c | 45 ++++++++++++++++++++++++++++++++++++---------
1 file changed, 36 insertions(+), 9 deletions(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 53c0241c66b7..d5ccb49a048f 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1077,17 +1077,25 @@ static inline bool zone_device_page_init_optimization_enabled(void)
return !page_ref_tracepoint_active(page_ref_set);
}
-static inline void zone_device_template_page_init(struct page *template,
- unsigned long pfn,
- unsigned long zone_idx,
- int nid,
- struct dev_pagemap *pgmap)
+static inline void zone_device_template_head_page_init(struct page *template,
+ unsigned long pfn, unsigned long zone_idx, int nid,
+ struct dev_pagemap *pgmap)
{
__zone_device_page_init(template, pfn, zone_idx, nid, pgmap);
if (!pagemap_resets_refcount(pgmap))
set_page_count(template, 0);
}
+static inline void zone_device_template_tail_page_init(struct page *template,
+ unsigned long pfn, unsigned long zone_idx, int nid,
+ struct dev_pagemap *pgmap, const struct page *head,
+ unsigned int order)
+{
+ __zone_device_page_init(template, pfn, zone_idx, nid, pgmap);
+ prep_compound_tail(template, head, order);
+ set_page_count(template, 0);
+}
+
/*
* 'template' is a reusable page prototype rather than a strictly immutable
* object. Most ZONE_DEVICE fields stay constant across the pages covered by
@@ -1139,10 +1147,12 @@ static void __ref memmap_init_compound(struct page *head,
unsigned long head_pfn,
unsigned long zone_idx, int nid,
struct dev_pagemap *pgmap,
- unsigned long nr_pages)
+ unsigned long nr_pages,
+ bool use_template)
{
unsigned long pfn, end_pfn = head_pfn + nr_pages;
unsigned int order = pgmap->vmemmap_shift;
+ struct page template;
/*
* We have to initialize the pages, including setting up page links.
@@ -1151,9 +1161,25 @@ static void __ref memmap_init_compound(struct page *head,
* the pages in the same go.
*/
__SetPageHead(head);
+
+ /*
+ * All tails of the same compound page share the state established by
+ * prep_compound_tail(). Reuse one tail template for the whole range and
+ * refresh only the PFN-dependent fields in that template before each copy.
+ */
+ if (use_template)
+ zone_device_template_tail_page_init(&template, head_pfn + 1,
+ zone_idx, nid, pgmap,
+ head, order);
+
for (pfn = head_pfn + 1; pfn < end_pfn; pfn++) {
struct page *page = pfn_to_page(pfn);
+ if (use_template) {
+ zone_device_page_init_from_template(page, pfn,
+ &template);
+ continue;
+ }
zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
prep_compound_tail(page, head, order);
set_page_count(page, 0);
@@ -1190,8 +1216,8 @@ void __ref memmap_init_zone_device(struct zone *zone,
}
if (use_template)
- zone_device_template_page_init(&template, start_pfn, zone_idx,
- nid, pgmap);
+ zone_device_template_head_page_init(&template, start_pfn,
+ zone_idx, nid, pgmap);
for (pfn = start_pfn; pfn < end_pfn; pfn += pfns_per_compound) {
struct page *page = pfn_to_page(pfn);
@@ -1207,7 +1233,8 @@ void __ref memmap_init_zone_device(struct zone *zone,
continue;
memmap_init_compound(page, pfn, zone_idx, nid, pgmap,
- compound_nr_pages(altmap, pgmap));
+ compound_nr_pages(altmap, pgmap),
+ use_template);
}
pr_debug("%s initialised %lu pages in %ums\n", __func__,
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 6/8] string: introduce memcpy_streaming() helpers
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (4 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 5/8] mm: extend the template fast path to zone-device compound tails Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 7/8] x86/string: extend memcpy_flushcache() fixed-size fastpaths Li Zhe
` (2 subsequent siblings)
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
Introduce a generic memcpy_streaming() interface for write-once copy
sites that can fall back to memcpy() when no architecture-specific
optimization is available, or when an architecture-specific backend
cannot safely handle a given transfer.
Add memcpy_streaming_drain() alongside it so callers can separate the
copy primitive from any required ordering point. On x86, use
memcpy_flushcache() and sfence only for aligned transfers that can stay
entirely on the non-temporal store path; otherwise fall back to memcpy()
so the generic API does not expose flushcache semantics on cached
head/tail fragments.
Callers are responsible for invoking memcpy_streaming_drain() before
later normal stores that must be ordered after the streaming copy.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
arch/x86/include/asm/string_64.h | 40 ++++++++++++++++++++++++++++++++
include/linux/string.h | 20 ++++++++++++++++
2 files changed, 60 insertions(+)
diff --git a/arch/x86/include/asm/string_64.h b/arch/x86/include/asm/string_64.h
index 4635616863f5..0b57e9e6f3db 100644
--- a/arch/x86/include/asm/string_64.h
+++ b/arch/x86/include/asm/string_64.h
@@ -100,6 +100,46 @@ static __always_inline void memcpy_flushcache(void *dst, const void *src, size_t
}
__memcpy_flushcache(dst, src, cnt);
}
+
+/*
+ * Only reuse memcpy_flushcache() for transfers that can stay entirely
+ * on its non-temporal store path. Fall back to memcpy() for zero-length
+ * copies and for unaligned transfers so the generic streaming API does
+ * not expose flushcache semantics on cached head/tail fragments.
+ */
+static __always_inline int memcpy_flushcache_nt_safe(const void *dst,
+ const void *src,
+ size_t cnt)
+{
+ unsigned long d = (unsigned long)dst;
+ unsigned long s = (unsigned long)src;
+
+ if (!cnt)
+ return 0;
+
+ if (cnt >= 8)
+ return !(d & 7) && !(s & 7) && !(cnt & 7);
+
+ return cnt == 4 && !(d & 3) && !(s & 3);
+}
+
+#define __HAVE_ARCH_MEMCPY_STREAMING 1
+static __always_inline void memcpy_streaming(void *dst, const void *src,
+ size_t cnt)
+{
+ if (!cnt)
+ return;
+
+ if (memcpy_flushcache_nt_safe(dst, src, cnt))
+ memcpy_flushcache(dst, src, cnt);
+ else
+ memcpy(dst, src, cnt);
+}
+
+static __always_inline void memcpy_streaming_drain(void)
+{
+ asm volatile("sfence" : : : "memory");
+}
#endif
#endif /* __KERNEL__ */
diff --git a/include/linux/string.h b/include/linux/string.h
index b850bd91b3d8..a4c2d4347f58 100644
--- a/include/linux/string.h
+++ b/include/linux/string.h
@@ -281,6 +281,26 @@ static inline void memcpy_flushcache(void *dst, const void *src, size_t cnt)
}
#endif
+#ifndef __HAVE_ARCH_MEMCPY_STREAMING
+/*
+ * memcpy_streaming() is for write-once copy sites that may use
+ * non-temporal stores on some architectures. Callers must follow it
+ * with memcpy_streaming_drain() before later normal stores that need to
+ * be ordered after the streaming copy. Implementations may fall back to
+ * memcpy() when a specialized backend cannot safely handle the given
+ * transfer, and backends that use regular cached stores can make the
+ * drain a no-op.
+ */
+static inline void memcpy_streaming(void *dst, const void *src, size_t cnt)
+{
+ memcpy(dst, src, cnt);
+}
+
+static inline void memcpy_streaming_drain(void)
+{
+}
+#endif
+
void *memchr_inv(const void *s, int c, size_t n);
char *strreplace(char *str, char old, char new);
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 7/8] x86/string: extend memcpy_flushcache() fixed-size fastpaths
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (5 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 6/8] string: introduce memcpy_streaming() helpers Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 3:36 ` [PATCH v3 8/8] mm: use memcpy_streaming() in zone-device template copies Li Zhe
2026-05-27 20:46 ` [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Andrew Morton
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
Small constant-sized flushcache copies currently fall back to
__memcpy_flushcache() unless they are exactly 4, 8, or 16 bytes.
Factor the existing inline movnti sequences into small helpers and
extend the fixed-size fastpath coverage to 24..96 bytes for naturally
aligned transfers. This keeps common struct-page-sized copies on the
inline path for the upcoming memcpy_streaming() user, while still
falling back to __memcpy_flushcache() for unaligned or uncommon sizes.
Zero-length copies return immediately.
Issue the fixed-size stores in ascending address order so
write-combining sees a forward stream.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
arch/x86/include/asm/string_64.h | 125 ++++++++++++++++++++++++++-----
1 file changed, 107 insertions(+), 18 deletions(-)
diff --git a/arch/x86/include/asm/string_64.h b/arch/x86/include/asm/string_64.h
index 0b57e9e6f3db..8e6fca0185ee 100644
--- a/arch/x86/include/asm/string_64.h
+++ b/arch/x86/include/asm/string_64.h
@@ -82,24 +82,6 @@ int strcmp(const char *cs, const char *ct);
#ifdef CONFIG_ARCH_HAS_UACCESS_FLUSHCACHE
#define __HAVE_ARCH_MEMCPY_FLUSHCACHE 1
void __memcpy_flushcache(void *dst, const void *src, size_t cnt);
-static __always_inline void memcpy_flushcache(void *dst, const void *src, size_t cnt)
-{
- if (__builtin_constant_p(cnt)) {
- switch (cnt) {
- case 4:
- asm ("movntil %1, %0" : "=m"(*(u32 *)dst) : "r"(*(u32 *)src));
- return;
- case 8:
- asm ("movntiq %1, %0" : "=m"(*(u64 *)dst) : "r"(*(u64 *)src));
- return;
- case 16:
- asm ("movntiq %1, %0" : "=m"(*(u64 *)dst) : "r"(*(u64 *)src));
- asm ("movntiq %1, %0" : "=m"(*(u64 *)(dst + 8)) : "r"(*(u64 *)(src + 8)));
- return;
- }
- }
- __memcpy_flushcache(dst, src, cnt);
-}
/*
* Only reuse memcpy_flushcache() for transfers that can stay entirely
@@ -123,6 +105,113 @@ static __always_inline int memcpy_flushcache_nt_safe(const void *dst,
return cnt == 4 && !(d & 3) && !(s & 3);
}
+static __always_inline void memcpy_flushcache_4(void *dst, const void *src)
+{
+ asm volatile("movntil %1, %0"
+ : "=m"(*(u32 *)dst)
+ : "r"(*(const u32 *)src)
+ : "memory");
+}
+
+static __always_inline void memcpy_flushcache_8(void *dst, const void *src)
+{
+ asm volatile("movntiq %1, %0"
+ : "=m"(*(u64 *)dst)
+ : "r"(*(const u64 *)src)
+ : "memory");
+}
+
+static __always_inline void memcpy_flushcache_16(void *dst,
+ const void *src)
+{
+ memcpy_flushcache_8(dst, src);
+ memcpy_flushcache_8(dst + 8, src + 8);
+}
+
+static __always_inline void memcpy_flushcache_32(void *dst,
+ const void *src)
+{
+ memcpy_flushcache_16(dst, src);
+ memcpy_flushcache_16(dst + 16, src + 16);
+}
+
+static __always_inline void memcpy_flushcache_64(void *dst,
+ const void *src)
+{
+ memcpy_flushcache_32(dst, src);
+ memcpy_flushcache_32(dst + 32, src + 32);
+}
+
+/*
+ * Keep common fixed-size copies on the inline movnti path when they can
+ * stay entirely on aligned non-temporal stores. Issue the stores in
+ * ascending address order so write-combining sees a forward stream.
+ */
+static __always_inline int memcpy_flushcache_small(void *dst,
+ const void *src,
+ size_t cnt)
+{
+ char *d = dst;
+ const char *s = src;
+
+ if (!memcpy_flushcache_nt_safe(dst, src, cnt))
+ return 0;
+
+ switch (cnt) {
+ case 4:
+ memcpy_flushcache_4(d, s);
+ return 1;
+ case 8:
+ memcpy_flushcache_8(d, s);
+ return 1;
+ }
+
+ if (cnt & 8) {
+ memcpy_flushcache_8(d, s);
+ d += 8;
+ s += 8;
+ cnt -= 8;
+ }
+
+ switch (cnt) {
+ case 16:
+ memcpy_flushcache_16(d, s);
+ return 1;
+ case 32:
+ memcpy_flushcache_32(d, s);
+ return 1;
+ case 48:
+ memcpy_flushcache_32(d, s);
+ memcpy_flushcache_16(d + 32, s + 32);
+ return 1;
+ case 64:
+ memcpy_flushcache_64(d, s);
+ return 1;
+ case 80:
+ memcpy_flushcache_64(d, s);
+ memcpy_flushcache_16(d + 64, s + 64);
+ return 1;
+ case 96:
+ memcpy_flushcache_64(d, s);
+ memcpy_flushcache_32(d + 64, s + 64);
+ return 1;
+ }
+
+ return 0;
+}
+
+static __always_inline void memcpy_flushcache(void *dst, const void *src,
+ size_t cnt)
+{
+ if (!cnt)
+ return;
+
+ if (__builtin_constant_p(cnt) && memcpy_flushcache_small(dst, src, cnt))
+ return;
+
+ __memcpy_flushcache(dst, src, cnt);
+}
+
#define __HAVE_ARCH_MEMCPY_STREAMING 1
static __always_inline void memcpy_streaming(void *dst, const void *src,
size_t cnt)
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 8/8] mm: use memcpy_streaming() in zone-device template copies
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (6 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 7/8] x86/string: extend memcpy_flushcache() fixed-size fastpaths Li Zhe
@ 2026-05-27 3:36 ` Li Zhe
2026-05-27 20:46 ` [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Andrew Morton
8 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-27 3:36 UTC (permalink / raw)
To: akpm, apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx
Cc: linux-arch, linux-hardening, linux-kernel, linux-mm, x86, Li Zhe
The template fast path still leaves the actual copy sequence up to the
compiler. Use the streaming-copy helpers introduced in the previous
patches for the ZONE_DEVICE template-copy path so common mm code can
request a write-once copy primitive without embedding arch-specific
store layout in the generic layer.
ZONE_DEVICE memmap initialization is a write-once path: each struct page
is populated once and is not expected to be reused from cache
immediately afterwards. A regular cached copy can therefore incur
write-allocate traffic and pollute the cache without much benefit.
Using memcpy_streaming() lets this path use an architecture-optimized
streaming copy where available, while still degrading to memcpy() on
architectures that do not provide a specialized implementation.
Keep pageblock-aligned PFNs on memcpy() so pageblock initialization can
immediately read back page metadata without introducing a
read-after-streaming dependency. For the remaining PFNs, use
memcpy_streaming() so the hot path can avoid write-allocate traffic
while still leaving unsupported or unsuitable cases to the fallback
implementation.
When the streaming backend uses non-temporal stores, order them before
entering memmap_init_compound(), before prep_compound_head() updates the
overlapping compound metadata, and before returning from
memmap_init_zone_device().
Keep sanitized builds on the slow path so KASAN/KMSAN retain their
instrumented stores.
Tested in a VM with a 100 GB fsdax namespace device configured with
map=dev and a 100 GB devdax namespace (align=2097152) on Intel Ice Lake
server.
Test procedure:
Rebind the nd_pmem and dax_pmem driver 30 times and collect the memmap
initialization time from the pr_debug() output of
memmap_init_zone_device().
Base(v7.1-rc3):
First binding for nd_pmem driver: 1486 ms
Average of subsequent rebinds: 273.52 ms
First binding for dax_pmem driver: 1515 ms
Average of subsequent rebinds: 313.45 ms
With this series:
First binding for nd_pmem driver: 1285 ms
Average of subsequent rebinds: 114.31 ms
First binding for dax_pmem driver: 1331 ms
Average of subsequent rebinds: 99.37 ms
This reduces the average rebind time by about 58.2% for nd_pmem and
68.3% for dax_pmem.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
mm/mm_init.c | 47 +++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 45 insertions(+), 2 deletions(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index d5ccb49a048f..1f56765b92e1 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1070,11 +1070,21 @@ static void __ref zone_device_page_init_slow(struct page *page,
static inline bool zone_device_page_init_optimization_enabled(void)
{
+ /*
+ * Keep sanitized builds on the slow path so their stores stay
+ * instrumented.
+ */
+ if (IS_ENABLED(CONFIG_KASAN) || IS_ENABLED(CONFIG_KMSAN))
+ return false;
+
/*
* The template fast path copies a preinitialized struct page image.
* Skip it when the page_ref_set tracepoint is enabled.
*/
- return !page_ref_tracepoint_active(page_ref_set);
+ if (page_ref_tracepoint_active(page_ref_set))
+ return false;
+
+ return true;
}
static inline void zone_device_template_head_page_init(struct page *template,
@@ -1120,9 +1130,19 @@ static void zone_device_page_init_from_template(struct page *page,
* 'template' carries the invariant portion of a ZONE_DEVICE struct
* page. Update the PFN-dependent fields in place before copying it
* to the destination page.
+ *
+ * pageblock-aligned pages immediately feed
+ * init_pageblock_migratetype(), which reads back page metadata via
+ * helpers like page_zone(page). Avoid a read-after-streaming
+ * dependency for these rare pages by using regular cached stores
+ * instead of non-temporal ones.
*/
zone_device_page_update_template(template, pfn);
- memcpy(page, template, sizeof(*page));
+ if (unlikely(pageblock_aligned(pfn)))
+ memcpy(page, template, sizeof(*page));
+ else
+ memcpy_streaming(page, template, sizeof(*page));
+
zone_device_page_init_pageblock(page, pfn);
}
@@ -1184,6 +1204,15 @@ static void __ref memmap_init_compound(struct page *head,
prep_compound_tail(page, head, order);
set_page_count(page, 0);
}
+
+ /*
+ * prep_compound_head() updates compound metadata in struct folio fields
+ * that alias the first tail-page descriptors. When the tail pages above
+ * were populated with non-temporal stores, order those writes before the
+ * overlapping metadata updates below.
+ */
+ if (use_template)
+ memcpy_streaming_drain();
prep_compound_head(head, order);
}
@@ -1232,10 +1261,24 @@ void __ref memmap_init_zone_device(struct zone *zone,
if (pfns_per_compound == 1)
continue;
+ /*
+ * memmap_init_compound() immediately updates compound-head
+ * metadata. If the head-page template copy above used
+ * non-temporal stores, order them before entering the
+ * compound setup path.
+ */
+ if (use_template)
+ memcpy_streaming_drain();
+
memmap_init_compound(page, pfn, zone_idx, nid, pgmap,
compound_nr_pages(altmap, pgmap),
use_template);
}
+ /*
+ * Drain any remaining non-temporal stores before returning.
+ */
+ if (use_template)
+ memcpy_streaming_drain();
pr_debug("%s initialised %lu pages in %ums\n", __func__,
nr_pages, jiffies_to_msecs(jiffies - start));
--
2.20.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
` (7 preceding siblings ...)
2026-05-27 3:36 ` [PATCH v3 8/8] mm: use memcpy_streaming() in zone-device template copies Li Zhe
@ 2026-05-27 20:46 ` Andrew Morton
2026-05-28 7:06 ` Li Zhe
8 siblings, 1 reply; 11+ messages in thread
From: Andrew Morton @ 2026-05-27 20:46 UTC (permalink / raw)
To: Li Zhe
Cc: apopple, arnd, bp, dave.hansen, david, kees, mingo, rppt, tglx,
linux-arch, linux-hardening, linux-kernel, linux-mm, x86
On Wed, 27 May 2026 11:36:28 +0800 "Li Zhe" <lizhe.67@bytedance.com> wrote:
> This series reduces latency in paths which create or recreate
> ZONE_DEVICE memmaps. Pmem hotplug is a concrete example.
>
> memmap_init_zone_device() can spend a substantial amount of time
> initializing large ZONE_DEVICE ranges because it repeats nearly
> identical struct page setup for every PFN.
>
> The series addresses that overhead in eight patches.
Thanks. I'll take no action at this time, as it's getting late in the
-rc cycle and the patchset's review status is rather thin.
AI review had a number of things to say:
https://sashiko.dev/#/patchset/20260527033636.28231-1-lizhe.67@bytedance.com
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization
2026-05-27 20:46 ` [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Andrew Morton
@ 2026-05-28 7:06 ` Li Zhe
0 siblings, 0 replies; 11+ messages in thread
From: Li Zhe @ 2026-05-28 7:06 UTC (permalink / raw)
To: akpm
Cc: apopple, arnd, bp, dave.hansen, david, kees, linux-arch,
linux-hardening, linux-kernel, linux-mm, lizhe.67, mingo, rppt,
tglx, x86
On Wed, 27 May 2026 13:46:37 -0700, akpm@linux-foundation.org wrote:
> AI review had a number of things to say:
> https://sashiko.dev/#/patchset/20260527033636.28231-1-lizhe.67@bytedance.com
Thanks for the review. All four points are valid.
1. The head-page template setup should not run page initialization and
refcount helpers on a stack-resident struct page.
2. The compound-tail template path has the same issue and should avoid
running compound-tail/refcount setup on a stack object.
3. The x86 memcpy_streaming() eligibility check is too permissive, so some
transfers that cannot stay entirely on the non-temporal path may still
go through memcpy_flushcache().
4. memcpy_flushcache_small() should reject unsupported sizes before issuing
any store, otherwise fallback can observe partial writes.
I have prepared corresponding fixups on top of v3.
Thanks,
Zhe
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-05-28 7:06 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-05-27 3:36 [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Li Zhe
2026-05-27 3:36 ` [PATCH v3 1/8] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
2026-05-27 3:36 ` [PATCH v3 2/8] mm: factor zone-device page init helpers out of __init_zone_device_page Li Zhe
2026-05-27 3:36 ` [PATCH v3 3/8] mm: add a set_page_section_from_pfn() helper Li Zhe
2026-05-27 3:36 ` [PATCH v3 4/8] mm: add a template-based fast path for zone-device page init Li Zhe
2026-05-27 3:36 ` [PATCH v3 5/8] mm: extend the template fast path to zone-device compound tails Li Zhe
2026-05-27 3:36 ` [PATCH v3 6/8] string: introduce memcpy_streaming() helpers Li Zhe
2026-05-27 3:36 ` [PATCH v3 7/8] x86/string: extend memcpy_flushcache() fixed-size fastpaths Li Zhe
2026-05-27 3:36 ` [PATCH v3 8/8] mm: use memcpy_streaming() in zone-device template copies Li Zhe
2026-05-27 20:46 ` [PATCH v3 0/8] mm: speed up ZONE_DEVICE memmap initialization Andrew Morton
2026-05-28 7:06 ` Li Zhe
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®