mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/2] mm/migrate: speed up move_pages() node queries
@ 2026-10-02  1:25 Qiliang Yuan
  2026-10-02  1:25 ` [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array() Qiliang Yuan
  2026-10-02  1:25 ` [PATCH v2 2/2] mm/migrate: raise the do_pages_stat() chunk to 512 pages Qiliang Yuan
  0 siblings, 2 replies; 4+ messages in thread
From: Qiliang Yuan @ 2026-10-02  1:25 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Zi Yan, Matthew Brost,
	Joshua Hahn, Byungchul Park, Gregory Price, Ying Huang,
	Alistair Popple
  Cc: linux-mm, linux-kernel, Qiliang Yuan

move_pages() with a NULL node list reports the node of each page, and
RDMA and KV-cache engines call it on every page of buffers spanning
hundreds of gigabytes. do_pages_stat_array() walks the page tables from
the top for every address, at about 105 ns per page.

Patch 1 walks runs of consecutive pages with walk_page_range(). Patch 2
raises the do_pages_stat() chunk from 16 to 512 pages so that a run can
cover a whole PTE table. Together they bring the per-page cost from
104 ns to 12.3 ns with 4K pages and from 90.6 ns to 2.0 ns with THP,
and a 16 GiB buffer of 4K pages from 486 ms to 52 ms.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
V1 -> V2:
- Walk the runs with walk_page_range() and a pmd_entry callback instead
  of extending folio_walk_start() by hand (Zi Yan)
- Raise DO_PAGES_STAT_CHUNK_NR from 16 to 512 in a separate patch, with
  numbers for each chunk size (Zi Yan)
- Re-measure everything in one run on a VM that now has two NUMA nodes,
  with the test bound to CPU 0, so the baselines differ from v1 (16 GiB
  of 4K pages: 486 ms instead of 440 ms)

v1: https://lore.kernel.org/r/20261001-bug-mm-move-pages-stat-batch-v1-1-255b7e915744@gmail.com

---
Qiliang Yuan (2):
      mm/migrate: walk runs of consecutive pages in do_pages_stat_array()
      mm/migrate: raise the do_pages_stat() chunk to 512 pages

 mm/migrate.c | 206 ++++++++++++++++++++++++++++++++++++++++++++++++++---------
 1 file changed, 177 insertions(+), 29 deletions(-)
---
base-commit: 551c722f40809618230001baccf219193e22fc5a
change-id: 20261001-bug-mm-move-pages-stat-batch-f62a29ea5867

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array()
  2026-10-02  1:25 [PATCH v2 0/2] mm/migrate: speed up move_pages() node queries Qiliang Yuan
@ 2026-10-02  1:25 ` Qiliang Yuan
  2026-10-02 18:53   ` David Hildenbrand (Arm)
  2026-10-02  1:25 ` [PATCH v2 2/2] mm/migrate: raise the do_pages_stat() chunk to 512 pages Qiliang Yuan
  1 sibling, 1 reply; 4+ messages in thread
From: Qiliang Yuan @ 2026-10-02  1:25 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Zi Yan, Matthew Brost,
	Joshua Hahn, Byungchul Park, Gregory Price, Ying Huang,
	Alistair Popple
  Cc: linux-mm, linux-kernel, Qiliang Yuan

move_pages() with a NULL node list reports the node of each page. RDMA
and KV-cache transfer engines use it to find where large registered
buffers live, querying every 4K page of buffers that span hundreds of
gigabytes.

do_pages_stat_array() looks up the VMA and walks the page tables from
the top for every address, taking and dropping the PTE lock each time.
That costs about 105 ns per page, so a 16 GiB buffer takes 486 ms.

Callers almost always pass consecutive addresses. Group them into runs
and walk each run with walk_page_range(), which looks up each VMA and
PTE table once and answers every page under it while holding the lock.
Report pages as folio_walk_start() with FW_ZEROPAGE found them: the node
of a normal folio, -EFAULT for the zero page or an address outside any
VMA, and -ENOENT otherwise. Handle PUD and hugetlb leaves in their own
callbacks so that the walk never splits them.

On 7.3-rc5 in a 16-vCPU VM, querying every page of a populated 4 GiB
buffer:

                before    after
  4K pages      105 ns    31.9 ns
  THP           90 ns     23.1 ns

Suggested-by: Zi Yan <ziy@nvidia.com>
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 mm/migrate.c | 188 ++++++++++++++++++++++++++++++++++++++++++++++++++---------
 1 file changed, 162 insertions(+), 26 deletions(-)

diff --git a/mm/migrate.c b/mm/migrate.c
index 15b45832bcfa7..f4d8bfb9b7b4a 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -2451,44 +2451,180 @@ static int do_pages_move(struct mm_struct *mm, nodemask_t task_nodes,
 	return err;
 }
 
+struct pages_stat_walk {
+	unsigned long start;
+	int *status;
+};
+
+static void pages_stat_set(struct pages_stat_walk *psw, unsigned long addr,
+			   unsigned long end, int stat)
+{
+	int *status = psw->status + ((addr - psw->start) >> PAGE_SHIFT);
+
+	/* end wraps to 0 for the last page of the address space */
+	for (; addr != end; addr += PAGE_SIZE)
+		*status++ = stat;
+}
+
+static int folio_stat(struct folio *folio)
+{
+	if (is_zero_folio(folio) || is_huge_zero_folio(folio))
+		return -EFAULT;
+	if (folio_is_zone_device(folio))
+		return -ENOENT;
+	return folio_nid(folio);
+}
+
+/* Report pages the same way folio_walk_start() with FW_ZEROPAGE finds them. */
+static int pages_stat_pud_entry(pud_t *pudp, unsigned long addr,
+				unsigned long end, struct mm_walk *walk)
+{
+	struct page *page;
+	spinlock_t *ptl;
+	pud_t pud;
+	int stat;
+
+	if (!IS_ENABLED(CONFIG_PGTABLE_HAS_HUGE_LEAVES))
+		return 0;
+	pud = pudp_get(pudp);
+	if (pud_present(pud) && !pud_leaf(pud))
+		return 0;
+
+	ptl = pud_lock(walk->mm, pudp);
+	pud = pudp_get(pudp);
+	if (pud_present(pud) && !pud_leaf(pud)) {
+		spin_unlock(ptl);
+		return 0;
+	}
+	stat = -ENOENT;
+	if (pud_present(pud)) {
+		page = vm_normal_page_pud(walk->vma, addr, pud);
+		if (page)
+			stat = folio_stat(page_folio(page));
+	}
+	pages_stat_set(walk->private, addr, end, stat);
+	spin_unlock(ptl);
+	walk->action = ACTION_CONTINUE;
+	return 0;
+}
+
+static int pages_stat_pmd_entry(pmd_t *pmdp, unsigned long addr,
+				unsigned long end, struct mm_walk *walk)
+{
+	struct vm_area_struct *vma = walk->vma;
+	struct page *page;
+	spinlock_t *ptl;
+	pte_t *ptep;
+	pmd_t pmd;
+	int stat;
+
+	pmd = pmdp_get_lockless(pmdp);
+	if (IS_ENABLED(CONFIG_PGTABLE_HAS_HUGE_LEAVES) &&
+	    (!pmd_present(pmd) || pmd_leaf(pmd))) {
+		ptl = pmd_lock(walk->mm, pmdp);
+		pmd = pmdp_get(pmdp);
+		if (pmd_present(pmd) && !pmd_leaf(pmd)) {
+			spin_unlock(ptl);
+			goto pte_table;
+		}
+		stat = -ENOENT;
+		if (pmd_present(pmd)) {
+			page = vm_normal_page_pmd(vma, addr, pmd);
+			if (page)
+				stat = folio_stat(page_folio(page));
+			else if (is_huge_zero_pmd(pmd))
+				stat = -EFAULT;
+		}
+		pages_stat_set(walk->private, addr, end, stat);
+		spin_unlock(ptl);
+		return 0;
+	}
+
+pte_table:
+	ptep = pte_offset_map_lock(walk->mm, pmdp, addr, &ptl);
+	if (!ptep) {
+		walk->action = ACTION_AGAIN;
+		return 0;
+	}
+	for (; addr < end; addr += PAGE_SIZE, ptep++) {
+		pte_t pte = ptep_get(ptep);
+
+		stat = -ENOENT;
+		if (pte_present(pte)) {
+			page = vm_normal_page(vma, addr, pte);
+			if (page)
+				stat = folio_stat(page_folio(page));
+			else if (is_zero_pfn(pte_pfn(pte)))
+				stat = -EFAULT;
+		}
+		pages_stat_set(walk->private, addr, addr + PAGE_SIZE, stat);
+	}
+	pte_unmap_unlock(ptep - 1, ptl);
+	return 0;
+}
+
+static int pages_stat_hugetlb_entry(pte_t *ptep, unsigned long hmask,
+				    unsigned long addr, unsigned long end,
+				    struct mm_walk *walk)
+{
+#ifdef CONFIG_HUGETLB_PAGE
+	spinlock_t *ptl;
+	pte_t pte;
+	int stat = -ENOENT;
+
+	ptl = huge_pte_lock(hstate_vma(walk->vma), walk->mm, ptep);
+	pte = huge_ptep_get(walk->mm, addr, ptep);
+	if (pte_present(pte))
+		stat = folio_stat(pfn_folio(pte_pfn(pte)));
+	pages_stat_set(walk->private, addr, end, stat);
+	spin_unlock(ptl);
+#endif
+	return 0;
+}
+
+static int pages_stat_pte_hole(unsigned long addr, unsigned long end,
+			       int depth, struct mm_walk *walk)
+{
+	/* No VMA at all is -EFAULT, a VMA without the page is -ENOENT */
+	pages_stat_set(walk->private, addr, end, walk->vma ? -ENOENT : -EFAULT);
+	return 0;
+}
+
+static const struct mm_walk_ops pages_stat_walk_ops = {
+	.pud_entry	= pages_stat_pud_entry,
+	.pmd_entry	= pages_stat_pmd_entry,
+	.hugetlb_entry	= pages_stat_hugetlb_entry,
+	.pte_hole	= pages_stat_pte_hole,
+	.walk_lock	= PGWALK_RDLOCK,
+};
+
 /*
  * Determine the nodes of an array of pages and store it in an array of status.
  */
 static void do_pages_stat_array(struct mm_struct *mm, unsigned long nr_pages,
 				const void __user **pages, int *status)
 {
-	unsigned long i;
+	unsigned long i, n;
 
 	mmap_read_lock(mm);
 
-	for (i = 0; i < nr_pages; i++) {
-		unsigned long addr = (unsigned long)(*pages);
-		struct vm_area_struct *vma;
-		struct folio_walk fw;
-		struct folio *folio;
-		int err = -EFAULT;
+	for (i = 0; i < nr_pages; i += n) {
+		unsigned long addr = (unsigned long)pages[i] & PAGE_MASK;
+		struct pages_stat_walk psw = {
+			.start = addr,
+			.status = status + i,
+		};
 
-		vma = vma_lookup(mm, addr);
-		if (!vma)
-			goto set_status;
+		/* Walk runs of consecutive pages in one go */
+		for (n = 1; i + n < nr_pages; n++) {
+			unsigned long next = (unsigned long)pages[i + n] & PAGE_MASK;
 
-		folio = folio_walk_start(&fw, vma, addr, FW_ZEROPAGE);
-		if (folio) {
-			if (is_zero_folio(folio) || is_huge_zero_folio(folio))
-				err = -EFAULT;
-			else if (folio_is_zone_device(folio))
-				err = -ENOENT;
-			else
-				err = folio_nid(folio);
-			folio_walk_end(&fw, vma);
-		} else {
-			err = -ENOENT;
+			if (next != addr + n * PAGE_SIZE || next < addr)
+				break;
 		}
-set_status:
-		*status = err;
-
-		pages++;
-		status++;
+		if (walk_page_range(mm, addr, addr + n * PAGE_SIZE,
+				    &pages_stat_walk_ops, &psw))
+			pages_stat_set(&psw, addr, addr + n * PAGE_SIZE, -EFAULT);
 	}
 
 	mmap_read_unlock(mm);

-- 
2.43.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH v2 2/2] mm/migrate: raise the do_pages_stat() chunk to 512 pages
  2026-10-02  1:25 [PATCH v2 0/2] mm/migrate: speed up move_pages() node queries Qiliang Yuan
  2026-10-02  1:25 ` [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array() Qiliang Yuan
@ 2026-10-02  1:25 ` Qiliang Yuan
  1 sibling, 0 replies; 4+ messages in thread
From: Qiliang Yuan @ 2026-10-02  1:25 UTC (permalink / raw)
  To: Andrew Morton, David Hildenbrand, Zi Yan, Matthew Brost,
	Joshua Hahn, Byungchul Park, Gregory Price, Ying Huang,
	Alistair Popple
  Cc: linux-mm, linux-kernel, Qiliang Yuan

do_pages_stat() copies the addresses in and the status out in chunks of
16 on the stack, and do_pages_stat_array() walks each chunk on its own.
With runs of consecutive pages walked in one go, the walk setup paid
every 16 pages now dominates the cost.

Raise the chunk to 512 pages, the number of pages a PTE table maps with
4K pages on x86-64, and allocate the two arrays rather than keeping
6 KiB on the stack. A failed allocation returns -ENOMEM, which
move_pages() already documents.

On 7.3-rc5 in a 16-vCPU VM, querying every page of a populated 4 GiB
buffer, per page and by chunk size:

                16        64        256       512       1024
  4K pages      31.9 ns   15.0 ns   11.8 ns   11.2 ns   10.9 ns
  THP           23.1 ns   6.6 ns    2.5 ns    1.8 ns    1.6 ns

Suggested-by: Zi Yan <ziy@nvidia.com>
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 mm/migrate.c | 18 +++++++++++++++---
 1 file changed, 15 insertions(+), 3 deletions(-)

diff --git a/mm/migrate.c b/mm/migrate.c
index f4d8bfb9b7b4a..9c77b5e5cf1d5 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -2656,10 +2656,20 @@ static int do_pages_stat(struct mm_struct *mm, unsigned long nr_pages,
 			 const void __user * __user *pages,
 			 int __user *status)
 {
-#define DO_PAGES_STAT_CHUNK_NR 16UL
-	const void __user *chunk_pages[DO_PAGES_STAT_CHUNK_NR];
-	int chunk_status[DO_PAGES_STAT_CHUNK_NR];
+#define DO_PAGES_STAT_CHUNK_NR 512UL
+	const void __user **chunk_pages;
 	unsigned long chunk_offset = 0;
+	int *chunk_status;
+
+	chunk_pages = kmalloc_array(DO_PAGES_STAT_CHUNK_NR,
+				    sizeof(*chunk_pages), GFP_KERNEL);
+	chunk_status = kmalloc_array(DO_PAGES_STAT_CHUNK_NR,
+				     sizeof(*chunk_status), GFP_KERNEL);
+	if (!chunk_pages || !chunk_status) {
+		kfree(chunk_pages);
+		kfree(chunk_status);
+		return -ENOMEM;
+	}
 
 	while (nr_pages) {
 		unsigned long chunk_nr = min(nr_pages, DO_PAGES_STAT_CHUNK_NR);
@@ -2683,6 +2693,8 @@ static int do_pages_stat(struct mm_struct *mm, unsigned long nr_pages,
 		chunk_offset += chunk_nr;
 		nr_pages -= chunk_nr;
 	}
+	kfree(chunk_pages);
+	kfree(chunk_status);
 	return nr_pages ? -EFAULT : 0;
 }
 

-- 
2.43.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array()
  2026-10-02  1:25 ` [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array() Qiliang Yuan
@ 2026-10-02 18:53   ` David Hildenbrand (Arm)
  0 siblings, 0 replies; 4+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-02 18:53 UTC (permalink / raw)
  To: Qiliang Yuan, Andrew Morton, Zi Yan, Matthew Brost, Joshua Hahn,
	Byungchul Park, Gregory Price, Ying Huang, Alistair Popple
  Cc: linux-mm, linux-kernel

On 10/2/26 03:25, Qiliang Yuan wrote:
> move_pages() with a NULL node list reports the node of each page. RDMA
> and KV-cache transfer engines use it to find where large registered
> buffers live, querying every 4K page of buffers that span hundreds of
> gigabytes.
> 
> do_pages_stat_array() looks up the VMA and walks the page tables from
> the top for every address, taking and dropping the PTE lock each time.
> That costs about 105 ns per page, so a 16 GiB buffer takes 486 ms.
> 
> Callers almost always pass consecutive addresses. Group them into runs
> and walk each run with walk_page_range(), which looks up each VMA and
> PTE table once and answers every page under it while holding the lock.
> Report pages as folio_walk_start() with FW_ZEROPAGE found them: the node
> of a normal folio, -EFAULT for the zero page or an address outside any
> VMA, and -ENOENT otherwise. Handle PUD and hugetlb leaves in their own
> callbacks so that the walk never splits them.
> 
> On 7.3-rc5 in a 16-vCPU VM, querying every page of a populated 4 GiB
> buffer:
> 
>                 before    after
>   4K pages      105 ns    31.9 ns
>   THP           90 ns     23.1 ns
> 
> Suggested-by: Zi Yan <ziy@nvidia.com>
> Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
> ---
>  mm/migrate.c | 188 ++++++++++++++++++++++++++++++++++++++++++++++++++---------
>  1 file changed, 162 insertions(+), 26 deletions(-)

That's a lot of churn ... which is really a shame, because all we want to walk
is folio ranges.

Oscar was working on a better page table walker API (but I was too busy to
provide review so far :( ), which sounds strategically like the better long-term
solution.

Is there particular need to optimize this in the ns range for 4 KiB of memory
immediately?

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-02 18:53 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02  1:25 [PATCH v2 0/2] mm/migrate: speed up move_pages() node queries Qiliang Yuan
2026-10-02  1:25 ` [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array() Qiliang Yuan
2026-10-02 18:53   ` David Hildenbrand (Arm)
2026-10-02  1:25 ` [PATCH v2 2/2] mm/migrate: raise the do_pages_stat() chunk to 512 pages Qiliang Yuan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®