From: Wen Jiang <jiangwenxiaomi@gmail.com>
To: akpm@linux-foundation.org, catalin.marinas@arm.com,
linux-mm@kvack.org, urezki@gmail.com, will@kernel.org
Cc: Xueyuan.chen21@gmail.com, ajd@linux.ibm.com,
anshuman.khandual@arm.com, baohua@kernel.org, chleroy@kernel.org,
david@kernel.org, dev.jain@arm.com, jiangwen6@xiaomi.com,
leo.yan@arm.com, linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
maddy@linux.ibm.com, mpe@ellerman.id.au, npiggin@gmail.com,
rppt@kernel.org, ryan.roberts@arm.com,
Xueyuan Chen <xueyuan.chen21@gmail.com>
Subject: [PATCH v8 09/10] mm/vmalloc: map contiguous pages in batches for vmap() if possible
Date: Thu, 17 Sep 2026 13:29:32 +0800 [thread overview]
Message-ID: <20260917052933.188679-10-jiangwenxiaomi@gmail.com> (raw)
In-Reply-To: <20260917052933.188679-1-jiangwenxiaomi@gmail.com>
From: "Barry Song (Xiaomi)" <baohua@kernel.org>
In many cases, the pages passed to vmap() may include high-order
pages. For example, the systemheap often allocates pages in descending
order: order 8, then 4, then 0. Currently, vmap() iterates over every
page individually—even pages inside a high-order block are handled
one by one.
This patch detects physically contiguous pages (regardless of whether
they are compound or non-compound) by scanning with
num_pages_contiguous(), and maps them as a single contiguous block
whenever possible. The mapping order is determined by taking the
minimum of the contiguous page count and the pfn alignment, allowing
graceful degradation when pfn alignment is less than the contiguous
range.
Pages with the same page_shift are coalesced and mapped via
vmap_pages_range_noflush_walk() to avoid page table rewalk.
As users typically allocate memory in descending orders (e.g.
8 → 4 → 0), once an order-0 page is encountered, we stop scanning
for contiguous pages since subsequent pages are likely order-0 as well.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Co-developed-by: Dev Jain <dev.jain@arm.com>
Signed-off-by: Dev Jain <dev.jain@arm.com>
Signed-off-by: Wen Jiang <jiangwen6@xiaomi.com>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
Tested-by: Leo Yan <leo.yan@arm.com>
---
mm/vmalloc.c | 82 ++++++++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 80 insertions(+), 2 deletions(-)
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index ed868afd8194d..bce2d2c09f546 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -3540,6 +3540,85 @@ static inline unsigned int vm_shift(pgprot_t prot, unsigned long size)
return arch_vmap_pte_supported_shift(size);
}
+static inline int get_vmap_batch_order(struct page **pages,
+ pgprot_t prot, unsigned int nr_pages)
+{
+ unsigned long pfn;
+ unsigned int nr_contig;
+ int order;
+
+ if (!IS_ENABLED(CONFIG_HAVE_ARCH_HUGE_VMAP))
+ return 0;
+
+ /* Limit nr_pages by pfn alignment */
+ pfn = page_to_pfn(*pages);
+ if (pfn > 0)
+ nr_pages = min_t(size_t, nr_pages, 1UL << __ffs(pfn));
+
+ nr_contig = num_pages_contiguous(pages, nr_pages);
+ if (nr_contig < 2)
+ return 0;
+
+ order = ilog2(nr_contig);
+
+ if (vm_shift(prot, PAGE_SIZE << order) == PAGE_SHIFT)
+ return 0;
+
+ return order;
+}
+
+static int vmap_pages_range_batched(unsigned long addr, unsigned long end,
+ pgprot_t prot, struct page **pages)
+{
+ const unsigned int nr_pages = (end - addr) >> PAGE_SHIFT;
+ unsigned int prev_shift = 0, batch_idx = 0;
+ unsigned long batch_start = addr, batch_end = addr;
+ int err;
+
+ err = kmsan_vmap_pages_range_noflush(addr, end, prot, pages,
+ PAGE_SHIFT, GFP_KERNEL);
+
+ if (err)
+ goto out;
+
+ for (unsigned int i = 0; i < nr_pages; ) {
+ unsigned int shift = PAGE_SHIFT +
+ get_vmap_batch_order(pages + i, prot, nr_pages - i);
+
+ if (!i)
+ prev_shift = shift;
+
+ if (shift != prev_shift) {
+ err = vmap_pages_range_noflush_walk(batch_start, batch_end,
+ prot, pages + batch_idx, prev_shift);
+ if (err)
+ goto out;
+ prev_shift = shift;
+ batch_start = batch_end;
+ batch_idx = i;
+ }
+
+ /*
+ * Once we fail to batch pages, we expect to fail batching
+ * for all remaining pages, so just give up.
+ */
+ if (shift == PAGE_SHIFT)
+ break;
+
+ batch_end += 1UL << shift;
+ i += 1U << (shift - PAGE_SHIFT);
+ }
+
+ /* Remaining */
+ if (batch_start < end)
+ err = vmap_pages_range_noflush_walk(batch_start, end, prot,
+ pages + batch_idx, prev_shift);
+
+out:
+ flush_cache_vmap(addr, end);
+ return err;
+}
+
/**
* vmap - map an array of pages into virtually contiguous space
* @pages: array of page pointers
@@ -3583,8 +3662,7 @@ void *vmap(struct page **pages, unsigned int count,
return NULL;
addr = (unsigned long)area->addr;
- if (vmap_pages_range(addr, addr + size, pgprot_nx(prot),
- pages, PAGE_SHIFT) < 0) {
+ if (vmap_pages_range_batched(addr, addr + size, pgprot_nx(prot), pages) < 0) {
vunmap(area->addr);
return NULL;
}
--
2.34.1
next prev parent reply other threads:[~2026-09-17 5:30 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 5:29 [PATCH v8 00/10] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory Wen Jiang
2026-09-17 5:29 ` [PATCH v8 01/10] arm64/mm: add pte_set_huge() and pte_clear_huge() Wen Jiang
2026-09-17 23:22 ` Barry Song
2026-09-17 5:29 ` [PATCH v8 02/10] powerpc/8xx: add pte_set_huge() Wen Jiang
2026-09-17 23:28 ` Barry Song
2026-09-18 6:14 ` Christophe Leroy (CS GROUP)
2026-09-17 5:29 ` [PATCH v8 03/10] mm/vmalloc: use pte_set_huge()/pte_clear_huge() for PTE-level block mappings Wen Jiang
2026-09-17 14:00 ` Christophe Leroy (CS GROUP)
2026-09-17 14:40 ` Wen Jiang
2026-09-17 21:44 ` Barry Song
2026-09-18 6:05 ` Christophe Leroy (CS GROUP)
2026-09-18 6:14 ` Barry Song
2026-09-17 5:29 ` [PATCH v8 04/10] arm64/hugetlb: drop the init_mm special case in clear_flush() Wen Jiang
2026-09-17 5:29 ` [PATCH v8 05/10] arm64/vmalloc: allow arch_vmap_pte_range_map_size() to batch multiple CONT_PTE Wen Jiang
2026-09-17 5:29 ` [PATCH v8 06/10] mm/vmalloc: extract vmap_set_ptes() to consolidate PTE mapping logic Wen Jiang
2026-09-17 5:29 ` [PATCH v8 07/10] mm/vmalloc: extend page table walk to support larger page_shift sizes and eliminate page table rewalk Wen Jiang
2026-09-17 5:29 ` [PATCH v8 08/10] mm/vmalloc: extract vm_shift() to consolidate mapping shift selection Wen Jiang
2026-09-17 5:29 ` Wen Jiang [this message]
2026-09-17 5:29 ` [PATCH v8 10/10] mm/vmalloc: align vm_area so vmap() can batch mappings Wen Jiang
2026-09-17 12:17 ` [PATCH v8 00/10] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory Christophe Leroy (CS GROUP)
2026-09-17 13:55 ` Christophe Leroy (CS GROUP)
2026-09-17 14:30 ` Wen Jiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917052933.188679-10-jiangwenxiaomi@gmail.com \
--to=jiangwenxiaomi@gmail.com \
--cc=Xueyuan.chen21@gmail.com \
--cc=ajd@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=anshuman.khandual@arm.com \
--cc=baohua@kernel.org \
--cc=catalin.marinas@arm.com \
--cc=chleroy@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jiangwen6@xiaomi.com \
--cc=leo.yan@arm.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=urezki@gmail.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®