From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E0CBD175A9D for ; Thu, 17 Sep 2026 05:29:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789622988; cv=none; b=YmKyO6+vzfz29JLg8sK09zOCSHUiZ5S9A9nfE63MUSw7h7nJPFN4R5u4P00Qv7MMrpq/gkUMuwQNL5SmXgg/Fd4OyFL6qIu2bYCLtlC4452qg00YCff3VqYRcPzWOcdHMREfW/7NAwCtoIipg/69s5kYmmHryBNjE4WxjNab7LI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789622988; c=relaxed/simple; bh=TH7YWc7d+oj49FybemXW2CasK4FRw1bRQKCWGIldqAs=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version:Content-Type; b=qkegqkh9uScvIApYzmhEK0O6vgHh1NuZeL7/rAp5SM+WmULfnNphg4+UiIEE1Lm+zAp4zHqVoeEN/tuGcrouZdzp2dd7YdkLmfvZLf4PWJCAQ7FLUpJcK5eXMaPCtTdu/QAnK3/A7a+GxbaKSEFt/G5RGT5BPz5qYwa+ss8GniU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pkaiN5jg; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pkaiN5jg" Received: by mail-pj2-f12.google.com with SMTP id 98e67ed59e1d1-396cccbba92so359096a91.0 for ; Wed, 16 Sep 2026 22:29:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789622986; x=1790227786; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:message-id:date :subject:cc:to:from:from:to:cc:subject:date:message-id:reply-to :content-type; bh=r+W2BgF0JXvbhEnB5mATNx7NuZIAwz407+caeqGwqIQ=; b=pkaiN5jgfcOL9Km4pRzGzQz627noXV6wfwbdOBLMOwE48reOQLlAP+gBG/wMlyU9s0 eFRy6vMkMyZ/4aezN/5tBjQ9omI2AcqP9BgKdQ17HVlzb939NgPidLcQI6zF1G9sUOte HYyb6Le++l3cQC+1fmdx8yngke0Z9m5cisWHYiZmer8RzY4oacNqfFnqudFp/c1YOBKT XpWJB4WbVYqiGy8E/nv2tYYjUIzV217VIq0Lh8IlttD5EVIFDi2wnI40o9HMQxSJV2LO GeADFmv/hmS00+A697BYGxV+BRc+CSQEwH5x9UB01T6h/Ql4gqGECZuIWLc2sv7j+GxF m4oA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789622986; x=1790227786; h=content-transfer-encoding:content-type:mime-version:message-id:date :subject:cc:to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject :date:message-id:reply-to:content-type; bh=r+W2BgF0JXvbhEnB5mATNx7NuZIAwz407+caeqGwqIQ=; b=FqJUIVKsY/3DmeLE12XlkNRuR6p57XxuA0KqxqCkUFM0ubcEzW8pzfPGsfQv1BhxQ8 M4g6hQ+XlvaSms6gNzXajNqKPv3i2cCovxnRo8khYzlexcBqln78lemkEN3iEIEYbPQy wuuoSUWk98PQyMNJW2pHzGLwzjh2jyVF07SBHJGA0ctMInJLrfa4jCxkhyKDwyHi4R8R UlVIiYrjCyiu01OuCA6xkkaNFJsQ6qinJ9v/gm9Qcwm3EyqddT4sF4prrLmW1hX5qwqY YhAbLwsU4pszG1b5go6OYT2WdkxDkSE/XzquPCxf9JoImziYi4xnjrfQ61P0JvD0DyZs mw3w== X-Forwarded-Encrypted: i=1; AKwUvBwDe7ox6sW0cHJIx8PnFU7PkPRMjGINBNjG1u6DljszcOT6Yw94I3ZeKMzhh7jI6vbaKB4lABAu/SCUc68=@vger.kernel.org X-Gm-Message-State: AFuF++kBlUNU2tmQa0ZGHgQXewEIMuU2dwX7gcZEOY+V1riqT082R04p xPlteHmzr1Iv9yslZXocKEuZV/C+XD5VXrF0QCsoIYFchZVxTn+hOfdn X-Gm-Gg: AYBFou122Y+9X7rl+nRq5bqM6i9wx6EN2Wyj1qhdzBeP3/9m41Th2V+pfs8VmTAIMfb G0jv/uO0NKSTGXtKolKP7eagbsKTmz/zaoHxvN8Ib2elvIPWGZXeROGvF5YC9mtU3Blklw3uzMa TCIPspSBlxci+cB2/xTag0Npk6yATtOMc++Qhs5eKUrrHSv78jEyPuNlmsHtubJD3FomgS+qksU x8ZOVAoV8M9BA2dJHL8L4hUSTfCHywOA8HD8OHC+Ssng97px/QzJUJ8NoV2zh1yhx6MA/r4nExP 5B26NOaUsvtDuaD0cbDZEaoiCarAgA2hJ2q0qCuX1a5OKhPKSXOnyzMs9toRlBxuCh5h3glUaUo 1nOHYbR+1LzRX1V//6jOWbGNrqjUCrPXfMloIfMoj54yTqyrnOFnQgVOCaG/IgJGoDeW6Ba80Mq xmsrnv7OaEfp/G+ty9GFZUpqG2kJm0Kr6UHIiG3S34jh7U014JbUzhWH0hfXLdnLJ66Rr/uPTbC L+Cz2PJ293366AZWnaQuDxSTCg= X-Received: by 2002:a17:90b:280a:b0:39d:fcbe:fdcf with SMTP id 98e67ed59e1d1-39e1e2659a8mr13162798a91.1.1789622986129; Wed, 16 Sep 2026 22:29:46 -0700 (PDT) Received: from mi-OptiPlex-7060.mioffice.cn ([43.224.245.234]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-39e3abada87sm1955836a91.14.2026.09.16.22.29.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 16 Sep 2026 22:29:44 -0700 (PDT) From: Wen Jiang To: akpm@linux-foundation.org, catalin.marinas@arm.com, linux-mm@kvack.org, urezki@gmail.com, will@kernel.org Cc: Xueyuan.chen21@gmail.com, ajd@linux.ibm.com, anshuman.khandual@arm.com, baohua@kernel.org, chleroy@kernel.org, david@kernel.org, dev.jain@arm.com, jiangwen6@xiaomi.com, leo.yan@arm.com, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, maddy@linux.ibm.com, mpe@ellerman.id.au, npiggin@gmail.com, rppt@kernel.org, ryan.roberts@arm.com Subject: [PATCH v8 00/10] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory Date: Thu, 17 Sep 2026 13:29:23 +0800 Message-Id: <20260917052933.188679-1-jiangwenxiaomi@gmail.com> X-Mailer: git-send-email 2.34.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit From: Wen Jiang This patchset accelerates ioremap, vmalloc, and vmap when the memory is physically fully or partially contiguous. Two techniques are used: 1. Avoid page table rewalk when setting PTEs/PMDs for multiple memory segments 2. Use batched mappings wherever possible in both vmalloc and ARM64 layers Besides accelerating the mapping path, this also enables large mappings (PMD and cont-PTE) for vmap, which are currently not supported. Patches 1-4 decouple the PTE-level block mapping path from HugeTLB. Previously vmap_pte_range() installed cont-PTE mappings by reusing set_huge_pte_at() and huge_ptep_get_and_clear(), which are HugeTLB helpers gated by CONFIG_HUGETLB_PAGE. This made the feature silently unavailable on CONFIG_HUGETLB_PAGE=n kernels, and coupled mm/vmalloc.c to HugeTLB internals. Patches 1 and 2 add the arm64 and powerpc/8xx implementations of pte_set_huge()/pte_clear_huge(), which join the existing pmd/pud_set_huge() family. Patch 3 adds the generic fallbacks and converts mm/vmalloc.c over. Patch 4 then removes the now-dead init_mm special case from arm64's clear_flush(). Patch 5 extends ARM64 arch_vmap_pte_range_map_size() to batch multiple CONT_PTE blocks in one call instead of one at a time. Patch 6 extracts a common helper vmap_set_ptes() that consolidates PTE mapping logic for the ioremap and vmalloc/vmap paths, handling both CONT_PTE and regular PTE mappings. This prepares for the next patch. Patch 7 extends the page table walk path to support page shifts other than PAGE_SHIFT and eliminates the page table rewalk for huge vmalloc mappings. The function is renamed from vmap_small_pages_range_noflush() to vmap_pages_range_noflush_walk(). Patch 8 extracts vm_shift() to consolidate vmalloc mapping shift selection for reuse in the batching path. Patches 9-10 add huge vmap support for contiguous pages, including support for non-compound pages with pfn alignment verification. On the RK3588 8-core ARM64 SoC, with tasks pinned to a little core and the performance CPUfreq policy enabled, benchmark results: * ioremap(1 MB): 1.35x faster (3407 ns -> 2526 ns) * vmalloc(1 MB) mapping time (excluding allocation) with VM_ALLOW_HUGE_VMAP: 1.42x faster (5.00 us -> 3.53 us) * vmap(100MB) with order-8 pages: 8.3x faster (1235 us -> 149 us) Many thanks to Xueyuan Chen for his testing efforts on RK3588 boards. Large vmap() mappings were also tested by Leo Yan with ARM trace buffer units, including TRBE and SPE. These units use the CPU page tables for address translation when writing trace data to DRAM, so using larger vmap() mapping granules can reduce TLB pressure on the trace writer. The TRBE test used a 1G CoreSight ETM AUX buffer. Across five runs on an isolated CPU, the average results were: * dtlb_walk: 68.4 -> 59.4 (-13.16%) * l1d_tlb_refill: 155.8 -> 119.6 (-23.23%) * l2d_tlb_refill: 161435.8 -> 495.0 (-99.69%) The SPE test used a 512M ARM SPE AUX buffer. Across five runs on an isolated CPU, the average results were: * dtlb_walk: 1710.4 -> 1315.6 (-23.08%) * l1d_tlb_refill: 16000.0 -> 15950.2 (-0.31%) * l2d_tlb_refill: 4796.0 -> 2931.2 (-38.88%) These results show that enabling larger vmap() mappings can materially reduce page table walks and TLB refills for large trace buffers. Many thanks to Leo Yan for his testing efforts on ARM trace buffers. Changes since v7: - v7's patch 1 (which extended the hugetlb helpers in arch/arm64/mm/hugetlbpage.c) is replaced by pte_set_huge()/ pte_clear_huge(), split across patches 1-3 so that the arm64, powerpc/8xx and generic changes can be reviewed and acked independently: patch 1 is arm64 only, patch 2 is powerpc/8xx only, patch 3 is the generic fallbacks plus the mm/vmalloc.c conversion. hugetlbpage.c no longer gains the CONT_PTE batching hooks, and patch 4 additionally removes its now-dead init_mm special case. mm/vmalloc.c no longer includes . - The generic pte_set_huge()/pte_clear_huge() fallbacks WARN_ON_ONCE() instead of silently doing nothing. They only exist to keep the build working on architectures without PTE-level block mappings, where they are unreachable (patch 3). - Patch 5 (v7 patch 2): use round_down(size, CONT_PTE_SIZE) instead of rounddown_pow_of_two(size), since pte_set_huge() takes the size directly without an ilog2() roundtrip. This lets a single call span several CONT_PTE blocks: a 768K (512K + 256K) ioremap is now 1.25x faster than v7. - Patch 6 (v7 patch 3): vmap_set_ptes() now calls pte_set_huge() instead of the hugetlb path. - Patch 9 (v7 patch 6): get_vmap_batch_order() takes pages+i instead of a separate idx argument, and applies pfn alignment limit to scan length rather than to the resulting order. Renamed idx/map_addr to batch_idx/batch_start. Dropped Dev's Reviewed-by. Changes since v6: - Add a clarifying comment about the reuse of hugetlb helpers by non-hugetlbfs(vmalloc) mm code (patch 1) - Expand the arm64/vmalloc commit message and comment to clarify that multi-CONT_PTE_SIZE values are vmalloc mapping spans, not HugeTLB hstate sizes (patch 2) - Move the local steps variable change in vmap_pte_range() into the vmap_set_ptes() extraction patch (patch 3) - Propagate vmap_pages_pte_range() errors through the upper vmap_pages_*() levels instead of returning -ENOMEM for all failures (patch 4) - Add a preparatory vm_shift() helper patch before the batching patch (patch 5) - Guard the PFN alignment clamp in get_vmap_batch_order() against PFN 0 before calling __ffs() (patch 6) - Fix kmsan_vmap_pages_range_noflush() indentation in the batching path (patch 6) Changes since v5: - No code changes. - Pick up Reviewed-by and Tested-by tags from Dev, Leo and Uladzislau. Many thanks! - Add TRBE/SPE large vmap() test results from Leo Yan to the cover letter. Changes since v4: - Move pgsize update before contig_ptes check (patch 1) - Use rounddown_pow_of_two instead of __fls in arch_vmap_pte_range_map_size (patch 2) - Reword comment to avoid mentioning cont_pte and remove if in vmap_set_ptes (patch 3) - Rename vmap_batched() to vmap_pages_range_batched() (patch 5) - Use batch_end as the batching cursor to avoid an unused start variable (patch 5) - Check arch_vmap_pmd_supported before PMD mapping (patch 6) Changes since v3: - Squash vmap_pte_range() loop variable fix into patch 4 (patch 3, 4) - Use shift >= PMD_SHIFT and fix *nr increment in vmap_pages_pmd_range() (patch 4) - Pass page_shift directly without capping at PMD_SHIFT (patch 4, 5) - Add vm_shift() helper and pass pgprot_t to get_vmap_batch_order() (patch 5) - Use min(order, __ffs(pfn)) for graceful pfn alignment degradation, replacing IS_ALIGNED check (patch 5) - Remove irrelevant ioremap_max_page_shift early-exit (patch 5) - Add __get_vm_area_node_aligned_caller() wrapper, rename to vmap_get_aligned_vm_area() (patch 6) Changes since v2: - Use __fls instead of fls in arch_vmap_pte_range_map_size (patch 2) - Add WARN_ON checks in vmap_pages_pmd_range (patch 4) - Fix flush_cache_vmap to use saved start address instead of the already-advanced addr (patch 5) - Rename __vmap_huge() to vmap_batched() (patch 5) - Add caller parameter and unroll while(1) loop (patch 5) - Squash patch 7 into patch 5 (stop scanning for compound pages after encountering small pages) Changes since v1: - Fix condition order and use PMD_SIZE instead of CONT_PMD_SIZE in patch 1 (Dev Jain) - Squash patch 3+4 and patch 5+7 (Dev Jain) - Replace "zigzag" with "page table rewalk" in commit messages (Dev Jain) - Rename vmap_small_pages_range_noflush() to vmap_pages_range_noflush_walk() (Dev Jain) - Extract vmap_set_ptes() as a new patch to consolidate PTE mapping logic between vmap_pte_range() and vmap_pages_pte_range(), handling both CONT_PTE and regular mappings (Mike Rapoport) - Support non-compound pages in get_vmap_batch_order() by falling back to physical contiguity scanning with pfn alignment check (Dev Jain, Uladzislau Rezki) - In get_vmap_batch_order(), filter out orders that the architecture cannot batch by checking arch_vmap_pte_supported_shift() directly. This avoids overhead for orders 1-3 on ARM64 CONT_PTE with 4K pages. (patch 5) Barry Song (Xiaomi) (4): arm64/vmalloc: allow arch_vmap_pte_range_map_size() to batch multiple CONT_PTE mm/vmalloc: extend page table walk to support larger page_shift sizes and eliminate page table rewalk mm/vmalloc: map contiguous pages in batches for vmap() if possible mm/vmalloc: align vm_area so vmap() can batch mappings Wen Jiang (6): arm64/mm: add pte_set_huge() and pte_clear_huge() powerpc/8xx: add pte_set_huge() mm/vmalloc: use pte_set_huge()/pte_clear_huge() for PTE-level block mappings arm64/hugetlb: drop the init_mm special case in clear_flush() mm/vmalloc: extract vmap_set_ptes() to consolidate PTE mapping logic mm/vmalloc: extract vm_shift() to consolidate mapping shift selection arch/arm64/include/asm/pgtable.h | 6 + arch/arm64/include/asm/vmalloc.h | 8 +- arch/arm64/mm/hugetlbpage.c | 5 +- arch/arm64/mm/mmu.c | 19 ++ arch/powerpc/include/asm/nohash/32/pte-8xx.h | 4 + arch/powerpc/mm/nohash/8xx.c | 29 ++ include/linux/pgtable.h | 29 ++ mm/vmalloc.c | 268 ++++++++++++++----- 8 files changed, 300 insertions(+), 68 deletions(-) -- 2.34.1