From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f52.google.com (mail-pj1-f52.google.com [209.85.216.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D421A3314D0 for ; Fri, 22 May 2026 05:31:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.52 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779427915; cv=none; b=CEoVwIn2LAdyWZ8ifLTiFP1j5Vj53/bqneD37iFxl34NInNqXz0/Mz3KB5URi6v37j/WgC3jgvOupndz2U91itI19AmTjl7MCe60YCb73kvP+klPdQxDLvYqpqkNLgFZAJL+1IPZkkHjBGZFX0DIZjVx3JlAkreDn+KWxQ3G4To= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779427915; c=relaxed/simple; bh=zgsOI8tJVWOK8nhM+ThPhuE7NIVk17Dnl3WdWgFMN2s=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version:Content-Type; b=JmAsHMggRaYLPf2Cq4IhIAJnytmNJoT10X+WNtiQd0rJ4u/+BX8ySV1NK6j6GP6FBTKLVq/BJTIKYtsIrXDCjOtiD0GiCbGIA6lH5j1Yd/PPmRG360SLG5s4aLWbKvLasJa88u+MmNsFCfa6pwjw7oGo0AbtLFIQNy1yXnTWPtU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BVpCCPMh; arc=none smtp.client-ip=209.85.216.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BVpCCPMh" Received: by mail-pj1-f52.google.com with SMTP id 98e67ed59e1d1-36a35e4eefeso1380143a91.1 for ; Thu, 21 May 2026 22:31:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1779427913; x=1780032713; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=WSsZqaAWHj+v1vpsCxKzKu7IVzLHCQIKTgP9zE3Hky4=; b=BVpCCPMh5068dH4viOioeAnEHgg6Jf6NXkfJRlk8VYxcR1VZ23b+TnBvd3eLi25yXb tuZHUda/v+ntBEloPbvbnlgREFWWg7fT3u/Lo1xn62tjMjI6xd4gUW4UdHKd7IUiOO/T ClbySei8/8pwrsDr4C+u4lti279NnTI4x7RWpRHpPfbJ++FG7lEQH23ypBPewTq+F6Jq 519EtAFRB5XkrOV86cNw/iTv3LMI/YZBy6dQcdzLEJal+zSY4vumcx6Izxk3T1oQyZ1V 6ar6/PrSqZwwtOOCQWgoX2l1idHl+v1M4WrpgkEhvmEaeOz3s5WbBTcBYCpn7xHXMLUr /kgA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1779427913; x=1780032713; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=WSsZqaAWHj+v1vpsCxKzKu7IVzLHCQIKTgP9zE3Hky4=; b=RH3LVHI707l5uACO9f4B8mmBNNHubCKeGqPNpkMaAs7HFz1P5BL9uTPlIQCHiDLS0l Z8e3SJ2MJGAEUKnjFIssbANvcR4s5xn7Mk0+4lV3O/14VdZ34taKpISatoaeoMl5ry1z 7rWZ0q2CAii4yzpeUem2kdI6LQfr7IEtcS1gPT2jzl6lhibsg7YdhY8OX0TssRuCAzXO Y2h1e0TSg26/VpVSmn1NgkYtLM6nuFDze1/YQW4XV/wASUaD//N1Ra+wh3FamBY/Q3/+ F1mVKomws8CwxV39WihJj7kbLYUh1wwoXAquLYBNr24mJuZqLntCQ6KeiJPynKPbKbn+ 8I+Q== X-Forwarded-Encrypted: i=1; AFNElJ8qvxMduXuiNrDvqpTBa1hEl2TcY/3/YqjnJheli44OXLGpeRGTpiqJxfuSM+LuvoQB8PZXpVIUQjzceDw=@vger.kernel.org X-Gm-Message-State: AOJu0YwjZxU5lmu7fhCg6iJjtTkiyYE5VtxMGfLCeigvoL6aLJnTQnq7 FxFy8/S+RlANfvIC2ce7zN5vWNLh/t7MzXTK18aCLaEsHgidUNp3Z/Dg X-Gm-Gg: Acq92OEj5+1M3/IPgAUJp2ojtZab/f0OZon3HxwxdSuystQmIC9RmwYDG4MUml8oiZ+ hW2ToSuXxI3IaSP0zU/fG8akd4YOkxibh80oXH7QVwGoG3vHa0jPgem5zZKgR8pGWuwMGjK1WJo 1htsWOD1e7Ite3kVUTeifOX5PkC0OGV0Dovz/VbrDvAjTFiQo+r58xvEIFTUbtbg6uIgkCZwhdU EnOlzEenMbarLXsPczKbHaAVOhfzGz+jlwWrnCGNOhwk/iDoURIlJQUBuBdn5QrE1+RDxnQWU04 jWrPZG9BGL0Gg77/HWPnisuwmMtug4pbOaEB2ZMiu7iN4kIzpmjy82diMGoe8EVE0RAMpdKjyI9 V9arelbl4pvAHvZZuJBIbs4B+JKANWNXKFlM43xfmpPLljacV/3cmFdV3P3g6TwhhtFyeebKnb4 pgMkUnBrvqwTvMqXHeeCpQGuzYa7LNC1osndXY+K/uGb0oraiahhY= X-Received: by 2002:a17:903:2b0f:b0:2bd:63dc:b7ad with SMTP id d9443c01a7336-2beb05a227emr21130695ad.2.1779427913092; Thu, 21 May 2026 22:31:53 -0700 (PDT) Received: from mi-OptiPlex-7060.mioffice.cn ([43.224.245.234]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2beb56d68adsm4782665ad.32.2026.05.21.22.31.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 21 May 2026 22:31:52 -0700 (PDT) From: Wen Jiang To: linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, catalin.marinas@arm.com, will@kernel.org, akpm@linux-foundation.org, urezki@gmail.com Cc: baohua@kernel.org, Xueyuan.chen21@gmail.com, dev.jain@arm.com, rppt@kernel.org, david@kernel.org, ryan.roberts@arm.com, anshuman.khandual@arm.com, ajd@linux.ibm.com, linux-kernel@vger.kernel.org, jiangwen6@xiaomi.com Subject: [PATCH v3 0/6] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory Date: Fri, 22 May 2026 13:31:40 +0800 Message-Id: <20260522053146.83209-1-jiangwenxiaomi@gmail.com> X-Mailer: git-send-email 2.34.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit From: jiangwen6 This patchset accelerates ioremap, vmalloc, and vmap when the memory is physically fully or partially contiguous. Two techniques are used: 1. Avoid page table rewalk when setting PTEs/PMDs for multiple memory segments 2. Use batched mappings wherever possible in both vmalloc and ARM64 layers Besides accelerating the mapping path, this also enables large mappings (PMD and cont-PTE) for vmap, which are currently not supported. Patches 1-2 extend ARM64 vmalloc CONT-PTE mapping to support multiple CONT-PTE regions instead of just one. Patch 3 extracts a common helper vmap_set_ptes() that consolidates PTE mapping logic between the ioremap and vmalloc/vmap paths, handling both CONT_PTE and regular PTE mappings. This prepares for the next patch. Patch 4 extends the page table walk path to support page shifts other than PAGE_SHIFT and eliminates the page table rewalk for huge vmalloc mappings. The function is renamed from vmap_small_pages_range_noflush() to vmap_pages_range_noflush_walk(). Patches 5-6 add huge vmap support for contiguous pages, including support for non-compound pages with pfn alignment verification. On the RK3588 8-core ARM64 SoC, with tasks pinned to a little core and the performance CPUfreq policy enabled, benchmark results: * ioremap(1 MB): 1.35x faster (3407 ns -> 2526 ns) * vmalloc(1 MB) mapping time (excluding allocation) with VM_ALLOW_HUGE_VMAP: 1.42x faster (5.00 us -> 3.53us) * vmap(100MB) with order-8 pages: 8.3x faster (1235 us -> 149 us) Many thanks to Xueyuan Chen for his testing efforts on RK3588 boards. Changes since v2: - Use __fls instead of fls in arch_vmap_pte_range_map_size (patch 2) - Add WARN_ON checks in vmap_pages_pmd_range (patch 4) - Fix flush_cache_vmap to use saved start address instead of the already-advanced addr (patch 5) - Rename __vmap_huge() to vmap_batched() (patch 5) - Add caller parameter and unroll while(1) loop (patch 5) - Squash patch 7 into patch 5 (stop scanning for compound pages after encountering small pages) Changes since v1: - Fix condition order and use PMD_SIZE instead of CONT_PMD_SIZE in patch 1 (Dev Jain) - Squash patch 3+4 and patch 5+7 (Dev Jain) - Replace "zigzag" with "page table rewalk" in commit messages (Dev Jain) - Rename vmap_small_pages_range_noflush() to vmap_pages_range_noflush_walk() (Dev Jain) - Extract vmap_set_ptes() as a new patch to consolidate PTE mapping logic between vmap_pte_range() and vmap_pages_pte_range(), handling both CONT_PTE and regular mappings (Mike Rapoport) - Support non-compound pages in get_vmap_batch_order() by falling back to physical contiguity scanning with pfn alignment check (Dev Jain, Uladzislau Rezki) - In get_vmap_batch_order(), filter out orders that the architecture cannot batch by checking arch_vmap_pte_supported_shift() directly. This avoids overhead for orders 1-3 on ARM64 CONT_PTE with 4K pages. (patch 5) Barry Song (Xiaomi) (5): arm64/hugetlb: Extend batching of multiple CONT_PTE in a single PTE setup arm64/vmalloc: Allow arch_vmap_pte_range_map_size to batch multiple CONT_PTE mm/vmalloc: Extend page table walk to support larger page_shift sizes and eliminate page table rewalk mm/vmalloc: map contiguous pages in batches for vmap() if possible mm/vmalloc: align vm_area so vmap() can batch mappings Wen Jiang (1): mm/vmalloc: Extract vmap_set_ptes() to consolidate PTE mapping logic arch/arm64/include/asm/vmalloc.h | 6 +- arch/arm64/mm/hugetlbpage.c | 10 ++ mm/vmalloc.c | 235 ++++++++++++++++++++++++------- 3 files changed, 201 insertions(+), 50 deletions(-) -- 2.34.1