From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 40CCA1E9B37 for ; Thu, 25 Jun 2026 06:37:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782369449; cv=none; b=CtHvpjJtnjv5ZZRvzTy00xH1YAR1r6SKBEDAAPmdp91PtirAJsUZS8hPleyHWRlwNDdQ1p4nEkX+vwrNHicuAMcGk7fBOuSJ10wJueWykKn4Er/oE+RjB6Vy7HCdVnaYz3PJWpbN+fjCiaQljDPufJ1XLqKvASynvlzaBpLLM1A= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782369449; c=relaxed/simple; bh=EEs9c0oR0ySt7qs8lyDve3TQdfPsDi4b1RDmHK97N2s=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=iriPojVpMVDZ6owx69IvAmH1Mtl+NjslR3W5KQ6H0mhSOjdEbAqe8Mv+diM/UevBRNL9Wnh7GthNAmCV5euaFrbn4zA2yJSmqIik2rQi795MCBUiTOkle+TItOZUu+B3/aAtT4YXa+awiRoiQ65fHowW8W9jP+DwzDiXu2MVZxk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=cvB+/xs6; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="cvB+/xs6" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 27C862BCE; Wed, 24 Jun 2026 23:37:14 -0700 (PDT) Received: from [10.164.19.14] (unknown [10.164.19.14]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 2B7363F836; Wed, 24 Jun 2026 23:37:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1782369438; bh=EEs9c0oR0ySt7qs8lyDve3TQdfPsDi4b1RDmHK97N2s=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=cvB+/xs6BeUfI5Q+MqjEYE+/jyUJaCsPo3JqLGcjDJIEOO5eq8V6e62FPOC4la5MU ASiXMlUF+QsKsIpEGNo1qlVHXUrPqK6NtC6W1d/wl03Setc1gjYV9m+gEx1WS8pqKF E86vo+p2VLEx/Yb5YOwHQG8YIsV6M43EzueVm21Q= Message-ID: Date: Thu, 25 Jun 2026 12:07:11 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4 0/6] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory To: Wen Jiang , linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, catalin.marinas@arm.com, will@kernel.org, akpm@linux-foundation.org, urezki@gmail.com Cc: baohua@kernel.org, Xueyuan.chen21@gmail.com, rppt@kernel.org, david@kernel.org, ryan.roberts@arm.com, anshuman.khandual@arm.com, ajd@linux.ibm.com, linux-kernel@vger.kernel.org, jiangwen6@xiaomi.com, shanghaoqiang@xiaomi.com, Ard Biesheuvel References: <20260618084726.1070022-1-jiangwen6@xiaomi.com> Content-Language: en-US From: Dev Jain In-Reply-To: <20260618084726.1070022-1-jiangwen6@xiaomi.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 18/06/26 2:17 pm, Wen Jiang wrote: > This patchset accelerates ioremap, vmalloc, and vmap when the memory > is physically fully or partially contiguous. Two techniques are used: > > 1. Avoid page table rewalk when setting PTEs/PMDs for multiple memory > segments > 2. Use batched mappings wherever possible in both vmalloc and ARM64 > layers > > Besides accelerating the mapping path, this also enables large > mappings (PMD and cont-PTE) for vmap, which are currently not > supported. > > Patches 1-2 extend ARM64 vmalloc CONT-PTE mapping to support multiple > CONT-PTE regions instead of just one. > > Patch 3 extracts a common helper vmap_set_ptes() that consolidates PTE > mapping logic between the ioremap and vmalloc/vmap paths, handling both > CONT_PTE and regular PTE mappings. This prepares for the next patch. > > Patch 4 extends the page table walk path to support page shifts other > than PAGE_SHIFT and eliminates the page table rewalk for huge vmalloc > mappings. The function is renamed from vmap_small_pages_range_noflush() > to vmap_pages_range_noflush_walk(). > > Patches 5-6 add huge vmap support for contiguous pages, including > support for non-compound pages with pfn alignment verification. > > On the RK3588 8-core ARM64 SoC, with tasks pinned to a little core and > the performance CPUfreq policy enabled, benchmark results: > > * ioremap(1 MB): 1.35x faster (3407 ns -> 2526 ns) > * vmalloc(1 MB) mapping time (excluding allocation) with > VM_ALLOW_HUGE_VMAP: 1.42x faster (5.00 us -> 3.53us) > * vmap(100MB) with order-8 pages: 8.3x faster (1235 us -> 149 us) > > Many thanks to Xueyuan Chen for his testing efforts on RK3588 boards. > I am still a little nervous about doing vmap-huge by default. We can play set_memory_* games on a vmap huge mapping partially, thus forcing a pgtable split, and not all arches can handle a kernel pgtable split. For arm64, we can handle that with BBML2_NOABORT, but interestingly, in change_memory_common, arch/arm64/mm/pageattr.c: area = find_vm_area((void *)addr); if (!area || ((unsigned long)kasan_reset_tag((void *)end) > (unsigned long)kasan_reset_tag(area->addr) + area->size) || ((area->flags & (VM_ALLOC | VM_ALLOW_HUGE_VMAP)) != VM_ALLOC)) return -EINVAL; Even before my change fcf8dda8cc48, we were bailing out on !(area->flags & VM_ALLOC)) So on arm64 we haven't been supporting set_memory_* for vmap memory at all, because it has VM_MAP set and not VM_ALLOC. Although we have a contradictory comment above this code so not sure if this was intentional: "Let's restrict ourselves to mappings created by vmalloc (or vmap)." So either there is no user in the kernel doing vmap + set_memory_* (looks like it by doing an LLM scan), or it is not fatal for set_memory_* to fail. But even if no one does it now, technically the API allows it. >