From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6A1FA42DA53 for ; Tue, 28 Jul 2026 12:09:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785240545; cv=none; b=gVnCbrHFiZpoxkC2h/r7G5gmnEmIL92ZffQVcjt/7eAt/Y0a8xiRvi0cwZ679a1vam5eq8EzblCY0/DNGhEhgZTeHzXdwGcg2DtI9/1M0tnOB52/uhTWxJhYlP8ooUXvFasuo59PYcJGR3sx7HXbCKigTdSnQ0DDwWtLDBeVdTs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785240545; c=relaxed/simple; bh=QZEQepEepItHyILI+7973yL9nc7gXy7xeKv1xHjxA0Q=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=BUw5A7x9Q3+2a6rlMLVv9VLsBImahVZeC1fB8BT9UFb42OvQx7kGGmJrCcQbnWZ9A94N/ZkXoL13PhgMoUUA1edeaIbBqJ4QymRFZylDPd+96XImz7g1ecy79d+pYAPg5/OP03+0b5TkLsYgCyu0JbpEjiit2LyfFXmrahcxaU0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=eGXgV9YQ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="eGXgV9YQ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C4BC31F000E9; Tue, 28 Jul 2026 12:09:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785240544; bh=coBQCYguauEUFlE+zRehCazJ7YtpgsT9g1InCofXSTM=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=eGXgV9YQEy+iglnVsOeKzzWKZ51pMguT48dsQ8WNzvcPPyWqhIhse1zFPQy9NWsIF qR86lFypyCG8ra+bC1hW5gF+wH7jsWK+lPfg1wkaIotLjtodukmtZKhhq2BTpfMRqr b6nawRofbM/q0NSkooh8JKp2tFvaigsqZ+dWMTjlIt4LBQNzlxyMoGVcVH690kopwH UgyTOkip3FZucPLm+wg1g8QucTaKogZr3lUhBbGPAfFYJhINOM6jVtCQ2tdXv4+/ye 7aVrMzkj2hvtxI5Ig/pW5pIHk/4IhA3fmIR44zXT3CU1PIARgFC1Qbfu4pB2rE1PtI aWsWTI5hPg5ig== Date: Tue, 28 Jul 2026 13:08:57 +0100 From: Will Deacon To: Wen Jiang Cc: akpm@linux-foundation.org, catalin.marinas@arm.com, linux-mm@kvack.org, urezki@gmail.com, Xueyuan.chen21@gmail.com, ajd@linux.ibm.com, anshuman.khandual@arm.com, david@kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, rppt@kernel.org, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, Wen Jiang , Leo Yan Subject: Re: [PATCH v7 2/7] arm64/vmalloc: Allow arch_vmap_pte_range_map_size to batch multiple CONT_PTE Message-ID: References: <20260715120813.3609949-1-jiangwen6@xiaomi.com> <20260715120813.3609949-3-jiangwen6@xiaomi.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260715120813.3609949-3-jiangwen6@xiaomi.com> On Wed, Jul 15, 2026 at 08:08:08PM +0800, Wen Jiang wrote: > From: "Barry Song (Xiaomi)" > > Allow arch_vmap_pte_range_map_size to batch across multiple CONT_PTE > blocks, reducing both PTE setup and TLB flush iterations. > > For CONT_PTE_SIZE-aligned ranges, return a power-of-two mapping size that > may cover multiple CONT_PTE blocks, capped below PMD_SIZE. These sizes > are vmalloc mapping spans, not HugeTLB hstate sizes. > > Signed-off-by: Barry Song (Xiaomi) > Signed-off-by: Wen Jiang > Tested-by: Xueyuan Chen > Tested-by: Leo Yan > Reviewed-by: Dev Jain > --- > arch/arm64/include/asm/vmalloc.h | 8 +++++++- > 1 file changed, 7 insertions(+), 1 deletion(-) > > diff --git a/arch/arm64/include/asm/vmalloc.h b/arch/arm64/include/asm/vmalloc.h > index 4ec1acd3c1b34..d665f9d687422 100644 > --- a/arch/arm64/include/asm/vmalloc.h > +++ b/arch/arm64/include/asm/vmalloc.h > @@ -23,10 +23,14 @@ static inline unsigned long arch_vmap_pte_range_map_size(unsigned long addr, > unsigned long end, u64 pfn, > unsigned int max_page_shift) > { > + unsigned long size; > + > /* > * If the block is at least CONT_PTE_SIZE in size, and is naturally > * aligned in both virtual and physical space, then we can pte-map the > * block using the PTE_CONT bit for more efficient use of the TLB. > + * The returned mapping size may cover multiple CONT_PTE_SIZE blocks, > + * capped below PMD_SIZE. > */ > if (max_page_shift < CONT_PTE_SHIFT) > return PAGE_SIZE; > @@ -40,7 +44,9 @@ static inline unsigned long arch_vmap_pte_range_map_size(unsigned long addr, > if (!IS_ALIGNED(PFN_PHYS(pfn), CONT_PTE_SIZE)) > return PAGE_SIZE; > > - return CONT_PTE_SIZE; > + size = min3(end - addr, 1UL << max_page_shift, PMD_SIZE >> 1); > + size = rounddown_pow_of_two(size); > + return size; Why does this have to be a power of two? We should be able to work with regions where the start and end are suitably aligned. Is it because the hugetlb code works in terms of shifts? Will