From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from out30-110.freemail.mail.aliyun.com (out30-110.freemail.mail.aliyun.com [115.124.30.110]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F411335541 for ; Tue, 16 Dec 2025 05:49:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=115.124.30.110 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765864157; cv=none; b=E1R3Ch34sAe5wJnG3wtPUoZwtiNuD7f8tnDGsUOJWepfbadXzL3iJmbfD55oMgcp3b3avIelbap9laEkokf4hWQestUDZ3aSM3qoMd84dYFEQsVv8IWk390X7b6M2jJ9Zil45ymgLMvQ+2GDz5U6btNF9EtfVEdz6BgU0u5ILOA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765864157; c=relaxed/simple; bh=XaOvbCV7XvqRS827l2TFd8jLtD4kcHy7Ux1ZtNqFgGA=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=abfquvKa7Jcuwv3i28+//JWGCExujAUkWWoeYONKTOKteTJdWTnOIYExie0RJjWJhVEb0IQkpBksBxGM2sp4SWtAOAdfK96y1Es8gz4OW5DNmgxYmjraEU7VbiWaabx5dXqeMFO8ab1axolQrfXbfu6Cp4Kwcrp2YBVcgnKVGE0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com; spf=pass smtp.mailfrom=linux.alibaba.com; dkim=pass (1024-bit key) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.b=RnvLh1N0; arc=none smtp.client-ip=115.124.30.110 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.alibaba.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.b="RnvLh1N0" DKIM-Signature:v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.alibaba.com; s=default; t=1765864135; h=Message-ID:Date:MIME-Version:Subject:To:From:Content-Type; bh=tRhxoTZ+3IIdYwTnJoWC2Y2iVLI35ww6bqSlIUl0XxQ=; b=RnvLh1N05nqsdW8RyN4dLQK4QHp3GfPYXhmNfatl8WXhSSWliqbl6ha7l0dax1NXRf40cVzB17gqe8dxDvpTI/siaQETZLh32mGyrF3vXsvkwSSo2lyMXSEmsRrLftNyv52HoT3ZadzTeVCllsn9jGa3Eex0lhCIynU+iU61GDI= Received: from 30.74.144.116(mailfrom:baolin.wang@linux.alibaba.com fp:SMTPD_---0Wuxvaan_1765864133 cluster:ay36) by smtp.aliyun-inc.com; Tue, 16 Dec 2025 13:48:53 +0800 Message-ID: Date: Tue, 16 Dec 2025 13:48:52 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v2 3/3] mm: rmap: support batched unmapping for file large folios To: Lorenzo Stoakes Cc: akpm@linux-foundation.org, david@kernel.org, catalin.marinas@arm.com, will@kernel.org, ryan.roberts@arm.com, Liam.Howlett@oracle.com, vbabka@suse.cz, rppt@kernel.org, surenb@google.com, mhocko@suse.com, riel@surriel.com, harry.yoo@oracle.com, jannh@google.com, willy@infradead.org, baohua@kernel.org, linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org References: <70d1dedc-b4fc-4eb6-baf6-9e54b6a62249@lucifer.local> From: Baolin Wang In-Reply-To: <70d1dedc-b4fc-4eb6-baf6-9e54b6a62249@lucifer.local> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 2025/12/15 20:38, Lorenzo Stoakes wrote: > On Thu, Dec 11, 2025 at 04:16:56PM +0800, Baolin Wang wrote: >> Similar to folio_referenced_one(), we can apply batched unmapping for file >> large folios to optimize the performance of file folios reclamation. >> >> Performance testing: >> Allocate 10G clean file-backed folios by mmap() in a memory cgroup, and try to >> reclaim 8G file-backed folios via the memory.reclaim interface. I can observe >> 75% performance improvement on my Arm64 32-core server. > > Again, you must test on non-arm64 architectures and report the numbers for this > also. Yes, I've tested on the x86 machine, and will add the data in the commit message. >> W/o patch: >> real 0m1.018s >> user 0m0.000s >> sys 0m1.018s >> >> W/ patch: >> real 0m0.249s >> user 0m0.000s >> sys 0m0.249s >> >> Signed-off-by: Baolin Wang >> --- >> mm/rmap.c | 7 ++++--- >> 1 file changed, 4 insertions(+), 3 deletions(-) >> >> diff --git a/mm/rmap.c b/mm/rmap.c >> index ec232165c47d..4c9d5777c8da 100644 >> --- a/mm/rmap.c >> +++ b/mm/rmap.c >> @@ -1855,9 +1855,10 @@ static inline unsigned int folio_unmap_pte_batch(struct folio *folio, >> end_addr = pmd_addr_end(addr, vma->vm_end); >> max_nr = (end_addr - addr) >> PAGE_SHIFT; >> >> - /* We only support lazyfree batching for now ... */ >> - if (!folio_test_anon(folio) || folio_test_swapbacked(folio)) >> + /* We only support lazyfree or file folios batching for now ... */ >> + if (folio_test_anon(folio) && folio_test_swapbacked(folio)) > > Why is it now ok to support file-backed batched unmapping when it wasn't in > Barry's series (see [0])? You don't seem to be justifying this? Barry's series[0] is merely aimed at optimizing lazyfree anonymous large folios and does not continue to optimize anonymous large folios or file-backed large folios at that point. Subsequently, Barry sent out a new patch (see [1]) to optimize anonymous large folios. As for file-backed large folios, the batched unmapping support is relatively simple, since we only need to clear the PTE entries for file-backed large folios. > [0]:https://lore.kernel.org/all/20250214093015.51024-4-21cnbao@gmail.com/T/#u [1] https://lore.kernel.org/all/20250513084620.58231-1-21cnbao@gmail.com/ >> return 1; >> + >> if (pte_unused(pte)) >> return 1; >> >> @@ -2223,7 +2224,7 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma, >> * >> * See Documentation/mm/mmu_notifier.rst >> */ >> - dec_mm_counter(mm, mm_counter_file(folio)); >> + add_mm_counter(mm, mm_counter_file(folio), -nr_pages); > > Was this just a bug before? Nope. Before this patch, we never supported batched unmapping for file-backed large folios, so the 'nr_pages' was always 1. After this patch, we should use the number of pages in this file-backed large folio.