From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from canpmsgout03.his.huawei.com (canpmsgout03.his.huawei.com [113.46.200.218]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 664E2267B05 for ; Wed, 22 Apr 2026 15:03:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=113.46.200.218 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776870207; cv=none; b=mh+yV+7g/1CZ4mIEP/UdJk4N7ZpL7vDnEDWecvbxVsAT6S6YCdriFLMmqCepua/zI+MZYXmjexBPytF/A9m1PpLbfZ70oe+i6vNd3+kx6eQ9KGsEl1rNp9kHsJIPoYls1FzcZoD+8eq0T9jEWjpRlKErh/W2JFyBV+iKcVQ6H8s= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776870207; c=relaxed/simple; bh=z1V5WIg/B9Kq7WozUDwuXVPWgtMGQPPZLdRZ4onggTU=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=cSFZnOstuvCLml7R4zA19LQEucw9SZjk2JZ8JlLEg9/LJJwZg6WvYhEJQMt47Z95ZHBzYGFnqBiOqkqGr77ElRhSiRfEr7O+sFW8z7YPHMyVkZSMSN/jFjO4ivGZKzes64eVEwiYVcd5tvscipFkAP94k7QHbZhXj0ax6KFoMFo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=mCxq2pFV; arc=none smtp.client-ip=113.46.200.218 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="mCxq2pFV" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=xgW+47Jv/ylGNBGjB8siOtex/NYpwe6kH9Qi/e5jtU8=; b=mCxq2pFVcDbRoLzB4u+FHEoJSVYkCatnbfup/6NCAJ9o92CbkIYDPUx0CCQEdX1Ir+btXrNsi sNblvTMnprDTcvTKpw8zJPqylh5h1axy2wb/7eNAcFAlZPt/KBbQn6qwakFc4eHHsbVoNRK1F8O 2/dMEm9p/1+4Cw/inuhPCZc= Received: from mail.maildlp.com (unknown [172.19.162.144]) by canpmsgout03.his.huawei.com (SkyGuard) with ESMTPS id 4g12P93jxCzpStt; Wed, 22 Apr 2026 22:56:45 +0800 (CST) Received: from dggpemf100008.china.huawei.com (unknown [7.185.36.138]) by mail.maildlp.com (Postfix) with ESMTPS id 9380D40538; Wed, 22 Apr 2026 23:03:10 +0800 (CST) Received: from [10.174.177.243] (10.174.177.243) by dggpemf100008.china.huawei.com (7.185.36.138) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Wed, 22 Apr 2026 23:03:09 +0800 Message-ID: <12bdade5-b239-4456-bb5a-f2648c867db8@huawei.com> Date: Wed, 22 Apr 2026 23:03:08 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3] mm: shmem: always support large folios for internal shmem mount To: Baolin Wang , "David Hildenbrand (Arm)" , , CC: , , , , , , Dave Hansen References: <26f954be62348591e720c4e8b7a9099b74dc1d6d.1776331555.git.baolin.wang@linux.alibaba.com> <1b3c0401-6d10-4a28-97c8-8e3858d8dc3d@kernel.org> <015de194-99b9-4f9e-8c89-d35807c6fd08@linux.alibaba.com> <07e26d39-6155-4661-b3df-c2419535ed43@kernel.org> <116df9f9-4db7-40d4-a4a4-30a87c0feffa@linux.alibaba.com> Content-Language: en-US From: Kefeng Wang In-Reply-To: <116df9f9-4db7-40d4-a4a4-30a87c0feffa@linux.alibaba.com> Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 8bit X-ClientProxiedBy: kwepems100002.china.huawei.com (7.221.188.206) To dggpemf100008.china.huawei.com (7.185.36.138) On 4/22/2026 2:28 PM, Baolin Wang wrote: > CC Kefeng, > > On 4/21/26 9:39 PM, David Hildenbrand (Arm) wrote: >> On 4/21/26 08:27, Baolin Wang wrote: >>> >>> >>> On 4/21/26 3:00 AM, David Hildenbrand (Arm) wrote: >>>> On 4/17/26 14:45, Baolin Wang wrote: >>>>> >>>>> >>>>> >>>>> Indeed. Good point. >>>>> >>>>> >>>>> Not really. There could be files created before remount whose mappings >>>>> don't support large folios (with 'huge=never' option), while files >>>>> created after remount will have mappings that support large folios (if >>>>> remounted with 'huge=always' option). >>>>> >>>>> It looks like the previous commit 5a90c155defa was also >>>>> problematic. The >>>>> huge mount option has introduced a lot of tricky issues:( >>>>> >>>>> Now I think Zi's previous suggestion should be able to clean up this >>>>> mess? That is, calling mapping_set_large_folios() unconditionally for >>>>> all shmem mounts, and revisiting Kefeng's first version to fix the >>>>> performance issue. >>>> >>>> Okay, so you'll send a patch to just set mapping_set_large_folios() >>>> unconditionally? >>> >>> I'm still hesitating on this. If we set mapping_set_large_folios() >>> unconditionally, we need to re-fix the performance regression that was >>> addressed by commit 5a90c155defa. >> >> Just so I can follow: where is the test for large folios that we would >> unlock large folios and cause a regression? > > I spent some time investigating the performance regression that was > addressed by commit 5a90c155defa ("tmpfs: don't enable large folios if > not supported"). From my testing, I found that the performance issue no > longer exists on upstream: > > mount tmpfs -t tmpfs -o size=50G /mnt/tmpfs > > Base: > dd if=/dev/zero of=/mnt/tmpfs/test bs=400K count=10485 (3.2 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=800K count=5242 (3.2 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=1600K count=2621 (3.1 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=2200K count=1906 (3.0 GB/s ) > dd if=/dev/zero of=/mnt/tmpfs/test bs=3000K count=1398 (3.0 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=4500K count=932 (3.1 GB/s) > > Base + revert 5a90c155defa: > dd if=/dev/zero of=/mnt/tmpfs/test bs=400K count=10485 (3.3 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=800K count=5242 (3.3 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=1600K count=2621 (3.2 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=2200K count=1906 (3.1 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/testbs=3000K count=1398 (3.0 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=4500K count=932 (3.1 GB/s) > > The data is basically consistent with minor fluctuation noise. > > Later, I continued investigating and found that commit 665575cff098b > ("filemap: move prefaulting out of hot write path") fixed the write > operation performance. > > Base + revert 665575cff098b + revert 5a90c155defa: > dd if=/dev/zero of=/mnt/tmpfs/test bs=400K count=10485 (3.0 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=800K count=5242 (2.9 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=1600K count=2621 (2.6 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=2200K count=1906 (2.6 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=3000K count=1398 (2.5 GB/s) > dd if=/dev/zero of=/mnt/tmpfs/test bs=4500K count=932 (2.5 GB/s) > > We can see that after reverting commit 665575cff098b, there is a > noticeable drop in write performance for tmpfs files. > > So my conclusion is that we can now safely revert commit 5a90c155defa to > set mapping_set_large_folios() for all shmem mounts unconditionally. > > Kefeng, please correct me if I missed anything. Hi Baolin,I found my testcases "bonnie Block/Re Write" ./bonnie -d /tmp -s Size (size is from 100,256,512,1024,2048,4096). But the dd test is similar as well, and as commit 4e527d5841e2 ("iomap: fault in smaller chunks for non-large folio mappings") said, the issue is, "If chunk is 2MB, total 512 pages need to be handled finally. During this period, fault_in_iov_iter_readable() is called to check iov_iter readable validity. Since only 4KB will be handled each time, below address space will be checked over and over again" But after 665575cff098b, fault_in_iov_iter_readable() is moved, so the issue should be fixed. +CC Dave, Since 665575cff098b is works well in generic_perform_write(), I think we could do the same optimization in iomap_write_iter()? but it seems maintainer forget pickup them[1]. [1] https://lore.kernel.org/all/20250129181753.3927F212@davehans-spike.ostc.intel.com/