From: "Zi Yan" <ziy@nvidia.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
"Andrew Morton" <akpm@linux-foundation.org>,
"Muchun Song" <muchun.song@linux.dev>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Mike Rapoport" <rppt@kernel.org>,
"Suren Baghdasaryan" <surenb@google.com>,
"Michal Hocko" <mhocko@suse.com>,
"Baolin Wang" <baolin.wang@linux.alibaba.com>,
"Nico Pache" <nico.pache@linux.dev>,
"Ryan Roberts" <ryan.roberts@arm.com>,
"Dev Jain" <dev.jain@arm.com>, "Barry Song" <baohua@kernel.org>,
"Lance Yang" <lance.yang@linux.dev>,
"Usama Arif" <usama.arif@linux.dev>,
"Gregory Price" <gourry@gourry.net>,
"Ying Huang" <ying.huang@linux.alibaba.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Johannes Weiner" <hannes@cmpxchg.org>,
"Qi Zheng" <qi.zheng@linux.dev>,
"Shakeel Butt" <shakeel.butt@linux.dev>,
"Kairui Song" <kasong@tencent.com>
Cc: <linux-mm@kvack.org>, <linux-kernel@vger.kernel.org>,
"Gao Xiang" <xiang@kernel.org>, "Chao Yu" <chao@kernel.org>,
"Jan Kara" <jack@suse.cz>, "Yue Hu" <zbestahu@gmail.com>,
"Jeffle Xu" <jefflexu@linux.alibaba.com>,
"Sandeep Dhavale" <dhavale@google.com>,
"Hongbo Li" <hongbohbli@tencent.com>,
"Chunhai Guo" <guochunhai@vivo.com>,
<linux-erofs@lists.ozlabs.org>, <linux-fsdevel@vger.kernel.org>
Subject: Re: [PATCH v3 07/14] erofs: mm/pagemap: add readahead_folio_last() to avoid folio->private
Date: Tue, 08 Sep 2026 13:05:45 -0400 [thread overview]
Message-ID: <DLA3KHC5R6G3.22TAPOVQ1589Q@nvidia.com> (raw)
In-Reply-To: <eefbff2b-d619-496e-a977-52062da90f05@kernel.org>
On Tue Sep 8, 2026 at 12:04 PM EDT, David Hildenbrand (Arm) wrote:
> On 9/8/26 04:56, Zi Yan wrote:
>> erofs needs to traverse readahead folios in reverse order to achieve
>> maximum performance by
>> 1. reading all folios from readahead_folio();
>> 2. storing the prior folio pointer in folio->private;
>> 3. traverse from the last folio to the first one.
>>
>> Add readahead_folio_last() to achieve the same function without using
>> folio->private. __readahead_advance() helper shares readahead_control
>> adjustment code among __readahead_folio(), readahead_folio_last(), and
>> __readahead_batch() by checking new private member, _forward, of
>> readahead_control.
>>
>> It prepares for a future commit that replaces PG_private checks with
>> !folio->private checks. After switching the checks, erofs's use of
>> folio->private without bumping folio refcount can cause unexpected
>> outcomes, e.g., in filemap_release_folio(), try_to_free_buffers() becomes
>> reachable.
>
> Ah, I was just about to ask. So it's really about folios never using
> folio->private manually (without the attach/detach).
>
>>
>> No functional change intended.
>>
>> Assisted-by: Claude:claude-opus-4-8
>> Assisted-by: Codex:gpt-5
>> Signed-off-by: Zi Yan <ziy@nvidia.com>
>> To: Gao Xiang <xiang@kernel.org>
>> To: Chao Yu <chao@kernel.org>
>> To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
>> To: Jan Kara <jack@suse.cz>
>> Cc: Yue Hu <zbestahu@gmail.com>
>> Cc: Jeffle Xu <jefflexu@linux.alibaba.com>
>> Cc: Sandeep Dhavale <dhavale@google.com>
>> Cc: Hongbo Li <hongbohbli@tencent.com>
>> Cc: Chunhai Guo <guochunhai@vivo.com>
>> Cc: linux-erofs@lists.ozlabs.org
>> Cc: linux-kernel@vger.kernel.org
>> Cc: linux-fsdevel@vger.kernel.org
>> Cc: linux-mm@kvack.org
>> ---
>> fs/erofs/zdata.c | 13 +++---------
>> include/linux/pagemap.h | 56 ++++++++++++++++++++++++++++++++++++++++++-------
>> 2 files changed, 51 insertions(+), 18 deletions(-)
>>
>> diff --git a/fs/erofs/zdata.c b/fs/erofs/zdata.c
>> index e1e25ca0d1904..78fd7d980e957 100644
>> --- a/fs/erofs/zdata.c
>> +++ b/fs/erofs/zdata.c
>> @@ -1898,21 +1898,14 @@ static void z_erofs_readahead(struct readahead_control *rac)
>> struct inode *realinode = erofs_real_inode(sharedinode, &need_iput);
>> Z_EROFS_DEFINE_FRONTEND(f, realinode, sharedinode, readahead_pos(rac));
>> unsigned int nrpages = readahead_count(rac);
>> - struct folio *head = NULL, *folio;
>> + struct folio *folio;
>> int err;
>>
>> trace_erofs_readahead(realinode, readahead_index(rac), nrpages, false);
>> z_erofs_pcluster_readmore(&f, rac, true);
>> - while ((folio = readahead_folio(rac))) {
>> - folio->private = head;
>> - head = folio;
>> - }
>> -
>> - /* traverse in reverse order for best metadata I/O performance */
>> - while (head) {
>> - folio = head;
>> - head = folio_get_private(folio);
>>
>> + /* traverse from last to first for best metadata I/O performance */
>> + while ((folio = readahead_folio_last(rac))) {
>
> Intuitively, this should be called readahead_folio_reverse /
> readahead_folio_reversed, thinking of list_for_each_entry_reverse()?
>
> list_for_each_entry_reverse - iterate backwards over list of given type.
>
> or maybe readahead_folio_backwards (which matches the forward below)
>
> But I'm not a readahead expert :)
Jan suggested the name[1]. It can be readahead_folio_reverse() if you
prefer it, like Jan said.
[1] https://lore.kernel.org/all/332rknj4vo3cfhvfhhlf6pvg37s3lbrnzbbnv4swa6gctsiu6a@ndotgokvnglc/
>
>> err = z_erofs_scan_folio(&f, folio, true);
>> if (err && err != -EINTR)
>> erofs_err(realinode->i_sb, "readahead error at folio %lu @ nid %llu",
>> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
>> index 939f3a5e973f6..2257df004305e 100644
>> --- a/include/linux/pagemap.h
>> +++ b/include/linux/pagemap.h
>> @@ -1415,6 +1415,7 @@ struct readahead_control {
>> bool dropbehind;
>> bool _workingset;
>> unsigned long _pflags;
>> + bool _forward;
>> };
>>
>> #define DEFINE_READAHEAD(ractl, f, r, m, i) \
>> @@ -1479,18 +1480,25 @@ void page_cache_async_readahead(struct address_space *mapping,
>> page_cache_async_ra(&ractl, folio, req_count);
>> }
>>
>> +static inline void __readahead_advance(struct readahead_control *rac)
>> +{
>> + if (rac->_forward)
>> + rac->_index += rac->_batch_count;
>
> No expert, but shouldn't we decrement the _index somewhere in the other case? Or
> where is that done? A comment might help :)
A readahead folio comes from [_index, _index + _nr_pages), so for
last/reverse/backwards case, the code only needs to decrease _nr_pages.
Will add a comment about this.
--
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-09-08 17:07 UTC|newest]
Thread overview: 60+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-08 2:56 [PATCH v3 00/14] Remove PG_private by using page/folio->private checks instead Zi Yan
2026-09-08 2:56 ` [PATCH v3 01/14] mm/zsmalloc: replace PG_private with pointer comparison Zi Yan
2026-09-08 4:11 ` Sergey Senozhatsky
2026-09-08 15:07 ` David Hildenbrand (Arm)
2026-09-08 15:21 ` Zi Yan
2026-09-09 14:07 ` David Hildenbrand (Arm)
2026-09-08 2:56 ` [PATCH v3 02/14] perf/ring_buffer: stop using PG_private as AUX page high-order marker Zi Yan
2026-09-08 15:09 ` David Hildenbrand (Arm)
2026-09-08 2:56 ` [PATCH v3 03/14] xen/grant-table: stop setting PG_private on pages for grant mapping Zi Yan
2026-09-08 15:16 ` David Hildenbrand (Arm)
2026-09-08 15:44 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 04/14] fscrypt: stop setting PG_private on bounce page Zi Yan
2026-09-08 15:16 ` David Hildenbrand (Arm)
2026-09-08 2:56 ` [PATCH v3 05/14] mm/hugetlb: use direct assignment instead of folio_change_private() Zi Yan
2026-09-08 15:26 ` David Hildenbrand (Arm)
2026-09-08 15:28 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 06/14] f2fs: stop using PG_private Zi Yan
2026-09-08 15:47 ` David Hildenbrand (Arm)
2026-09-08 18:20 ` Tal Zussman
2026-09-09 12:48 ` David Hildenbrand (Arm)
2026-09-09 17:47 ` Tal Zussman
2026-09-10 7:35 ` David Hildenbrand (Arm)
2026-09-10 12:21 ` Zi Yan
2026-09-10 21:11 ` Zi Yan
2026-09-10 2:41 ` Chao Yu
2026-09-10 7:34 ` David Hildenbrand (Arm)
2026-09-10 8:43 ` Chao Yu
2026-09-10 9:15 ` David Hildenbrand (Arm)
2026-09-10 14:15 ` Chao Yu
2026-09-08 2:56 ` [PATCH v3 07/14] erofs: mm/pagemap: add readahead_folio_last() to avoid folio->private Zi Yan
2026-09-08 16:04 ` David Hildenbrand (Arm)
2026-09-08 17:05 ` Zi Yan [this message]
2026-09-09 14:10 ` David Hildenbrand (Arm)
2026-09-08 2:56 ` [PATCH v3 08/14] erofs: use folio_attach/detach_private() instead of direct assignment Zi Yan
2026-09-08 16:13 ` David Hildenbrand (Arm)
2026-09-08 17:19 ` Zi Yan
2026-09-09 13:27 ` David Hildenbrand (Arm)
2026-09-10 2:06 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 09/14] mm/page-flags: check page/folio->private instead of PG_private Zi Yan
2026-09-08 16:56 ` David Hildenbrand (Arm)
2026-09-10 2:08 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 10/14] mm/page-flags: introduce folio_test_fs_private() Zi Yan
2026-09-08 16:59 ` David Hildenbrand (Arm)
2026-09-08 17:22 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 11/14] treewide: remove folio_set/clear_private() Zi Yan
2026-09-08 17:01 ` David Hildenbrand (Arm)
2026-09-08 17:22 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 12/14] treewide: replace PagePrivate() with page_private() Zi Yan
2026-09-09 14:18 ` David Hildenbrand (Arm)
2026-09-09 14:23 ` David Hildenbrand (Arm)
2026-09-10 2:09 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 13/14] treewide: adjust comments on PagePrivate and PG_private Zi Yan
2026-09-09 14:20 ` David Hildenbrand (Arm)
2026-09-09 14:27 ` David Hildenbrand (Arm)
2026-09-10 2:12 ` Zi Yan
2026-09-08 2:56 ` [PATCH v3 14/14] mm/page-flags: remove PG_private Zi Yan
2026-09-09 14:31 ` David Hildenbrand (Arm)
2026-09-09 14:48 ` Zi Yan
2026-09-09 14:50 ` David Hildenbrand (Arm)
2026-09-09 14:51 ` Zi Yan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLA3KHC5R6G3.22TAPOVQ1589Q@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=chao@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=dhavale@google.com \
--cc=gourry@gourry.net \
--cc=guochunhai@vivo.com \
--cc=hannes@cmpxchg.org \
--cc=hongbohbli@tencent.com \
--cc=jack@suse.cz \
--cc=jefflexu@linux.alibaba.com \
--cc=kasong@tencent.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-erofs@lists.ozlabs.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nico.pache@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=xiang@kernel.org \
--cc=ying.huang@linux.alibaba.com \
--cc=zbestahu@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®