From: Zhang Yi <yizhang089@gmail.com>
To: Ojaswin Mujoo <ojaswin@linux.ibm.com>,
Zhang Yi <yi.zhang@huaweicloud.com>
Cc: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org,
linux-kernel@vger.kernel.org, tytso@mit.edu,
adilger.kernel@dilger.ca, libaokun@linux.alibaba.com,
jack@suse.cz, ritesh.list@gmail.com, djwong@kernel.org,
hch@infradead.org, yi.zhang@huawei.com, chengzhihao1@huawei.com,
yangerkun@huawei.com, wangkefeng.wang@huawei.com,
yukuai@fnnas.com
Subject: Re: [PATCH v6 05/31] ext4: recheck extent status tree before block allocation
Date: Mon, 28 Sep 2026 15:15:21 +0800 [thread overview]
Message-ID: <7993c203-3172-455c-a3ac-a27f195fc1b5@gmail.com> (raw)
In-Reply-To: <arUj-7QI2U12y1i-@li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com>
On 9/24/2026 9:22 PM, Ojaswin Mujoo wrote:
> On Thu, Sep 03, 2026 at 08:35:17PM +0800, Zhang Yi wrote:
>> From: Zhang Yi <yi.zhang@huawei.com>
>>
>> After acquiring i_data_sem in write mode, recheck that the mapping
>> found via the extent status tree or disk query has not changed. A
>> racing truncate may have trimmed the extent between the earlier lookup
>> and the write lock acquisition, since writeback does not hold i_rwsem
>> or the folio locks covering the full extent. This could cause
>> ext4_map_create_blocks() to allocate blocks beyond the truncated range,
>> potentially leading to quota leaks in the upcomming iomap buffered
>> writeback path since the iomap writeback infrastructure caches extents
>> beyond the folio range.
>>
>> Therefore, if we find a valid extent and the sequence number has
>> changed, retry the entire lookup to obtain the correct trimmed mapping.
>>
>> Suggested-by: Jan Kara <jack@suse.cz>
>> Signed-off-by: Zhang Yi <yi.zhang@huawei.com>
>> ---
>> fs/ext4/inode.c | 16 ++++++++++++++++
>> 1 file changed, 16 insertions(+)
>>
>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>> index 84991fe99071..b71b1d2588ae 100644
>> --- a/fs/ext4/inode.c
>> +++ b/fs/ext4/inode.c
>> @@ -734,6 +734,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
>> else
>> ext4_check_map_extents_env(inode);
>>
>> +create_retry:
>> /* Lookup extent status tree firstly */
>> if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, &map->m_seq)) {
>> if (ext4_es_is_written(&es) || ext4_es_is_unwritten(&es)) {
>> @@ -784,6 +785,8 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
>> down_read(&EXT4_I(inode)->i_data_sem);
>> retval = ext4_map_query_blocks(handle, inode, map, flags);
>> up_read((&EXT4_I(inode)->i_data_sem));
>> + if (retval < 0)
>> + return retval;
>
> Hey Zhang, I think there's a slight change in behavior that we return if
> the query fails now. Earlier we would have still tried the create_blocks
> call and, in case of transient errors like ENOMEM, might have suceeded.
> I don't think it's that big of a deal though.
>
> Rest looks fine and I agree with your reply to Sashiko, I think we
> can tweak those things if we ever encounter a livelock.
>
> Feel free to add:
>
> Reviewed-by: Ojaswin Mujoo <ojaswin@linux.ibm.com>
>
> Also, if you don't mind can you add
>
> Link: https://lore.kernel.org/linux-ext4/b0781809-4759-4e12-be17-71555b764f48@gmail.com/
>
> (or the call trace) to the next version as it explains the issue pretty well.
> Looking at these things later takes some time to recall everything and
> such traces help a lot.
>
Hi, Ojaswin!
Thank you for the review and suggestion, this makes sense to me. I will
add this link in the next iteration.
Thanks,
Yi.
> Regards,
> ojaswin
>
>>
>> found:
>> if (retval > 0 && map->m_flags & EXT4_MAP_MAPPED) {
>> @@ -820,6 +823,19 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
>> * with create == 1 flag.
>> */
>> down_write(&EXT4_I(inode)->i_data_sem);
>> +
>> + /*
>> + * Check the validity of the mapping found via the extent status
>> + * tree or the disk query. A racing truncate may have changed the
>> + * extent, since writeback does not hold i_rwsem or the folio locks
>> + * covering the full extent.
>> + */
>> + if (map->m_seq != READ_ONCE(EXT4_I(inode)->i_es_seq)) {
>> + up_write(&EXT4_I(inode)->i_data_sem);
>> + map->m_flags = 0;
>> + map->m_len = orig_mlen;
>> + goto create_retry;
>> + }
>> retval = ext4_map_create_blocks(handle, inode, map, flags);
>> up_write((&EXT4_I(inode)->i_data_sem));
>>
>> --
>> 2.52.0
>>
next prev parent reply other threads:[~2026-09-28 7:15 UTC|newest]
Thread overview: 41+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 12:35 [PATCH v6 00/31] ext4: use iomap for regular file's buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 01/31] ext4: simplify size updating in ext4_setattr() Zhang Yi
2026-09-03 12:35 ` [PATCH v6 02/31] ext4: factor out ext4_truncate_[up|down]() Zhang Yi
2026-09-03 12:35 ` [PATCH v6 03/31] ext4: skip ordered I/O wait when zeroing beyond i_disksize block Zhang Yi
2026-09-03 12:35 ` [PATCH v6 04/31] ext4: set EXT4_MAP_NEW flag for delayed allocated blocks Zhang Yi
2026-09-24 11:42 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 05/31] ext4: recheck extent status tree before block allocation Zhang Yi
2026-09-24 13:22 ` Ojaswin Mujoo
2026-09-28 7:15 ` Zhang Yi [this message]
2026-09-03 12:35 ` [PATCH v6 06/31] ext4: fix orig_mlen initialization in ext4_map_blocks() Zhang Yi
2026-09-24 13:24 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 07/31] ext4: allow ext4_map_blocks() to start its own transaction handle Zhang Yi
2026-09-28 10:19 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 08/31] ext4: avoid unnecessary transaction in ext4_map_blocks() for unwritten extents Zhang Yi
2026-09-28 10:21 ` Ojaswin Mujoo
2026-09-28 11:57 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 09/31] ext4: skip block allocation for holes in the data submission path Zhang Yi
2026-09-28 10:36 ` Ojaswin Mujoo
2026-09-28 12:10 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 10/31] ext4: add iomap address space operations for buffered I/O Zhang Yi
2026-09-03 12:35 ` [PATCH v6 11/31] ext4: implement buffered read path using iomap Zhang Yi
2026-09-03 12:35 ` [PATCH v6 12/31] ext4: pass out extent seq counter when mapping da blocks Zhang Yi
2026-09-03 12:35 ` [PATCH v6 13/31] ext4: do not use data=ordered mode for inodes using buffered iomap path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 14/31] ext4: implement buffered write path using iomap Zhang Yi
2026-09-03 12:35 ` [PATCH v6 15/31] ext4: implement writeback " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 16/31] ext4: implement mmap " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 17/31] ext4: implement partial block zero range " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 18/31] ext4: drain writeback before removing extents on the iomap path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 19/31] ext4: add block mapping tracepoints for iomap buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 20/31] ext4: disable online defrag when inode using " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 21/31] ext4: add EXT4_STATE_DISKSIZE_GROW_PENDING state bit and helpers Zhang Yi
2026-09-03 12:35 ` [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback Zhang Yi
2026-09-03 12:35 ` [PATCH v6 23/31] ext4: advance i_disksize to i_size upon disksize-grow I/O completion Zhang Yi
2026-09-03 12:35 ` [PATCH v6 24/31] ext4: defer i_disksize update while DISKSIZE_GROW_PENDING is set Zhang Yi
2026-09-03 12:35 ` [PATCH v6 25/31] ext4: submit and wait for disksize-grow I/O in fallocate paths Zhang Yi
2026-09-03 12:35 ` [PATCH v6 26/31] ext4: clear DISKSIZE_GROW_PENDING on truncate or error Zhang Yi
2026-09-03 12:35 ` [PATCH v6 27/31] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block Zhang Yi
2026-09-03 12:35 ` [PATCH v6 28/31] ext4: add tracepoints for DISKSIZE_GROW_PENDING set, clear, and wait Zhang Yi
2026-09-03 12:40 ` [PATCH v6 29/31] ext4: add tracepoints for EOF block zeroing and disksize-grow I/O Zhang Yi
2026-09-03 12:40 ` [PATCH v6 30/31] ext4: partially enable iomap for the buffered I/O path of regular files Zhang Yi
2026-09-03 12:40 ` [PATCH v6 31/31] ext4: introduce a mount option for iomap buffered I/O path Zhang Yi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7993c203-3172-455c-a3ac-a27f195fc1b5@gmail.com \
--to=yizhang089@gmail.com \
--cc=adilger.kernel@dilger.ca \
--cc=chengzhihao1@huawei.com \
--cc=djwong@kernel.org \
--cc=hch@infradead.org \
--cc=jack@suse.cz \
--cc=libaokun@linux.alibaba.com \
--cc=linux-ext4@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ojaswin@linux.ibm.com \
--cc=ritesh.list@gmail.com \
--cc=tytso@mit.edu \
--cc=wangkefeng.wang@huawei.com \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yi.zhang@huaweicloud.com \
--cc=yukuai@fnnas.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®