From: Zhang Yi <yizhang089@gmail.com>
To: Ojaswin Mujoo <ojaswin@linux.ibm.com>,
Zhang Yi <yi.zhang@huaweicloud.com>
Cc: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org,
linux-kernel@vger.kernel.org, tytso@mit.edu,
adilger.kernel@dilger.ca, libaokun@linux.alibaba.com,
jack@suse.cz, ritesh.list@gmail.com, djwong@kernel.org,
hch@infradead.org, yi.zhang@huawei.com, chengzhihao1@huawei.com,
yangerkun@huawei.com, wangkefeng.wang@huawei.com,
yukuai@fnnas.com
Subject: Re: [PATCH v6 09/31] ext4: skip block allocation for holes in the data submission path
Date: Mon, 28 Sep 2026 20:10:06 +0800 [thread overview]
Message-ID: <e3b38169-dc4a-4a42-8a42-b017ecff73d4@gmail.com> (raw)
In-Reply-To: <arpDP43MIAGjaWYx@li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com>
On 9/28/2026 6:36 PM, Ojaswin Mujoo wrote:
> On Thu, Sep 03, 2026 at 08:35:21PM +0800, Zhang Yi wrote:
>> From: Zhang Yi <yi.zhang@huawei.com>
>>
>> When ext4_map_blocks() is called from the data submission path and I/O
>> end extent conversion path (EXT4_GET_BLOCKS_IO_SUBMIT), it should not
>> allocate blocks if the lookup returns a hole.
>>
>> The writeback path can legitimately encounter dirty ranges that map to
>> holes. For example, when a folio straddles i_size and the tail beyond
>> i_size is dirtied via a mmap write. Allocating blocks for such ranges is
>> wrong because there is no data to write back, the dirty bits should
>> simply be discarded without submitting I/O. This prepares for the
>> buffered iomap writeback conversion, mirrors the existing buffer_head
>> writeback path, where mpage_add_bh_to_extent() skip unmapped buffers and
>> ext4_bio_write_folio() clears their dirty bits.
>
> Looks good, checking the iomap vs bh head logic. If we consider
> a 1k bs FS with 4k page size, with hte following ops:
>
> ftruncate(fd, 0);
> pwrite(fd, buf, 1024, 0);
> map = mmap(NULL, 1024, PROT_WRITE, MAP_SHARED, fd, 0);
> map[0] = 'a'; ----> mkwrite allocates 1 block and dirties the whole
> page
> ftruncate(fd, 10000);
>
> Here both non-iomap path (ext4) and iomap path (xfs) will allocate a
> single block but both will mark the whole page range dirty in bh/ifs.
>
> In both iomap and ext4, while writeback, we skip bhs which are holes,
> and we clear dirty on them as well.. The only difference is that in
> iomap xfs_map_blocks() shall detect a hole and inform about the hole but
> in ext4, we filter holes out before the map block call.
Yeah. In the writeback path, almost all of the work is done by a single
call to ext4_map_blocks() now. This makes the writeback code look very
simple and clear. :)
>
> So with this change we will be closer to iomap's behavior. Looks good
> althought it's unfortunate that IO_SUBMIT has yet another implicit
> behavior added to it's scope :)
>
> Feel free to add:
>
> Reviewed-by: Ojaswin Mujoo <ojaswin@linux.ibm.com>
>
>>
>> In the ioend extent conversion path, holes are also not expected because
>> we should wait for folio writeback before punching hole. If one is
>> encountered, it likely indicates a failure in the concurrency
>> protection, so ext4_map_blocks() returns zero, and we keep warning and
>> bail out with -EINVAL to surface the failure rather than continuing
>> conversion on torn data. Atomic writes in
>> ext4_convert_unwritten_extents_atomic() already bail out similarly.
>
> Ahh that's right we do bail out but we don't seem to be returning an
> error in the atomic write path. I think ideally that would be the right
> thing to do. I'll fix it, thanks!
Indeed, thanks for fixing this.
Yi.
>
> Regards,
> ojaswin
>
>>
>> Signed-off-by: Zhang Yi <yi.zhang@huawei.com>
>> ---
>> fs/ext4/inode.c | 7 +++++++
>> 1 file changed, 7 insertions(+)
>>
>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>> index 7a5c74af8ff3..2881596bff8b 100644
>> --- a/fs/ext4/inode.c
>> +++ b/fs/ext4/inode.c
>> @@ -825,6 +825,13 @@ int ext4_map_blocks(handle_t *handle, struct inode *inode,
>> map->m_flags |= EXT4_MAP_MAPPED;
>> goto out_handle;
>> }
>> + } else if (retval == 0) {
>> + /*
>> + * Do not allocate blocks for holes in the context of
>> + * data submission path.
>> + */
>> + if (!map->m_flags && (flags & EXT4_GET_BLOCKS_IO_SUBMIT))
>> + goto out_handle;
>> }
>>
>> if (!handle) {
>> --
>> 2.52.0
>>
next prev parent reply other threads:[~2026-09-28 12:10 UTC|newest]
Thread overview: 41+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 12:35 [PATCH v6 00/31] ext4: use iomap for regular file's buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 01/31] ext4: simplify size updating in ext4_setattr() Zhang Yi
2026-09-03 12:35 ` [PATCH v6 02/31] ext4: factor out ext4_truncate_[up|down]() Zhang Yi
2026-09-03 12:35 ` [PATCH v6 03/31] ext4: skip ordered I/O wait when zeroing beyond i_disksize block Zhang Yi
2026-09-03 12:35 ` [PATCH v6 04/31] ext4: set EXT4_MAP_NEW flag for delayed allocated blocks Zhang Yi
2026-09-24 11:42 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 05/31] ext4: recheck extent status tree before block allocation Zhang Yi
2026-09-24 13:22 ` Ojaswin Mujoo
2026-09-28 7:15 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 06/31] ext4: fix orig_mlen initialization in ext4_map_blocks() Zhang Yi
2026-09-24 13:24 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 07/31] ext4: allow ext4_map_blocks() to start its own transaction handle Zhang Yi
2026-09-28 10:19 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 08/31] ext4: avoid unnecessary transaction in ext4_map_blocks() for unwritten extents Zhang Yi
2026-09-28 10:21 ` Ojaswin Mujoo
2026-09-28 11:57 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 09/31] ext4: skip block allocation for holes in the data submission path Zhang Yi
2026-09-28 10:36 ` Ojaswin Mujoo
2026-09-28 12:10 ` Zhang Yi [this message]
2026-09-03 12:35 ` [PATCH v6 10/31] ext4: add iomap address space operations for buffered I/O Zhang Yi
2026-09-03 12:35 ` [PATCH v6 11/31] ext4: implement buffered read path using iomap Zhang Yi
2026-09-03 12:35 ` [PATCH v6 12/31] ext4: pass out extent seq counter when mapping da blocks Zhang Yi
2026-09-03 12:35 ` [PATCH v6 13/31] ext4: do not use data=ordered mode for inodes using buffered iomap path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 14/31] ext4: implement buffered write path using iomap Zhang Yi
2026-09-03 12:35 ` [PATCH v6 15/31] ext4: implement writeback " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 16/31] ext4: implement mmap " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 17/31] ext4: implement partial block zero range " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 18/31] ext4: drain writeback before removing extents on the iomap path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 19/31] ext4: add block mapping tracepoints for iomap buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 20/31] ext4: disable online defrag when inode using " Zhang Yi
2026-09-03 12:35 ` [PATCH v6 21/31] ext4: add EXT4_STATE_DISKSIZE_GROW_PENDING state bit and helpers Zhang Yi
2026-09-03 12:35 ` [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback Zhang Yi
2026-09-03 12:35 ` [PATCH v6 23/31] ext4: advance i_disksize to i_size upon disksize-grow I/O completion Zhang Yi
2026-09-03 12:35 ` [PATCH v6 24/31] ext4: defer i_disksize update while DISKSIZE_GROW_PENDING is set Zhang Yi
2026-09-03 12:35 ` [PATCH v6 25/31] ext4: submit and wait for disksize-grow I/O in fallocate paths Zhang Yi
2026-09-03 12:35 ` [PATCH v6 26/31] ext4: clear DISKSIZE_GROW_PENDING on truncate or error Zhang Yi
2026-09-03 12:35 ` [PATCH v6 27/31] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block Zhang Yi
2026-09-03 12:35 ` [PATCH v6 28/31] ext4: add tracepoints for DISKSIZE_GROW_PENDING set, clear, and wait Zhang Yi
2026-09-03 12:40 ` [PATCH v6 29/31] ext4: add tracepoints for EOF block zeroing and disksize-grow I/O Zhang Yi
2026-09-03 12:40 ` [PATCH v6 30/31] ext4: partially enable iomap for the buffered I/O path of regular files Zhang Yi
2026-09-03 12:40 ` [PATCH v6 31/31] ext4: introduce a mount option for iomap buffered I/O path Zhang Yi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e3b38169-dc4a-4a42-8a42-b017ecff73d4@gmail.com \
--to=yizhang089@gmail.com \
--cc=adilger.kernel@dilger.ca \
--cc=chengzhihao1@huawei.com \
--cc=djwong@kernel.org \
--cc=hch@infradead.org \
--cc=jack@suse.cz \
--cc=libaokun@linux.alibaba.com \
--cc=linux-ext4@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ojaswin@linux.ibm.com \
--cc=ritesh.list@gmail.com \
--cc=tytso@mit.edu \
--cc=wangkefeng.wang@huawei.com \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yi.zhang@huaweicloud.com \
--cc=yukuai@fnnas.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®