From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f174.google.com (mail-pf1-f174.google.com [209.85.210.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C9E3F3859EB for ; Tue, 10 Feb 2026 15:57:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.174 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770739037; cv=none; b=hpFz/PKRxlyuk4PscmQ1cXE1HG/Jm3PLmup+HFu+YXI+0U8iKBbP/n1qRxUWWPLZmL58l8VNSOfj5dPgLn2JeSJ99c8fqBJN4WZzVKV1aAKkJPtF4ZrUyG/KSvd6vlrSLOP+00PR3RlMNYVpka4rbw7NDxUT/LAJo2ZSAQ+8ByI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770739037; c=relaxed/simple; bh=ZjnvTuiYbrA1uIqajO+7i2psmuxO7DSmNzcrr8AuSt8=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=HvkCByba1IFxonKdpGTbLGFpNpqYRlaBC/lcQnn26NWXHbd5uynvwR6e8ZYXjhvwOi4XJFHYUpSVPoeA+yBoXLuUByqOWUmsDu38y2kP2Uq6l7aU40hChaf2wWLIjuJEhIGOrUPHZlAkpl+yM0ekd7kBEe3rtFUkj95Ooq+8UBs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=InO6h68u; arc=none smtp.client-ip=209.85.210.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="InO6h68u" Received: by mail-pf1-f174.google.com with SMTP id d2e1a72fcca58-823075fed75so3416106b3a.1 for ; Tue, 10 Feb 2026 07:57:15 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1770739035; x=1771343835; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=bVZRNpm49w3gc4RCu4RyEyFBgJ9/rzWmQll910jJSYo=; b=InO6h68uWUDRBnUYwcDbR8iSmyskLWlyz5hkOJUrdBa2BWgXXeNAIDrnd4V92UfFQZ JGmrdf4RoAyzERFL0pUnFEBnOUE1XTfDKF0SnlIi52lcg4HBrqzN1gAP2Tg/aB0glI7m yLTJe+1r5fQHIowahCj/XkdWwGPBBuQJ0H9KXHxHRBEgvBXI3wZrL1+6Q/+nKl/E1ZoF JgZiqEOqTflHTy26jDt5hZYQC8udOY9Ib44zpi5pZRCIkoTnHHYT4gNW4qSTF6BAWY/J taq4eR1+4hXSGSN/esWXzuSbNL+KsLMTnVotyO1fYEZeiMkZFYXH5ad7JnQdouPijuhe b7fw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1770739035; x=1771343835; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=bVZRNpm49w3gc4RCu4RyEyFBgJ9/rzWmQll910jJSYo=; b=pCM2qAA+hSu1N4BfvtJkeGqQExJiZwLosP8WSikE/7IzKi9O3YRxiBeC5QWrYv8n5b 79T6M2i1F/ZyC3TGiENdcwIqce15Shbq9OrRk8vlfT8Sc6ARA2Toc7oVLX28UhLZ8W82 4jZxlP+skAEAjN9HnghYNJlB+RD4g8uc+7/shWZ0xhTsvKBSizP0Lq9iz6XsPqxa2DnP 0F8COisSv7LzJz+DYPvJqUh2KLT5qc8455zMgPnl3KIwZVSlyy9aHanyaTRenkhBE72Z 0d/m1J3T76cBx9XZ9A+O5kvpQ4tUwZHUZW0fn6rrClQXrVZ9QWE4NtkzZXUS4+v7F6SR JCtg== X-Forwarded-Encrypted: i=1; AJvYcCVhzD5Wf/TsgYxKnl7GVMcVAfctx7fcYqE9vx7fjITczsbRURxFPujHRLS1xPfdQU57UdZj1l55t0/YWa0=@vger.kernel.org X-Gm-Message-State: AOJu0YwglGwLbnFaFrUM2r/9mwuFVtHNPDQQyQY6Twb3GOaZ7oDfv5ws eYHuiXFEErLO25ImPKVlgrs0DoB0Ma9MbXQ2UtjQjUXxB9YDEl924kvh X-Gm-Gg: AZuq6aL2/zKbV57ZuU+JLZBXgKc+t4Lahe4fZJKLdUIVCcpeVqeeYhtSjH7Ifz+qGWT yYUZU/6iThC65noe3Eo9G+MX9HQqW2658ROadu4WcljP2XRPyi7rSVPV8dxGSvFB+Rai1insJpt wOCev6AX+FpnI5cI8mpAhvILpaMhb4eZXhnq1ybrdIigmoPdRYpSLmiRSW76Yrq8IdC9rujXsx8 NaGysftwuUxm5Y1tNF0HiqB8Hc4Wpsqbl5i14FurQRkQ0pSjeu9iGITtwbNywOQu7Y02U55yrdY +JZRSssTq2vupjk1vnULHySyCFj5JkTQTg6ElC0ZtyJSCBwVjGL+wBhDP7IJpGwlnyM7HX7N503 MmKxSrXRO1jPj/NTT1sgA0XLY35OfBwHUZJ+iZ65cREhcB/V6SxEOEfrW994onwX0AYQBQZY+kh g3hy/izK87yBBLogYlFN8tG0rzAGliZ25ouk9R+PagRTCpKofmmVmlcPCARLJU3wbKML3VSB2FP Q== X-Received: by 2002:a05:6a00:399b:b0:81f:4708:b46e with SMTP id d2e1a72fcca58-824877cafd2mr2557597b3a.20.1770739035023; Tue, 10 Feb 2026 07:57:15 -0800 (PST) Received: from ?IPV6:240e:390:a90:6d21:e579:6116:b665:1484? ([240e:390:a90:6d21:e579:6116:b665:1484]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-8244168fdf5sm14013049b3a.17.2026.02.10.07.57.10 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 10 Feb 2026 07:57:14 -0800 (PST) Message-ID: <04b0a510-0a97-464f-a6d3-8410fff9243d@gmail.com> Date: Tue, 10 Feb 2026 23:57:03 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH -next v2 03/22] ext4: only order data when partially block truncating down To: Ojaswin Mujoo , Zhang Yi Cc: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, jack@suse.cz, ritesh.list@gmail.com, hch@infradead.org, djwong@kernel.org, yi.zhang@huaweicloud.com, libaokun1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com References: <20260203062523.3869120-1-yi.zhang@huawei.com> <20260203062523.3869120-4-yi.zhang@huawei.com> Content-Language: en-US From: Zhang Yi In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 2/10/2026 3:05 PM, Ojaswin Mujoo wrote: > On Tue, Feb 03, 2026 at 02:25:03PM +0800, Zhang Yi wrote: >> Currently, __ext4_block_zero_page_range() is called in the following >> four cases to zero out the data in partial blocks: >> >> 1. Truncate down. >> 2. Truncate up. >> 3. Perform block allocation (e.g., fallocate) or append writes across a >> range extending beyond the end of the file (EOF). >> 4. Partial block punch hole. >> >> If the default ordered data mode is used, __ext4_block_zero_page_range() >> will write back the zeroed data to the disk through the order mode after >> zeroing out. >> >> Among the cases 1,2 and 3 described above, only case 1 actually requires >> this ordered write. Assuming no one intentionally bypasses the file >> system to write directly to the disk. When performing a truncate down >> operation, ensuring that the data beyond the EOF is zeroed out before >> updating i_disksize is sufficient to prevent old data from being exposed >> when the file is later extended. In other words, as long as the on-disk >> data in case 1 can be properly zeroed out, only the data in memory needs >> to be zeroed out in cases 2 and 3, without requiring ordered data. >> >> Case 4 does not require ordered data because the entire punch hole >> operation does not provide atomicity guarantees. Therefore, it's safe to >> move the ordered data operation from __ext4_block_zero_page_range() to >> ext4_truncate(). >> >> It should be noted that after this change, we can only determine whether >> to perform ordered data operations based on whether the target block has >> been zeroed, rather than on the state of the buffer head. Consequently, >> unnecessary ordered data operations may occur when truncating an >> unwritten dirty block. However, this scenario is relatively rare, so the >> overall impact is minimal. >> >> This is prepared for the conversion to the iomap infrastructure since it >> doesn't use ordered data mode and requires active writeback, which >> reduces the complexity of the conversion. > > Hi Yi, > > Took me quite some time to understand what we are doing here, I'll > just add my understanding here to confirm/document :) Hi, Ojaswin! Thank you for review and test this series. > > So your argument is that currently all paths that change the i_size take > care of zeroing the (newsize, eof block boundary) before i_size change > is seen by users: > - dio does it in iomap_dio_bio_iter if IOMAP_UNWRITTEN (true for first allocation) > - buffered IO/mmap write does it in ext4_da_write_begin() -> > ext4_block_write_begin() for buffer_new (true for first allocation) > - falloc doesn't zero the new eof block but it allocates an unwrit > extent so no stale data issue. When an allocation happens from the > above 2 methods then we anyways will zero it. These two zeroing operations mentioned above are mainly used to initialize newly allocated blocks, which is not the main focus of this discussion. The focus of this discussion is how to clear the portions of allocated blocks that extend beyond the EOF. > - truncate down also takes care of this via ext4_truncate() -> > ext4_block_truncate_page() > > Now, parallely there are also codepaths that say grow the i_size but > then also zero the (old_size, block boundary) range before the i_size > commits. This is so that they want to be sure the newly visible range > doesn't expose stale data. > For example: > - truncate up from 2kb to 8kb will zero (2kb,4kb) via ext4_block_truncate_page() > - with i_size = 2kb, buffered IO at 6kb would zero 2kb,4kb in ext4_da_write_end() Yes, you are right. > - I'm unable to see if/where we do it via dio path. I don't see it too, so I think this is also a problem. > > You originally proposed that we can remove the logic to zeroout > (old_size, block_boundary) in data=ordered fashion, ie we don't need to > trigger the zeroout IO before the i_size change commits, we can just zero the > range in memory because we would have already zeroed them earlier when > we had allocated at old_isize, or truncated down to old_isize. Yes. > > To this Jan pointed out that although we take care to zeroout (new_size, > block_boundary) its not enough because we could still end up with data > past eof: > > 1. race of buffered write vs mmap write past eof. i_size = 2kb, > we write (2kb, 3kb). > 2. The write goes through but we crash before i_size=3kb txn can commit. > Again we have data past 2kb ie the eof block. > Yes. > Now, Im still looking into this part but the reason we want to get rid of > this data=ordered IO is so that we don't trigger a writeback due to > journal commit which tries to acquire folio_lock of a folio already > locked by iomap. Yes, and iomap will start a new transaction under the folio lock, which may also wait the current committing transaction to finish. > However we will now try an alternate way to get past > this. > > Is my understanding correct? Yes. Cheers, Yi. > > Regards, > ojaswin > > PS: -g auto tests are passing (no regressions) with 64k and 4k bs on > powerpc 64k pagesize box so thats nice :D > >> >> Signed-off-by: Zhang Yi >> --- >> fs/ext4/inode.c | 32 +++++++++++++++++++------------- >> 1 file changed, 19 insertions(+), 13 deletions(-) >> >> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c >> index f856ea015263..20b60abcf777 100644 >> --- a/fs/ext4/inode.c >> +++ b/fs/ext4/inode.c >> @@ -4106,19 +4106,10 @@ static int __ext4_block_zero_page_range(handle_t *handle, >> folio_zero_range(folio, offset, length); >> BUFFER_TRACE(bh, "zeroed end of block"); >> >> - if (ext4_should_journal_data(inode)) { >> + if (ext4_should_journal_data(inode)) >> err = ext4_dirty_journalled_data(handle, bh); >> - } else { >> + else >> mark_buffer_dirty(bh); >> - /* >> - * Only the written block requires ordered data to prevent >> - * exposing stale data. >> - */ >> - if (!buffer_unwritten(bh) && !buffer_delay(bh) && >> - ext4_should_order_data(inode)) >> - err = ext4_jbd2_inode_add_write(handle, inode, from, >> - length); >> - } >> if (!err && did_zero) >> *did_zero = true; >> >> @@ -4578,8 +4569,23 @@ int ext4_truncate(struct inode *inode) >> goto out_trace; >> } >> >> - if (inode->i_size & (inode->i_sb->s_blocksize - 1)) >> - ext4_block_truncate_page(handle, mapping, inode->i_size); >> + if (inode->i_size & (inode->i_sb->s_blocksize - 1)) { >> + unsigned int zero_len; >> + >> + zero_len = ext4_block_truncate_page(handle, mapping, >> + inode->i_size); >> + if (zero_len < 0) { >> + err = zero_len; >> + goto out_stop; >> + } >> + if (zero_len && !IS_DAX(inode) && >> + ext4_should_order_data(inode)) { >> + err = ext4_jbd2_inode_add_write(handle, inode, >> + inode->i_size, zero_len); >> + if (err) >> + goto out_stop; >> + } >> + } >> >> /* >> * We add the inode to the orphan list, so that if this >> -- >> 2.52.0 >>