mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3 0/3] fs: drain in-flight DIO before buffered write fallback
@ 2026-09-28  8:54 Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 1/3] ext2: " Jiale Yao
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Jiale Yao @ 2026-09-28  8:54 UTC (permalink / raw)
  To: Namjae Jeon, Sungjong Seo, Yuezhang Mo, Jan Kara, Hyunchul Lee,
	Ritesh Harjani (IBM),
	Darrick J. Wong, exfat, linux-kernel, linux-ext4, ntfs
  Cc: Jiale Yao

An asynchronous direct write can remain in flight after its submitting
thread releases the inode lock.  If a buffered write dirties page cache
in the meantime, the direct write can complete post-I/O invalidation after
the pages become dirty.  The invalidation then reports a page cache
invalidation failure and records -EIO in the mapping error sequence.  A
later fsync() returns -EIO.

Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
in-flight DIO before buffered write fallback") fixed this race in ext4.
The same ordering is missing from the buffered fallback paths in ext2,
NTFS, and exFAT, and from the regular buffered-write path in exFAT.

This series adds inode_dio_wait() before these paths dirty page cache.
Each patch fixes one filesystem and remains independently buildable.  The
NTFS patch also avoids entering the blocking fallback for IOCB_NOWAIT
requests: it returns -EAGAIN if DIO made no progress and preserves a
positive result after a partial direct write.

A reproducer using concurrent AIO direct writes and buffered fallback
triggered the following warning on all three filesystems and made a
subsequent fsync() return -EIO:

  Page cache invalidation failure on direct I/O.  Possible data corruption
  due to collision with buffered I/O!

Changes in v3:
- Add inode_dio_wait() to both the regular and fallback buffered-write
  paths in exFAT, and move the explanatory comment to
  exfat_file_write_iter(), as requested by Chi Zhiling.
- Add Baolin Liu's Reviewed-by tag to the NTFS patch.

Changes in v2:
- Handle IOCB_NOWAIT before the potentially blocking NTFS fallback,
  returning -EAGAIN before any bytes are written and preserving a positive
  short-write result otherwise, as requested in review.

Jiale Yao (3):
  ext2: drain in-flight DIO before buffered write fallback
  ntfs: drain in-flight DIO before buffered write fallback
  exfat: drain in-flight DIO before buffered writes

 fs/exfat/file.c | 12 ++++++++++--
 fs/ext2/file.c  |  7 +++++++
 fs/ntfs/file.c  | 13 +++++++++++++
 3 files changed, 30 insertions(+), 2 deletions(-)

-- 
2.34.1


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 1/3] ext2: drain in-flight DIO before buffered write fallback
  2026-09-28  8:54 [PATCH v3 0/3] fs: drain in-flight DIO before buffered write fallback Jiale Yao
@ 2026-09-28  8:54 ` Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 2/3] ntfs: " Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes Jiale Yao
  2 siblings, 0 replies; 6+ messages in thread
From: Jiale Yao @ 2026-09-28  8:54 UTC (permalink / raw)
  To: Jan Kara, Ritesh Harjani (IBM), linux-ext4, linux-kernel; +Cc: Jiale Yao

An asynchronous direct write can remain in flight after the inode lock is
released.  If another direct write falls back to buffered I/O while the
first write is still pending, generic_perform_write() can dirty pages
before the first write completes its post-I/O page cache invalidation.
The invalidation then finds dirty pages, reports a page cache invalidation
failure, and records -EIO in the mapping error sequence.  A later fsync()
therefore returns -EIO.

Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
in-flight DIO before buffered write fallback") fixed the same race in
ext4.  Ext2 has an equivalent fallback after iomap_dio_rw() returns
-ENOTBLK or a short write, but does not drain other in-flight DIO before
dirtying the page cache.

Wait for in-flight DIO before calling generic_perform_write() in the
fallback path.

A reproducer using concurrent AIO direct writes and buffered fallback
triggered the following warning and made a subsequent fsync() return
-EIO:

  Page cache invalidation failure on direct I/O.  Possible data corruption
  due to collision with buffered I/O!

Fixes: fb5de4358e1a ("ext2: Move direct-io to use iomap")
Link: https://lore.kernel.org/r/20260629113827.4074335-3-libaokun@linux.alibaba.com
Signed-off-by: Jiale Yao <yaojiale02@163.com>
---
 fs/ext2/file.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/fs/ext2/file.c b/fs/ext2/file.c
index b9020df7d89e..67fe423c3828 100644
--- a/fs/ext2/file.c
+++ b/fs/ext2/file.c
@@ -135,6 +135,13 @@ static ssize_t ext2_dio_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		int ret2;
 
 		iocb->ki_flags &= ~IOCB_DIRECT;
+
+		/*
+		 * Prevent concurrent direct I/O and buffered I/O to the same file
+		 * range. Wait for in-flight DIO to finish before dirtying pages.
+		 */
+		inode_dio_wait(inode);
+
 		pos = iocb->ki_pos;
 		status = generic_perform_write(iocb, from);
 		if (unlikely(status < 0)) {
-- 
2.34.1


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 2/3] ntfs: drain in-flight DIO before buffered write fallback
  2026-09-28  8:54 [PATCH v3 0/3] fs: drain in-flight DIO before buffered write fallback Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 1/3] ext2: " Jiale Yao
@ 2026-09-28  8:54 ` Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes Jiale Yao
  2 siblings, 0 replies; 6+ messages in thread
From: Jiale Yao @ 2026-09-28  8:54 UTC (permalink / raw)
  To: Namjae Jeon, Hyunchul Lee, ntfs, linux-kernel; +Cc: Jiale Yao, Baolin Liu

An asynchronous direct write can remain in flight after the inode lock is
released.  If another direct write falls back to buffered I/O while the
first write is still pending, iomap_file_buffered_write() can dirty pages
before the first write completes its post-I/O page cache invalidation.
The invalidation then finds dirty pages, reports a page cache invalidation
failure, and records -EIO in the mapping error sequence.  A later fsync()
therefore returns -EIO.

Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
in-flight DIO before buffered write fallback") fixed the same race in
ext4.  NTFS has an equivalent fallback after iomap_dio_rw() returns
-ENOTBLK or a short write, but does not drain other in-flight DIO before
dirtying the page cache.

Wait for in-flight DIO before calling iomap_file_buffered_write() in the
fallback path.  Since NTFS supports IOCB_NOWAIT, do not enter the blocking
fallback for such requests.  Return -EAGAIN if no bytes were written, or
preserve the positive short-write result if the direct write made partial
progress.

A reproducer using concurrent AIO direct writes and buffered fallback
triggered the following warning and made a subsequent fsync() return
-EIO:

  Page cache invalidation failure on direct I/O.  Possible data corruption
  due to collision with buffered I/O!

Fixes: 9c87959601e8 ("ntfs: update file operations")
Link: https://lore.kernel.org/r/20260629113827.4074335-3-libaokun@linux.alibaba.com
Reviewed-by: Baolin Liu <liubaolin@kylinos.cn>
Signed-off-by: Jiale Yao <yaojiale02@163.com>
---
 fs/ntfs/file.c | 13 +++++++++++++
 1 file changed, 13 insertions(+)

diff --git a/fs/ntfs/file.c b/fs/ntfs/file.c
index 007d1614b9ac..8bbfa842889a 100644
--- a/fs/ntfs/file.c
+++ b/fs/ntfs/file.c
@@ -525,8 +525,21 @@ static ssize_t ntfs_dio_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		ssize_t written;
 		int ret2;
 
+		if (iocb->ki_flags & IOCB_NOWAIT) {
+			if (!ret)
+				ret = -EAGAIN;
+			goto out;
+		}
+
 		offset = iocb->ki_pos;
 		iocb->ki_flags &= ~IOCB_DIRECT;
+
+		/*
+		 * Prevent concurrent direct I/O and buffered I/O to the same file
+		 * range. Wait for in-flight DIO to finish before dirtying pages.
+		 */
+		inode_dio_wait(file_inode(iocb->ki_filp));
+
 		written = iomap_file_buffered_write(iocb, from,
 				&ntfs_write_iomap_ops, &ntfs_iomap_folio_ops,
 				NULL);
-- 
2.34.1


^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes
  2026-09-28  8:54 [PATCH v3 0/3] fs: drain in-flight DIO before buffered write fallback Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 1/3] ext2: " Jiale Yao
  2026-09-28  8:54 ` [PATCH v3 2/3] ntfs: " Jiale Yao
@ 2026-09-28  8:54 ` Jiale Yao
  2026-10-03  1:17   ` Chi Zhiling
  2 siblings, 1 reply; 6+ messages in thread
From: Jiale Yao @ 2026-09-28  8:54 UTC (permalink / raw)
  To: Namjae Jeon, Sungjong Seo, Yuezhang Mo, Darrick J. Wong, exfat,
	linux-kernel
  Cc: Jiale Yao

An asynchronous direct write can remain in flight after the inode lock is
released.  If a buffered write dirties page cache while the direct write
is still pending, the direct write's post-I/O invalidation can find the
dirty pages, report a page cache invalidation failure, and record -EIO in
the mapping error sequence.  A later fsync() therefore returns -EIO.

Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
in-flight DIO before buffered write fallback") fixed the same race in
ext4.  ExFAT does not drain in-flight DIO before either a regular buffered
write or the buffered fallback after iomap_dio_rw() returns -ENOTBLK or a
short write.

Wait for in-flight DIO before calling iomap_file_buffered_write() in both
paths.

A reproducer using concurrent AIO direct writes and buffered fallback
triggered the following warning and made a subsequent fsync() return
-EIO:

  Page cache invalidation failure on direct I/O.  Possible data corruption
  due to collision with buffered I/O!

Fixes: 867b9c96dc83 ("exfat: add iomap direct I/O support")
Link: https://lore.kernel.org/r/20260629113827.4074335-3-libaokun@linux.alibaba.com
Signed-off-by: Jiale Yao <yaojiale02@163.com>
---
 fs/exfat/file.c | 12 ++++++++++--
 1 file changed, 10 insertions(+), 2 deletions(-)

diff --git a/fs/exfat/file.c b/fs/exfat/file.c
index a2a9ee1a2004..28811210663f 100644
--- a/fs/exfat/file.c
+++ b/fs/exfat/file.c
@@ -807,6 +807,8 @@ static ssize_t exfat_fallback_buffered_write(struct kiocb *iocb,
 
 	iocb->ki_flags &= ~IOCB_DIRECT;
 
+	inode_dio_wait(file_inode(iocb->ki_filp));
+
 	written = iomap_file_buffered_write(iocb, from, &exfat_write_iomap_ops,
 			NULL, NULL);
 	if (written < 0)
@@ -889,11 +891,17 @@ static ssize_t exfat_file_write_iter(struct kiocb *iocb, struct iov_iter *iter)
 			goto unlock;
 	}
 
-	if (iocb->ki_flags & IOCB_DIRECT)
+	if (iocb->ki_flags & IOCB_DIRECT) {
 		ret = exfat_dio_write_iter(iocb, iter);
-	else
+	} else {
+		/*
+		 * Prevent concurrent direct I/O and buffered I/O to the same file
+		 * range. Wait for in-flight DIO to finish before dirtying pages.
+		 */
+		inode_dio_wait(inode);
 		ret = iomap_file_buffered_write(iocb, iter,
 				&exfat_write_iomap_ops, NULL, NULL);
+	}
 	if (ret < 0)
 		goto unlock;
 
-- 
2.34.1


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes
  2026-09-28  8:54 ` [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes Jiale Yao
@ 2026-10-03  1:17   ` Chi Zhiling
  2026-10-03  9:32     ` jiale yao
  0 siblings, 1 reply; 6+ messages in thread
From: Chi Zhiling @ 2026-10-03  1:17 UTC (permalink / raw)
  To: Jiale Yao, Namjae Jeon, Sungjong Seo, Yuezhang Mo,
	Darrick J. Wong, exfat, linux-kernel

On 9/28/26 4:54 PM, Jiale Yao wrote:
> An asynchronous direct write can remain in flight after the inode lock is
> released.  If a buffered write dirties page cache while the direct write
> is still pending, the direct write's post-I/O invalidation can find the
> dirty pages, report a page cache invalidation failure, and record -EIO in
> the mapping error sequence.  A later fsync() therefore returns -EIO.
> 
> Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
> in-flight DIO before buffered write fallback") fixed the same race in
> ext4.  ExFAT does not drain in-flight DIO before either a regular buffered
> write or the buffered fallback after iomap_dio_rw() returns -ENOTBLK or a
> short write.
> 
> Wait for in-flight DIO before calling iomap_file_buffered_write() in both
> paths.
> 
> A reproducer using concurrent AIO direct writes and buffered fallback
> triggered the following warning and made a subsequent fsync() return
> -EIO:
> 
>    Page cache invalidation failure on direct I/O.  Possible data corruption
>    due to collision with buffered I/O!
> 
> Fixes: 867b9c96dc83 ("exfat: add iomap direct I/O support")
> Link: https://lore.kernel.org/r/20260629113827.4074335-3-libaokun@linux.alibaba.com
> Signed-off-by: Jiale Yao <yaojiale02@163.com>
> ---
>   fs/exfat/file.c | 12 ++++++++++--
>   1 file changed, 10 insertions(+), 2 deletions(-)
> 
> diff --git a/fs/exfat/file.c b/fs/exfat/file.c
> index a2a9ee1a2004..28811210663f 100644
> --- a/fs/exfat/file.c
> +++ b/fs/exfat/file.c
> @@ -807,6 +807,8 @@ static ssize_t exfat_fallback_buffered_write(struct kiocb *iocb,
>   
>   	iocb->ki_flags &= ~IOCB_DIRECT;
>   
> +	inode_dio_wait(file_inode(iocb->ki_filp));
> +
>   	written = iomap_file_buffered_write(iocb, from, &exfat_write_iomap_ops,
>   			NULL, NULL);
>   	if (written < 0)
> @@ -889,11 +891,17 @@ static ssize_t exfat_file_write_iter(struct kiocb *iocb, struct iov_iter *iter)
>   			goto unlock;
>   	}
>   
> -	if (iocb->ki_flags & IOCB_DIRECT)
> +	if (iocb->ki_flags & IOCB_DIRECT) {
>   		ret = exfat_dio_write_iter(iocb, iter);
> -	else
> +	} else {
> +		/*
> +		 * Prevent concurrent direct I/O and buffered I/O to the same file
> +		 * range. Wait for in-flight DIO to finish before dirtying pages.
> +		 */
> +		inode_dio_wait(inode);
>   		ret = iomap_file_buffered_write(iocb, iter,
>   				&exfat_write_iomap_ops, NULL, NULL);
> +	}
>   	if (ret < 0)
>   		goto unlock;
>   

Also, if you don't mind, could you add inode_dio_wait() to 
exfat_extend_valid_size() as well?

Thanks.


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re:Re: [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes
  2026-10-03  1:17   ` Chi Zhiling
@ 2026-10-03  9:32     ` jiale yao
  0 siblings, 0 replies; 6+ messages in thread
From: jiale yao @ 2026-10-03  9:32 UTC (permalink / raw)
  To: Chi Zhiling
  Cc: Namjae Jeon, Sungjong Seo, Yuezhang Mo, Darrick J. Wong, exfat,
	linux-kernel

At 2026-10-03 09:17:18, "Chi Zhiling" <chizhiling@163.com> wrote:
>On 9/28/26 4:54 PM, Jiale Yao wrote:
>> An asynchronous direct write can remain in flight after the inode lock is
>> released.  If a buffered write dirties page cache while the direct write
>> is still pending, the direct write's post-I/O invalidation can find the
>> dirty pages, report a page cache invalidation failure, and record -EIO in
>> the mapping error sequence.  A later fsync() therefore returns -EIO.
>> 
>> Commit 15cdefd0c0522f9d5e12d947fa04f4c11649b699 ("ext4: drain
>> in-flight DIO before buffered write fallback") fixed the same race in
>> ext4.  ExFAT does not drain in-flight DIO before either a regular buffered
>> write or the buffered fallback after iomap_dio_rw() returns -ENOTBLK or a
>> short write.
>> 
>> Wait for in-flight DIO before calling iomap_file_buffered_write() in both
>> paths.
>> 
>> A reproducer using concurrent AIO direct writes and buffered fallback
>> triggered the following warning and made a subsequent fsync() return
>> -EIO:
>> 
>>    Page cache invalidation failure on direct I/O.  Possible data corruption
>>    due to collision with buffered I/O!
>> 
>> Fixes: 867b9c96dc83 ("exfat: add iomap direct I/O support")
>> Link: https://lore.kernel.org/r/20260629113827.4074335-3-libaokun@linux.alibaba.com
>> Signed-off-by: Jiale Yao <yaojiale02@163.com>
>> ---
>>   fs/exfat/file.c | 12 ++++++++++--
>>   1 file changed, 10 insertions(+), 2 deletions(-)
>> 
>> diff --git a/fs/exfat/file.c b/fs/exfat/file.c
>> index a2a9ee1a2004..28811210663f 100644
>> --- a/fs/exfat/file.c
>> +++ b/fs/exfat/file.c
>> @@ -807,6 +807,8 @@ static ssize_t exfat_fallback_buffered_write(struct kiocb *iocb,
>>   
>>   	iocb->ki_flags &= ~IOCB_DIRECT;
>>   
>> +	inode_dio_wait(file_inode(iocb->ki_filp));
>> +
>>   	written = iomap_file_buffered_write(iocb, from, &exfat_write_iomap_ops,
>>   			NULL, NULL);
>>   	if (written < 0)
>> @@ -889,11 +891,17 @@ static ssize_t exfat_file_write_iter(struct kiocb *iocb, struct iov_iter *iter)
>>   			goto unlock;
>>   	}
>>   
>> -	if (iocb->ki_flags & IOCB_DIRECT)
>> +	if (iocb->ki_flags & IOCB_DIRECT) {
>>   		ret = exfat_dio_write_iter(iocb, iter);
>> -	else
>> +	} else {
>> +		/*
>> +		 * Prevent concurrent direct I/O and buffered I/O to the same file
>> +		 * range. Wait for in-flight DIO to finish before dirtying pages.
>> +		 */
>> +		inode_dio_wait(inode);
>>   		ret = iomap_file_buffered_write(iocb, iter,
>>   				&exfat_write_iomap_ops, NULL, NULL);
>> +	}
>>   	if (ret < 0)
>>   		goto unlock;
>>   
>
>Also, if you don't mind, could you add inode_dio_wait() to 
>exfat_extend_valid_size() as well?
Sure, v4 in https://lore.kernel.org/all/20261003093035.532916-1-yaojiale02@163.com/
>
>Thanks.

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-10-03  9:32 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28  8:54 [PATCH v3 0/3] fs: drain in-flight DIO before buffered write fallback Jiale Yao
2026-09-28  8:54 ` [PATCH v3 1/3] ext2: " Jiale Yao
2026-09-28  8:54 ` [PATCH v3 2/3] ntfs: " Jiale Yao
2026-09-28  8:54 ` [PATCH v3 3/3] exfat: drain in-flight DIO before buffered writes Jiale Yao
2026-10-03  1:17   ` Chi Zhiling
2026-10-03  9:32     ` jiale yao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®