mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v4 00/13] Add RWF_WRITETHROUGH support to iomap & xfs
@ 2026-09-28 12:03 Ojaswin Mujoo
  2026-09-28 12:03 ` [RFC PATCH v4 01/12] xfs: reexpand iter to original count when upgrading to excl ILOCK Ojaswin Mujoo
                   ` (11 more replies)
  0 siblings, 12 replies; 13+ messages in thread
From: Ojaswin Mujoo @ 2026-09-28 12:03 UTC (permalink / raw)
  To: Christian Brauner, linux-fsdevel
  Cc: Darrick J . Wong, Carlos Maiolino, Alexander Viro, Jan Kara,
	Matthew Wilcox, Andrew Morton, Ritesh Harjani, Zhang Yi,
	Christoph Hellwig, Dave Chinner, Daniel Gomez, Pankaj Raghav,
	Theodore Tso, linux-xfs, linux-kernel, linux-mm, Andres Freund

This is the v4 RFC of RWF_WRITETHROUGH patches. Mostly the design is the
same as v3 with some changes based on Pankaj's review (thanks) and
Sashiko's comments. Further, the error handling is improved and more
defined. This series survives xfstests -g quick and fsx/fsstress
stressing, I'll do more rigorous testing once the design is stablized.

I'll quote part of the original cover:

  Hi all,

  This patchset implements buffered write-through IO in linux.

  This idea mainly picked up traction to enable RWF_ATOMIC buffered IO,
  however write-through path can have many use cases beyond atomic writes,
  - such as enabling truly async AIO buffered I/O when issued with O_DSYNC
  - better scalability for buffered I/O

==============================================
** Changes since rfc v3 ** [4]
==============================================

  Patch 1: xfs: reexpand iter to original count when upgrading to excl ILOCK
  * Small fix for xfs write path found while working on similar code in
    writethrough
  
  Patch 5: iomap: Add initial support for buffered RWF_WRITETHROUGH
  * iomap: Prevent duplicate bvec additions for same block across short
    copies in writethrough
  * Error handling changes to ensure address space error is set when we have 
    inconsistent folio state
  
  Patch 6: xfs: Add RWF_WRITETHROUGH support to xfs
  * Add a comment on lockless ip->i_disk_size usage
  * Proper error handling in xfs_writethrough_end_io
  * Restart writethrough checks and re-expand iter when upgrading to exclusive 
    lock
  
  Patch 7: iomap: Add aio support to RWF_WRITETHROUGH
  * Propagate IOCB_NOWAIT correctly
  
  Patch 9: Introduce RWF_NOSERIAL flag to indicate parallel reads/writes
  * NOSERIAL can be passed by user and isn't default on with
    writethrough. It needs to be explicitly passed (as per discussion in
    v3)
  
  Patch 11:  iomap: Avoid folio dirtying in case of RWF_WRITETHROUGH
  * We use iomap_start/end_writeback() instead of just setting the writeback
    bit like before, because we need the writeback xarray tags.


Adding some snippets from the original v3 cover letter (updated) below
as it goes through some important bits of the design:

1. Introduce RWF_NOSERIAL to allow parallel writes
--------------------------------------------------
As per our discussions in LSFMM 2026, we noticed that writethrough (v2
design) suffered from a big regression (~65%) in workloads with multiple
writers writing to a single file. This is because writethrough submits
IO within the inode lock and since buffered IO has an exclusive lock in
write, this hurt massively. In LSFMM, we discussed 3 approaches:
  a) Defer IO submission outside inode lock
  b) Avoid folio dirty - clear cycle
  c) Shared lock for writes

We tried a) by only staging prepared folios in a list under inode lock and then
submitting them outside. However, the initial implementation still showed
around ~30% regression even though the code complexity was significantly
higher. As the ROI was not worth it, we dropped this idea (if there is
interest I can share a github link for these patches). So we finally
decided to go with b) and c).

With b), we can avoid cycling folio through dirty and clear as we immediately
submit the IO. This cuts back on xa_lock contention.

With c), the NOSERIAL flag allows us to do writes under a shared lock if
possible. This violates guarantees that XFS has historically provided however
its expected to be an advanced feature that should be used by applications who
know what they are doing. If used correctly, it can give a very good
performance boost to parallel workloads. This is something that was also
discussed at LSFMM 2026 [3]. Based on previous discussions we've kept
the RWF_NOSERIAL flag serparate to RWF_WRITETHROUGH but, for now, it is only
supported with RWF_WRITETHROUGH.

With b & c, we are able to see a good performance improvement in almost all of
the cases that were regressing. More details and performance numbers
specific to the NOSERIAL flag can be found in the respective patches.


2. Use REQ_SYNC | REQ_IDLE like dio
-----------------------------------
Writethrough doesn't go via the writeback mechanism and it's IO characteristics
are similar to dio. Hence we pass REQ_SYNC | REQ_IDLE in the bio, just like
dio, which allows us to bypass writeback throttling.


3. Move inode i_size update from completion path to write path
--------------------------------------------------------------
As per Dave's suggestion we originally wanted to update isize in
completion like dio however this resulted in a big regression in
extending IO because we end up holding the exclusive lock throughout the
IO. To avoid this, we can just take the buffered IO approach of updating
i_size in write path so that we can safely drop the inode lock and allow
completion to finish outside the lock.  This brings back the append IO
performance in par with buffered IO.

4. There's a deadock in v2 that is fixed in the last patch. If needed
this can be squashed in but I've kept it separate for now for easier
review.

===================================================================
Performance Comparison Tables (with fio snippet)
(Writethrough IO = RWF_WRITETHROUGH + RWF_NOSERIAL)
===================================================================

Table 1: Extending writes using (libaio + O_DSYNC) - All writers on single file
+----------+-------------------+---------------------+
| numjobs  | Buffered IO       | Writethrough IO     |
+----------+-------------------+---------------------+
|    1     | 133 MiB/s         | 133 MiB/s (+0.0%)   |
|    2     | 179 MiB/s         | 243 MiB/s (+35.8%)  |
|    4     | 253 MiB/s         | 358 MiB/s (+41.5%)  |
|    8     | 366 MiB/s         | 376 MiB/s (+2.7%)   |
|   16     | 474 MiB/s         | 449 MiB/s (-5.3%)   |
+----------+-------------------+---------------------+
(fio --ioengine=libaio --writethrough=0/1 --bs=4k --rw=write \
     --iodepth=32 --sync=dsync--file_append=1)


Table 2: Random pure overwrites (libaio + O_DSYNC) - All writers on single file
+----------+-------------------+---------------------+
| numjobs  | Buffered IO       | Writethrough IO     |
+----------+-------------------+---------------------+
|    1     | 131 MiB/s         | 391 MiB/s (+198.5%) |
|    2     | 377 MiB/s         | 791 MiB/s (+109.8%) |
|    4     | 695 MiB/s         | 1591 MiB/s (+128.9%)|
|    8     | 1217 MiB/s        | 1846 MiB/s (+51.7%) |
|   16     | 1197 MiB/s        | 1844 MiB/s (+54.1%) |
+----------+-------------------+---------------------+
(fio --ioengine=libaio --bs=4k --size=2G --sync=dsync --writethrough=0/1 \
    --overwrite=1--rw=randwrite --iodepth=32)


Table 3: Random pure overwrites (libaio + O_DSYNC) - Each write writes own file
+----------+-------------------+---------------------+
| numjobs  | Buffered IO       | Writethrough IO     |
+----------+-------------------+---------------------+
|    1     | 187 MiB/s         | 389 MiB/s (+108.0%) |
|    2     | 373 MiB/s         | 781 MiB/s (+109.4%) |
|    4     | 706 MiB/s         | 1568 MiB/s (+122.1%)|
|    8     | 1222 MiB/s        | 1796 MiB/s (+47.0%) |
+----------+-------------------+---------------------+
(fio --ioengine=libaio --bs=4k --size=2G --filename_format=.test_file.\$jobnum \
     --sync=dsync --writethrough=0/1 --overwrite=1 --rw=randwrite --iodepth=32)


Table 4: Random 16KB overwrites with sync_file_range:16 - Single file
(Roughly mimics postgresql IO pattern)
+----------+-------------------+---------------------+
| numjobs  | Buffered IO       | Writethrough IO     |
+----------+-------------------+---------------------+
|    1     | 1323 MiB/s        | 1275 MiB/s (-3.6%)  |
|    2     | 1779 MiB/s        | 2434 MiB/s (+36.8%) |
|    4     | 2092 MiB/s        | 2519 MiB/s (+20.4%) |
|    8     | 2272 MiB/s        | 2517 MiB/s (+10.8%) |
|   16     | 2328 MiB/s        | 2525 MiB/s (+8.5%)  |
+----------+-------------------+---------------------+
(fio --ioengine=libaio --bs=16k --size=5G --writethrough=0/1 --overwrite=1\
     --sync_file_range=wait_before,write:16 --rw=randwrite --iodepth=32)

* Environment details *

CPU        : IBM Power 11 LPAR
Memory     : 62Gi
Storage    : Samsung PM173-series Enterprise NVMe SSD
Kernel     : Linux 7.2-rc1
Filesystem : XFS (4k block size)


Thoughts and suggestions welcome!

Regards,
ojaswin

[1] https://lore.kernel.org/linux-xfs/cover.1775658795.git.ojaswin@linux.ibm.com/
[2] https://github.com/OjaswinM/xfstests/tree/iomap-buf-writethrough2
[3] https://lwn.net/Articles/1072019
[4] https://lore.kernel.org/linux-xfs/cover.1785908600.git.ojaswin@linux.ibm.com/

Ojaswin Mujoo (12):
  xfs: reexpand iter to original count when upgrading to excl ILOCK
  fs: Add counter to track inflight writes that need stable pages
  mm: Refactor folio_clear_dirty_for_io()
  iomap: Add helper to revert iomap iter
  iomap: Add initial support for buffered RWF_WRITETHROUGH
  xfs: Add RWF_WRITETHROUGH support to xfs
  iomap: Add aio support to RWF_WRITETHROUGH
  iomap: Add DSYNC support to RWF_WRITETHROUGH
  fs: Introduce RWF_NOSERIAL flag to indicate parallel reads/writes
  xfs: Implement RWF_NOSERIAL to parallelize RWF_WRITETHROUGH writes
  iomap: Avoid folio dirtying in case of RWF_WRITETHROUGH
  iomap: Handle deadlock due to repeating folios in RWF_WRITETHROUGH

 fs/iomap/buffered-io.c        | 658 +++++++++++++++++++++++++++++++++-
 fs/iomap/iter.c               |  10 +
 fs/xfs/xfs_file.c             | 155 +++++++-
 include/linux/fs.h            |  22 +-
 include/linux/iomap.h         |  51 +++
 include/linux/pagemap.h       |  15 +-
 include/uapi/linux/fs.h       |   9 +-
 mm/page-writeback.c           |  73 +++-
 tools/include/uapi/linux/fs.h |   8 +-
 9 files changed, 964 insertions(+), 37 deletions(-)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-09-28 12:05 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 12:03 [RFC PATCH v4 00/13] Add RWF_WRITETHROUGH support to iomap & xfs Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 01/12] xfs: reexpand iter to original count when upgrading to excl ILOCK Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 02/12] fs: Add counter to track inflight writes that need stable pages Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 03/12] mm: Refactor folio_clear_dirty_for_io() Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 04/12] iomap: Add helper to revert iomap iter Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 05/12] iomap: Add initial support for buffered RWF_WRITETHROUGH Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 06/12] xfs: Add RWF_WRITETHROUGH support to xfs Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 07/12] iomap: Add aio support to RWF_WRITETHROUGH Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 08/12] iomap: Add DSYNC " Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 09/12] fs: Introduce RWF_NOSERIAL flag to indicate parallel reads/writes Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 10/12] xfs: Implement RWF_NOSERIAL to parallelize RWF_WRITETHROUGH writes Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 11/12] iomap: Avoid folio dirtying in case of RWF_WRITETHROUGH Ojaswin Mujoo
2026-09-28 12:03 ` [RFC PATCH v4 12/12] iomap: Handle deadlock due to repeating folios in RWF_WRITETHROUGH Ojaswin Mujoo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®