* [PATCH 1/8] xfs: map pNFS layouts to the end of the extent again
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap Daejun Park via B4 Relay
` (6 subsequent siblings)
7 siblings, 0 replies; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, stable, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
Since commit 36ca6f11424a ("xfs: fix overlapping extents returned for
pNFS LAYOUTGET"), xfs_fs_map_blocks() maps only the range that nfsd asks
for. For O_DIRECT, the Linux block layout client asks for the range of
the I/O at hand, so it now needs a LAYOUTGET for every O_DIRECT read or
write to a part of a file it has no layout for yet, where the first
LAYOUTGET used to return the whole extent. Each LAYOUTGET is a round
trip, and xfs_fs_map_blocks() takes the iolock exclusively and writes
back and invalidates the page cache of the file for it.
On three QEMU VMs (an NVMe/TCP target, the server with nfsd and XFS, and
a client) running the same kernel without KASAN or lock debugging, a
client doing O_DIRECT I/O over the block layout of NFSv4.2 for 30
seconds to a 1 GiB file of one written extent, in order unless noted and
with an fsync every 16 MiB written, gets this many I/Os done (and
LAYOUTGETs counted at the server), one run each:
unpatched ENTIRE, no trim this patch
4 KiB reads 41207 (41207) 334931 (1) 343749 (1)
4 KiB random 36670 (34148) 259527 (1) 263829 (15)
4 KiB overwrite 38556 (38556) 325533 (1) 307105 (1)
64 KiB reads 70241 (16384) 80829 (1) 91196 (1)
1 MiB reads 6857 (1024) 6759 (1) 6706 (1)
"ENTIRE, no trim" is a test kernel that maps with XFS_BMAPI_ENTIRE and
does not trim. The random reads need a LAYOUTGET each time they go
before the lowest offset read so far, as the part of the extent before
the offset asked for is not mapped. 1 MiB overwrites and writes to an
unwritten extent also go from 1024 LAYOUTGETs to 1, with no clear change
in I/Os done. Appending to a new file needs a LAYOUTGET per 1 MiB
appended either way, as XFS allocates only the range asked for on a
filesystem without a stripe unit or an extent size hint.
The overlap fixed by that commit comes from XFS_BMAPI_ENTIRE reaching
back. nfsd4_block_proc_layoutget() calls ->map_blocks once per extent
of a LAYOUTGET, each time for the range left after the previous extent.
An allocation in one call can merge the extent that the next call starts
in with the extents before it, and the whole extent then starts before
the offset of that call and overlaps extents already in the layout. The
Linux client rejects such a layout (verify_extent() returns -EIO), and
the I/O that needed it fails.
In the thread of that commit, Christoph Hellwig asked for
XFS_BMAPI_ENTIRE to be dropped to stop the overlap, Darrick J. Wong
agreed, and it was said that the flag makes no difference on the first
call, which is for the whole range the client asked for. The difference
is past that range: the Linux client asks for the range of each O_DIRECT
I/O and uses the rest of a longer layout for the I/Os that follow, which
RFC 8881 allows (Table 22 sets only a minimum length). A client could
ask for more for a read layout, but for a write layout
xfs_fs_map_blocks() allocates any hole in the range asked for.
Dave Chinner suggested keeping the flag for the first call and trimming
the mappings of the calls that follow. ->map_blocks cannot tell the
first call from the others, and Christoph preferred to keep such a
choice out of that interface, so trim every mapping to start at the
offset asked for instead. No mapping can then overlap the one before
it, since each call starts where the previous extent ended. Unlike
Dave's suggestion, the first mapping loses the part of the extent before
the offset, and the last mapping keeps the part past the end of the
range, which the loop in nfsd4_block_proc_layoutget() already handles.
Trimming in that loop instead would keep his suggestion exactly, but
would change nfsd as well, while trimming in XFS keeps every mapping it
returns free of overlap whatever the caller does.
Unlike before that commit, don't let the mapping reach past EOF beyond
the range asked for. On an inode without XFS_DIFLAG_PREALLOC or
XFS_DIFLAG_APPEND, xfs_free_eofblocks() can free blocks past EOF, such
as speculative preallocation, without breaking the layout, so a client
that did not ask for them should not get them. What a client gets for a
range past EOF that it asks for does not change.
The aio group, the fsx tests and generic/013 (fsstress) of fstests over
the block layout, and the fsx tests and generic/013 over the SCSI
layout, give the same results with and without this patch. generic/091
and 263 fail with the ENTIRE, no trim kernel and pass with this patch,
and so does the pynfs test BLOCK5, which checks a write layout over a
hole between two allocated blocks against three rules of RFC 5663
section 2.3.1.
Fixes: 36ca6f11424a ("xfs: fix overlapping extents returned for pNFS LAYOUTGET")
Cc: stable@vger.kernel.org
Suggested-by: Dave Chinner <dgc@kernel.org>
Link: https://lore.kernel.org/r/ageSguSyf2kBY33a@dread
Link: https://lore.kernel.org/r/agwDhixPAAA0-cTa@infradead.org
Link: https://lore.kernel.org/r/agqfBPRWXQDR2ImG@infradead.org
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_pnfs.c | 19 +++++++++++++++++--
1 file changed, 17 insertions(+), 2 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index f8535ecde2..ab38561700 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -183,13 +183,28 @@ xfs_fs_map_blocks(
offset_fsb = XFS_B_TO_FSBT(mp, offset);
lock_flags = xfs_ilock_data_map_shared(ip);
- /* request mappings for the specified range only */
+ /*
+ * Map to the end of the extent that covers the start of the range,
+ * so that a client doing I/O in pieces gets a layout it can use for
+ * the pieces that follow. Never map anything before the start of
+ * the range: nfsd calls in here once per extent of a LAYOUTGET, for
+ * the range that is left after the previous extent, and the mapping
+ * can change in between, so a mapping that reaches back can overlap
+ * one already in the layout. Don't extend the mapping past EOF
+ * beyond the range either: xfs_free_eofblocks() can free blocks past
+ * EOF without breaking the layout.
+ */
error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb,
- &imap, &nimaps, 0);
+ &imap, &nimaps, XFS_BMAPI_ENTIRE);
if (error) {
xfs_iunlock(ip, lock_flags);
goto out_unlock;
}
+ if (nimaps)
+ xfs_trim_extent(&imap, offset_fsb,
+ max_t(xfs_fileoff_t, end_fsb,
+ XFS_B_TO_FSB(mp, XFS_ISIZE(ip))) -
+ offset_fsb);
seq = xfs_iomap_inode_sequence(ip, 0);
ASSERT(!nimaps || imap.br_startblock != DELAYSTARTBLOCK);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 1/8] xfs: map pNFS layouts to the end of the extent again Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 15:40 ` Darrick J. Wong
2026-10-08 1:39 ` [PATCH 3/8] xfs: clamp the pNFS layout range to the maximum file size Daejun Park via B4 Relay
` (5 subsequent siblings)
7 siblings, 1 reply; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, stable, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
xfs_iomap_write_direct() maps one extent, but it reserves blocks for the
whole count it is given, in an unsigned int. A count of about 2^32
blocks or more wraps the reservation. If it wraps to a few blocks, the
allocation uses more blocks than the transaction reserved, and
xfs_trans_mod_sb() shuts the filesystem down. If it wraps to nearly 2^32
blocks, the allocation fails with -ENOSPC.
Direct I/O, DAX and buffered writes with an extent size hint never ask
for that much, as xfs_direct_write_iomap_begin() limits them to
1024 pages to keep the count below 32 bits. xfs_fs_map_blocks() has no
such limit: it asks for everything from the start of a hole to the end
of the range of a pNFS layout, which a client chooses, and an RW
LAYOUTGET at offset 0 of an empty file for 16 TiB does it.
Fail such a count with -ENOSPC before anything is reserved, also on a
filesystem with that much free space, as the transaction cannot take
it. A count that fits is reserved in full, as before, so a hole longer
than the free space still fails before it is allocated.
With pynfs as the client and a 32 GiB XFS, an RW LAYOUTGET at offset 0
of an empty file for 16 TiB shut the filesystem down: "Corruption of
in-memory data (0x8) detected at xfs_trans_mod_sb". With this patch it
gets NFS4ERR_NOSPC and nothing is allocated, as with 100 GiB, and one
for 16 GiB, which fits, is still granted, in three extents.
Fixes: 527851124d10 ("xfs: implement pNFS export operations")
Cc: stable@vger.kernel.org
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_iomap.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 7c6238fed6..1917e49166 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -287,6 +287,13 @@ xfs_iomap_write_direct(
resaligned = xfs_aligned_fsb_count(offset_fsb, count_fsb,
xfs_get_extsz_hint(ip));
+ /*
+ * The transaction takes the block reservation as an unsigned int.
+ * Refuse a count that does not fit rather than reserve too little;
+ * only a pNFS layout for a huge range asks for that much.
+ */
+ if (resaligned > UINT_MAX - XFS_DIOSTRAT_SPACE_RES(mp, 0))
+ return -ENOSPC;
if (unlikely(XFS_IS_REALTIME_INODE(ip))) {
dblocks = XFS_DIOSTRAT_SPACE_RES(mp, 0);
rblocks = resaligned;
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap
2026-10-08 1:39 ` [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap Daejun Park via B4 Relay
@ 2026-10-08 15:40 ` Darrick J. Wong
0 siblings, 0 replies; 13+ messages in thread
From: Darrick J. Wong @ 2026-10-08 15:40 UTC (permalink / raw)
To: daejun7.park
Cc: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein,
Dave Chinner, Sergey Bashirov, Christian Brauner, linux-nfs,
linux-xfs, linux-fsdevel, linux-kernel, stable
On Thu, Oct 08, 2026 at 10:39:31AM +0900, Daejun Park via B4 Relay wrote:
> From: Daejun Park <daejun7.park@samsung.com>
>
> xfs_iomap_write_direct() maps one extent, but it reserves blocks for the
> whole count it is given, in an unsigned int. A count of about 2^32
> blocks or more wraps the reservation. If it wraps to a few blocks, the
> allocation uses more blocks than the transaction reserved, and
> xfs_trans_mod_sb() shuts the filesystem down. If it wraps to nearly 2^32
> blocks, the allocation fails with -ENOSPC.
>
> Direct I/O, DAX and buffered writes with an extent size hint never ask
> for that much, as xfs_direct_write_iomap_begin() limits them to
> 1024 pages to keep the count below 32 bits. xfs_fs_map_blocks() has no
> such limit: it asks for everything from the start of a hole to the end
> of the range of a pNFS layout, which a client chooses, and an RW
> LAYOUTGET at offset 0 of an empty file for 16 TiB does it.
>
> Fail such a count with -ENOSPC before anything is reserved, also on a
> filesystem with that much free space, as the transaction cannot take
> it. A count that fits is reserved in full, as before, so a hole longer
> than the free space still fails before it is allocated.
>
> With pynfs as the client and a 32 GiB XFS, an RW LAYOUTGET at offset 0
> of an empty file for 16 TiB shut the filesystem down: "Corruption of
> in-memory data (0x8) detected at xfs_trans_mod_sb". With this patch it
> gets NFS4ERR_NOSPC and nothing is allocated, as with 100 GiB, and one
> for 16 GiB, which fits, is still granted, in three extents.
>
> Fixes: 527851124d10 ("xfs: implement pNFS export operations")
> Cc: stable@vger.kernel.org
> Signed-off-by: Daejun Park <daejun7.park@samsung.com>
> ---
> fs/xfs/xfs_iomap.c | 7 +++++++
> 1 file changed, 7 insertions(+)
>
> diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
> index 7c6238fed6..1917e49166 100644
> --- a/fs/xfs/xfs_iomap.c
> +++ b/fs/xfs/xfs_iomap.c
> @@ -287,6 +287,13 @@ xfs_iomap_write_direct(
>
> resaligned = xfs_aligned_fsb_count(offset_fsb, count_fsb,
> xfs_get_extsz_hint(ip));
> + /*
> + * The transaction takes the block reservation as an unsigned int.
> + * Refuse a count that does not fit rather than reserve too little;
> + * only a pNFS layout for a huge range asks for that much.
> + */
> + if (resaligned > UINT_MAX - XFS_DIOSTRAT_SPACE_RES(mp, 0))
> + return -ENOSPC;
Uh... seeing as extent records can only map 2^21 blocks maximum and
iomap/pnfs can handle xfs returning a mapping of whatever length we
want, why don't we constrain count_fsb to XFS_BMBT_MAX_EXTLEN?
--D
> if (unlikely(XFS_IS_REALTIME_INODE(ip))) {
> dblocks = XFS_DIOSTRAT_SPACE_RES(mp, 0);
> rblocks = resaligned;
>
> --
> 2.43.0
>
>
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH 3/8] xfs: clamp the pNFS layout range to the maximum file size
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 1/8] xfs: map pNFS layouts to the end of the extent again Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 15:41 ` Darrick J. Wong
2026-10-08 1:39 ` [PATCH 4/8] exportfs: let ->map_blocks return more than one mapping Daejun Park via B4 Relay
` (4 subsequent siblings)
7 siblings, 1 reply; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, stable, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
xfs_fs_map_blocks() checks the range of a layout with
if (offset > limit)
goto out_unlock;
if (offset > limit - length)
length = limit - offset;
limit - length is computed in u64, so a length larger than the limit
wraps and the length is not clamped. nfsd passes the LAYOUTGET range
through, and RFC 8881 lets a client ask for NFS4_UINT64_MAX bytes, for
the rest of the file. offset + length then wraps to offset - 1, end_fsb
rounds back to offset_fsb, and xfs_bmapi_read() maps nothing. For a read
layout, xfs_bmbt_to_iomap() then gets an uninitialized imap: with a
zeroed stack it logs "Access to block zero", marks the data fork sick
and fails with -EFSCORRUPTED, and otherwise the mapping is whatever was
on the stack and can reach the client. For a write layout,
xfs_iomap_write_direct() is called with a count of zero, which trips an
ASSERT with XFS_DEBUG. An offset of 2^63 or more, which nfsd accepts
too, becomes negative in the loff_t argument and passes the first check.
Reject a negative offset and an offset at the limit, and compare the
length with what is left up to the limit, which cannot wrap once the
offset is below it.
With pynfs as the client, a READ LAYOUTGET for NFS4_UINT64_MAX bytes at
offset 0 got NFS4ERR_IO, as did an RW one, which also made nfsd warn
about the non-standard errno -63, and a READ LAYOUTGET at offset 2^63
got NFS4ERR_IO as well. With this patch and the previous one, the one at
2^63 gets NFS4ERR_INVAL, the RW one gets NFS4ERR_NOSPC before anything
is allocated, as the range up to the maximum file size cannot be
reserved, and the server logs nothing. The READ one gets NFS4ERR_INVAL
as well: nfsd, which maps a layout one extent at a time since
commit cc6c40e09d7b ("NFSD/blocklayout: Support multiple extents per
LAYOUTGET"), asks again from the maximum file size on. An RW one on a
file whose first 64 KiB are allocated, with a loga_maxcount that has
room for one extent, is granted 0+65536.
This depends on the previous patch. Without it, an RW LAYOUTGET to the
end of the file, whose length is now clamped instead of wrapping, asks
xfs_iomap_write_direct() for every block up to the maximum file size,
which wraps the block reservation and shuts the filesystem down where it
got NFS4ERR_IO before.
Fixes: 527851124d10 ("xfs: implement pNFS export operations")
Cc: stable@vger.kernel.org
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_pnfs.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index ab38561700..88d9c3c043 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -167,9 +167,9 @@ xfs_fs_map_blocks(
if (!write)
limit = max(limit, round_up(i_size_read(inode),
inode->i_sb->s_blocksize));
- if (offset > limit)
+ if (offset < 0 || offset >= limit)
goto out_unlock;
- if (offset > limit - length)
+ if (length > limit - offset)
length = limit - offset;
error = filemap_write_and_wait(inode->i_mapping);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH 3/8] xfs: clamp the pNFS layout range to the maximum file size
2026-10-08 1:39 ` [PATCH 3/8] xfs: clamp the pNFS layout range to the maximum file size Daejun Park via B4 Relay
@ 2026-10-08 15:41 ` Darrick J. Wong
0 siblings, 0 replies; 13+ messages in thread
From: Darrick J. Wong @ 2026-10-08 15:41 UTC (permalink / raw)
To: daejun7.park
Cc: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein,
Dave Chinner, Sergey Bashirov, Christian Brauner, linux-nfs,
linux-xfs, linux-fsdevel, linux-kernel, stable
On Thu, Oct 08, 2026 at 10:39:32AM +0900, Daejun Park via B4 Relay wrote:
> From: Daejun Park <daejun7.park@samsung.com>
>
> xfs_fs_map_blocks() checks the range of a layout with
>
> if (offset > limit)
> goto out_unlock;
> if (offset > limit - length)
> length = limit - offset;
>
> limit - length is computed in u64, so a length larger than the limit
> wraps and the length is not clamped. nfsd passes the LAYOUTGET range
> through, and RFC 8881 lets a client ask for NFS4_UINT64_MAX bytes, for
> the rest of the file. offset + length then wraps to offset - 1, end_fsb
> rounds back to offset_fsb, and xfs_bmapi_read() maps nothing. For a read
> layout, xfs_bmbt_to_iomap() then gets an uninitialized imap: with a
> zeroed stack it logs "Access to block zero", marks the data fork sick
> and fails with -EFSCORRUPTED, and otherwise the mapping is whatever was
> on the stack and can reach the client. For a write layout,
> xfs_iomap_write_direct() is called with a count of zero, which trips an
> ASSERT with XFS_DEBUG. An offset of 2^63 or more, which nfsd accepts
> too, becomes negative in the loff_t argument and passes the first check.
>
> Reject a negative offset and an offset at the limit, and compare the
> length with what is left up to the limit, which cannot wrap once the
> offset is below it.
>
> With pynfs as the client, a READ LAYOUTGET for NFS4_UINT64_MAX bytes at
> offset 0 got NFS4ERR_IO, as did an RW one, which also made nfsd warn
> about the non-standard errno -63, and a READ LAYOUTGET at offset 2^63
> got NFS4ERR_IO as well. With this patch and the previous one, the one at
> 2^63 gets NFS4ERR_INVAL, the RW one gets NFS4ERR_NOSPC before anything
> is allocated, as the range up to the maximum file size cannot be
> reserved, and the server logs nothing. The READ one gets NFS4ERR_INVAL
> as well: nfsd, which maps a layout one extent at a time since
> commit cc6c40e09d7b ("NFSD/blocklayout: Support multiple extents per
> LAYOUTGET"), asks again from the maximum file size on. An RW one on a
> file whose first 64 KiB are allocated, with a loga_maxcount that has
> room for one extent, is granted 0+65536.
>
> This depends on the previous patch. Without it, an RW LAYOUTGET to the
> end of the file, whose length is now clamped instead of wrapping, asks
> xfs_iomap_write_direct() for every block up to the maximum file size,
> which wraps the block reservation and shuts the filesystem down where it
> got NFS4ERR_IO before.
>
> Fixes: 527851124d10 ("xfs: implement pNFS export operations")
> Cc: stable@vger.kernel.org
> Signed-off-by: Daejun Park <daejun7.park@samsung.com>
Looks fine to me
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> fs/xfs/xfs_pnfs.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
> index ab38561700..88d9c3c043 100644
> --- a/fs/xfs/xfs_pnfs.c
> +++ b/fs/xfs/xfs_pnfs.c
> @@ -167,9 +167,9 @@ xfs_fs_map_blocks(
> if (!write)
> limit = max(limit, round_up(i_size_read(inode),
> inode->i_sb->s_blocksize));
> - if (offset > limit)
> + if (offset < 0 || offset >= limit)
> goto out_unlock;
> - if (offset > limit - length)
> + if (length > limit - offset)
> length = limit - offset;
>
> error = filemap_write_and_wait(inode->i_mapping);
>
> --
> 2.43.0
>
>
>
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH 4/8] exportfs: let ->map_blocks return more than one mapping
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
` (2 preceding siblings ...)
2026-10-08 1:39 ` [PATCH 3/8] xfs: clamp the pNFS layout range to the maximum file size Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 5/8] xfs: factor the mapping of one pNFS extent out of xfs_fs_map_blocks() Daejun Park via B4 Relay
` (3 subsequent siblings)
7 siblings, 0 replies; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
Since commit cc6c40e09d7b ("NFSD/blocklayout: Support multiple extents
per LAYOUTGET"), nfsd calls ->map_blocks once per extent of a block
layout, each time for the range left after the extent before. In XFS,
each call locks the inode, flushes and invalidates its page cache, maps
one extent and unlocks the inode again, and the mapping of the file can
change between the calls. When XFS still mapped whole extents, that
gave layouts whose extents overlapped, see commit 36ca6f11424a ("xfs:
fix overlapping extents returned for pNFS LAYOUTGET"), and nfsd now
trims the start of each extent after the first ("nfsd: do not return
overlapping extents in a block layout").
Let ->map_blocks fill an array of mappings, so that a filesystem can map
the whole range of a layout in one go, under one lock, and make the
mappings fit together: the first one contains the offset asked for, and
each further one starts where the one before it ends. Document as well
that the offset and the length come from the client unchecked, and what
the array holds on error. Only the interface changes here: XFS fills one
mapping, and nfsd asks for one at a time.
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Suggested-by: Christoph Hellwig <hch@lst.de>
Link: https://lore.kernel.org/r/20260519145949.GH9555@frogsfrogsfrogs
Link: https://lore.kernel.org/r/20261007131133.GB30647@lst.de
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/nfsd/blocklayout.c | 4 +++-
fs/xfs/xfs_pnfs.c | 6 ++++--
include/linux/exportfs_block.h | 14 ++++++++++----
3 files changed, 17 insertions(+), 7 deletions(-)
diff --git a/fs/nfsd/blocklayout.c b/fs/nfsd/blocklayout.c
index 1aabae0033..cedaf0e15d 100644
--- a/fs/nfsd/blocklayout.c
+++ b/fs/nfsd/blocklayout.c
@@ -30,11 +30,13 @@ nfsd4_block_map_extent(struct inode *inode, const struct svc_fh *fhp,
{
struct super_block *sb = inode->i_sb;
struct iomap iomap;
+ unsigned int nr_iomaps = 1;
u32 device_generation = 0;
int error;
error = sb->s_export_op->block_ops->map_blocks(inode, offset, length,
- &iomap, iomode != IOMODE_READ, &device_generation);
+ &iomap, &nr_iomaps, iomode != IOMODE_READ,
+ &device_generation);
if (error) {
if (error == -ENXIO)
return nfserr_layoutunavailable;
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index 88d9c3c043..055e6fb299 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -121,7 +121,8 @@ xfs_fs_map_blocks(
struct inode *inode,
loff_t offset,
u64 length,
- struct iomap *iomap,
+ struct iomap *iomaps,
+ unsigned int *nr_iomaps,
bool write,
u32 *device_generation)
{
@@ -238,7 +239,8 @@ xfs_fs_map_blocks(
}
xfs_iunlock(ip, XFS_IOLOCK_EXCL);
- error = xfs_bmbt_to_iomap(ip, iomap, &imap, 0, 0, seq);
+ error = xfs_bmbt_to_iomap(ip, iomaps, &imap, 0, 0, seq);
+ *nr_iomaps = 1;
*device_generation = mp->m_generation;
return error;
out_unlock:
diff --git a/include/linux/exportfs_block.h b/include/linux/exportfs_block.h
index 21e61fc012..1d0b0abbd3 100644
--- a/include/linux/exportfs_block.h
+++ b/include/linux/exportfs_block.h
@@ -44,12 +44,18 @@ struct exportfs_block_ops {
/*
* Map blocks for direct block access.
* If @write is %true, also allocate the blocks for the range if needed.
- * The mapping returned must contain @offset. It may start before
- * @offset and may end before or after @offset + @len.
+ * @offset and @len come from the client: @offset may be negative or at
+ * or past the maximum file size, and @offset + @len may overflow.
+ * Fill in at most *@nr_iomaps mappings, which is at least one, and on
+ * success set *@nr_iomaps to the number filled in, at least one. On
+ * error, the mappings and *@nr_iomaps are undefined. The first mapping
+ * must contain @offset and may start before it. Each further mapping
+ * must start where the one before it ends. The last one may end before
+ * or after @offset + @len.
*/
int (*map_blocks)(struct inode *inode, loff_t offset, u64 len,
- struct iomap *iomap, bool write,
- u32 *device_generation);
+ struct iomap *iomaps, unsigned int *nr_iomaps,
+ bool write, u32 *device_generation);
/*
* Commit blocks previously handed out by ->map_blocks and written to by
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH 5/8] xfs: factor the mapping of one pNFS extent out of xfs_fs_map_blocks()
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
` (3 preceding siblings ...)
2026-10-08 1:39 ` [PATCH 4/8] exportfs: let ->map_blocks return more than one mapping Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 6/8] xfs: take the invalidate lock while mapping a pNFS layout Daejun Park via B4 Relay
` (2 subsequent siblings)
7 siblings, 0 replies; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
Move the part of xfs_fs_map_blocks() that maps one extent, and allocates
it first if it is a hole in a write layout, into a new helper,
xfs_fs_map_extent(), so that the next patch can call it for each extent
of a layout. The inode update and the log force after an allocation stay
in xfs_fs_map_blocks().
No change in behavior.
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_pnfs.c | 119 +++++++++++++++++++++++++++++++++---------------------
1 file changed, 72 insertions(+), 47 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index 055e6fb299..81e1ae0351 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -113,6 +113,73 @@ xfs_fs_map_update_inode(
return xfs_trans_commit(tp);
}
+/*
+ * Map the extent of a pNFS layout that starts at @offset_fsb, allocating it
+ * first if it is a hole in a write layout. The mapping may end before or
+ * after @end. Called with the iolock held.
+ */
+static int
+xfs_fs_map_extent(
+ struct xfs_inode *ip,
+ xfs_fileoff_t offset_fsb,
+ loff_t end,
+ bool write,
+ struct xfs_bmbt_irec *imap,
+ u64 *seq,
+ bool *allocated)
+{
+ struct xfs_mount *mp = ip->i_mount;
+ xfs_fileoff_t end_fsb = XFS_B_TO_FSB(mp, (xfs_ufsize_t)end);
+ int nimaps = 1;
+ uint lock_flags;
+ int error;
+
+ lock_flags = xfs_ilock_data_map_shared(ip);
+ /*
+ * Map to the end of the extent that covers the start of the range,
+ * so that a client doing I/O in pieces gets a layout it can use for
+ * the pieces that follow. Never map anything before the start of
+ * the range: nfsd calls in here once per extent of a LAYOUTGET, for
+ * the range that is left after the previous extent, and the mapping
+ * can change in between, so a mapping that reaches back can overlap
+ * one already in the layout. Don't extend the mapping past EOF
+ * beyond the range either: xfs_free_eofblocks() can free blocks past
+ * EOF without breaking the layout.
+ */
+ error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb,
+ imap, &nimaps, XFS_BMAPI_ENTIRE);
+ if (error) {
+ xfs_iunlock(ip, lock_flags);
+ return error;
+ }
+ if (nimaps)
+ xfs_trim_extent(imap, offset_fsb,
+ max_t(xfs_fileoff_t, end_fsb,
+ XFS_B_TO_FSB(mp, XFS_ISIZE(ip))) -
+ offset_fsb);
+ *seq = xfs_iomap_inode_sequence(ip, 0);
+
+ ASSERT(!nimaps || imap->br_startblock != DELAYSTARTBLOCK);
+
+ if (!write || (nimaps && imap->br_startblock != HOLESTARTBLOCK)) {
+ xfs_iunlock(ip, lock_flags);
+ return 0;
+ }
+
+ if (end > XFS_ISIZE(ip))
+ end_fsb = xfs_iomap_eof_align_last_fsb(ip, end_fsb);
+ else if (nimaps)
+ end_fsb = min(end_fsb, imap->br_startoff + imap->br_blockcount);
+ xfs_iunlock(ip, lock_flags);
+
+ error = xfs_iomap_write_direct(ip, offset_fsb, end_fsb - offset_fsb, 0,
+ imap, seq);
+ if (error)
+ return error;
+ *allocated = true;
+ return 0;
+}
+
/*
* Get a layout for the pNFS client.
*/
@@ -129,10 +196,8 @@ xfs_fs_map_blocks(
struct xfs_inode *ip = XFS_I(inode);
struct xfs_mount *mp = ip->i_mount;
struct xfs_bmbt_irec imap;
- xfs_fileoff_t offset_fsb, end_fsb;
+ bool allocated = false;
loff_t limit;
- int nimaps = 1;
- uint lock_flags;
int error = 0;
u64 seq;
@@ -180,49 +245,12 @@ xfs_fs_map_blocks(
if (WARN_ON_ONCE(error))
goto out_unlock;
- end_fsb = XFS_B_TO_FSB(mp, (xfs_ufsize_t)offset + length);
- offset_fsb = XFS_B_TO_FSBT(mp, offset);
-
- lock_flags = xfs_ilock_data_map_shared(ip);
- /*
- * Map to the end of the extent that covers the start of the range,
- * so that a client doing I/O in pieces gets a layout it can use for
- * the pieces that follow. Never map anything before the start of
- * the range: nfsd calls in here once per extent of a LAYOUTGET, for
- * the range that is left after the previous extent, and the mapping
- * can change in between, so a mapping that reaches back can overlap
- * one already in the layout. Don't extend the mapping past EOF
- * beyond the range either: xfs_free_eofblocks() can free blocks past
- * EOF without breaking the layout.
- */
- error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb,
- &imap, &nimaps, XFS_BMAPI_ENTIRE);
- if (error) {
- xfs_iunlock(ip, lock_flags);
+ error = xfs_fs_map_extent(ip, XFS_B_TO_FSBT(mp, offset),
+ offset + length, write, &imap, &seq, &allocated);
+ if (error)
goto out_unlock;
- }
- if (nimaps)
- xfs_trim_extent(&imap, offset_fsb,
- max_t(xfs_fileoff_t, end_fsb,
- XFS_B_TO_FSB(mp, XFS_ISIZE(ip))) -
- offset_fsb);
- seq = xfs_iomap_inode_sequence(ip, 0);
-
- ASSERT(!nimaps || imap.br_startblock != DELAYSTARTBLOCK);
-
- if (write && (!nimaps || imap.br_startblock == HOLESTARTBLOCK)) {
- if (offset + length > XFS_ISIZE(ip))
- end_fsb = xfs_iomap_eof_align_last_fsb(ip, end_fsb);
- else if (nimaps && imap.br_startblock == HOLESTARTBLOCK)
- end_fsb = min(end_fsb, imap.br_startoff +
- imap.br_blockcount);
- xfs_iunlock(ip, lock_flags);
-
- error = xfs_iomap_write_direct(ip, offset_fsb,
- end_fsb - offset_fsb, 0, &imap, &seq);
- if (error)
- goto out_unlock;
+ if (allocated) {
/*
* Ensure the next transaction is committed synchronously so
* that the blocks allocated and handed out to the client are
@@ -233,9 +261,6 @@ xfs_fs_map_blocks(
error = xfs_log_force_inode(ip);
if (error)
goto out_unlock;
-
- } else {
- xfs_iunlock(ip, lock_flags);
}
xfs_iunlock(ip, XFS_IOLOCK_EXCL);
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH 6/8] xfs: take the invalidate lock while mapping a pNFS layout
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
` (4 preceding siblings ...)
2026-10-08 1:39 ` [PATCH 5/8] xfs: factor the mapping of one pNFS extent out of xfs_fs_map_blocks() Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 7/8] xfs: map the whole range of a pNFS layout in one ->map_blocks call Daejun Park via B4 Relay
2026-10-08 1:39 ` [PATCH 8/8] nfsd: get all extents of a block " Daejun Park via B4 Relay
7 siblings, 0 replies; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
xfs_fs_map_blocks() flushes and invalidates the page cache of the file
under the iolock and then maps the range, but a write fault takes only
the invalidate lock (XFS_MMAPLOCK_SHARED in __xfs_write_fault()), not
the iolock. A fault on a shared mapping of the file can therefore add a
delalloc extent to the range between the flush and the mapping, which
trips the ASSERT on DELAYSTARTBLOCK, or is handed to nfsd, which warns
about a filesystem that returned a delalloc extent and refuses the
layout.
Take the invalidate lock in exclusive mode together with the iolock, as
__xfs_file_fallocate() does, so that page faults wait until the layout
has been mapped. The next patch maps all extents of a layout under these
locks, which makes the window between the flush and the last mapping
wider.
Not marked for stable: the race needs a process on the server that
writes to a shared mapping of the file while a client gets a layout for
the same range, and the comment above xfs_break_leased_layouts() already
leaves such writers unsynchronized with pNFS clients.
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Link: https://lore.kernel.org/r/20260519145949.GH9555@frogsfrogsfrogs
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_pnfs.c | 11 +++++++----
1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index 81e1ae0351..bd6f7d5093 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -116,7 +116,7 @@ xfs_fs_map_update_inode(
/*
* Map the extent of a pNFS layout that starts at @offset_fsb, allocating it
* first if it is a hole in a write layout. The mapping may end before or
- * after @end. Called with the iolock held.
+ * after @end. Called with the iolock and the invalidate lock held.
*/
static int
xfs_fs_map_extent(
@@ -225,8 +225,11 @@ xfs_fs_map_blocks(
* similar to direct I/O, except that the synchronization is much more
* complicated. See the comment near xfs_break_leased_layouts
* for a detailed explanation.
+ *
+ * Take the invalidate lock as well, so that page faults cannot add
+ * delalloc extents to the range while it is being mapped.
*/
- xfs_ilock(ip, XFS_IOLOCK_EXCL);
+ xfs_ilock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
error = -EINVAL;
limit = mp->m_super->s_maxbytes;
@@ -262,14 +265,14 @@ xfs_fs_map_blocks(
if (error)
goto out_unlock;
}
- xfs_iunlock(ip, XFS_IOLOCK_EXCL);
+ xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
error = xfs_bmbt_to_iomap(ip, iomaps, &imap, 0, 0, seq);
*nr_iomaps = 1;
*device_generation = mp->m_generation;
return error;
out_unlock:
- xfs_iunlock(ip, XFS_IOLOCK_EXCL);
+ xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
return error;
}
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH 7/8] xfs: map the whole range of a pNFS layout in one ->map_blocks call
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
` (5 preceding siblings ...)
2026-10-08 1:39 ` [PATCH 6/8] xfs: take the invalidate lock while mapping a pNFS layout Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 16:37 ` Chuck Lever
2026-10-08 1:39 ` [PATCH 8/8] nfsd: get all extents of a block " Daejun Park via B4 Relay
7 siblings, 1 reply; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
xfs_fs_map_blocks() maps one extent per call, and nfsd calls it again
for the rest of the range. Each call takes the iolock and the
invalidate lock, flushes and invalidates the page cache, and, if it
allocated blocks for a write layout, updates the inode and forces the
log.
Map extent after extent until the range is covered or the array of
mappings is full, all under one hold of the iolock and the invalidate
lock, with one flush and invalidation and at most one inode update and
log force. Each mapping starts where the one before it ends.
xfs_fs_map_extent() never maps anything before the offset it is given,
which matters here: allocating a hole merges it with the extents around
it, and a mapping of the next offset would otherwise reach back over the
hole.
If an allocation fails after others in the same call, the error is
returned without the inode update and the log force. nfsd then fails the
LAYOUTGET, so none of the blocks allocated is handed out; they stay
unwritten, and blocks past EOF can be freed again unless the inode
already has XFS_DIFLAG_PREALLOC or the file is empty.
With the nfsd change that follows, a LAYOUTGET for the first 256 KiB of
a file in which 32 blocks of 4 KiB written with O_DIRECT are each
followed by a 4 KiB hole gets 64 extents in one call. For an RW layout,
which allocates the 32 holes, the log is forced once instead of 32
times, and nfsd4_block_proc_layoutget() takes 2.2 ms instead of 166 ms,
most of which were the 32 synchronous log forces (median of 20
LAYOUTGETs each; QEMU VMs, XFS on an NVMe/TCP namespace, pynfs as the
client, time from a kprobe, KASAN and lockdep off).
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/xfs/xfs_pnfs.c | 47 +++++++++++++++++++++++++++++++----------------
1 file changed, 31 insertions(+), 16 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index bd6f7d5093..60c552c511 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -138,13 +138,12 @@ xfs_fs_map_extent(
/*
* Map to the end of the extent that covers the start of the range,
* so that a client doing I/O in pieces gets a layout it can use for
- * the pieces that follow. Never map anything before the start of
- * the range: nfsd calls in here once per extent of a LAYOUTGET, for
- * the range that is left after the previous extent, and the mapping
- * can change in between, so a mapping that reaches back can overlap
- * one already in the layout. Don't extend the mapping past EOF
- * beyond the range either: xfs_free_eofblocks() can free blocks past
- * EOF without breaking the layout.
+ * the pieces that follow. Never map anything before @offset_fsb:
+ * allocating the hole before it can have merged that hole with the
+ * extents around it, and a mapping that reaches back would overlap
+ * the previous one. Don't extend the mapping past EOF beyond the
+ * range either: xfs_free_eofblocks() can free blocks past EOF without
+ * breaking the layout.
*/
error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb,
imap, &nimaps, XFS_BMAPI_ENTIRE);
@@ -182,6 +181,10 @@ xfs_fs_map_extent(
/*
* Get a layout for the pNFS client.
+ *
+ * Map the range, or as much of it as fits into @iomaps, while holding the
+ * iolock and the invalidate lock, so that the mappings fit together: each
+ * one starts where the one before it ends.
*/
static int
xfs_fs_map_blocks(
@@ -195,14 +198,16 @@ xfs_fs_map_blocks(
{
struct xfs_inode *ip = XFS_I(inode);
struct xfs_mount *mp = ip->i_mount;
- struct xfs_bmbt_irec imap;
+ xfs_fileoff_t offset_fsb, end_fsb;
+ unsigned int nr = 0;
bool allocated = false;
loff_t limit;
int error = 0;
- u64 seq;
if (xfs_is_shutdown(mp))
return -EIO;
+ if (WARN_ON_ONCE(!*nr_iomaps))
+ return -EINVAL;
/*
* We can't export inodes residing on the realtime device. The realtime
@@ -248,10 +253,21 @@ xfs_fs_map_blocks(
if (WARN_ON_ONCE(error))
goto out_unlock;
- error = xfs_fs_map_extent(ip, XFS_B_TO_FSBT(mp, offset),
- offset + length, write, &imap, &seq, &allocated);
- if (error)
- goto out_unlock;
+ offset_fsb = XFS_B_TO_FSBT(mp, offset);
+ end_fsb = XFS_B_TO_FSB(mp, (xfs_ufsize_t)offset + length);
+ while (nr < *nr_iomaps && offset_fsb < end_fsb) {
+ struct xfs_bmbt_irec imap;
+ u64 seq;
+
+ error = xfs_fs_map_extent(ip, offset_fsb, offset + length,
+ write, &imap, &seq, &allocated);
+ if (error)
+ goto out_unlock;
+ error = xfs_bmbt_to_iomap(ip, &iomaps[nr++], &imap, 0, 0, seq);
+ if (error)
+ goto out_unlock;
+ offset_fsb = imap.br_startoff + imap.br_blockcount;
+ }
if (allocated) {
/*
@@ -267,10 +283,9 @@ xfs_fs_map_blocks(
}
xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
- error = xfs_bmbt_to_iomap(ip, iomaps, &imap, 0, 0, seq);
- *nr_iomaps = 1;
+ *nr_iomaps = nr;
*device_generation = mp->m_generation;
- return error;
+ return 0;
out_unlock:
xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
return error;
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH 7/8] xfs: map the whole range of a pNFS layout in one ->map_blocks call
2026-10-08 1:39 ` [PATCH 7/8] xfs: map the whole range of a pNFS layout in one ->map_blocks call Daejun Park via B4 Relay
@ 2026-10-08 16:37 ` Chuck Lever
0 siblings, 0 replies; 13+ messages in thread
From: Chuck Lever @ 2026-10-08 16:37 UTC (permalink / raw)
To: daejun7.park
Cc: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey,
Christoph Hellwig, Carlos Maiolino, Amir Goldstein,
Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel
Daejun Park <daejun7.park@samsung.com> wrote:
> If an allocation fails after others in the same call, the error is
> returned without the inode update and the log force. nfsd then fails the
> LAYOUTGET, so none of the blocks allocated is handed out; they stay
> unwritten, and blocks past EOF can be freed again unless the inode
> already has XFS_DIFLAG_PREALLOC or the file is empty.
The same freeing can happen to blocks a client holds a write layout
for, AFAICS. But that is not new. Commit 527851124d10 handed out the
same blocks.
This is however the patch where a fix could go.
> + while (nr < *nr_iomaps && offset_fsb < end_fsb) {
> + struct xfs_bmbt_irec imap;
> + u64 seq;
> +
> + error = xfs_fs_map_extent(ip, offset_fsb, offset + length,
> + write, &imap, &seq, &allocated);
> + if (error)
> + goto out_unlock;
> + error = xfs_bmbt_to_iomap(ip, &iomaps[nr++], &imap, 0, 0, seq);
> + if (error)
> + goto out_unlock;
> + offset_fsb = imap.br_startoff + imap.br_blockcount;
> + }
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH 8/8] nfsd: get all extents of a block layout in one ->map_blocks call
2026-10-08 1:39 [PATCH 0/8] xfs, nfsd: map a whole pNFS block layout in one ->map_blocks call Daejun Park via B4 Relay
` (6 preceding siblings ...)
2026-10-08 1:39 ` [PATCH 7/8] xfs: map the whole range of a pNFS layout in one ->map_blocks call Daejun Park via B4 Relay
@ 2026-10-08 1:39 ` Daejun Park via B4 Relay
2026-10-08 16:44 ` Chuck Lever
7 siblings, 1 reply; 13+ messages in thread
From: Daejun Park via B4 Relay @ 2026-10-08 1:39 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
Tom Talpey, Christoph Hellwig, Carlos Maiolino, Amir Goldstein
Cc: Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel, Daejun Park
From: Daejun Park <daejun7.park@samsung.com>
nfsd4_block_proc_layoutget() calls ->map_blocks once per extent, each
time for the range left after the extent before, and trims each extent
so that it starts where the one before it ended, as the mapping can
change between the calls.
Ask for all extents in one call, which ->map_blocks now allows, so that
the filesystem maps them together. The array of mappings is bounded by
the same loga_maxcount limit as the layout. Instead of trimming, check
that the first extent contains the offset asked for and that each one
after it starts where the one before it ends, and warn and answer
NFS4ERR_LAYOUTUNAVAILABLE if not, as nfsd already does for a first
extent that does not contain the offset: with the whole range mapped
under one lock, a mapping that does not fit is a filesystem bug, which
trimming would hide. nfsd also warns if the filesystem returns no
mapping, or more than it has room for. nfsd4_block_map_extent() becomes
nfsd4_block_iomap_to_extent(), which only converts a mapping, and the
device ID is set up once per layout instead of once per extent.
With XFS, a READ LAYOUTGET for the first 256 KiB of a file in which 32
blocks of 4 KiB are each followed by a 4 KiB hole calls ->map_blocks
once instead of 64 times, gets the same 64 extents, and
nfsd4_block_proc_layoutget() takes 9 us instead of 26 us (median of 20
LAYOUTGETs each; QEMU VMs, XFS on an NVMe/TCP namespace, pynfs as the
client, time from a kprobe, KASAN and lockdep off). A READ LAYOUTGET for
NFS4_UINT64_MAX bytes, which got NFS4ERR_INVAL when nfsd asked again
from the maximum file size on, now gets a layout to 2^63. An RW
LAYOUTGET with a loga_minlength of zero, which nfsd refuses once it
meets an unwritten extent, now has all holes in its range allocated
before it is refused, instead of only the first one; the Linux client
asks for at least one page.
pynfs BLOCK1, BLOCK2, BLOCK4 and BLOCK5 pass as before (five runs
before, three after, with KASAN and lockdep). On a test kernel whose XFS
does not trim a mapping to the offset asked for, BLOCK5 gets
NFS4ERR_LAYOUTUNAVAILABLE and nfsd warns that the filesystem "returned
extent 0+12288 for offset 8192", where nfsd used to trim that extent
instead. fstests generic/075, 091 and 263 over the block and the SCSI
layout give the same results as before: 075 fails with the fsx size
error that it also gets without this series, and bl_alloc_lseg() on the
client returns no error.
Signed-off-by: Daejun Park <daejun7.park@samsung.com>
---
fs/nfsd/blocklayout.c | 136 ++++++++++++++++++++++++++------------------------
1 file changed, 70 insertions(+), 66 deletions(-)
diff --git a/fs/nfsd/blocklayout.c b/fs/nfsd/blocklayout.c
index cedaf0e15d..e5e0a43ade 100644
--- a/fs/nfsd/blocklayout.c
+++ b/fs/nfsd/blocklayout.c
@@ -20,43 +20,19 @@
/*
- * Get an extent from the file system that contains offset. It may start
- * below offset and may be shorter than the requested length.
+ * Turn a mapping from the filesystem into an extent of the layout.
*/
static __be32
-nfsd4_block_map_extent(struct inode *inode, const struct svc_fh *fhp,
- u64 offset, u64 length, u32 iomode, u64 minlength,
- struct pnfs_block_extent *bex)
+nfsd4_block_iomap_to_extent(const struct iomap *iomap, u32 iomode,
+ u64 minlength, struct pnfs_block_extent *bex)
{
- struct super_block *sb = inode->i_sb;
- struct iomap iomap;
- unsigned int nr_iomaps = 1;
- u32 device_generation = 0;
- int error;
-
- error = sb->s_export_op->block_ops->map_blocks(inode, offset, length,
- &iomap, &nr_iomaps, iomode != IOMODE_READ,
- &device_generation);
- if (error) {
- if (error == -ENXIO)
- return nfserr_layoutunavailable;
- return nfserrno(error);
- }
-
- if (WARN_ONCE(iomap.offset > offset ||
- offset - iomap.offset >= iomap.length,
- "pnfsd: %s ino %llu: filesystem returned extent %lld+%llu for offset %llu\n",
- sb->s_id, inode->i_ino, iomap.offset, iomap.length,
- offset))
- return nfserr_layoutunavailable;
-
- switch (iomap.type) {
+ switch (iomap->type) {
case IOMAP_MAPPED:
if (iomode == IOMODE_READ)
bex->es = PNFS_BLOCK_READ_DATA;
else
bex->es = PNFS_BLOCK_READWRITE_DATA;
- bex->soff = iomap.addr;
+ bex->soff = iomap->addr;
break;
case IOMAP_UNWRITTEN:
if (iomode & IOMODE_RW) {
@@ -69,7 +45,7 @@ nfsd4_block_map_extent(struct inode *inode, const struct svc_fh *fhp,
}
bex->es = PNFS_BLOCK_INVALID_DATA;
- bex->soff = iomap.addr;
+ bex->soff = iomap->addr;
break;
}
fallthrough;
@@ -81,16 +57,12 @@ nfsd4_block_map_extent(struct inode *inode, const struct svc_fh *fhp,
fallthrough;
case IOMAP_DELALLOC:
default:
- WARN(1, "pnfsd: filesystem returned %d extent\n", iomap.type);
+ WARN(1, "pnfsd: filesystem returned %d extent\n", iomap->type);
return nfserr_layoutunavailable;
}
- error = nfsd4_set_deviceid(&bex->vol_id, fhp, device_generation);
- if (error)
- return nfserrno(error);
-
- bex->foff = iomap.offset;
- bex->len = iomap.length;
+ bex->foff = iomap->offset;
+ bex->len = iomap->length;
return nfs_ok;
}
@@ -99,11 +71,17 @@ nfsd4_block_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode,
const struct svc_fh *fhp, struct nfsd4_layoutget *args)
{
struct nfsd4_layout_seg *seg = &args->lg_seg;
+ struct super_block *sb = inode->i_sb;
struct pnfs_block_layout *bl;
struct pnfs_block_extent *first_bex, *last_bex;
+ struct nfsd4_deviceid vol_id;
+ struct iomap *iomaps = NULL;
u64 offset = seg->offset, length = seg->length;
u32 i, nr_extents_max, block_size = i_blocksize(inode);
+ u32 device_generation = 0;
+ unsigned int nr_iomaps;
__be32 nfserr;
+ int error;
if (locks_in_grace(SVC_NET(rqstp)))
return nfserr_grace;
@@ -146,42 +124,67 @@ nfsd4_block_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode,
goto out_error;
bl->nr_extents = nr_extents_max;
args->lg_content = bl;
+ iomaps = kmalloc_array(nr_extents_max, sizeof(*iomaps), GFP_KERNEL);
+ if (!iomaps)
+ goto out_error;
- for (i = 0; i < bl->nr_extents; i++) {
- struct pnfs_block_extent *bex = bl->extents + i;
- u64 bex_length;
+ /*
+ * Get all extents in one call, so that the filesystem maps them
+ * together and they fit.
+ */
+ nr_iomaps = nr_extents_max;
+ error = sb->s_export_op->block_ops->map_blocks(inode, offset, length,
+ iomaps, &nr_iomaps, seg->iomode != IOMODE_READ,
+ &device_generation);
+ if (error) {
+ if (error == -ENXIO)
+ nfserr = nfserr_layoutunavailable;
+ else
+ nfserr = nfserrno(error);
+ goto out_error;
+ }
- nfserr = nfsd4_block_map_extent(inode, fhp, offset, length,
- seg->iomode, args->lg_minlength, bex);
- if (nfserr != nfs_ok)
- goto out_error;
+ nfserr = nfserr_layoutunavailable;
+ if (WARN_ONCE(!nr_iomaps || nr_iomaps > nr_extents_max,
+ "pnfsd: %s ino %llu: filesystem returned %u extents for %u\n",
+ sb->s_id, inode->i_ino, nr_iomaps, nr_extents_max))
+ goto out_error;
- /*
- * Each extent after the first was mapped for the range that
- * starts where the previous extent ends, but the filesystem
- * may return a mapping that starts below that point. Trim
- * it, as RFC 5663 section 2.3.1 does not allow extents to
- * overlap. nfsd4_block_map_extent() made sure the mapping
- * contains offset. NONE_DATA extents have no volume offset.
- */
- if (i > 0 && bex->foff < offset) {
- u64 skip = offset - bex->foff;
+ error = nfsd4_set_deviceid(&vol_id, fhp, device_generation);
+ if (error) {
+ nfserr = nfserrno(error);
+ goto out_error;
+ }
- bex->foff = offset;
- bex->len -= skip;
- if (bex->es != PNFS_BLOCK_NONE_DATA)
- bex->soff += skip;
- }
+ for (i = 0; i < nr_iomaps; i++) {
+ const struct iomap *iomap = iomaps + i;
+ struct pnfs_block_extent *bex = bl->extents + i;
- bex_length = bex->len - (offset - bex->foff);
- if (bex_length >= length) {
- bl->nr_extents = i + 1;
- break;
- }
+ /*
+ * The first extent contains the offset asked for, and each
+ * extent after it starts where the one before it ends: RFC
+ * 5663 section 2.3.1 requires the extents to be logically
+ * contiguous.
+ */
+ nfserr = nfserr_layoutunavailable;
+ if (WARN_ONCE(iomap->offset > offset ||
+ offset - iomap->offset >= iomap->length ||
+ (i > 0 && iomap->offset != offset),
+ "pnfsd: %s ino %llu: filesystem returned extent %lld+%llu for offset %llu\n",
+ sb->s_id, inode->i_ino, iomap->offset,
+ iomap->length, offset))
+ goto out_error;
- offset = bex->foff + bex->len;
- length -= bex_length;
+ nfserr = nfsd4_block_iomap_to_extent(iomap, seg->iomode,
+ args->lg_minlength, bex);
+ if (nfserr != nfs_ok)
+ goto out_error;
+ bex->vol_id = vol_id;
+ offset = iomap->offset + iomap->length;
}
+ bl->nr_extents = nr_iomaps;
+ kfree(iomaps);
+ iomaps = NULL;
first_bex = bl->extents;
last_bex = bl->extents + bl->nr_extents - 1;
@@ -198,6 +201,7 @@ nfsd4_block_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode,
return nfs_ok;
out_error:
+ kfree(iomaps);
seg->length = 0;
return nfserr;
}
--
2.43.0
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH 8/8] nfsd: get all extents of a block layout in one ->map_blocks call
2026-10-08 1:39 ` [PATCH 8/8] nfsd: get all extents of a block " Daejun Park via B4 Relay
@ 2026-10-08 16:44 ` Chuck Lever
0 siblings, 0 replies; 13+ messages in thread
From: Chuck Lever @ 2026-10-08 16:44 UTC (permalink / raw)
To: daejun7.park
Cc: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey,
Christoph Hellwig, Carlos Maiolino, Amir Goldstein,
Darrick J. Wong, Dave Chinner, Sergey Bashirov,
Christian Brauner, linux-nfs, linux-xfs, linux-fsdevel,
linux-kernel
Daejun Park <daejun7.park@samsung.com> wrote:
> @@ -146,42 +124,67 @@ nfsd4_block_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode,
> goto out_error;
> bl->nr_extents = nr_extents_max;
> args->lg_content = bl;
> + iomaps = kmalloc_array(nr_extents_max, sizeof(*iomaps), GFP_KERNEL);
> + if (!iomaps)
> + goto out_error;
nr_extents_max is bounded by PAGE_SIZE so that the per-request buffer
stays small. But a struct iomap is nearly twice the size of an encoded
extent, so this array is nearly twice the layout buffer it sits next
to. On 64 KiB pages that is well over 100 KiB of physically contiguous
memory per LAYOUTGET. A failure here turns into NFS4ERR_DELAY for
every request under fragmentation.
Using kvmalloc_array() would cover the large case, or you could cap
nr_extents_max independently of PAGE_SIZE.
> + /*
> + * Get all extents in one call, so that the filesystem maps them
> + * together and they fit.
> + */
> + nr_iomaps = nr_extents_max;
> + error = sb->s_export_op->block_ops->map_blocks(inode, offset, length,
> + iomaps, &nr_iomaps, seg->iomode != IOMODE_READ,
> + &device_generation);
An RW LAYOUTGET with loga_minlength of 0 is refused by
nfsd4_block_iomap_to_extent() on the first unwritten extent, so
currently no write mapping is ever handed out for it. Yet this call
still passes write, so XFS allocates every hole in the range, updates
the inode and forces the log before the result is thrown away. With
one call per request this looks like up to nr_extents_max allocation
transactions instead of one.
Passing write only when loga_minlength is non-zero, and having
nfsd4_block_iomap_to_extent() refuse IOMAP_HOLE for IOMODE_RW the way
it already refuses IOMAP_UNWRITTEN, gives the client the same answer
without allocating anything. Admittedly this is a pre-existing
issue -- perhaps a pre-requisite fix for this series is best, so
that the fix can be backported to LTS.
A note about process: the series touches XFS and exportfs. It will
need to go through the VFS tree, I think, and might conflict with
what's already in nfsd-testing.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 13+ messages in thread