mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding
@ 2026-10-08 22:28 Lucheng Bao
  2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao
  2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao
  0 siblings, 2 replies; 4+ messages in thread
From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw)
  To: linux-nfs, Trond Myklebust, Anna Schumaker
  Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao

When a layout requires LAYOUTCOMMIT, the size reported by the MDS may stay
stale until the commit is processed. The client already keeps i_size
from shrinking while writes are outstanding. This series extends that
protection to an outstanding LAYOUTCOMMIT.

Patch 1 rejects smaller sizes while a LAYOUTCOMMIT is pending or in
flight.

Patch 2 defers a LAYOUTRETURN whose OLD_STATEID refresh fails because a
newer segment is busy. Today that return is treated as complete and the
pending LAYOUTCOMMIT is discarded.

Without this series, a stale OPEN or CLOSE reply can lower i_size, and
the next O_APPEND write then lands at the old offset.

Lucheng Bao (2):
  NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding
  NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh

 fs/nfs/inode.c    | 8 ++++++--
 fs/nfs/nfs4proc.c | 3 ++-
 2 files changed, 8 insertions(+), 3 deletions(-)

-- 
2.34.1


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding
  2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao
@ 2026-10-08 22:28 ` Lucheng Bao
  2026-10-09  4:43   ` Trond Myklebust
  2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao
  1 sibling, 1 reply; 4+ messages in thread
From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw)
  To: linux-nfs, Trond Myklebust, Anna Schumaker
  Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao, stable

Commit ac46bd374c9a ("pNFS: Ensure we layoutcommit before revalidating
attributes") replaced outstanding-layoutcommit attribute filtering with
synchronization before explicit revalidation. For layouts requiring
LAYOUTCOMMIT, OPEN attributes can still shrink i_size before metadata
synchronization completes.

Commit d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS
LAYOUTCOMMIT is outstanding") moved LAYOUTCOMMIT post-op attribute
processing after cleanup, leaving another unprotected window after
NFS_INO_LAYOUTCOMMITTING is cleared.

Extend the existing writeback checks to reject smaller sizes while a
LAYOUTCOMMIT is pending or in flight, preserving size increases and WCC
checks. Process the reply's attributes and establish their generation
barrier before cleanup removes that protection.

Fixes: d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS LAYOUTCOMMIT is outstanding")
Fixes: ac46bd374c9a ("pNFS: Ensure we layoutcommit before revalidating attributes")
Cc: stable@vger.kernel.org
Signed-off-by: Lucheng Bao <lubao@everpuredata.com>
---
 fs/nfs/inode.c    | 8 ++++++--
 fs/nfs/nfs4proc.c | 2 +-
 2 files changed, 7 insertions(+), 3 deletions(-)

diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c
index 3022454f7698..11432f521b04 100644
--- a/fs/nfs/inode.c
+++ b/fs/nfs/inode.c
@@ -1636,7 +1636,9 @@ static void nfs_wcc_update_inode(struct inode *inode, struct nfs_fattr *fattr)
 	if ((fattr->valid & NFS_ATTR_FATTR_PRESIZE)
 			&& (fattr->valid & NFS_ATTR_FATTR_SIZE)
 			&& i_size_read(inode) == nfs_size_to_loff_t(fattr->pre_size)
-			&& !nfs_have_writebacks(inode)) {
+			&& !nfs_have_writebacks(inode)
+			&& (nfs_size_to_loff_t(fattr->size) >= i_size_read(inode)
+				|| !pnfs_layoutcommit_outstanding(inode))) {
 		trace_nfs_size_wcc(inode, fattr->size);
 		i_size_write(inode, nfs_size_to_loff_t(fattr->size));
 	}
@@ -2392,7 +2394,9 @@ static int nfs_update_inode(struct inode *inode, struct nfs_fattr *fattr)
 		if (new_isize != cur_isize && !have_delegation) {
 			/* Do we perhaps have any outstanding writes, or has
 			 * the file grown beyond our last write? */
-			if (!nfs_have_writebacks(inode) || new_isize > cur_isize) {
+			if ((!nfs_have_writebacks(inode) &&
+			     !pnfs_layoutcommit_outstanding(inode)) ||
+			    new_isize > cur_isize) {
 				trace_nfs_size_update(inode, new_isize);
 				i_size_write(inode, new_isize);
 				if (!have_writers)
diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c
index 04b1987115d5..1ae047a062b8 100644
--- a/fs/nfs/nfs4proc.c
+++ b/fs/nfs/nfs4proc.c
@@ -10093,9 +10093,9 @@ static void nfs4_layoutcommit_release(void *calldata)
 {
 	struct nfs4_layoutcommit_data *data = calldata;
 
-	pnfs_cleanup_layoutcommit(data);
 	nfs_post_op_update_inode_force_wcc(data->args.inode,
 					   data->res.fattr);
+	pnfs_cleanup_layoutcommit(data);
 	put_cred(data->cred);
 	nfs_iput_and_deactive(data->inode);
 	kfree(data);
-- 
2.34.1


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh
  2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao
  2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao
@ 2026-10-08 22:28 ` Lucheng Bao
  1 sibling, 0 replies; 4+ messages in thread
From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw)
  To: linux-nfs, Trond Myklebust, Anna Schumaker
  Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao, stable

Commit c16467dc03db ("pnfs: Fix handling of NFS4ERR_OLD_STATEID replies
to layoutreturn") made stateid refresh fail while newer layout segments
remain busy. The standalone LAYOUTRETURN handler still treats this as a
completed return, potentially discarding a pending LAYOUTCOMMIT.

Set rpc_status to -EAGAIN before the existing completion handling so the
release callback defers the return for a live inode. Leave successful
refreshes and inode-teardown behavior unchanged.

On 6.8.0, this change applies without commit bbbff6d5edd1 ("NFSv4/pNFS:
Retry the layout return later in case of a timeout or reboot"), but that
commit and its prerequisites are required for the deferred-release
behavior.

Fixes: c16467dc03db ("pnfs: Fix handling of NFS4ERR_OLD_STATEID replies to layoutreturn")
Cc: stable@vger.kernel.org # See commit message for prerequisites
Signed-off-by: Lucheng Bao <lubao@everpuredata.com>
---
 fs/nfs/nfs4proc.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c
index 1ae047a062b8..9993e4d29d72 100644
--- a/fs/nfs/nfs4proc.c
+++ b/fs/nfs/nfs4proc.c
@@ -9867,6 +9867,7 @@ static void nfs4_layoutreturn_done(struct rpc_task *task, void *calldata)
 					&lrp->args.range,
 					lrp->args.inode))
 			goto out_restart;
+		lrp->rpc_status = -EAGAIN;
 		fallthrough;
 	default:
 		task->tk_status = 0;
-- 
2.34.1


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding
  2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao
@ 2026-10-09  4:43   ` Trond Myklebust
  0 siblings, 0 replies; 4+ messages in thread
From: Trond Myklebust @ 2026-10-09  4:43 UTC (permalink / raw)
  To: Lucheng Bao, linux-nfs, Anna Schumaker
  Cc: linux-kernel, tmenninger, jcurley, stable

On Thu, 2026-10-08 at 22:28 +0000, Lucheng Bao wrote:
> Commit ac46bd374c9a ("pNFS: Ensure we layoutcommit before
> revalidating
> attributes") replaced outstanding-layoutcommit attribute filtering
> with
> synchronization before explicit revalidation. For layouts requiring
> LAYOUTCOMMIT, OPEN attributes can still shrink i_size before metadata
> synchronization completes.

The client can't police OPEN. It has no idea which file will be
affected (particularly when doing an open-by-filename) so unlike the
case of an explicit truncate() call, it can't serialise with the
truncate.
It is therefore up to the server to resolve any ambiguity, either by
recalling or revoking the layout before allowing the OPEN truncate to
proceed.

> 
> Commit d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS
> LAYOUTCOMMIT is outstanding") moved LAYOUTCOMMIT post-op attribute
> processing after cleanup, leaving another unprotected window after
> NFS_INO_LAYOUTCOMMITTING is cleared.

This is unrelated to the above problem, so fixes for this issue need to
be in a separate patch.

> 
> Extend the existing writeback checks to reject smaller sizes while a
> LAYOUTCOMMIT is pending or in flight, preserving size increases and
> WCC
> checks. Process the reply's attributes and establish their generation
> barrier before cleanup removes that protection.
> 
> Fixes: d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS
> LAYOUTCOMMIT is outstanding")
> Fixes: ac46bd374c9a ("pNFS: Ensure we layoutcommit before
> revalidating attributes")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lucheng Bao <lubao@everpuredata.com>
> ---
>  fs/nfs/inode.c    | 8 ++++++--
>  fs/nfs/nfs4proc.c | 2 +-
>  2 files changed, 7 insertions(+), 3 deletions(-)
> 
> diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c
> index 3022454f7698..11432f521b04 100644
> --- a/fs/nfs/inode.c
> +++ b/fs/nfs/inode.c
> @@ -1636,7 +1636,9 @@ static void nfs_wcc_update_inode(struct inode
> *inode, struct nfs_fattr *fattr)
>  	if ((fattr->valid & NFS_ATTR_FATTR_PRESIZE)
>  			&& (fattr->valid & NFS_ATTR_FATTR_SIZE)
>  			&& i_size_read(inode) ==
> nfs_size_to_loff_t(fattr->pre_size)
> -			&& !nfs_have_writebacks(inode)) {
> +			&& !nfs_have_writebacks(inode)
> +			&& (nfs_size_to_loff_t(fattr->size) >=
> i_size_read(inode)
> +				||
> !pnfs_layoutcommit_outstanding(inode))) {

Under what circumstance is the client legally supposed to be able to
call nfs_wcc_update_inode() with NFS_ATTR_FATTR_PRESIZE set, while
there is an outstanding layoutcommit?

>  		trace_nfs_size_wcc(inode, fattr->size);
>  		i_size_write(inode, nfs_size_to_loff_t(fattr-
> >size));
>  	}
> @@ -2392,7 +2394,9 @@ static int nfs_update_inode(struct inode
> *inode, struct nfs_fattr *fattr)
>  		if (new_isize != cur_isize && !have_delegation) {
>  			/* Do we perhaps have any outstanding
> writes, or has
>  			 * the file grown beyond our last write? */
> -			if (!nfs_have_writebacks(inode) || new_isize
> > cur_isize) {
> +			if ((!nfs_have_writebacks(inode) &&
> +			     !pnfs_layoutcommit_outstanding(inode))
> ||
> +			    new_isize > cur_isize) {
>  				trace_nfs_size_update(inode,
> new_isize);
>  				i_size_write(inode, new_isize);
>  				if (!have_writers)
> diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c
> index 04b1987115d5..1ae047a062b8 100644
> --- a/fs/nfs/nfs4proc.c
> +++ b/fs/nfs/nfs4proc.c
> @@ -10093,9 +10093,9 @@ static void nfs4_layoutcommit_release(void
> *calldata)
>  {
>  	struct nfs4_layoutcommit_data *data = calldata;
>  
> -	pnfs_cleanup_layoutcommit(data);
>  	nfs_post_op_update_inode_force_wcc(data->args.inode,
>  					   data->res.fattr);
> +	pnfs_cleanup_layoutcommit(data);
>  	put_cred(data->cred);
>  	nfs_iput_and_deactive(data->inode);
>  	kfree(data);

-- 
Trond Myklebust
Linux NFS client maintainer, Hammerspace
trondmy@kernel.org, trond.myklebust@hammerspace.com

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-09  4:43 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao
2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao
2026-10-09  4:43   ` Trond Myklebust
2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®