* [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding @ 2026-10-08 22:28 Lucheng Bao 2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao 2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao 0 siblings, 2 replies; 4+ messages in thread From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw) To: linux-nfs, Trond Myklebust, Anna Schumaker Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao When a layout requires LAYOUTCOMMIT, the size reported by the MDS may stay stale until the commit is processed. The client already keeps i_size from shrinking while writes are outstanding. This series extends that protection to an outstanding LAYOUTCOMMIT. Patch 1 rejects smaller sizes while a LAYOUTCOMMIT is pending or in flight. Patch 2 defers a LAYOUTRETURN whose OLD_STATEID refresh fails because a newer segment is busy. Today that return is treated as complete and the pending LAYOUTCOMMIT is discarded. Without this series, a stale OPEN or CLOSE reply can lower i_size, and the next O_APPEND write then lands at the old offset. Lucheng Bao (2): NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh fs/nfs/inode.c | 8 ++++++-- fs/nfs/nfs4proc.c | 3 ++- 2 files changed, 8 insertions(+), 3 deletions(-) -- 2.34.1 ^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding 2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao @ 2026-10-08 22:28 ` Lucheng Bao 2026-10-09 4:43 ` Trond Myklebust 2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao 1 sibling, 1 reply; 4+ messages in thread From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw) To: linux-nfs, Trond Myklebust, Anna Schumaker Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao, stable Commit ac46bd374c9a ("pNFS: Ensure we layoutcommit before revalidating attributes") replaced outstanding-layoutcommit attribute filtering with synchronization before explicit revalidation. For layouts requiring LAYOUTCOMMIT, OPEN attributes can still shrink i_size before metadata synchronization completes. Commit d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS LAYOUTCOMMIT is outstanding") moved LAYOUTCOMMIT post-op attribute processing after cleanup, leaving another unprotected window after NFS_INO_LAYOUTCOMMITTING is cleared. Extend the existing writeback checks to reject smaller sizes while a LAYOUTCOMMIT is pending or in flight, preserving size increases and WCC checks. Process the reply's attributes and establish their generation barrier before cleanup removes that protection. Fixes: d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS LAYOUTCOMMIT is outstanding") Fixes: ac46bd374c9a ("pNFS: Ensure we layoutcommit before revalidating attributes") Cc: stable@vger.kernel.org Signed-off-by: Lucheng Bao <lubao@everpuredata.com> --- fs/nfs/inode.c | 8 ++++++-- fs/nfs/nfs4proc.c | 2 +- 2 files changed, 7 insertions(+), 3 deletions(-) diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 3022454f7698..11432f521b04 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -1636,7 +1636,9 @@ static void nfs_wcc_update_inode(struct inode *inode, struct nfs_fattr *fattr) if ((fattr->valid & NFS_ATTR_FATTR_PRESIZE) && (fattr->valid & NFS_ATTR_FATTR_SIZE) && i_size_read(inode) == nfs_size_to_loff_t(fattr->pre_size) - && !nfs_have_writebacks(inode)) { + && !nfs_have_writebacks(inode) + && (nfs_size_to_loff_t(fattr->size) >= i_size_read(inode) + || !pnfs_layoutcommit_outstanding(inode))) { trace_nfs_size_wcc(inode, fattr->size); i_size_write(inode, nfs_size_to_loff_t(fattr->size)); } @@ -2392,7 +2394,9 @@ static int nfs_update_inode(struct inode *inode, struct nfs_fattr *fattr) if (new_isize != cur_isize && !have_delegation) { /* Do we perhaps have any outstanding writes, or has * the file grown beyond our last write? */ - if (!nfs_have_writebacks(inode) || new_isize > cur_isize) { + if ((!nfs_have_writebacks(inode) && + !pnfs_layoutcommit_outstanding(inode)) || + new_isize > cur_isize) { trace_nfs_size_update(inode, new_isize); i_size_write(inode, new_isize); if (!have_writers) diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c index 04b1987115d5..1ae047a062b8 100644 --- a/fs/nfs/nfs4proc.c +++ b/fs/nfs/nfs4proc.c @@ -10093,9 +10093,9 @@ static void nfs4_layoutcommit_release(void *calldata) { struct nfs4_layoutcommit_data *data = calldata; - pnfs_cleanup_layoutcommit(data); nfs_post_op_update_inode_force_wcc(data->args.inode, data->res.fattr); + pnfs_cleanup_layoutcommit(data); put_cred(data->cred); nfs_iput_and_deactive(data->inode); kfree(data); -- 2.34.1 ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit is outstanding 2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao @ 2026-10-09 4:43 ` Trond Myklebust 0 siblings, 0 replies; 4+ messages in thread From: Trond Myklebust @ 2026-10-09 4:43 UTC (permalink / raw) To: Lucheng Bao, linux-nfs, Anna Schumaker Cc: linux-kernel, tmenninger, jcurley, stable On Thu, 2026-10-08 at 22:28 +0000, Lucheng Bao wrote: > Commit ac46bd374c9a ("pNFS: Ensure we layoutcommit before > revalidating > attributes") replaced outstanding-layoutcommit attribute filtering > with > synchronization before explicit revalidation. For layouts requiring > LAYOUTCOMMIT, OPEN attributes can still shrink i_size before metadata > synchronization completes. The client can't police OPEN. It has no idea which file will be affected (particularly when doing an open-by-filename) so unlike the case of an explicit truncate() call, it can't serialise with the truncate. It is therefore up to the server to resolve any ambiguity, either by recalling or revoking the layout before allowing the OPEN truncate to proceed. > > Commit d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS > LAYOUTCOMMIT is outstanding") moved LAYOUTCOMMIT post-op attribute > processing after cleanup, leaving another unprotected window after > NFS_INO_LAYOUTCOMMITTING is cleared. This is unrelated to the above problem, so fixes for this issue need to be in a separate patch. > > Extend the existing writeback checks to reject smaller sizes while a > LAYOUTCOMMIT is pending or in flight, preserving size increases and > WCC > checks. Process the reply's attributes and establish their generation > barrier before cleanup removes that protection. > > Fixes: d8c951c313ed ("NFSv4.1: Don't trust attributes if a pNFS > LAYOUTCOMMIT is outstanding") > Fixes: ac46bd374c9a ("pNFS: Ensure we layoutcommit before > revalidating attributes") > Cc: stable@vger.kernel.org > Signed-off-by: Lucheng Bao <lubao@everpuredata.com> > --- > fs/nfs/inode.c | 8 ++++++-- > fs/nfs/nfs4proc.c | 2 +- > 2 files changed, 7 insertions(+), 3 deletions(-) > > diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c > index 3022454f7698..11432f521b04 100644 > --- a/fs/nfs/inode.c > +++ b/fs/nfs/inode.c > @@ -1636,7 +1636,9 @@ static void nfs_wcc_update_inode(struct inode > *inode, struct nfs_fattr *fattr) > if ((fattr->valid & NFS_ATTR_FATTR_PRESIZE) > && (fattr->valid & NFS_ATTR_FATTR_SIZE) > && i_size_read(inode) == > nfs_size_to_loff_t(fattr->pre_size) > - && !nfs_have_writebacks(inode)) { > + && !nfs_have_writebacks(inode) > + && (nfs_size_to_loff_t(fattr->size) >= > i_size_read(inode) > + || > !pnfs_layoutcommit_outstanding(inode))) { Under what circumstance is the client legally supposed to be able to call nfs_wcc_update_inode() with NFS_ATTR_FATTR_PRESIZE set, while there is an outstanding layoutcommit? > trace_nfs_size_wcc(inode, fattr->size); > i_size_write(inode, nfs_size_to_loff_t(fattr- > >size)); > } > @@ -2392,7 +2394,9 @@ static int nfs_update_inode(struct inode > *inode, struct nfs_fattr *fattr) > if (new_isize != cur_isize && !have_delegation) { > /* Do we perhaps have any outstanding > writes, or has > * the file grown beyond our last write? */ > - if (!nfs_have_writebacks(inode) || new_isize > > cur_isize) { > + if ((!nfs_have_writebacks(inode) && > + !pnfs_layoutcommit_outstanding(inode)) > || > + new_isize > cur_isize) { > trace_nfs_size_update(inode, > new_isize); > i_size_write(inode, new_isize); > if (!have_writers) > diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c > index 04b1987115d5..1ae047a062b8 100644 > --- a/fs/nfs/nfs4proc.c > +++ b/fs/nfs/nfs4proc.c > @@ -10093,9 +10093,9 @@ static void nfs4_layoutcommit_release(void > *calldata) > { > struct nfs4_layoutcommit_data *data = calldata; > > - pnfs_cleanup_layoutcommit(data); > nfs_post_op_update_inode_force_wcc(data->args.inode, > data->res.fattr); > + pnfs_cleanup_layoutcommit(data); > put_cred(data->cred); > nfs_iput_and_deactive(data->inode); > kfree(data); -- Trond Myklebust Linux NFS client maintainer, Hammerspace trondmy@kernel.org, trond.myklebust@hammerspace.com ^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh 2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao 2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao @ 2026-10-08 22:28 ` Lucheng Bao 1 sibling, 0 replies; 4+ messages in thread From: Lucheng Bao @ 2026-10-08 22:28 UTC (permalink / raw) To: linux-nfs, Trond Myklebust, Anna Schumaker Cc: linux-kernel, tmenninger, jcurley, Lucheng Bao, stable Commit c16467dc03db ("pnfs: Fix handling of NFS4ERR_OLD_STATEID replies to layoutreturn") made stateid refresh fail while newer layout segments remain busy. The standalone LAYOUTRETURN handler still treats this as a completed return, potentially discarding a pending LAYOUTCOMMIT. Set rpc_status to -EAGAIN before the existing completion handling so the release callback defers the return for a live inode. Leave successful refreshes and inode-teardown behavior unchanged. On 6.8.0, this change applies without commit bbbff6d5edd1 ("NFSv4/pNFS: Retry the layout return later in case of a timeout or reboot"), but that commit and its prerequisites are required for the deferred-release behavior. Fixes: c16467dc03db ("pnfs: Fix handling of NFS4ERR_OLD_STATEID replies to layoutreturn") Cc: stable@vger.kernel.org # See commit message for prerequisites Signed-off-by: Lucheng Bao <lubao@everpuredata.com> --- fs/nfs/nfs4proc.c | 1 + 1 file changed, 1 insertion(+) diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c index 1ae047a062b8..9993e4d29d72 100644 --- a/fs/nfs/nfs4proc.c +++ b/fs/nfs/nfs4proc.c @@ -9867,6 +9867,7 @@ static void nfs4_layoutreturn_done(struct rpc_task *task, void *calldata) &lrp->args.range, lrp->args.inode)) goto out_restart; + lrp->rpc_status = -EAGAIN; fallthrough; default: task->tk_status = 0; -- 2.34.1 ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-09 4:43 UTC | newest] Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-10-08 22:28 [PATCH 0/2] NFSv4/pNFS: don't trust inode size from MDS while LAYOUTCOMMIT is outstanding Lucheng Bao 2026-10-08 22:28 ` [PATCH 1/2] NFSv4/pNFS: prevent i_size regression while layoutcommit " Lucheng Bao 2026-10-09 4:43 ` Trond Myklebust 2026-10-08 22:28 ` [PATCH 2/2] NFSv4/pNFS: defer LAYOUTRETURN after an unsuccessful OLD_STATEID refresh Lucheng Bao
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®