mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
@ 2026-10-09 18:15 Tim Menninger
  2026-10-09 19:40 ` Chuck Lever
  0 siblings, 1 reply; 2+ messages in thread
From: Tim Menninger @ 2026-10-09 18:15 UTC (permalink / raw)
  To: Chuck Lever, Trond Myklebust, Anna Schumaker
  Cc: Olga Kornievskaia, Tom Talpey, linux-nfs, netdev, linux-kernel,
	Eric Badger, Jon Curley, Yongjian Mu, stable

From: Yongjian Mu <ymu@everpuredata.com>

The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
reply header size. For RPCSEC_GSS that estimate includes
auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
2 (25 XDR words, or 100 bytes) and which is corrected to the real
verifier size only after the first reply on that rpc_auth has been
unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
request encoded before the first reply completes has a head that is 64
bytes too long.

On TCP this is harmless: xdr_realign_pages() moves the excess head
bytes, which are genuine reply data, into the page list. With
RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
DDP is allowed), the payload has already been placed directly into the
page list, and the receive buffer holds only the inline part of the
reply. rpcrdma_inline_fixup() clamps its local copy of the head length
to the inline length but leaves head.iov_len untouched, so
xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
past the inline reply in the receive buffer into the front of the
directly-placed payload:

  rpc_xdr_recvfrom:  head=[...,196] page=524288 tail=[...,64]
  rpc_xdr_alignment: nfsv4 READ offset=132 copied=64

The first 64 bytes returned by each such READ are incorrect (zeroed, or
shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
implementation). This is readily hit with pNFS flexfiles over RDMA and
sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
large READ. Reads following the first completed reply return the
expected data.

Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
rpcrdma_inline_fixup()") removed the head-length correction while fixing
the handling of pure-inline krb5p replies.

When a Write chunk conveyed the payload, trim head.iov_len (in both
rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
the inline length actually received so that the XDR layer does not
realign the page list. Pure inline replies and Reply chunk replies are
unchanged.

Fixes: cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in rpcrdma_inline_fixup()")
Cc: stable@vger.kernel.org
Signed-off-by: Yongjian Mu <ymu@everpuredata.com>
Signed-off-by: Tim Menninger <tmenninger@everpuredata.com>
---
 net/sunrpc/xprtrdma/rpc_rdma.c | 23 +++++++++++++++++++----
 1 file changed, 19 insertions(+), 4 deletions(-)

diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c
index 1285f04cdac1..bebc9a20b28e 100644
--- a/net/sunrpc/xprtrdma/rpc_rdma.c
+++ b/net/sunrpc/xprtrdma/rpc_rdma.c
@@ -984,7 +984,7 @@ void rpcrdma_reset_cwnd(struct rpcrdma_xprt *r_xprt)
  * @rqst: controlling RPC request
  * @srcp: points to RPC message payload in receive buffer
  * @copy_len: remaining length of receive buffer content
- * @pad: Write chunk pad bytes needed (zero for pure inline)
+ * @writelist: bytes conveyed by the Write chunk (zero for pure inline)
  *
  * The upper layer has set the maximum number of bytes it can
  * receive in each component of rq_rcv_buf. These values are set in
@@ -998,13 +998,16 @@ void rpcrdma_reset_cwnd(struct rpcrdma_xprt *r_xprt)
  * Returns the count of bytes which had to be memcopied.
  */
 static unsigned long
-rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len, int pad)
+rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len,
+		     u32 writelist)
 {
 	unsigned long fixup_copy_count;
 	int i, npages, curlen;
 	char *destp;
 	struct page **ppages;
 	int page_base;
+	int pad = writelist & 3;
+	unsigned int delta;
 
 	/* The head iovec is redirected to the RPC reply message
 	 * in the receive buffer, to avoid a memcopy.
@@ -1016,8 +1019,20 @@ rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len, int pad)
 	 * head.iov_len bytes are copied into the page list.
 	 */
 	curlen = rqst->rq_rcv_buf.head[0].iov_len;
-	if (curlen > copy_len)
+	if (curlen > copy_len) {
+		/* These READ replies have no inline fields after the payload.
+		 * Trim the head without moving the directly-placed data.
+		 */
+		if (writelist) {
+			delta = curlen - copy_len;
+
+			rqst->rq_rcv_buf.head[0].iov_len = copy_len;
+			rqst->rq_private_buf.head[0].iov_len = copy_len;
+			rqst->rq_rcv_buf.buflen -= delta;
+			rqst->rq_private_buf.buflen -= delta;
+		}
 		curlen = copy_len;
+	}
 	srcp += curlen;
 	copy_len -= curlen;
 
@@ -1245,7 +1260,7 @@ rpcrdma_decode_msg(struct rpcrdma_xprt *r_xprt, struct rpcrdma_rep *rep,
 	base = (char *)xdr_inline_decode(xdr, 0);
 	rpclen = xdr_stream_remaining(xdr);
 	r_xprt->rx_stats.fixup_copy_count +=
-		rpcrdma_inline_fixup(rqst, base, rpclen, writelist & 3);
+		rpcrdma_inline_fixup(rqst, base, rpclen, writelist);
 
 	r_xprt->rx_stats.total_rdma_reply += writelist;
 	return rpclen + xdr_align_size(writelist);
-- 
2.34.1


^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
  2026-10-09 18:15 [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk Tim Menninger
@ 2026-10-09 19:40 ` Chuck Lever
  0 siblings, 0 replies; 2+ messages in thread
From: Chuck Lever @ 2026-10-09 19:40 UTC (permalink / raw)
  To: Tim Menninger, Trond Myklebust, Anna Schumaker
  Cc: Olga Kornievskaia, Tom Talpey, linux-nfs, netdev, linux-kernel,
	Eric Badger, Jon Curley, Yongjian Mu, stable



On Fri, Oct 9, 2026, at 2:15 PM, Tim Menninger wrote:
> From: Yongjian Mu <ymu@everpuredata.com>
>
> The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
> reply header size. For RPCSEC_GSS that estimate includes
> auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
> 2 (25 XDR words, or 100 bytes) and which is corrected to the real
> verifier size only after the first reply on that rpc_auth has been
> unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
> request encoded before the first reply completes has a head that is 64
> bytes too long.
>
> On TCP this is harmless: xdr_realign_pages() moves the excess head
> bytes, which are genuine reply data, into the page list. With
> RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
> DDP is allowed), the payload has already been placed directly into the
> page list, and the receive buffer holds only the inline part of the
> reply. rpcrdma_inline_fixup() clamps its local copy of the head length
> to the inline length but leaves head.iov_len untouched, so
> xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
> past the inline reply in the receive buffer into the front of the
> directly-placed payload:
>
>   rpc_xdr_recvfrom:  head=[...,196] page=524288 tail=[...,64]
>   rpc_xdr_alignment: nfsv4 READ offset=132 copied=64
>
> The first 64 bytes returned by each such READ are incorrect (zeroed, or
> shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
> implementation). This is readily hit with pNFS flexfiles over RDMA and
> sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
> large READ. Reads following the first completed reply return the
> expected data.
>
> Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
> rpcrdma_inline_fixup()") removed the head-length correction while fixing
> the handling of pure-inline krb5p replies.
>
> When a Write chunk conveyed the payload, trim head.iov_len (in both
> rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
> the inline length actually received so that the XDR layer does not
> realign the page list. Pure inline replies and Reply chunk replies are
> unchanged.

pNFS flexfiles exposes this issue because a per-DS rpc_clnt gets a fresh
rpc_auth, and its first RPC can be a large READ.

On the MDS mount (and on non-pNFS mounts) the first GSS reply is a small
non-DDP operation. The RPC client corrects the auth slack value before
any READ can occur.

Reviewed-by: Chuck Lever <cel@kernel.org>

Nit: The kernel-doc still says the upper layer's per-component maximums
live in head.iov_len and buflen, implying the transport leaves them alone.
That rule was set by commit cb0ae1fbb2f5.

But the inline fixup function now lowers both in rq_rcv_buf and
rq_private_buf when a Write chunk is present. The commit message should
mention that, and/or the kdoc should explain why Write chunks are a safe
exception to that rule.

Anna/Trond, does the NFS client or XDR code rely on the reply-side kvec
lengths staying as set during send preparation, or might there be an
in-tree reader of head[0].iov_len on the post-receive path that I missed?


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-09 19:41 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 18:15 [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk Tim Menninger
2026-10-09 19:40 ` Chuck Lever

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®