* [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
@ 2026-10-09 18:15 Tim Menninger
2026-10-09 19:40 ` Chuck Lever
0 siblings, 1 reply; 2+ messages in thread
From: Tim Menninger @ 2026-10-09 18:15 UTC (permalink / raw)
To: Chuck Lever, Trond Myklebust, Anna Schumaker
Cc: Olga Kornievskaia, Tom Talpey, linux-nfs, netdev, linux-kernel,
Eric Badger, Jon Curley, Yongjian Mu, stable
From: Yongjian Mu <ymu@everpuredata.com>
The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
reply header size. For RPCSEC_GSS that estimate includes
auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
2 (25 XDR words, or 100 bytes) and which is corrected to the real
verifier size only after the first reply on that rpc_auth has been
unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
request encoded before the first reply completes has a head that is 64
bytes too long.
On TCP this is harmless: xdr_realign_pages() moves the excess head
bytes, which are genuine reply data, into the page list. With
RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
DDP is allowed), the payload has already been placed directly into the
page list, and the receive buffer holds only the inline part of the
reply. rpcrdma_inline_fixup() clamps its local copy of the head length
to the inline length but leaves head.iov_len untouched, so
xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
past the inline reply in the receive buffer into the front of the
directly-placed payload:
rpc_xdr_recvfrom: head=[...,196] page=524288 tail=[...,64]
rpc_xdr_alignment: nfsv4 READ offset=132 copied=64
The first 64 bytes returned by each such READ are incorrect (zeroed, or
shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
implementation). This is readily hit with pNFS flexfiles over RDMA and
sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
large READ. Reads following the first completed reply return the
expected data.
Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
rpcrdma_inline_fixup()") removed the head-length correction while fixing
the handling of pure-inline krb5p replies.
When a Write chunk conveyed the payload, trim head.iov_len (in both
rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
the inline length actually received so that the XDR layer does not
realign the page list. Pure inline replies and Reply chunk replies are
unchanged.
Fixes: cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in rpcrdma_inline_fixup()")
Cc: stable@vger.kernel.org
Signed-off-by: Yongjian Mu <ymu@everpuredata.com>
Signed-off-by: Tim Menninger <tmenninger@everpuredata.com>
---
net/sunrpc/xprtrdma/rpc_rdma.c | 23 +++++++++++++++++++----
1 file changed, 19 insertions(+), 4 deletions(-)
diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c
index 1285f04cdac1..bebc9a20b28e 100644
--- a/net/sunrpc/xprtrdma/rpc_rdma.c
+++ b/net/sunrpc/xprtrdma/rpc_rdma.c
@@ -984,7 +984,7 @@ void rpcrdma_reset_cwnd(struct rpcrdma_xprt *r_xprt)
* @rqst: controlling RPC request
* @srcp: points to RPC message payload in receive buffer
* @copy_len: remaining length of receive buffer content
- * @pad: Write chunk pad bytes needed (zero for pure inline)
+ * @writelist: bytes conveyed by the Write chunk (zero for pure inline)
*
* The upper layer has set the maximum number of bytes it can
* receive in each component of rq_rcv_buf. These values are set in
@@ -998,13 +998,16 @@ void rpcrdma_reset_cwnd(struct rpcrdma_xprt *r_xprt)
* Returns the count of bytes which had to be memcopied.
*/
static unsigned long
-rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len, int pad)
+rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len,
+ u32 writelist)
{
unsigned long fixup_copy_count;
int i, npages, curlen;
char *destp;
struct page **ppages;
int page_base;
+ int pad = writelist & 3;
+ unsigned int delta;
/* The head iovec is redirected to the RPC reply message
* in the receive buffer, to avoid a memcopy.
@@ -1016,8 +1019,20 @@ rpcrdma_inline_fixup(struct rpc_rqst *rqst, char *srcp, int copy_len, int pad)
* head.iov_len bytes are copied into the page list.
*/
curlen = rqst->rq_rcv_buf.head[0].iov_len;
- if (curlen > copy_len)
+ if (curlen > copy_len) {
+ /* These READ replies have no inline fields after the payload.
+ * Trim the head without moving the directly-placed data.
+ */
+ if (writelist) {
+ delta = curlen - copy_len;
+
+ rqst->rq_rcv_buf.head[0].iov_len = copy_len;
+ rqst->rq_private_buf.head[0].iov_len = copy_len;
+ rqst->rq_rcv_buf.buflen -= delta;
+ rqst->rq_private_buf.buflen -= delta;
+ }
curlen = copy_len;
+ }
srcp += curlen;
copy_len -= curlen;
@@ -1245,7 +1260,7 @@ rpcrdma_decode_msg(struct rpcrdma_xprt *r_xprt, struct rpcrdma_rep *rep,
base = (char *)xdr_inline_decode(xdr, 0);
rpclen = xdr_stream_remaining(xdr);
r_xprt->rx_stats.fixup_copy_count +=
- rpcrdma_inline_fixup(rqst, base, rpclen, writelist & 3);
+ rpcrdma_inline_fixup(rqst, base, rpclen, writelist);
r_xprt->rx_stats.total_rdma_reply += writelist;
return rpclen + xdr_align_size(writelist);
--
2.34.1
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
2026-10-09 18:15 [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk Tim Menninger
@ 2026-10-09 19:40 ` Chuck Lever
0 siblings, 0 replies; 2+ messages in thread
From: Chuck Lever @ 2026-10-09 19:40 UTC (permalink / raw)
To: Tim Menninger, Trond Myklebust, Anna Schumaker
Cc: Olga Kornievskaia, Tom Talpey, linux-nfs, netdev, linux-kernel,
Eric Badger, Jon Curley, Yongjian Mu, stable
On Fri, Oct 9, 2026, at 2:15 PM, Tim Menninger wrote:
> From: Yongjian Mu <ymu@everpuredata.com>
>
> The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
> reply header size. For RPCSEC_GSS that estimate includes
> auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
> 2 (25 XDR words, or 100 bytes) and which is corrected to the real
> verifier size only after the first reply on that rpc_auth has been
> unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
> request encoded before the first reply completes has a head that is 64
> bytes too long.
>
> On TCP this is harmless: xdr_realign_pages() moves the excess head
> bytes, which are genuine reply data, into the page list. With
> RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
> DDP is allowed), the payload has already been placed directly into the
> page list, and the receive buffer holds only the inline part of the
> reply. rpcrdma_inline_fixup() clamps its local copy of the head length
> to the inline length but leaves head.iov_len untouched, so
> xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
> past the inline reply in the receive buffer into the front of the
> directly-placed payload:
>
> rpc_xdr_recvfrom: head=[...,196] page=524288 tail=[...,64]
> rpc_xdr_alignment: nfsv4 READ offset=132 copied=64
>
> The first 64 bytes returned by each such READ are incorrect (zeroed, or
> shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
> implementation). This is readily hit with pNFS flexfiles over RDMA and
> sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
> large READ. Reads following the first completed reply return the
> expected data.
>
> Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
> rpcrdma_inline_fixup()") removed the head-length correction while fixing
> the handling of pure-inline krb5p replies.
>
> When a Write chunk conveyed the payload, trim head.iov_len (in both
> rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
> the inline length actually received so that the XDR layer does not
> realign the page list. Pure inline replies and Reply chunk replies are
> unchanged.
pNFS flexfiles exposes this issue because a per-DS rpc_clnt gets a fresh
rpc_auth, and its first RPC can be a large READ.
On the MDS mount (and on non-pNFS mounts) the first GSS reply is a small
non-DDP operation. The RPC client corrects the auth slack value before
any READ can occur.
Reviewed-by: Chuck Lever <cel@kernel.org>
Nit: The kernel-doc still says the upper layer's per-component maximums
live in head.iov_len and buflen, implying the transport leaves them alone.
That rule was set by commit cb0ae1fbb2f5.
But the inline fixup function now lowers both in rq_rcv_buf and
rq_private_buf when a Write chunk is present. The commit message should
mention that, and/or the kdoc should explain why Write chunks are a safe
exception to that rule.
Anna/Trond, does the NFS client or XDR code rely on the reply-side kvec
lengths staying as set during send preparation, or might there be an
in-tree reader of head[0].iov_len on the post-receive path that I missed?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-09 19:41 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 18:15 [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk Tim Menninger
2026-10-09 19:40 ` Chuck Lever
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®