mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Chuck Lever" <cel@kernel.org>
To: "Tim Menninger" <tmenninger@everpuredata.com>,
	"Trond Myklebust" <trondmy@kernel.org>,
	"Anna Schumaker" <anna@kernel.org>
Cc: "Olga Kornievskaia" <okorniev@redhat.com>,
	"Tom Talpey" <tom@talpey.com>,
	linux-nfs@vger.kernel.org, netdev@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	"Eric Badger" <ebadger@everpuredata.com>,
	"Jon Curley" <jcurley@everpuredata.com>,
	"Yongjian Mu" <ymu@everpuredata.com>,
	stable@vger.kernel.org
Subject: Re: [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
Date: Fri, 09 Oct 2026 15:40:46 -0400	[thread overview]
Message-ID: <e62ff502-be18-454e-8a64-80b366d477f4@app.fastmail.com> (raw)
In-Reply-To: <20261009181554.2280374-1-tmenninger@everpuredata.com>



On Fri, Oct 9, 2026, at 2:15 PM, Tim Menninger wrote:
> From: Yongjian Mu <ymu@everpuredata.com>
>
> The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
> reply header size. For RPCSEC_GSS that estimate includes
> auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
> 2 (25 XDR words, or 100 bytes) and which is corrected to the real
> verifier size only after the first reply on that rpc_auth has been
> unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
> request encoded before the first reply completes has a head that is 64
> bytes too long.
>
> On TCP this is harmless: xdr_realign_pages() moves the excess head
> bytes, which are genuine reply data, into the page list. With
> RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
> DDP is allowed), the payload has already been placed directly into the
> page list, and the receive buffer holds only the inline part of the
> reply. rpcrdma_inline_fixup() clamps its local copy of the head length
> to the inline length but leaves head.iov_len untouched, so
> xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
> past the inline reply in the receive buffer into the front of the
> directly-placed payload:
>
>   rpc_xdr_recvfrom:  head=[...,196] page=524288 tail=[...,64]
>   rpc_xdr_alignment: nfsv4 READ offset=132 copied=64
>
> The first 64 bytes returned by each such READ are incorrect (zeroed, or
> shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
> implementation). This is readily hit with pNFS flexfiles over RDMA and
> sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
> large READ. Reads following the first completed reply return the
> expected data.
>
> Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
> rpcrdma_inline_fixup()") removed the head-length correction while fixing
> the handling of pure-inline krb5p replies.
>
> When a Write chunk conveyed the payload, trim head.iov_len (in both
> rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
> the inline length actually received so that the XDR layer does not
> realign the page list. Pure inline replies and Reply chunk replies are
> unchanged.

pNFS flexfiles exposes this issue because a per-DS rpc_clnt gets a fresh
rpc_auth, and its first RPC can be a large READ.

On the MDS mount (and on non-pNFS mounts) the first GSS reply is a small
non-DDP operation. The RPC client corrects the auth slack value before
any READ can occur.

Reviewed-by: Chuck Lever <cel@kernel.org>

Nit: The kernel-doc still says the upper layer's per-component maximums
live in head.iov_len and buflen, implying the transport leaves them alone.
That rule was set by commit cb0ae1fbb2f5.

But the inline fixup function now lowers both in rq_rcv_buf and
rq_private_buf when a Write chunk is present. The commit message should
mention that, and/or the kdoc should explain why Write chunks are a safe
exception to that rule.

Anna/Trond, does the NFS client or XDR code rely on the reply-side kvec
lengths staying as set during send preparation, or might there be an
in-tree reader of head[0].iov_len on the post-receive path that I missed?


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

      reply	other threads:[~2026-10-09 19:41 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09 18:15 Tim Menninger
2026-10-09 19:40 ` Chuck Lever [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=e62ff502-be18-454e-8a64-80b366d477f4@app.fastmail.com \
    --to=cel@kernel.org \
    --cc=anna@kernel.org \
    --cc=ebadger@everpuredata.com \
    --cc=jcurley@everpuredata.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=okorniev@redhat.com \
    --cc=stable@vger.kernel.org \
    --cc=tmenninger@everpuredata.com \
    --cc=tom@talpey.com \
    --cc=trondmy@kernel.org \
    --cc=ymu@everpuredata.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®