Re: [PATCH] xprtrdma: trim head iovec when the payload arrives via a Write chunk
From: Chuck Lever
Date: Fri Oct 09 2026 - 15:43:47 EST
On Fri, Oct 9, 2026, at 2:15 PM, Tim Menninger wrote:
> From: Yongjian Mu <ymu@xxxxxxxxxxxxxxxx>
>
> The upper layer sets rq_rcv_buf.head[0].iov_len from its estimate of the
> reply header size. For RPCSEC_GSS that estimate includes
> auth->au_ralign, which gss_create_new() initializes to GSS_VERF_SLACK >>
> 2 (25 XDR words, or 100 bytes) and which is corrected to the real
> verifier size only after the first reply on that rpc_auth has been
> unwrapped. In this case, the krb5 MIC verifier is 36 bytes, so every
> request encoded before the first reply completes has a head that is 64
> bytes too long.
>
> On TCP this is harmless: xdr_realign_pages() moves the excess head
> bytes, which are genuine reply data, into the page list. With
> RPC-over-RDMA and a Write chunk (sec=krb5, i.e. RPC_GSS_SVC_NONE, where
> DDP is allowed), the payload has already been placed directly into the
> page list, and the receive buffer holds only the inline part of the
> reply. rpcrdma_inline_fixup() clamps its local copy of the head length
> to the inline length but leaves head.iov_len untouched, so
> xdr_realign_pages() still sees iov_len > cur and shifts whatever lies
> past the inline reply in the receive buffer into the front of the
> directly-placed payload:
>
> rpc_xdr_recvfrom: head=[...,196] page=524288 tail=[...,64]
> rpc_xdr_alignment: nfsv4 READ offset=132 copied=64
>
> The first 64 bytes returned by each such READ are incorrect (zeroed, or
> shifted right by 64 bytes depending on the kernel's xdr_shrink_bufhead()
> implementation). This is readily hit with pNFS flexfiles over RDMA and
> sec=krb5: the per-DS rpc_clnt gets a fresh rpc_auth whose first RPC is a
> large READ. Reads following the first completed reply return the
> expected data.
>
> Commit cb0ae1fbb2f5 ("xprtrdma: Do not update {head, tail}.iov_len in
> rpcrdma_inline_fixup()") removed the head-length correction while fixing
> the handling of pure-inline krb5p replies.
>
> When a Write chunk conveyed the payload, trim head.iov_len (in both
> rq_rcv_buf and rq_private_buf, which call_decode() expects to match) to
> the inline length actually received so that the XDR layer does not
> realign the page list. Pure inline replies and Reply chunk replies are
> unchanged.
pNFS flexfiles exposes this issue because a per-DS rpc_clnt gets a fresh
rpc_auth, and its first RPC can be a large READ.
On the MDS mount (and on non-pNFS mounts) the first GSS reply is a small
non-DDP operation. The RPC client corrects the auth slack value before
any READ can occur.
Reviewed-by: Chuck Lever <cel@xxxxxxxxxx>
Nit: The kernel-doc still says the upper layer's per-component maximums
live in head.iov_len and buflen, implying the transport leaves them alone.
That rule was set by commit cb0ae1fbb2f5.
But the inline fixup function now lowers both in rq_rcv_buf and
rq_private_buf when a Write chunk is present. The commit message should
mention that, and/or the kdoc should explain why Write chunks are a safe
exception to that rule.
Anna/Trond, does the NFS client or XDR code rely on the reply-side kvec
lengths staying as set during send preparation, or might there be an
in-tree reader of head[0].iov_len on the post-receive path that I missed?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)