* [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
` (6 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
The NFS client will soon pass PCI peer-to-peer DMA (P2PDMA) pages
to the transport as READ or WRITE payload. Only a transport that
can DMA directly to and from that memory may move it.
Introduce XDRBUF_P2PDMA and a helper, xprt_rqst_has_p2pdma(), to
mark and detect requests whose pages are P2PDMA memory. The TCP
(including TLS), UDP and AF_LOCAL transports return -EREMOTEIO for
such a request before touching the socket. iov_iter_extract_pages()
returns the same error today for P2PDMA pages the caller did not allow.
Thus, users see no new error code.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
include/linux/sunrpc/xdr.h | 1 +
include/linux/sunrpc/xprt.h | 5 +++++
net/sunrpc/xprtsock.c | 12 ++++++++++++
3 files changed, 18 insertions(+)
diff --git a/include/linux/sunrpc/xdr.h b/include/linux/sunrpc/xdr.h
index b102b4f21e6b..79526f47be2d 100644
--- a/include/linux/sunrpc/xdr.h
+++ b/include/linux/sunrpc/xdr.h
@@ -65,6 +65,7 @@ struct xdr_buf {
#define XDRBUF_READ 0x01 /* target of file read */
#define XDRBUF_WRITE 0x02 /* source of file write */
#define XDRBUF_SPARSE_PAGES 0x04 /* Page array is sparse */
+#define XDRBUF_P2PDMA 0x08 /* Pages are P2PDMA memory */
unsigned int buflen, /* Total length of storage buffer */
len; /* Length of XDR encoded message */
diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h
index a82045804d34..08ab04fa4509 100644
--- a/include/linux/sunrpc/xprt.h
+++ b/include/linux/sunrpc/xprt.h
@@ -134,6 +134,11 @@ static inline int xprt_rqst_add_seqno(struct rpc_rqst *req, u32 seqno)
return 0;
}
+static inline bool xprt_rqst_has_p2pdma(const struct rpc_rqst *req)
+{
+ return (req->rq_snd_buf.flags | req->rq_rcv_buf.flags) & XDRBUF_P2PDMA;
+}
+
/* RPC transport layer security policies */
enum xprtsec_policies {
RPC_XPRTSEC_NONE = 0,
diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
index f5a5136327ba..9fa6d7ec207e 100644
--- a/net/sunrpc/xprtsock.c
+++ b/net/sunrpc/xprtsock.c
@@ -962,6 +962,10 @@ static int xs_local_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
/* Close the stream if the previous transmission was incomplete */
if (xs_send_request_was_aborted(transport, req)) {
xprt_force_disconnect(xprt);
@@ -1031,6 +1035,10 @@ static int xs_udp_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
xs_pktdump("packet data:",
req->rq_svec->iov_base,
req->rq_svec->iov_len);
@@ -1118,6 +1126,10 @@ static int xs_tcp_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
/* Close the stream if the previous transmission was incomplete */
if (xs_send_request_was_aborted(transport, req)) {
if (transport->sock != NULL)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
` (5 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
A transport returns -EREMOTEIO from ->send_request() for a request
whose P2PDMA pages it cannot move. Retrying doesn't help, but today
the error stops xprt_transmit(), finds whichever task is draining
the transmit queue, and leaves the refused request queued.
Dequeue the request, on -EREMOTEIO, cancel only its own task with
rpc_task_try_cancel(), and keep draining. Setting tk_status instead
would not work, a dequeued task looks transmitted and would wait for
a reply that never comes.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprt.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c
index 48a3618cbb29..084d5a508357 100644
--- a/net/sunrpc/xprt.c
+++ b/net/sunrpc/xprt.c
@@ -1579,6 +1579,15 @@ xprt_request_transmit(struct rpc_rqst *req, struct rpc_task *snd_task)
if (status != 0) {
req->rq_ntrans--;
trace_xprt_transmit(req, status);
+ if (status == -EREMOTEIO) {
+ /*
+ * This transport will never send req. Fail only
+ * its task, and let the caller keep draining.
+ */
+ xprt_request_dequeue_transmit(task);
+ rpc_task_try_cancel(task, status);
+ return 0;
+ }
return status;
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
` (4 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
Only chunks move a P2PDMA payload by DMA directly between the NIC
and the device memory. The inline paths copy the payload via CPU
accesses or DMA-map it with calls that do not handle P2PDMA pages.
Always send a P2PDMA WRITE payload in a Read chunk, and receive a
P2PDMA READ payload in a Write chunk, whatever its size. Return
-EREMOTEIO when that is not possible: the NIC cannot DMA to PCI
peer-to-peer memory (e.g. rxe, siw), the GSS service forbids direct
data placement (krb5i, krb5p), or the rest of the READ reply does
not fit inline.
If a server returns READ data inline anyway, fail the RPC with
-EREMOTEIO instead of copying the data into the P2PDMA pages.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprtrdma/rpc_rdma.c | 40 ++++++++++++++++++++++++++++++++--
1 file changed, 38 insertions(+), 2 deletions(-)
diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c
index 1285f04cdac1..065ab9a7edc9 100644
--- a/net/sunrpc/xprtrdma/rpc_rdma.c
+++ b/net/sunrpc/xprtrdma/rpc_rdma.c
@@ -175,6 +175,22 @@ rpcrdma_nonpayload_inline(const struct rpcrdma_xprt *r_xprt,
r_xprt->rx_ep->re_max_inline_recv;
}
+/* A P2PDMA payload moves only by DMA between the NIC and device
+ * memory, in its own Read or Write chunk. That requires a NIC that
+ * can DMA to PCI peer-to-peer memory and, for a READ, a Reply whose
+ * non-payload part fits inline.
+ */
+static bool
+rpcrdma_p2pdma_allowed(const struct rpcrdma_xprt *r_xprt,
+ const struct rpc_rqst *rqst)
+{
+ if (!ib_dma_pci_p2p_dma_supported(r_xprt->rx_ep->re_id->device))
+ return false;
+ if (rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA)
+ return rpcrdma_nonpayload_inline(r_xprt, rqst);
+ return true;
+}
+
/* ACL likes to be lazy in allocating pages. For TCP, these
* pages can be allocated during receive processing. Not true
* for RDMA, which must always provision receive buffers
@@ -815,6 +831,7 @@ inline int rpcrdma_prepare_send_sges(struct rpcrdma_xprt *r_xprt,
* %-EAGAIN if the caller should call again with the same arguments,
* %-ENOBUFS if the caller should call again after a delay,
* %-EMSGSIZE if the transport header is too small,
+ * %-EREMOTEIO if the device cannot move the request's P2PDMA pages,
* %-EIO if a permanent problem occurred while marshaling.
*/
int
@@ -854,16 +871,25 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
ddp_allowed = !test_bit(RPCAUTH_AUTH_DATATOUCH,
&rqst->rq_cred->cr_auth->au_flags);
+ if (xprt_rqst_has_p2pdma(rqst) &&
+ (!ddp_allowed || !rpcrdma_p2pdma_allowed(r_xprt, rqst))) {
+ ret = -EREMOTEIO;
+ goto out_err;
+ }
+
/*
* Chunks needed for results?
*
+ * o A P2PDMA read payload always returns in a write chunk.
* o If the expected result is under the inline threshold, all ops
* return as inline.
* o Large read ops return data as write chunk(s), header as
* inline.
* o Large non-read ops return as a single reply chunk.
*/
- if (rpcrdma_results_inline(r_xprt, rqst))
+ if (rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA)
+ wtype = rpcrdma_writech;
+ else if (rpcrdma_results_inline(r_xprt, rqst))
wtype = rpcrdma_noch;
else if ((ddp_allowed && rqst->rq_rcv_buf.flags & XDRBUF_READ) &&
rpcrdma_nonpayload_inline(r_xprt, rqst))
@@ -874,6 +900,7 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
/*
* Chunks needed for arguments?
*
+ * o A P2PDMA write payload is always sent as a read chunk.
* o If the total request is under the inline threshold, all ops
* are sent as inline.
* o Large write ops transmit data as read chunk(s), header as
@@ -885,7 +912,10 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
* that both has a data payload, and whose non-data arguments
* by themselves are larger than the inline threshold.
*/
- if (rpcrdma_args_inline(r_xprt, rqst)) {
+ if (buf->flags & XDRBUF_P2PDMA) {
+ *p++ = rdma_msg;
+ rtype = rpcrdma_readch;
+ } else if (rpcrdma_args_inline(r_xprt, rqst)) {
*p++ = rdma_msg;
rtype = buf->len < rdmab_length(req->rl_sendbuf) ?
rpcrdma_noch_pullup : rpcrdma_noch_mapped;
@@ -1244,6 +1274,12 @@ rpcrdma_decode_msg(struct rpcrdma_xprt *r_xprt, struct rpcrdma_rep *rep,
/* Build the RPC reply's Payload stream in rqst->rq_rcv_buf */
base = (char *)xdr_inline_decode(xdr, 0);
rpclen = xdr_stream_remaining(xdr);
+
+ /* Never copy inline reply data into P2PDMA pages */
+ if (unlikely(rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA &&
+ rpclen > rqst->rq_rcv_buf.head[0].iov_len))
+ return -EREMOTEIO;
+
r_xprt->rx_stats.fixup_copy_count +=
rpcrdma_inline_fixup(rqst, base, rpclen, writelist & 3);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (2 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
` (3 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
ib_dma_map_sg() returns 0 on any failure, so frwr_map() cannot tell
why a mapping failed and reports -EIO. Use ib_dma_map_sgtable_attrs()
instead, which returns the reason. Pass up -EREMOTEIO, which means
the device cannot reach the P2PDMA memory, so the RPC fails with the
same error as other P2PDMA refusals. Every other mapping failure
still returns -EIO.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprtrdma/frwr_ops.c | 15 ++++++++++-----
1 file changed, 10 insertions(+), 5 deletions(-)
diff --git a/net/sunrpc/xprtrdma/frwr_ops.c b/net/sunrpc/xprtrdma/frwr_ops.c
index e83cef19e656..a55d5e11c0e2 100644
--- a/net/sunrpc/xprtrdma/frwr_ops.c
+++ b/net/sunrpc/xprtrdma/frwr_ops.c
@@ -293,7 +293,8 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
bool sg_gaps = ep->re_mrtype == IB_MR_TYPE_SG_GAPS;
unsigned int max_depth = ep->re_max_fr_depth;
struct ib_reg_wr *reg_wr;
- int i, n, dma_nents;
+ struct sg_table sgt;
+ int i, n, dma_nents, ret;
struct ib_mr *ibmr;
u8 key;
@@ -381,10 +382,13 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
mr->mr_dir = rpcrdma_data_dir(writing);
mr->mr_nents = i;
- dma_nents = ib_dma_map_sg(ep->re_id->device, mr->mr_sg, mr->mr_nents,
- mr->mr_dir);
- if (!dma_nents)
+ /* Unlike ib_dma_map_sg(), this reports why a mapping failed */
+ sgt.sgl = mr->mr_sg;
+ sgt.orig_nents = mr->mr_nents;
+ ret = ib_dma_map_sgtable_attrs(ep->re_id->device, &sgt, mr->mr_dir, 0);
+ if (ret || !sgt.nents)
goto out_dmamap_err;
+ dma_nents = sgt.nents;
mr->mr_device = ep->re_id->device;
ibmr = mr->mr_ibmr;
@@ -413,7 +417,8 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
out_dmamap_err:
trace_xprtrdma_frwr_sgerr(mr, i);
- return -EIO;
+ /* Return -EREMOTEIO if the device cannot reach the P2PDMA memory */
+ return ret == -EREMOTEIO ? ret : -EIO;
out_mapmr_err:
trace_xprtrdma_frwr_maperr(mr, n);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (3 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
` (2 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
A transport must know when an RPC's payload pages are P2PDMA memory,
so it can move them by DMA or refuse the RPC.
Check each page as nfs_generic_pgio() builds the page array, and
record the result in nfs_pgio_args. The READ and WRITE encoders of
every NFS version then set XDRBUF_P2PDMA on the buffer holding the
payload.
Additionally, avoid using READ_PLUS when the pages are P2PDMA memory.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/nfs2xdr.c | 4 ++++
fs/nfs/nfs3xdr.c | 4 ++++
fs/nfs/nfs4proc.c | 3 +++
fs/nfs/nfs4xdr.c | 4 ++++
fs/nfs/pagelist.c | 2 ++
include/linux/nfs_xdr.h | 1 +
6 files changed, 18 insertions(+)
diff --git a/fs/nfs/nfs2xdr.c b/fs/nfs/nfs2xdr.c
index 9eff09158518..5a8f2eccd068 100644
--- a/fs/nfs/nfs2xdr.c
+++ b/fs/nfs/nfs2xdr.c
@@ -628,6 +628,8 @@ static void nfs2_xdr_enc_readargs(struct rpc_rqst *req,
rpc_prepare_reply_pages(req, args->pages, args->pgbase, args->count,
NFS_readres_sz - NFS_pagepad_sz);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
}
/*
@@ -668,6 +670,8 @@ static void nfs2_xdr_enc_writeargs(struct rpc_rqst *req,
encode_writeargs(xdr, args);
xdr->buf->flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ xdr->buf->flags |= XDRBUF_P2PDMA;
}
/*
diff --git a/fs/nfs/nfs3xdr.c b/fs/nfs/nfs3xdr.c
index e745e78faab0..e39c58883a2a 100644
--- a/fs/nfs/nfs3xdr.c
+++ b/fs/nfs/nfs3xdr.c
@@ -946,6 +946,8 @@ static void nfs3_xdr_enc_read3args(struct rpc_rqst *req,
rpc_prepare_reply_pages(req, args->pages, args->pgbase,
args->count, replen);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
}
/*
@@ -988,6 +990,8 @@ static void nfs3_xdr_enc_write3args(struct rpc_rqst *req,
encode_write3args(xdr, args);
xdr->buf->flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ xdr->buf->flags |= XDRBUF_P2PDMA;
}
/*
diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c
index beb659744760..443079f5adc9 100644
--- a/fs/nfs/nfs4proc.c
+++ b/fs/nfs/nfs4proc.c
@@ -5721,6 +5721,9 @@ static int nfs4_read_done(struct rpc_task *task, struct nfs_pgio_header *hdr)
static bool nfs42_read_plus_support(struct nfs_pgio_header *hdr,
struct rpc_message *msg)
{
+ /* P2PDMA payloads move only by DMA */
+ if (hdr->args.p2pdma)
+ return false;
/* Note: We don't use READ_PLUS with pNFS yet */
if (nfs_server_capable(hdr->inode, NFS_CAP_READ_PLUS) && !hdr->ds_clp) {
msg->rpc_proc = &nfs4_procedures[NFSPROC4_CLNT_READ_PLUS];
diff --git a/fs/nfs/nfs4xdr.c b/fs/nfs/nfs4xdr.c
index 8b3d96b4a0f0..04765b56de28 100644
--- a/fs/nfs/nfs4xdr.c
+++ b/fs/nfs/nfs4xdr.c
@@ -2621,6 +2621,8 @@ static void nfs4_xdr_enc_read(struct rpc_rqst *req, struct xdr_stream *xdr,
rpc_prepare_reply_pages(req, args->pages, args->pgbase,
args->count, hdr.replen - pagepad_maxsz);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
encode_nops(&hdr);
}
@@ -2686,6 +2688,8 @@ static void nfs4_xdr_enc_write(struct rpc_rqst *req, struct xdr_stream *xdr,
encode_putfh(xdr, args->fh, &hdr);
encode_write(xdr, args, &hdr);
req->rq_snd_buf.flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ req->rq_snd_buf.flags |= XDRBUF_P2PDMA;
if (args->bitmask)
encode_getfattr(xdr, args->bitmask, &hdr);
encode_nops(&hdr);
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index 71f0ce2bc4ea..3ec4e1e5fe1f 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -970,6 +970,8 @@ int nfs_generic_pgio(struct nfs_pageio_descriptor *desc,
if (pageused > pagecount)
goto full;
*pages++ = last_page = page;
+ if (is_pci_p2pdma_page(page))
+ hdr->args.p2pdma = true;
}
}
}
diff --git a/include/linux/nfs_xdr.h b/include/linux/nfs_xdr.h
index c0e29b4dfa62..85cdee107559 100644
--- a/include/linux/nfs_xdr.h
+++ b/include/linux/nfs_xdr.h
@@ -685,6 +685,7 @@ struct nfs_pgio_args {
__u32 count;
unsigned int pgbase;
struct page ** pages;
+ bool p2pdma; /* pages include P2PDMA memory */
union {
unsigned int replen; /* used by read */
struct {
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (4 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
krb5i checksums and krb5p encrypts the payload via CPU accesses
while the RPC is encoded, before the transport can refuse it.
Fail such READs and WRITEs with -EREMOTEIO in nfs_pgio_prepare(),
before encoding. Check the task's RPC client, which for a WRITE may
be the krb5i client that SP4_MACH_CRED swaps in.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/pagelist.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index 3ec4e1e5fe1f..ec535ed3051d 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -771,6 +771,14 @@ static void nfs_pgio_prepare(struct rpc_task *task, void *calldata)
{
struct nfs_pgio_header *hdr = calldata;
int err;
+
+ /* P2PDMA payloads move only by DMA */
+ if (hdr->args.p2pdma &&
+ test_bit(RPCAUTH_AUTH_DATATOUCH,
+ &task->tk_client->cl_auth->au_flags)) {
+ rpc_exit(task, -EREMOTEIO);
+ return;
+ }
err = NFS_PROTO(hdr->inode)->pgio_rpc_prepare(task, hdr);
if (err)
rpc_exit(task, err);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (5 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
LOCALIO hands READs and WRITEs to the local filesystem, which can take
P2PDMA pages only if its disk can DMA to them. Use LOCALIO for such I/O
only if the filesystem's block device supports P2PDMA. Otherwise send a
regular READ or WRITE RPC, as if LOCALIO were off.
When the I/O is not a single DIO-aligned segment, LOCALIO falls back to
buffered I/O, which copies the payload through the page cache via CPU
accesses. Fail such I/O with -EINVAL instead.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/internal.h | 8 ++++++++
fs/nfs/localio.c | 40 ++++++++++++++++++++++++++++++++++++++++
fs/nfs/pagelist.c | 3 +++
3 files changed, 51 insertions(+)
diff --git a/fs/nfs/internal.h b/fs/nfs/internal.h
index d6ea41a3f9b4..9338b5e0896a 100644
--- a/fs/nfs/internal.h
+++ b/fs/nfs/internal.h
@@ -476,6 +476,7 @@ extern struct nfsd_file *nfs_local_open_fh(struct nfs_client *,
struct nfs_fh *,
struct nfs_file_localio *,
const fmode_t);
+struct nfsd_file *nfs_local_p2pdma_check(struct nfsd_file *localio);
extern int nfs_local_doio(struct nfs_client *,
struct nfsd_file *,
struct nfs_pgio_header *,
@@ -495,6 +496,13 @@ nfs_local_open_fh(struct nfs_client *clp, const struct cred *cred,
{
return NULL;
}
+
+static inline struct nfsd_file *
+nfs_local_p2pdma_check(struct nfsd_file *localio)
+{
+ return NULL;
+}
+
static inline int nfs_local_doio(struct nfs_client *clp,
struct nfsd_file *localio,
struct nfs_pgio_header *hdr,
diff --git a/fs/nfs/localio.c b/fs/nfs/localio.c
index f42b6112a613..6c8f2d393c35 100644
--- a/fs/nfs/localio.c
+++ b/fs/nfs/localio.c
@@ -19,6 +19,7 @@
#include <linux/nfs_common.h>
#include <linux/nfslocalio.h>
#include <linux/bvec.h>
+#include <linux/blkdev.h>
#include <linux/nfs.h>
#include <linux/nfs_fs.h>
@@ -290,6 +291,22 @@ nfs_local_open_fh(struct nfs_client *clp, const struct cred *cred,
}
EXPORT_SYMBOL_GPL(nfs_local_open_fh);
+/*
+ * LOCALIO moves P2PDMA pages only by DMA to and from the local disk.
+ * If that disk cannot do P2PDMA, put @localio and return NULL so the
+ * I/O is sent as a regular RPC instead.
+ */
+struct nfsd_file *nfs_local_p2pdma_check(struct nfsd_file *localio)
+{
+ struct file *file = nfs_to->nfsd_file_file(localio);
+ struct block_device *bdev = file_inode(file)->i_sb->s_bdev;
+
+ if (bdev && blk_queue_pci_p2pdma(bdev_get_queue(bdev)))
+ return localio;
+ nfs_local_file_put(localio);
+ return NULL;
+}
+
/*
* Ensure all page cache allocations are done from GFP_NOFS context to
* prevent direct reclaim recursion back into NFS via nfs_writepages.
@@ -511,6 +528,17 @@ nfs_local_iters_init(struct nfs_local_kiocb *iocb, int rw)
iov_iter_bvec(&iocb->iters[0], rw, iocb->bvec, v, len);
}
+/*
+ * P2PDMA payloads move only by DMA, never through the page cache, so the
+ * whole I/O must be a single DIO-aligned segment.
+ */
+static bool nfs_local_p2pdma_misaligned(struct nfs_local_kiocb *iocb)
+{
+ return iocb->hdr->args.p2pdma &&
+ (atomic_read(&iocb->n_iters) != 1 ||
+ !iocb->iter_is_dio_aligned[0]);
+}
+
static void
nfs_local_hdr_release(struct nfs_pgio_header *hdr,
const struct rpc_call_ops *call_ops)
@@ -732,6 +760,12 @@ static void nfs_local_do_read(struct nfs_local_kiocb *iocb,
nfs_local_pgio_init(hdr, call_ops);
hdr->res.eof = false;
+ if (nfs_local_p2pdma_misaligned(iocb)) {
+ nfs_local_pgio_done(iocb, -EINVAL);
+ nfs_local_pgio_release(iocb);
+ return;
+ }
+
INIT_WORK(&iocb->work, nfs_local_call_read);
if (nfs_local_defer_io())
queue_work(nfslocaliod_workqueue, &iocb->work);
@@ -953,6 +987,12 @@ static void nfs_local_do_write(struct nfs_local_kiocb *iocb,
nfs_set_local_verifier(hdr->inode, hdr->res.verf, hdr->args.stable);
+ if (nfs_local_p2pdma_misaligned(iocb)) {
+ nfs_local_pgio_done(iocb, -EINVAL);
+ nfs_local_pgio_release(iocb);
+ return;
+ }
+
INIT_WORK(&iocb->work, nfs_local_call_write);
if (nfs_local_defer_io())
queue_work(nfslocaliod_workqueue, &iocb->work);
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index ec535ed3051d..2d1d07c460fd 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -1023,6 +1023,9 @@ static int nfs_generic_pg_pgios(struct nfs_pageio_descriptor *desc)
&hdr->args.context->nfl,
hdr->args.context->mode);
+ if (localio && hdr->args.p2pdma)
+ localio = nfs_local_p2pdma_check(localio);
+
if (NFS_SERVER(hdr->inode)->nfs_client->cl_minorversion)
task_flags = RPC_TASK_MOVEABLE;
ret = nfs_initiate_pgio(NFS_CLIENT(hdr->inode),
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread* [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (6 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
Pass ITER_ALLOW_P2PDMA when extracting O_DIRECT pages on mounts that
use the RDMA transport and not pNFS. The RDMA transport moves such pages
only by DMA, and fails the RPC with -EREMOTEIO if the device cannot
reach them.
On other mounts, extracting P2PDMA pages still fails with -EREMOTEIO.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/direct.c | 10 +++++++++-
1 file changed, 9 insertions(+), 1 deletion(-)
diff --git a/fs/nfs/direct.c b/fs/nfs/direct.c
index a3c6e8f4ea06..dbeb4f8274e3 100644
--- a/fs/nfs/direct.c
+++ b/fs/nfs/direct.c
@@ -157,14 +157,22 @@ static ssize_t nfs_direct_extract_pages(struct nfs_direct_req *dreq,
size_t size, loff_t *pos,
struct list_head *list)
{
+ struct nfs_server *server = NFS_SERVER(dreq->inode);
bool pinned = iov_iter_extract_will_pin(iter);
+ iov_iter_extraction_t flags = 0;
struct page **pagevec = NULL;
ssize_t result, bytes = 0;
int err = 0;
unsigned int npages, i;
size_t pgbase;
- result = iov_iter_extract_pages(iter, &pagevec, size, ~0U, 0, &pgbase);
+ /* Allow P2PDMA pages only on RDMA mounts that do not use pNFS */
+ if (server->nfs_client->cl_proto == XPRT_TRANSPORT_RDMA &&
+ !pnfs_enabled_sb(server))
+ flags = ITER_ALLOW_P2PDMA;
+
+ result = iov_iter_extract_pages(iter, &pagevec, size, ~0U, flags,
+ &pgbase);
if (result <= 0)
return result;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread