* [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA
@ 2026-10-06 23:32 Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
` (7 more replies)
0 siblings, 8 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
Enable NFS O_DIRECT to and from PCI peer-to-peer DMA (P2PDMA) memory
on RDMA mounts. This replaces the initial RFC [1] and builds on the
Direct I/O modernization series [2].
P2PDMA memory can be an NVMe Controller Memory Buffer, or the BAR of a
vfio-pci device exported as a ZONE_DEVICE-backed DMABUF, as proposed
in [3]. Today, NFS O_DIRECT fails with -EREMOTEIO when the user buffer
is such memory, since iov_iter_extract_pages() refuses P2PDMA pages
unless the caller passes ITER_ALLOW_P2PDMA. Buffered I/O is not
affected, as the CPU copies the data into the page cache.
P2PDMA memory may not be safe to access with the CPU. Thus, the series
follows one rule: a P2PDMA payload moves only by DMA between the device
memory and the NIC (or the local disk with LOCALIO). Any path that
would touch the payload via CPU accesses refuses the I/O instead of
copying it.
Design
======
O_DIRECT, user buffer in P2PDMA memory
|
nfs_direct_extract_pages() ITER_ALLOW_P2PDMA on RDMA, !pNFS mounts
| (other mounts: -EREMOTEIO, as today)
nfs_generic_pgio() is_pci_p2pdma_page() -> args.p2pdma
|
+-- LOCALIO: disk lacks BLK_FEAT_PCI_P2PDMA -> regular RPC
| not one DIO-aligned segment -> -EINVAL
|
+-- RPC: krb5i / krb5p -> -EREMOTEIO
encoders set XDRBUF_P2PDMA
|
+-- xprtsock -> -EREMOTEIO
+-- xprtrdma: NIC can't do P2P -> -EREMOTEIO
NIC can't reach it -> -EREMOTEIO
otherwise: Read/Write chunks only
1. Per-RPC decision
ITER_ALLOW_P2PDMA at extraction is only a hint. NFS tags each READ and
WRITE whose pages are P2PDMA memory with XDRBUF_P2PDMA, and the
transport serving that RPC decides whether it can move them. With
nconnect, reconnects and migration, the transport and its device can
change between extraction and transmission, so a mount-time capability
cannot be trusted. LOCALIO does not use the NIC at all.
2. Socket transports
TCP (including TLS), UDP and AF_LOCAL move data via CPU accesses, so
they return -EREMOTEIO for a tagged RPC before touching the socket.
xprt_transmit() now fails only the refused RPC and keeps draining the
transmit queue.
3. RPC-over-RDMA
The inline paths copy the payload via CPU accesses or DMA-map it with
calls that do not handle P2PDMA pages. Thus, xprtrdma always sends a
P2PDMA WRITE payload in a Read chunk and receives a P2PDMA READ payload
in a Write chunk, whatever its size. It returns -EREMOTEIO when:
- the NIC cannot DMA to PCI peer-to-peer memory (e.g. rxe, siw),
- the NIC cannot reach the P2PDMA memory (DMA mapping fails),
- the GSS service forbids direct data placement (krb5i, krb5p), or
- the rest of the READ reply does not fit inline.
If a server returns READ data inline despite the Write chunk, which
RFC 8166 forbids, the RPC fails instead of copying into the pages.
Small P2PDMA I/O therefore always pays for memory registration.
4. GSS
krb5i checksums and krb5p encrypts the payload via CPU accesses while
the RPC is encoded, before the transport sees it. nfs_pgio_prepare()
fails such READs and WRITEs with -EREMOTEIO before encoding.
5. LOCALIO
LOCALIO is used for P2PDMA pages only if the filesystem's block device
supports P2PDMA. Otherwise, the I/O goes out as a regular RPC. I/O that
is not a single DIO-aligned segment would fall back to buffered I/O,
so it fails with -EINVAL instead.
6. READ_PLUS
READ_PLUS decoding moves data and zero-fills holes via CPU accesses, so
it is never used for P2PDMA pages. RDMA mounts already disable it.
All refusals return -EREMOTEIO, the error iov_iter_extract_pages()
returns today, except misaligned LOCALIO I/O, which returns -EINVAL.
Refused I/O fails at once, even on hard mounts.
Call for review
===============
Feedback is appreciated on:
1. Layering: "RDMA transport, !pNFS" is only a hint, and each RPC
re-checks its transport. Should the transport advertise P2PDMA
capability up front instead? That would also allow LOCALIO on TCP
mounts, which this series refuses.
2. nconnect: fail an RPC on a transport without P2PDMA, or pick a
capable transport?
3. Migration / reconnect: the path can move to a NIC without P2PDMA,
and I/O that worked starts failing with -EREMOTEIO. Acceptable?
4. Hard mounts: refused I/O fails at once instead of retrying.
Acceptable?
5. LOCALIO: the check uses sb->s_bdev, which misses multi-device
filesystems. Is there a better way to ask whether a file can do
P2PDMA? Should misaligned I/O fall back to an RPC instead of
failing with -EINVAL?
6. pNFS: excluded for now. Route P2PDMA I/O through the MDS, or let
each layout driver opt in?
Known issue: the pNFS check happens at extraction. If a mount gains
pNFS afterwards (e.g. after migration), a rescheduled WRITE could reach
a layout driver. I plan to decide once per O_DIRECT call and force the
MDS path when P2PDMA pages are allowed.
Thanks,
Praan
[1] https://lore.kernel.org/all/20260401194501.2269200-1-praan@google.com/
[2] https://lore.kernel.org/all/20260814143255.861084-1-praan@google.com/
[3] https://lore.kernel.org/all/20260804185050.2053672-1-praan@google.com/
Changes since the initial RFC [1]:
- Split the Direct I/O modernization into its own series [2].
- Drop NFS_CAP_P2PDMA, the 64-bit caps expansion and the
->supports_p2pdma xprt op. Decide per RPC instead, as Chuck
suggested.
- Refuse P2PDMA pages in the socket transports, and fail only the
refused RPC.
- Move P2PDMA payloads only via chunks in xprtrdma, and return
-EREMOTEIO when the NIC cannot reach the memory.
- Refuse krb5i and krb5p, and allow LOCALIO only as aligned direct
I/O to a P2PDMA-capable disk.
Pranjal Shrivastava (8):
sunrpc: introduce XDRBUF_P2PDMA flag
sunrpc: fail only the RPC a transport refuses
xprtrdma: move P2PDMA payloads only via chunks
xprtrdma: return -EREMOTEIO on P2PDMA map failure
nfs: tag READ/WRITE RPCs carrying P2PDMA pages
nfs: refuse P2PDMA pages with krb5i and krb5p
nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O
nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts
fs/nfs/direct.c | 10 ++++++++-
fs/nfs/internal.h | 8 +++++++
fs/nfs/localio.c | 40 ++++++++++++++++++++++++++++++++++
fs/nfs/nfs2xdr.c | 4 ++++
fs/nfs/nfs3xdr.c | 4 ++++
fs/nfs/nfs4proc.c | 3 +++
fs/nfs/nfs4xdr.c | 4 ++++
fs/nfs/pagelist.c | 13 +++++++++++
include/linux/nfs_xdr.h | 1 +
include/linux/sunrpc/xdr.h | 1 +
include/linux/sunrpc/xprt.h | 5 +++++
net/sunrpc/xprt.c | 9 ++++++++
net/sunrpc/xprtrdma/frwr_ops.c | 15 ++++++++-----
net/sunrpc/xprtrdma/rpc_rdma.c | 40 ++++++++++++++++++++++++++++++++--
net/sunrpc/xprtsock.c | 12 ++++++++++
15 files changed, 161 insertions(+), 8 deletions(-)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
` (6 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
The NFS client will soon pass PCI peer-to-peer DMA (P2PDMA) pages
to the transport as READ or WRITE payload. Only a transport that
can DMA directly to and from that memory may move it.
Introduce XDRBUF_P2PDMA and a helper, xprt_rqst_has_p2pdma(), to
mark and detect requests whose pages are P2PDMA memory. The TCP
(including TLS), UDP and AF_LOCAL transports return -EREMOTEIO for
such a request before touching the socket. iov_iter_extract_pages()
returns the same error today for P2PDMA pages the caller did not allow.
Thus, users see no new error code.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
include/linux/sunrpc/xdr.h | 1 +
include/linux/sunrpc/xprt.h | 5 +++++
net/sunrpc/xprtsock.c | 12 ++++++++++++
3 files changed, 18 insertions(+)
diff --git a/include/linux/sunrpc/xdr.h b/include/linux/sunrpc/xdr.h
index b102b4f21e6b..79526f47be2d 100644
--- a/include/linux/sunrpc/xdr.h
+++ b/include/linux/sunrpc/xdr.h
@@ -65,6 +65,7 @@ struct xdr_buf {
#define XDRBUF_READ 0x01 /* target of file read */
#define XDRBUF_WRITE 0x02 /* source of file write */
#define XDRBUF_SPARSE_PAGES 0x04 /* Page array is sparse */
+#define XDRBUF_P2PDMA 0x08 /* Pages are P2PDMA memory */
unsigned int buflen, /* Total length of storage buffer */
len; /* Length of XDR encoded message */
diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h
index a82045804d34..08ab04fa4509 100644
--- a/include/linux/sunrpc/xprt.h
+++ b/include/linux/sunrpc/xprt.h
@@ -134,6 +134,11 @@ static inline int xprt_rqst_add_seqno(struct rpc_rqst *req, u32 seqno)
return 0;
}
+static inline bool xprt_rqst_has_p2pdma(const struct rpc_rqst *req)
+{
+ return (req->rq_snd_buf.flags | req->rq_rcv_buf.flags) & XDRBUF_P2PDMA;
+}
+
/* RPC transport layer security policies */
enum xprtsec_policies {
RPC_XPRTSEC_NONE = 0,
diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
index f5a5136327ba..9fa6d7ec207e 100644
--- a/net/sunrpc/xprtsock.c
+++ b/net/sunrpc/xprtsock.c
@@ -962,6 +962,10 @@ static int xs_local_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
/* Close the stream if the previous transmission was incomplete */
if (xs_send_request_was_aborted(transport, req)) {
xprt_force_disconnect(xprt);
@@ -1031,6 +1035,10 @@ static int xs_udp_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
xs_pktdump("packet data:",
req->rq_svec->iov_base,
req->rq_svec->iov_len);
@@ -1118,6 +1126,10 @@ static int xs_tcp_send_request(struct rpc_rqst *req)
unsigned int sent;
int status;
+ /* Sockets move data via CPU accesses, so refuse P2PDMA pages */
+ if (xprt_rqst_has_p2pdma(req))
+ return -EREMOTEIO;
+
/* Close the stream if the previous transmission was incomplete */
if (xs_send_request_was_aborted(transport, req)) {
if (transport->sock != NULL)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
` (5 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
A transport returns -EREMOTEIO from ->send_request() for a request
whose P2PDMA pages it cannot move. Retrying doesn't help, but today
the error stops xprt_transmit(), finds whichever task is draining
the transmit queue, and leaves the refused request queued.
Dequeue the request, on -EREMOTEIO, cancel only its own task with
rpc_task_try_cancel(), and keep draining. Setting tk_status instead
would not work, a dequeued task looks transmitted and would wait for
a reply that never comes.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprt.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/net/sunrpc/xprt.c b/net/sunrpc/xprt.c
index 48a3618cbb29..084d5a508357 100644
--- a/net/sunrpc/xprt.c
+++ b/net/sunrpc/xprt.c
@@ -1579,6 +1579,15 @@ xprt_request_transmit(struct rpc_rqst *req, struct rpc_task *snd_task)
if (status != 0) {
req->rq_ntrans--;
trace_xprt_transmit(req, status);
+ if (status == -EREMOTEIO) {
+ /*
+ * This transport will never send req. Fail only
+ * its task, and let the caller keep draining.
+ */
+ xprt_request_dequeue_transmit(task);
+ rpc_task_try_cancel(task, status);
+ return 0;
+ }
return status;
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
` (4 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
Only chunks move a P2PDMA payload by DMA directly between the NIC
and the device memory. The inline paths copy the payload via CPU
accesses or DMA-map it with calls that do not handle P2PDMA pages.
Always send a P2PDMA WRITE payload in a Read chunk, and receive a
P2PDMA READ payload in a Write chunk, whatever its size. Return
-EREMOTEIO when that is not possible: the NIC cannot DMA to PCI
peer-to-peer memory (e.g. rxe, siw), the GSS service forbids direct
data placement (krb5i, krb5p), or the rest of the READ reply does
not fit inline.
If a server returns READ data inline anyway, fail the RPC with
-EREMOTEIO instead of copying the data into the P2PDMA pages.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprtrdma/rpc_rdma.c | 40 ++++++++++++++++++++++++++++++++--
1 file changed, 38 insertions(+), 2 deletions(-)
diff --git a/net/sunrpc/xprtrdma/rpc_rdma.c b/net/sunrpc/xprtrdma/rpc_rdma.c
index 1285f04cdac1..065ab9a7edc9 100644
--- a/net/sunrpc/xprtrdma/rpc_rdma.c
+++ b/net/sunrpc/xprtrdma/rpc_rdma.c
@@ -175,6 +175,22 @@ rpcrdma_nonpayload_inline(const struct rpcrdma_xprt *r_xprt,
r_xprt->rx_ep->re_max_inline_recv;
}
+/* A P2PDMA payload moves only by DMA between the NIC and device
+ * memory, in its own Read or Write chunk. That requires a NIC that
+ * can DMA to PCI peer-to-peer memory and, for a READ, a Reply whose
+ * non-payload part fits inline.
+ */
+static bool
+rpcrdma_p2pdma_allowed(const struct rpcrdma_xprt *r_xprt,
+ const struct rpc_rqst *rqst)
+{
+ if (!ib_dma_pci_p2p_dma_supported(r_xprt->rx_ep->re_id->device))
+ return false;
+ if (rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA)
+ return rpcrdma_nonpayload_inline(r_xprt, rqst);
+ return true;
+}
+
/* ACL likes to be lazy in allocating pages. For TCP, these
* pages can be allocated during receive processing. Not true
* for RDMA, which must always provision receive buffers
@@ -815,6 +831,7 @@ inline int rpcrdma_prepare_send_sges(struct rpcrdma_xprt *r_xprt,
* %-EAGAIN if the caller should call again with the same arguments,
* %-ENOBUFS if the caller should call again after a delay,
* %-EMSGSIZE if the transport header is too small,
+ * %-EREMOTEIO if the device cannot move the request's P2PDMA pages,
* %-EIO if a permanent problem occurred while marshaling.
*/
int
@@ -854,16 +871,25 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
ddp_allowed = !test_bit(RPCAUTH_AUTH_DATATOUCH,
&rqst->rq_cred->cr_auth->au_flags);
+ if (xprt_rqst_has_p2pdma(rqst) &&
+ (!ddp_allowed || !rpcrdma_p2pdma_allowed(r_xprt, rqst))) {
+ ret = -EREMOTEIO;
+ goto out_err;
+ }
+
/*
* Chunks needed for results?
*
+ * o A P2PDMA read payload always returns in a write chunk.
* o If the expected result is under the inline threshold, all ops
* return as inline.
* o Large read ops return data as write chunk(s), header as
* inline.
* o Large non-read ops return as a single reply chunk.
*/
- if (rpcrdma_results_inline(r_xprt, rqst))
+ if (rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA)
+ wtype = rpcrdma_writech;
+ else if (rpcrdma_results_inline(r_xprt, rqst))
wtype = rpcrdma_noch;
else if ((ddp_allowed && rqst->rq_rcv_buf.flags & XDRBUF_READ) &&
rpcrdma_nonpayload_inline(r_xprt, rqst))
@@ -874,6 +900,7 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
/*
* Chunks needed for arguments?
*
+ * o A P2PDMA write payload is always sent as a read chunk.
* o If the total request is under the inline threshold, all ops
* are sent as inline.
* o Large write ops transmit data as read chunk(s), header as
@@ -885,7 +912,10 @@ rpcrdma_marshal_req(struct rpcrdma_xprt *r_xprt, struct rpc_rqst *rqst)
* that both has a data payload, and whose non-data arguments
* by themselves are larger than the inline threshold.
*/
- if (rpcrdma_args_inline(r_xprt, rqst)) {
+ if (buf->flags & XDRBUF_P2PDMA) {
+ *p++ = rdma_msg;
+ rtype = rpcrdma_readch;
+ } else if (rpcrdma_args_inline(r_xprt, rqst)) {
*p++ = rdma_msg;
rtype = buf->len < rdmab_length(req->rl_sendbuf) ?
rpcrdma_noch_pullup : rpcrdma_noch_mapped;
@@ -1244,6 +1274,12 @@ rpcrdma_decode_msg(struct rpcrdma_xprt *r_xprt, struct rpcrdma_rep *rep,
/* Build the RPC reply's Payload stream in rqst->rq_rcv_buf */
base = (char *)xdr_inline_decode(xdr, 0);
rpclen = xdr_stream_remaining(xdr);
+
+ /* Never copy inline reply data into P2PDMA pages */
+ if (unlikely(rqst->rq_rcv_buf.flags & XDRBUF_P2PDMA &&
+ rpclen > rqst->rq_rcv_buf.head[0].iov_len))
+ return -EREMOTEIO;
+
r_xprt->rx_stats.fixup_copy_count +=
rpcrdma_inline_fixup(rqst, base, rpclen, writelist & 3);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (2 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
` (3 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
ib_dma_map_sg() returns 0 on any failure, so frwr_map() cannot tell
why a mapping failed and reports -EIO. Use ib_dma_map_sgtable_attrs()
instead, which returns the reason. Pass up -EREMOTEIO, which means
the device cannot reach the P2PDMA memory, so the RPC fails with the
same error as other P2PDMA refusals. Every other mapping failure
still returns -EIO.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
net/sunrpc/xprtrdma/frwr_ops.c | 15 ++++++++++-----
1 file changed, 10 insertions(+), 5 deletions(-)
diff --git a/net/sunrpc/xprtrdma/frwr_ops.c b/net/sunrpc/xprtrdma/frwr_ops.c
index e83cef19e656..a55d5e11c0e2 100644
--- a/net/sunrpc/xprtrdma/frwr_ops.c
+++ b/net/sunrpc/xprtrdma/frwr_ops.c
@@ -293,7 +293,8 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
bool sg_gaps = ep->re_mrtype == IB_MR_TYPE_SG_GAPS;
unsigned int max_depth = ep->re_max_fr_depth;
struct ib_reg_wr *reg_wr;
- int i, n, dma_nents;
+ struct sg_table sgt;
+ int i, n, dma_nents, ret;
struct ib_mr *ibmr;
u8 key;
@@ -381,10 +382,13 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
mr->mr_dir = rpcrdma_data_dir(writing);
mr->mr_nents = i;
- dma_nents = ib_dma_map_sg(ep->re_id->device, mr->mr_sg, mr->mr_nents,
- mr->mr_dir);
- if (!dma_nents)
+ /* Unlike ib_dma_map_sg(), this reports why a mapping failed */
+ sgt.sgl = mr->mr_sg;
+ sgt.orig_nents = mr->mr_nents;
+ ret = ib_dma_map_sgtable_attrs(ep->re_id->device, &sgt, mr->mr_dir, 0);
+ if (ret || !sgt.nents)
goto out_dmamap_err;
+ dma_nents = sgt.nents;
mr->mr_device = ep->re_id->device;
ibmr = mr->mr_ibmr;
@@ -413,7 +417,8 @@ int frwr_map(struct rpcrdma_xprt *r_xprt,
out_dmamap_err:
trace_xprtrdma_frwr_sgerr(mr, i);
- return -EIO;
+ /* Return -EREMOTEIO if the device cannot reach the P2PDMA memory */
+ return ret == -EREMOTEIO ? ret : -EIO;
out_mapmr_err:
trace_xprtrdma_frwr_maperr(mr, n);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (3 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
` (2 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
A transport must know when an RPC's payload pages are P2PDMA memory,
so it can move them by DMA or refuse the RPC.
Check each page as nfs_generic_pgio() builds the page array, and
record the result in nfs_pgio_args. The READ and WRITE encoders of
every NFS version then set XDRBUF_P2PDMA on the buffer holding the
payload.
Additionally, avoid using READ_PLUS when the pages are P2PDMA memory.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/nfs2xdr.c | 4 ++++
fs/nfs/nfs3xdr.c | 4 ++++
fs/nfs/nfs4proc.c | 3 +++
fs/nfs/nfs4xdr.c | 4 ++++
fs/nfs/pagelist.c | 2 ++
include/linux/nfs_xdr.h | 1 +
6 files changed, 18 insertions(+)
diff --git a/fs/nfs/nfs2xdr.c b/fs/nfs/nfs2xdr.c
index 9eff09158518..5a8f2eccd068 100644
--- a/fs/nfs/nfs2xdr.c
+++ b/fs/nfs/nfs2xdr.c
@@ -628,6 +628,8 @@ static void nfs2_xdr_enc_readargs(struct rpc_rqst *req,
rpc_prepare_reply_pages(req, args->pages, args->pgbase, args->count,
NFS_readres_sz - NFS_pagepad_sz);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
}
/*
@@ -668,6 +670,8 @@ static void nfs2_xdr_enc_writeargs(struct rpc_rqst *req,
encode_writeargs(xdr, args);
xdr->buf->flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ xdr->buf->flags |= XDRBUF_P2PDMA;
}
/*
diff --git a/fs/nfs/nfs3xdr.c b/fs/nfs/nfs3xdr.c
index e745e78faab0..e39c58883a2a 100644
--- a/fs/nfs/nfs3xdr.c
+++ b/fs/nfs/nfs3xdr.c
@@ -946,6 +946,8 @@ static void nfs3_xdr_enc_read3args(struct rpc_rqst *req,
rpc_prepare_reply_pages(req, args->pages, args->pgbase,
args->count, replen);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
}
/*
@@ -988,6 +990,8 @@ static void nfs3_xdr_enc_write3args(struct rpc_rqst *req,
encode_write3args(xdr, args);
xdr->buf->flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ xdr->buf->flags |= XDRBUF_P2PDMA;
}
/*
diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c
index beb659744760..443079f5adc9 100644
--- a/fs/nfs/nfs4proc.c
+++ b/fs/nfs/nfs4proc.c
@@ -5721,6 +5721,9 @@ static int nfs4_read_done(struct rpc_task *task, struct nfs_pgio_header *hdr)
static bool nfs42_read_plus_support(struct nfs_pgio_header *hdr,
struct rpc_message *msg)
{
+ /* P2PDMA payloads move only by DMA */
+ if (hdr->args.p2pdma)
+ return false;
/* Note: We don't use READ_PLUS with pNFS yet */
if (nfs_server_capable(hdr->inode, NFS_CAP_READ_PLUS) && !hdr->ds_clp) {
msg->rpc_proc = &nfs4_procedures[NFSPROC4_CLNT_READ_PLUS];
diff --git a/fs/nfs/nfs4xdr.c b/fs/nfs/nfs4xdr.c
index 8b3d96b4a0f0..04765b56de28 100644
--- a/fs/nfs/nfs4xdr.c
+++ b/fs/nfs/nfs4xdr.c
@@ -2621,6 +2621,8 @@ static void nfs4_xdr_enc_read(struct rpc_rqst *req, struct xdr_stream *xdr,
rpc_prepare_reply_pages(req, args->pages, args->pgbase,
args->count, hdr.replen - pagepad_maxsz);
req->rq_rcv_buf.flags |= XDRBUF_READ;
+ if (args->p2pdma)
+ req->rq_rcv_buf.flags |= XDRBUF_P2PDMA;
encode_nops(&hdr);
}
@@ -2686,6 +2688,8 @@ static void nfs4_xdr_enc_write(struct rpc_rqst *req, struct xdr_stream *xdr,
encode_putfh(xdr, args->fh, &hdr);
encode_write(xdr, args, &hdr);
req->rq_snd_buf.flags |= XDRBUF_WRITE;
+ if (args->p2pdma)
+ req->rq_snd_buf.flags |= XDRBUF_P2PDMA;
if (args->bitmask)
encode_getfattr(xdr, args->bitmask, &hdr);
encode_nops(&hdr);
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index 71f0ce2bc4ea..3ec4e1e5fe1f 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -970,6 +970,8 @@ int nfs_generic_pgio(struct nfs_pageio_descriptor *desc,
if (pageused > pagecount)
goto full;
*pages++ = last_page = page;
+ if (is_pci_p2pdma_page(page))
+ hdr->args.p2pdma = true;
}
}
}
diff --git a/include/linux/nfs_xdr.h b/include/linux/nfs_xdr.h
index c0e29b4dfa62..85cdee107559 100644
--- a/include/linux/nfs_xdr.h
+++ b/include/linux/nfs_xdr.h
@@ -685,6 +685,7 @@ struct nfs_pgio_args {
__u32 count;
unsigned int pgbase;
struct page ** pages;
+ bool p2pdma; /* pages include P2PDMA memory */
union {
unsigned int replen; /* used by read */
struct {
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (4 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
krb5i checksums and krb5p encrypts the payload via CPU accesses
while the RPC is encoded, before the transport can refuse it.
Fail such READs and WRITEs with -EREMOTEIO in nfs_pgio_prepare(),
before encoding. Check the task's RPC client, which for a WRITE may
be the krb5i client that SP4_MACH_CRED swaps in.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/pagelist.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index 3ec4e1e5fe1f..ec535ed3051d 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -771,6 +771,14 @@ static void nfs_pgio_prepare(struct rpc_task *task, void *calldata)
{
struct nfs_pgio_header *hdr = calldata;
int err;
+
+ /* P2PDMA payloads move only by DMA */
+ if (hdr->args.p2pdma &&
+ test_bit(RPCAUTH_AUTH_DATATOUCH,
+ &task->tk_client->cl_auth->au_flags)) {
+ rpc_exit(task, -EREMOTEIO);
+ return;
+ }
err = NFS_PROTO(hdr->inode)->pgio_rpc_prepare(task, hdr);
if (err)
rpc_exit(task, err);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (5 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
LOCALIO hands READs and WRITEs to the local filesystem, which can take
P2PDMA pages only if its disk can DMA to them. Use LOCALIO for such I/O
only if the filesystem's block device supports P2PDMA. Otherwise send a
regular READ or WRITE RPC, as if LOCALIO were off.
When the I/O is not a single DIO-aligned segment, LOCALIO falls back to
buffered I/O, which copies the payload through the page cache via CPU
accesses. Fail such I/O with -EINVAL instead.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/internal.h | 8 ++++++++
fs/nfs/localio.c | 40 ++++++++++++++++++++++++++++++++++++++++
fs/nfs/pagelist.c | 3 +++
3 files changed, 51 insertions(+)
diff --git a/fs/nfs/internal.h b/fs/nfs/internal.h
index d6ea41a3f9b4..9338b5e0896a 100644
--- a/fs/nfs/internal.h
+++ b/fs/nfs/internal.h
@@ -476,6 +476,7 @@ extern struct nfsd_file *nfs_local_open_fh(struct nfs_client *,
struct nfs_fh *,
struct nfs_file_localio *,
const fmode_t);
+struct nfsd_file *nfs_local_p2pdma_check(struct nfsd_file *localio);
extern int nfs_local_doio(struct nfs_client *,
struct nfsd_file *,
struct nfs_pgio_header *,
@@ -495,6 +496,13 @@ nfs_local_open_fh(struct nfs_client *clp, const struct cred *cred,
{
return NULL;
}
+
+static inline struct nfsd_file *
+nfs_local_p2pdma_check(struct nfsd_file *localio)
+{
+ return NULL;
+}
+
static inline int nfs_local_doio(struct nfs_client *clp,
struct nfsd_file *localio,
struct nfs_pgio_header *hdr,
diff --git a/fs/nfs/localio.c b/fs/nfs/localio.c
index f42b6112a613..6c8f2d393c35 100644
--- a/fs/nfs/localio.c
+++ b/fs/nfs/localio.c
@@ -19,6 +19,7 @@
#include <linux/nfs_common.h>
#include <linux/nfslocalio.h>
#include <linux/bvec.h>
+#include <linux/blkdev.h>
#include <linux/nfs.h>
#include <linux/nfs_fs.h>
@@ -290,6 +291,22 @@ nfs_local_open_fh(struct nfs_client *clp, const struct cred *cred,
}
EXPORT_SYMBOL_GPL(nfs_local_open_fh);
+/*
+ * LOCALIO moves P2PDMA pages only by DMA to and from the local disk.
+ * If that disk cannot do P2PDMA, put @localio and return NULL so the
+ * I/O is sent as a regular RPC instead.
+ */
+struct nfsd_file *nfs_local_p2pdma_check(struct nfsd_file *localio)
+{
+ struct file *file = nfs_to->nfsd_file_file(localio);
+ struct block_device *bdev = file_inode(file)->i_sb->s_bdev;
+
+ if (bdev && blk_queue_pci_p2pdma(bdev_get_queue(bdev)))
+ return localio;
+ nfs_local_file_put(localio);
+ return NULL;
+}
+
/*
* Ensure all page cache allocations are done from GFP_NOFS context to
* prevent direct reclaim recursion back into NFS via nfs_writepages.
@@ -511,6 +528,17 @@ nfs_local_iters_init(struct nfs_local_kiocb *iocb, int rw)
iov_iter_bvec(&iocb->iters[0], rw, iocb->bvec, v, len);
}
+/*
+ * P2PDMA payloads move only by DMA, never through the page cache, so the
+ * whole I/O must be a single DIO-aligned segment.
+ */
+static bool nfs_local_p2pdma_misaligned(struct nfs_local_kiocb *iocb)
+{
+ return iocb->hdr->args.p2pdma &&
+ (atomic_read(&iocb->n_iters) != 1 ||
+ !iocb->iter_is_dio_aligned[0]);
+}
+
static void
nfs_local_hdr_release(struct nfs_pgio_header *hdr,
const struct rpc_call_ops *call_ops)
@@ -732,6 +760,12 @@ static void nfs_local_do_read(struct nfs_local_kiocb *iocb,
nfs_local_pgio_init(hdr, call_ops);
hdr->res.eof = false;
+ if (nfs_local_p2pdma_misaligned(iocb)) {
+ nfs_local_pgio_done(iocb, -EINVAL);
+ nfs_local_pgio_release(iocb);
+ return;
+ }
+
INIT_WORK(&iocb->work, nfs_local_call_read);
if (nfs_local_defer_io())
queue_work(nfslocaliod_workqueue, &iocb->work);
@@ -953,6 +987,12 @@ static void nfs_local_do_write(struct nfs_local_kiocb *iocb,
nfs_set_local_verifier(hdr->inode, hdr->res.verf, hdr->args.stable);
+ if (nfs_local_p2pdma_misaligned(iocb)) {
+ nfs_local_pgio_done(iocb, -EINVAL);
+ nfs_local_pgio_release(iocb);
+ return;
+ }
+
INIT_WORK(&iocb->work, nfs_local_call_write);
if (nfs_local_defer_io())
queue_work(nfslocaliod_workqueue, &iocb->work);
diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c
index ec535ed3051d..2d1d07c460fd 100644
--- a/fs/nfs/pagelist.c
+++ b/fs/nfs/pagelist.c
@@ -1023,6 +1023,9 @@ static int nfs_generic_pg_pgios(struct nfs_pageio_descriptor *desc)
&hdr->args.context->nfl,
hdr->args.context->mode);
+ if (localio && hdr->args.p2pdma)
+ localio = nfs_local_p2pdma_check(localio);
+
if (NFS_SERVER(hdr->inode)->nfs_client->cl_minorversion)
task_flags = RPC_TASK_MOVEABLE;
ret = nfs_initiate_pgio(NFS_CLIENT(hdr->inode),
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
` (6 preceding siblings ...)
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
@ 2026-10-06 23:32 ` Pranjal Shrivastava
7 siblings, 0 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
To: linux-nfs, Trond Myklebust, Anna Schumaker
Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
Pranjal Shrivastava
Pass ITER_ALLOW_P2PDMA when extracting O_DIRECT pages on mounts that
use the RDMA transport and not pNFS. The RDMA transport moves such pages
only by DMA, and fails the RPC with -EREMOTEIO if the device cannot
reach them.
On other mounts, extracting P2PDMA pages still fails with -EREMOTEIO.
Signed-off-by: Pranjal Shrivastava <praan@google.com>
---
fs/nfs/direct.c | 10 +++++++++-
1 file changed, 9 insertions(+), 1 deletion(-)
diff --git a/fs/nfs/direct.c b/fs/nfs/direct.c
index a3c6e8f4ea06..dbeb4f8274e3 100644
--- a/fs/nfs/direct.c
+++ b/fs/nfs/direct.c
@@ -157,14 +157,22 @@ static ssize_t nfs_direct_extract_pages(struct nfs_direct_req *dreq,
size_t size, loff_t *pos,
struct list_head *list)
{
+ struct nfs_server *server = NFS_SERVER(dreq->inode);
bool pinned = iov_iter_extract_will_pin(iter);
+ iov_iter_extraction_t flags = 0;
struct page **pagevec = NULL;
ssize_t result, bytes = 0;
int err = 0;
unsigned int npages, i;
size_t pgbase;
- result = iov_iter_extract_pages(iter, &pagevec, size, ~0U, 0, &pgbase);
+ /* Allow P2PDMA pages only on RDMA mounts that do not use pNFS */
+ if (server->nfs_client->cl_proto == XPRT_TRANSPORT_RDMA &&
+ !pnfs_enabled_sb(server))
+ flags = ITER_ALLOW_P2PDMA;
+
+ result = iov_iter_extract_pages(iter, &pagevec, size, ~0U, flags,
+ &pgbase);
if (result <= 0)
return result;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-10-06 23:33 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®