mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA
@ 2026-10-06 23:32 Pranjal Shrivastava
  2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
                   ` (7 more replies)
  0 siblings, 8 replies; 9+ messages in thread
From: Pranjal Shrivastava @ 2026-10-06 23:32 UTC (permalink / raw)
  To: linux-nfs, Trond Myklebust, Anna Schumaker
  Cc: Chuck Lever, Jeff Layton, linux-kernel, Christoph Hellwig,
	Logan Gunthorpe, Jason Gunthorpe, linux-pci, linux-rdma,
	Shivaji Kant, Tom Talpey, Leon Romanovsky, Unnati Sachan,
	Pranjal Shrivastava

Enable NFS O_DIRECT to and from PCI peer-to-peer DMA (P2PDMA) memory
on RDMA mounts. This replaces the initial RFC [1] and builds on the
Direct I/O modernization series [2].

P2PDMA memory can be an NVMe Controller Memory Buffer, or the BAR of a
vfio-pci device exported as a ZONE_DEVICE-backed DMABUF, as proposed
in [3]. Today, NFS O_DIRECT fails with -EREMOTEIO when the user buffer
is such memory, since iov_iter_extract_pages() refuses P2PDMA pages
unless the caller passes ITER_ALLOW_P2PDMA. Buffered I/O is not
affected, as the CPU copies the data into the page cache.

P2PDMA memory may not be safe to access with the CPU. Thus, the series
follows one rule: a P2PDMA payload moves only by DMA between the device
memory and the NIC (or the local disk with LOCALIO). Any path that
would touch the payload via CPU accesses refuses the I/O instead of
copying it.

Design
======

  O_DIRECT, user buffer in P2PDMA memory
      |
  nfs_direct_extract_pages()  ITER_ALLOW_P2PDMA on RDMA, !pNFS mounts
      |                       (other mounts: -EREMOTEIO, as today)
  nfs_generic_pgio()          is_pci_p2pdma_page() -> args.p2pdma
      |
      +-- LOCALIO: disk lacks BLK_FEAT_PCI_P2PDMA  -> regular RPC
      |            not one DIO-aligned segment     -> -EINVAL
      |
      +-- RPC: krb5i / krb5p                       -> -EREMOTEIO
               encoders set XDRBUF_P2PDMA
                 |
                 +-- xprtsock                      -> -EREMOTEIO
                 +-- xprtrdma: NIC can't do P2P    -> -EREMOTEIO
                               NIC can't reach it  -> -EREMOTEIO
                               otherwise: Read/Write chunks only

1. Per-RPC decision
ITER_ALLOW_P2PDMA at extraction is only a hint. NFS tags each READ and
WRITE whose pages are P2PDMA memory with XDRBUF_P2PDMA, and the
transport serving that RPC decides whether it can move them. With
nconnect, reconnects and migration, the transport and its device can
change between extraction and transmission, so a mount-time capability
cannot be trusted. LOCALIO does not use the NIC at all.

2. Socket transports
TCP (including TLS), UDP and AF_LOCAL move data via CPU accesses, so
they return -EREMOTEIO for a tagged RPC before touching the socket.
xprt_transmit() now fails only the refused RPC and keeps draining the
transmit queue.

3. RPC-over-RDMA
The inline paths copy the payload via CPU accesses or DMA-map it with
calls that do not handle P2PDMA pages. Thus, xprtrdma always sends a
P2PDMA WRITE payload in a Read chunk and receives a P2PDMA READ payload
in a Write chunk, whatever its size. It returns -EREMOTEIO when:
 - the NIC cannot DMA to PCI peer-to-peer memory (e.g. rxe, siw),
 - the NIC cannot reach the P2PDMA memory (DMA mapping fails),
 - the GSS service forbids direct data placement (krb5i, krb5p), or
 - the rest of the READ reply does not fit inline.
If a server returns READ data inline despite the Write chunk, which
RFC 8166 forbids, the RPC fails instead of copying into the pages.
Small P2PDMA I/O therefore always pays for memory registration.

4. GSS
krb5i checksums and krb5p encrypts the payload via CPU accesses while
the RPC is encoded, before the transport sees it. nfs_pgio_prepare()
fails such READs and WRITEs with -EREMOTEIO before encoding.

5. LOCALIO
LOCALIO is used for P2PDMA pages only if the filesystem's block device
supports P2PDMA. Otherwise, the I/O goes out as a regular RPC. I/O that
is not a single DIO-aligned segment would fall back to buffered I/O,
so it fails with -EINVAL instead.

6. READ_PLUS
READ_PLUS decoding moves data and zero-fills holes via CPU accesses, so
it is never used for P2PDMA pages. RDMA mounts already disable it.

All refusals return -EREMOTEIO, the error iov_iter_extract_pages()
returns today, except misaligned LOCALIO I/O, which returns -EINVAL.
Refused I/O fails at once, even on hard mounts.

Call for review
===============
Feedback is appreciated on:

1. Layering: "RDMA transport, !pNFS" is only a hint, and each RPC
   re-checks its transport. Should the transport advertise P2PDMA
   capability up front instead? That would also allow LOCALIO on TCP
   mounts, which this series refuses.
2. nconnect: fail an RPC on a transport without P2PDMA, or pick a
   capable transport?
3. Migration / reconnect: the path can move to a NIC without P2PDMA,
   and I/O that worked starts failing with -EREMOTEIO. Acceptable?
4. Hard mounts: refused I/O fails at once instead of retrying.
   Acceptable?
5. LOCALIO: the check uses sb->s_bdev, which misses multi-device
   filesystems. Is there a better way to ask whether a file can do
   P2PDMA? Should misaligned I/O fall back to an RPC instead of
   failing with -EINVAL?
6. pNFS: excluded for now. Route P2PDMA I/O through the MDS, or let
   each layout driver opt in?

Known issue: the pNFS check happens at extraction. If a mount gains
pNFS afterwards (e.g. after migration), a rescheduled WRITE could reach
a layout driver. I plan to decide once per O_DIRECT call and force the
MDS path when P2PDMA pages are allowed.

Thanks,
Praan

[1] https://lore.kernel.org/all/20260401194501.2269200-1-praan@google.com/
[2] https://lore.kernel.org/all/20260814143255.861084-1-praan@google.com/
[3] https://lore.kernel.org/all/20260804185050.2053672-1-praan@google.com/

Changes since the initial RFC [1]:
 - Split the Direct I/O modernization into its own series [2].
 - Drop NFS_CAP_P2PDMA, the 64-bit caps expansion and the
   ->supports_p2pdma xprt op. Decide per RPC instead, as Chuck
   suggested.
 - Refuse P2PDMA pages in the socket transports, and fail only the
   refused RPC.
 - Move P2PDMA payloads only via chunks in xprtrdma, and return
   -EREMOTEIO when the NIC cannot reach the memory.
 - Refuse krb5i and krb5p, and allow LOCALIO only as aligned direct
   I/O to a P2PDMA-capable disk.

Pranjal Shrivastava (8):
  sunrpc: introduce XDRBUF_P2PDMA flag
  sunrpc: fail only the RPC a transport refuses
  xprtrdma: move P2PDMA payloads only via chunks
  xprtrdma: return -EREMOTEIO on P2PDMA map failure
  nfs: tag READ/WRITE RPCs carrying P2PDMA pages
  nfs: refuse P2PDMA pages with krb5i and krb5p
  nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O
  nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts

 fs/nfs/direct.c                | 10 ++++++++-
 fs/nfs/internal.h              |  8 +++++++
 fs/nfs/localio.c               | 40 ++++++++++++++++++++++++++++++++++
 fs/nfs/nfs2xdr.c               |  4 ++++
 fs/nfs/nfs3xdr.c               |  4 ++++
 fs/nfs/nfs4proc.c              |  3 +++
 fs/nfs/nfs4xdr.c               |  4 ++++
 fs/nfs/pagelist.c              | 13 +++++++++++
 include/linux/nfs_xdr.h        |  1 +
 include/linux/sunrpc/xdr.h     |  1 +
 include/linux/sunrpc/xprt.h    |  5 +++++
 net/sunrpc/xprt.c              |  9 ++++++++
 net/sunrpc/xprtrdma/frwr_ops.c | 15 ++++++++-----
 net/sunrpc/xprtrdma/rpc_rdma.c | 40 ++++++++++++++++++++++++++++++++--
 net/sunrpc/xprtsock.c          | 12 ++++++++++
 15 files changed, 161 insertions(+), 8 deletions(-)

-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-06 23:33 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 23:32 [RFC PATCH v1 0/8] nfs: PCI P2PDMA for O_DIRECT over RDMA Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 1/8] sunrpc: introduce XDRBUF_P2PDMA flag Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 2/8] sunrpc: fail only the RPC a transport refuses Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 3/8] xprtrdma: move P2PDMA payloads only via chunks Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 4/8] xprtrdma: return -EREMOTEIO on P2PDMA map failure Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 5/8] nfs: tag READ/WRITE RPCs carrying P2PDMA pages Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 6/8] nfs: refuse P2PDMA pages with krb5i and krb5p Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 7/8] nfs: allow LOCALIO for P2PDMA pages only as aligned direct I/O Pranjal Shrivastava
2026-10-06 23:32 ` [RFC PATCH v1 8/8] nfs: allow P2PDMA pages for O_DIRECT on RDMA mounts Pranjal Shrivastava

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®