mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Mina Almasry <almasrymina@google.com>
To: netdev@vger.kernel.org, linux-doc@vger.kernel.org,
	 linux-kernel@vger.kernel.org, bpf@vger.kernel.org
Cc: "Mina Almasry" <almasrymina@google.com>,
	"David S. Miller" <davem@davemloft.net>,
	"Eric Dumazet" <edumazet@kernel.org>,
	"Jakub Kicinski" <kuba@kernel.org>,
	"Paolo Abeni" <pabeni@redhat.com>,
	"Simon Horman" <horms@kernel.org>,
	"Jonathan Corbet" <corbet@lwn.net>,
	"Shuah Khan" <skhan@linuxfoundation.org>,
	"Randy Dunlap" <rdunlap@infradead.org>,
	"Jesper Dangaard Brouer" <hawk@kernel.org>,
	"Ilias Apalodimas" <ilias.apalodimas@linaro.org>,
	"Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"John Fastabend" <john.fastabend@gmail.com>,
	"Stanislav Fomichev" <sdf@fomichev.me>,
	"Luigi Rizzo" <lrizzo@google.com>,
	"Björn Töpel" <bjorn@kernel.org>,
	"Pavel Begunkov" <asml.silence@gmail.com>
Subject: [PATCH net-next v3 1/2] net: netmem: document netmem and memory provider design in comments
Date: Sat, 10 Oct 2026 03:36:08 +0000	[thread overview]
Message-ID: <20261010033630.1171692-2-almasrymina@google.com> (raw)
In-Reply-To: <20261010033630.1171692-1-almasrymina@google.com>

Clarify the netmem, memory provider, page_pool, and skb fragment design
principles in header and code comments:

- Memory providers allocate struct net_iov or struct page, cast them to
  netmem_ref, and pass them to page_pool; page_pool, drivers, and the
  core stack operate on netmem_ref and must not downcast to page or
  net_iov outside dedicated netmem helpers.
- While current memory providers supply struct net_iov, they are not
  architecturally restricted to net_iov and may supply page-backed
  netmems.
- While current net_iov types are CPU-unreadable, net_iov has no
  inherent restrictions and may be readable or unreadable.
- New code should generalize existing limitations as much as possible to
  match these design principles.
- Per-provider logic belongs in memory_provider_ops, and per-netmem-type
  logic belongs in netmem helpers.
- All frags in an skb must either all be struct page-backed or all
  belong to the same memory provider instance; mixing them (including
  via coalescing) is not allowed.

Cc: Luigi Rizzo <lrizzo@google.com>
Cc: Pavel Begunkov <asml.silence@gmail.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
Reviewed-by: Björn Töpel <bjorn@kernel.org>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
v3:
- Collect Björn's Reviewed-by and Stanislav's Acked-by tags.
- Clarify skb fragment comments: all frags must be struct page-backed or
  belong to the same memory provider instance (Pavel Begunkov).
- Link to v2: https://lore.kernel.org/netdev/20261008023030.1089616-1-almasrymina@google.com/
v2:
- Add Jesper's Acked-by tag.
- Note current implementation status (memory providers currently supply
  net_iov, and net_iov is currently unreadable) alongside target design
  principles and expectation for new code to generalize existing
  limitations (Stanislav Fomichev).
- Link to v1: https://lore.kernel.org/netdev/20261005004958.3603059-1-almasrymina@google.com/
---
 include/linux/skbuff.h                  |  6 +++++
 include/net/netmem.h                    | 30 ++++++++++++++++---------
 include/net/page_pool/helpers.h         | 18 ++++++++++-----
 include/net/page_pool/memory_provider.h | 10 +++++++++
 include/net/page_pool/types.h           |  6 ++---
 net/core/skbuff.c                       |  4 ++++
 6 files changed, 56 insertions(+), 18 deletions(-)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 27ec1e38c8283..af1cbc2799b77 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -358,6 +358,12 @@ struct sk_buff;
  */
 #define GSO_BY_FRAGS	0xFFFF
 
+/* All frags in an skb must either all be struct page-backed or all belong to
+ * the same memory provider instance (e.g. the same devmem binding or io_uring
+ * area). Mixing page and memory-provider frags, or mixing frags from different
+ * memory providers within a single skb (including via coalescing), is not
+ * allowed.
+ */
 typedef struct skb_frag {
 	netmem_ref netmem;
 	unsigned int len;
diff --git a/include/net/netmem.h b/include/net/netmem.h
index 0cc8572f62caf..c4beb6fbc1306 100644
--- a/include/net/netmem.h
+++ b/include/net/netmem.h
@@ -70,16 +70,18 @@ enum net_iov_type {
 	NET_IOV_IOURING,
 };
 
-/* A memory descriptor representing abstract networking I/O vectors,
- * generally for non-pages memory that doesn't have its corresponding
- * struct page and needs to be explicitly allocated through slab.
+/* A memory descriptor representing abstract networking I/O vectors.
  *
  * net_iovs are allocated and used by networking code, and the size of
  * the chunk is PAGE_SIZE.
  *
- * This memory can be any form of non-struct paged memory.  Examples
- * include imported dmabuf memory and imported io_uring memory.  See
- * net_iov_type for all the supported types.
+ * Examples include imported dmabuf memory and imported io_uring memory. See
+ * net_iov_type for all the supported types. While current net_iov types are
+ * unreadable by the CPU, net_iov has no inherent restrictions and may be
+ * CPU-readable or unreadable. New code must not assume net_iov implies
+ * unreadable memory (check readability via netmem_address() or
+ * skb_frags_readable() instead) and should, as much as possible, generalize
+ * existing limitations to match the design principles.
  *
  * @pp_magic:	pp field, similar to the one in struct page/struct
  *		netmem_desc.
@@ -131,8 +133,17 @@ static inline void net_iov_init(struct net_iov *niov,
  * network memory.
  *
  * A netmem_ref can be a struct page* or a struct net_iov* underneath.
+ * Memory providers (or the default page_pool allocator) allocate struct
+ * net_iov or struct page, cast them to netmem_ref, and hand them to
+ * page_pool.
  *
- * Use the supplied helpers to obtain the underlying memory pointer and fields.
+ * The page_pool, drivers, and core networking stack should operate on
+ * netmem_ref rather than struct page or struct net_iov. Downcasting
+ * netmem_ref via netmem_to_page() or netmem_to_net_iov() in callers is
+ * not allowed unless a code path strictly requires a specific backing
+ * type (e.g., kmap_local_page()). In such cases, add a netmem helper here
+ * that handles both page and net_iov cases and returns an error if the
+ * underlying type cannot support the operation.
  */
 typedef unsigned long __bitwise netmem_ref;
 
@@ -297,9 +308,8 @@ static inline atomic_long_t *netmem_get_pp_ref_count_ref(netmem_ref netmem)
 
 static inline bool netmem_is_pref_nid(netmem_ref netmem, int pref_nid)
 {
-	/* NUMA node preference only makes sense if we're allocating
-	 * system memory. Memory providers (which give us net_iovs)
-	 * choose for us.
+	/* NUMA node preference only applies to struct page; net_iovs are
+	 * managed by their memory provider.
 	 */
 	if (netmem_is_net_iov(netmem))
 		return true;
diff --git a/include/net/page_pool/helpers.h b/include/net/page_pool/helpers.h
index cd021832c3fa3..28635ce9454e8 100644
--- a/include/net/page_pool/helpers.h
+++ b/include/net/page_pool/helpers.h
@@ -8,12 +8,20 @@
 /**
  * DOC: page_pool allocator
  *
- * The page_pool allocator is optimized for recycling page or page fragment used
- * by skb packet and xdp frame.
+ * The page_pool allocator is optimized for recycling network memory
+ * (netmem_ref) or fragments used by skb packets and xdp frames.
  *
- * Basic use involves replacing any alloc_pages() calls with page_pool_alloc(),
- * which allocate memory with or without page splitting depending on the
- * requested memory size.
+ * page_pool natively operates on netmem_ref, which abstracts the underlying
+ * memory type (struct page or struct net_iov) supplied by the page allocator
+ * or a memory provider. Drivers and core networking code should use the
+ * netmem-based APIs (e.g. page_pool_alloc_netmem(), page_pool_put_netmem()).
+ * The struct page-based APIs (e.g. page_pool_alloc(), page_pool_alloc_pages(),
+ * page_pool_put_page()) are legacy compatibility wrappers for drivers not yet
+ * converted to netmem.
+ *
+ * Basic use involves replacing any alloc_pages() calls with
+ * page_pool_alloc_netmem() (or legacy page_pool_alloc()), which allocate memory
+ * with or without splitting depending on the requested memory size.
  *
  * If the driver knows that it always requires full pages or its allocations are
  * always smaller than half a page, it can use one of the more specific API
diff --git a/include/net/page_pool/memory_provider.h b/include/net/page_pool/memory_provider.h
index 255ce4cfd9755..137cfc50833ac 100644
--- a/include/net/page_pool/memory_provider.h
+++ b/include/net/page_pool/memory_provider.h
@@ -9,6 +9,16 @@ struct netdev_rx_queue;
 struct netlink_ext_ack;
 struct sk_buff;
 
+/* Memory providers allocate underlying memory (struct net_iov or struct page),
+ * cast it to netmem_ref, and supply it to page_pool. While current memory
+ * providers only return struct net_iov, they are not architecturally limited to
+ * net_iov; a provider returning page-backed netmems is allowed. New code must
+ * not assume a memory provider implies net_iov and should, as much as possible,
+ * generalize existing limitations to match the design principles.
+ *
+ * Per-provider custom logic must be delegated to memory_provider_ops rather
+ * than handled directly in page_pool core code.
+ */
 struct memory_provider_ops {
 	netmem_ref (*alloc_netmems)(struct page_pool *pool, gfp_t gfp);
 	bool (*release_netmem)(struct page_pool *pool, netmem_ref netmem);
diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
index 03da138722f58..6d543076a0c0b 100644
--- a/include/net/page_pool/types.h
+++ b/include/net/page_pool/types.h
@@ -22,9 +22,9 @@
 					*/
 #define PP_FLAG_SYSTEM_POOL	BIT(2) /* Global system page_pool */
 
-/* Allow unreadable (net_iov backed) netmem in this page_pool. Drivers setting
- * this must be able to support unreadable netmem, where netmem_address() would
- * return NULL. This flag should not be set for header page_pools.
+/* Allow unreadable netmem in this page_pool. Drivers setting this must be able
+ * to support unreadable netmem, where netmem_address() returns NULL. This flag
+ * should not be set for header page_pools.
  *
  * If the driver sets PP_FLAG_ALLOW_UNREADABLE_NETMEM, it should also set
  * page_pool_params.slow.queue_idx.
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index d56f4f0102f75..69d3d27e29d97 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -6210,6 +6210,10 @@ bool skb_try_coalesce(struct sk_buff *to, struct sk_buff *from,
 	if (to->pp_recycle != from->pp_recycle)
 		return false;
 
+	/* All frags in an skb must either all be struct page-backed or all
+	 * belong to the same memory provider instance; do not coalesce skbs
+	 * that would mix them.
+	 */
 	if (skb_frags_readable(from) != skb_frags_readable(to))
 		return false;
 
-- 
2.56.0.385.gd3acb90ef8-goog


  reply	other threads:[~2026-10-10  3:36 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-10  3:36 [PATCH net-next v3 0/2] net: netmem: document design principles and intended direction Mina Almasry
2026-10-10  3:36 ` Mina Almasry [this message]
2026-10-11  3:51   ` [PATCH net-next v3 1/2] net: netmem: document netmem and memory provider design in comments netdev-bot+sashiko
2026-10-10  3:36 ` [PATCH net-next v3 2/2] docs: netmem: document netmem and memory provider design principles Mina Almasry
2026-10-11  3:51   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261010033630.1171692-2-almasrymina@google.com \
    --to=almasrymina@google.com \
    --cc=asml.silence@gmail.com \
    --cc=ast@kernel.org \
    --cc=bjorn@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=corbet@lwn.net \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=edumazet@kernel.org \
    --cc=hawk@kernel.org \
    --cc=horms@kernel.org \
    --cc=ilias.apalodimas@linaro.org \
    --cc=john.fastabend@gmail.com \
    --cc=kuba@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lrizzo@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=rdunlap@infradead.org \
    --cc=sdf@fomichev.me \
    --cc=skhan@linuxfoundation.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®