mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v2 0/2] net: netmem: document design principles and intended direction
@ 2026-10-08  2:30 Mina Almasry
  2026-10-08  2:30 ` [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments Mina Almasry
  2026-10-08  2:30 ` [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles Mina Almasry
  0 siblings, 2 replies; 5+ messages in thread
From: Mina Almasry @ 2026-10-08  2:30 UTC (permalink / raw)
  To: netdev, linux-doc, linux-kernel, bpf
  Cc: Mina Almasry, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Jesper Dangaard Brouer, Ilias Apalodimas,
	Alexei Starovoitov, Daniel Borkmann, John Fastabend,
	Stanislav Fomichev, Luigi Rizzo, Björn Töpel,
	Pavel Begunkov

Because the networking stack and drivers are only partially converted to
netmem_ref and current memory providers only supply unreadable net_iovs,
automated code review and analysis tools (such as LLMs) frequently infer
the wrong architectural invariants from existing code. Specifically,
they often assume that page_pool's struct page APIs are primary rather
than legacy wrappers, that memory providers always imply net_iov, that
net_iov is inherently unreadable, or that callers should branch on or
downcast netmem_ref directly.

Document the intended netmem, memory provider, page_pool, and skb
fragment design principles concisely in the relevant headers, code
comments, and netmem documentation so both developers and automated
tools follow the intended abstractions. This series adds comments and
documentation rather than performing a large refactor all at once, so
that future incremental changes nudge the codebase in the intended
direction.

This reflects my mental model of the netmem, page_pool, and memory
provider architecture; I welcome feedback and disagreements,
particularly from major contributors to page_pool and memory providers.

Cc: Luigi Rizzo <lrizzo@google.com>
Cc: Björn Töpel <bjorn@kernel.org>
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Pavel Begunkov <asml.silence@gmail.com>
---
v2:
- Document both current implementation status (memory providers currently
  supply net_iov, and net_iov is currently unreadable) and target design
  principles, and clarify that new code should generalize existing
  limitations as much as possible (Stanislav Fomichev).
- Link to v1: https://lore.kernel.org/netdev/20261005004958.3603059-1-almasrymina@google.com/

Mina Almasry (2):
  net: netmem: document netmem and memory provider design in comments
  docs: netmem: document netmem and memory provider design principles

 Documentation/networking/netmem.rst     | 52 +++++++++++++++++++++++++
 include/linux/skbuff.h                  |  4 ++
 include/net/netmem.h                    | 30 +++++++++-----
 include/net/page_pool/helpers.h         | 18 ++++++---
 include/net/page_pool/memory_provider.h | 10 +++++
 include/net/page_pool/types.h           |  6 +--
 net/core/skbuff.c                       |  3 ++
 7 files changed, 105 insertions(+), 18 deletions(-)


base-commit: 8df0638138d3e0344fd1fb36cf2d1ca1cf5028f0
-- 
2.56.0.385.gd3acb90ef8-goog


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments
  2026-10-08  2:30 [PATCH net-next v2 0/2] net: netmem: document design principles and intended direction Mina Almasry
@ 2026-10-08  2:30 ` Mina Almasry
  2026-10-08 11:29   ` Björn Töpel
  2026-10-08  2:30 ` [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles Mina Almasry
  1 sibling, 1 reply; 5+ messages in thread
From: Mina Almasry @ 2026-10-08  2:30 UTC (permalink / raw)
  To: netdev, linux-doc, linux-kernel, bpf
  Cc: Mina Almasry, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Jesper Dangaard Brouer, Ilias Apalodimas,
	Alexei Starovoitov, Daniel Borkmann, John Fastabend,
	Stanislav Fomichev, Luigi Rizzo, Björn Töpel,
	Pavel Begunkov

Clarify the netmem, memory provider, page_pool, and skb fragment design
principles in header and code comments:

- Memory providers allocate struct net_iov or struct page, cast them to
  netmem_ref, and pass them to page_pool; page_pool, drivers, and the
  core stack operate on netmem_ref and must not downcast to page or
  net_iov outside dedicated netmem helpers.
- While current memory providers supply struct net_iov, they are not
  architecturally restricted to net_iov and may supply page-backed
  netmems.
- While current net_iov types are CPU-unreadable, net_iov has no
  inherent restrictions and may be readable or unreadable.
- New code should generalize existing limitations as much as possible to
  match these design principles.
- Per-provider logic belongs in memory_provider_ops, and per-netmem-type
  logic belongs in netmem helpers.
- All frags in an skb must share the same backing netmem memory type,
  and skbs with different frag memory types must not be coalesced.

Cc: Luigi Rizzo <lrizzo@google.com>
Cc: Björn Töpel <bjorn@kernel.org>
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Pavel Begunkov <asml.silence@gmail.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Acked-by: Jesper Dangaard Brouer <hawk@kernel.org>
---
v2:
- Add Jesper's Acked-by tag.
- Note current implementation status (memory providers currently supply
  net_iov, and net_iov is currently unreadable) alongside target design
  principles and expectation for new code to generalize existing
  limitations (Stanislav Fomichev).
- Link to v1: https://lore.kernel.org/netdev/20261005004958.3603059-1-almasrymina@google.com/
---
 include/linux/skbuff.h                  |  4 ++++
 include/net/netmem.h                    | 30 ++++++++++++++++---------
 include/net/page_pool/helpers.h         | 18 ++++++++++-----
 include/net/page_pool/memory_provider.h | 10 +++++++++
 include/net/page_pool/types.h           |  6 ++---
 net/core/skbuff.c                       |  3 +++
 6 files changed, 53 insertions(+), 18 deletions(-)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 27ec1e38c8283..c022e5fd3124f 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -358,6 +358,10 @@ struct sk_buff;
  */
 #define GSO_BY_FRAGS	0xFFFF
 
+/* All fragments in an skb (skb_shinfo(skb)->frags[]) must be backed by
+ * netmems of the same memory type. Mixing fragments of different memory types
+ * within a single skb (including via skb coalescing) is not allowed.
+ */
 typedef struct skb_frag {
 	netmem_ref netmem;
 	unsigned int len;
diff --git a/include/net/netmem.h b/include/net/netmem.h
index 0cc8572f62caf..c4beb6fbc1306 100644
--- a/include/net/netmem.h
+++ b/include/net/netmem.h
@@ -70,16 +70,18 @@ enum net_iov_type {
 	NET_IOV_IOURING,
 };
 
-/* A memory descriptor representing abstract networking I/O vectors,
- * generally for non-pages memory that doesn't have its corresponding
- * struct page and needs to be explicitly allocated through slab.
+/* A memory descriptor representing abstract networking I/O vectors.
  *
  * net_iovs are allocated and used by networking code, and the size of
  * the chunk is PAGE_SIZE.
  *
- * This memory can be any form of non-struct paged memory.  Examples
- * include imported dmabuf memory and imported io_uring memory.  See
- * net_iov_type for all the supported types.
+ * Examples include imported dmabuf memory and imported io_uring memory. See
+ * net_iov_type for all the supported types. While current net_iov types are
+ * unreadable by the CPU, net_iov has no inherent restrictions and may be
+ * CPU-readable or unreadable. New code must not assume net_iov implies
+ * unreadable memory (check readability via netmem_address() or
+ * skb_frags_readable() instead) and should, as much as possible, generalize
+ * existing limitations to match the design principles.
  *
  * @pp_magic:	pp field, similar to the one in struct page/struct
  *		netmem_desc.
@@ -131,8 +133,17 @@ static inline void net_iov_init(struct net_iov *niov,
  * network memory.
  *
  * A netmem_ref can be a struct page* or a struct net_iov* underneath.
+ * Memory providers (or the default page_pool allocator) allocate struct
+ * net_iov or struct page, cast them to netmem_ref, and hand them to
+ * page_pool.
  *
- * Use the supplied helpers to obtain the underlying memory pointer and fields.
+ * The page_pool, drivers, and core networking stack should operate on
+ * netmem_ref rather than struct page or struct net_iov. Downcasting
+ * netmem_ref via netmem_to_page() or netmem_to_net_iov() in callers is
+ * not allowed unless a code path strictly requires a specific backing
+ * type (e.g., kmap_local_page()). In such cases, add a netmem helper here
+ * that handles both page and net_iov cases and returns an error if the
+ * underlying type cannot support the operation.
  */
 typedef unsigned long __bitwise netmem_ref;
 
@@ -297,9 +308,8 @@ static inline atomic_long_t *netmem_get_pp_ref_count_ref(netmem_ref netmem)
 
 static inline bool netmem_is_pref_nid(netmem_ref netmem, int pref_nid)
 {
-	/* NUMA node preference only makes sense if we're allocating
-	 * system memory. Memory providers (which give us net_iovs)
-	 * choose for us.
+	/* NUMA node preference only applies to struct page; net_iovs are
+	 * managed by their memory provider.
 	 */
 	if (netmem_is_net_iov(netmem))
 		return true;
diff --git a/include/net/page_pool/helpers.h b/include/net/page_pool/helpers.h
index cd021832c3fa3..28635ce9454e8 100644
--- a/include/net/page_pool/helpers.h
+++ b/include/net/page_pool/helpers.h
@@ -8,12 +8,20 @@
 /**
  * DOC: page_pool allocator
  *
- * The page_pool allocator is optimized for recycling page or page fragment used
- * by skb packet and xdp frame.
+ * The page_pool allocator is optimized for recycling network memory
+ * (netmem_ref) or fragments used by skb packets and xdp frames.
  *
- * Basic use involves replacing any alloc_pages() calls with page_pool_alloc(),
- * which allocate memory with or without page splitting depending on the
- * requested memory size.
+ * page_pool natively operates on netmem_ref, which abstracts the underlying
+ * memory type (struct page or struct net_iov) supplied by the page allocator
+ * or a memory provider. Drivers and core networking code should use the
+ * netmem-based APIs (e.g. page_pool_alloc_netmem(), page_pool_put_netmem()).
+ * The struct page-based APIs (e.g. page_pool_alloc(), page_pool_alloc_pages(),
+ * page_pool_put_page()) are legacy compatibility wrappers for drivers not yet
+ * converted to netmem.
+ *
+ * Basic use involves replacing any alloc_pages() calls with
+ * page_pool_alloc_netmem() (or legacy page_pool_alloc()), which allocate memory
+ * with or without splitting depending on the requested memory size.
  *
  * If the driver knows that it always requires full pages or its allocations are
  * always smaller than half a page, it can use one of the more specific API
diff --git a/include/net/page_pool/memory_provider.h b/include/net/page_pool/memory_provider.h
index 255ce4cfd9755..137cfc50833ac 100644
--- a/include/net/page_pool/memory_provider.h
+++ b/include/net/page_pool/memory_provider.h
@@ -9,6 +9,16 @@ struct netdev_rx_queue;
 struct netlink_ext_ack;
 struct sk_buff;
 
+/* Memory providers allocate underlying memory (struct net_iov or struct page),
+ * cast it to netmem_ref, and supply it to page_pool. While current memory
+ * providers only return struct net_iov, they are not architecturally limited to
+ * net_iov; a provider returning page-backed netmems is allowed. New code must
+ * not assume a memory provider implies net_iov and should, as much as possible,
+ * generalize existing limitations to match the design principles.
+ *
+ * Per-provider custom logic must be delegated to memory_provider_ops rather
+ * than handled directly in page_pool core code.
+ */
 struct memory_provider_ops {
 	netmem_ref (*alloc_netmems)(struct page_pool *pool, gfp_t gfp);
 	bool (*release_netmem)(struct page_pool *pool, netmem_ref netmem);
diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
index 03da138722f58..6d543076a0c0b 100644
--- a/include/net/page_pool/types.h
+++ b/include/net/page_pool/types.h
@@ -22,9 +22,9 @@
 					*/
 #define PP_FLAG_SYSTEM_POOL	BIT(2) /* Global system page_pool */
 
-/* Allow unreadable (net_iov backed) netmem in this page_pool. Drivers setting
- * this must be able to support unreadable netmem, where netmem_address() would
- * return NULL. This flag should not be set for header page_pools.
+/* Allow unreadable netmem in this page_pool. Drivers setting this must be able
+ * to support unreadable netmem, where netmem_address() returns NULL. This flag
+ * should not be set for header page_pools.
  *
  * If the driver sets PP_FLAG_ALLOW_UNREADABLE_NETMEM, it should also set
  * page_pool_params.slow.queue_idx.
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 43ebe61c7fc48..2c42a218dcc00 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -6206,6 +6206,9 @@ bool skb_try_coalesce(struct sk_buff *to, struct sk_buff *from,
 	if (to->pp_recycle != from->pp_recycle)
 		return false;
 
+	/* All frags in an skb must have the same backing netmem memory type;
+	 * do not coalesce skbs with different frag memory types.
+	 */
 	if (skb_frags_readable(from) != skb_frags_readable(to))
 		return false;
 
-- 
2.56.0.385.gd3acb90ef8-goog


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles
  2026-10-08  2:30 [PATCH net-next v2 0/2] net: netmem: document design principles and intended direction Mina Almasry
  2026-10-08  2:30 ` [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments Mina Almasry
@ 2026-10-08  2:30 ` Mina Almasry
  2026-10-08 11:33   ` Björn Töpel
  1 sibling, 1 reply; 5+ messages in thread
From: Mina Almasry @ 2026-10-08  2:30 UTC (permalink / raw)
  To: netdev, linux-doc, linux-kernel, bpf
  Cc: Mina Almasry, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Jesper Dangaard Brouer, Ilias Apalodimas,
	Alexei Starovoitov, Daniel Borkmann, John Fastabend,
	Stanislav Fomichev, Luigi Rizzo, Björn Töpel,
	Pavel Begunkov

Add a Design Principles section to Documentation/networking/netmem.rst
covering the netmem_ref abstraction, the prohibition on direct
downcasting in callers, decoupling memory providers from net_iov,
decoupling net_iov from unreadability, delegating provider/type logic to
memory_provider_ops and netmem helpers, and the homogeneous skb fragment
memory type invariant.

Cc: Luigi Rizzo <lrizzo@google.com>
Cc: Björn Töpel <bjorn@kernel.org>
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Pavel Begunkov <asml.silence@gmail.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
v2:
- Document both current implementation status (mp returns net_iov,
  net_iov is unreadable) and target design principles in items 2 & 3,
  and note that new code should generalize existing limitations as much
  as possible (Stanislav Fomichev).
- Link to v1: https://lore.kernel.org/netdev/20261005004958.3603059-1-almasrymina@google.com/
---
 Documentation/networking/netmem.rst | 52 +++++++++++++++++++++++++++++
 1 file changed, 52 insertions(+)

diff --git a/Documentation/networking/netmem.rst b/Documentation/networking/netmem.rst
index 217869d1108dd..e023f4c69d2a6 100644
--- a/Documentation/networking/netmem.rst
+++ b/Documentation/networking/netmem.rst
@@ -19,6 +19,58 @@ Benefits of Netmem :
 * Simplified Development: Drivers interact with a consistent API,
   regardless of the underlying memory implementation.
 
+Design Principles
+=================
+
+Memory providers (or the default ``page_pool`` allocator) allocate underlying
+memory (``struct net_iov`` or ``struct page``), cast it to ``netmem_ref``, and
+supply it to ``page_pool``. The ``page_pool``, drivers, and networking stack
+operate on ``netmem_ref`` as the abstract type. Existing ``page_pool`` APIs
+that allocate or free ``struct page`` are legacy compatibility wrappers for
+drivers that do not yet support ``netmem_ref``. Code that is not yet
+``netmem``-aware should be converted to ``netmem_ref`` unless it will never
+need to support ``netmem``.
+
+1. **Operate on netmem_ref, do not downcast**: ``page_pool``, drivers, and the
+   core networking stack should deal with ``netmem_ref`` rather than
+   ``struct net_iov`` or ``struct page``. Downcasting ``netmem_ref`` to
+   ``struct net_iov`` or ``struct page`` is not allowed unless a code path
+   strictly cannot function without knowing the underlying memory type (for
+   example, ``kmap_local_page()``). In those cases, to keep call sites simple,
+   add a ``netmem`` helper that performs the operation on behalf of the caller,
+   cleanly handles all ``net_iov`` and ``page`` cases, and returns an error if
+   the ``netmem`` type cannot support the requested operation.
+
+2. **Decouple memory providers from net_iov**: Memory providers are not
+   architecturally limited to ``struct net_iov``; a memory provider that returns
+   ``struct page``-backed ``netmem_ref``\ s to upper layers is allowed. Today,
+   in-tree memory providers only supply ``struct net_iov`` and some existing
+   code still reflects that limitation, but new code must not assume that using
+   a memory provider implies ``net_iov`` memory and should, as much as possible,
+   generalize existing limitations to match the design principles.
+
+3. **Decouple net_iov from unreadability**: ``struct net_iov`` is flexible and
+   has no inherent restrictions; it may represent either CPU-readable or
+   unreadable memory. Today, in-tree ``net_iov`` implementations are unreadable
+   by the CPU (``netmem_address()`` returns ``NULL``) and some existing code
+   still reflects that limitation, but new code must not assume ``net_iov``
+   implies unreadable memory (check readability via ``netmem_address()`` or
+   ``skb_frags_readable()`` instead) and should, as much as possible, generalize
+   existing limitations to match the design principles.
+
+4. **Delegate complexity to the lowest layer**: Each layer must respect its
+   abstraction boundary. ``page_pool`` must not implement per-memory-provider
+   custom logic in its main code; instead, it delegates provider-specific
+   handling to ``struct memory_provider_ops``. Similarly, core networking code
+   should avoid per-``netmem``-type branching and instead delegate operations
+   to ``netmem`` helpers that handle the underlying memory type.
+
+5. **Homogeneous skb fragment memory types**: An ``sk_buff``'s ``frags[]`` are
+   always backed by ``netmem_ref``\ s of the same memory type. Mixing fragments
+   from different memory types within a single ``sk_buff`` is not allowed,
+   keeping ``sk_buff`` handling simple. Consequently, coalescing ``sk_buff``\ s
+   with different fragment memory types must not happen.
+
 Driver RX Requirements
 ======================
 
-- 
2.56.0.385.gd3acb90ef8-goog


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments
  2026-10-08  2:30 ` [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments Mina Almasry
@ 2026-10-08 11:29   ` Björn Töpel
  0 siblings, 0 replies; 5+ messages in thread
From: Björn Töpel @ 2026-10-08 11:29 UTC (permalink / raw)
  To: Mina Almasry, netdev, linux-doc, linux-kernel, bpf
  Cc: Mina Almasry, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Jesper Dangaard Brouer, Ilias Apalodimas,
	Alexei Starovoitov, Daniel Borkmann, John Fastabend,
	Stanislav Fomichev, Luigi Rizzo, Pavel Begunkov

Mina Almasry <almasrymina@google.com> writes:

> Clarify the netmem, memory provider, page_pool, and skb fragment design
> principles in header and code comments:

Thanks, Mina. This is helpful!

Reviewed-by: Björn Töpel <bjorn@kernel.org>


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles
  2026-10-08  2:30 ` [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles Mina Almasry
@ 2026-10-08 11:33   ` Björn Töpel
  0 siblings, 0 replies; 5+ messages in thread
From: Björn Töpel @ 2026-10-08 11:33 UTC (permalink / raw)
  To: Mina Almasry, netdev, linux-doc, linux-kernel, bpf
  Cc: Mina Almasry, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Jesper Dangaard Brouer, Ilias Apalodimas,
	Alexei Starovoitov, Daniel Borkmann, John Fastabend,
	Stanislav Fomichev, Luigi Rizzo, Pavel Begunkov

Mina Almasry <almasrymina@google.com> writes:

> Add a Design Principles section to Documentation/networking/netmem.rst
> covering the netmem_ref abstraction, the prohibition on direct
> downcasting in callers, decoupling memory providers from net_iov,
> decoupling net_iov from unreadability, delegating provider/type logic to
> memory_provider_ops and netmem helpers, and the homogeneous skb fragment
> memory type invariant.
>
> Cc: Luigi Rizzo <lrizzo@google.com>
> Cc: Björn Töpel <bjorn@kernel.org>
> Cc: Stanislav Fomichev <sdf@fomichev.me>
> Cc: Pavel Begunkov <asml.silence@gmail.com>
> Signed-off-by: Mina Almasry <almasrymina@google.com>

This helps (helped me at least!) -- thanks for the clarifications!

Reviewed-by: Björn Töpel <bjorn@kernel.org>

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-08 11:33 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08  2:30 [PATCH net-next v2 0/2] net: netmem: document design principles and intended direction Mina Almasry
2026-10-08  2:30 ` [PATCH net-next v2 1/2] net: netmem: document netmem and memory provider design in comments Mina Almasry
2026-10-08 11:29   ` Björn Töpel
2026-10-08  2:30 ` [PATCH net-next v2 2/2] docs: netmem: document netmem and memory provider design principles Mina Almasry
2026-10-08 11:33   ` Björn Töpel

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®