* [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting
@ 2026-10-10 8:38 Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov() Mina Almasry
` (6 more replies)
0 siblings, 7 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
netmems (pages and net_iovs) have two independent refcounts: a non-pp
refcount (page._refcount or binding->ref) governing backing memory
lifetime, and pp_ref_count governing recycling within a page_pool while
page_pool holds a single backing non-pp reference. SKBs with
pp_recycle=1 own pp_ref_count references on page_pool fragments; SKBs
with pp_recycle=0 own non-pp references.
Core helpers currently mix these two refcounts asymmetrically:
skb_frag_ref() always increments the non-pp refcount via get_netmem(),
whereas skb_frag_unref() decrements pp_ref_count when skb->pp_recycle
is set. Any path that refs a fragment on a pp_recycle=1 SKB (or unrefs
with a hardcoded recycle=false) increments one counter and decrements
the other, leaking the backing memory while underflowing pp_ref_count
(e.g., IP-TFS [1], skb_split(), skb_shift(), and cloned SKB uncloning).
Unclear core APIs also led drivers and ULPs to open-code pp_ref_count
manipulations or add one-off helpers like skb_pp_frag_ref().
Unify fragment refcounting into symmetric pairs where each layer clearly
specifies which refcount it touches:
- Non-PP refcount: get/put_page(), get/put_net_iov(), get/put_netmem()
- PP refcount: napi_pp_get/put_page()
- SKB netmem: skb_netmem_ref/unref(netmem, recycle)
- SKB fragment: skb_frag_ref/unref(skb, f)
[1] https://lore.kernel.org/netdev/xfrm-iptfs-pp_ref_count-underflow-v4-1-912fa72106f0@secunet.com/
Mina Almasry (6):
netmem: rename __get/__put_netmem() to get/put_net_iov()
net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf()
net: skbuff: replace skb_page_unref() and __skb_frag_unref() with
skb_netmem_unref()
net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref()
net: skbuff: use skb_frag_ref() in skb_try_coalesce()
net: kunit: test netmem, page_pool, and skb frag refcounting
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
drivers/net/ethernet/sun/cassini.c | 4 +-
drivers/net/veth.c | 2 +-
include/linux/skbuff_ref.h | 64 +-
include/net/netmem.h | 24 +-
net/core/net_test.c | 747 ++++++++++++++++++
net/core/skbuff.c | 145 ++--
net/ipv4/esp4.c | 4 +-
net/ipv4/tcp.c | 6 +-
net/ipv6/esp6.c | 4 +-
net/tls/tls_device.c | 2 +-
net/tls/tls_device_fallback.c | 2 +-
net/tls/tls_strp.c | 2 +-
net/xfrm/xfrm_iptfs.c | 3 +-
16 files changed, 900 insertions(+), 115 deletions(-)
base-commit: d8674294aefef02266c4d47ad10131f1bffbe534
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov()
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 2/6] net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf() Mina Almasry
` (5 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
netmems carry both a non-pp backing refcount and a page_pool
pp_ref_count. Make the non-pp layer explicit and symmetric across both
netmem types: get/put_netmem() dispatches to get/put_page() for pages
and get/put_net_iov() for net_iovs, always acquiring or releasing a
non-pp reference regardless of page_pool membership.
Warn via DEBUG_NET_WARN_ON_ONCE() for net_iov types that do not support
non-pp references.
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
include/net/netmem.h | 24 ++++++++++++++++++++----
net/core/skbuff.c | 36 ++++++++++++++++++++++++++----------
2 files changed, 46 insertions(+), 14 deletions(-)
diff --git a/include/net/netmem.h b/include/net/netmem.h
index 0cc8572f62caf..697d50c51a5ee 100644
--- a/include/net/netmem.h
+++ b/include/net/netmem.h
@@ -399,21 +399,37 @@ static inline bool net_is_devmem_iov(const struct net_iov *niov)
}
#endif
-void __get_netmem(netmem_ref netmem);
-void __put_netmem(netmem_ref netmem);
+void get_net_iov(struct net_iov *niov);
+void put_net_iov(struct net_iov *niov);
+/**
+ * get_netmem - acquire a non-page_pool reference on a netmem
+ * @netmem: netmem to reference
+ *
+ * Acquires a non-pp backing reference (page._refcount via get_page() or
+ * binding->ref via get_net_iov()) regardless of page_pool membership.
+ * Counterpart to put_netmem().
+ */
static __always_inline void get_netmem(netmem_ref netmem)
{
if (netmem_is_net_iov(netmem))
- __get_netmem(netmem);
+ get_net_iov(netmem_to_net_iov(netmem));
else
get_page(netmem_to_page(netmem));
}
+/**
+ * put_netmem - release a non-page_pool reference on a netmem
+ * @netmem: netmem to unreference
+ *
+ * Drops a non-pp backing reference (page._refcount via put_page() or
+ * binding->ref via put_net_iov()) regardless of page_pool membership.
+ * Counterpart to get_netmem().
+ */
static __always_inline void put_netmem(netmem_ref netmem)
{
if (netmem_is_net_iov(netmem))
- __put_netmem(netmem);
+ put_net_iov(netmem_to_net_iov(netmem));
else
put_page(netmem_to_page(netmem));
}
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index d56f4f0102f75..b459264ba5180 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -7551,23 +7551,39 @@ bool csum_and_copy_from_iter_full(void *addr, size_t bytes,
}
EXPORT_SYMBOL(csum_and_copy_from_iter_full);
-void __get_netmem(netmem_ref netmem)
+/**
+ * get_net_iov - acquire a non-page_pool reference on a net_iov
+ * @niov: net_iov to reference
+ *
+ * Acquires a non-pp backing reference (binding->ref) on @niov regardless of
+ * page_pool membership, warning if @niov does not support non-pp references.
+ * Counterpart to put_net_iov().
+ */
+void get_net_iov(struct net_iov *niov)
{
- struct net_iov *niov = netmem_to_net_iov(netmem);
-
if (net_is_devmem_iov(niov))
- net_devmem_get_net_iov(netmem_to_net_iov(netmem));
+ net_devmem_get_net_iov(niov);
+ else
+ DEBUG_NET_WARN_ON_ONCE(true);
}
-EXPORT_SYMBOL(__get_netmem);
+EXPORT_SYMBOL(get_net_iov);
-void __put_netmem(netmem_ref netmem)
+/**
+ * put_net_iov - release a non-page_pool reference on a net_iov
+ * @niov: net_iov to unreference
+ *
+ * Drops a non-pp backing reference (binding->ref) on @niov regardless of
+ * page_pool membership, warning if @niov does not support non-pp references.
+ * Counterpart to get_net_iov().
+ */
+void put_net_iov(struct net_iov *niov)
{
- struct net_iov *niov = netmem_to_net_iov(netmem);
-
if (net_is_devmem_iov(niov))
- net_devmem_put_net_iov(netmem_to_net_iov(netmem));
+ net_devmem_put_net_iov(niov);
+ else
+ DEBUG_NET_WARN_ON_ONCE(true);
}
-EXPORT_SYMBOL(__put_netmem);
+EXPORT_SYMBOL(put_net_iov);
struct vlan_type_depth __vlan_get_protocol_offset(const struct sk_buff *skb,
__be16 type,
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 2/6] net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf()
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov() Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 3/6] net: skbuff: replace skb_page_unref() and __skb_frag_unref() with skb_netmem_unref() Mina Almasry
` (4 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
While napi_pp_put_page() releases a page_pool pp_ref_count reference on
a netmem, no symmetric helper existed to acquire one, forcing callers
like tcp_recvmsg_dmabuf() to open-code atomic_long_inc() on
niov->desc.pp_ref_count.
Add napi_pp_get_page() to pair symmetrically with napi_pp_put_page(),
and use WARN_ON_ONCE(!napi_pp_get_page(netmem)) in tcp_recvmsg_dmabuf()
to match WARN_ON_ONCE(!napi_pp_put_page(netmem)) in
tcp_release_user_frags().
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
include/linux/skbuff_ref.h | 2 ++
net/core/skbuff.c | 31 +++++++++++++++++++++++++++++++
net/ipv4/tcp.c | 6 ++++--
3 files changed, 37 insertions(+), 2 deletions(-)
diff --git a/include/linux/skbuff_ref.h b/include/linux/skbuff_ref.h
index 05c8486bafac8..0441c220fdbdf 100644
--- a/include/linux/skbuff_ref.h
+++ b/include/linux/skbuff_ref.h
@@ -9,6 +9,8 @@
#include <linux/skbuff.h>
+bool napi_pp_get_page(netmem_ref netmem);
+
/**
* __skb_frag_ref - take an addition reference on a paged fragment.
* @frag: the paged fragment
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index b459264ba5180..8a8be746828db 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -1026,6 +1026,15 @@ int skb_cow_data_for_xdp(struct page_pool *pool, struct sk_buff **pskb,
EXPORT_SYMBOL(skb_cow_data_for_xdp);
#if IS_ENABLED(CONFIG_PAGE_POOL)
+/**
+ * napi_pp_put_page - release a page_pool reference on a netmem
+ * @netmem: netmem to unreference
+ *
+ * Drops a page_pool reference (pp_ref_count) on @netmem if it belongs to a
+ * page_pool. Counterpart to napi_pp_get_page().
+ *
+ * Return: true if @netmem belongs to a page_pool, false otherwise.
+ */
bool napi_pp_put_page(netmem_ref netmem)
{
netmem = netmem_compound_head(netmem);
@@ -1038,6 +1047,28 @@ bool napi_pp_put_page(netmem_ref netmem)
return true;
}
EXPORT_SYMBOL(napi_pp_put_page);
+
+/**
+ * napi_pp_get_page - acquire a page_pool reference on a netmem
+ * @netmem: netmem to reference
+ *
+ * Acquires a page_pool reference (pp_ref_count) on @netmem if it belongs to a
+ * page_pool. Counterpart to napi_pp_put_page().
+ *
+ * Return: true if @netmem belongs to a page_pool, false otherwise.
+ */
+bool napi_pp_get_page(netmem_ref netmem)
+{
+ netmem = netmem_compound_head(netmem);
+
+ if (unlikely(!netmem_is_pp(netmem)))
+ return false;
+
+ page_pool_ref_netmem(netmem);
+
+ return true;
+}
+EXPORT_SYMBOL(napi_pp_get_page);
#endif
static bool skb_pp_recycle(struct sk_buff *skb, void *data)
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index f4c9101e5996a..99d7237a623d8 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -254,6 +254,7 @@
#include <linux/init.h>
#include <linux/fs.h>
#include <linux/skbuff.h>
+#include <linux/skbuff_ref.h>
#include <linux/splice.h>
#include <linux/net.h>
#include <linux/socket.h>
@@ -2559,6 +2560,7 @@ static int tcp_recvmsg_dmabuf(struct sock *sk, const struct sk_buff *skb,
*/
for (i = 0; i < skb_shinfo(skb)->nr_frags; i++) {
skb_frag_t *frag = &skb_shinfo(skb)->frags[i];
+ netmem_ref netmem = skb_frag_netmem(frag);
struct net_iov *niov;
u64 frag_offset;
int end;
@@ -2612,8 +2614,8 @@ static int tcp_recvmsg_dmabuf(struct sock *sk, const struct sk_buff *skb,
if (err)
goto out;
- atomic_long_inc(&niov->desc.pp_ref_count);
- tcp_xa_pool.netmems[tcp_xa_pool.idx++] = skb_frag_netmem(frag);
+ WARN_ON_ONCE(!napi_pp_get_page(netmem));
+ tcp_xa_pool.netmems[tcp_xa_pool.idx++] = netmem;
sent += copy;
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 3/6] net: skbuff: replace skb_page_unref() and __skb_frag_unref() with skb_netmem_unref()
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 2/6] net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf() Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 4/6] net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref() Mina Almasry
` (3 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
Unreferencing a page_pool fragment with recycle=false decrements the
non-pp backing refcount instead of pp_ref_count, underflowing the page
refcount when the fragment was owned via page_pool (as on mlx4's RX
error path). Having both skb_page_unref() and __skb_frag_unref(frag,
recycle) made it easy for callers to pass hardcoded recycle=false on
SKB-owned fragments.
Replace skb_page_unref() and __skb_frag_unref() with
skb_netmem_unref(netmem, recycle), which dispatches to
napi_pp_put_page() when recycle is true and put_netmem() otherwise:
- Use skb_frag_unref(skb, ...) on SKB-owned fragments in mlx4, sky2,
tls_strp, and skb_shift() so unrefs always match skb->pp_recycle.
- Use put_netmem() directly in tls_device's destroy_record(), which
releases raw non-pp page references outside an SKB.
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
drivers/net/ethernet/marvell/sky2.c | 2 +-
drivers/net/ethernet/mellanox/mlx4/en_rx.c | 2 +-
include/linux/skbuff_ref.h | 34 +++++++++++-----------
net/core/skbuff.c | 7 +++--
net/ipv4/esp4.c | 4 +--
net/ipv6/esp6.c | 4 +--
net/tls/tls_device.c | 2 +-
net/tls/tls_strp.c | 2 +-
8 files changed, 29 insertions(+), 28 deletions(-)
diff --git a/drivers/net/ethernet/marvell/sky2.c b/drivers/net/ethernet/marvell/sky2.c
index 30facdece3e29..1d42a82db8f2b 100644
--- a/drivers/net/ethernet/marvell/sky2.c
+++ b/drivers/net/ethernet/marvell/sky2.c
@@ -2501,7 +2501,7 @@ static void skb_put_frags(struct sk_buff *skb, unsigned int hdr_space,
if (length == 0) {
/* don't need this page */
- __skb_frag_unref(frag, false);
+ skb_frag_unref(skb, i);
--skb_shinfo(skb)->nr_frags;
} else {
size = min(length, (unsigned) PAGE_SIZE);
diff --git a/drivers/net/ethernet/mellanox/mlx4/en_rx.c b/drivers/net/ethernet/mellanox/mlx4/en_rx.c
index 96f97fa0f5489..fe18606c2ab55 100644
--- a/drivers/net/ethernet/mellanox/mlx4/en_rx.c
+++ b/drivers/net/ethernet/mellanox/mlx4/en_rx.c
@@ -497,7 +497,7 @@ static int mlx4_en_complete_rx_desc(struct mlx4_en_priv *priv,
fail:
while (nr > 0) {
nr--;
- __skb_frag_unref(skb_shinfo(skb)->frags + nr, false);
+ skb_frag_unref(skb, nr);
}
return 0;
}
diff --git a/include/linux/skbuff_ref.h b/include/linux/skbuff_ref.h
index 0441c220fdbdf..04114c64dad99 100644
--- a/include/linux/skbuff_ref.h
+++ b/include/linux/skbuff_ref.h
@@ -36,7 +36,16 @@ static __always_inline void skb_frag_ref(struct sk_buff *skb, int f)
bool napi_pp_put_page(netmem_ref netmem);
-static __always_inline void skb_page_unref(netmem_ref netmem, bool recycle)
+/**
+ * skb_netmem_unref - release a fragment reference on a netmem
+ * @netmem: netmem to unreference
+ * @recycle: whether page_pool recycling is enabled (e.g. skb->pp_recycle)
+ *
+ * Drops pp_ref_count via napi_pp_put_page() if @recycle is true and @netmem
+ * belongs to a page_pool; otherwise drops a non-pp backing reference via
+ * put_netmem(). Counterpart to skb_netmem_ref().
+ */
+static __always_inline void skb_netmem_unref(netmem_ref netmem, bool recycle)
{
#ifdef CONFIG_PAGE_POOL
if (recycle && napi_pp_put_page(netmem))
@@ -46,31 +55,22 @@ static __always_inline void skb_page_unref(netmem_ref netmem, bool recycle)
}
/**
- * __skb_frag_unref - release a reference on a paged fragment.
- * @frag: the paged fragment
- * @recycle: recycle the page if allocated via page_pool
- *
- * Releases a reference on the paged fragment @frag
- * or recycles the page via the page_pool API.
- */
-static __always_inline void __skb_frag_unref(skb_frag_t *frag, bool recycle)
-{
- skb_page_unref(skb_frag_netmem(frag), recycle);
-}
-
-/**
- * skb_frag_unref - release a reference on a paged fragment of an skb.
+ * skb_frag_unref - release a reference on a paged fragment of an skb
* @skb: the buffer
* @f: the fragment offset
*
- * Releases a reference on the @f'th paged fragment of @skb.
+ * Drops a reference on the @f'th fragment of @skb via skb_netmem_unref()
+ * (pp_ref_count if skb->pp_recycle and page_pool-backed, else non-pp
+ * backing refcount), skipping managed zerocopy fragments. Counterpart to
+ * skb_frag_ref().
*/
static __always_inline void skb_frag_unref(struct sk_buff *skb, int f)
{
struct skb_shared_info *shinfo = skb_shinfo(skb);
if (!skb_zcopy_managed(skb))
- __skb_frag_unref(&shinfo->frags[f], skb->pp_recycle);
+ skb_netmem_unref(skb_frag_netmem(&shinfo->frags[f]),
+ skb->pp_recycle);
}
#endif /* _LINUX_SKBUFF_REF_H */
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 8a8be746828db..46da61ad1f9eb 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -1144,7 +1144,8 @@ static void skb_release_data(struct sk_buff *skb, enum skb_drop_reason reason)
}
for (i = 0; i < shinfo->nr_frags; i++)
- __skb_frag_unref(&shinfo->frags[i], skb->pp_recycle);
+ skb_netmem_unref(skb_frag_netmem(&shinfo->frags[i]),
+ skb->pp_recycle);
free_head:
if (shinfo->frag_list)
@@ -1154,7 +1155,7 @@ static void skb_release_data(struct sk_buff *skb, enum skb_drop_reason reason)
exit:
/* When we clone an SKB we copy the reycling bit. The pp_recycle
* bit is only set on the head though, so in order to avoid races
- * while trying to recycle fragments on __skb_frag_unref() we need
+ * while trying to recycle fragments on skb_netmem_unref() we need
* to make one SKB responsible for triggering the recycle path.
* So disable the recycling bit if an SKB is cloned and we have
* additional references to the fragmented part of the SKB.
@@ -4412,7 +4413,7 @@ int skb_shift(struct sk_buff *tgt, struct sk_buff *skb, int shiftlen)
fragto = &skb_shinfo(tgt)->frags[merge];
skb_frag_size_add(fragto, skb_frag_size(fragfrom));
- __skb_frag_unref(fragfrom, skb->pp_recycle);
+ skb_frag_unref(skb, 0);
}
/* Reposition in the original skb */
diff --git a/net/ipv4/esp4.c b/net/ipv4/esp4.c
index e76db5817e78e..ae701561d9f12 100644
--- a/net/ipv4/esp4.c
+++ b/net/ipv4/esp4.c
@@ -117,8 +117,8 @@ static void esp_ssg_unref(struct xfrm_state *x, void *tmp, struct sk_buff *skb,
struct scatterlist *src = already_unref ? esp_req_sg(aead, req) : req->src;
for (sg = sg_next(src); sg; sg = sg_next(sg))
- skb_page_unref(page_to_netmem(sg_page(sg)),
- skb->pp_recycle);
+ skb_netmem_unref(page_to_netmem(sg_page(sg)),
+ skb->pp_recycle);
}
}
diff --git a/net/ipv6/esp6.c b/net/ipv6/esp6.c
index b1c9b36f76dc4..d654f0f442388 100644
--- a/net/ipv6/esp6.c
+++ b/net/ipv6/esp6.c
@@ -134,8 +134,8 @@ static void esp_ssg_unref(struct xfrm_state *x, void *tmp, struct sk_buff *skb,
struct scatterlist *src = already_unref ? esp_req_sg(aead, req) : req->src;
for (sg = sg_next(src); sg; sg = sg_next(sg))
- skb_page_unref(page_to_netmem(sg_page(sg)),
- skb->pp_recycle);
+ skb_netmem_unref(page_to_netmem(sg_page(sg)),
+ skb->pp_recycle);
}
}
diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index f11d0528fc431..b333a9a72cdae 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -143,7 +143,7 @@ static void destroy_record(struct tls_record_info *record)
int i;
for (i = 0; i < record->num_frags; i++)
- __skb_frag_unref(&record->frags[i], false);
+ put_netmem(skb_frag_netmem(&record->frags[i]));
kfree(record);
}
diff --git a/net/tls/tls_strp.c b/net/tls/tls_strp.c
index 6cc222008d95c..53d3256d0114a 100644
--- a/net/tls/tls_strp.c
+++ b/net/tls/tls_strp.c
@@ -197,7 +197,7 @@ static void tls_strp_flush_anchor_copy(struct tls_strparser *strp)
DEBUG_NET_WARN_ON_ONCE(atomic_read(&shinfo->dataref) != 1);
for (i = 0; i < shinfo->nr_frags; i++)
- __skb_frag_unref(&shinfo->frags[i], false);
+ skb_frag_unref(strp->anchor, i);
shinfo->nr_frags = 0;
if (strp->copy_mode) {
kfree_skb_list(shinfo->frag_list);
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 4/6] net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref()
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
` (2 preceding siblings ...)
2026-10-10 8:38 ` [PATCH net-next v1 3/6] net: skbuff: replace skb_page_unref() and __skb_frag_unref() with skb_netmem_unref() Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 5/6] net: skbuff: use skb_frag_ref() in skb_try_coalesce() Mina Almasry
` (2 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski,
Stanislav Fomichev, Daniel Borkmann, Alexei Starovoitov,
Jesper Dangaard Brouer
Previously, __skb_frag_ref() and skb_frag_ref() unconditionally
incremented the non-pp backing refcount via get_netmem(), while
skb_frag_unref() decremented pp_ref_count whenever skb->pp_recycle was
set. Ref-ing a page_pool fragment into a pp_recycle=1 SKB (or clearing
pp_recycle=0 on an uncloned SKB while clones still held pp_recycle=1)
leaked the backing non-pp refcount and underflowed pp_ref_count across
IP-TFS, skb_split(), skb_shift(), and cloned unclone paths.
Replace __skb_frag_ref() with skb_netmem_ref(netmem, recycle) as the
exact counterpart to skb_netmem_unref(netmem, recycle):
- Keep skb->pp_recycle intact in skb_release_data() so uncloning paths
(pskb_expand_head() and pskb_carve_inside_{header,nonlinear}()) keep
trading pp_ref_count references symmetrically across clones.
- Propagate pp_recycle and call skb_frag_ref() on the destination SKB in
__pskb_copy_fclone(), skb_zerocopy(), skb_split(), skb_shift(), and
skb_segment().
- Use skb_frag_ref() in xfrm_iptfs, chcr_ktls, and cassini, and call
get_netmem() directly in veth and tls_device_fallback where callers
explicitly take a raw non-pp reference.
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Daniel Borkmann <daniel@iogearbox.net>
Cc: Alexei Starovoitov <ast@kernel.org>
Cc: Jesper Dangaard Brouer <hawk@kernel.org>
Cc: bpf@vger.kernel.org
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
.../chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 2 +-
drivers/net/ethernet/sun/cassini.c | 4 +-
drivers/net/veth.c | 2 +-
include/linux/skbuff_ref.h | 28 +++++++++-----
net/core/skbuff.c | 38 ++++++++-----------
net/tls/tls_device_fallback.c | 2 +-
net/xfrm/xfrm_iptfs.c | 3 +-
7 files changed, 41 insertions(+), 38 deletions(-)
diff --git a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
index dbec95b497773..4d645771596a8 100644
--- a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
+++ b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
@@ -1660,7 +1660,7 @@ static void chcr_ktls_copy_record_in_skb(struct sk_buff *nskb,
for (i = 0; i < record->num_frags; i++) {
skb_shinfo(nskb)->frags[i] = record->frags[i];
/* increase the frag ref count */
- __skb_frag_ref(&skb_shinfo(nskb)->frags[i]);
+ skb_frag_ref(nskb, i);
}
skb_shinfo(nskb)->nr_frags = record->num_frags;
diff --git a/drivers/net/ethernet/sun/cassini.c b/drivers/net/ethernet/sun/cassini.c
index 9a8b5cf2d6485..02158a1bb2f70 100644
--- a/drivers/net/ethernet/sun/cassini.c
+++ b/drivers/net/ethernet/sun/cassini.c
@@ -1998,7 +1998,7 @@ static int cas_rx_process_pkt(struct cas *cp, struct cas_rx_comp *rxc,
skb->len += hlen - swivel;
skb_frag_fill_page_desc(frag, page->buffer, off, hlen - swivel);
- __skb_frag_ref(frag);
+ skb_frag_ref(skb, 0);
/* any more data? */
if ((words[0] & RX_COMP1_SPLIT_PKT) && ((dlen -= hlen) > 0)) {
@@ -2022,7 +2022,7 @@ static int cas_rx_process_pkt(struct cas *cp, struct cas_rx_comp *rxc,
frag++;
skb_frag_fill_page_desc(frag, page->buffer, 0, hlen);
- __skb_frag_ref(frag);
+ skb_frag_ref(skb, 1);
RX_USED_ADD(page, hlen + cp->crc_size);
}
diff --git a/drivers/net/veth.c b/drivers/net/veth.c
index 643b97dc52456..90a333ebab683 100644
--- a/drivers/net/veth.c
+++ b/drivers/net/veth.c
@@ -746,7 +746,7 @@ static void veth_xdp_get(struct xdp_buff *xdp)
return;
for (i = 0; i < sinfo->nr_frags; i++)
- __skb_frag_ref(&sinfo->frags[i]);
+ get_netmem(skb_frag_netmem(&sinfo->frags[i]));
}
static int veth_convert_skb_to_xdp_buff(struct veth_rq *rq,
diff --git a/include/linux/skbuff_ref.h b/include/linux/skbuff_ref.h
index 04114c64dad99..f768a3fa1d606 100644
--- a/include/linux/skbuff_ref.h
+++ b/include/linux/skbuff_ref.h
@@ -12,26 +12,36 @@
bool napi_pp_get_page(netmem_ref netmem);
/**
- * __skb_frag_ref - take an addition reference on a paged fragment.
- * @frag: the paged fragment
+ * skb_netmem_ref - acquire a fragment reference on a netmem
+ * @netmem: netmem to reference
+ * @recycle: whether page_pool recycling is enabled (e.g. skb->pp_recycle)
*
- * Takes an additional reference on the paged fragment @frag.
+ * Acquires pp_ref_count via napi_pp_get_page() if @recycle is true and
+ * @netmem belongs to a page_pool; otherwise acquires a non-pp backing
+ * reference via get_netmem(). Counterpart to skb_netmem_unref().
*/
-static __always_inline void __skb_frag_ref(skb_frag_t *frag)
+static __always_inline void skb_netmem_ref(netmem_ref netmem, bool recycle)
{
- get_netmem(skb_frag_netmem(frag));
+#ifdef CONFIG_PAGE_POOL
+ if (recycle && napi_pp_get_page(netmem))
+ return;
+#endif
+ get_netmem(netmem);
}
/**
- * skb_frag_ref - take an addition reference on a paged fragment of an skb.
+ * skb_frag_ref - acquire a reference on a paged fragment of an skb
* @skb: the buffer
- * @f: the fragment offset.
+ * @f: the fragment offset
*
- * Takes an additional reference on the @f'th paged fragment of @skb.
+ * Acquires a reference on the @f'th fragment of @skb via skb_netmem_ref()
+ * (pp_ref_count if skb->pp_recycle and page_pool-backed, else non-pp
+ * backing refcount). Counterpart to skb_frag_unref().
*/
static __always_inline void skb_frag_ref(struct sk_buff *skb, int f)
{
- __skb_frag_ref(&skb_shinfo(skb)->frags[f]);
+ skb_netmem_ref(skb_frag_netmem(&skb_shinfo(skb)->frags[f]),
+ skb->pp_recycle);
}
bool napi_pp_put_page(netmem_ref netmem);
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 46da61ad1f9eb..f03ce8d5ae585 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -1133,7 +1133,7 @@ static void skb_release_data(struct sk_buff *skb, enum skb_drop_reason reason)
int i;
if (!skb_data_unref(skb, shinfo))
- goto exit;
+ return;
if (skb_zcopy(skb)) {
bool skip_unref = shinfo->flags & SKBFL_MANAGED_FRAG_REFS;
@@ -1152,17 +1152,6 @@ static void skb_release_data(struct sk_buff *skb, enum skb_drop_reason reason)
kfree_skb_list_reason(shinfo->frag_list, reason);
skb_free_head(skb);
-exit:
- /* When we clone an SKB we copy the reycling bit. The pp_recycle
- * bit is only set on the head though, so in order to avoid races
- * while trying to recycle fragments on skb_netmem_unref() we need
- * to make one SKB responsible for triggering the recycle path.
- * So disable the recycling bit if an SKB is cloned and we have
- * additional references to the fragmented part of the SKB.
- * Eventually the last SKB will have the recycling bit set and it's
- * dataref set to 0, which will trigger the recycling
- */
- skb->pp_recycle = 0;
}
/*
@@ -2287,9 +2276,10 @@ struct sk_buff *__pskb_copy_fclone(struct sk_buff *skb, int headroom,
n = NULL;
goto out;
}
+ n->pp_recycle = skb->pp_recycle;
for (i = 0; i < skb_shinfo(skb)->nr_frags; i++) {
skb_shinfo(n)->frags[i] = skb_shinfo(skb)->frags[i];
- skb_frag_ref(skb, i);
+ skb_frag_ref(n, i);
}
skb_shinfo(n)->nr_frags = i;
skb_shinfo(n)->flags |= skb_shinfo(skb)->flags & SKBFL_SHARED_FRAG;
@@ -3927,6 +3917,7 @@ skb_zerocopy(struct sk_buff *to, struct sk_buff *from, int len, int hlen)
if (len <= skb_tailroom(to))
return skb_copy_bits(from, 0, skb_put(to, len), len);
+ to->pp_recycle = from->pp_recycle;
if (hlen) {
ret = skb_copy_bits(from, 0, skb_put(to, hlen), hlen);
if (unlikely(ret))
@@ -3939,14 +3930,14 @@ skb_zerocopy(struct sk_buff *to, struct sk_buff *from, int len, int hlen)
offset = from->data - (unsigned char *)page_address(page);
__skb_fill_netmem_desc(to, 0, page_to_netmem(page),
offset, plen);
- get_page(page);
+ skb_frag_ref(to, 0);
j = 1;
len -= plen;
}
}
if (!skb_frags_readable(from) && j > 0 && len) {
- put_page(page);
+ skb_frag_unref(to, 0);
return -EFAULT;
}
@@ -3954,7 +3945,7 @@ skb_zerocopy(struct sk_buff *to, struct sk_buff *from, int len, int hlen)
if (unlikely(skb_orphan_frags(from, GFP_ATOMIC))) {
if (j > 0)
- put_page(page);
+ skb_frag_unref(to, 0);
return -ENOMEM;
}
skb_zerocopy_clone(to, from, GFP_ATOMIC);
@@ -4255,7 +4246,7 @@ static inline void skb_split_no_header(struct sk_buff *skb,
* where splitting is expensive.
* 2. Split is accurately. We make this.
*/
- skb_frag_ref(skb, i);
+ skb_frag_ref(skb1, k);
skb_frag_off_add(&skb_shinfo(skb1)->frags[0], len - pos);
skb_frag_size_sub(&skb_shinfo(skb1)->frags[0], len - pos);
skb_frag_size_set(&skb_shinfo(skb)->frags[i], len - pos);
@@ -4286,6 +4277,7 @@ void skb_split(struct sk_buff *skb, struct sk_buff *skb1, const u32 len)
skb_shinfo(skb1)->flags |= skb_shinfo(skb)->flags & zc_flags;
skb_zerocopy_clone(skb1, skb, 0);
+ skb1->pp_recycle = skb->pp_recycle;
if (len < pos) /* Split line is inside header. */
skb_split_inside_header(skb, skb1, len, pos);
else /* Second chunk has no header, nothing to copy. */
@@ -4343,8 +4335,8 @@ int skb_shift(struct sk_buff *tgt, struct sk_buff *skb, int shiftlen)
/* Actual merge is delayed until the point when we know we can
* commit all, so that we don't have to undo partial changes
*/
- if (!skb_can_coalesce(tgt, to, skb_frag_page(fragfrom),
- skb_frag_off(fragfrom))) {
+ if (!skb_can_coalesce_netmem(tgt, to, skb_frag_netmem(fragfrom),
+ skb_frag_off(fragfrom))) {
merge = -1;
} else {
merge = to - 1;
@@ -4391,8 +4383,8 @@ int skb_shift(struct sk_buff *tgt, struct sk_buff *skb, int shiftlen)
to++;
} else {
- __skb_frag_ref(fragfrom);
skb_frag_page_copy(fragto, fragfrom);
+ skb_frag_ref(tgt, to);
skb_frag_off_copy(fragto, fragfrom);
skb_frag_size_set(fragto, todo);
@@ -5087,8 +5079,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
goto err;
}
+ if (!skb_shinfo(nskb)->nr_frags)
+ nskb->pp_recycle = frag_skb->pp_recycle;
*nskb_frag = (i < 0) ? skb_head_frag_to_page_desc(frag_skb) : *frag;
- __skb_frag_ref(nskb_frag);
+ skb_frag_ref(nskb, skb_shinfo(nskb)->nr_frags);
size = skb_frag_size(nskb_frag);
if (pos < offset) {
@@ -6303,7 +6297,7 @@ bool skb_try_coalesce(struct sk_buff *to, struct sk_buff *from,
*/
if (skb_pp_frag_ref(from)) {
for (i = 0; i < from_shinfo->nr_frags; i++)
- __skb_frag_ref(&from_shinfo->frags[i]);
+ get_netmem(skb_frag_netmem(&from_shinfo->frags[i]));
}
to->truesize += delta;
diff --git a/net/tls/tls_device_fallback.c b/net/tls/tls_device_fallback.c
index 3b7d0ab2bcf17..35cc510db4321 100644
--- a/net/tls/tls_device_fallback.c
+++ b/net/tls/tls_device_fallback.c
@@ -256,7 +256,7 @@ static int fill_sg_in(struct scatterlist *sg_in,
for (i = 0; remaining > 0; i++) {
skb_frag_t *frag = &record->frags[i];
- __skb_frag_ref(frag);
+ get_netmem(skb_frag_netmem(frag));
sg_set_page(sg_in + i, skb_frag_page(frag),
skb_frag_size(frag), skb_frag_off(frag));
diff --git a/net/xfrm/xfrm_iptfs.c b/net/xfrm/xfrm_iptfs.c
index 6920940a35b49..103cf7c5dbfb9 100644
--- a/net/xfrm/xfrm_iptfs.c
+++ b/net/xfrm/xfrm_iptfs.c
@@ -486,8 +486,7 @@ static int iptfs_skb_add_frags(struct sk_buff *skb,
tofrag->len -= offset;
offset = 0;
}
- __skb_frag_ref(tofrag);
- shinfo->nr_frags++;
+ skb_frag_ref(skb, shinfo->nr_frags++);
shinfo->flags |= SKBFL_SHARED_FRAG;
/* see if we are done */
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 5/6] net: skbuff: use skb_frag_ref() in skb_try_coalesce()
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
` (3 preceding siblings ...)
2026-10-10 8:38 ` [PATCH net-next v1 4/6] net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref() Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 6/6] net: kunit: test netmem, page_pool, and skb frag refcounting Mina Almasry
2026-10-10 8:45 ` [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting netdev-bot+sinfo
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
skb_pp_frag_ref() was introduced solely to work around skb_frag_ref()
ignoring skb->pp_recycle and incrementing the non-pp backing refcount.
Now that skb_frag_ref() handles both pp_recycle=1 and pp_recycle=0
symmetrically, call skb_frag_ref() directly in skb_try_coalesce() and
remove skb_pp_frag_ref().
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
net/core/skbuff.c | 37 ++-----------------------------------
1 file changed, 2 insertions(+), 35 deletions(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index f03ce8d5ae585..554af794bd7cd 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -1078,37 +1078,6 @@ static bool skb_pp_recycle(struct sk_buff *skb, void *data)
return napi_pp_put_page(page_to_netmem(virt_to_page(data)));
}
-/**
- * skb_pp_frag_ref() - Increase fragment references of a page pool aware skb
- * @skb: page pool aware skb
- *
- * Increase the fragment reference count (pp_ref_count) of a skb. This is
- * intended to gain fragment references only for page pool aware skbs,
- * i.e. when skb->pp_recycle is true, and not for fragments in a
- * non-pp-recycling skb. It has a fallback to increase references on normal
- * pages, as page pool aware skbs may also have normal page fragments.
- */
-static int skb_pp_frag_ref(struct sk_buff *skb)
-{
- struct skb_shared_info *shinfo;
- netmem_ref head_netmem;
- int i;
-
- if (!skb->pp_recycle)
- return -EINVAL;
-
- shinfo = skb_shinfo(skb);
-
- for (i = 0; i < shinfo->nr_frags; i++) {
- head_netmem = netmem_compound_head(shinfo->frags[i].netmem);
- if (likely(netmem_is_pp(head_netmem)))
- page_pool_ref_netmem(head_netmem);
- else
- page_ref_inc(netmem_to_page(head_netmem));
- }
- return 0;
-}
-
static void skb_kfree_head(void *head)
{
kfree(head);
@@ -6295,10 +6264,8 @@ bool skb_try_coalesce(struct sk_buff *to, struct sk_buff *from,
/* if the skb is not cloned this does nothing
* since we set nr_frags to 0.
*/
- if (skb_pp_frag_ref(from)) {
- for (i = 0; i < from_shinfo->nr_frags; i++)
- get_netmem(skb_frag_netmem(&from_shinfo->frags[i]));
- }
+ for (i = 0; i < from_shinfo->nr_frags; i++)
+ skb_frag_ref(from, i);
to->truesize += delta;
to->len += len;
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH net-next v1 6/6] net: kunit: test netmem, page_pool, and skb frag refcounting
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
` (4 preceding siblings ...)
2026-10-10 8:38 ` [PATCH net-next v1 5/6] net: skbuff: use skb_frag_ref() in skb_try_coalesce() Mina Almasry
@ 2026-10-10 8:38 ` Mina Almasry
2026-10-10 8:45 ` [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting netdev-bot+sinfo
6 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:38 UTC (permalink / raw)
To: netdev, linux-kernel, linux-rdma, bpf
Cc: Mina Almasry, Ayush Sawal, Andrew Lunn, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Tariq Toukan,
Simon Horman, Steffen Klassert, Herbert Xu, Neal Cardwell,
Kuniyuki Iwashima, John Fastabend, Sabrina Dubroca, Eric Biggers,
Kees Cook, Michael Grzeschik,
Uwe Kleine-König (The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
Verify that non-pp (page._refcount / binding->ref) and page_pool
(pp_ref_count) references stay balanced without cross-counter leaks or
underflows across regular pages, page_pool pages, non-pp net_iovs, and
page_pool net_iovs with both pp_recycle=0 and pp_recycle=1.
Cover get/put_net_iov(), get/put_netmem(), napi_pp_get/put_page(),
skb_netmem_ref/unref(), skb_frag_ref/unref(), pskb_copy(),
skb_zerocopy(), skb_shift(), skb_split(), skb_release_data() (clones,
unclone ordering, pskb_extract(), and skb_morph()), and skb_segment().
Signed-off-by: Mina Almasry <almasrymina@google.com>
---
net/core/net_test.c | 747 ++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 747 insertions(+)
diff --git a/net/core/net_test.c b/net/core/net_test.c
index 9c3a590865d26..f9f50f9159902 100644
--- a/net/core/net_test.c
+++ b/net/core/net_test.c
@@ -370,10 +370,757 @@ static void ip_tunnel_flags_test_run(struct kunit *test)
KUNIT_ASSERT_TRUE(test, __ipt_flag_op(bitmap_equal, exp, out));
}
+/* Netmem & SKB fragment refcounting */
+
+#include <linux/skbuff_ref.h>
+#include <net/netmem.h>
+#include <net/netdev_rx_queue.h>
+#include <net/page_pool/helpers.h>
+#include <net/page_pool/memory_provider.h>
+#include "devmem.h"
+#include "netmem_priv.h"
+
+static void netmem_ref_test_page(struct kunit *test)
+{
+ struct sk_buff *skb;
+ netmem_ref netmem;
+ struct page *page;
+ int base_ref;
+
+ page = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ netmem = page_to_netmem(page);
+ base_ref = page_ref_count(page);
+
+ /* Layer 1: get_netmem / put_netmem use page non-pp refcount */
+ get_netmem(netmem);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ put_netmem(netmem);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+#ifdef CONFIG_PAGE_POOL
+ /* Layer 2: napi_pp_get/put_page return false on non-PP page */
+ KUNIT_EXPECT_FALSE(test, napi_pp_get_page(netmem));
+ KUNIT_EXPECT_FALSE(test, napi_pp_put_page(netmem));
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+#endif
+
+ /* Layer 3: skb_netmem_ref / unref fall back to non-pp on non-PP page */
+ skb_netmem_ref(netmem, false);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ skb_netmem_unref(netmem, false);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ skb_netmem_ref(netmem, true);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ skb_netmem_unref(netmem, true);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ /* Layer 4: skb_frag_ref / skb_frag_unref on non-PP page */
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+
+ skb->pp_recycle = 1;
+ skb_frag_ref(skb, 0);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ skb_frag_unref(skb, 0);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ skb_shinfo(skb)->nr_frags = 0;
+ consume_skb(skb);
+ __free_page(page);
+}
+
+#ifdef CONFIG_PAGE_POOL
+static void netmem_ref_test_pp_page(struct kunit *test)
+{
+ struct page_pool_params pp_params = {
+ .order = 0,
+ .pool_size = 4,
+ .nid = NUMA_NO_NODE,
+ };
+ struct page_pool *pool;
+ struct sk_buff *skb, *copy, *clone, *split;
+ netmem_ref netmem;
+ struct page *page;
+ long base_pp_ref;
+ int base_ref;
+
+ pool = page_pool_create(&pp_params);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool);
+
+ page = page_pool_dev_alloc_pages(pool);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ netmem = page_to_netmem(page);
+ base_ref = page_ref_count(page);
+ base_pp_ref = atomic_long_read(&page->pp_ref_count);
+
+ /* Layer 1: get_netmem / put_netmem always use non-pp refcount */
+ get_netmem(netmem);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+ put_netmem(netmem);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ /* Layer 2: napi_pp_get/put_page use pp_ref_count */
+ KUNIT_EXPECT_TRUE(test, napi_pp_get_page(netmem));
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_TRUE(test, napi_pp_put_page(netmem));
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* Layer 3: skb_netmem_ref / unref obey recycle flag */
+ skb_netmem_ref(netmem, false);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+ skb_netmem_unref(netmem, false);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ skb_netmem_ref(netmem, true);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ skb_netmem_unref(netmem, true);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* Layer 4: skb_frag_ref / skb_frag_unref match on pp_recycle=1 & 0 */
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len = 64;
+ skb->data_len = 64;
+
+ skb->pp_recycle = 1;
+ skb_frag_ref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ skb_frag_unref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* pskb_copy on pp_recycle=1 skb propagates pp_recycle & pp_ref_count */
+ copy = pskb_copy(skb, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, copy);
+ KUNIT_EXPECT_TRUE(test, copy->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(copy);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* skb_zerocopy on pp_recycle=1 skb propagates pp_ref_count */
+ copy = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, copy);
+ skb_put(copy, skb_tailroom(copy));
+ skb->head_frag = 1;
+ KUNIT_ASSERT_EQ(test, skb_zerocopy(copy, skb, skb->len, 0), 0);
+ skb->head_frag = 0;
+ KUNIT_EXPECT_TRUE(test, copy->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(copy);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* skb_clone + pskb_expand_head preserves pp_recycle and pp_ref_count */
+ clone = skb_clone(skb, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, clone);
+ KUNIT_ASSERT_EQ(test, pskb_expand_head(clone, 32, 0, GFP_KERNEL), 0);
+ KUNIT_EXPECT_TRUE(test, clone->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(clone);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* skb_split on page_pool frag propagates pp_recycle & balances refs */
+ split = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, split);
+ skb_split(skb, split, 32);
+ KUNIT_EXPECT_TRUE(test, split->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(split);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* skb_segment on pp_recycle=1 skb propagates pp_recycle to segments */
+ skb->protocol = htons(ETH_P_IP);
+ skb_reset_network_header(skb);
+ skb_shinfo(skb)->gso_size = 16;
+ split = skb_segment(skb, NETIF_F_SG | NETIF_F_HW_CSUM);
+ KUNIT_ASSERT_FALSE(test, IS_ERR_OR_NULL(split));
+ KUNIT_EXPECT_TRUE(test, split->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 2);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ kfree_skb_list(split);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* skb_shift balances pp_ref_count on split and merge */
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len = 64;
+ skb->data_len = 64;
+ split = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, split);
+ split->pp_recycle = 1;
+ KUNIT_EXPECT_EQ(test, skb_shift(split, skb, 16), 16);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ KUNIT_EXPECT_EQ(test, skb_shift(split, skb, 48), 48);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+ skb_shinfo(split)->nr_frags = 0;
+ consume_skb(split);
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len = 64;
+ skb->data_len = 64;
+
+ skb->pp_recycle = 0;
+ skb_frag_ref(skb, 0);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref + 1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+ skb_frag_unref(skb, 0);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ skb_shinfo(skb)->nr_frags = 0;
+ consume_skb(skb);
+ page_pool_put_full_page(pool, page, false);
+ page_pool_destroy(pool);
+}
+#endif
+
+#if defined(CONFIG_PAGE_POOL) && defined(CONFIG_NET_DEVMEM)
+static void dummy_percpu_ref_release(struct percpu_ref *ref)
+{
+}
+
+static void netmem_ref_test_net_iov(struct kunit *test)
+{
+ struct net_devmem_dmabuf_binding binding = {};
+ struct page_pool_params pp_params = {
+ .order = 0,
+ .pool_size = 4,
+ .nid = NUMA_NO_NODE,
+ };
+ struct sk_buff *skb, *clone, *copy;
+ struct page_pool *pool;
+ struct net_iov niov = {};
+ netmem_ref netmem;
+ long base_ref;
+ int err;
+
+ err = percpu_ref_init(&binding.ref, dummy_percpu_ref_release,
+ PERCPU_REF_INIT_ATOMIC, GFP_KERNEL);
+ KUNIT_ASSERT_EQ(test, err, 0);
+
+ net_iov_init(&niov, &binding.area, NET_IOV_DMABUF);
+ netmem = net_iov_to_netmem(&niov);
+ base_ref = atomic_long_read(&binding.ref.data->count);
+
+ /* Non-PP net_iov: get/put_net_iov & get/put_netmem */
+ get_net_iov(&niov);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref + 1);
+ put_net_iov(&niov);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+
+ KUNIT_EXPECT_FALSE(test, napi_pp_get_page(netmem));
+ KUNIT_EXPECT_FALSE(test, napi_pp_put_page(netmem));
+
+ skb_netmem_ref(netmem, true);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref + 1);
+ skb_netmem_unref(netmem, true);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+
+ /* Attach net_iov to a page_pool to test PP net_iov behavior */
+ pool = page_pool_create(&pp_params);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool);
+ netmem_set_pp(netmem, pool);
+ netmem_or_pp_magic(netmem, PP_SIGNATURE);
+ atomic_long_set(&niov.desc.pp_ref_count, 2);
+
+ /* Layer 0 & 1: get_net_iov / get_netmem grab non-pp ref even on PP */
+ get_net_iov(&niov);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref + 1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+ put_net_iov(&niov);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+
+ /* Layer 2: napi_pp_get/put_page use pp_ref_count */
+ KUNIT_EXPECT_TRUE(test, napi_pp_get_page(netmem));
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ KUNIT_EXPECT_TRUE(test, napi_pp_put_page(netmem));
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ /* Layer 4: skb_frag_ref / skb_frag_unref on PP net_iov */
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len = 64;
+ skb->data_len = 64;
+
+ skb->pp_recycle = 1;
+ skb_frag_ref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ skb_frag_unref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ /* pskb_copy on pp_recycle=1 net_iov skb keeps pp_ref_count */
+ copy = pskb_copy(skb, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, copy);
+ KUNIT_EXPECT_TRUE(test, copy->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ consume_skb(copy);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ /* skb_zerocopy on pp_recycle=1 net_iov skb keeps pp_ref_count */
+ copy = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, copy);
+ skb_put(copy, skb_tailroom(copy));
+ skb->head_frag = 1;
+ skb->unreadable = 1;
+ KUNIT_ASSERT_EQ(test, skb_zerocopy(copy, skb, skb->len, 0), 0);
+ skb->head_frag = 0;
+ KUNIT_EXPECT_TRUE(test, copy->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ consume_skb(copy);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ /* skb_shift balances pp_ref_count on split and merge */
+ copy = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, copy);
+ copy->pp_recycle = 1;
+ KUNIT_EXPECT_EQ(test, skb_shift(copy, skb, 16), 16);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ KUNIT_EXPECT_EQ(test, skb_shift(copy, skb, 48), 48);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+ skb_shinfo(copy)->nr_frags = 0;
+ consume_skb(copy);
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len = 64;
+ skb->data_len = 64;
+
+ /* Uncloning a cloned pp_recycle=1 net_iov skb keeps pp_ref_count */
+ clone = skb_clone(skb, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, clone);
+ KUNIT_ASSERT_EQ(test, pskb_expand_head(skb, 32, 0, GFP_KERNEL), 0);
+ KUNIT_EXPECT_TRUE(test, skb->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+ consume_skb(clone);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ skb->pp_recycle = 0;
+ skb_frag_ref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref + 1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+ skb_frag_unref(skb, 0);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_ref);
+
+ skb_shinfo(skb)->nr_frags = 0;
+ consume_skb(skb);
+ netmem_clear_pp_magic(netmem);
+ netmem_set_pp(netmem, NULL);
+ page_pool_destroy(pool);
+ percpu_ref_exit(&binding.ref);
+}
+
+static void netmem_ref_test_release_data(struct kunit *test)
+{
+ struct page_pool_params pp_params = {
+ .order = 0,
+ .pool_size = 4,
+ .nid = NUMA_NO_NODE,
+ };
+ struct sk_buff *skb, *clone1, *clone2, *ext;
+ struct net_devmem_dmabuf_binding binding = {};
+ struct net_iov niov = {};
+ struct page_pool *pool;
+ struct page *page;
+ netmem_ref netmem;
+ long base_pp_ref;
+ int base_ref;
+ int i;
+
+ pool = page_pool_create(&pp_params);
+ KUNIT_ASSERT_FALSE(test, IS_ERR(pool));
+ page = page_pool_dev_alloc_pages(pool);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ netmem = page_to_netmem(page);
+ page_pool_ref_netmem(netmem);
+
+ /* 3-way clone chain: only last clone drops pp_ref_count */
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_put(skb, 32);
+ skb->pp_recycle = 1;
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len += 64;
+ skb->data_len = 64;
+ base_ref = page_ref_count(page);
+ base_pp_ref = atomic_long_read(&page->pp_ref_count);
+
+ page_pool_ref_netmem(netmem);
+ clone1 = skb_clone(skb, GFP_KERNEL);
+ clone2 = skb_clone(clone1, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, clone1);
+ KUNIT_ASSERT_NOT_NULL(test, clone2);
+
+ /* Freeing head skb first leaves clone1/clone2 pp_recycle=1 intact */
+ consume_skb(skb);
+ KUNIT_EXPECT_TRUE(test, clone1->pp_recycle);
+ KUNIT_EXPECT_TRUE(test, clone2->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ consume_skb(clone2);
+ KUNIT_EXPECT_TRUE(test, clone1->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ consume_skb(clone1);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+
+ /* All 4 combinations of (unclone skb vs clone) x (free order) */
+ for (i = 0; i < 4; i++) {
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_put(skb, 32);
+ skb->pp_recycle = 1;
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len += 64;
+ skb->data_len = 64;
+ page_pool_ref_netmem(netmem);
+
+ clone1 = skb_clone(skb, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, clone1);
+ KUNIT_ASSERT_EQ(test,
+ pskb_expand_head((i & 1) ? clone1 : skb,
+ 32, 0, GFP_KERNEL), 0);
+ KUNIT_EXPECT_TRUE(test, skb->pp_recycle);
+ KUNIT_EXPECT_TRUE(test, clone1->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 2);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+
+ if (i & 2) {
+ consume_skb(clone1);
+ KUNIT_EXPECT_EQ(test,
+ atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ consume_skb(skb);
+ } else {
+ consume_skb(skb);
+ KUNIT_EXPECT_EQ(test,
+ atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ consume_skb(clone1);
+ }
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ }
+
+ /* pskb_extract carving inside header & inside nonlinear frags */
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_put(skb, 32);
+ skb->pp_recycle = 1;
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len += 64;
+ skb->data_len = 64;
+ page_pool_ref_netmem(netmem);
+
+ /* Carve inside header keeping frag -> pp_ref_count +1 until free */
+ ext = pskb_extract(skb, 16, skb->len - 16, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, ext);
+ KUNIT_EXPECT_TRUE(test, ext->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 2);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(ext);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+
+ /* Carve + trim + skb_condense ref/unrefs pp_ref_count cleanly */
+ ext = pskb_extract(skb, 16, 48, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, ext);
+ KUNIT_EXPECT_TRUE(test, ext->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(ext);
+
+ /* Carve inside nonlinear keeping frag -> pp_ref_count +1 until free */
+ ext = pskb_extract(skb, 48, skb->len - 48, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, ext);
+ KUNIT_EXPECT_TRUE(test, ext->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 2);
+ KUNIT_EXPECT_EQ(test, page_ref_count(page), base_ref);
+ consume_skb(ext);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+
+ /* skb_morph drops old skb data via skb_release_data */
+ page_pool_ref_netmem(netmem);
+ clone1 = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, clone1);
+ clone1->pp_recycle = 1;
+ skb_fill_netmem_desc(clone1, 0, netmem, 0, 64);
+ clone1->len = 64;
+ clone1->data_len = 64;
+ skb_morph(clone1, skb);
+ KUNIT_EXPECT_TRUE(test, clone1->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&page->pp_ref_count),
+ base_pp_ref + 1);
+ consume_skb(clone1);
+ kfree_skb(skb);
+ page_pool_put_full_page(pool, page, false);
+ page_pool_put_full_page(pool, page, false);
+
+ /* Also verify pskb_extract on page_pool net_iov */
+ KUNIT_ASSERT_EQ(test,
+ percpu_ref_init(&binding.ref, dummy_percpu_ref_release,
+ PERCPU_REF_INIT_ATOMIC, GFP_KERNEL),
+ 0);
+ net_iov_init(&niov, &binding.area, NET_IOV_DMABUF);
+ netmem = net_iov_to_netmem(&niov);
+ netmem_set_pp(netmem, pool);
+ netmem_or_pp_magic(netmem, PP_SIGNATURE);
+ atomic_long_set(&niov.desc.pp_ref_count, 2);
+
+ skb = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, skb);
+ skb_put(skb, 32);
+ skb->pp_recycle = 1;
+ skb_fill_netmem_desc(skb, 0, netmem, 0, 64);
+ skb->len += 64;
+ skb->data_len = 64;
+
+ ext = pskb_extract(skb, 48, skb->len - 48, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, ext);
+ KUNIT_EXPECT_TRUE(test, ext->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 3L);
+ consume_skb(ext);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+
+ skb_shinfo(skb)->nr_frags = 0;
+ consume_skb(skb);
+ netmem_clear_pp_magic(netmem);
+ netmem_set_pp(netmem, NULL);
+ page_pool_destroy(pool);
+ percpu_ref_exit(&binding.ref);
+}
+
+static void netmem_ref_test_skb_segment(struct kunit *test)
+{
+ struct page_pool_params pp_params = {
+ .order = 0,
+ .pool_size = 4,
+ .nid = NUMA_NO_NODE,
+ };
+ struct sk_buff *head, *list, *segs, *seg;
+ struct net_devmem_dmabuf_binding binding = {};
+ struct net_iov niov = {};
+ struct page *pp_page, *reg_page;
+ netmem_ref pp_nm, niov_nm;
+ struct page_pool *pool;
+ long base_pp_ref, base_iov_ref;
+ int base_pp_page_ref, base_reg_ref;
+
+ pool = page_pool_create(&pp_params);
+ KUNIT_ASSERT_FALSE(test, IS_ERR(pool));
+ pp_page = page_pool_dev_alloc_pages(pool);
+ KUNIT_ASSERT_NOT_NULL(test, pp_page);
+ pp_nm = page_to_netmem(pp_page);
+ page_pool_ref_netmem(pp_nm);
+ reg_page = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, reg_page);
+
+ KUNIT_ASSERT_EQ(test,
+ percpu_ref_init(&binding.ref, dummy_percpu_ref_release,
+ PERCPU_REF_INIT_ATOMIC, GFP_KERNEL),
+ 0);
+ net_iov_init(&niov, &binding.area, NET_IOV_DMABUF);
+ niov_nm = net_iov_to_netmem(&niov);
+ netmem_set_pp(niov_nm, pool);
+ netmem_or_pp_magic(niov_nm, PP_SIGNATURE);
+ atomic_long_set(&niov.desc.pp_ref_count, 2);
+
+ base_pp_ref = atomic_long_read(&pp_page->pp_ref_count);
+ base_pp_page_ref = page_ref_count(pp_page);
+ base_reg_ref = page_ref_count(reg_page);
+ base_iov_ref = atomic_long_read(&binding.ref.data->count);
+
+ /* 1. Single GSO skb with page_pool net_iov frags */
+ head = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, head);
+ head->pp_recycle = 1;
+ head->protocol = htons(ETH_P_IP);
+ skb_reset_mac_header(head);
+ skb_reset_network_header(head);
+ skb_fill_netmem_desc(head, 0, niov_nm, 0, 64);
+ head->len = 64;
+ head->data_len = 64;
+ skb_shinfo(head)->gso_size = 16;
+
+ segs = skb_segment(head, NETIF_F_SG | NETIF_F_HW_CSUM);
+ KUNIT_ASSERT_FALSE(test, IS_ERR_OR_NULL(segs));
+ for (seg = segs; seg; seg = seg->next)
+ KUNIT_EXPECT_TRUE(test, seg->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 6L);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&binding.ref.data->count),
+ base_iov_ref);
+ kfree_skb_list(segs);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&niov.desc.pp_ref_count), 2L);
+ skb_shinfo(head)->nr_frags = 0;
+ consume_skb(head);
+
+ /* 2. GRO frag_list fast path (skb_clone of pp_recycle=1 list_skb) */
+ head = alloc_skb(64, GFP_KERNEL);
+ list = alloc_skb(64, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, head);
+ KUNIT_ASSERT_NOT_NULL(test, list);
+ head->pp_recycle = 1;
+ list->pp_recycle = 1;
+ head->protocol = htons(ETH_P_IP);
+ skb_reset_mac_header(head);
+ skb_reset_network_header(head);
+ skb_put(head, 16);
+ skb_fill_netmem_desc(head, 0, pp_nm, 0, 16);
+ head->len += 16;
+ head->data_len = 16;
+ page_pool_ref_netmem(pp_nm);
+
+ skb_put(list, 16);
+ skb_fill_netmem_desc(list, 0, pp_nm, 16, 16);
+ list->len += 16;
+ list->data_len = 16;
+ page_pool_ref_netmem(pp_nm);
+
+ skb_shinfo(head)->frag_list = list;
+ head->len += list->len;
+ head->data_len += list->len;
+ skb_shinfo(head)->gso_size = 32;
+
+ segs = skb_segment(head, NETIF_F_SG | NETIF_F_HW_CSUM);
+ KUNIT_ASSERT_FALSE(test, IS_ERR_OR_NULL(segs));
+ for (seg = segs; seg; seg = seg->next)
+ KUNIT_EXPECT_TRUE(test, seg->pp_recycle);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref + 4);
+ KUNIT_EXPECT_EQ(test, page_ref_count(pp_page), base_pp_page_ref);
+ kfree_skb_list(segs);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref + 2);
+ consume_skb(head);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref);
+
+ /* 3. Mixed GRO frag_list: head pp_page + list reg_page */
+ head = alloc_skb(64, GFP_KERNEL);
+ list = alloc_skb(0, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, head);
+ KUNIT_ASSERT_NOT_NULL(test, list);
+ head->pp_recycle = 1;
+ list->pp_recycle = 0;
+ head->protocol = htons(ETH_P_IP);
+ skb_reset_mac_header(head);
+ skb_reset_network_header(head);
+ skb_fill_netmem_desc(head, 0, pp_nm, 0, 24);
+ head->len = 24;
+ head->data_len = 24;
+ page_pool_ref_netmem(pp_nm);
+
+ skb_fill_page_desc(list, 0, reg_page, 0, 24);
+ get_page(reg_page);
+ list->len = 24;
+ list->data_len = 24;
+
+ skb_shinfo(head)->frag_list = list;
+ head->len += list->len;
+ head->data_len += list->len;
+ skb_shinfo(head)->gso_size = 16;
+
+ /* Segment 0: [pp_page:16], Segment 1: [pp_page:8 + reg_page:8],
+ * Segment 2: [reg_page:16]. Freeing all segments restores exact refs!
+ */
+ segs = skb_segment(head, NETIF_F_SG | NETIF_F_HW_CSUM);
+ KUNIT_ASSERT_FALSE(test, IS_ERR_OR_NULL(segs));
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref + 3);
+ KUNIT_EXPECT_EQ(test, page_ref_count(pp_page), base_pp_page_ref);
+ KUNIT_EXPECT_EQ(test, page_ref_count(reg_page), base_reg_ref + 3);
+ kfree_skb_list(segs);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref + 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(pp_page), base_pp_page_ref);
+ KUNIT_EXPECT_EQ(test, page_ref_count(reg_page), base_reg_ref + 1);
+
+ consume_skb(head);
+ KUNIT_EXPECT_EQ(test, atomic_long_read(&pp_page->pp_ref_count),
+ base_pp_ref);
+ page_pool_put_full_page(pool, pp_page, false);
+ page_pool_put_full_page(pool, pp_page, false);
+ __free_page(reg_page);
+ netmem_clear_pp_magic(niov_nm);
+ netmem_set_pp(niov_nm, NULL);
+ page_pool_destroy(pool);
+ percpu_ref_exit(&binding.ref);
+}
+#endif
+
static struct kunit_case net_test_cases[] = {
KUNIT_CASE_PARAM(gso_test_func, gso_test_gen_params),
KUNIT_CASE_PARAM(ip_tunnel_flags_test_run,
ip_tunnel_flags_test_gen_params),
+ KUNIT_CASE(netmem_ref_test_page),
+#ifdef CONFIG_PAGE_POOL
+ KUNIT_CASE(netmem_ref_test_pp_page),
+#endif
+#if defined(CONFIG_PAGE_POOL) && defined(CONFIG_NET_DEVMEM)
+ KUNIT_CASE(netmem_ref_test_net_iov),
+ KUNIT_CASE(netmem_ref_test_release_data),
+ KUNIT_CASE(netmem_ref_test_skb_segment),
+#endif
{ },
};
--
2.56.0.385.gd3acb90ef8-goog
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
` (5 preceding siblings ...)
2026-10-10 8:38 ` [PATCH net-next v1 6/6] net: kunit: test netmem, page_pool, and skb frag refcounting Mina Almasry
@ 2026-10-10 8:45 ` netdev-bot+sinfo
2026-10-10 8:51 ` Mina Almasry
6 siblings, 1 reply; 9+ messages in thread
From: netdev-bot+sinfo @ 2026-10-10 8:45 UTC (permalink / raw)
To: Mina Almasry
Cc: netdev, linux-kernel, linux-rdma, bpf, Ayush Sawal, Andrew Lunn,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Tariq Toukan, Simon Horman, Steffen Klassert, Herbert Xu,
Neal Cardwell, Kuniyuki Iwashima, John Fastabend,
Sabrina Dubroca, Eric Biggers, Kees Cook, Michael Grzeschik,
Uwe Kleine-König(The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
Hi!
This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:
- How the issue was discovered, e.g. hit in production, hit during
development, syzbot report, manual code inspection, LLM or static
analysis tool scan.
- Whether the issue was actually triggered, or is only theoretical
(e.g. found by code inspection). If it was triggered please include
the symptoms, like the stack trace or error messages.
- What hardware the change was tested on. For driver fixes please
mention the device (and if relevant firmware version) used for
testing, or say that the change was not tested on real hardware.
Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.
The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting
2026-10-10 8:45 ` [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting netdev-bot+sinfo
@ 2026-10-10 8:51 ` Mina Almasry
0 siblings, 0 replies; 9+ messages in thread
From: Mina Almasry @ 2026-10-10 8:51 UTC (permalink / raw)
To: netdev-bot+sinfo
Cc: netdev, linux-kernel, linux-rdma, bpf, Ayush Sawal, Andrew Lunn,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Tariq Toukan, Simon Horman, Steffen Klassert, Herbert Xu,
Neal Cardwell, Kuniyuki Iwashima, John Fastabend,
Sabrina Dubroca, Eric Biggers, Kees Cook, Michael Grzeschik,
Uwe Kleine-König(The Capable Hub),
Petr Machata, Arend van Spriel, Jakub Raczynski
On Sat, Oct 10, 2026 at 1:45 AM <netdev-bot+sinfo@kernel.org> wrote:
>
> Hi!
>
> This is an automated message. This series looks like a fix, but its
> commit messages seem to be missing some information:
>
> - How the issue was discovered, e.g. hit in production, hit during
> development, syzbot report, manual code inspection, LLM or static
> analysis tool scan.
>
> - Whether the issue was actually triggered, or is only theoretical
> (e.g. found by code inspection). If it was triggered please include
> the symptoms, like the stack trace or error messages.
>
> - What hardware the change was tested on. For driver fixes please
> mention the device (and if relevant firmware version) used for
> testing, or say that the change was not tested on real hardware.
>
> Please do not repost the series just to address the above. Instead,
> reply to this email with the missing information, so that reviewers
> can take it into account. If the series needs another revision for
> other reasons, please include the information in the commit messages
> then.
>
> The evaluation is done by an LLM so it may be wrong, if you think
> that is the case please reply and explain.
I was reminded what I complicated mess skb frag accounting currently
is when I looked at this patch:
[1] https://lore.kernel.org/netdev/xfrm-iptfs-pp_ref_count-underflow-v4-1-912fa72106f0@secunet.com/
And saw that the author was trying to write their own skb frag helper
inside the driver because core didn't provide one that made sense :/
code inspection + LLM written tests for accounting confirmed the
issues. I worked with the LLM to fix the discovered issues. Not hit in
production or reproduced organically. Only code inspection and
artificial tests.
--
Thanks,
Mina
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-10-10 8:51 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 8:38 [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 2/6] net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 3/6] net: skbuff: replace skb_page_unref() and __skb_frag_unref() with skb_netmem_unref() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 4/6] net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 5/6] net: skbuff: use skb_frag_ref() in skb_try_coalesce() Mina Almasry
2026-10-10 8:38 ` [PATCH net-next v1 6/6] net: kunit: test netmem, page_pool, and skb frag refcounting Mina Almasry
2026-10-10 8:45 ` [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting netdev-bot+sinfo
2026-10-10 8:51 ` Mina Almasry
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®