mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Mina Almasry <almasrymina@google.com>
To: netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
	 linux-rdma@vger.kernel.org, bpf@vger.kernel.org
Cc: "Mina Almasry" <almasrymina@google.com>,
	"Ayush Sawal" <ayush.sawal@chelsio.com>,
	"Andrew Lunn" <andrew+netdev@lunn.ch>,
	"David S. Miller" <davem@davemloft.net>,
	"Eric Dumazet" <edumazet@kernel.org>,
	"Jakub Kicinski" <kuba@kernel.org>,
	"Paolo Abeni" <pabeni@redhat.com>,
	"Tariq Toukan" <tariqt@nvidia.com>,
	"Simon Horman" <horms@kernel.org>,
	"Steffen Klassert" <steffen.klassert@secunet.com>,
	"Herbert Xu" <herbert@gondor.apana.org.au>,
	"Neal Cardwell" <ncardwell@google.com>,
	"Kuniyuki Iwashima" <kuniyu@google.com>,
	"John Fastabend" <john.fastabend@gmail.com>,
	"Sabrina Dubroca" <sd@queasysnail.net>,
	"Eric Biggers" <ebiggers@kernel.org>,
	"Kees Cook" <kees@kernel.org>,
	"Michael Grzeschik" <mgr@kernel.org>,
	"Uwe Kleine-König (The Capable Hub)"
	<u.kleine-koenig@baylibre.com>, "Petr Machata" <petrm@nvidia.com>,
	"Arend van Spriel" <arend.vanspriel@broadcom.com>,
	"Jakub Raczynski" <j.raczynski@samsung.com>
Subject: [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting
Date: Sat, 10 Oct 2026 08:38:41 +0000	[thread overview]
Message-ID: <20261010083935.3274178-1-almasrymina@google.com> (raw)

netmems (pages and net_iovs) have two independent refcounts: a non-pp
refcount (page._refcount or binding->ref) governing backing memory
lifetime, and pp_ref_count governing recycling within a page_pool while
page_pool holds a single backing non-pp reference. SKBs with
pp_recycle=1 own pp_ref_count references on page_pool fragments; SKBs
with pp_recycle=0 own non-pp references.

Core helpers currently mix these two refcounts asymmetrically:
skb_frag_ref() always increments the non-pp refcount via get_netmem(),
whereas skb_frag_unref() decrements pp_ref_count when skb->pp_recycle
is set. Any path that refs a fragment on a pp_recycle=1 SKB (or unrefs
with a hardcoded recycle=false) increments one counter and decrements
the other, leaking the backing memory while underflowing pp_ref_count
(e.g., IP-TFS [1], skb_split(), skb_shift(), and cloned SKB uncloning).
Unclear core APIs also led drivers and ULPs to open-code pp_ref_count
manipulations or add one-off helpers like skb_pp_frag_ref().

Unify fragment refcounting into symmetric pairs where each layer clearly
specifies which refcount it touches:

- Non-PP refcount: get/put_page(), get/put_net_iov(), get/put_netmem()
- PP refcount:     napi_pp_get/put_page()
- SKB netmem:      skb_netmem_ref/unref(netmem, recycle)
- SKB fragment:    skb_frag_ref/unref(skb, f)

[1] https://lore.kernel.org/netdev/xfrm-iptfs-pp_ref_count-underflow-v4-1-912fa72106f0@secunet.com/

Mina Almasry (6):
  netmem: rename __get/__put_netmem() to get/put_net_iov()
  net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf()
  net: skbuff: replace skb_page_unref() and __skb_frag_unref() with
    skb_netmem_unref()
  net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref()
  net: skbuff: use skb_frag_ref() in skb_try_coalesce()
  net: kunit: test netmem, page_pool, and skb frag refcounting

 .../chelsio/inline_crypto/ch_ktls/chcr_ktls.c |   2 +-
 drivers/net/ethernet/marvell/sky2.c           |   2 +-
 drivers/net/ethernet/mellanox/mlx4/en_rx.c    |   2 +-
 drivers/net/ethernet/sun/cassini.c            |   4 +-
 drivers/net/veth.c                            |   2 +-
 include/linux/skbuff_ref.h                    |  64 +-
 include/net/netmem.h                          |  24 +-
 net/core/net_test.c                           | 747 ++++++++++++++++++
 net/core/skbuff.c                             | 145 ++--
 net/ipv4/esp4.c                               |   4 +-
 net/ipv4/tcp.c                                |   6 +-
 net/ipv6/esp6.c                               |   4 +-
 net/tls/tls_device.c                          |   2 +-
 net/tls/tls_device_fallback.c                 |   2 +-
 net/tls/tls_strp.c                            |   2 +-
 net/xfrm/xfrm_iptfs.c                         |   3 +-
 16 files changed, 900 insertions(+), 115 deletions(-)


base-commit: d8674294aefef02266c4d47ad10131f1bffbe534
-- 
2.56.0.385.gd3acb90ef8-goog


             reply	other threads:[~2026-10-10  8:39 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-10  8:38 Mina Almasry [this message]
2026-10-10  8:38 ` [PATCH net-next v1 1/6] netmem: rename __get/__put_netmem() to get/put_net_iov() Mina Almasry
2026-10-10  8:38 ` [PATCH net-next v1 2/6] net: skbuff: add napi_pp_get_page() and use it in tcp_recvmsg_dmabuf() Mina Almasry
2026-10-10  8:38 ` [PATCH net-next v1 3/6] net: skbuff: replace skb_page_unref() and __skb_frag_unref() with skb_netmem_unref() Mina Almasry
2026-10-10  8:38 ` [PATCH net-next v1 4/6] net: skbuff: replace __skb_frag_ref() with symmetric skb_netmem_ref() Mina Almasry
2026-10-10  8:38 ` [PATCH net-next v1 5/6] net: skbuff: use skb_frag_ref() in skb_try_coalesce() Mina Almasry
2026-10-10  8:38 ` [PATCH net-next v1 6/6] net: kunit: test netmem, page_pool, and skb frag refcounting Mina Almasry
2026-10-10  8:45 ` [PATCH net-next v1 0/6] net: unify symmetric netmem and page_pool refcounting netdev-bot+sinfo
2026-10-10  8:51   ` Mina Almasry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261010083935.3274178-1-almasrymina@google.com \
    --to=almasrymina@google.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=arend.vanspriel@broadcom.com \
    --cc=ayush.sawal@chelsio.com \
    --cc=bpf@vger.kernel.org \
    --cc=davem@davemloft.net \
    --cc=ebiggers@kernel.org \
    --cc=edumazet@kernel.org \
    --cc=herbert@gondor.apana.org.au \
    --cc=horms@kernel.org \
    --cc=j.raczynski@samsung.com \
    --cc=john.fastabend@gmail.com \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=mgr@kernel.org \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=petrm@nvidia.com \
    --cc=sd@queasysnail.net \
    --cc=steffen.klassert@secunet.com \
    --cc=tariqt@nvidia.com \
    --cc=u.kleine-koenig@baylibre.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®