* [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations
@ 2026-09-11 15:45 Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
Replace devmem's gen_pool based fixed-size allocator with a binding-level
freelist similar to the one used by io_uring zero-copy receive.
This is motivated by allocation latency observed in the NAPI receive path:
[ 1036.228913] ? gen_pool_create+0x90/0x90
[ 1036.228915] net_devmem_alloc_dmabuf+0x1f/0x60
[ 1036.228918] mp_dmabuf_devmem_alloc_netmems+0x17/0x80
[ 1036.228920] mlx5e_post_rx_mpwqes+0xdbe/0xdd0
[ 1036.228926] mlx5e_napi_poll+0x113/0x830
[ 1036.228928] ? sched_clock+0x5/0x10
[ 1036.228931] ? wake_up_process+0x778/0x14b0
[ 1036.228933] net_rx_action+0x15d/0x570
[ 1036.228934] ? update_rq_clock+0x31/0x240
[ 1036.228937] ? __napi_schedule+0x55/0xa0
[ 1036.228938] ? mlx5_eq_comp_int+0x137/0x230
[ 1036.228940] ? atomic_notifier_call_chain+0x36/0x90
[ 1036.228943] ? sched_clock+0x5/0x10
[ 1036.228944] ? sched_clock_cpu+0xc/0x170
[ 1036.228947] irq_exit_rcu+0x12b/0x370
[ 1036.228950] common_interrupt+0x85/0x90
udmabuf can create a very large number of SG entries. In the worst case,
devmem ends up adding one gen_pool chunk for each net_iov allocation
unit backed by those entries. The gen_pool allocation path then has to
traverse a linked list that can become too long for this hot path.
Patch 1 removes the gen_pool and replaces it with a simple freelist of
net_iov pointers protected by the same spin_lock_bh() pattern used by
io_uring zcrx. Patch 2 removes the now-unnecessary chunk owner wrapper by
embedding the net_iov_area directly in the dma-buf binding. Patch 3
batches freelist allocations.
= Performance:
kperf/client ... \
--num-rx-queues 4 \
--dmabuf-rx-size-mb 2048 \
--dmabuf-tx-size-mb 2048 \
--validate no \
--time 60 \
--read-size 67108864 \
--write-size 67108864 \
--num-connections 4 \
--tcp-cc dctcp \
--pin-off 4 \
--devmem-rx \
--devmem-tx \
--devmem-rx-memory cuda \
--devmem-tx-memory cuda
With 4 queues, 4 flows, 2GB BB, cuda for both rx and tx I see no difference
in throughput or cpu utilization (see selective runs below).
== Before
10 runs: 206.031 243.546 293.931 319.015 319.923 323.189 324.321 325.499 327.134 349.450 Gbps
Sample:
client: == Source <redacted>
client: Tx 48.170 Gbps (361716776960 bytes in 60072872 usec)
client: Tx101.256 Gbps (760343429120 bytes in 60072872 usec)
client: Tx101.077 Gbps (759001251840 bytes in 60072872 usec)
client: Tx 69.440 Gbps (521435873280 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Rx 48.158 Gbps (361638901920 bytes in 60074846 usec)
client: Rx101.253 Gbps (760343429120 bytes in 60074846 usec)
client: Rx101.074 Gbps (759001251840 bytes in 60074846 usec)
client: Rx 69.438 Gbps (521435873280 bytes in 60074846 usec)
client: net CPU 1: usr: 0.00% sys: 0.01% idle:39.42% iow: 0.00% irq: 0.64% sirq:59.91%
client: app CPU 5: usr: 1.21% sys:97.71% idle: 0.09% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 2: usr: 0.00% sys: 0.00% idle:38.19% iow: 0.00% irq: 0.84% sirq:60.96%
client: app CPU 6: usr: 1.24% sys:97.73% idle: 0.04% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 0: usr: 0.05% sys: 0.27% idle:69.92% iow: 0.00% irq: 7.71% sirq:22.04%
client: app CPU 4: usr: 1.78% sys:80.71% idle:16.79% iow: 0.00% irq: 0.38% sirq: 0.32%
client: net CPU 3: usr: 0.00% sys: 0.00% idle: 0.00% iow: 0.00% irq: 0.39% sirq:99.60%
client: app CPU 7: usr: 0.21% sys: 6.64% idle:90.75% iow: 0.00% irq: 0.05% sirq: 2.32%
== After
10 runs: 213.032 226.471 235.052 257.113 323.233 331.597 342.893 348.188 351.866 354.311 Gbps
Sample:
client: == Source <redacted>
client: Tx104.829 Gbps (786515886080 bytes in 60022934 usec)
client: Tx105.455 Gbps (791213506560 bytes in 60022934 usec)
client: Tx 61.717 Gbps (463051161600 bytes in 60022934 usec)
client: Tx 51.430 Gbps (385875968000 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Rx104.767 Gbps (786515886080 bytes in 60058283 usec)
client: Rx105.393 Gbps (791213506560 bytes in 60058283 usec)
client: Rx 61.673 Gbps (462998508768 bytes in 60058283 usec)
client: Rx 51.400 Gbps (385875968000 bytes in 60058283 usec)
client: net CPU 0: usr: 0.01% sys: 0.21% idle:65.30% iow: 0.00% irq: 9.15% sirq:25.29%
client: app CPU 4: usr: 1.19% sys:98.00% idle: 0.03% iow: 0.00% irq: 0.21% sirq: 0.54%
client: net CPU 2: usr: 0.00% sys: 0.00% idle: 0.03% iow: 0.00% irq: 0.44% sirq:99.51%
client: app CPU 6: usr: 2.14% sys:77.48% idle:19.92% iow: 0.00% irq: 0.36% sirq: 0.07%
client: net CPU 3: usr: 0.00% sys: 0.00% idle:43.01% iow: 0.00% irq: 0.69% sirq:56.28%
client: app CPU 7: usr: 2.03% sys:71.04% idle:26.45% iow: 0.00% irq: 0.37% sirq: 0.08%
client: net CPU 1: usr: 0.00% sys: 0.01% idle:46.66% iow: 0.00% irq: 0.56% sirq:52.75%
client: app CPU 5: usr: 0.05% sys: 0.08% idle:99.21% iow: 0.00% irq: 0.03% sirq: 0.61%
== Comparison, over 10 runs
Median Target RX: 321.556 Gbps vs 327.415 Gbps
Mean Target RX: 303.204 Gbps vs 298.376 Gbps
Range: 206.031-349.450 Gbps vs 213.032-354.311 Gbps
v2:
- xmas tree (Jakub)
- batching (Mina)
- perf numbers (Mina & Jakub)
Stanislav Fomichev (3):
net: devmem: replace gen_pool with freelist
net: devmem: embed net_iov_area in binding
net: devmem: batch net_iov allocations into the page_pool cache
net/Kconfig | 1 -
net/core/devmem.c | 208 +++++++++++++++++++++-------------------------
net/core/devmem.h | 46 +++-------
3 files changed, 105 insertions(+), 150 deletions(-)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2026-09-14 21:46 ` Mina Almasry
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
2 siblings, 1 reply; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
devmem only needs fixed-size net_iov allocations for each dma-buf binding.
The gen_pool tracks the same free set indirectly through DMA addresses,
which makes devmem depend on the generic allocator even though the users
are fixed-size net_iov chunks.
Mirror the io_uring zcrx model more closely by keeping a binding-level
freelist protected by spin_lock_bh(). Use a single net_iov_area owner for
the binding, populate each net_iov's DMA address while walking the SG
table, and check at teardown that all net_iovs have returned to the
freelist.
Drop the NET_DEVMEM select of GENERIC_ALLOCATOR now that devmem no longer
calls gen_pool APIs.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/Kconfig | 1 -
net/core/devmem.c | 172 +++++++++++++++++++++-------------------------
net/core/devmem.h | 16 ++---
3 files changed, 85 insertions(+), 104 deletions(-)
diff --git a/net/Kconfig b/net/Kconfig
index e38477393551..76ab44aa439a 100644
--- a/net/Kconfig
+++ b/net/Kconfig
@@ -68,7 +68,6 @@ config SKB_EXTENSIONS
config NET_DEVMEM
def_bool y
- select GENERIC_ALLOCATOR
depends on DMA_SHARED_BUFFER
depends on PAGE_POOL
diff --git a/net/core/devmem.c b/net/core/devmem.c
index f4d60654ce7f..4883eb7f3a95 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -8,7 +8,6 @@
*/
#include <linux/dma-buf.h>
-#include <linux/genalloc.h>
#include <linux/mm.h>
#include <linux/netdevice.h>
#include <linux/types.h>
@@ -30,23 +29,13 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
static const struct memory_provider_ops dmabuf_devmem_ops;
-static void net_devmem_dmabuf_free_chunk_owner(struct gen_pool *genpool,
- struct gen_pool_chunk *chunk,
- void *not_used)
+static void
+net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
{
- struct dmabuf_genpool_chunk_owner *owner = chunk->owner;
-
- kvfree(owner->area.niovs);
- kfree(owner);
-}
-
-static dma_addr_t net_devmem_get_dma_addr(const struct net_iov *niov)
-{
- struct dmabuf_genpool_chunk_owner *owner;
-
- owner = net_devmem_iov_to_chunk_owner(niov);
- return owner->base_dma_addr +
- ((dma_addr_t)net_iov_idx(niov) << owner->binding->niov_shift);
+ if (owner) {
+ kvfree(owner->area.niovs);
+ kfree(owner);
+ }
}
static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
@@ -62,24 +51,18 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
{
struct net_devmem_dmabuf_binding *binding = container_of(wq, typeof(*binding), unbind_w);
- size_t size, avail;
-
- gen_pool_for_each_chunk(binding->chunk_pool,
- net_devmem_dmabuf_free_chunk_owner, NULL);
-
- size = gen_pool_size(binding->chunk_pool);
- avail = gen_pool_avail(binding->chunk_pool);
-
- if (!WARN(size != avail, "can't destroy genpool. size=%zu, avail=%zu",
- size, avail))
- gen_pool_destroy(binding->chunk_pool);
+ WARN(binding->free_count != binding->total_niovs,
+ "can't destroy dmabuf binding. total=%zu, free=%zu",
+ binding->total_niovs, binding->free_count);
+ net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
binding->direction);
dma_buf_detach(binding->dmabuf, binding->attachment);
dma_buf_put(binding->dmabuf);
xa_destroy(&binding->bound_rxqs);
percpu_ref_exit(&binding->ref);
+ kvfree(binding->freelist);
kvfree(binding->tx_vec);
kfree(binding);
}
@@ -87,21 +70,16 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
struct net_iov *
net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
{
- struct dmabuf_genpool_chunk_owner *owner;
- unsigned long dma_addr;
struct net_iov *niov;
- ssize_t offset;
- ssize_t index;
-
- dma_addr = gen_pool_alloc_owner(binding->chunk_pool,
- 1UL << binding->niov_shift,
- (void **)&owner);
- if (!dma_addr)
+ spin_lock_bh(&binding->freelist_lock);
+ if (unlikely(!binding->free_count)) {
+ spin_unlock_bh(&binding->freelist_lock);
return NULL;
+ }
- offset = dma_addr - owner->base_dma_addr;
- index = offset >> binding->niov_shift;
- niov = &owner->area.niovs[index];
+ niov = binding->freelist[--binding->free_count];
+ binding->freelist[binding->free_count] = NULL;
+ spin_unlock_bh(&binding->freelist_lock);
niov->desc.pp_magic = 0;
niov->desc.pp = NULL;
@@ -113,14 +91,15 @@ net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
void net_devmem_free_dmabuf(struct net_iov *niov)
{
struct net_devmem_dmabuf_binding *binding = net_devmem_iov_binding(niov);
- unsigned long dma_addr = net_devmem_get_dma_addr(niov);
- size_t niov_size = 1UL << binding->niov_shift;
- if (WARN_ON(!gen_pool_has_addr(binding->chunk_pool, dma_addr,
- niov_size)))
+ spin_lock_bh(&binding->freelist_lock);
+ if (WARN_ON_ONCE(binding->free_count >= binding->total_niovs)) {
+ spin_unlock_bh(&binding->freelist_lock);
return;
+ }
- gen_pool_free(binding->chunk_pool, dma_addr, niov_size);
+ binding->freelist[binding->free_count++] = niov;
+ spin_unlock_bh(&binding->freelist_lock);
}
void net_devmem_unbind_dmabuf(struct net_devmem_dmabuf_binding *binding)
@@ -194,12 +173,15 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
struct netlink_ext_ack *extack)
{
struct net_devmem_dmabuf_binding *binding;
+ struct dmabuf_genpool_chunk_owner *owner;
size_t niov_size = 1UL << niov_shift;
static u32 id_alloc_next;
struct scatterlist *sg;
struct dma_buf *dmabuf;
- unsigned int sg_idx, i;
- unsigned long virtual;
+ unsigned int sg_idx;
+ size_t total_niovs;
+ size_t niov_idx;
+ size_t i;
int err;
if (!dma_dev) {
@@ -230,6 +212,7 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
goto err_free_binding;
mutex_init(&binding->lock);
+ spin_lock_init(&binding->freelist_lock);
binding->dmabuf = dmabuf;
binding->direction = direction;
@@ -262,20 +245,10 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
goto err_unmap;
}
}
-
- binding->chunk_pool = gen_pool_create(niov_shift,
- dev_to_node(&dev->dev));
- if (!binding->chunk_pool) {
- err = -ENOMEM;
- goto err_tx_vec;
- }
-
- virtual = 0;
+ total_niovs = 0;
for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
dma_addr_t dma_addr = sg_dma_address(sg);
- struct dmabuf_genpool_chunk_owner *owner;
size_t len = sg_dma_len(sg);
- struct net_iov *niov;
if (!IS_ALIGNED(dma_addr, niov_size) ||
!IS_ALIGNED(len, niov_size)) {
@@ -283,63 +256,74 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
NL_SET_ERR_MSG_FMT(extack,
"dmabuf sg entry (addr=%pad, len=%zu) not aligned to niov size %zu",
&dma_addr, len, niov_size);
- goto err_free_chunks;
+ goto err_tx_vec;
}
- owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
- dev_to_node(&dev->dev));
- if (!owner) {
- err = -ENOMEM;
- goto err_free_chunks;
- }
+ total_niovs += len >> niov_shift;
+ }
- owner->area.base_virtual = virtual;
- owner->base_dma_addr = dma_addr;
- owner->area.num_niovs = len >> niov_shift;
- owner->binding = binding;
+ binding->freelist = kvmalloc_array(total_niovs,
+ sizeof(binding->freelist[0]),
+ GFP_KERNEL);
+ if (!binding->freelist) {
+ err = -ENOMEM;
+ goto err_tx_vec;
+ }
+ binding->total_niovs = total_niovs;
- err = gen_pool_add_owner(binding->chunk_pool, dma_addr,
- dma_addr, len, dev_to_node(&dev->dev),
- owner);
- if (err) {
- kfree(owner);
- err = -EINVAL;
- goto err_free_chunks;
- }
+ owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
+ dev_to_node(&dev->dev));
+ if (!owner) {
+ err = -ENOMEM;
+ goto err_free_freelist;
+ }
- owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
- owner->area.num_niovs);
- if (!owner->area.niovs) {
- err = -ENOMEM;
- goto err_free_chunks;
- }
+ owner->area.num_niovs = total_niovs;
+ owner->binding = binding;
+ owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
+ owner->area.num_niovs);
+ if (!owner->area.niovs) {
+ err = -ENOMEM;
+ goto err_free_owner;
+ }
+ binding->chunk_owner = owner;
+
+ niov_idx = 0;
+ for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
+ dma_addr_t dma_addr = sg_dma_address(sg);
+ size_t len = sg_dma_len(sg);
+ struct net_iov *niov;
+ size_t nr_niovs = len >> niov_shift;
- for (i = 0; i < owner->area.num_niovs; i++) {
- niov = &owner->area.niovs[i];
+ for (i = 0; i < nr_niovs; i++, niov_idx++) {
+ niov = &owner->area.niovs[niov_idx];
net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
- net_devmem_get_dma_addr(niov));
+ dma_addr);
if (direction == DMA_TO_DEVICE)
- binding->tx_vec[owner->area.base_virtual / PAGE_SIZE + i] = niov;
+ binding->tx_vec[niov_idx] = niov;
+ binding->freelist[binding->free_count++] = niov;
+ dma_addr += niov_size;
}
-
- virtual += len;
}
err = xa_alloc_cyclic(&net_devmem_dmabuf_bindings, &binding->id,
binding, xa_limit_32b, &id_alloc_next,
GFP_KERNEL);
if (err < 0)
- goto err_free_chunks;
+ goto err_free_chunk_owner;
list_add(&binding->list, &priv->bindings);
return binding;
-err_free_chunks:
- gen_pool_for_each_chunk(binding->chunk_pool,
- net_devmem_dmabuf_free_chunk_owner, NULL);
- gen_pool_destroy(binding->chunk_pool);
+err_free_chunk_owner:
+ net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
+ goto err_free_freelist;
+err_free_owner:
+ kfree(owner);
+err_free_freelist:
+ kvfree(binding->freelist);
err_tx_vec:
kvfree(binding->tx_vec);
err_unmap:
diff --git a/net/core/devmem.h b/net/core/devmem.h
index 4a293a7d1149..a5ee2d8d9169 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -14,6 +14,7 @@
#include <net/netdev_netlink.h>
struct netlink_ext_ack;
+struct dmabuf_genpool_chunk_owner;
struct net_devmem_dmabuf_binding {
struct dma_buf *dmabuf;
@@ -26,7 +27,7 @@ struct net_devmem_dmabuf_binding {
* dereferenced.
*/
void *vdev;
- struct gen_pool *chunk_pool;
+ struct dmabuf_genpool_chunk_owner *chunk_owner;
/* Protect dev */
struct mutex lock;
@@ -57,6 +58,11 @@ struct net_devmem_dmabuf_binding {
/* rxq's this binding is active on. */
struct xarray bound_rxqs;
+ spinlock_t freelist_lock ____cacheline_aligned_in_smp;
+ size_t free_count;
+ size_t total_niovs;
+ struct net_iov **freelist;
+
/* ID of this binding. Globally unique to all bindings currently
* active.
*/
@@ -77,17 +83,9 @@ struct net_devmem_dmabuf_binding {
};
#if defined(CONFIG_NET_DEVMEM)
-/* Owner of the dma-buf chunks inserted into the gen pool. Each scatterlist
- * entry from the dmabuf is inserted into the genpool as a chunk, and needs
- * this owner struct to keep track of some metadata necessary to create
- * allocations from this chunk.
- */
struct dmabuf_genpool_chunk_owner {
struct net_iov_area area;
struct net_devmem_dmabuf_binding *binding;
-
- /* dma_addr of the start of the chunk. */
- dma_addr_t base_dma_addr;
};
void __net_devmem_dmabuf_binding_free(struct work_struct *wq);
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2026-09-14 19:37 ` Stanislav Fomichev
2026-09-14 21:53 ` Mina Almasry
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
2 siblings, 2 replies; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
After replacing the gen_pool with a binding-level freelist, devmem no
longer needs a separate chunk owner object. There is only one
net_iov_area for the binding, so store it directly in struct
net_devmem_dmabuf_binding.
Derive the binding from net_iov_owner() with container_of(), matching the
pattern used by io_uring zcrx. This removes the leftover
dmabuf_genpool_chunk_owner wrapper and its allocation/free path.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/core/devmem.c | 41 ++++++++++-------------------------------
net/core/devmem.h | 26 +++++++-------------------
2 files changed, 17 insertions(+), 50 deletions(-)
diff --git a/net/core/devmem.c b/net/core/devmem.c
index 4883eb7f3a95..7949f8425bcd 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
static const struct memory_provider_ops dmabuf_devmem_ops;
-static void
-net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
-{
- if (owner) {
- kvfree(owner->area.niovs);
- kfree(owner);
- }
-}
-
static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
{
struct net_devmem_dmabuf_binding *binding =
@@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
"can't destroy dmabuf binding. total=%zu, free=%zu",
binding->total_niovs, binding->free_count);
- net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
+ kvfree(binding->area.niovs);
dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
binding->direction);
dma_buf_detach(binding->dmabuf, binding->attachment);
@@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
}
binding->total_niovs = total_niovs;
- owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
- dev_to_node(&dev->dev));
- if (!owner) {
+ binding->area.num_niovs = total_niovs;
+ binding->area.niovs = kvmalloc_objs(*binding->area.niovs,
+ binding->area.num_niovs);
+ if (!binding->area.niovs) {
err = -ENOMEM;
goto err_free_freelist;
}
- owner->area.num_niovs = total_niovs;
- owner->binding = binding;
- owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
- owner->area.num_niovs);
- if (!owner->area.niovs) {
- err = -ENOMEM;
- goto err_free_owner;
- }
- binding->chunk_owner = owner;
-
niov_idx = 0;
for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
dma_addr_t dma_addr = sg_dma_address(sg);
@@ -296,8 +278,8 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
size_t nr_niovs = len >> niov_shift;
for (i = 0; i < nr_niovs; i++, niov_idx++) {
- niov = &owner->area.niovs[niov_idx];
- net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
+ niov = &binding->area.niovs[niov_idx];
+ net_iov_init(niov, &binding->area, NET_IOV_DMABUF);
page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
dma_addr);
if (direction == DMA_TO_DEVICE)
@@ -311,17 +293,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
binding, xa_limit_32b, &id_alloc_next,
GFP_KERNEL);
if (err < 0)
- goto err_free_chunk_owner;
+ goto err_free_niovs;
list_add(&binding->list, &priv->bindings);
return binding;
-err_free_chunk_owner:
- net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
- goto err_free_freelist;
-err_free_owner:
- kfree(owner);
+err_free_niovs:
+ kvfree(binding->area.niovs);
err_free_freelist:
kvfree(binding->freelist);
err_tx_vec:
diff --git a/net/core/devmem.h b/net/core/devmem.h
index a5ee2d8d9169..20a3eb90ea7f 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -14,9 +14,9 @@
#include <net/netdev_netlink.h>
struct netlink_ext_ack;
-struct dmabuf_genpool_chunk_owner;
struct net_devmem_dmabuf_binding {
+ struct net_iov_area area;
struct dma_buf *dmabuf;
struct dma_buf_attachment *attachment;
struct sg_table *sgt;
@@ -27,7 +27,6 @@ struct net_devmem_dmabuf_binding {
* dereferenced.
*/
void *vdev;
- struct dmabuf_genpool_chunk_owner *chunk_owner;
/* Protect dev */
struct mutex lock;
@@ -83,11 +82,6 @@ struct net_devmem_dmabuf_binding {
};
#if defined(CONFIG_NET_DEVMEM)
-struct dmabuf_genpool_chunk_owner {
- struct net_iov_area area;
- struct net_devmem_dmabuf_binding *binding;
-};
-
void __net_devmem_dmabuf_binding_free(struct work_struct *wq);
struct net_devmem_dmabuf_binding *
net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
@@ -102,18 +96,12 @@ int net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx,
struct net_devmem_dmabuf_binding *binding,
struct netlink_ext_ack *extack);
-static inline struct dmabuf_genpool_chunk_owner *
-net_devmem_iov_to_chunk_owner(const struct net_iov *niov)
-{
- struct net_iov_area *owner = net_iov_owner(niov);
-
- return container_of(owner, struct dmabuf_genpool_chunk_owner, area);
-}
-
static inline struct net_devmem_dmabuf_binding *
net_devmem_iov_binding(const struct net_iov *niov)
{
- return net_devmem_iov_to_chunk_owner(niov)->binding;
+ struct net_iov_area *owner = net_iov_owner(niov);
+
+ return container_of(owner, struct net_devmem_dmabuf_binding, area);
}
static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov)
@@ -123,11 +111,11 @@ static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov)
static inline unsigned long net_iov_virtual_addr(const struct net_iov *niov)
{
- struct dmabuf_genpool_chunk_owner *co =
- net_devmem_iov_to_chunk_owner(niov);
+ struct net_devmem_dmabuf_binding *binding =
+ net_devmem_iov_binding(niov);
return net_iov_owner(niov)->base_virtual +
- ((unsigned long)net_iov_idx(niov) << co->binding->niov_shift);
+ ((unsigned long)net_iov_idx(niov) << binding->niov_shift);
}
static inline bool
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2026-09-14 21:58 ` Mina Almasry
2 siblings, 1 reply; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
Rename net_devmem_alloc_dmabuf() into net_devmem_alloc_dmabuf_bulk() and
make it refill page pool with up to PP_ALLOC_CACHE_REFILL NIOVs,
similar to io_pp_zc_alloc_netmems(). That should amortize recently
introduced freelist_lock.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/core/devmem.c | 59 +++++++++++++++++++++++++++++------------------
net/core/devmem.h | 8 -------
2 files changed, 37 insertions(+), 30 deletions(-)
diff --git a/net/core/devmem.c b/net/core/devmem.c
index 7949f8425bcd..a0dcc896dd12 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -58,25 +58,25 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
kfree(binding);
}
-struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
+static unsigned int
+net_devmem_alloc_dmabuf_bulk(struct net_devmem_dmabuf_binding *binding,
+ netmem_ref *netmems, unsigned int count)
{
- struct net_iov *niov;
+ unsigned int i;
+
spin_lock_bh(&binding->freelist_lock);
- if (unlikely(!binding->free_count)) {
- spin_unlock_bh(&binding->freelist_lock);
- return NULL;
+
+ count = min_t(size_t, count, binding->free_count);
+ for (i = 0; i < count; i++) {
+ struct net_iov *niov = binding->freelist[--binding->free_count];
+
+ binding->freelist[binding->free_count] = NULL;
+ netmems[i] = net_iov_to_netmem(niov);
}
- niov = binding->freelist[--binding->free_count];
- binding->freelist[binding->free_count] = NULL;
spin_unlock_bh(&binding->freelist_lock);
- niov->desc.pp_magic = 0;
- niov->desc.pp = NULL;
- atomic_long_set(&niov->desc.pp_ref_count, 0);
-
- return niov;
+ return count;
}
void net_devmem_free_dmabuf(struct net_iov *niov)
@@ -434,20 +434,35 @@ int mp_dmabuf_devmem_init(struct page_pool *pool)
netmem_ref mp_dmabuf_devmem_alloc_netmems(struct page_pool *pool, gfp_t gfp)
{
struct net_devmem_dmabuf_binding *binding = pool->mp_priv;
- struct net_iov *niov;
- netmem_ref netmem;
+ netmem_ref *netmems = pool->alloc.cache;
+ unsigned int allocated, i;
+
+ if (WARN_ON_ONCE(pool->alloc.count))
+ return 0;
- niov = net_devmem_alloc_dmabuf(binding);
- if (!niov)
+ allocated = net_devmem_alloc_dmabuf_bulk(binding, netmems,
+ PP_ALLOC_CACHE_REFILL);
+ if (unlikely(!allocated))
return 0;
- netmem = net_iov_to_netmem(niov);
+ for (i = 0; i < allocated; i++) {
+ struct net_iov *niov = netmem_to_net_iov(netmems[i]);
- page_pool_set_pp_info(pool, netmem);
+ niov->desc.pp_magic = 0;
+ niov->desc.pp = NULL;
+ atomic_long_set(&niov->desc.pp_ref_count, 0);
+
+ page_pool_set_pp_info(pool, netmems[i]);
+
+ pool->pages_state_hold_cnt++;
+ trace_page_pool_state_hold(pool, netmems[i],
+ pool->pages_state_hold_cnt);
+ }
- pool->pages_state_hold_cnt++;
- trace_page_pool_state_hold(pool, netmem, pool->pages_state_hold_cnt);
- return netmem;
+ /* Return the last one, the rest stay in the page_pool cache. */
+ allocated--;
+ pool->alloc.count = allocated;
+ return netmems[allocated];
}
void mp_dmabuf_devmem_destroy(struct page_pool *pool)
diff --git a/net/core/devmem.h b/net/core/devmem.h
index 20a3eb90ea7f..7195769b8bd1 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -133,8 +133,6 @@ net_devmem_dmabuf_binding_put(struct net_devmem_dmabuf_binding *binding)
void net_devmem_get_net_iov(struct net_iov *niov);
void net_devmem_put_net_iov(struct net_iov *niov);
-struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding);
void net_devmem_free_dmabuf(struct net_iov *ppiov);
@@ -191,12 +189,6 @@ net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx,
return -EOPNOTSUPP;
}
-static inline struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
-{
- return NULL;
-}
-
static inline void net_devmem_free_dmabuf(struct net_iov *ppiov)
{
}
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
@ 2026-09-14 19:37 ` Stanislav Fomichev
2026-09-14 21:53 ` Mina Almasry
1 sibling, 0 replies; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-14 19:37 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
On 09/11, Stanislav Fomichev wrote:
> After replacing the gen_pool with a binding-level freelist, devmem no
> longer needs a separate chunk owner object. There is only one
> net_iov_area for the binding, so store it directly in struct
> net_devmem_dmabuf_binding.
>
> Derive the binding from net_iov_owner() with container_of(), matching the
> pattern used by io_uring zcrx. This removes the leftover
> dmabuf_genpool_chunk_owner wrapper and its allocation/free path.
>
> Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
> ---
> net/core/devmem.c | 41 ++++++++++-------------------------------
> net/core/devmem.h | 26 +++++++-------------------
> 2 files changed, 17 insertions(+), 50 deletions(-)
>
> diff --git a/net/core/devmem.c b/net/core/devmem.c
> index 4883eb7f3a95..7949f8425bcd 100644
> --- a/net/core/devmem.c
> +++ b/net/core/devmem.c
> @@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
>
> static const struct memory_provider_ops dmabuf_devmem_ops;
>
> -static void
> -net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
> -{
> - if (owner) {
> - kvfree(owner->area.niovs);
> - kfree(owner);
> - }
> -}
> -
> static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
> {
> struct net_devmem_dmabuf_binding *binding =
> @@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
> "can't destroy dmabuf binding. total=%zu, free=%zu",
> binding->total_niovs, binding->free_count);
>
> - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
> + kvfree(binding->area.niovs);
> dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
> binding->direction);
> dma_buf_detach(binding->dmabuf, binding->attachment);
> @@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> }
> binding->total_niovs = total_niovs;
>
> - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
> - dev_to_node(&dev->dev));
There is a leftover owner var, the build fails, will repost later this week
(to give some time to review for whoever is interested).
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
@ 2026-09-14 21:46 ` Mina Almasry
0 siblings, 0 replies; 8+ messages in thread
From: Mina Almasry @ 2026-09-14 21:46 UTC (permalink / raw)
To: Stanislav Fomichev
Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
linux-kernel
On Fri, Sep 11, 2026 at 8:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> devmem only needs fixed-size net_iov allocations for each dma-buf binding.
> The gen_pool tracks the same free set indirectly through DMA addresses,
> which makes devmem depend on the generic allocator even though the users
> are fixed-size net_iov chunks.
>
> Mirror the io_uring zcrx model more closely by keeping a binding-level
> freelist protected by spin_lock_bh(). Use a single net_iov_area owner for
> the binding, populate each net_iov's DMA address while walking the SG
> table, and check at teardown that all net_iovs have returned to the
> freelist.
>
> Drop the NET_DEVMEM select of GENERIC_ALLOCATOR now that devmem no longer
> calls gen_pool APIs.
>
> Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
Approach is great, some suggested improvements.
> ---
> net/Kconfig | 1 -
> net/core/devmem.c | 172 +++++++++++++++++++++-------------------------
> net/core/devmem.h | 16 ++---
> 3 files changed, 85 insertions(+), 104 deletions(-)
>
> diff --git a/net/Kconfig b/net/Kconfig
> index e38477393551..76ab44aa439a 100644
> --- a/net/Kconfig
> +++ b/net/Kconfig
> @@ -68,7 +68,6 @@ config SKB_EXTENSIONS
>
> config NET_DEVMEM
> def_bool y
> - select GENERIC_ALLOCATOR
> depends on DMA_SHARED_BUFFER
> depends on PAGE_POOL
>
> diff --git a/net/core/devmem.c b/net/core/devmem.c
> index f4d60654ce7f..4883eb7f3a95 100644
> --- a/net/core/devmem.c
> +++ b/net/core/devmem.c
> @@ -8,7 +8,6 @@
> */
>
> #include <linux/dma-buf.h>
> -#include <linux/genalloc.h>
> #include <linux/mm.h>
> #include <linux/netdevice.h>
> #include <linux/types.h>
> @@ -30,23 +29,13 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
>
> static const struct memory_provider_ops dmabuf_devmem_ops;
>
> -static void net_devmem_dmabuf_free_chunk_owner(struct gen_pool *genpool,
> - struct gen_pool_chunk *chunk,
> - void *not_used)
> +static void
> +net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
> {
> - struct dmabuf_genpool_chunk_owner *owner = chunk->owner;
> -
> - kvfree(owner->area.niovs);
> - kfree(owner);
> -}
> -
> -static dma_addr_t net_devmem_get_dma_addr(const struct net_iov *niov)
> -{
> - struct dmabuf_genpool_chunk_owner *owner;
> -
> - owner = net_devmem_iov_to_chunk_owner(niov);
> - return owner->base_dma_addr +
> - ((dma_addr_t)net_iov_idx(niov) << owner->binding->niov_shift);
> + if (owner) {
> + kvfree(owner->area.niovs);
> + kfree(owner);
> + }
> }
>
> static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
> @@ -62,24 +51,18 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
> {
> struct net_devmem_dmabuf_binding *binding = container_of(wq, typeof(*binding), unbind_w);
>
> - size_t size, avail;
> -
> - gen_pool_for_each_chunk(binding->chunk_pool,
> - net_devmem_dmabuf_free_chunk_owner, NULL);
> -
> - size = gen_pool_size(binding->chunk_pool);
> - avail = gen_pool_avail(binding->chunk_pool);
> -
> - if (!WARN(size != avail, "can't destroy genpool. size=%zu, avail=%zu",
> - size, avail))
> - gen_pool_destroy(binding->chunk_pool);
> + WARN(binding->free_count != binding->total_niovs,
> + "can't destroy dmabuf binding. total=%zu, free=%zu",
> + binding->total_niovs, binding->free_count);
>
You're warning here that you can't destroy the dmabuf binding but
you're destroying it anyway. Something is off here. Do we want an
early return or something else?
> + net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
> dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
> binding->direction);
> dma_buf_detach(binding->dmabuf, binding->attachment);
> dma_buf_put(binding->dmabuf);
> xa_destroy(&binding->bound_rxqs);
> percpu_ref_exit(&binding->ref);
> + kvfree(binding->freelist);
> kvfree(binding->tx_vec);
> kfree(binding);
> }
> @@ -87,21 +70,16 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
> struct net_iov *
> net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
> {
> - struct dmabuf_genpool_chunk_owner *owner;
> - unsigned long dma_addr;
> struct net_iov *niov;
> - ssize_t offset;
> - ssize_t index;
> -
> - dma_addr = gen_pool_alloc_owner(binding->chunk_pool,
> - 1UL << binding->niov_shift,
> - (void **)&owner);
> - if (!dma_addr)
> + spin_lock_bh(&binding->freelist_lock);
> + if (unlikely(!binding->free_count)) {
> + spin_unlock_bh(&binding->freelist_lock);
> return NULL;
> + }
>
> - offset = dma_addr - owner->base_dma_addr;
> - index = offset >> binding->niov_shift;
> - niov = &owner->area.niovs[index];
> + niov = binding->freelist[--binding->free_count];
> + binding->freelist[binding->free_count] = NULL;
The LLM thinks this NULL store in unnecassary. IDK if it will help
anything in practice to remove it :-)
> + spin_unlock_bh(&binding->freelist_lock);
>
> niov->desc.pp_magic = 0;
> niov->desc.pp = NULL;
> @@ -113,14 +91,15 @@ net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
> void net_devmem_free_dmabuf(struct net_iov *niov)
> {
> struct net_devmem_dmabuf_binding *binding = net_devmem_iov_binding(niov);
> - unsigned long dma_addr = net_devmem_get_dma_addr(niov);
> - size_t niov_size = 1UL << binding->niov_shift;
>
> - if (WARN_ON(!gen_pool_has_addr(binding->chunk_pool, dma_addr,
> - niov_size)))
> + spin_lock_bh(&binding->freelist_lock);
> + if (WARN_ON_ONCE(binding->free_count >= binding->total_niovs)) {
> + spin_unlock_bh(&binding->freelist_lock);
> return;
> + }
>
> - gen_pool_free(binding->chunk_pool, dma_addr, niov_size);
> + binding->freelist[binding->free_count++] = niov;
> + spin_unlock_bh(&binding->freelist_lock);
> }
>
> void net_devmem_unbind_dmabuf(struct net_devmem_dmabuf_binding *binding)
> @@ -194,12 +173,15 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> struct netlink_ext_ack *extack)
> {
> struct net_devmem_dmabuf_binding *binding;
> + struct dmabuf_genpool_chunk_owner *owner;
> size_t niov_size = 1UL << niov_shift;
> static u32 id_alloc_next;
> struct scatterlist *sg;
> struct dma_buf *dmabuf;
> - unsigned int sg_idx, i;
> - unsigned long virtual;
> + unsigned int sg_idx;
> + size_t total_niovs;
> + size_t niov_idx;
> + size_t i;
> int err;
>
> if (!dma_dev) {
> @@ -230,6 +212,7 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> goto err_free_binding;
>
> mutex_init(&binding->lock);
> + spin_lock_init(&binding->freelist_lock);
We don't need freelists on tx right? We should probably not allocate them then?
>
> binding->dmabuf = dmabuf;
> binding->direction = direction;
> @@ -262,20 +245,10 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> goto err_unmap;
> }
> }
> -
> - binding->chunk_pool = gen_pool_create(niov_shift,
> - dev_to_node(&dev->dev));
> - if (!binding->chunk_pool) {
> - err = -ENOMEM;
> - goto err_tx_vec;
> - }
> -
> - virtual = 0;
> + total_niovs = 0;
Do we really need a secondary for_each_sgtable_dma_sg loop just to
calculate the total_niovs? In what edge case is the total_niovs not
just dmabuf_len / niov_len? We do a bunch of alignment checks to make
sure it all works out to that no?
> for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
> dma_addr_t dma_addr = sg_dma_address(sg);
> - struct dmabuf_genpool_chunk_owner *owner;
> size_t len = sg_dma_len(sg);
> - struct net_iov *niov;
>
> if (!IS_ALIGNED(dma_addr, niov_size) ||
> !IS_ALIGNED(len, niov_size)) {
> @@ -283,63 +256,74 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> NL_SET_ERR_MSG_FMT(extack,
> "dmabuf sg entry (addr=%pad, len=%zu) not aligned to niov size %zu",
> &dma_addr, len, niov_size);
> - goto err_free_chunks;
> + goto err_tx_vec;
> }
>
> - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
> - dev_to_node(&dev->dev));
> - if (!owner) {
> - err = -ENOMEM;
> - goto err_free_chunks;
> - }
> + total_niovs += len >> niov_shift;
> + }
>
> - owner->area.base_virtual = virtual;
> - owner->base_dma_addr = dma_addr;
> - owner->area.num_niovs = len >> niov_shift;
> - owner->binding = binding;
> + binding->freelist = kvmalloc_array(total_niovs,
> + sizeof(binding->freelist[0]),
> + GFP_KERNEL);
> + if (!binding->freelist) {
> + err = -ENOMEM;
> + goto err_tx_vec;
> + }
> + binding->total_niovs = total_niovs;
binding->total_niovs and binding->area.num_niovs seem the same thing
always. please get rid of one, probably binding->total_niovs.
I wonder if now that both zcrx and devmem use a freelist if the
freelist should be part of the net_iov_area. The point of that field
was to hold the common stuff actually, but I'm guessing there are
micro-implementation differences that will make converging annoying.
I'm fine either way. :shrug:
>
> - err = gen_pool_add_owner(binding->chunk_pool, dma_addr,
> - dma_addr, len, dev_to_node(&dev->dev),
> - owner);
> - if (err) {
> - kfree(owner);
> - err = -EINVAL;
> - goto err_free_chunks;
> - }
> + owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
> + dev_to_node(&dev->dev));
> + if (!owner) {
> + err = -ENOMEM;
> + goto err_free_freelist;
> + }
>
> - owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
> - owner->area.num_niovs);
> - if (!owner->area.niovs) {
> - err = -ENOMEM;
> - goto err_free_chunks;
> - }
> + owner->area.num_niovs = total_niovs;
> + owner->binding = binding;
> + owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
> + owner->area.num_niovs);
> + if (!owner->area.niovs) {
> + err = -ENOMEM;
> + goto err_free_owner;
> + }
> + binding->chunk_owner = owner;
> +
> + niov_idx = 0;
> + for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
> + dma_addr_t dma_addr = sg_dma_address(sg);
> + size_t len = sg_dma_len(sg);
len is referenced once now; not worth a local var.
> + struct net_iov *niov;
> + size_t nr_niovs = len >> niov_shift;
>
> - for (i = 0; i < owner->area.num_niovs; i++) {
> - niov = &owner->area.niovs[i];
> + for (i = 0; i < nr_niovs; i++, niov_idx++) {
> + niov = &owner->area.niovs[niov_idx];
> net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
> page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
> - net_devmem_get_dma_addr(niov));
> + dma_addr);
> if (direction == DMA_TO_DEVICE)
> - binding->tx_vec[owner->area.base_virtual / PAGE_SIZE + i] = niov;
> + binding->tx_vec[niov_idx] = niov;
> + binding->freelist[binding->free_count++] = niov;
if feels somewhat simple to exclude freelist from TX. Something like:
if (direction == dma_to_device)
<store into tx_vec>
else
<store in freelist>
--
Thanks,
Mina
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
2026-09-14 19:37 ` Stanislav Fomichev
@ 2026-09-14 21:53 ` Mina Almasry
1 sibling, 0 replies; 8+ messages in thread
From: Mina Almasry @ 2026-09-14 21:53 UTC (permalink / raw)
To: Stanislav Fomichev
Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
linux-kernel
On Fri, Sep 11, 2026 at 8:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> After replacing the gen_pool with a binding-level freelist, devmem no
> longer needs a separate chunk owner object. There is only one
> net_iov_area for the binding, so store it directly in struct
> net_devmem_dmabuf_binding.
>
> Derive the binding from net_iov_owner() with container_of(), matching the
> pattern used by io_uring zcrx. This removes the leftover
> dmabuf_genpool_chunk_owner wrapper and its allocation/free path.
>
> Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
> ---
> net/core/devmem.c | 41 ++++++++++-------------------------------
> net/core/devmem.h | 26 +++++++-------------------
> 2 files changed, 17 insertions(+), 50 deletions(-)
>
> diff --git a/net/core/devmem.c b/net/core/devmem.c
> index 4883eb7f3a95..7949f8425bcd 100644
> --- a/net/core/devmem.c
> +++ b/net/core/devmem.c
> @@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
>
> static const struct memory_provider_ops dmabuf_devmem_ops;
>
> -static void
> -net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
> -{
> - if (owner) {
> - kvfree(owner->area.niovs);
> - kfree(owner);
> - }
> -}
> -
> static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
> {
> struct net_devmem_dmabuf_binding *binding =
> @@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
> "can't destroy dmabuf binding. total=%zu, free=%zu",
> binding->total_niovs, binding->free_count);
>
> - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
> + kvfree(binding->area.niovs);
> dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
> binding->direction);
> dma_buf_detach(binding->dmabuf, binding->attachment);
> @@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> }
> binding->total_niovs = total_niovs;
>
> - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
> - dev_to_node(&dev->dev));
> - if (!owner) {
> + binding->area.num_niovs = total_niovs;
> + binding->area.niovs = kvmalloc_objs(*binding->area.niovs,
> + binding->area.num_niovs);
> + if (!binding->area.niovs) {
> err = -ENOMEM;
> goto err_free_freelist;
> }
>
> - owner->area.num_niovs = total_niovs;
> - owner->binding = binding;
> - owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
> - owner->area.num_niovs);
> - if (!owner->area.niovs) {
> - err = -ENOMEM;
> - goto err_free_owner;
> - }
> - binding->chunk_owner = owner;
> -
> niov_idx = 0;
> for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
> dma_addr_t dma_addr = sg_dma_address(sg);
> @@ -296,8 +278,8 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> size_t nr_niovs = len >> niov_shift;
>
> for (i = 0; i < nr_niovs; i++, niov_idx++) {
> - niov = &owner->area.niovs[niov_idx];
> - net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
> + niov = &binding->area.niovs[niov_idx];
> + net_iov_init(niov, &binding->area, NET_IOV_DMABUF);
> page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
> dma_addr);
> if (direction == DMA_TO_DEVICE)
> @@ -311,17 +293,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> binding, xa_limit_32b, &id_alloc_next,
> GFP_KERNEL);
> if (err < 0)
> - goto err_free_chunk_owner;
> + goto err_free_niovs;
>
> list_add(&binding->list, &priv->bindings);
>
> return binding;
>
> -err_free_chunk_owner:
> - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
> - goto err_free_freelist;
> -err_free_owner:
> - kfree(owner);
> +err_free_niovs:
> + kvfree(binding->area.niovs);
> err_free_freelist:
> kvfree(binding->freelist);
> err_tx_vec:
> diff --git a/net/core/devmem.h b/net/core/devmem.h
> index a5ee2d8d9169..20a3eb90ea7f 100644
> --- a/net/core/devmem.h
> +++ b/net/core/devmem.h
> @@ -14,9 +14,9 @@
> #include <net/netdev_netlink.h>
>
> struct netlink_ext_ack;
> -struct dmabuf_genpool_chunk_owner;
>
> struct net_devmem_dmabuf_binding {
> + struct net_iov_area area;
> struct dma_buf *dmabuf;
> struct dma_buf_attachment *attachment;
> struct sg_table *sgt;
> @@ -27,7 +27,6 @@ struct net_devmem_dmabuf_binding {
> * dereferenced.
> */
> void *vdev;
> - struct dmabuf_genpool_chunk_owner *chunk_owner;
> /* Protect dev */
> struct mutex lock;
>
> @@ -83,11 +82,6 @@ struct net_devmem_dmabuf_binding {
> };
>
> #if defined(CONFIG_NET_DEVMEM)
> -struct dmabuf_genpool_chunk_owner {
> - struct net_iov_area area;
> - struct net_devmem_dmabuf_binding *binding;
> -};
> -
> void __net_devmem_dmabuf_binding_free(struct work_struct *wq);
> struct net_devmem_dmabuf_binding *
> net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
> @@ -102,18 +96,12 @@ int net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx,
> struct net_devmem_dmabuf_binding *binding,
> struct netlink_ext_ack *extack);
>
> -static inline struct dmabuf_genpool_chunk_owner *
> -net_devmem_iov_to_chunk_owner(const struct net_iov *niov)
> -{
> - struct net_iov_area *owner = net_iov_owner(niov);
> -
> - return container_of(owner, struct dmabuf_genpool_chunk_owner, area);
> -}
> -
> static inline struct net_devmem_dmabuf_binding *
> net_devmem_iov_binding(const struct net_iov *niov)
> {
> - return net_devmem_iov_to_chunk_owner(niov)->binding;
> + struct net_iov_area *owner = net_iov_owner(niov);
> +
Not worth a local var anymore tbh.
Reviewed-by: Mina Almasry <almasrymina@google.com>
--
Thanks,
Mina
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
@ 2026-09-14 21:58 ` Mina Almasry
0 siblings, 0 replies; 8+ messages in thread
From: Mina Almasry @ 2026-09-14 21:58 UTC (permalink / raw)
To: Stanislav Fomichev
Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
linux-kernel
On Fri, Sep 11, 2026 at 8:46 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> Rename net_devmem_alloc_dmabuf() into net_devmem_alloc_dmabuf_bulk() and
> make it refill page pool with up to PP_ALLOC_CACHE_REFILL NIOVs,
> similar to io_pp_zc_alloc_netmems(). That should amortize recently
> introduced freelist_lock.
>
> Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
> ---
> net/core/devmem.c | 59 +++++++++++++++++++++++++++++------------------
> net/core/devmem.h | 8 -------
> 2 files changed, 37 insertions(+), 30 deletions(-)
>
> diff --git a/net/core/devmem.c b/net/core/devmem.c
> index 7949f8425bcd..a0dcc896dd12 100644
> --- a/net/core/devmem.c
> +++ b/net/core/devmem.c
> @@ -58,25 +58,25 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
> kfree(binding);
> }
>
> -struct net_iov *
> -net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
> +static unsigned int
> +net_devmem_alloc_dmabuf_bulk(struct net_devmem_dmabuf_binding *binding,
> + netmem_ref *netmems, unsigned int count)
> {
> - struct net_iov *niov;
> + unsigned int i;
> +
> spin_lock_bh(&binding->freelist_lock);
> - if (unlikely(!binding->free_count)) {
> - spin_unlock_bh(&binding->freelist_lock);
> - return NULL;
> +
> + count = min_t(size_t, count, binding->free_count);
> + for (i = 0; i < count; i++) {
> + struct net_iov *niov = binding->freelist[--binding->free_count];
> +
> + binding->freelist[binding->free_count] = NULL;
Nulling is probably unnecessary?
Reviewed-by: Mina Almasry <almasrymina@google.com>
--
Thanks,
Mina
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-14 21:58 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
2026-09-14 21:46 ` Mina Almasry
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
2026-09-14 19:37 ` Stanislav Fomichev
2026-09-14 21:53 ` Mina Almasry
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
2026-09-14 21:58 ` Mina Almasry
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®