* [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations
@ 2026-09-11 15:45 Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
Replace devmem's gen_pool based fixed-size allocator with a binding-level
freelist similar to the one used by io_uring zero-copy receive.
This is motivated by allocation latency observed in the NAPI receive path:
[ 1036.228913] ? gen_pool_create+0x90/0x90
[ 1036.228915] net_devmem_alloc_dmabuf+0x1f/0x60
[ 1036.228918] mp_dmabuf_devmem_alloc_netmems+0x17/0x80
[ 1036.228920] mlx5e_post_rx_mpwqes+0xdbe/0xdd0
[ 1036.228926] mlx5e_napi_poll+0x113/0x830
[ 1036.228928] ? sched_clock+0x5/0x10
[ 1036.228931] ? wake_up_process+0x778/0x14b0
[ 1036.228933] net_rx_action+0x15d/0x570
[ 1036.228934] ? update_rq_clock+0x31/0x240
[ 1036.228937] ? __napi_schedule+0x55/0xa0
[ 1036.228938] ? mlx5_eq_comp_int+0x137/0x230
[ 1036.228940] ? atomic_notifier_call_chain+0x36/0x90
[ 1036.228943] ? sched_clock+0x5/0x10
[ 1036.228944] ? sched_clock_cpu+0xc/0x170
[ 1036.228947] irq_exit_rcu+0x12b/0x370
[ 1036.228950] common_interrupt+0x85/0x90
udmabuf can create a very large number of SG entries. In the worst case,
devmem ends up adding one gen_pool chunk for each net_iov allocation
unit backed by those entries. The gen_pool allocation path then has to
traverse a linked list that can become too long for this hot path.
Patch 1 removes the gen_pool and replaces it with a simple freelist of
net_iov pointers protected by the same spin_lock_bh() pattern used by
io_uring zcrx. Patch 2 removes the now-unnecessary chunk owner wrapper by
embedding the net_iov_area directly in the dma-buf binding. Patch 3
batches freelist allocations.
= Performance:
kperf/client ... \
--num-rx-queues 4 \
--dmabuf-rx-size-mb 2048 \
--dmabuf-tx-size-mb 2048 \
--validate no \
--time 60 \
--read-size 67108864 \
--write-size 67108864 \
--num-connections 4 \
--tcp-cc dctcp \
--pin-off 4 \
--devmem-rx \
--devmem-tx \
--devmem-rx-memory cuda \
--devmem-tx-memory cuda
With 4 queues, 4 flows, 2GB BB, cuda for both rx and tx I see no difference
in throughput or cpu utilization (see selective runs below).
== Before
10 runs: 206.031 243.546 293.931 319.015 319.923 323.189 324.321 325.499 327.134 349.450 Gbps
Sample:
client: == Source <redacted>
client: Tx 48.170 Gbps (361716776960 bytes in 60072872 usec)
client: Tx101.256 Gbps (760343429120 bytes in 60072872 usec)
client: Tx101.077 Gbps (759001251840 bytes in 60072872 usec)
client: Tx 69.440 Gbps (521435873280 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Rx 48.158 Gbps (361638901920 bytes in 60074846 usec)
client: Rx101.253 Gbps (760343429120 bytes in 60074846 usec)
client: Rx101.074 Gbps (759001251840 bytes in 60074846 usec)
client: Rx 69.438 Gbps (521435873280 bytes in 60074846 usec)
client: net CPU 1: usr: 0.00% sys: 0.01% idle:39.42% iow: 0.00% irq: 0.64% sirq:59.91%
client: app CPU 5: usr: 1.21% sys:97.71% idle: 0.09% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 2: usr: 0.00% sys: 0.00% idle:38.19% iow: 0.00% irq: 0.84% sirq:60.96%
client: app CPU 6: usr: 1.24% sys:97.73% idle: 0.04% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 0: usr: 0.05% sys: 0.27% idle:69.92% iow: 0.00% irq: 7.71% sirq:22.04%
client: app CPU 4: usr: 1.78% sys:80.71% idle:16.79% iow: 0.00% irq: 0.38% sirq: 0.32%
client: net CPU 3: usr: 0.00% sys: 0.00% idle: 0.00% iow: 0.00% irq: 0.39% sirq:99.60%
client: app CPU 7: usr: 0.21% sys: 6.64% idle:90.75% iow: 0.00% irq: 0.05% sirq: 2.32%
== After
10 runs: 213.032 226.471 235.052 257.113 323.233 331.597 342.893 348.188 351.866 354.311 Gbps
Sample:
client: == Source <redacted>
client: Tx104.829 Gbps (786515886080 bytes in 60022934 usec)
client: Tx105.455 Gbps (791213506560 bytes in 60022934 usec)
client: Tx 61.717 Gbps (463051161600 bytes in 60022934 usec)
client: Tx 51.430 Gbps (385875968000 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Rx104.767 Gbps (786515886080 bytes in 60058283 usec)
client: Rx105.393 Gbps (791213506560 bytes in 60058283 usec)
client: Rx 61.673 Gbps (462998508768 bytes in 60058283 usec)
client: Rx 51.400 Gbps (385875968000 bytes in 60058283 usec)
client: net CPU 0: usr: 0.01% sys: 0.21% idle:65.30% iow: 0.00% irq: 9.15% sirq:25.29%
client: app CPU 4: usr: 1.19% sys:98.00% idle: 0.03% iow: 0.00% irq: 0.21% sirq: 0.54%
client: net CPU 2: usr: 0.00% sys: 0.00% idle: 0.03% iow: 0.00% irq: 0.44% sirq:99.51%
client: app CPU 6: usr: 2.14% sys:77.48% idle:19.92% iow: 0.00% irq: 0.36% sirq: 0.07%
client: net CPU 3: usr: 0.00% sys: 0.00% idle:43.01% iow: 0.00% irq: 0.69% sirq:56.28%
client: app CPU 7: usr: 2.03% sys:71.04% idle:26.45% iow: 0.00% irq: 0.37% sirq: 0.08%
client: net CPU 1: usr: 0.00% sys: 0.01% idle:46.66% iow: 0.00% irq: 0.56% sirq:52.75%
client: app CPU 5: usr: 0.05% sys: 0.08% idle:99.21% iow: 0.00% irq: 0.03% sirq: 0.61%
== Comparison, over 10 runs
Median Target RX: 321.556 Gbps vs 327.415 Gbps
Mean Target RX: 303.204 Gbps vs 298.376 Gbps
Range: 206.031-349.450 Gbps vs 213.032-354.311 Gbps
v2:
- xmas tree (Jakub)
- batching (Mina)
- perf numbers (Mina & Jakub)
Stanislav Fomichev (3):
net: devmem: replace gen_pool with freelist
net: devmem: embed net_iov_area in binding
net: devmem: batch net_iov allocations into the page_pool cache
net/Kconfig | 1 -
net/core/devmem.c | 208 +++++++++++++++++++++-------------------------
net/core/devmem.h | 46 +++-------
3 files changed, 105 insertions(+), 150 deletions(-)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist 2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev @ 2026-09-11 15:45 ` Stanislav Fomichev 2026-09-14 21:46 ` Mina Almasry 2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev 2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev 2 siblings, 1 reply; 8+ messages in thread From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw) To: netdev Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, almasrymina, linux-kernel devmem only needs fixed-size net_iov allocations for each dma-buf binding. The gen_pool tracks the same free set indirectly through DMA addresses, which makes devmem depend on the generic allocator even though the users are fixed-size net_iov chunks. Mirror the io_uring zcrx model more closely by keeping a binding-level freelist protected by spin_lock_bh(). Use a single net_iov_area owner for the binding, populate each net_iov's DMA address while walking the SG table, and check at teardown that all net_iovs have returned to the freelist. Drop the NET_DEVMEM select of GENERIC_ALLOCATOR now that devmem no longer calls gen_pool APIs. Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> --- net/Kconfig | 1 - net/core/devmem.c | 172 +++++++++++++++++++++------------------------- net/core/devmem.h | 16 ++--- 3 files changed, 85 insertions(+), 104 deletions(-) diff --git a/net/Kconfig b/net/Kconfig index e38477393551..76ab44aa439a 100644 --- a/net/Kconfig +++ b/net/Kconfig @@ -68,7 +68,6 @@ config SKB_EXTENSIONS config NET_DEVMEM def_bool y - select GENERIC_ALLOCATOR depends on DMA_SHARED_BUFFER depends on PAGE_POOL diff --git a/net/core/devmem.c b/net/core/devmem.c index f4d60654ce7f..4883eb7f3a95 100644 --- a/net/core/devmem.c +++ b/net/core/devmem.c @@ -8,7 +8,6 @@ */ #include <linux/dma-buf.h> -#include <linux/genalloc.h> #include <linux/mm.h> #include <linux/netdevice.h> #include <linux/types.h> @@ -30,23 +29,13 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1); static const struct memory_provider_ops dmabuf_devmem_ops; -static void net_devmem_dmabuf_free_chunk_owner(struct gen_pool *genpool, - struct gen_pool_chunk *chunk, - void *not_used) +static void +net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner) { - struct dmabuf_genpool_chunk_owner *owner = chunk->owner; - - kvfree(owner->area.niovs); - kfree(owner); -} - -static dma_addr_t net_devmem_get_dma_addr(const struct net_iov *niov) -{ - struct dmabuf_genpool_chunk_owner *owner; - - owner = net_devmem_iov_to_chunk_owner(niov); - return owner->base_dma_addr + - ((dma_addr_t)net_iov_idx(niov) << owner->binding->niov_shift); + if (owner) { + kvfree(owner->area.niovs); + kfree(owner); + } } static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref) @@ -62,24 +51,18 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) { struct net_devmem_dmabuf_binding *binding = container_of(wq, typeof(*binding), unbind_w); - size_t size, avail; - - gen_pool_for_each_chunk(binding->chunk_pool, - net_devmem_dmabuf_free_chunk_owner, NULL); - - size = gen_pool_size(binding->chunk_pool); - avail = gen_pool_avail(binding->chunk_pool); - - if (!WARN(size != avail, "can't destroy genpool. size=%zu, avail=%zu", - size, avail)) - gen_pool_destroy(binding->chunk_pool); + WARN(binding->free_count != binding->total_niovs, + "can't destroy dmabuf binding. total=%zu, free=%zu", + binding->total_niovs, binding->free_count); + net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt, binding->direction); dma_buf_detach(binding->dmabuf, binding->attachment); dma_buf_put(binding->dmabuf); xa_destroy(&binding->bound_rxqs); percpu_ref_exit(&binding->ref); + kvfree(binding->freelist); kvfree(binding->tx_vec); kfree(binding); } @@ -87,21 +70,16 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) struct net_iov * net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) { - struct dmabuf_genpool_chunk_owner *owner; - unsigned long dma_addr; struct net_iov *niov; - ssize_t offset; - ssize_t index; - - dma_addr = gen_pool_alloc_owner(binding->chunk_pool, - 1UL << binding->niov_shift, - (void **)&owner); - if (!dma_addr) + spin_lock_bh(&binding->freelist_lock); + if (unlikely(!binding->free_count)) { + spin_unlock_bh(&binding->freelist_lock); return NULL; + } - offset = dma_addr - owner->base_dma_addr; - index = offset >> binding->niov_shift; - niov = &owner->area.niovs[index]; + niov = binding->freelist[--binding->free_count]; + binding->freelist[binding->free_count] = NULL; + spin_unlock_bh(&binding->freelist_lock); niov->desc.pp_magic = 0; niov->desc.pp = NULL; @@ -113,14 +91,15 @@ net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) void net_devmem_free_dmabuf(struct net_iov *niov) { struct net_devmem_dmabuf_binding *binding = net_devmem_iov_binding(niov); - unsigned long dma_addr = net_devmem_get_dma_addr(niov); - size_t niov_size = 1UL << binding->niov_shift; - if (WARN_ON(!gen_pool_has_addr(binding->chunk_pool, dma_addr, - niov_size))) + spin_lock_bh(&binding->freelist_lock); + if (WARN_ON_ONCE(binding->free_count >= binding->total_niovs)) { + spin_unlock_bh(&binding->freelist_lock); return; + } - gen_pool_free(binding->chunk_pool, dma_addr, niov_size); + binding->freelist[binding->free_count++] = niov; + spin_unlock_bh(&binding->freelist_lock); } void net_devmem_unbind_dmabuf(struct net_devmem_dmabuf_binding *binding) @@ -194,12 +173,15 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, struct netlink_ext_ack *extack) { struct net_devmem_dmabuf_binding *binding; + struct dmabuf_genpool_chunk_owner *owner; size_t niov_size = 1UL << niov_shift; static u32 id_alloc_next; struct scatterlist *sg; struct dma_buf *dmabuf; - unsigned int sg_idx, i; - unsigned long virtual; + unsigned int sg_idx; + size_t total_niovs; + size_t niov_idx; + size_t i; int err; if (!dma_dev) { @@ -230,6 +212,7 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, goto err_free_binding; mutex_init(&binding->lock); + spin_lock_init(&binding->freelist_lock); binding->dmabuf = dmabuf; binding->direction = direction; @@ -262,20 +245,10 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, goto err_unmap; } } - - binding->chunk_pool = gen_pool_create(niov_shift, - dev_to_node(&dev->dev)); - if (!binding->chunk_pool) { - err = -ENOMEM; - goto err_tx_vec; - } - - virtual = 0; + total_niovs = 0; for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { dma_addr_t dma_addr = sg_dma_address(sg); - struct dmabuf_genpool_chunk_owner *owner; size_t len = sg_dma_len(sg); - struct net_iov *niov; if (!IS_ALIGNED(dma_addr, niov_size) || !IS_ALIGNED(len, niov_size)) { @@ -283,63 +256,74 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, NL_SET_ERR_MSG_FMT(extack, "dmabuf sg entry (addr=%pad, len=%zu) not aligned to niov size %zu", &dma_addr, len, niov_size); - goto err_free_chunks; + goto err_tx_vec; } - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, - dev_to_node(&dev->dev)); - if (!owner) { - err = -ENOMEM; - goto err_free_chunks; - } + total_niovs += len >> niov_shift; + } - owner->area.base_virtual = virtual; - owner->base_dma_addr = dma_addr; - owner->area.num_niovs = len >> niov_shift; - owner->binding = binding; + binding->freelist = kvmalloc_array(total_niovs, + sizeof(binding->freelist[0]), + GFP_KERNEL); + if (!binding->freelist) { + err = -ENOMEM; + goto err_tx_vec; + } + binding->total_niovs = total_niovs; - err = gen_pool_add_owner(binding->chunk_pool, dma_addr, - dma_addr, len, dev_to_node(&dev->dev), - owner); - if (err) { - kfree(owner); - err = -EINVAL; - goto err_free_chunks; - } + owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, + dev_to_node(&dev->dev)); + if (!owner) { + err = -ENOMEM; + goto err_free_freelist; + } - owner->area.niovs = kvmalloc_objs(*owner->area.niovs, - owner->area.num_niovs); - if (!owner->area.niovs) { - err = -ENOMEM; - goto err_free_chunks; - } + owner->area.num_niovs = total_niovs; + owner->binding = binding; + owner->area.niovs = kvmalloc_objs(*owner->area.niovs, + owner->area.num_niovs); + if (!owner->area.niovs) { + err = -ENOMEM; + goto err_free_owner; + } + binding->chunk_owner = owner; + + niov_idx = 0; + for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { + dma_addr_t dma_addr = sg_dma_address(sg); + size_t len = sg_dma_len(sg); + struct net_iov *niov; + size_t nr_niovs = len >> niov_shift; - for (i = 0; i < owner->area.num_niovs; i++) { - niov = &owner->area.niovs[i]; + for (i = 0; i < nr_niovs; i++, niov_idx++) { + niov = &owner->area.niovs[niov_idx]; net_iov_init(niov, &owner->area, NET_IOV_DMABUF); page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov), - net_devmem_get_dma_addr(niov)); + dma_addr); if (direction == DMA_TO_DEVICE) - binding->tx_vec[owner->area.base_virtual / PAGE_SIZE + i] = niov; + binding->tx_vec[niov_idx] = niov; + binding->freelist[binding->free_count++] = niov; + dma_addr += niov_size; } - - virtual += len; } err = xa_alloc_cyclic(&net_devmem_dmabuf_bindings, &binding->id, binding, xa_limit_32b, &id_alloc_next, GFP_KERNEL); if (err < 0) - goto err_free_chunks; + goto err_free_chunk_owner; list_add(&binding->list, &priv->bindings); return binding; -err_free_chunks: - gen_pool_for_each_chunk(binding->chunk_pool, - net_devmem_dmabuf_free_chunk_owner, NULL); - gen_pool_destroy(binding->chunk_pool); +err_free_chunk_owner: + net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); + goto err_free_freelist; +err_free_owner: + kfree(owner); +err_free_freelist: + kvfree(binding->freelist); err_tx_vec: kvfree(binding->tx_vec); err_unmap: diff --git a/net/core/devmem.h b/net/core/devmem.h index 4a293a7d1149..a5ee2d8d9169 100644 --- a/net/core/devmem.h +++ b/net/core/devmem.h @@ -14,6 +14,7 @@ #include <net/netdev_netlink.h> struct netlink_ext_ack; +struct dmabuf_genpool_chunk_owner; struct net_devmem_dmabuf_binding { struct dma_buf *dmabuf; @@ -26,7 +27,7 @@ struct net_devmem_dmabuf_binding { * dereferenced. */ void *vdev; - struct gen_pool *chunk_pool; + struct dmabuf_genpool_chunk_owner *chunk_owner; /* Protect dev */ struct mutex lock; @@ -57,6 +58,11 @@ struct net_devmem_dmabuf_binding { /* rxq's this binding is active on. */ struct xarray bound_rxqs; + spinlock_t freelist_lock ____cacheline_aligned_in_smp; + size_t free_count; + size_t total_niovs; + struct net_iov **freelist; + /* ID of this binding. Globally unique to all bindings currently * active. */ @@ -77,17 +83,9 @@ struct net_devmem_dmabuf_binding { }; #if defined(CONFIG_NET_DEVMEM) -/* Owner of the dma-buf chunks inserted into the gen pool. Each scatterlist - * entry from the dmabuf is inserted into the genpool as a chunk, and needs - * this owner struct to keep track of some metadata necessary to create - * allocations from this chunk. - */ struct dmabuf_genpool_chunk_owner { struct net_iov_area area; struct net_devmem_dmabuf_binding *binding; - - /* dma_addr of the start of the chunk. */ - dma_addr_t base_dma_addr; }; void __net_devmem_dmabuf_binding_free(struct work_struct *wq); -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist 2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev @ 2026-09-14 21:46 ` Mina Almasry 0 siblings, 0 replies; 8+ messages in thread From: Mina Almasry @ 2026-09-14 21:46 UTC (permalink / raw) To: Stanislav Fomichev Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, linux-kernel On Fri, Sep 11, 2026 at 8:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote: > > devmem only needs fixed-size net_iov allocations for each dma-buf binding. > The gen_pool tracks the same free set indirectly through DMA addresses, > which makes devmem depend on the generic allocator even though the users > are fixed-size net_iov chunks. > > Mirror the io_uring zcrx model more closely by keeping a binding-level > freelist protected by spin_lock_bh(). Use a single net_iov_area owner for > the binding, populate each net_iov's DMA address while walking the SG > table, and check at teardown that all net_iovs have returned to the > freelist. > > Drop the NET_DEVMEM select of GENERIC_ALLOCATOR now that devmem no longer > calls gen_pool APIs. > > Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> Approach is great, some suggested improvements. > --- > net/Kconfig | 1 - > net/core/devmem.c | 172 +++++++++++++++++++++------------------------- > net/core/devmem.h | 16 ++--- > 3 files changed, 85 insertions(+), 104 deletions(-) > > diff --git a/net/Kconfig b/net/Kconfig > index e38477393551..76ab44aa439a 100644 > --- a/net/Kconfig > +++ b/net/Kconfig > @@ -68,7 +68,6 @@ config SKB_EXTENSIONS > > config NET_DEVMEM > def_bool y > - select GENERIC_ALLOCATOR > depends on DMA_SHARED_BUFFER > depends on PAGE_POOL > > diff --git a/net/core/devmem.c b/net/core/devmem.c > index f4d60654ce7f..4883eb7f3a95 100644 > --- a/net/core/devmem.c > +++ b/net/core/devmem.c > @@ -8,7 +8,6 @@ > */ > > #include <linux/dma-buf.h> > -#include <linux/genalloc.h> > #include <linux/mm.h> > #include <linux/netdevice.h> > #include <linux/types.h> > @@ -30,23 +29,13 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1); > > static const struct memory_provider_ops dmabuf_devmem_ops; > > -static void net_devmem_dmabuf_free_chunk_owner(struct gen_pool *genpool, > - struct gen_pool_chunk *chunk, > - void *not_used) > +static void > +net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner) > { > - struct dmabuf_genpool_chunk_owner *owner = chunk->owner; > - > - kvfree(owner->area.niovs); > - kfree(owner); > -} > - > -static dma_addr_t net_devmem_get_dma_addr(const struct net_iov *niov) > -{ > - struct dmabuf_genpool_chunk_owner *owner; > - > - owner = net_devmem_iov_to_chunk_owner(niov); > - return owner->base_dma_addr + > - ((dma_addr_t)net_iov_idx(niov) << owner->binding->niov_shift); > + if (owner) { > + kvfree(owner->area.niovs); > + kfree(owner); > + } > } > > static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref) > @@ -62,24 +51,18 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) > { > struct net_devmem_dmabuf_binding *binding = container_of(wq, typeof(*binding), unbind_w); > > - size_t size, avail; > - > - gen_pool_for_each_chunk(binding->chunk_pool, > - net_devmem_dmabuf_free_chunk_owner, NULL); > - > - size = gen_pool_size(binding->chunk_pool); > - avail = gen_pool_avail(binding->chunk_pool); > - > - if (!WARN(size != avail, "can't destroy genpool. size=%zu, avail=%zu", > - size, avail)) > - gen_pool_destroy(binding->chunk_pool); > + WARN(binding->free_count != binding->total_niovs, > + "can't destroy dmabuf binding. total=%zu, free=%zu", > + binding->total_niovs, binding->free_count); > You're warning here that you can't destroy the dmabuf binding but you're destroying it anyway. Something is off here. Do we want an early return or something else? > + net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); > dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt, > binding->direction); > dma_buf_detach(binding->dmabuf, binding->attachment); > dma_buf_put(binding->dmabuf); > xa_destroy(&binding->bound_rxqs); > percpu_ref_exit(&binding->ref); > + kvfree(binding->freelist); > kvfree(binding->tx_vec); > kfree(binding); > } > @@ -87,21 +70,16 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) > struct net_iov * > net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) > { > - struct dmabuf_genpool_chunk_owner *owner; > - unsigned long dma_addr; > struct net_iov *niov; > - ssize_t offset; > - ssize_t index; > - > - dma_addr = gen_pool_alloc_owner(binding->chunk_pool, > - 1UL << binding->niov_shift, > - (void **)&owner); > - if (!dma_addr) > + spin_lock_bh(&binding->freelist_lock); > + if (unlikely(!binding->free_count)) { > + spin_unlock_bh(&binding->freelist_lock); > return NULL; > + } > > - offset = dma_addr - owner->base_dma_addr; > - index = offset >> binding->niov_shift; > - niov = &owner->area.niovs[index]; > + niov = binding->freelist[--binding->free_count]; > + binding->freelist[binding->free_count] = NULL; The LLM thinks this NULL store in unnecassary. IDK if it will help anything in practice to remove it :-) > + spin_unlock_bh(&binding->freelist_lock); > > niov->desc.pp_magic = 0; > niov->desc.pp = NULL; > @@ -113,14 +91,15 @@ net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) > void net_devmem_free_dmabuf(struct net_iov *niov) > { > struct net_devmem_dmabuf_binding *binding = net_devmem_iov_binding(niov); > - unsigned long dma_addr = net_devmem_get_dma_addr(niov); > - size_t niov_size = 1UL << binding->niov_shift; > > - if (WARN_ON(!gen_pool_has_addr(binding->chunk_pool, dma_addr, > - niov_size))) > + spin_lock_bh(&binding->freelist_lock); > + if (WARN_ON_ONCE(binding->free_count >= binding->total_niovs)) { > + spin_unlock_bh(&binding->freelist_lock); > return; > + } > > - gen_pool_free(binding->chunk_pool, dma_addr, niov_size); > + binding->freelist[binding->free_count++] = niov; > + spin_unlock_bh(&binding->freelist_lock); > } > > void net_devmem_unbind_dmabuf(struct net_devmem_dmabuf_binding *binding) > @@ -194,12 +173,15 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > struct netlink_ext_ack *extack) > { > struct net_devmem_dmabuf_binding *binding; > + struct dmabuf_genpool_chunk_owner *owner; > size_t niov_size = 1UL << niov_shift; > static u32 id_alloc_next; > struct scatterlist *sg; > struct dma_buf *dmabuf; > - unsigned int sg_idx, i; > - unsigned long virtual; > + unsigned int sg_idx; > + size_t total_niovs; > + size_t niov_idx; > + size_t i; > int err; > > if (!dma_dev) { > @@ -230,6 +212,7 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > goto err_free_binding; > > mutex_init(&binding->lock); > + spin_lock_init(&binding->freelist_lock); We don't need freelists on tx right? We should probably not allocate them then? > > binding->dmabuf = dmabuf; > binding->direction = direction; > @@ -262,20 +245,10 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > goto err_unmap; > } > } > - > - binding->chunk_pool = gen_pool_create(niov_shift, > - dev_to_node(&dev->dev)); > - if (!binding->chunk_pool) { > - err = -ENOMEM; > - goto err_tx_vec; > - } > - > - virtual = 0; > + total_niovs = 0; Do we really need a secondary for_each_sgtable_dma_sg loop just to calculate the total_niovs? In what edge case is the total_niovs not just dmabuf_len / niov_len? We do a bunch of alignment checks to make sure it all works out to that no? > for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { > dma_addr_t dma_addr = sg_dma_address(sg); > - struct dmabuf_genpool_chunk_owner *owner; > size_t len = sg_dma_len(sg); > - struct net_iov *niov; > > if (!IS_ALIGNED(dma_addr, niov_size) || > !IS_ALIGNED(len, niov_size)) { > @@ -283,63 +256,74 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > NL_SET_ERR_MSG_FMT(extack, > "dmabuf sg entry (addr=%pad, len=%zu) not aligned to niov size %zu", > &dma_addr, len, niov_size); > - goto err_free_chunks; > + goto err_tx_vec; > } > > - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, > - dev_to_node(&dev->dev)); > - if (!owner) { > - err = -ENOMEM; > - goto err_free_chunks; > - } > + total_niovs += len >> niov_shift; > + } > > - owner->area.base_virtual = virtual; > - owner->base_dma_addr = dma_addr; > - owner->area.num_niovs = len >> niov_shift; > - owner->binding = binding; > + binding->freelist = kvmalloc_array(total_niovs, > + sizeof(binding->freelist[0]), > + GFP_KERNEL); > + if (!binding->freelist) { > + err = -ENOMEM; > + goto err_tx_vec; > + } > + binding->total_niovs = total_niovs; binding->total_niovs and binding->area.num_niovs seem the same thing always. please get rid of one, probably binding->total_niovs. I wonder if now that both zcrx and devmem use a freelist if the freelist should be part of the net_iov_area. The point of that field was to hold the common stuff actually, but I'm guessing there are micro-implementation differences that will make converging annoying. I'm fine either way. :shrug: > > - err = gen_pool_add_owner(binding->chunk_pool, dma_addr, > - dma_addr, len, dev_to_node(&dev->dev), > - owner); > - if (err) { > - kfree(owner); > - err = -EINVAL; > - goto err_free_chunks; > - } > + owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, > + dev_to_node(&dev->dev)); > + if (!owner) { > + err = -ENOMEM; > + goto err_free_freelist; > + } > > - owner->area.niovs = kvmalloc_objs(*owner->area.niovs, > - owner->area.num_niovs); > - if (!owner->area.niovs) { > - err = -ENOMEM; > - goto err_free_chunks; > - } > + owner->area.num_niovs = total_niovs; > + owner->binding = binding; > + owner->area.niovs = kvmalloc_objs(*owner->area.niovs, > + owner->area.num_niovs); > + if (!owner->area.niovs) { > + err = -ENOMEM; > + goto err_free_owner; > + } > + binding->chunk_owner = owner; > + > + niov_idx = 0; > + for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { > + dma_addr_t dma_addr = sg_dma_address(sg); > + size_t len = sg_dma_len(sg); len is referenced once now; not worth a local var. > + struct net_iov *niov; > + size_t nr_niovs = len >> niov_shift; > > - for (i = 0; i < owner->area.num_niovs; i++) { > - niov = &owner->area.niovs[i]; > + for (i = 0; i < nr_niovs; i++, niov_idx++) { > + niov = &owner->area.niovs[niov_idx]; > net_iov_init(niov, &owner->area, NET_IOV_DMABUF); > page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov), > - net_devmem_get_dma_addr(niov)); > + dma_addr); > if (direction == DMA_TO_DEVICE) > - binding->tx_vec[owner->area.base_virtual / PAGE_SIZE + i] = niov; > + binding->tx_vec[niov_idx] = niov; > + binding->freelist[binding->free_count++] = niov; if feels somewhat simple to exclude freelist from TX. Something like: if (direction == dma_to_device) <store into tx_vec> else <store in freelist> -- Thanks, Mina ^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding 2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev 2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev @ 2026-09-11 15:45 ` Stanislav Fomichev 2026-09-14 19:37 ` Stanislav Fomichev 2026-09-14 21:53 ` Mina Almasry 2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev 2 siblings, 2 replies; 8+ messages in thread From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw) To: netdev Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, almasrymina, linux-kernel After replacing the gen_pool with a binding-level freelist, devmem no longer needs a separate chunk owner object. There is only one net_iov_area for the binding, so store it directly in struct net_devmem_dmabuf_binding. Derive the binding from net_iov_owner() with container_of(), matching the pattern used by io_uring zcrx. This removes the leftover dmabuf_genpool_chunk_owner wrapper and its allocation/free path. Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> --- net/core/devmem.c | 41 ++++++++++------------------------------- net/core/devmem.h | 26 +++++++------------------- 2 files changed, 17 insertions(+), 50 deletions(-) diff --git a/net/core/devmem.c b/net/core/devmem.c index 4883eb7f3a95..7949f8425bcd 100644 --- a/net/core/devmem.c +++ b/net/core/devmem.c @@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1); static const struct memory_provider_ops dmabuf_devmem_ops; -static void -net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner) -{ - if (owner) { - kvfree(owner->area.niovs); - kfree(owner); - } -} - static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref) { struct net_devmem_dmabuf_binding *binding = @@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) "can't destroy dmabuf binding. total=%zu, free=%zu", binding->total_niovs, binding->free_count); - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); + kvfree(binding->area.niovs); dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt, binding->direction); dma_buf_detach(binding->dmabuf, binding->attachment); @@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, } binding->total_niovs = total_niovs; - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, - dev_to_node(&dev->dev)); - if (!owner) { + binding->area.num_niovs = total_niovs; + binding->area.niovs = kvmalloc_objs(*binding->area.niovs, + binding->area.num_niovs); + if (!binding->area.niovs) { err = -ENOMEM; goto err_free_freelist; } - owner->area.num_niovs = total_niovs; - owner->binding = binding; - owner->area.niovs = kvmalloc_objs(*owner->area.niovs, - owner->area.num_niovs); - if (!owner->area.niovs) { - err = -ENOMEM; - goto err_free_owner; - } - binding->chunk_owner = owner; - niov_idx = 0; for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { dma_addr_t dma_addr = sg_dma_address(sg); @@ -296,8 +278,8 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, size_t nr_niovs = len >> niov_shift; for (i = 0; i < nr_niovs; i++, niov_idx++) { - niov = &owner->area.niovs[niov_idx]; - net_iov_init(niov, &owner->area, NET_IOV_DMABUF); + niov = &binding->area.niovs[niov_idx]; + net_iov_init(niov, &binding->area, NET_IOV_DMABUF); page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov), dma_addr); if (direction == DMA_TO_DEVICE) @@ -311,17 +293,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, binding, xa_limit_32b, &id_alloc_next, GFP_KERNEL); if (err < 0) - goto err_free_chunk_owner; + goto err_free_niovs; list_add(&binding->list, &priv->bindings); return binding; -err_free_chunk_owner: - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); - goto err_free_freelist; -err_free_owner: - kfree(owner); +err_free_niovs: + kvfree(binding->area.niovs); err_free_freelist: kvfree(binding->freelist); err_tx_vec: diff --git a/net/core/devmem.h b/net/core/devmem.h index a5ee2d8d9169..20a3eb90ea7f 100644 --- a/net/core/devmem.h +++ b/net/core/devmem.h @@ -14,9 +14,9 @@ #include <net/netdev_netlink.h> struct netlink_ext_ack; -struct dmabuf_genpool_chunk_owner; struct net_devmem_dmabuf_binding { + struct net_iov_area area; struct dma_buf *dmabuf; struct dma_buf_attachment *attachment; struct sg_table *sgt; @@ -27,7 +27,6 @@ struct net_devmem_dmabuf_binding { * dereferenced. */ void *vdev; - struct dmabuf_genpool_chunk_owner *chunk_owner; /* Protect dev */ struct mutex lock; @@ -83,11 +82,6 @@ struct net_devmem_dmabuf_binding { }; #if defined(CONFIG_NET_DEVMEM) -struct dmabuf_genpool_chunk_owner { - struct net_iov_area area; - struct net_devmem_dmabuf_binding *binding; -}; - void __net_devmem_dmabuf_binding_free(struct work_struct *wq); struct net_devmem_dmabuf_binding * net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, @@ -102,18 +96,12 @@ int net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx, struct net_devmem_dmabuf_binding *binding, struct netlink_ext_ack *extack); -static inline struct dmabuf_genpool_chunk_owner * -net_devmem_iov_to_chunk_owner(const struct net_iov *niov) -{ - struct net_iov_area *owner = net_iov_owner(niov); - - return container_of(owner, struct dmabuf_genpool_chunk_owner, area); -} - static inline struct net_devmem_dmabuf_binding * net_devmem_iov_binding(const struct net_iov *niov) { - return net_devmem_iov_to_chunk_owner(niov)->binding; + struct net_iov_area *owner = net_iov_owner(niov); + + return container_of(owner, struct net_devmem_dmabuf_binding, area); } static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov) @@ -123,11 +111,11 @@ static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov) static inline unsigned long net_iov_virtual_addr(const struct net_iov *niov) { - struct dmabuf_genpool_chunk_owner *co = - net_devmem_iov_to_chunk_owner(niov); + struct net_devmem_dmabuf_binding *binding = + net_devmem_iov_binding(niov); return net_iov_owner(niov)->base_virtual + - ((unsigned long)net_iov_idx(niov) << co->binding->niov_shift); + ((unsigned long)net_iov_idx(niov) << binding->niov_shift); } static inline bool -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding 2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev @ 2026-09-14 19:37 ` Stanislav Fomichev 2026-09-14 21:53 ` Mina Almasry 1 sibling, 0 replies; 8+ messages in thread From: Stanislav Fomichev @ 2026-09-14 19:37 UTC (permalink / raw) To: netdev Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, almasrymina, linux-kernel On 09/11, Stanislav Fomichev wrote: > After replacing the gen_pool with a binding-level freelist, devmem no > longer needs a separate chunk owner object. There is only one > net_iov_area for the binding, so store it directly in struct > net_devmem_dmabuf_binding. > > Derive the binding from net_iov_owner() with container_of(), matching the > pattern used by io_uring zcrx. This removes the leftover > dmabuf_genpool_chunk_owner wrapper and its allocation/free path. > > Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> > --- > net/core/devmem.c | 41 ++++++++++------------------------------- > net/core/devmem.h | 26 +++++++------------------- > 2 files changed, 17 insertions(+), 50 deletions(-) > > diff --git a/net/core/devmem.c b/net/core/devmem.c > index 4883eb7f3a95..7949f8425bcd 100644 > --- a/net/core/devmem.c > +++ b/net/core/devmem.c > @@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1); > > static const struct memory_provider_ops dmabuf_devmem_ops; > > -static void > -net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner) > -{ > - if (owner) { > - kvfree(owner->area.niovs); > - kfree(owner); > - } > -} > - > static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref) > { > struct net_devmem_dmabuf_binding *binding = > @@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) > "can't destroy dmabuf binding. total=%zu, free=%zu", > binding->total_niovs, binding->free_count); > > - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); > + kvfree(binding->area.niovs); > dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt, > binding->direction); > dma_buf_detach(binding->dmabuf, binding->attachment); > @@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > } > binding->total_niovs = total_niovs; > > - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, > - dev_to_node(&dev->dev)); There is a leftover owner var, the build fails, will repost later this week (to give some time to review for whoever is interested). ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding 2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev 2026-09-14 19:37 ` Stanislav Fomichev @ 2026-09-14 21:53 ` Mina Almasry 1 sibling, 0 replies; 8+ messages in thread From: Mina Almasry @ 2026-09-14 21:53 UTC (permalink / raw) To: Stanislav Fomichev Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, linux-kernel On Fri, Sep 11, 2026 at 8:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote: > > After replacing the gen_pool with a binding-level freelist, devmem no > longer needs a separate chunk owner object. There is only one > net_iov_area for the binding, so store it directly in struct > net_devmem_dmabuf_binding. > > Derive the binding from net_iov_owner() with container_of(), matching the > pattern used by io_uring zcrx. This removes the leftover > dmabuf_genpool_chunk_owner wrapper and its allocation/free path. > > Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> > --- > net/core/devmem.c | 41 ++++++++++------------------------------- > net/core/devmem.h | 26 +++++++------------------- > 2 files changed, 17 insertions(+), 50 deletions(-) > > diff --git a/net/core/devmem.c b/net/core/devmem.c > index 4883eb7f3a95..7949f8425bcd 100644 > --- a/net/core/devmem.c > +++ b/net/core/devmem.c > @@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1); > > static const struct memory_provider_ops dmabuf_devmem_ops; > > -static void > -net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner) > -{ > - if (owner) { > - kvfree(owner->area.niovs); > - kfree(owner); > - } > -} > - > static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref) > { > struct net_devmem_dmabuf_binding *binding = > @@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) > "can't destroy dmabuf binding. total=%zu, free=%zu", > binding->total_niovs, binding->free_count); > > - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); > + kvfree(binding->area.niovs); > dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt, > binding->direction); > dma_buf_detach(binding->dmabuf, binding->attachment); > @@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > } > binding->total_niovs = total_niovs; > > - owner = kzalloc_node(sizeof(*owner), GFP_KERNEL, > - dev_to_node(&dev->dev)); > - if (!owner) { > + binding->area.num_niovs = total_niovs; > + binding->area.niovs = kvmalloc_objs(*binding->area.niovs, > + binding->area.num_niovs); > + if (!binding->area.niovs) { > err = -ENOMEM; > goto err_free_freelist; > } > > - owner->area.num_niovs = total_niovs; > - owner->binding = binding; > - owner->area.niovs = kvmalloc_objs(*owner->area.niovs, > - owner->area.num_niovs); > - if (!owner->area.niovs) { > - err = -ENOMEM; > - goto err_free_owner; > - } > - binding->chunk_owner = owner; > - > niov_idx = 0; > for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) { > dma_addr_t dma_addr = sg_dma_address(sg); > @@ -296,8 +278,8 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > size_t nr_niovs = len >> niov_shift; > > for (i = 0; i < nr_niovs; i++, niov_idx++) { > - niov = &owner->area.niovs[niov_idx]; > - net_iov_init(niov, &owner->area, NET_IOV_DMABUF); > + niov = &binding->area.niovs[niov_idx]; > + net_iov_init(niov, &binding->area, NET_IOV_DMABUF); > page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov), > dma_addr); > if (direction == DMA_TO_DEVICE) > @@ -311,17 +293,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > binding, xa_limit_32b, &id_alloc_next, > GFP_KERNEL); > if (err < 0) > - goto err_free_chunk_owner; > + goto err_free_niovs; > > list_add(&binding->list, &priv->bindings); > > return binding; > > -err_free_chunk_owner: > - net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner); > - goto err_free_freelist; > -err_free_owner: > - kfree(owner); > +err_free_niovs: > + kvfree(binding->area.niovs); > err_free_freelist: > kvfree(binding->freelist); > err_tx_vec: > diff --git a/net/core/devmem.h b/net/core/devmem.h > index a5ee2d8d9169..20a3eb90ea7f 100644 > --- a/net/core/devmem.h > +++ b/net/core/devmem.h > @@ -14,9 +14,9 @@ > #include <net/netdev_netlink.h> > > struct netlink_ext_ack; > -struct dmabuf_genpool_chunk_owner; > > struct net_devmem_dmabuf_binding { > + struct net_iov_area area; > struct dma_buf *dmabuf; > struct dma_buf_attachment *attachment; > struct sg_table *sgt; > @@ -27,7 +27,6 @@ struct net_devmem_dmabuf_binding { > * dereferenced. > */ > void *vdev; > - struct dmabuf_genpool_chunk_owner *chunk_owner; > /* Protect dev */ > struct mutex lock; > > @@ -83,11 +82,6 @@ struct net_devmem_dmabuf_binding { > }; > > #if defined(CONFIG_NET_DEVMEM) > -struct dmabuf_genpool_chunk_owner { > - struct net_iov_area area; > - struct net_devmem_dmabuf_binding *binding; > -}; > - > void __net_devmem_dmabuf_binding_free(struct work_struct *wq); > struct net_devmem_dmabuf_binding * > net_devmem_bind_dmabuf(struct net_device *dev, void *vdev, > @@ -102,18 +96,12 @@ int net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx, > struct net_devmem_dmabuf_binding *binding, > struct netlink_ext_ack *extack); > > -static inline struct dmabuf_genpool_chunk_owner * > -net_devmem_iov_to_chunk_owner(const struct net_iov *niov) > -{ > - struct net_iov_area *owner = net_iov_owner(niov); > - > - return container_of(owner, struct dmabuf_genpool_chunk_owner, area); > -} > - > static inline struct net_devmem_dmabuf_binding * > net_devmem_iov_binding(const struct net_iov *niov) > { > - return net_devmem_iov_to_chunk_owner(niov)->binding; > + struct net_iov_area *owner = net_iov_owner(niov); > + Not worth a local var anymore tbh. Reviewed-by: Mina Almasry <almasrymina@google.com> -- Thanks, Mina ^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache 2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev 2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev 2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev @ 2026-09-11 15:45 ` Stanislav Fomichev 2026-09-14 21:58 ` Mina Almasry 2 siblings, 1 reply; 8+ messages in thread From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw) To: netdev Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, almasrymina, linux-kernel Rename net_devmem_alloc_dmabuf() into net_devmem_alloc_dmabuf_bulk() and make it refill page pool with up to PP_ALLOC_CACHE_REFILL NIOVs, similar to io_pp_zc_alloc_netmems(). That should amortize recently introduced freelist_lock. Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> --- net/core/devmem.c | 59 +++++++++++++++++++++++++++++------------------ net/core/devmem.h | 8 ------- 2 files changed, 37 insertions(+), 30 deletions(-) diff --git a/net/core/devmem.c b/net/core/devmem.c index 7949f8425bcd..a0dcc896dd12 100644 --- a/net/core/devmem.c +++ b/net/core/devmem.c @@ -58,25 +58,25 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) kfree(binding); } -struct net_iov * -net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) +static unsigned int +net_devmem_alloc_dmabuf_bulk(struct net_devmem_dmabuf_binding *binding, + netmem_ref *netmems, unsigned int count) { - struct net_iov *niov; + unsigned int i; + spin_lock_bh(&binding->freelist_lock); - if (unlikely(!binding->free_count)) { - spin_unlock_bh(&binding->freelist_lock); - return NULL; + + count = min_t(size_t, count, binding->free_count); + for (i = 0; i < count; i++) { + struct net_iov *niov = binding->freelist[--binding->free_count]; + + binding->freelist[binding->free_count] = NULL; + netmems[i] = net_iov_to_netmem(niov); } - niov = binding->freelist[--binding->free_count]; - binding->freelist[binding->free_count] = NULL; spin_unlock_bh(&binding->freelist_lock); - niov->desc.pp_magic = 0; - niov->desc.pp = NULL; - atomic_long_set(&niov->desc.pp_ref_count, 0); - - return niov; + return count; } void net_devmem_free_dmabuf(struct net_iov *niov) @@ -434,20 +434,35 @@ int mp_dmabuf_devmem_init(struct page_pool *pool) netmem_ref mp_dmabuf_devmem_alloc_netmems(struct page_pool *pool, gfp_t gfp) { struct net_devmem_dmabuf_binding *binding = pool->mp_priv; - struct net_iov *niov; - netmem_ref netmem; + netmem_ref *netmems = pool->alloc.cache; + unsigned int allocated, i; + + if (WARN_ON_ONCE(pool->alloc.count)) + return 0; - niov = net_devmem_alloc_dmabuf(binding); - if (!niov) + allocated = net_devmem_alloc_dmabuf_bulk(binding, netmems, + PP_ALLOC_CACHE_REFILL); + if (unlikely(!allocated)) return 0; - netmem = net_iov_to_netmem(niov); + for (i = 0; i < allocated; i++) { + struct net_iov *niov = netmem_to_net_iov(netmems[i]); - page_pool_set_pp_info(pool, netmem); + niov->desc.pp_magic = 0; + niov->desc.pp = NULL; + atomic_long_set(&niov->desc.pp_ref_count, 0); + + page_pool_set_pp_info(pool, netmems[i]); + + pool->pages_state_hold_cnt++; + trace_page_pool_state_hold(pool, netmems[i], + pool->pages_state_hold_cnt); + } - pool->pages_state_hold_cnt++; - trace_page_pool_state_hold(pool, netmem, pool->pages_state_hold_cnt); - return netmem; + /* Return the last one, the rest stay in the page_pool cache. */ + allocated--; + pool->alloc.count = allocated; + return netmems[allocated]; } void mp_dmabuf_devmem_destroy(struct page_pool *pool) diff --git a/net/core/devmem.h b/net/core/devmem.h index 20a3eb90ea7f..7195769b8bd1 100644 --- a/net/core/devmem.h +++ b/net/core/devmem.h @@ -133,8 +133,6 @@ net_devmem_dmabuf_binding_put(struct net_devmem_dmabuf_binding *binding) void net_devmem_get_net_iov(struct net_iov *niov); void net_devmem_put_net_iov(struct net_iov *niov); -struct net_iov * -net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding); void net_devmem_free_dmabuf(struct net_iov *ppiov); @@ -191,12 +189,6 @@ net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx, return -EOPNOTSUPP; } -static inline struct net_iov * -net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) -{ - return NULL; -} - static inline void net_devmem_free_dmabuf(struct net_iov *ppiov) { } -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache 2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev @ 2026-09-14 21:58 ` Mina Almasry 0 siblings, 0 replies; 8+ messages in thread From: Mina Almasry @ 2026-09-14 21:58 UTC (permalink / raw) To: Stanislav Fomichev Cc: netdev, davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman, linux-kernel On Fri, Sep 11, 2026 at 8:46 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote: > > Rename net_devmem_alloc_dmabuf() into net_devmem_alloc_dmabuf_bulk() and > make it refill page pool with up to PP_ALLOC_CACHE_REFILL NIOVs, > similar to io_pp_zc_alloc_netmems(). That should amortize recently > introduced freelist_lock. > > Signed-off-by: Stanislav Fomichev <sdf@fomichev.me> > --- > net/core/devmem.c | 59 +++++++++++++++++++++++++++++------------------ > net/core/devmem.h | 8 ------- > 2 files changed, 37 insertions(+), 30 deletions(-) > > diff --git a/net/core/devmem.c b/net/core/devmem.c > index 7949f8425bcd..a0dcc896dd12 100644 > --- a/net/core/devmem.c > +++ b/net/core/devmem.c > @@ -58,25 +58,25 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq) > kfree(binding); > } > > -struct net_iov * > -net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding) > +static unsigned int > +net_devmem_alloc_dmabuf_bulk(struct net_devmem_dmabuf_binding *binding, > + netmem_ref *netmems, unsigned int count) > { > - struct net_iov *niov; > + unsigned int i; > + > spin_lock_bh(&binding->freelist_lock); > - if (unlikely(!binding->free_count)) { > - spin_unlock_bh(&binding->freelist_lock); > - return NULL; > + > + count = min_t(size_t, count, binding->free_count); > + for (i = 0; i < count; i++) { > + struct net_iov *niov = binding->freelist[--binding->free_count]; > + > + binding->freelist[binding->free_count] = NULL; Nulling is probably unnecessary? Reviewed-by: Mina Almasry <almasrymina@google.com> -- Thanks, Mina ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-14 21:58 UTC | newest] Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev 2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev 2026-09-14 21:46 ` Mina Almasry 2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev 2026-09-14 19:37 ` Stanislav Fomichev 2026-09-14 21:53 ` Mina Almasry 2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev 2026-09-14 21:58 ` Mina Almasry
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®