mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Narcisa Vasile <narcisav.kernel@gmail.com>
To: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
Cc: netdev@vger.kernel.org, "K. Y. Srinivasan" <kys@microsoft.com>,
	Haiyang Zhang <haiyangz@microsoft.com>,
	Wei Liu <wei.liu@kernel.org>, Dexuan Cui <decui@microsoft.com>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Alexei Starovoitov <ast@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Jesper Dangaard Brouer <hawk@kernel.org>,
	John Fastabend <john.fastabend@gmail.com>,
	Stanislav Fomichev <sdf@fomichev.me>,
	Simon Horman <horms@kernel.org>,
	Erni Sri Satya Vennela <ernis@linux.microsoft.com>,
	Dipayaan Roy <dipayanroy@linux.microsoft.com>,
	Aditya Garg <gargaditya@linux.microsoft.com>,
	Jacob Keller <jacob.e.keller@intel.com>,
	Saurabh Sengar <ssengar@linux.microsoft.com>,
	linux-hyperv@vger.kernel.org, bpf@vger.kernel.org,
	linux-kernel@vger.kernel.org, stable@vger.kernel.org
Subject: Re: [PATCH net] net: mana: reserve RX buffer headroom to fix forwarding performance
Date: Sat, 26 Sep 2026 22:14:55 -0700	[thread overview]
Message-ID: <arimAGRnmHEAviqi@gmail.com> (raw)
In-Reply-To: <20260923144500.4073380-1-hamzamahfooz@linux.microsoft.com>

On Wed, Sep 23, 2026 at 10:45:00AM -0400, Hamza Mahfooz wrote:
> Commit 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers
> instead of full pages to improve memory efficiency.") started handing
> out RX buffers with zero headroom so that two buffers fit into one page
> at the default MTU.
> 
> The MANA TX path, however, stores the per scatter-gather entry DMA
> mappings in `struct mana_skb_head` at skb->head, and mana_start_xmit()
> therefore calls skb_cow_head(skb, MANA_HEADROOM). The port advertises
> this requirement as ndev->needed_headroom = MANA_HEADROOM.
> 
> As a result every packet that is received and then forwarded out of a
> MANA port fails the skb_cow() in ip_forward() and gets reallocated and
> copied by pskb_expand_head(). This is invisible to a plain RX or TX
> workload, but it puts a full skb reallocation plus memcpy on the hot
> path of every single forwarded packet, which is exactly what a
> router/NVA workload does.
> 
> Restore the headroom. Note that reserving MANA_HEADROOM (232) is not
> enough: ip_forward() asks for LL_RESERVED_SPACE(dev), which rounds
> hard_header_len + needed_headroom up to HH_DATA_MOD and is 256 bytes on
> ethernet. Use LL_RESERVED_SPACE() directly so the value keeps tracking
> both constants.
> 
> At the default MTU on a 4K page this means a buffer no longer fits twice
> into a page (SKB_DATA_ALIGN(1500 + MANA_RXBUF_PAD + 256) = 2112), so the
> frag-vs-single decision is now made by computing the real buffer size
> instead of comparing the MTU against PAGE_SIZE / 2. The page_pool
> fragment path is still used wherever at least two buffers genuinely fit,
> e.g. on 16K and 64K page sizes.
> 
> Measured on an Azure VM with a MANA NIC acting as a forwarding NVA (UDP,
> 1400 byte payload, 4 streams, 8 Gbps offered, only the forwarding
> node's kernel differs), 8 runs each, median:
> 
>                   forwarded pps    throughput
>   before              272,830       3.06 Gbps
>   after               390,560       4.37 Gbps   (+43%)
> 
> perf on the forwarding node, same workload:
> 
>                   memset_orig   __pi_memcpy   pskb_expand_head
>   before             10.07%         3.96%         present
>   after               0.94%         0.64%         gone
> 
> Cc: stable@vger.kernel.org
> Fixes: 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers instead of full pages to improve memory efficiency.")
> Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
> ---
>  drivers/net/ethernet/microsoft/mana/mana_en.c | 58 ++++++++++++++-----
>  1 file changed, 44 insertions(+), 14 deletions(-)
> 
> diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
> index 591fb4191d90d..e3f3b33ba9062 100644
> --- a/drivers/net/ethernet/microsoft/mana/mana_en.c
> +++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
> @@ -758,6 +758,36 @@ static void *mana_get_rxbuf_pre(struct mana_rxq *rxq, dma_addr_t *da)
>  	return va;
>  }
>  
> +/* RX buffers must be allocated with enough headroom for the TX path:
> + * mana_start_xmit() stores the SGE DMA mappings in struct mana_skb_head at
> + * skb->head, which is why the port advertises ndev->needed_headroom =
> + * MANA_HEADROOM.
> + *
> + * An skb that is forwarded out of a MANA port has to satisfy
> + * skb_cow(skb, LL_RESERVED_SPACE(dev) + ...) in ip_forward(), so reserve
> + * LL_RESERVED_SPACE() here rather than just MANA_HEADROOM - it rounds
> + * hard_header_len + needed_headroom up to HH_DATA_MOD and is therefore
> + * larger. Reserving less makes every forwarded packet get reallocated and
> + * copied by pskb_expand_head().
> + */

nit: maybe trim this comment to make it easier to read. For example, 
"Reserve enough headroom to satisfy the skb_cow() call in ip_forward()
and avoid reallocation."

> +static u32 mana_get_rxbuf_headroom(struct mana_port_context *apc)
> +{
> +	u32 headroom = LL_RESERVED_SPACE(apc->ndev);
> +
> +	if (mana_xdp_get(apc))
> +		return max_t(u32, headroom, XDP_PACKET_HEADROOM);
> +
This implies that the headroom could now be greater than XDP_PACKET_HEADROOM,
in XDP case. Should we then use the actual headroom value, instead
of assuming XDP_PACKET_HEADROOM, in mana_run_xdp() when preparing the buffer:

  line 94:	xdp_prepare_buff(xdp, buf_va, XDP_PACKET_HEADROOM, pkt_len, true);

Does MANA_XDP_MTU_MAX need to be updated?

> +	return headroom;
> +}
> +

  reply	other threads:[~2026-09-27  5:15 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-23 14:45 Hamza Mahfooz
2026-09-27  5:14 ` Narcisa Vasile [this message]
2026-09-27 15:27 ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arimAGRnmHEAviqi@gmail.com \
    --to=narcisav.kernel@gmail.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=decui@microsoft.com \
    --cc=dipayanroy@linux.microsoft.com \
    --cc=edumazet@google.com \
    --cc=ernis@linux.microsoft.com \
    --cc=gargaditya@linux.microsoft.com \
    --cc=haiyangz@microsoft.com \
    --cc=hamzamahfooz@linux.microsoft.com \
    --cc=hawk@kernel.org \
    --cc=horms@kernel.org \
    --cc=jacob.e.keller@intel.com \
    --cc=john.fastabend@gmail.com \
    --cc=kuba@kernel.org \
    --cc=kys@microsoft.com \
    --cc=linux-hyperv@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sdf@fomichev.me \
    --cc=ssengar@linux.microsoft.com \
    --cc=stable@vger.kernel.org \
    --cc=wei.liu@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®