mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v10] net: reduce ARFS flow updates by checking LLC affinity
@ 2026-10-08 13:32 Chuang Wang
  2026-10-08 14:14 ` Neal Cardwell
  0 siblings, 1 reply; 2+ messages in thread
From: Chuang Wang @ 2026-10-08 13:32 UTC (permalink / raw)
  Cc: Chuang Wang, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Kuniyuki Iwashima, Stanislav Fomichev,
	Hangbin Liu, Samiullah Khawaja, Neal Cardwell, Roman Gushchin,
	netdev, linux-kernel

The current implementation of rps_record_sock_flow() updates the flow
table every time a socket is processed on a different CPU. In high-load
scenarios, especially with Accelerated RFS (ARFS), this triggers
frequent flow steering updates via ndo_rx_flow_steer.

For drivers like mlx5 that implement hardware flow steering, these
constant updates lead to significant contention on internal driver locks
(e.g., arfs_lock). In high-load scenarios, this contention often
becomes a performance bottleneck that outweighs the steering benefits.

This patch introduces a cache-aware update strategy: the hardware
steering update is skipped when the flow's previous target CPU and the
new one share the same Last Level Cache (LLC), as cache locality for
the application is preserved either way. The per-queue flow bookkeeping
(rflow->cpu) is still updated, so RFS keeps steering packets to the CPU
the application is running on, and subsequent packets take the fast
path in get_rps_cpu() (tcpu == next_cpu). A new sysctl,
net.core.rps_feat_llc_affinity, is added to toggle this feature
(default: off).

Performance Test Results:
The patch was tested in a K8s environment (AMD CPU 128*2, 16-core Pod
with CPU pinning, mlx5 NIC) using brpc[1] echo_server and rpc_press.

rpc_press Commands:

  for i in {1..8}; do
    ./rpc_press -proto=./echo.proto -method=example.EchoService.Echo
    -server=<IP>:8000 -input='{"message":"hello"}'
    -qps=0 -thread_num=512 -connection_type=pooled &
  done

Monitor mlx5e_rx_flow_steer frequency:

  /usr/share/bcc/tools/funccount -i 1 mlx5e_rx_flow_steer

Frequency of mlx5e_rx_flow_steer (via funccount[2]):

  Before: ~200,000 counts/sec
  After:       ~10 counts/sec (reduced by ~99%)

These results demonstrate that filtering updates by LLC affinity
significantly reduces driver lock contention and improves overall
CPU efficiency under heavy network load.

[1] https://github.com/apache/brpc/
[2] https://github.com/iovisor/bcc/blob/master/tools/funccount.py

Signed-off-by: Chuang Wang <nashuiliang@gmail.com>
---
v9 - v10:
- Move the LLC check from rps_record_sock_flow() into set_rps_cpu()
  by Eric Dumazet
- Drop the sock_rps_record_flow_hash()/sock_rps_record_flow() exports,
  which are no longer needed.
v6 -> v9:
- simplify code and fix errors in AI submissions by Simon Horman
v5 -> v6:
- remove the multi-check 'old_val == new_val' by Xuan Zhuo
- fix 'modpost: "sock_rps_record_flow_hash" [drivers/net/tun.ko] undefined!' by kernel
  test robot
- fix 'tcp.c:(.text+0x3e90): undefined reference to `sock_rps_record_flow'' by kernel test
  robot
v4 -> v5: fix 'modpost: "rps_llc_check" [net/sctp/sctp.ko] undefined!' by kernel test robot
v3 -> v4: add rps_llc_check by Eric Dumazet
v2 -> v3: patch net -> net-next by Jakub Kicinski
v1 -> v2: add rps_feat_llc_affinity; add brpc tests
 include/net/rps.h          |  1 +
 net/core/dev.c             | 23 +++++++++++++++++++++--
 net/core/sysctl_net_core.c |  7 +++++++
 3 files changed, 29 insertions(+), 2 deletions(-)

diff --git a/include/net/rps.h b/include/net/rps.h
index e33c6a2fa8bb..65d1fc43f8b5 100644
--- a/include/net/rps.h
+++ b/include/net/rps.h
@@ -12,6 +12,7 @@
 
 extern struct static_key_false rps_needed;
 extern struct static_key_false rfs_needed;
+extern struct static_key_false rps_feat_llc_affinity;
 
 /*
  * This structure holds an RPS map which can be of variable length.  The
diff --git a/net/core/dev.c b/net/core/dev.c
index f587645e930a..8b046100525a 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -5122,6 +5122,7 @@ struct static_key_false rps_needed __read_mostly;
 EXPORT_SYMBOL(rps_needed);
 struct static_key_false rfs_needed __read_mostly;
 EXPORT_SYMBOL(rfs_needed);
+struct static_key_false rps_feat_llc_affinity __read_mostly;
 
 static u32 rfs_slot(u32 hash, rps_tag_ptr tag_ptr)
 {
@@ -5160,7 +5161,7 @@ static bool rps_flow_is_active(struct rps_dev_flow *rflow,
 
 static struct rps_dev_flow *
 set_rps_cpu(struct net_device *dev, struct sk_buff *skb,
-	    struct rps_dev_flow *rflow, u16 next_cpu, u32 hash)
+	    struct rps_dev_flow *rflow, u16 old_cpu, u16 next_cpu, u32 hash)
 {
 	if (next_cpu < nr_cpu_ids) {
 		u32 head;
@@ -5179,6 +5180,24 @@ set_rps_cpu(struct net_device *dev, struct sk_buff *skb,
 		if (!skb_rx_queue_recorded(skb) || !dev->rx_cpu_rmap ||
 		    !(dev->features & NETIF_F_NTUPLE))
 			goto out;
+
+		/*
+		 * RPS LLC Affinity Feature:
+		 * Reduce RFS/ARFS flow updates by checking LLC affinity.
+		 *
+		 * Frequent flow table updates can trigger constant hardware steering
+		 * reconfigurations (e.g., ndo_rx_flow_steer), leading to significant
+		 * contention on driver internal locks (like mlx5's arfs_lock).
+		 *
+		 * This strategy only updates the flow record if it migrates across LLC
+		 * boundaries. This minimizes expensive hardware updates while
+		 * preserving cache locality for the application.
+		 */
+		if (static_branch_unlikely(&rps_feat_llc_affinity) &&
+		    old_cpu < nr_cpu_ids && cpu_online(old_cpu) &&
+			cpus_share_cache(old_cpu, next_cpu))
+			goto out;
+
 		rxq_index = cpu_rmap_lookup_index(dev->rx_cpu_rmap, next_cpu);
 		if (rxq_index == skb_get_rx_queue(skb))
 			goto out;
@@ -5308,8 +5327,8 @@ static int get_rps_cpu(struct net_device *dev, struct sk_buff *skb,
 		    (tcpu >= nr_cpu_ids || !cpu_online(tcpu) ||
 		     ((int)(READ_ONCE(per_cpu(softnet_data, tcpu).input_queue_head) -
 		      rflow->last_qtail)) >= 0)) {
+			rflow = set_rps_cpu(dev, skb, rflow, tcpu, next_cpu, hash);
 			tcpu = next_cpu;
-			rflow = set_rps_cpu(dev, skb, rflow, next_cpu, hash);
 		}
 
 		if (tcpu < nr_cpu_ids && cpu_online(tcpu)) {
diff --git a/net/core/sysctl_net_core.c b/net/core/sysctl_net_core.c
index 473c162736d6..6043ccb52cff 100644
--- a/net/core/sysctl_net_core.c
+++ b/net/core/sysctl_net_core.c
@@ -560,6 +560,13 @@ static struct ctl_table net_core_table[] = {
 		.mode		= 0644,
 		.proc_handler	= rps_sock_flow_sysctl
 	},
+	{
+		.procname	= "rps_feat_llc_affinity",
+		.data		= &rps_feat_llc_affinity.key,
+		.maxlen		= sizeof(rps_feat_llc_affinity.key),
+		.mode		= 0644,
+		.proc_handler	= proc_do_static_key
+	},
 #endif
 #ifdef CONFIG_NET_FLOW_LIMIT
 	{
-- 
2.47.3


^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH net-next v10] net: reduce ARFS flow updates by checking LLC affinity
  2026-10-08 13:32 [PATCH net-next v10] net: reduce ARFS flow updates by checking LLC affinity Chuang Wang
@ 2026-10-08 14:14 ` Neal Cardwell
  0 siblings, 0 replies; 2+ messages in thread
From: Neal Cardwell @ 2026-10-08 14:14 UTC (permalink / raw)
  To: Chuang Wang
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Kuniyuki Iwashima, Stanislav Fomichev, Hangbin Liu,
	Samiullah Khawaja, Roman Gushchin, netdev, linux-kernel

On Thu, Oct 8, 2026 at 9:33 AM Chuang Wang <nashuiliang@gmail.com> wrote:
>
...
> --- a/net/core/sysctl_net_core.c
> +++ b/net/core/sysctl_net_core.c
> @@ -560,6 +560,13 @@ static struct ctl_table net_core_table[] = {
>                 .mode           = 0644,
>                 .proc_handler   = rps_sock_flow_sysctl
>         },
> +       {
> +               .procname       = "rps_feat_llc_affinity",
> +               .data           = &rps_feat_llc_affinity.key,
> +               .maxlen         = sizeof(rps_feat_llc_affinity.key),
> +               .mode           = 0644,
> +               .proc_handler   = proc_do_static_key
> +       },

Having "feat_" in the name of a sysctl seems redundant, since most
sysctls are features.

Also, the name seems overly broad, as it could also apply to other
RPS/RFS mechanisms leveraging LLC affinity, for which we might want
separate sysctl control knobs. I'd suggest perhaps something more
specific, like "rps_llc_affinity_for_set_cpu".

neal

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-08 14:14 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 13:32 [PATCH net-next v10] net: reduce ARFS flow updates by checking LLC affinity Chuang Wang
2026-10-08 14:14 ` Neal Cardwell

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®