mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net] net: clear stale traffic class mappings when num_tc shrinks
@ 2026-09-08 18:17 Norbert Szetei
  2026-09-10 18:19 ` netdev-bot+sashiko
  0 siblings, 1 reply; 3+ messages in thread
From: Norbert Szetei @ 2026-09-08 18:17 UTC (permalink / raw)
  To: netdev
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Kees Cook, Kuniyuki Iwashima, Alexander Duyck,
	linux-kernel

Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
lowers dev->num_tc without touching either array.

netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
entry is handed out as a traffic class >= dev->num_tc. Taking that
class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
negative one and indexes an XPS map sized for dev->num_tc classes:

	tci = j * num_tc + tc;
	RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);

Any caller that lowers num_tc leaves such entries behind, and
mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
still describe txq 1..7. The splat is from an XPS write to txq 2:

  BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue+0x1eb9/0x2440
  Write of size 8 at addr ffff888110e978d8 by task xps_oob/573
   __netif_set_xps_queue+0x1eb9/0x2440
   xps_rxqs_store+0x24d/0x360
   netdev_queue_attr_store+0x61/0x90
  Allocated by task 573:
   __kmalloc_noprof+0x246/0x6c0
   __netif_set_xps_queue+0x8ca/0x2440
  The buggy address is located 0 bytes to the right of
   allocated 88-byte region [ffff888110e97880, ffff888110e978d8)

Clear the entries the new num_tc no longer covers, before publishing
num_tc so that a lockless reader cannot observe the new num_tc together
with the old mappings. For num_tc == 0 this leaves the same state as
netdev_reset_tc().

Fixes: 184c449f91fe ("net: Add support for XPS with QoS via traffic classes")
Assisted-by: LLM
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
---
A follow-up for net-next can factor this into a helper shared with
netdev_reset_tc() and netdev_unbind_sb_channel().

I can share the reproducer on request.

 net/core/dev.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/net/core/dev.c b/net/core/dev.c
index ecfbd72..e63e36a 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3145,6 +3145,8 @@ EXPORT_SYMBOL(netdev_set_tc_queue);
 
 int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
 {
+	int i;
+
 	if (num_tc > TC_MAX_QUEUE)
 		return -EINVAL;
 
@@ -3153,6 +3155,14 @@ int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
 #endif
 	netdev_unbind_all_sb_channels(dev);
 
+	/* Drop the mappings the new num_tc no longer covers. */
+	for (i = num_tc; i < TC_MAX_QUEUE; i++)
+		WRITE_ONCE(dev->tc_to_txq[i].combined, 0);
+	for (i = 0; i <= TC_BITMASK; i++) {
+		if (READ_ONCE(dev->prio_tc_map[i]) >= num_tc)
+			WRITE_ONCE(dev->prio_tc_map[i], 0);
+	}
+
 	WRITE_ONCE(dev->num_tc, num_tc);
 	return 0;
 }
-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH net] net: clear stale traffic class mappings when num_tc shrinks
  2026-09-08 18:17 [PATCH net] net: clear stale traffic class mappings when num_tc shrinks Norbert Szetei
@ 2026-09-10 18:19 ` netdev-bot+sashiko
  2026-09-12  5:47   ` Norbert Szetei
  0 siblings, 1 reply; 3+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 18:19 UTC (permalink / raw)
  To: norbert
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, kees, kuniyu,
	alexander.duyck, linux-kernel

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider:

Critical: 0 · High: 1 · Medium: 1 · Low: 0

- [High] The patch keeps the 'tc < num_tc' invariant only at a distance
  (scrubbing the source arrays) and never validates it at the point of…
- [Medium] netdev_set_num_tc() now zeroes
  dev->tc_to_txq[num_tc..TC_MAX_QUEUE-1] and the out-of-range…

--

net: clear stale traffic class mappings when num_tc shrinks

netdev_set_num_tc() lowers dev->num_tc without touching dev->tc_to_txq[]
or dev->prio_tc_map[], so stale entries naming a class >= dev->num_tc are
handed back out by netdev_txq_to_tc() and netdev_get_prio_tc_map().  The
patch clears the entries the new num_tc no longer covers before publishing
the new num_tc.

> diff --git a/net/core/dev.c b/net/core/dev.c
> index ecfbd72d5d1a..e63e36a7c4d4 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3145,6 +3145,8 @@ EXPORT_SYMBOL(netdev_set_tc_queue);
>  
>  int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
>  {
> +	int i;
> +
>  	if (num_tc > TC_MAX_QUEUE)
>  		return -EINVAL;
>  
> @@ -3153,6 +3155,14 @@ int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
>  #endif
>  	netdev_unbind_all_sb_channels(dev);
>  
> +	/* Drop the mappings the new num_tc no longer covers. */
> +	for (i = num_tc; i < TC_MAX_QUEUE; i++)
> +		WRITE_ONCE(dev->tc_to_txq[i].combined, 0);
> +	for (i = 0; i <= TC_BITMASK; i++) {
> +		if (READ_ONCE(dev->prio_tc_map[i]) >= num_tc)
> +			WRITE_ONCE(dev->prio_tc_map[i], 0);
> +	}
> +

[Severity: Medium]
These two loops run while the old, larger dev->num_tc is still published.
Can a concurrent lockless transmitter observe the old num_tc together with
already-cleared mappings?

netdev_pick_tx() -> skb_tx_hash() gates on num_tc and then reads
tc_to_txq[] with no lock:

net/core/dev.c:skb_tx_hash() {
	if (READ_ONCE(dev->num_tc)) {
		u8 tc = netdev_get_prio_tc_map(dev, skb->priority);
		struct netdev_tc_txq res;

		res.combined = READ_ONCE(sb_dev->tc_to_txq[tc].combined);
		qoffset = res.offset;
		qcount = res.count;
		if (unlikely(!qcount)) {
			net_warn_ratelimited("%s: invalid qcount, qoffset %u for tc %u\n",
					     sb_dev->name, qoffset, tc);
			qoffset = 0;
			qcount = dev->real_num_tx_queues;
		}
	}

During the window the entry is already zero, so qcount == 0, the
"invalid qcount" warning is emitted and the packet is hashed over the whole
device queue range instead of the class range that num_tc still advertises.

netdev_txq_to_tc() has the same exposure and returns -1 while num_tc is
still non-zero:

net/core/dev.c:netdev_txq_to_tc() {
	if (READ_ONCE(dev->num_tc)) {
		...
		/* didn't find it, just return -1 to indicate no match */
		return -1;
	}

DSA taggers store that result in a u8 (ocelot_xmit() in
net/dsa/tag_ocelot_8021q.c, and the equivalents in net/dsa/tag_ksz.c and
net/dsa/tag_sja1105.c), which turns into a priority of 255.

The peer teardown helper netdev_reset_tc() uses the opposite store order,
publishing num_tc first and clearing the arrays afterwards:

net/core/dev.c:netdev_reset_tc() {
	WRITE_ONCE(dev->num_tc, 0);
	for (i = 0; i < TC_MAX_QUEUE; i++)
		WRITE_ONCE(dev->tc_to_txq[i].combined, 0);
	for (i = 0; i <= TC_BITMASK; i++)
		WRITE_ONCE(dev->prio_tc_map[i], 0);
}

Since every reader gates map use on num_tc, would matching that order here
avoid the intermediate state, and should the commit message mention the
inverse inconsistency this order creates?

>  	WRITE_ONCE(dev->num_tc, num_tc);
>  	return 0;
>  }

[Severity: High]
Does scrubbing the source arrays close the out-of-bounds write quoted in
the commit message for XPS callers that do not hold RTNL?

__netif_set_xps_queue() takes two independent unsynchronized snapshots and
only rejects a negative class:

net/core/dev.c:__netif_set_xps_queue() {
	dev_num_tc = READ_ONCE(dev->num_tc);
	if (dev_num_tc) {
		num_tc = dev_num_tc;
		...
		tc = netdev_txq_to_tc(dev, index);
		if (tc < 0)
			return -EINVAL;
	}

	mutex_lock(&xps_map_mutex);
	...
		tci = j * num_tc + tc;
		...
		RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);

xps_map_mutex is taken only after both loads and does not protect
dev->num_tc, and the allocation is sized with
XPS_CPU_DEV_MAPS_SIZE(num_tc) / XPS_RXQ_DEV_MAPS_SIZE(num_tc, nr_rx), so
any tc >= the snapshotted num_tc writes attr_map[] past the allocation.

Two paths look like they remain:

Raising num_tc is not covered at all.  A reader that snapshotted
num_tc == 1 and is then preempted can pick up tc == 3..7 after another CPU
ran mqprio_init() -> netdev_set_num_tc(dev, 8) plus
netdev_set_tc_queue(dev, 3, ...) in net/sched/sch_mqprio.c, and then writes
attr_map[j * 1 + 3] into a map sized for one class.

The ordering the commit message relies on is not enforced.  The message
says the entries are cleared "before publishing num_tc so that a lockless
reader cannot observe the new num_tc together with the old mappings", but
the clearing stores and WRITE_ONCE(dev->num_tc, num_tc) are plain
WRITE_ONCE() with no smp_wmb()/smp_store_release(), and netdev_txq_to_tc(),
skb_tx_hash() and netdev_get_prio_tc_map() use bare READ_ONCE() with no
smp_rmb()/acquire.  On a weakly ordered CPU, can a non-RTNL reader still
see the new smaller num_tc together with mappings that have not been
cleared yet?

A non-RTNL XPS caller exists:

drivers/net/virtio_net.c
  virtnet_cpu_online() / virtnet_cpu_dead()
    virtnet_set_affinity()
      __netif_set_xps_queue(vi->dev, cpumask_bits(mask), i, XPS_CPUS);

which runs with only cpus_read_lock() held while netdev_set_num_tc() holds
only RTNL.  The sysfs callers xps_cpus_store()/xps_rxqs_store() in
net/core/net-sysfs.c are serialized by sysfs_rtnl_lock() and are not
affected.

Would bounding the class at the point of use, for example

	tc = netdev_txq_to_tc(dev, index);
	if (tc < 0 || tc >= num_tc)
		return -EINVAL;

after the snapshot that sizes the map, together with a release/acquire pair
or shared locking for the num_tc publish, be a more direct fix than
scrubbing the arrays?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/16E3A318-5532-4B5E-8D03-86D21B463A2D%40doyensec.com

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH net] net: clear stale traffic class mappings when num_tc shrinks
  2026-09-10 18:19 ` netdev-bot+sashiko
@ 2026-09-12  5:47   ` Norbert Szetei
  0 siblings, 0 replies; 3+ messages in thread
From: Norbert Szetei @ 2026-09-12  5:47 UTC (permalink / raw)
  To: netdev-bot+sashiko
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, kees, kuniyu,
	alexander.duyck, linux-kernel

On Sep 10, 2026, at 20:19, netdev-bot+sashiko@kernel.org wrote:
> 
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider:
> 
> Critical: 0 · High: 1 · Medium: 1 · Low: 0
> 
> - [High] The patch keeps the 'tc < num_tc' invariant only at a distance
>  (scrubbing the source arrays) and never validates it at the point of…
> - [Medium] netdev_set_num_tc() now zeroes
>  dev->tc_to_txq[num_tc..TC_MAX_QUEUE-1] and the out-of-range…
> 
> --
> 
> net: clear stale traffic class mappings when num_tc shrinks
> 
> netdev_set_num_tc() lowers dev->num_tc without touching dev->tc_to_txq[]
> or dev->prio_tc_map[], so stale entries naming a class >= dev->num_tc are
> handed back out by netdev_txq_to_tc() and netdev_get_prio_tc_map().  The
> patch clears the entries the new num_tc no longer covers before publishing
> the new num_tc.
> 
>> diff --git a/net/core/dev.c b/net/core/dev.c
>> index ecfbd72d5d1a..e63e36a7c4d4 100644
>> --- a/net/core/dev.c
>> +++ b/net/core/dev.c
>> @@ -3145,6 +3145,8 @@ EXPORT_SYMBOL(netdev_set_tc_queue);
>> 
>> int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
>> {
>> + int i;
>> +
>> if (num_tc > TC_MAX_QUEUE)
>> return -EINVAL;
>> 
>> @@ -3153,6 +3155,14 @@ int netdev_set_num_tc(struct net_device *dev, u8 num_tc)
>> #endif
>> netdev_unbind_all_sb_channels(dev);
>> 
>> + /* Drop the mappings the new num_tc no longer covers. */
>> + for (i = num_tc; i < TC_MAX_QUEUE; i++)
>> + WRITE_ONCE(dev->tc_to_txq[i].combined, 0);
>> + for (i = 0; i <= TC_BITMASK; i++) {
>> + if (READ_ONCE(dev->prio_tc_map[i]) >= num_tc)
>> + WRITE_ONCE(dev->prio_tc_map[i], 0);
>> + }
>> +
> 
> [Severity: Medium]
> These two loops run while the old, larger dev->num_tc is still published.
> Can a concurrent lockless transmitter observe the old num_tc together with
> already-cleared mappings?
> 
> netdev_pick_tx() -> skb_tx_hash() gates on num_tc and then reads
> tc_to_txq[] with no lock:
> 
> net/core/dev.c:skb_tx_hash() {
> if (READ_ONCE(dev->num_tc)) {
> u8 tc = netdev_get_prio_tc_map(dev, skb->priority);
> struct netdev_tc_txq res;
> 
> res.combined = READ_ONCE(sb_dev->tc_to_txq[tc].combined);
> qoffset = res.offset;
> qcount = res.count;
> if (unlikely(!qcount)) {
> net_warn_ratelimited("%s: invalid qcount, qoffset %u for tc %u\n",
>     sb_dev->name, qoffset, tc);
> qoffset = 0;
> qcount = dev->real_num_tx_queues;
> }
> }
> 
> During the window the entry is already zero, so qcount == 0, the
> "invalid qcount" warning is emitted and the packet is hashed over the whole
> device queue range instead of the class range that num_tc still advertises.
> 
> netdev_txq_to_tc() has the same exposure and returns -1 while num_tc is
> still non-zero:
> 
> net/core/dev.c:netdev_txq_to_tc() {
> if (READ_ONCE(dev->num_tc)) {
> ...
> /* didn't find it, just return -1 to indicate no match */
> return -1;
> }
> 
> DSA taggers store that result in a u8 (ocelot_xmit() in
> net/dsa/tag_ocelot_8021q.c, and the equivalents in net/dsa/tag_ksz.c and
> net/dsa/tag_sja1105.c), which turns into a priority of 255.
> 
> The peer teardown helper netdev_reset_tc() uses the opposite store order,
> publishing num_tc first and clearing the arrays afterwards:
> 
> net/core/dev.c:netdev_reset_tc() {
> WRITE_ONCE(dev->num_tc, 0);
> for (i = 0; i < TC_MAX_QUEUE; i++)
> WRITE_ONCE(dev->tc_to_txq[i].combined, 0);
> for (i = 0; i <= TC_BITMASK; i++)
> WRITE_ONCE(dev->prio_tc_map[i], 0);
> }
> 
> Since every reader gates map use on num_tc, would matching that order here
> avoid the intermediate state, and should the commit message mention the
> inverse inconsistency this order creates?

Agreed with that, I am dropping the loops. v2 follows, I'll retitle it since
the fix moved into the XPS path.

> 
>> WRITE_ONCE(dev->num_tc, num_tc);
>> return 0;
>> }
> 
> [Severity: High]
> Does scrubbing the source arrays close the out-of-bounds write quoted in
> the commit message for XPS callers that do not hold RTNL?

No, and this is the right objection. v1 fixed the instance I could reach, 
but I missed virtnet_set_affinity() as a non-RTNL caller. v2 rejects a class 
the map has no room for instead.

> __netif_set_xps_queue() takes two independent unsynchronized snapshots and
> only rejects a negative class:
> 
> net/core/dev.c:__netif_set_xps_queue() {
> dev_num_tc = READ_ONCE(dev->num_tc);
> if (dev_num_tc) {
> num_tc = dev_num_tc;
> ...
> tc = netdev_txq_to_tc(dev, index);
> if (tc < 0)
> return -EINVAL;
> }
> 
> mutex_lock(&xps_map_mutex);
> ...
> tci = j * num_tc + tc;
> ...
> RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);
> 
> xps_map_mutex is taken only after both loads and does not protect
> dev->num_tc, and the allocation is sized with
> XPS_CPU_DEV_MAPS_SIZE(num_tc) / XPS_RXQ_DEV_MAPS_SIZE(num_tc, nr_rx), so
> any tc >= the snapshotted num_tc writes attr_map[] past the allocation.
> 
> Two paths look like they remain:
> 
> Raising num_tc is not covered at all.  A reader that snapshotted
> num_tc == 1 and is then preempted can pick up tc == 3..7 after another CPU
> ran mqprio_init() -> netdev_set_num_tc(dev, 8) plus
> netdev_set_tc_queue(dev, 3, ...) in net/sched/sch_mqprio.c, and then writes
> attr_map[j * 1 + 3] into a map sized for one class.
> 
> The ordering the commit message relies on is not enforced.  The message
> says the entries are cleared "before publishing num_tc so that a lockless
> reader cannot observe the new num_tc together with the old mappings", but
> the clearing stores and WRITE_ONCE(dev->num_tc, num_tc) are plain
> WRITE_ONCE() with no smp_wmb()/smp_store_release(), and netdev_txq_to_tc(),
> skb_tx_hash() and netdev_get_prio_tc_map() use bare READ_ONCE() with no
> smp_rmb()/acquire.  On a weakly ordered CPU, can a non-RTNL reader still
> see the new smaller num_tc together with mappings that have not been
> cleared yet?
> 
> A non-RTNL XPS caller exists:
> 
> drivers/net/virtio_net.c
>  virtnet_cpu_online() / virtnet_cpu_dead()
>    virtnet_set_affinity()
>      __netif_set_xps_queue(vi->dev, cpumask_bits(mask), i, XPS_CPUS);
> 
> which runs with only cpus_read_lock() held while netdev_set_num_tc() holds
> only RTNL.  The sysfs callers xps_cpus_store()/xps_rxqs_store() in
> net/core/net-sysfs.c are serialized by sysfs_rtnl_lock() and are not
> affected.
> 
> Would bounding the class at the point of use, for example
> 
> tc = netdev_txq_to_tc(dev, index);
> if (tc < 0 || tc >= num_tc)
> return -EINVAL;
> 
> after the snapshot that sizes the map, together with a release/acquire pair
> or shared locking for the num_tc publish, be a more direct fix than
> scrubbing the arrays?

I believe so. I left out the barrier and the extra locking, as I don't think
it's needed here.

N.

> -- 
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/16E3A318-5532-4B5E-8D03-86D21B463A2D%40doyensec.com


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-12  5:47 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-08 18:17 [PATCH net] net: clear stale traffic class mappings when num_tc shrinks Norbert Szetei
2026-09-10 18:19 ` netdev-bot+sashiko
2026-09-12  5:47   ` Norbert Szetei

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®