* [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking
@ 2026-09-28 6:42 Karl Mehltretter
2026-09-28 6:42 ` [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
` (4 more replies)
0 siblings, 5 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-28 6:42 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger, linux-kernel,
linux-rt-devel, stable
netpoll's deferred transmit path can acquire two sleeping locks with hard
interrupts disabled on PREEMPT_RT. One is the sk_buff_head lock used by
the sender and worker. The other is the device transmit lock used by the
worker. They produce independent atomic-sleep reports.
Patch 1 gives the deferred skb queue a dedicated raw lock.
Patch 2 makes the worker try the device transmit lock and defer the skb
again when the lock is busy. It uses the queue helper added by patch 1.
Testing:
- Arm64 and x86_64 PREEMPT_RT W=1 builds passed.
- An unpatched Pi 400 emitted the queue-lock report five times during a
one-hour run. Ten boots with the complete fix passed, including one
3,606-second run. Every trigger and post-trigger marker arrived, Wi-Fi
remained usable, and neither atomic-sleep warning occurred.
- An arm64 PREEMPT_RT QEMU control reproduced both warnings under forced
transmit backpressure. Three boots with the complete fix passed.
- An x86_64 QEMU control reproduced the queue-lock warning on PREEMPT_RT.
Three fixed RT boots passed. The non-RT control and treatment each
passed three boots.
Karl Mehltretter (2):
netpoll: use a raw lock for the deferred transmit queue
netpoll: avoid blocking on the transmit lock in queue_process
include/linux/netpoll.h | 1 +
net/core/netpoll.c | 83 +++++++++++++++++++++++++++++++++++++----
2 files changed, 77 insertions(+), 7 deletions(-)
base-commit: a7bfaba4823e3c165bb2004c74eff7c096672bc7
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
@ 2026-09-28 6:42 ` Karl Mehltretter
2026-09-30 21:43 ` netdev-bot+sashiko
2026-09-28 6:42 ` [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
` (3 subsequent siblings)
4 siblings, 1 reply; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-28 6:42 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger, linux-kernel,
linux-rt-devel, stable
netpoll_send_skb() calls __netpoll_send_skb() with hard interrupts
disabled. When direct transmission cannot complete, the latter queues
the skb with skb_queue_tail(). The sk_buff_head lock may sleep on
PREEMPT_RT:
BUG: sleeping function called from invalid context
in_atomic(): 0, irqs_disabled(): 1, non_block: 0
rt_spin_lock
skb_queue_tail
netpoll_send_skb
The delayed transmit worker has the same problem when it requeues a busy
skb with skb_queue_head() after disabling interrupts.
Add a dedicated raw spinlock and use the unlocked skb queue helpers under
it. Keep raw critical sections limited to queue operations. During
cleanup, splice the queue to a private list before freeing its skbs.
Fixes: b6cd27ed3388 ("netpoll per device txq")
Cc: stable@vger.kernel.org # 6.12+
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
include/linux/netpoll.h | 1 +
net/core/netpoll.c | 75 +++++++++++++++++++++++++++++++++++++----
2 files changed, 70 insertions(+), 6 deletions(-)
diff --git a/include/linux/netpoll.h b/include/linux/netpoll.h
index 1c6b1eec5efd6..e20e0592e9349 100644
--- a/include/linux/netpoll.h
+++ b/include/linux/netpoll.h
@@ -47,6 +47,7 @@ struct netpoll_info {
struct semaphore dev_lock;
struct sk_buff_head txq;
+ raw_spinlock_t txq_lock;
struct delayed_work tx_work;
diff --git a/net/core/netpoll.c b/net/core/netpoll.c
index fe1e0cda5d6bf..e0cfcb05468e2 100644
--- a/net/core/netpoll.c
+++ b/net/core/netpoll.c
@@ -79,6 +79,68 @@ static netdev_tx_t netpoll_start_xmit(struct sk_buff *skb,
return status;
}
+/*
+ * Transmit paths can access txq with hard IRQs disabled. Use a raw lock
+ * because the skb queue lock may sleep on PREEMPT_RT.
+ */
+static bool netpoll_txq_empty(struct netpoll_info *npinfo)
+{
+ unsigned long flags;
+ bool empty;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ empty = skb_queue_empty(&npinfo->txq);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ return empty;
+}
+
+static struct sk_buff *netpoll_txq_dequeue(struct netpoll_info *npinfo)
+{
+ unsigned long flags;
+ struct sk_buff *skb;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ skb = __skb_dequeue(&npinfo->txq);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ return skb;
+}
+
+static void netpoll_txq_queue_head(struct netpoll_info *npinfo,
+ struct sk_buff *skb)
+{
+ unsigned long flags;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ __skb_queue_head(&npinfo->txq, skb);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+}
+
+static void netpoll_txq_queue_tail(struct netpoll_info *npinfo,
+ struct sk_buff *skb)
+{
+ unsigned long flags;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ __skb_queue_tail(&npinfo->txq, skb);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+}
+
+static void netpoll_txq_purge(struct netpoll_info *npinfo)
+{
+ struct sk_buff_head purge;
+ unsigned long flags;
+
+ __skb_queue_head_init(&purge);
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ skb_queue_splice_init(&npinfo->txq, &purge);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ __skb_queue_purge(&purge);
+}
+
static void queue_process(struct work_struct *work)
{
struct netpoll_info *npinfo =
@@ -86,7 +148,7 @@ static void queue_process(struct work_struct *work)
struct sk_buff *skb;
unsigned long flags;
- while ((skb = skb_dequeue(&npinfo->txq))) {
+ while ((skb = netpoll_txq_dequeue(npinfo))) {
struct net_device *dev = skb->dev;
struct netdev_queue *txq;
unsigned int q_index;
@@ -107,7 +169,7 @@ static void queue_process(struct work_struct *work)
HARD_TX_LOCK(dev, txq, smp_processor_id());
if (netif_xmit_frozen_or_stopped(txq) ||
!dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
- skb_queue_head(&npinfo->txq, skb);
+ netpoll_txq_queue_head(npinfo, skb);
HARD_TX_UNLOCK(dev, txq);
local_irq_restore(flags);
@@ -282,7 +344,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb)
}
/* don't get messages out of order, and no recursion */
- if (skb_queue_len(&npinfo->txq) == 0 && !netpoll_owner_active(dev)) {
+ if (netpoll_txq_empty(npinfo) && !netpoll_owner_active(dev)) {
struct netdev_queue *txq;
txq = netdev_core_pick_tx(dev, skb, NULL);
@@ -314,7 +376,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb)
}
if (!dev_xmit_complete(status)) {
- skb_queue_tail(&npinfo->txq, skb);
+ netpoll_txq_queue_tail(npinfo, skb);
schedule_delayed_work(&npinfo->tx_work,0);
}
ret = NETDEV_TX_OK;
@@ -362,7 +424,8 @@ int __netpoll_setup(struct netpoll *np, struct net_device *ndev)
}
sema_init(&npinfo->dev_lock, 1);
- skb_queue_head_init(&npinfo->txq);
+ __skb_queue_head_init(&npinfo->txq);
+ raw_spin_lock_init(&npinfo->txq_lock);
INIT_DELAYED_WORK(&npinfo->tx_work, queue_process);
refcount_set(&npinfo->refcnt, 1);
@@ -397,7 +460,7 @@ static void rcu_cleanup_netpoll_info(struct rcu_head *rcu_head)
struct netpoll_info *npinfo =
container_of(rcu_head, struct netpoll_info, rcu);
- skb_queue_purge(&npinfo->txq);
+ netpoll_txq_purge(npinfo);
kfree(npinfo);
}
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
2026-09-28 6:42 ` [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
@ 2026-09-28 6:42 ` Karl Mehltretter
2026-09-29 6:43 ` sashiko-bot
2026-09-30 21:43 ` netdev-bot+sashiko
2026-09-30 19:21 ` [PATCH net v2 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
` (2 subsequent siblings)
4 siblings, 2 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-28 6:42 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger, linux-kernel,
linux-rt-devel, stable
queue_process() disables interrupts before taking the device transmit
lock with HARD_TX_LOCK(). The lock may sleep on PREEMPT_RT:
BUG: sleeping function called from invalid context
in_atomic(): 0, irqs_disabled(): 1, non_block: 0
Workqueue: events queue_process
rt_spin_lock
queue_process
process_one_work
Use HARD_TX_TRYLOCK() instead. If another CPU owns the lock, put the skb
back at the head of the deferred queue, restore interrupts and retry
after the existing HZ / 10 delay.
Fixes: 3640543df26f ("[PATCH] netpoll: fix netpoll lockup")
Cc: stable@vger.kernel.org # 6.12+
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
net/core/netpoll.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/net/core/netpoll.c b/net/core/netpoll.c
index e0cfcb05468e2..a3c94f0b3f764 100644
--- a/net/core/netpoll.c
+++ b/net/core/netpoll.c
@@ -166,7 +166,13 @@ static void queue_process(struct work_struct *work)
skb_set_queue_mapping(skb, q_index);
}
txq = netdev_get_tx_queue(dev, q_index);
- HARD_TX_LOCK(dev, txq, smp_processor_id());
+ if (!HARD_TX_TRYLOCK(dev, txq)) {
+ netpoll_txq_queue_head(npinfo, skb);
+ local_irq_restore(flags);
+
+ schedule_delayed_work(&npinfo->tx_work, HZ / 10);
+ return;
+ }
if (netif_xmit_frozen_or_stopped(txq) ||
!dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
netpoll_txq_queue_head(npinfo, skb);
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process
2026-09-28 6:42 ` [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
@ 2026-09-29 6:43 ` sashiko-bot
2026-09-30 18:38 ` Karl Mehltretter
2026-09-30 21:43 ` netdev-bot+sashiko
1 sibling, 1 reply; 10+ messages in thread
From: sashiko-bot @ 2026-09-29 6:43 UTC (permalink / raw)
To: Karl Mehltretter; +Cc: linux-kernel, linux-rt-devel
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] net: severe latency in netpoll queue_process on lock contention
--
commit eadeea48c72992b0b24239d8a823d8f836b84527
Author: Karl Mehltretter <kmehltretter@gmail.com>
netpoll: avoid blocking on the transmit lock in queue_process
This commit replaces HARD_TX_LOCK() with HARD_TX_TRYLOCK() in queue_process()
to fix a sleeping-in-atomic bug on PREEMPT_RT. If the lock is held, it requeues
the skb and schedules a delayed work retry after HZ / 10.
> diff --git a/net/core/netpoll.c b/net/core/netpoll.c
> index e0cfcb05468e2..a3c94f0b3f764 100644
> --- a/net/core/netpoll.c
> +++ b/net/core/netpoll.c
> @@ -166,7 +166,13 @@ static void queue_process(struct work_struct *work)
> skb_set_queue_mapping(skb, q_index);
> }
> txq = netdev_get_tx_queue(dev, q_index);
> - HARD_TX_LOCK(dev, txq, smp_processor_id());
> + if (!HARD_TX_TRYLOCK(dev, txq)) {
> + netpoll_txq_queue_head(npinfo, skb);
> + local_irq_restore(flags);
> +
> + schedule_delayed_work(&npinfo->tx_work, HZ / 10);
> + return;
[Severity: High]
Does this schedule an excessively long delay when encountering a transient
spinlock contention?
When queue_process() attempts to transmit a deferred skb, it uses
HARD_TX_TRYLOCK(). If another CPU is actively transmitting on the same queue,
the TX lock is transiently held and the trylock fails.
The TX spinlock is typically held for just a few microseconds during normal
network operations, but this introduces a 100ms (HZ / 10) delay upon lock
contention. Will this cause netpoll to stall for 100ms every time it
encounters a busy TX lock, dropping throughput drastically (e.g., to ~10
pkts/sec) and likely causing the deferred queue to overflow resulting in
lost kernel logs?
> + }
> if (netif_xmit_frozen_or_stopped(txq) ||
> !dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
> netpoll_txq_queue_head(npinfo, skb);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928064239.32456-1-kmehltretter@gmail.com?part=2
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process
2026-09-29 6:43 ` sashiko-bot
@ 2026-09-30 18:38 ` Karl Mehltretter
0 siblings, 0 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-30 18:38 UTC (permalink / raw)
To: sashiko-reviews
Cc: linux-kernel, linux-rt-devel, netdev, sashiko-bot,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Sebastian Andrzej Siewior, Clark Williams,
Steven Rostedt, Stephen Hemminger
On Tue, Sep 29, 2026 at 06:43:06AM +0100, sashiko-bot@kernel.org wrote:
> > + schedule_delayed_work(&npinfo->tx_work, HZ / 10);
> > + return;
>
> [Severity: High]
> Does this schedule an excessively long delay when encountering a transient
> spinlock contention?
Yes, HZ / 10 is too long for transient TX-lock contention. I've changed
it in v2 to retry after one jiffy. The existing HZ / 10 backoff remains
for a stopped transmit queue or NETDEV_TX_BUSY.
Karl
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH net v2 0/2] netpoll: fix PREEMPT_RT deferred transmit locking
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
2026-09-28 6:42 ` [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
2026-09-28 6:42 ` [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
@ 2026-09-30 19:21 ` Karl Mehltretter
2026-09-30 19:21 ` [PATCH net v2 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
2026-09-30 19:21 ` [PATCH net v2 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
4 siblings, 0 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-30 19:21 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger,
Arend van Spriel, linux-kernel, linux-rt-devel, stable
netpoll's deferred transmit path can acquire two sleeping locks with hard
interrupts disabled on PREEMPT_RT. One is the sk_buff_head lock used by
the sender and worker. The other is the device transmit lock used by the
worker. They produce independent atomic-sleep reports.
Patch 1 gives the deferred skb queue a dedicated raw lock.
Patch 2 makes the worker try the device transmit lock and defer the skb
again when the lock is busy. It uses the queue helper added by patch 1.
Testing:
- Arm64 and x86_64 PREEMPT_RT W=1 builds passed.
- An unpatched Pi 400 emitted the queue-lock report five times during a
one-hour run. Ten boots with the complete fix passed, including one
3,606-second run. Every trigger and post-trigger marker arrived, Wi-Fi
remained usable, and neither atomic-sleep warning occurred.
- An arm64 PREEMPT_RT QEMU control reproduced both warnings under forced
transmit backpressure. Three boots with the complete fix passed.
- An x86_64 QEMU control reproduced the queue-lock warning on PREEMPT_RT.
Three fixed RT boots passed; the non-RT control and treatment each
passed three boots.
- An arm64 QEMU A/B test held the device transmit lock while queuing 64
skbs through a synthetic CYW43455 SDIO device. It used HZ=250 and ran
with PREEMPT_RT enabled and disabled. After releasing the lock, the
HZ / 10 version completed in 128-178 ms on RT and 108-109 ms on
non-RT. The one-jiffy version completed in 59-114 ms on RT and
28-55 ms on non-RT. All 64 skbs completed in every run.
- The exact v2 series and pending brcmfmac fix passed 10/10 counted
physical boots across a Pi 400 and Pi 500+: three RT and two non-RT
boots per board. Every lock-contention run freed all 64 skbs without a
timeout. Each board and kernel flavor completed a 1,800-second stream,
ten TXHI stop/drain/wake cycles and a Wi-Fi reconnect. No counted boot
reported an atomic-sleep warning, new lockdep splat, stall or lockup.
The A/B test applied the same pending brcmfmac netpoll fix to both
variants. The driver recorded zero or one dropped packet per run, and
host capture lost the same number; both netpoll variants showed this.
v1: https://lore.kernel.org/r/20260928064239.32456-1-kmehltretter@gmail.com/
Changes in v2:
- Retry device transmit lock contention on the next tick instead of
after HZ / 10. Keep HZ / 10 for a stopped transmit queue or a driver
which returns busy.
- Add the arm64 QEMU A/B results and combined Pi 400/Pi 500+ RT and
non-RT hardware results.
Karl Mehltretter (2):
netpoll: use a raw lock for the deferred transmit queue
netpoll: avoid blocking on the transmit lock in queue_process
include/linux/netpoll.h | 1 +
net/core/netpoll.c | 83 +++++++++++++++++++++++++++++++++++++----
2 files changed, 77 insertions(+), 7 deletions(-)
Range-diff against v1:
1: b95097f73f29 = 1: b95097f73f29 netpoll: use a raw lock for the deferred transmit queue
2: bb9c0adcd8a9 ! 2: 30f5c80cea9c netpoll: avoid blocking on the transmit lock in queue_process
@@ Commit message
process_one_work
Use HARD_TX_TRYLOCK() instead. If another CPU owns the lock, put the skb
- back at the head of the deferred queue, restore interrupts and retry
- after the existing HZ / 10 delay.
+ back at the head of the deferred queue, restore interrupts and retry on
+ the next tick. Keep the longer HZ / 10 backoff for a stopped transmit
+ queue or a driver which returns busy.
Fixes: 3640543df26f ("[PATCH] netpoll: fix netpoll lockup")
Cc: stable@vger.kernel.org # 6.12+
@@ net/core/netpoll.c: static void queue_process(struct work_struct *work)
+ netpoll_txq_queue_head(npinfo, skb);
+ local_irq_restore(flags);
+
-+ schedule_delayed_work(&npinfo->tx_work, HZ / 10);
++ schedule_delayed_work(&npinfo->tx_work, 1);
+ return;
+ }
if (netif_xmit_frozen_or_stopped(txq) ||
base-commit: a7bfaba4823e3c165bb2004c74eff7c096672bc7
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH net v2 1/2] netpoll: use a raw lock for the deferred transmit queue
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
` (2 preceding siblings ...)
2026-09-30 19:21 ` [PATCH net v2 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
@ 2026-09-30 19:21 ` Karl Mehltretter
2026-09-30 19:21 ` [PATCH net v2 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
4 siblings, 0 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-30 19:21 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger,
Arend van Spriel, linux-kernel, linux-rt-devel, stable
netpoll_send_skb() calls __netpoll_send_skb() with hard interrupts
disabled. When direct transmission cannot complete, the latter queues
the skb with skb_queue_tail(). The sk_buff_head lock may sleep on
PREEMPT_RT:
BUG: sleeping function called from invalid context
in_atomic(): 0, irqs_disabled(): 1, non_block: 0
rt_spin_lock
skb_queue_tail
netpoll_send_skb
The delayed transmit worker has the same problem when it requeues a busy
skb with skb_queue_head() after disabling interrupts.
Add a dedicated raw spinlock and use the unlocked skb queue helpers under
it. Keep raw critical sections limited to queue operations. During
cleanup, splice the queue to a private list before freeing its skbs.
Fixes: b6cd27ed3388 ("netpoll per device txq")
Cc: stable@vger.kernel.org # 6.12+
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
include/linux/netpoll.h | 1 +
net/core/netpoll.c | 75 +++++++++++++++++++++++++++++++++++++----
2 files changed, 70 insertions(+), 6 deletions(-)
diff --git a/include/linux/netpoll.h b/include/linux/netpoll.h
index 1c6b1eec5efd6..e20e0592e9349 100644
--- a/include/linux/netpoll.h
+++ b/include/linux/netpoll.h
@@ -47,6 +47,7 @@ struct netpoll_info {
struct semaphore dev_lock;
struct sk_buff_head txq;
+ raw_spinlock_t txq_lock;
struct delayed_work tx_work;
diff --git a/net/core/netpoll.c b/net/core/netpoll.c
index fe1e0cda5d6bf..e0cfcb05468e2 100644
--- a/net/core/netpoll.c
+++ b/net/core/netpoll.c
@@ -79,6 +79,68 @@ static netdev_tx_t netpoll_start_xmit(struct sk_buff *skb,
return status;
}
+/*
+ * Transmit paths can access txq with hard IRQs disabled. Use a raw lock
+ * because the skb queue lock may sleep on PREEMPT_RT.
+ */
+static bool netpoll_txq_empty(struct netpoll_info *npinfo)
+{
+ unsigned long flags;
+ bool empty;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ empty = skb_queue_empty(&npinfo->txq);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ return empty;
+}
+
+static struct sk_buff *netpoll_txq_dequeue(struct netpoll_info *npinfo)
+{
+ unsigned long flags;
+ struct sk_buff *skb;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ skb = __skb_dequeue(&npinfo->txq);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ return skb;
+}
+
+static void netpoll_txq_queue_head(struct netpoll_info *npinfo,
+ struct sk_buff *skb)
+{
+ unsigned long flags;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ __skb_queue_head(&npinfo->txq, skb);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+}
+
+static void netpoll_txq_queue_tail(struct netpoll_info *npinfo,
+ struct sk_buff *skb)
+{
+ unsigned long flags;
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ __skb_queue_tail(&npinfo->txq, skb);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+}
+
+static void netpoll_txq_purge(struct netpoll_info *npinfo)
+{
+ struct sk_buff_head purge;
+ unsigned long flags;
+
+ __skb_queue_head_init(&purge);
+
+ raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
+ skb_queue_splice_init(&npinfo->txq, &purge);
+ raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
+
+ __skb_queue_purge(&purge);
+}
+
static void queue_process(struct work_struct *work)
{
struct netpoll_info *npinfo =
@@ -86,7 +148,7 @@ static void queue_process(struct work_struct *work)
struct sk_buff *skb;
unsigned long flags;
- while ((skb = skb_dequeue(&npinfo->txq))) {
+ while ((skb = netpoll_txq_dequeue(npinfo))) {
struct net_device *dev = skb->dev;
struct netdev_queue *txq;
unsigned int q_index;
@@ -107,7 +169,7 @@ static void queue_process(struct work_struct *work)
HARD_TX_LOCK(dev, txq, smp_processor_id());
if (netif_xmit_frozen_or_stopped(txq) ||
!dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
- skb_queue_head(&npinfo->txq, skb);
+ netpoll_txq_queue_head(npinfo, skb);
HARD_TX_UNLOCK(dev, txq);
local_irq_restore(flags);
@@ -282,7 +344,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb)
}
/* don't get messages out of order, and no recursion */
- if (skb_queue_len(&npinfo->txq) == 0 && !netpoll_owner_active(dev)) {
+ if (netpoll_txq_empty(npinfo) && !netpoll_owner_active(dev)) {
struct netdev_queue *txq;
txq = netdev_core_pick_tx(dev, skb, NULL);
@@ -314,7 +376,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb)
}
if (!dev_xmit_complete(status)) {
- skb_queue_tail(&npinfo->txq, skb);
+ netpoll_txq_queue_tail(npinfo, skb);
schedule_delayed_work(&npinfo->tx_work,0);
}
ret = NETDEV_TX_OK;
@@ -362,7 +424,8 @@ int __netpoll_setup(struct netpoll *np, struct net_device *ndev)
}
sema_init(&npinfo->dev_lock, 1);
- skb_queue_head_init(&npinfo->txq);
+ __skb_queue_head_init(&npinfo->txq);
+ raw_spin_lock_init(&npinfo->txq_lock);
INIT_DELAYED_WORK(&npinfo->tx_work, queue_process);
refcount_set(&npinfo->refcnt, 1);
@@ -397,7 +460,7 @@ static void rcu_cleanup_netpoll_info(struct rcu_head *rcu_head)
struct netpoll_info *npinfo =
container_of(rcu_head, struct netpoll_info, rcu);
- skb_queue_purge(&npinfo->txq);
+ netpoll_txq_purge(npinfo);
kfree(npinfo);
}
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH net v2 2/2] netpoll: avoid blocking on the transmit lock in queue_process
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
` (3 preceding siblings ...)
2026-09-30 19:21 ` [PATCH net v2 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
@ 2026-09-30 19:21 ` Karl Mehltretter
4 siblings, 0 replies; 10+ messages in thread
From: Karl Mehltretter @ 2026-09-30 19:21 UTC (permalink / raw)
To: netdev
Cc: Karl Mehltretter, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Sebastian Andrzej Siewior,
Clark Williams, Steven Rostedt, Stephen Hemminger,
Arend van Spriel, linux-kernel, linux-rt-devel, stable
queue_process() disables interrupts before taking the device transmit
lock with HARD_TX_LOCK(). The lock may sleep on PREEMPT_RT:
BUG: sleeping function called from invalid context
in_atomic(): 0, irqs_disabled(): 1, non_block: 0
Workqueue: events queue_process
rt_spin_lock
queue_process
process_one_work
Use HARD_TX_TRYLOCK() instead. If another CPU owns the lock, put the skb
back at the head of the deferred queue, restore interrupts and retry on
the next tick. Keep the longer HZ / 10 backoff for a stopped transmit
queue or a driver which returns busy.
Fixes: 3640543df26f ("[PATCH] netpoll: fix netpoll lockup")
Cc: stable@vger.kernel.org # 6.12+
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
---
net/core/netpoll.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/net/core/netpoll.c b/net/core/netpoll.c
index e0cfcb05468e2..8d8e3c33cfdfe 100644
--- a/net/core/netpoll.c
+++ b/net/core/netpoll.c
@@ -166,7 +166,13 @@ static void queue_process(struct work_struct *work)
skb_set_queue_mapping(skb, q_index);
}
txq = netdev_get_tx_queue(dev, q_index);
- HARD_TX_LOCK(dev, txq, smp_processor_id());
+ if (!HARD_TX_TRYLOCK(dev, txq)) {
+ netpoll_txq_queue_head(npinfo, skb);
+ local_irq_restore(flags);
+
+ schedule_delayed_work(&npinfo->tx_work, 1);
+ return;
+ }
if (netif_xmit_frozen_or_stopped(txq) ||
!dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
netpoll_txq_queue_head(npinfo, skb);
--
2.53.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue
2026-09-28 6:42 ` [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
@ 2026-09-30 21:43 ` netdev-bot+sashiko
0 siblings, 0 replies; 10+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 21:43 UTC (permalink / raw)
To: kmehltretter
Cc: netdev, davem, edumazet, kuba, pabeni, horms, bigeasy, clrkwllms,
rostedt, stephen, linux-kernel, linux-rt-devel, stable
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] At this commit, queue_process() in net/core/netpoll.c still calls
HARD_TX_LOCK(dev, txq, smp_processor_id()) (line 169) right after…
- [Low] The patch adds a raw_spin_lock_irqsave(&npinfo->txq_lock) to
every netpoll send, through netpoll_txq_empty()…
--
netpoll: use a raw lock for the deferred transmit queue
This patch adds a raw spinlock, txq_lock, to struct netpoll_info. All
accesses to npinfo->txq now go through small helpers that take this lock
and use the unlocked sk_buff_head operations, so the queue can be touched
with hard IRQs disabled on PREEMPT_RT. On cleanup, the queue is spliced
to a private list before its skbs are freed.
> diff --git a/net/core/netpoll.c b/net/core/netpoll.c
> index fe1e0cda5d6b..e0cfcb05468e 100644
> --- a/net/core/netpoll.c
> +++ b/net/core/netpoll.c
> @@ -79,6 +79,68 @@ static netdev_tx_t netpoll_start_xmit(struct sk_buff *skb,
> return status;
> }
>
> +/*
> + * Transmit paths can access txq with hard IRQs disabled. Use a raw lock
> + * because the skb queue lock may sleep on PREEMPT_RT.
> + */
> +static bool netpoll_txq_empty(struct netpoll_info *npinfo)
> +{
> + unsigned long flags;
> + bool empty;
> +
> + raw_spin_lock_irqsave(&npinfo->txq_lock, flags);
> + empty = skb_queue_empty(&npinfo->txq);
> + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags);
> +
> + return empty;
> +}
[Severity: Low]
Is this lock needed here? Before the patch, __netpoll_send_skb() checked
the queue without a lock:
if (skb_queue_len(&npinfo->txq) == 0 && !netpoll_owner_active(dev)) {
and the direct transmit fast path took no txq lock at all. With this
patch, every netpoll send takes txq_lock.
The empty check and the later netpoll_txq_queue_tail() are separate
critical sections, so the result of the check is only advisory in either
case. The writers (__skb_insert() and __skb_unlink()) already use
WRITE_ONCE().
Would skb_queue_empty_lockless() be enough here?
The new lock acquisition can also hang netconsole's atomic panic path.
Netconsole's write_atomic (CON_NBCON_ATOMIC_UNSAFE) is only called from
nbcon_atomic_flush_unsafe() at the end of vpanic(). That final flush
could spin forever on txq_lock in two cases:
- the panic runs on a CPU that an NMI interrupted while it held
txq_lock in netpoll_txq_dequeue(), netpoll_txq_queue_head() or
netpoll_txq_queue_tail()
- another CPU was NMI-stopped inside one of those critical sections,
for example by crash_smp_send_stop() on x86 with
crash_kexec_post_notifiers set and no kdump image loaded
In either case the panic log would not make it out over netconsole.
This does not block the panic_timeout reboot. In vpanic(),
emergency_restart() runs before nbcon_atomic_flush_unsafe(), so the hang
can only happen with panic_timeout == 0. The same atomic path already
takes other locks that are not NMI-safe: slab locks from
alloc_skb(GFP_ATOMIC), skb_dequeue() on nt->skb_pool, and driver locks
through netpoll_poll_dev().
[ ... ]
> @@ -107,7 +169,7 @@ static void queue_process(struct work_struct *work)
> HARD_TX_LOCK(dev, txq, smp_processor_id());
> if (netif_xmit_frozen_or_stopped(txq) ||
> !dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
> - skb_queue_head(&npinfo->txq, skb);
> + netpoll_txq_queue_head(npinfo, skb);
> HARD_TX_UNLOCK(dev, txq);
> local_irq_restore(flags);
[Severity: Low]
The commit message says:
The delayed transmit worker has the same problem when it requeues a busy
skb with skb_queue_head() after disabling interrupts.
Does the worker actually reach this requeue on PREEMPT_RT? queue_process()
calls local_irq_save() and then HARD_TX_LOCK() before the requeue. For
devices without lltx, that becomes:
queue_process()
local_irq_save(flags)
HARD_TX_LOCK(dev, txq, smp_processor_id())
__netif_tx_lock()
spin_lock(&txq->_xmit_lock)
On PREEMPT_RT this spin_lock() is rt_spin_lock() with IRQs disabled. So
the worker would seem to splat on the tx lock before it gets to
netpoll_txq_queue_head().
The next commit in the series, "netpoll: avoid blocking on the transmit
lock in queue_process", replaces HARD_TX_LOCK() with HARD_TX_TRYLOCK()
and requeues on failure. The final state of the series is fine, and this
patch is no worse than the baseline.
Could the commit message say that the worker hits HARD_TX_LOCK() first,
and that the follow-up patch fixes that?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928064239.32456-1-kmehltretter%40gmail.com
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process
2026-09-28 6:42 ` [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
2026-09-29 6:43 ` sashiko-bot
@ 2026-09-30 21:43 ` netdev-bot+sashiko
1 sibling, 0 replies; 10+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 21:43 UTC (permalink / raw)
To: kmehltretter
Cc: netdev, davem, edumazet, kuba, pabeni, horms, bigeasy, clrkwllms,
rostedt, stephen, linux-kernel, linux-rt-devel, stable
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 2 · Medium: 0 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] On PREEMPT_RT the patch does not fix the bug it claims to fix.
- [High] net: severe latency in netpoll queue_process on lock contention
--
netpoll: avoid blocking on the transmit lock in queue_process
This patch replaces HARD_TX_LOCK() with HARD_TX_TRYLOCK() in
queue_process() to avoid a sleeping function called from invalid context
splat on PREEMPT_RT. If the transmit lock is contended, the skb is put back
at the head of the deferred queue and the work is rescheduled after
HZ / 10.
> diff --git a/net/core/netpoll.c b/net/core/netpoll.c
> index e0cfcb05468e..a3c94f0b3f76 100644
> --- a/net/core/netpoll.c
> +++ b/net/core/netpoll.c
> @@ -166,7 +166,13 @@ static void queue_process(struct work_struct *work)
> skb_set_queue_mapping(skb, q_index);
> }
> txq = netdev_get_tx_queue(dev, q_index);
> - HARD_TX_LOCK(dev, txq, smp_processor_id());
> + if (!HARD_TX_TRYLOCK(dev, txq)) {
> + netpoll_txq_queue_head(npinfo, skb);
> + local_irq_restore(flags);
> +
> + schedule_delayed_work(&npinfo->tx_work, HZ / 10);
> + return;
> + }
> if (netif_xmit_frozen_or_stopped(txq) ||
> !dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) {
[Severity: High]
Is HZ / 10 the right backoff for a failed trylock?
The existing HZ/10 retry further down covers a frozen or stopped queue,
or a driver returning NETDEV_TX_BUSY. That usually means the ring is
full and will take some time to drain.
A failed HARD_TX_TRYLOCK() here means something much smaller: another
CPU is inside ndo_start_xmit() on the same txq. That lock is normally
held for only a few microseconds.
With this change, every time queue_process() races with regular
traffic on that queue, the deferred skb waits 100ms. So does everything
queued behind it in npinfo->txq.
__netpoll_send_skb() only transmits directly when npinfo->txq is empty.
That means new netconsole messages keep being appended to the deferred
queue for the whole backoff as well.
On a busy interface this can cut netconsole down to a handful of
packets per second per queue. The deferred queue keeps growing and
console output falls further and further behind. Anything still sitting
in npinfo->txq when the machine hangs is never sent, and those are
exactly the messages netconsole users care about most.
This also changes behaviour on !PREEMPT_RT kernels. There,
HARD_TX_LOCK() was not a problem and simply waited out the short
critical section.
Could the retry on trylock failure be much shorter, for example a delay
of 0 or 1 jiffy? Alternatively, would it be better to avoid the trylock
altogether by not running this path with hard IRQs disabled, so the
normal tx lock can be taken?
[Severity: High]
Does this fully fix the PREEMPT_RT splat described in the commit message?
The earlier local_irq_save(flags) in queue_process() is unchanged. That
means hard IRQs are still disabled when the driver's ndo_start_xmit() runs
after a successful trylock:
queue_process()
local_irq_save(flags);
HARD_TX_TRYLOCK(dev, txq)
netpoll_start_xmit()
netdev_start_xmit()
ndo_start_xmit()
Several non-lltx drivers take a spinlock_t in their xmit path. On
PREEMPT_RT a spinlock_t is an rtmutex. For example:
drivers/net/ethernet/broadcom/genet/bcmgenet.c:bcmgenet_xmit() {
...
spin_lock(&ring->lock);
...
}
rtl8139_start_xmit() takes spin_lock_irqsave(&tp->lock, flags), and
macb_start_xmit() takes spin_lock_irqsave(&queue->tx_ptr_lock, flags).
HARD_TX_TRYLOCK() takes no core lock for lltx drivers, and those drivers do
their own spinlock_t locking, so they have the same problem.
On PREEMPT_RT, rt_spin_lock()->__rt_spin_lock() calls
rtlock_might_resched(), and __might_resched() prints "BUG: sleeping
function called from invalid context" whenever irqs_disabled() is true.
Won't the same splat still fire, just one frame deeper?
rt_spin_lock
bcmgenet_xmit
netpoll_start_xmit
queue_process
If the driver lock is contended, for example by threaded NAPI TX reclaim on
RT, rtlock_slowlock() can actually block while hard IRQs are disabled.
The trylock only seems to silence the txq lock case, because
__rt_spin_trylock() skips rtlock_might_resched(). The sibling path
__netpoll_send_skb() already uses HARD_TX_TRYLOCK() under local_irq_save(),
and it still runs driver xmit and netpoll_poll_dev() with IRQs disabled.
Documentation/networking/netdevices.rst says ndo_start_xmit "will be called
with interrupts disabled by netconsole". On RT, that conflicts with drivers
that use a plain spinlock_t in xmit.
queue_process() runs from a workqueue in process context. Would it be
better to stop disabling hard IRQs here instead of changing the lock
primitive? For example, it could use local_bh_disable() plus the normal tx
lock, as the regular xmit path does.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928064239.32456-1-kmehltretter%40gmail.com
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2026-09-30 21:43 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 6:42 [PATCH net 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
2026-09-28 6:42 ` [PATCH net 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
2026-09-30 21:43 ` netdev-bot+sashiko
2026-09-28 6:42 ` [PATCH net 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
2026-09-29 6:43 ` sashiko-bot
2026-09-30 18:38 ` Karl Mehltretter
2026-09-30 21:43 ` netdev-bot+sashiko
2026-09-30 19:21 ` [PATCH net v2 0/2] netpoll: fix PREEMPT_RT deferred transmit locking Karl Mehltretter
2026-09-30 19:21 ` [PATCH net v2 1/2] netpoll: use a raw lock for the deferred transmit queue Karl Mehltretter
2026-09-30 19:21 ` [PATCH net v2 2/2] netpoll: avoid blocking on the transmit lock in queue_process Karl Mehltretter
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®