* [PATCH rtw-next 0/2] rtw88: recover a stopped TX queue, and notice a stalled ring
@ 2026-10-02 23:10 Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 1/2] wifi: rtw88: pci: wake the TX queues when the rings are reset Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 2/2] wifi: rtw88: pci: warn when a TX ring stops advancing Abdurrahman Karadag
0 siblings, 2 replies; 3+ messages in thread
From: Abdurrahman Karadag @ 2026-10-02 23:10 UTC (permalink / raw)
To: pkshih; +Cc: linux-wireless, linux-kernel, rtl8821cerfe2, abkarada
From: abkarada <abdurrahmankaradag19@gmail.com>
These come out of a long thread about an RTL8821CE whose TX dies while
the link stays associated:
https://lore.kernel.org/linux-wireless/20260826162514.80580-1-abdurrahmankaradag19@gmail.com/
Ping-Ke asked whether triggering the driver's recovery resolves the
stuck ring. Trying to answer that turned up patch 1, which is a real
bug and independent of whatever makes the hardware stop in the first
place.
Patch 1: for a queue stopped by the PCI TX ring-full path, the flag and
the stop reason are released only from the completion loop in
rtw_pci_tx_isr(). Resetting the rings drops the pending descriptors and
frees their skbs directly, so that loop never runs for them, and the
stop survives over an empty ring that will never complete anything
again. ieee80211_restart_hw() therefore cannot recover a device that
had stopped a queue - which is precisely when you would want it to. The
patch records the queue mappings that path stops and releases them both
from the completion loop and when the reset empties the ring.
I can produce that state on demand by pausing TX in hardware, which
freezes the read index while the driver keeps submitting, the same
shape the chip shows when it wedges by itself. Same script both ways,
only the patch differs:
without patch with patch
BE queue stopped after 1 s 1 s
doorbell rewrite no effect no effect
rtw_fw_recovery() ran ran
BE stop reason afterwards 0x1 0x0
traffic none in 90 s back within 2 s
In the failing run the station reassociated twice inside those 90 s, so
the link was up; only the queue was still stopped.
Patch 2 is diagnostic and unrelated to the above: there is currently no
sign anywhere when a TX ring stops advancing. I have five captures
where the hardware read index is frozen while the write index runs on,
with power save off and REG_TXPAUSE at 0x00, and in four of them the
ring had not even filled yet - traffic was simply gone, with nothing in
the log. This warns once per episode. Checked silent across 180 s of
saturated TX, about 600 MB.
Patch 2 is the detection half of a patch I sent and withdrew in
September; its recovery was a doorbell rewrite, which I measured as
ineffective and which is gone.
Both build clean with W=1 and pass checkpatch --strict.
Abdurrahman Karadag (2):
wifi: rtw88: pci: wake the TX queues when the rings are reset
wifi: rtw88: pci: warn when a TX ring stops advancing
drivers/net/wireless/realtek/rtw88/hci.h | 7 ++
drivers/net/wireless/realtek/rtw88/main.c | 2 +
drivers/net/wireless/realtek/rtw88/pci.c | 128 ++++++++++++++++++++--
drivers/net/wireless/realtek/rtw88/pci.h | 4 +
4 files changed, 134 insertions(+), 7 deletions(-)
base-commit: 7cde94dab0e74434ccc0a387d7ad373fb3becda0
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH rtw-next 1/2] wifi: rtw88: pci: wake the TX queues when the rings are reset
2026-10-02 23:10 [PATCH rtw-next 0/2] rtw88: recover a stopped TX queue, and notice a stalled ring Abdurrahman Karadag
@ 2026-10-02 23:10 ` Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 2/2] wifi: rtw88: pci: warn when a TX ring stops advancing Abdurrahman Karadag
1 sibling, 0 replies; 3+ messages in thread
From: Abdurrahman Karadag @ 2026-10-02 23:10 UTC (permalink / raw)
To: pkshih; +Cc: linux-wireless, linux-kernel, rtl8821cerfe2, Abdurrahman Karadag
For a queue stopped by the PCI TX ring-full path, ring->queue_stopped
is normally cleared and the stop reason released only from the
completion loop in rtw_pci_tx_isr(). When the rings are reset the
pending descriptors are dropped and their skbs are freed directly by
rtw_pci_free_tx_ring_skbs(), so that loop never runs for them. The flag
and the stop reason both survive the reset, and because the ring is now
empty no completion will ever arrive to clear them. Any queue stopped
that way stays stopped.
This makes ieee80211_restart_hw() unable to recover a device that
stopped a queue before the restart, which is the opposite of what the
recovery is for.
Reproduced on an RTL8821CE by pausing TX in hardware, which freezes the
read index while the driver keeps submitting, the same shape the chip
shows when it wedges on its own:
# echo "522 f 1" > /sys/kernel/debug/ieee80211/phy0/rtw88/write_reg
# ... push traffic until the ring fills ...
BE 0x3a8: 0x0081007f, avail_desc() 1, BE queue stop reason 0x1
# (call rtw_fw_recovery() from a debug build)
firmware crash, start reset and recover
ieee80211 phy1: Hardware restart was requested
wlan0: associated
REG_TXPAUSE 0x00, BE 0x3a8: 0x00000000, ring empty
BE queue stop reason 0x1, 100% packet loss, no recovery in 90 s
The station reassociated twice during those 90 s, so the link was fine;
only the queue was still stopped. With this patch the same sequence
clears the stop reason and traffic returns within 2 s.
Record the queue mappings this path stops and release them both from
the completion loop and when the reset empties the ring. It has to be a
set rather than one value: the stop runs after every submission that
leaves fewer than two descriptors, rtw_pci_tx_write_data() still
accepts a frame while one is left, and rtw_tx_queue_mapping() places
management frames on the MGMT ring and multicast on HI0 whatever their
skb queue mapping is, so one ring can stop two different queues before
it is emptied. Keeping only the last one would leave the other stopped
for good.
Recording the mappings also limits the wake to the queues this
ring-full path actually stopped, instead of waking every mac80211
queue.
Fixes: e3037485c68e ("rtw88: new Realtek 802.11ac driver")
Signed-off-by: Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
---
drivers/net/wireless/realtek/rtw88/pci.c | 45 ++++++++++++++++++++----
drivers/net/wireless/realtek/rtw88/pci.h | 1 +
2 files changed, 39 insertions(+), 7 deletions(-)
diff --git a/drivers/net/wireless/realtek/rtw88/pci.c b/drivers/net/wireless/realtek/rtw88/pci.c
index 66d2e5f..ff751bf 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.c
+++ b/drivers/net/wireless/realtek/rtw88/pci.c
@@ -473,9 +473,41 @@ static void rtw_pci_reset_buf_desc(struct rtw_dev *rtwdev)
BIT_CLR_H2CQ_HOST_IDX | BIT_CLR_H2CQ_HW_IDX);
}
+static void rtw_pci_wake_stopped_queues(struct rtw_dev *rtwdev,
+ struct rtw_pci_tx_ring *ring)
+{
+ unsigned long q;
+
+ for_each_set_bit(q, &ring->stopped_queues, rtwdev->hw->queues)
+ ieee80211_wake_queue(rtwdev->hw, q);
+
+ ring->stopped_queues = 0;
+ ring->queue_stopped = false;
+}
+
static void rtw_pci_reset_trx_ring(struct rtw_dev *rtwdev)
{
+ struct rtw_pci *rtwpci = (struct rtw_pci *)rtwdev->priv;
+ struct rtw_pci_tx_ring *ring;
+ enum rtw_tx_queue_type queue;
+
rtw_pci_reset_buf_desc(rtwdev);
+
+ /*
+ * The rings are empty again, so nothing is left whose completion
+ * could reach the wake in rtw_pci_tx_isr(). Release the queues this
+ * path stopped - the stop reasons it set are cleared nowhere else,
+ * and over an empty ring no completion will ever arrive to clear
+ * them.
+ */
+ for (queue = 0; queue < RTK_MAX_TX_QUEUE_NUM; queue++) {
+ ring = &rtwpci->tx_rings[queue];
+
+ if (!ring->queue_stopped)
+ continue;
+
+ rtw_pci_wake_stopped_queues(rtwdev, ring);
+ }
}
static void rtw_pci_enable_interrupt(struct rtw_dev *rtwdev,
@@ -930,7 +962,10 @@ static int rtw_pci_tx_write(struct rtw_dev *rtwdev,
ring = &rtwpci->tx_rings[queue];
spin_lock_bh(&rtwpci->irq_lock);
if (avail_desc(ring->r.wp, ring->r.rp, ring->r.len) < 2) {
- ieee80211_stop_queue(rtwdev->hw, skb_get_queue_mapping(skb));
+ u16 q_map = skb_get_queue_mapping(skb);
+
+ ieee80211_stop_queue(rtwdev->hw, q_map);
+ set_bit(q_map, &ring->stopped_queues);
ring->queue_stopped = true;
}
spin_unlock_bh(&rtwpci->irq_lock);
@@ -949,7 +984,6 @@ static void rtw_pci_tx_isr(struct rtw_dev *rtwdev, struct rtw_pci *rtwpci,
u32 count;
u32 bd_idx_addr;
u32 bd_idx, cur_rp, rp_idx;
- u16 q_map;
ring = &rtwpci->tx_rings[hw_queue];
@@ -981,11 +1015,8 @@ static void rtw_pci_tx_isr(struct rtw_dev *rtwdev, struct rtw_pci *rtwpci,
}
if (ring->queue_stopped &&
- avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
- q_map = skb_get_queue_mapping(skb);
- ieee80211_wake_queue(hw, q_map);
- ring->queue_stopped = false;
- }
+ avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4)
+ rtw_pci_wake_stopped_queues(rtwdev, ring);
if (++rp_idx >= ring->r.len)
rp_idx = 0;
diff --git a/drivers/net/wireless/realtek/rtw88/pci.h b/drivers/net/wireless/realtek/rtw88/pci.h
index 8ffdea1..04630d1 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.h
+++ b/drivers/net/wireless/realtek/rtw88/pci.h
@@ -188,6 +188,7 @@ struct rtw_pci_tx_ring {
struct rtw_pci_ring r;
struct sk_buff_head queue;
bool queue_stopped;
+ unsigned long stopped_queues;
};
struct rtw_pci_rx_buffer_desc {
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH rtw-next 2/2] wifi: rtw88: pci: warn when a TX ring stops advancing
2026-10-02 23:10 [PATCH rtw-next 0/2] rtw88: recover a stopped TX queue, and notice a stalled ring Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 1/2] wifi: rtw88: pci: wake the TX queues when the rings are reset Abdurrahman Karadag
@ 2026-10-02 23:10 ` Abdurrahman Karadag
1 sibling, 0 replies; 3+ messages in thread
From: Abdurrahman Karadag @ 2026-10-02 23:10 UTC (permalink / raw)
To: pkshih; +Cc: linux-wireless, linux-kernel, rtl8821cerfe2, Abdurrahman Karadag
On RTL8821CE the hardware read index of a TX ring can stop moving while
the driver keeps submitting descriptors. The link stays associated and
RX keeps working, so from userspace this is a dead network with no
explanation: ping and curl report 100% loss and nothing appears in the
log.
A watchdog dumping the TXBD registers caught five of these across three
boots, all with station power save off and REG_TXPAUSE at 0x00. Two
examples:
2026-08-28 22:05 BEQ 0x3a8: 0x003c003a, unchanged over four
samples; read index 0x3c, write index 0x3a, one
descriptor free, so avail_desc() < 2 had already
stopped the BE queue
2026-09-14 17:48 BEQ 0x3a8: 0x00c9004c -> 0x00c90068; read index
0xc9 frozen while the write index ran from 0x4c
to 0x68, 124 descriptors free falling to 96
The second is the more common shape: the ring is still mostly empty and
no limit has been reached, yet traffic is already gone. Rewriting the
host write index does not restart it, so simply reissuing the doorbell
is not sufficient to recover the condition.
This patch does not recover from it either; I do not have data showing
which recovery works. It only makes the condition visible: once per
watchdog round, compare each AC ring's hardware read index with the
driver's write index, and warn if it has not moved for three rounds
while descriptors are in flight.
Only the four AC rings are checked. BCN has no index register, and
MGMT, HI0 and H2C are drained on events of their own - HI0 in
particular is the after-DTIM group queue in AP mode and may
legitimately sit still for a whole DTIM interval. The hardware read
index and a snapshot of the host write index are taken together under
irq_lock, so a submission cannot change the write index between the
two. The check is not run while scanning, because rtw_watch_dog_work()
returns before reaching it. A ring with nothing in flight has read
index equal to write index and is skipped, so an idle device is never
reported. The recorded history is discarded whenever the ring is empty,
whenever REG_TXPAUSE is set and when the rings are reset, since samples
from before say nothing about the ones after. The warning is emitted
once per episode and rearms as soon as the index moves.
This is the detection half of a patch I sent and withdrew in September;
the doorbell rewrite it used as recovery turned out not to work, and is
gone.
Signed-off-by: Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
---
drivers/net/wireless/realtek/rtw88/hci.h | 7 ++
drivers/net/wireless/realtek/rtw88/main.c | 2 +
drivers/net/wireless/realtek/rtw88/pci.c | 83 +++++++++++++++++++++++
drivers/net/wireless/realtek/rtw88/pci.h | 3 +
4 files changed, 95 insertions(+)
diff --git a/drivers/net/wireless/realtek/rtw88/hci.h b/drivers/net/wireless/realtek/rtw88/hci.h
index d4bee9c..432333e 100644
--- a/drivers/net/wireless/realtek/rtw88/hci.h
+++ b/drivers/net/wireless/realtek/rtw88/hci.h
@@ -12,6 +12,7 @@ struct rtw_hci_ops {
struct sk_buff *skb);
void (*tx_kick_off)(struct rtw_dev *rtwdev);
void (*flush_queues)(struct rtw_dev *rtwdev, u32 queues, bool drop);
+ void (*tx_stall_check)(struct rtw_dev *rtwdev);
int (*setup)(struct rtw_dev *rtwdev);
int (*start)(struct rtw_dev *rtwdev);
void (*stop)(struct rtw_dev *rtwdev);
@@ -271,6 +272,12 @@ static inline enum rtw_hci_type rtw_hci_type(struct rtw_dev *rtwdev)
return rtwdev->hci.type;
}
+static inline void rtw_hci_tx_stall_check(struct rtw_dev *rtwdev)
+{
+ if (rtwdev->hci.ops->tx_stall_check)
+ rtwdev->hci.ops->tx_stall_check(rtwdev);
+}
+
static inline void rtw_hci_flush_queues(struct rtw_dev *rtwdev, u32 queues,
bool drop)
{
diff --git a/drivers/net/wireless/realtek/rtw88/main.c b/drivers/net/wireless/realtek/rtw88/main.c
index 0f23498..a6adbda 100644
--- a/drivers/net/wireless/realtek/rtw88/main.c
+++ b/drivers/net/wireless/realtek/rtw88/main.c
@@ -273,6 +273,8 @@ static void rtw_watch_dog_work(struct work_struct *work)
/* make sure BB/RF is working for dynamic mech */
rtw_leave_lps(rtwdev);
+
+ rtw_hci_tx_stall_check(rtwdev);
rtw_coex_wl_status_check(rtwdev);
rtw_coex_query_bt_hid_list(rtwdev);
rtw_coex_active_query_bt_info(rtwdev);
diff --git a/drivers/net/wireless/realtek/rtw88/pci.c b/drivers/net/wireless/realtek/rtw88/pci.c
index ff751bf..50c9fae 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.c
+++ b/drivers/net/wireless/realtek/rtw88/pci.c
@@ -473,6 +473,13 @@ static void rtw_pci_reset_buf_desc(struct rtw_dev *rtwdev)
BIT_CLR_H2CQ_HOST_IDX | BIT_CLR_H2CQ_HW_IDX);
}
+static void rtw_pci_tx_stall_reset(struct rtw_pci_tx_ring *ring)
+{
+ ring->last_rp = U32_MAX;
+ ring->stall_rounds = 0;
+ ring->stall_warned = false;
+}
+
static void rtw_pci_wake_stopped_queues(struct rtw_dev *rtwdev,
struct rtw_pci_tx_ring *ring)
{
@@ -503,6 +510,8 @@ static void rtw_pci_reset_trx_ring(struct rtw_dev *rtwdev)
for (queue = 0; queue < RTK_MAX_TX_QUEUE_NUM; queue++) {
ring = &rtwpci->tx_rings[queue];
+ rtw_pci_tx_stall_reset(ring);
+
if (!ring->queue_stopped)
continue;
@@ -819,6 +828,79 @@ static void rtw_pci_tx_kick_off_queue(struct rtw_dev *rtwdev,
spin_unlock_bh(&rtwpci->irq_lock);
}
+/*
+ * A TX ring with descriptors in flight advances its read index within
+ * milliseconds. When the hardware stops consuming them the ring just goes
+ * quiet: the link stays associated, RX keeps working, and nothing says that
+ * TX has died until the ring fills and the queue is stopped. Report it once
+ * per episode, where an episode ends as soon as the read index moves again.
+ *
+ * Only the four AC rings are checked. BCN has no index register, and MGMT,
+ * HI0 and H2C are drained on events of their own - HI0 in particular is the
+ * after-DTIM group queue in AP mode and may legitimately sit still for a
+ * whole DTIM interval.
+ */
+#define RTW_PCI_TX_STALL_ROUNDS 3
+
+static void rtw_pci_tx_stall_check(struct rtw_dev *rtwdev)
+{
+ struct rtw_pci *rtwpci = (struct rtw_pci *)rtwdev->priv;
+ struct rtw_pci_tx_ring *ring;
+ u32 bd_idx, cur_rp, wp;
+ bool paused;
+ u8 queue;
+
+ paused = rtw_read8(rtwdev, REG_TXPAUSE);
+
+ for (queue = RTW_TX_QUEUE_BK; queue <= RTW_TX_QUEUE_VO; queue++) {
+ ring = &rtwpci->tx_rings[queue];
+
+ /* a paused ring is not advancing on purpose, and samples
+ * taken before the pause say nothing about the ones after
+ */
+ if (paused) {
+ rtw_pci_tx_stall_reset(ring);
+ continue;
+ }
+
+ spin_lock_bh(&rtwpci->irq_lock);
+ bd_idx = rtw_read32(rtwdev, rtw_pci_tx_queue_idx_addr[queue]);
+ cur_rp = (bd_idx >> 16) & TRX_BD_IDX_MASK;
+ wp = ring->r.wp;
+ spin_unlock_bh(&rtwpci->irq_lock);
+
+ /* nothing in flight, so there is no observation to carry
+ * into the next round
+ */
+ if (cur_rp == wp) {
+ rtw_pci_tx_stall_reset(ring);
+ continue;
+ }
+
+ /* the ring is draining. last_rp is U32_MAX until the first
+ * sample with descriptors in flight, and no index can equal
+ * that, so that sample lands here and only records.
+ */
+ if (cur_rp != ring->last_rp) {
+ ring->last_rp = cur_rp;
+ ring->stall_rounds = 0;
+ ring->stall_warned = false;
+ continue;
+ }
+
+ if (ring->stall_warned)
+ continue;
+
+ if (++ring->stall_rounds >= RTW_PCI_TX_STALL_ROUNDS) {
+ ring->stall_warned = true;
+ rtw_warn(rtwdev,
+ "tx ring %u: read index unchanged at 0x%03x for %u watchdog rounds, write index 0x%03x, %d descriptors free\n",
+ queue, cur_rp, ring->stall_rounds, wp,
+ avail_desc(wp, cur_rp, ring->r.len));
+ }
+ }
+}
+
static void rtw_pci_tx_kick_off(struct rtw_dev *rtwdev)
{
struct rtw_pci *rtwpci = (struct rtw_pci *)rtwdev->priv;
@@ -1637,6 +1719,7 @@ static const struct rtw_hci_ops rtw_pci_ops = {
.tx_write = rtw_pci_tx_write,
.tx_kick_off = rtw_pci_tx_kick_off,
.flush_queues = rtw_pci_flush_queues,
+ .tx_stall_check = rtw_pci_tx_stall_check,
.setup = rtw_pci_setup,
.start = rtw_pci_start,
.stop = rtw_pci_stop,
diff --git a/drivers/net/wireless/realtek/rtw88/pci.h b/drivers/net/wireless/realtek/rtw88/pci.h
index 04630d1..c6c9a63 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.h
+++ b/drivers/net/wireless/realtek/rtw88/pci.h
@@ -189,6 +189,9 @@ struct rtw_pci_tx_ring {
struct sk_buff_head queue;
bool queue_stopped;
unsigned long stopped_queues;
+ u32 last_rp;
+ u8 stall_rounds;
+ bool stall_warned;
};
struct rtw_pci_rx_buffer_desc {
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-02 23:10 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 23:10 [PATCH rtw-next 0/2] rtw88: recover a stopped TX queue, and notice a stalled ring Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 1/2] wifi: rtw88: pci: wake the TX queues when the rings are reset Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 2/2] wifi: rtw88: pci: warn when a TX ring stops advancing Abdurrahman Karadag
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®