From: Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
To: pkshih@realtek.com
Cc: linux-wireless@vger.kernel.org, linux-kernel@vger.kernel.org,
rtl8821cerfe2@gmail.com,
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
Subject: [PATCH rtw-next 2/2] wifi: rtw88: pci: warn when a TX ring stops advancing
Date: Sat, 3 Oct 2026 02:10:13 +0300 [thread overview]
Message-ID: <20261002231013.11792-3-abdurrahmankaradag19@gmail.com> (raw)
In-Reply-To: <20261002231013.11792-1-abdurrahmankaradag19@gmail.com>
On RTL8821CE the hardware read index of a TX ring can stop moving while
the driver keeps submitting descriptors. The link stays associated and
RX keeps working, so from userspace this is a dead network with no
explanation: ping and curl report 100% loss and nothing appears in the
log.
A watchdog dumping the TXBD registers caught five of these across three
boots, all with station power save off and REG_TXPAUSE at 0x00. Two
examples:
2026-08-28 22:05 BEQ 0x3a8: 0x003c003a, unchanged over four
samples; read index 0x3c, write index 0x3a, one
descriptor free, so avail_desc() < 2 had already
stopped the BE queue
2026-09-14 17:48 BEQ 0x3a8: 0x00c9004c -> 0x00c90068; read index
0xc9 frozen while the write index ran from 0x4c
to 0x68, 124 descriptors free falling to 96
The second is the more common shape: the ring is still mostly empty and
no limit has been reached, yet traffic is already gone. Rewriting the
host write index does not restart it, so simply reissuing the doorbell
is not sufficient to recover the condition.
This patch does not recover from it either; I do not have data showing
which recovery works. It only makes the condition visible: once per
watchdog round, compare each AC ring's hardware read index with the
driver's write index, and warn if it has not moved for three rounds
while descriptors are in flight.
Only the four AC rings are checked. BCN has no index register, and
MGMT, HI0 and H2C are drained on events of their own - HI0 in
particular is the after-DTIM group queue in AP mode and may
legitimately sit still for a whole DTIM interval. The hardware read
index and a snapshot of the host write index are taken together under
irq_lock, so a submission cannot change the write index between the
two. The check is not run while scanning, because rtw_watch_dog_work()
returns before reaching it. A ring with nothing in flight has read
index equal to write index and is skipped, so an idle device is never
reported. The recorded history is discarded whenever the ring is empty,
whenever REG_TXPAUSE is set and when the rings are reset, since samples
from before say nothing about the ones after. The warning is emitted
once per episode and rearms as soon as the index moves.
This is the detection half of a patch I sent and withdrew in September;
the doorbell rewrite it used as recovery turned out not to work, and is
gone.
Signed-off-by: Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
---
drivers/net/wireless/realtek/rtw88/hci.h | 7 ++
drivers/net/wireless/realtek/rtw88/main.c | 2 +
drivers/net/wireless/realtek/rtw88/pci.c | 83 +++++++++++++++++++++++
drivers/net/wireless/realtek/rtw88/pci.h | 3 +
4 files changed, 95 insertions(+)
diff --git a/drivers/net/wireless/realtek/rtw88/hci.h b/drivers/net/wireless/realtek/rtw88/hci.h
index d4bee9c..432333e 100644
--- a/drivers/net/wireless/realtek/rtw88/hci.h
+++ b/drivers/net/wireless/realtek/rtw88/hci.h
@@ -12,6 +12,7 @@ struct rtw_hci_ops {
struct sk_buff *skb);
void (*tx_kick_off)(struct rtw_dev *rtwdev);
void (*flush_queues)(struct rtw_dev *rtwdev, u32 queues, bool drop);
+ void (*tx_stall_check)(struct rtw_dev *rtwdev);
int (*setup)(struct rtw_dev *rtwdev);
int (*start)(struct rtw_dev *rtwdev);
void (*stop)(struct rtw_dev *rtwdev);
@@ -271,6 +272,12 @@ static inline enum rtw_hci_type rtw_hci_type(struct rtw_dev *rtwdev)
return rtwdev->hci.type;
}
+static inline void rtw_hci_tx_stall_check(struct rtw_dev *rtwdev)
+{
+ if (rtwdev->hci.ops->tx_stall_check)
+ rtwdev->hci.ops->tx_stall_check(rtwdev);
+}
+
static inline void rtw_hci_flush_queues(struct rtw_dev *rtwdev, u32 queues,
bool drop)
{
diff --git a/drivers/net/wireless/realtek/rtw88/main.c b/drivers/net/wireless/realtek/rtw88/main.c
index 0f23498..a6adbda 100644
--- a/drivers/net/wireless/realtek/rtw88/main.c
+++ b/drivers/net/wireless/realtek/rtw88/main.c
@@ -273,6 +273,8 @@ static void rtw_watch_dog_work(struct work_struct *work)
/* make sure BB/RF is working for dynamic mech */
rtw_leave_lps(rtwdev);
+
+ rtw_hci_tx_stall_check(rtwdev);
rtw_coex_wl_status_check(rtwdev);
rtw_coex_query_bt_hid_list(rtwdev);
rtw_coex_active_query_bt_info(rtwdev);
diff --git a/drivers/net/wireless/realtek/rtw88/pci.c b/drivers/net/wireless/realtek/rtw88/pci.c
index ff751bf..50c9fae 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.c
+++ b/drivers/net/wireless/realtek/rtw88/pci.c
@@ -473,6 +473,13 @@ static void rtw_pci_reset_buf_desc(struct rtw_dev *rtwdev)
BIT_CLR_H2CQ_HOST_IDX | BIT_CLR_H2CQ_HW_IDX);
}
+static void rtw_pci_tx_stall_reset(struct rtw_pci_tx_ring *ring)
+{
+ ring->last_rp = U32_MAX;
+ ring->stall_rounds = 0;
+ ring->stall_warned = false;
+}
+
static void rtw_pci_wake_stopped_queues(struct rtw_dev *rtwdev,
struct rtw_pci_tx_ring *ring)
{
@@ -503,6 +510,8 @@ static void rtw_pci_reset_trx_ring(struct rtw_dev *rtwdev)
for (queue = 0; queue < RTK_MAX_TX_QUEUE_NUM; queue++) {
ring = &rtwpci->tx_rings[queue];
+ rtw_pci_tx_stall_reset(ring);
+
if (!ring->queue_stopped)
continue;
@@ -819,6 +828,79 @@ static void rtw_pci_tx_kick_off_queue(struct rtw_dev *rtwdev,
spin_unlock_bh(&rtwpci->irq_lock);
}
+/*
+ * A TX ring with descriptors in flight advances its read index within
+ * milliseconds. When the hardware stops consuming them the ring just goes
+ * quiet: the link stays associated, RX keeps working, and nothing says that
+ * TX has died until the ring fills and the queue is stopped. Report it once
+ * per episode, where an episode ends as soon as the read index moves again.
+ *
+ * Only the four AC rings are checked. BCN has no index register, and MGMT,
+ * HI0 and H2C are drained on events of their own - HI0 in particular is the
+ * after-DTIM group queue in AP mode and may legitimately sit still for a
+ * whole DTIM interval.
+ */
+#define RTW_PCI_TX_STALL_ROUNDS 3
+
+static void rtw_pci_tx_stall_check(struct rtw_dev *rtwdev)
+{
+ struct rtw_pci *rtwpci = (struct rtw_pci *)rtwdev->priv;
+ struct rtw_pci_tx_ring *ring;
+ u32 bd_idx, cur_rp, wp;
+ bool paused;
+ u8 queue;
+
+ paused = rtw_read8(rtwdev, REG_TXPAUSE);
+
+ for (queue = RTW_TX_QUEUE_BK; queue <= RTW_TX_QUEUE_VO; queue++) {
+ ring = &rtwpci->tx_rings[queue];
+
+ /* a paused ring is not advancing on purpose, and samples
+ * taken before the pause say nothing about the ones after
+ */
+ if (paused) {
+ rtw_pci_tx_stall_reset(ring);
+ continue;
+ }
+
+ spin_lock_bh(&rtwpci->irq_lock);
+ bd_idx = rtw_read32(rtwdev, rtw_pci_tx_queue_idx_addr[queue]);
+ cur_rp = (bd_idx >> 16) & TRX_BD_IDX_MASK;
+ wp = ring->r.wp;
+ spin_unlock_bh(&rtwpci->irq_lock);
+
+ /* nothing in flight, so there is no observation to carry
+ * into the next round
+ */
+ if (cur_rp == wp) {
+ rtw_pci_tx_stall_reset(ring);
+ continue;
+ }
+
+ /* the ring is draining. last_rp is U32_MAX until the first
+ * sample with descriptors in flight, and no index can equal
+ * that, so that sample lands here and only records.
+ */
+ if (cur_rp != ring->last_rp) {
+ ring->last_rp = cur_rp;
+ ring->stall_rounds = 0;
+ ring->stall_warned = false;
+ continue;
+ }
+
+ if (ring->stall_warned)
+ continue;
+
+ if (++ring->stall_rounds >= RTW_PCI_TX_STALL_ROUNDS) {
+ ring->stall_warned = true;
+ rtw_warn(rtwdev,
+ "tx ring %u: read index unchanged at 0x%03x for %u watchdog rounds, write index 0x%03x, %d descriptors free\n",
+ queue, cur_rp, ring->stall_rounds, wp,
+ avail_desc(wp, cur_rp, ring->r.len));
+ }
+ }
+}
+
static void rtw_pci_tx_kick_off(struct rtw_dev *rtwdev)
{
struct rtw_pci *rtwpci = (struct rtw_pci *)rtwdev->priv;
@@ -1637,6 +1719,7 @@ static const struct rtw_hci_ops rtw_pci_ops = {
.tx_write = rtw_pci_tx_write,
.tx_kick_off = rtw_pci_tx_kick_off,
.flush_queues = rtw_pci_flush_queues,
+ .tx_stall_check = rtw_pci_tx_stall_check,
.setup = rtw_pci_setup,
.start = rtw_pci_start,
.stop = rtw_pci_stop,
diff --git a/drivers/net/wireless/realtek/rtw88/pci.h b/drivers/net/wireless/realtek/rtw88/pci.h
index 04630d1..c6c9a63 100644
--- a/drivers/net/wireless/realtek/rtw88/pci.h
+++ b/drivers/net/wireless/realtek/rtw88/pci.h
@@ -189,6 +189,9 @@ struct rtw_pci_tx_ring {
struct sk_buff_head queue;
bool queue_stopped;
unsigned long stopped_queues;
+ u32 last_rp;
+ u8 stall_rounds;
+ bool stall_warned;
};
struct rtw_pci_rx_buffer_desc {
--
2.55.0
prev parent reply other threads:[~2026-10-02 23:10 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 23:10 [PATCH rtw-next 0/2] rtw88: recover a stopped TX queue, and notice a stalled ring Abdurrahman Karadag
2026-10-02 23:10 ` [PATCH rtw-next 1/2] wifi: rtw88: pci: wake the TX queues when the rings are reset Abdurrahman Karadag
2026-10-02 23:10 ` Abdurrahman Karadag [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002231013.11792-3-abdurrahmankaradag19@gmail.com \
--to=abdurrahmankaradag19@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-wireless@vger.kernel.org \
--cc=pkshih@realtek.com \
--cc=rtl8821cerfe2@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®