mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Tim JH Chen <tim770802@gmail.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, andrew+netdev@lunn.ch, horms@kernel.org,
	ilpo.jarvinen@linux.intel.com, johannes@sipsolutions.net,
	loic.poulain@oss.qualcomm.com, ryazanov.s.a@gmail.com,
	chandrashekar.devegowda@intel.com, haijun.liu@mediatek.com,
	ricardo.martinez@linux.intel.com, linux-kernel@vger.kernel.org,
	tim.jh.chen@wnc.com.tw, Chih.Hung.Huang@wnc.com.tw,
	Tim JH Chen <tim770802@gmail.com>
Subject: [PATCH net v6 4/4] net: wwan: t7xx: complete NAPI on the not-started RX poll early return
Date: Fri,  2 Oct 2026 09:46:38 +0800	[thread overview]
Message-ID: <20261002014638.47981-5-tim770802@gmail.com> (raw)
In-Reply-To: <20261002014638.47981-1-tim770802@gmail.com>

t7xx_dpmaif_napi_rx_poll() has an early return taken when the RX queue is
no longer started:

	if (!rxq->que_started) {
		atomic_set(&rxq->rx_processing, 0);
		pm_runtime_put_autosuspend(rxq->dpmaif_ctrl->dev);
		dev_err(..., "Work RXQ: %d has not been started\n", rxq->index);
		return work_done;
	}

work_done is 0 here, so the poll returns less than the budget without
calling napi_complete_done(). Per __napi_poll() (net/core/dev.c) a poll
that returns less than the budget is assumed to have completed the NAPI
itself, so the core does not reschedule it. NAPI_STATE_SCHED is therefore
left set and the NAPI is never polled again.

t7xx_dpmaif_rx_stop() clears que_started without completing the NAPI, on
the MD_STATE_EXCEPTION teardown path as well as on system suspend, so the
next poll takes this early return and strands NAPI_STATE_SCHED. A later
netdev close then hangs:

	t7xx_ccmni_close()
	  t7xx_ccmni_disable_napi()
	    napi_synchronize()	<- spins on NAPI_STATE_SCHED forever

napi_synchronize() runs with rtnl_lock held, so every subsequent rtnl
operation blocks and the network stack wedges (a hung task, not a soft
lockup).

Complete the NAPI on this early return so NAPI_STATE_SCHED is cleared and
napi_synchronize() can make progress. The sleep-lock retry branch just
below already calls napi_complete_done() before rescheduling, so only the
not-started branch needs the fix.

This is a pre-existing bug, independent of the TX/RX data path vs system
PM suspend race fixed earlier in this series. It was only reached once
that race stopped soft-locking the CPU before the teardown could run.

Fixes: d642b012df70 ("net: wwan: t7xx: Add data path interface")
Signed-off-by: Tim JH Chen <tim770802@gmail.com>
---
 drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c b/drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c
index 0e1174ee611d..6272956fe053 100644
--- a/drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c
+++ b/drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c
@@ -848,6 +848,12 @@ int t7xx_dpmaif_napi_rx_poll(struct napi_struct *napi, const int budget)
 		atomic_set(&rxq->rx_processing, 0);
 		pm_runtime_put_autosuspend(rxq->dpmaif_ctrl->dev);
 		dev_err(rxq->dpmaif_ctrl->dev, "Work RXQ: %d has not been started\n", rxq->index);
+		/* Returning work_done < budget without completing the NAPI would
+		 * leave NAPI_STATE_SCHED set, hanging a later napi_synchronize()
+		 * in t7xx_ccmni_disable_napi() (which holds rtnl_lock). Complete
+		 * it here so the queue is cleanly unscheduled after rx_stop().
+		 */
+		napi_complete_done(napi, work_done);
 		return work_done;
 	}
 
-- 
2.43.0


      parent reply	other threads:[~2026-10-02  1:47 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02  1:46 [PATCH net v6 0/4] net: wwan: t7xx: fix DPMAIF data path vs system PM suspend Tim JH Chen
2026-10-02  1:46 ` [PATCH net v6 1/4] net: wwan: t7xx: fix runtime PM usage count underflow on -EACCES Tim JH Chen
2026-10-02  1:46 ` [PATCH net v6 2/4] net: wwan: t7xx: do not exit the TX push kthread on resume failure Tim JH Chen
2026-10-02  1:46 ` [PATCH net v6 3/4] net: wwan: t7xx: fix race between TX/RX data path and system PM suspend Tim JH Chen
2026-10-02  1:46 ` Tim JH Chen [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261002014638.47981-5-tim770802@gmail.com \
    --to=tim770802@gmail.com \
    --cc=Chih.Hung.Huang@wnc.com.tw \
    --cc=andrew+netdev@lunn.ch \
    --cc=chandrashekar.devegowda@intel.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=haijun.liu@mediatek.com \
    --cc=horms@kernel.org \
    --cc=ilpo.jarvinen@linux.intel.com \
    --cc=johannes@sipsolutions.net \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=loic.poulain@oss.qualcomm.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=ricardo.martinez@linux.intel.com \
    --cc=ryazanov.s.a@gmail.com \
    --cc=tim.jh.chen@wnc.com.tw \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®