From: Tim JH Chen <tim770802@gmail.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, andrew+netdev@lunn.ch, horms@kernel.org,
ilpo.jarvinen@linux.intel.com, johannes@sipsolutions.net,
loic.poulain@oss.qualcomm.com, ryazanov.s.a@gmail.com,
chandrashekar.devegowda@intel.com, haijun.liu@mediatek.com,
ricardo.martinez@linux.intel.com, linux-kernel@vger.kernel.org,
tim.jh.chen@wnc.com.tw, Chih.Hung.Huang@wnc.com.tw,
Tim JH Chen <tim770802@gmail.com>
Subject: [PATCH net v6 0/4] net: wwan: t7xx: fix DPMAIF data path vs system PM suspend
Date: Fri, 2 Oct 2026 09:46:34 +0800 [thread overview]
Message-ID: <20261002014638.47981-1-tim770802@gmail.com> (raw)
This series fixes a race between the t7xx DPMAIF data path and system PM
suspend, plus three pre-existing bugs uncovered while fixing it.
With ASPM L1 enabled and repeated suspend/resume cycles, several DPMAIF
data-plane contexts take a runtime PM reference and then access device
registers. System suspend ignores that reference, so these contexts can
touch the hardware while the suspend callback is tearing it down. The
observable symptom is a CPU soft lockup in the TX push kthread:
watchdog: BUG: soft lockup - CPU#N stuck for 26s! [dpmaif_tx_hw_pu]
__pm_runtime_resume+0x5b/0x80
t7xx_dpmaif_tx_hw_push_thread+0xc4 [mtk_t7xx]
Patch 3 is the actual fix: it quiesces the DPMAIF data-plane contexts
across system suspend (a freezable TX push kthread plus an explicit drain
of the TX-done workers and the in-flight NAPI RX poll in the suspend
callback), rather than marking the data-plane workqueues freezable, which
the v5 review showed can deadlock suspend.
Patches 1, 2 and 4 are pre-existing bugs found while developing the fix,
each with its own Fixes: tag and independent of the main race:
1 - runtime PM usage-count underflow on the -EACCES path
2 - use-after-free from the TX push kthread exiting on resume failure
4 - a NAPI that is never completed on the "RX queue not started" early
return; a later napi_synchronize() under rtnl_lock then hangs the
network stack. Patch 4 is argued from the NAPI contract (a poll
returning < budget must call napi_complete_done()); it is not tied
to a specific reproducer.
Only the DPMAIF data path is touched; the CLDMA control path uses a
different mechanism and is out of scope (t7xx_hif_cldma.c is unchanged).
The data-path fix (patch 3) was tested with 500+ suspend/resume cycles
with a SIM registered and ASPM L1 enabled.
v5 -> v6:
- Drop the "freezer as a global quiesce" direction: remove every
WQ_FREEZABLE annotation added in v4/v5. Marking the data-plane
workqueues freezable can deadlock the suspend, because a
flush_work()/cancel_work_sync() issued from a context the freezer
does not freeze (the FSM kthread, or an unbind holding device_lock)
blocks until thaw_workqueues(). (Reported on v5 review.)
- Instead quiesce the DPMAIF data plane explicitly in the system
suspend callback: mask interrupts, cancel_work_sync() the TX-done
workers, call t7xx_dpmaif_rx_stop() to drain the in-flight NAPI RX
poll (the freezer cannot park a softirq), then cancel
bat_release_work and stop the hardware last.
- Keep the PM freezer only for the lone TX push kthread, and use
kthread_freezable_should_stop() instead of try_to_freeze(), so a
concurrent kthread_stop() is honoured while the thread is frozen.
- Fold in the pre-existing fixes previously deferred to a separate
series: the -EACCES runtime PM usage-count underflow (patch 1), the
TX push kthread self-exit use-after-free (patch 2), and a NAPI that
is never completed on the not-started RX poll early return, which
hangs a later napi_synchronize() under rtnl_lock (patch 4). This
revision is therefore a 4-patch series.
- Do not touch t7xx_hif_cldma.c. Restrict the commit messages to the
DPMAIF data path; the CLDMA control path uses a different mechanism
and is called out as out of scope.
v4 -> v5:
- Fix freeze deadlock in t7xx_do_tx_hw_push(): when the TX-done
workqueue (WQ_FREEZABLE) is frozen first it stops draining the DRB
ring; the kthread then loops indefinitely in the ring-full retry
branch and never reaches try_to_freeze(), causing a freezer timeout
and suspend abort. (Simon Horman)
- Extend WQ_FREEZABLE to the BAT-release and CLDMA TX/RX workqueues.
(Simon Horman) [reverted in v6, see above]
- Note the -EACCES underflow and the stale kthread pointer as
pre-existing issues to be addressed separately. [done in v6, patches 1-2]
v3 -> v4:
- Drop the tx_pm_lock / state-snapshot approach entirely and use the PM
freezer instead. The previous approach deadlocked through the runtime
PM wait queue and opened ISR windows by writing dpmaif_ctrl->state in
suspend/resume.
v2 -> v3: process fixes (Fixes tag, changelog placement).
v1 -> v2: save/restore pre-suspend state; wrap pm_runtime with a mutex.
Tim JH Chen (4):
net: wwan: t7xx: fix runtime PM usage count underflow on -EACCES
net: wwan: t7xx: do not exit the TX push kthread on resume failure
net: wwan: t7xx: fix race between TX/RX data path and system PM
suspend
net: wwan: t7xx: complete NAPI on the not-started RX poll early return
drivers/net/wwan/t7xx/t7xx_hif_dpmaif.c | 24 +++++++++-
drivers/net/wwan/t7xx/t7xx_hif_dpmaif_rx.c | 9 +++-
drivers/net/wwan/t7xx/t7xx_hif_dpmaif_tx.c | 55 ++++++++++++++++++----
3 files changed, 77 insertions(+), 11 deletions(-)
--
2.43.0
next reply other threads:[~2026-10-02 1:46 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 1:46 Tim JH Chen [this message]
2026-10-02 1:46 ` [PATCH net v6 1/4] net: wwan: t7xx: fix runtime PM usage count underflow on -EACCES Tim JH Chen
2026-10-02 1:46 ` [PATCH net v6 2/4] net: wwan: t7xx: do not exit the TX push kthread on resume failure Tim JH Chen
2026-10-02 1:46 ` [PATCH net v6 3/4] net: wwan: t7xx: fix race between TX/RX data path and system PM suspend Tim JH Chen
2026-10-02 1:46 ` [PATCH net v6 4/4] net: wwan: t7xx: complete NAPI on the not-started RX poll early return Tim JH Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002014638.47981-1-tim770802@gmail.com \
--to=tim770802@gmail.com \
--cc=Chih.Hung.Huang@wnc.com.tw \
--cc=andrew+netdev@lunn.ch \
--cc=chandrashekar.devegowda@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=haijun.liu@mediatek.com \
--cc=horms@kernel.org \
--cc=ilpo.jarvinen@linux.intel.com \
--cc=johannes@sipsolutions.net \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=loic.poulain@oss.qualcomm.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=ricardo.martinez@linux.intel.com \
--cc=ryazanov.s.a@gmail.com \
--cc=tim.jh.chen@wnc.com.tw \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®