mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver
@ 2026-09-30  7:46 Jack Wu via B4 Relay
  2026-09-30  7:46 ` [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core Jack Wu via B4 Relay
                   ` (5 more replies)
  0 siblings, 6 replies; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

T9XX is the PCIe host device driver for MediaTek's
t900 modem. The driver uses the WWAN framework
infrastructure to create the following control ports
and network interfaces for data transactions.
* /dev/wwan0at0 - Interface that supports AT commands.
* /dev/wwan0mbim0 - Interface conforming to the MBIM
  protocol.
* wwan0-X - Primary network interface for IP traffic.

The main blocks in the T9XX driver are:
* HW layer - Abstracts the hardware bus operations for
   the device, and provides generic interfaces for the
   transaction layer to get the device's information and
   control the device's behavior. It includes:

   * PCIe - Implements probe, removal and interrupt
     handling.
   * MHCCIF (Modem Host Cross-Core Interface) - Provides
     interrupt channels for bidirectional event
     notification such as handshake and port enumeration.

* Transaction layer - Implements data transactions for
   the control plane and the data plane. It includes:

   * DPMAIF (Data Plane Modem AP Interface) - Controls
     the hardware that provides uplink and downlink
     queues for the data path. The data exchange takes
     place using circular buffers to share data buffer
     addresses and metadata to describe the packets.
   * CLDMA (Cross Layer DMA) - Manages the hardware
     used by the port layer to send control messages to
     the device using MediaTek's CCCI (Cross-Core
     Communication Interface) protocol.
   * TX Services - Dispatch packets from the port layer
     to the device.
   * RX Services - Dispatch packets to the port layer
     when receiving packets from the device.

* Port layer - Provides control plane and data plane
   interfaces to userspace. It includes:

   * Control Plane - Provides device node interfaces
     for controlling data transactions.
   * Data Plane - Provides network link interfaces
     wwanX (0, 1, 2...) for IP data transactions.

* Device lifecycle management - Contains the logic that
   keeps the device working. It includes:

   * FSM (Finite State Machine) - Drives the T9xx
     device lifecycle (boot handshake, error recovery,
     removal), and notifies each module when the state
     changes.

The compilation of the T9XX driver is enabled by the
CONFIG_MTK_T9XX config option, which depends on
CONFIG_WWAN.

This submission covers the control plane only
(patches 1-6). The data plane will follow in a
separate series once the control plane is accepted.

---
Changes in v9:
- Patch 1 (Add PCIe core):
  - Dropped the priv->irq_type term from the mask/unmask/clear IRQ
    guards; the field was write-only and made the probe-failure path
    log a false "input irq_id=28" error
  - mtk_pci_pldr() no longer requests an ACPI_ALLOCATE_BUFFER output
    buffer from acpi_evaluate_object(): nothing read it and the
    failure branch leaked it
  - mtk_mhccif_exit() masks the MHCCIF source again after
    cancel_work_sync(), so a work item already past the unregister
    cannot leave remove() with an unmasked source and no callback
  - pcim_enable_device() -> pci_enable_device(), with explicit
    pci_clear_master()/pci_disable_device() in remove(): the devres
    action touched config space after the ACPI reset
  - The AER handler prints pci_channel_state_t with %u
  - Deleted struct mtk_dev_ops and the mtk_dev_* inline wrappers;
    every caller uses the mtk_pci_* implementation directly, and
    mtk_dev_get_dev_cfg()/mtk_pci_get_dev_cfg() are gone (no callers)
  - Single module: mtk_dev.c, pcie/Makefile and CONFIG_MTK_T9XX_PCI
    are gone, driver init/exit live in pcie/mtk_pci.c, and every
    EXPORT_SYMBOL_GPL is removed with the module boundary
  - Removed the PCI dispatch-by-hw_ver table; there is one
    implementation and it is called directly
  - Declined a BAR length sanity check: the driver binds two fixed PCI
    IDs and pcim_iomap_region() already fails probe for the
    generically detectable cases
  - Declined reordering the probe unwind to match remove(): probe
    unwinds in reverse acquisition order, remove() deviates
    deliberately and the reason is commented at the call site
  - No delay added after the reset: MRST._RST is an ACPI method, so
    acpi_evaluate_object() returns when the AML has finished, and it
    is the last device access in remove()
  - Removed the PCIe PM registers and the MHCCIF PM/DPMAIF/CLDMA
    channel bits; nothing in this series reads them
  - Removed the unused IRQ sources, MTK_USER_MIN/DATA, DEV_EVT_*_MAX,
    ATR_DST_AXIM_1-3 and the <linux/debugfs.h> include
- Patch 2 (Add control plane transaction layer):
  - Deleted struct mtk_ctrl_hif_ops and ctrl_blk->ops with the rest of
    the single-implementation dispatch layers; mtk_ctrl_blk now holds
    only mdev and trans
- Patch 3 (Add control DMA interface):
  - Deleted struct cldma_drv_ops, drv_info->drv_ops,
    cldma_drv_info_tbl[] and mtk_cldma_drv_m9xx.c/.h; the register
    table stays as data and is assigned directly in dev_init()
  - Deleted pcie/mtk_ctrl_cfg_m9xx.c: the queue, service and port
    tables are file-scope statics in pcie/mtk_trans_ctrl.c and the
    hw_ver lookup is gone
  - REG_CLDMA_INT_WF_MASK is no longer introduced here; it had no
    reader and Patch 5 used to remove it again
  - Removed CLDMA_GPD_FLAG_BPS, enum mtk_ip_busy_src, DIR_MAX, unused
    mtk_intr_type members, six reg_* fields and the *_done_cnt counters
  - mtk_cldma_txq_alloc()/rxq_alloc() take the queue_info the caller
    already resolved, dropping a repeated lookup that cannot fail
- Patch 4 (Add control port):
  - Deleted the stale-list and device-id-IDA machinery instead of
    hardening it: nothing here can reach the stale list and dev_id was
    only handed back to ida_free(), so both reported call sites are gone
  - mtk_port_tbl_destroy() uses the file's own radix-tree accessor
    (rcu_dereference_raw()); radix_tree_deref_slot() splats under
    CONFIG_PROVE_RCU in this lock-free teardown
  - Added port_mngr->port_tbl_mtx around the radix insert/delete and
    the port_cnt update, and struct mtk_port gained a rcu_head so
    mtk_port_release() frees with kfree_rcu()
  - Removed PORT_F_ALLOW_DROP (set, never read) and PORT_F_FORCE_SEND
    (read, never set, so the term was provably zero)
  - mtk_port_internal_disable() logs a failed mtk_port_ch_disable()
    with dev_warn(); teardown still continues and PORT_S_ENABLE is
    still cleared, or the port could never re-enable
  - clear_bit(PORT_S_WR) in mtk_port_internal_disable() is followed by
    wake_up_all(): the blocking writer loops back on timeout and
    otherwise only noticed one trb timeout later
  - mtk_ctrl_init() takes a struct mtk_port_layer_cfg * directly, the
    struct mtk_ctrl_cfg wrapper is gone, and mtk_port.c calls
    mtk_pcie_hif_submit_skb() instead of dispatching through ops
  - Removed the Q_MTU_*/Q_FRAG_* sizes no queue uses, MTK_PEER_ID_SAP
    and _MD, and PORT_S_DFLT/PORT_S_RD/PORT_S_STOP
- Patch 5 (Add FSM thread):
  - mtk_pcie_hif_exit() stops the TRB service kthreads before
    mtk_cldma_exit() frees the rings they walk; the two wake_up()
    sites skip a NULL trb_srv so the reorder cannot oops instead
  - One owner per direction on error: err_work sets txq->is_stopping
    under ring_lock and mtk_cldma_start_xfer() rechecks it there, and
    the RX side hands the queue to rx_done_work through need_restart
  - mtk_cldma_rx_done_work() consumes need_restart before honouring
    -ENXIO, so a queue left with no start address after a reset is
    restarted instead of skipped
  - The RX length is clamped against req->mtu (host state) instead of
    the GPD's data_allow_len, and an oversized frame is dropped with
    -EPROTO
  - fsm->hif_err became a per-producer array and the recorder stores
    unconditionally, so a successful handshake retry clears its own
    slot instead of refusing FSM_STATE_READY forever
  - The invalid dev_stage path of mtk_fsm_early_bootup_handler() uses
    goto exit, re-arming the boot-flow channel that was masked on
    entry
  - Both CLDMA dma_pools are created with align 16; with align 4 the
    24-byte descriptors were only 8-byte aligned every other block
  - cldma_drv_info[] is published with smp_store_release() and read
    with smp_load_acquire(); the removal sites and the txq/rxq slots
    use WRITE_ONCE(), completing the pairing already used elsewhere
  - struct runtime_feature_entry is __packed: offset advances by a
    device-supplied length, so a later entry's data_len can land
    unaligned. The layout is unchanged
  - Removed MTK_UEVENT_MINIDUMP and MTK_UEVENT_LOWPOWER, declared with
    no emitter anywhere in the series
  - Commit message rewritten to cover the CLDMA scope: the ISR, the
    QUEUE_ERROR worker and its -EPIPE completion policy, the
    ring_leaked rule, the teardown changes and the FSM uevent payload
  - Kept the uevent rather than reporting state through the wwan
    framework: wwan exposes ports and netdevs, neither carries a
    driver state machine, and adding one is a core change
  - Declined splitting the CLDMA work into its own patch: the ISR is
    registered by dev_init(), which only the FSM listener added here
    calls, so either split order leaves a non-functional commit
  - mtk_fsm.c calls mtk_pci_* directly and its six EXPORT_SYMBOL_GPL
    are removed
  - Removed FEATURE_TYPE_MUST/OPTIONAL, FSM_PRIO_0, MTK_UEVENT_UNDEF
    and _MAX, DEV_CFG_NORMAL, QUERY_RTFT_ID_MAX and the TAG define
- Patch 6 (Add AT & MBIM WWAN ports):
  - Dropped .tx_poll: it poll_wait()ed on a waitqueue inside struct
    mtk_port, which wwan_remove_port() releases while the fd stays
    open, and wake_up_pollfree() is not available to modules
  - TX back-pressure now reaches poll() through wwan_port_txon() and
    wwan_port_txoff() on the core's waitqueue, whose lifetime the open
    file pins; this also removes the errno/queue-full poll confusion
  - mtk_port_common_write() builds every CCCI packet onto a local
    queue before submitting any, so an allocation or copy failure can
    no longer truncate a message whose first packets are on the wire
  - caps.frag_len is tx_mtu minus the CCCI header, so one core
    fragment is exactly one CCCI packet instead of a full packet plus
    a 16-byte runt
  - w_lock now covers PORT_S_OPEN on the wwan ops: a receive that saw
    the port open completes wwan_port_rx() before stop() returns and
    the core purges its rx queue
  - clear_bit(PORT_S_WR) in mtk_port_wwan_disable() is followed by
    wake_up_all(), and a failed mtk_port_ch_disable() is logged with
    dev_warn() as in Patch 4
  - Commit message corrected: init() only prepares the port object,
    wwan_create_port() is in enable(), which FSM_STATE_READY reaches
    through the new mtk_port_enable_by_type(PORT_TBL_MD) hook
  - w_port is still published after PORT_S_ENABLE/PORT_S_WR: it is
    wwan_create_port()'s return value and the node becomes openable
    inside that call, so a comment now records the -ENXIO window
  - The AT/MBIM port and queue rows are added directly to the
    file-scope tables in pcie/mtk_trans_ctrl.c, and mtk_port_io.c
    calls mtk_pcie_hif_cmd_func() directly
- Link to v8: https://patch.msgid.link/20260914-t9xx_driver_v1-v8-0-5206c2e6bea0@compal.com

Changes in v8:
- Patch 1 (Add PCIe core):
  - Commit message: reworded the MHCCIF bullet (the driver masks all
    channels at init and does not clear pending status) and added a
    paragraph on the ACPI reset performed on removal
  - Added irq_cb_lock around IRQ callback register/unregister;
    unregister now masks the vector and calls synchronize_irq() before
    clearing the slot, so no callback runs after it returns
  - mtk_pci_pldr() evaluates MRST._RST on the device's own ACPI node
    instead of PXP._OFF/_ON on the upstream bridge: the bridge power
    resource is shared and its 500 ms off/on induced a link down/up
    that a hotplug-capable port reports as a remove/add
  - Removed mtk_pci_fldr(), enum mtk_reset_type, mtk_pci_dev_reset()
    and mtk_pci_reset(): dead here, they belong with the devlink
    firmware-update series that will use them
  - MHCCIF interrupts are now acked before dispatch, not after;
    mtk_mhccif_init() masks every channel but deliberately does not
    ack, as the status is one-shot latched by the device
  - Removed the write-only bar[] array and MTK_PCI_BAR_NUM
  - Switched to pcim_iomap_region(), replacing the deprecated
    pcim_iomap_regions()/pcim_iomap_table() pair
  - Dropped the BIT(30) exemption from the all-ones MMIO check; a
    stray bit is now logged with dev_err_ratelimited() and cleared
  - probe() masks all MHCCIF channels unconditionally; removed the
    dead configuration plumbing that used to select the mask
  - Check pci_save_state() and return -ENOMEM instead of -EFAULT
  - Reordered remove(): dev_exit, mask all 32 channels, mhccif_exit,
    free_irq, clear bus master, PLDR reset last
  - Removed the unused enum mtk_atr_type and the .type field
- Patch 2 (Add control plane transaction layer):
  - Converted the remaining EXPORT_SYMBOL() to EXPORT_SYMBOL_GPL()
    (mtk_ctrl_init/exit here, mtk_port_trb_free in Patch 4,
    mtk_fsm_start/evt_submit/init/exit in Patch 5) and deleted the
    unused mtk_fsm_notifier_register/unregister exports
- Patch 3 (Add control DMA interface):
  - Clamp the RX length reported by the GPD against data_allow_len and
    drop the packet with -EPROTO instead of trusting the device
  - Publish txq[]/rxq[] with smp_store_release(); the readers added in
    Patch 5 pair with smp_load_acquire()
  - mtk_cldma_stop_queue() returns -ENODEV on an all-ones read,
    escalates to cldma_drv_reset() on timeout and only leaks the ring
    if the queue still refuses to stop
  - Added NULL checks to all six radix_tree_lookup() call sites
  - mtk_cldma_rxq_free() splits ownership three ways (host-owned,
    device-owned, in-flight) instead of unmapping everything
  - New mtk_cldma_hw_recovery(); mtk_cldma_start_xfer() rewritten to
    run entirely under ring_lock so a resume cannot race a submit
  - Populate the control-queue entries of the queue info table here
    instead of in Patch 4, keeping each commit self-consistent
  - srv_cfg is const int (*)[HW_QUE_NUM] instead of int **, dropping
    the type-punning cast
  - Use skb_queue_len_lockless() where the list lock is not held
  - Hold the skb_list lock across every list read; the one place it
    must be dropped (sleeping submit) is documented
  - Check mtk_cldma_trb_process() return and add
    mtk_cldma_txq_flush() to complete pending requests on error
  - Initialise submit_lock and trans->available in
    mtk_trans_ctrl_init(), with a NULL srv guard
  - mtk_pcie_hif_exit() order: cldma_exit, then trb_srv_exit, then
    remove the radix tree
  - Evaluate CHECK_TX_FULL under submit_lock
  - Removed dead code found while auditing the above
  - Reject a non-linear skb with -EINVAL instead of WARN_ON_ONCE()
  - Commit message: probe only registers the transport plane; the
    CLDMA instances and the TRB service thread are started by the
    later "Add FSM thread" patch
  - A stop timeout in mtk_cldma_txq_free()/rxq_free() resets the whole
    IP, so the new mtk_cldma_rearm_queues() re-arms every queue still
    published on that instance, not just the one being freed
  - The RX re-arm is handed to mtk_cldma_rx_done_work() through a
    need_restart flag; it owns free_idx, which the re-arm used to read
    without synchronisation
  - mtk_cldma_submit_tx() rejects a NULL dev, closing the window
    between mtk_cldma_exit() clearing trans->dev and the TRB service
    thread being stopped
- Patch 5 (Add FSM thread):
  - New mtk_fsm_hif_err_record(): a transport init failure is sticky,
    so the STARTUP handler refuses the BOOTUP->READY promotion instead
    of reaching READY on a dead transport; DEV_ADD clears it
  - A feature the device reports as NOT_EXIST/NOT_SUPPORT now fails
    the handshake with -EPROTO when the host marks it MUST_SUPPORT
  - Both ctrl-msg handlers consume the skb on every path and return 0;
    the caller no longer frees behind the handler's back
  - The duplicate-HS2 and invalid-id paths no longer touch rt_data, so
    an epilogue cannot free data owned by a queued STARTUP event
  - mtk_fsm_idle_evt_handler() returns an error on a failed submit and
    unmasks its handshake channels only on success;
    mtk_fsm_early_bootup_handler() latches last_dev_state only after
    the handler chain succeeded, so a failure is retried
  - Re-arm BOOT_FLOW_SYNC in mtk_fsm_early_bootup_handler(): the
    channel was masked on entry and never unmasked, so a device still
    booting when the driver binds never delivered its DEV_STAGE_IDLE
    notification and the FSM stayed at FSM_STATE_ON with no ports
  - The in-flight runtime-data skb is owned by the STARTUP event it
    was submitted with and freed when that event completes; rt_data is
    demoted to a duplicate-HS2 guard and rt_data_len removed
  - mtk_fsm_init() unwinds through explicit labels and no longer leaks
    the kthread on any error path
  - mtk_fsm_exit() calls mtk_fsm_ctrl_ch_stop(), releasing the control
    ports while the port table they reference is still alive
  - Added the 0x0900 entry to cldma_drv_info_tbl[]; the primary
    MediaTek PCI ID no longer fails CLDMA dev_init with -EIO
  - QUEUE_ERROR is now cleared, re-unmasked and handled by a new
    err_work that stops the errored queues and completes their pending
    requests with -EPIPE, instead of being logged against a masked
    interrupt that is never re-armed
  - mtk_cldma_dev_exit() masks the L1 interrupt and calls
    synchronize_irq() before unregistering the callback
  - mtk_cldma_dev_exit() quiesces the IP (interrupt output disabled,
    cldma_drv_reset()) before releasing descriptor memory, and leaks
    the DMA pools as well if a ring was already leaked
  - New mtk_fsm_stop(), called by mtk_pci_dev_exit() after a possibly
    failed DEV_RM and before mtk_trans_ctrl_exit(), so the FSM thread
    cannot run on a dismantled transport plane; mtk_pcie_hif_exit() is
    idempotent, mtk_cldma_exit() latches trans->dev, and the FSM
    kthread is held with get_task_struct() so a thread that died
    abnormally cannot be stopped through a freed task_struct
  - mtk_cldma_txq_free()/rxq_free() flush err_work as well: it may
    already hold the queue they are about to free
  - err_work no longer completes and unmaps the requests of a queue
    that failed to stop; the device may still be reading them
  - The rt_data duplicate-HS2 guard is accessed with READ_ONCE() and
    WRITE_ONCE(), the rx path and the FSM kthread being concurrent
- Patch 6 (Add AT & MBIM WWAN ports):
  - Commit message: note the copy path now uses skb_copy_bits()
  - Removed union user_buf and mtk_port_copy_data_from(); the
    kernel-buffer branch had no user
  - Removed mtk_port_common_write_frag_skb() and the dead
    scatter-gather gate it was reached through
  - Fixed a heap overread: the TX path copied from a paged skb as if
    it were linear; use skb_copy_bits()
  - mtk_port_send_data() takes explicit blocking/force_send arguments,
    no longer mutates the caller's flags, and is all-or-nothing
  - The AT and MBIM TX callbacks are thin wrappers over a shared
    mtk_port_wwan_tx()
  - A port DISABLE is inserted ahead of pending ENABLEs (Patch 3) and
    mtk_port_internal_enable() unwinds on failure (Patch 4), so a
    close cannot be starved by a queued enable
  - Declined the reordering suggested in review, with a comment
    explaining the ordering constraint; w_port is read inside w_lock
  - wwan_port_rx() is called only when PORT_S_OPEN is set; otherwise
    the packet is dropped with -ENXIO and dev_dbg_ratelimited()
- Link to v7: https://patch.msgid.link/20260828-t9xx_driver_v1-v7-0-bf8f6074d88a@compal.com

Changes in v7:
- Cover letter: renamed "Core logic" to "Device lifecycle management".
  The FSM is T9xx-specific (same model as the in-tree t7xx driver) and
  stays in the driver rather than the WWAN framework.
- Patch 1 (Add PCIe core):
  - Removed RGU wording from commit message, no RGU code in the series
  - Removed `select NET_DEVLINK` from Kconfig, no devlink consumer
  - Reordered ATR setup: program TRSL_ADDR/TRSL_PARAM before enabling
    ATR in SRC_ADDR_LSB, with a readback, matching t7xx
  - Adopted the t7xx IRQ model: require exactly MTK_IRQ_CNT_MAX MSI-X
    vectors with one handler per vector; removed the merged-IRQ mode
    and scoped dispatch to each vector's own status bits
  - Fixed IRQ callback publish ordering: WRITE_ONCE + smp_wmb on
    register, READ_ONCE + smp_rmb in the interrupt handler
  - Added hw_bits == 0 guard in mtk_pci_send_ext_evt() to reject
    unmapped event channels
  - MHCCIF callbacks stay under spin_lock_bh by design; added a comment
    documenting the must-not-sleep callback contract
  - Added a drain loop in mtk_mhccif_exit() to free any remaining
    callback nodes
  - Moved cancel_work_sync(&priv->mhccif_work) before mtk_pci_pldr()
    in remove, preventing worker MMIO during the power cycle
  - Squashed the MAINTAINERS entry into this patch (was Patch 7); the
    series is now 6 patches
- Patch 2 (Add control plane transaction layer):
  - Commit message and Kconfig help text now name both modules
    (mtk_t9xx and mtk_t9xx_pcie)
  - Reworded mtk_ctrl_exit() kernel-doc: the allocation itself is
    devres-managed and freed on driver detach
- Patch 3 (Add control DMA interface):
  - RX re-map failure now keeps the buffer in its slot instead of
    freeing it and advancing free_idx, preventing a permanent ring
    stall from a stale GPD DMA address
  - mtk_cldma_stop_queue() returns -ETIMEDOUT on poll timeout; alloc
    paths abort on failure, free paths warn and proceed
  - Added READ_ONCE() for the free_idx and tx_started reads in
    mtk_cldma_start_xfer()
  - mtk_cldma_get_tx_start_addr() reads ADDRL+ADDRH as u64, so a >4GB
    DMA address is not mistaken for "queue not started"
  - mtk_cldma_open() decrements usr_cnt on failure. Beyond the emailed
    reply: mtk_pcie_hif_init() re-zeroes usr_cnt and ENABLE/DISABLE no
    longer gate on TX budget, so a channel reopens after devlink reload
  - mtk_cldma_close() completes the TRB with -EPIPE when drv_info is
    already gone, instead of stalling the submitter until timeout
  - mtk_cldma_txbuf_set() returns -ENOMEM instead of -EAGAIN on
    dma_map failure (no infinite retry); dev_err_ratelimited
  - Added the 0x0900 (MTK reference) entry to mtk_ctrl_info_tbl[];
    both PCI IDs are now present in all dispatch tables
  - Documented the single-consumer assumption of
    mtk_ctrl_trb_handler() in a comment
  - Added a submit_lock mutex closing the submit vs. hif_exit race on
    trans->available; hif_cmd_func() now checks available
  - Kept the empty queue table: queue_info entries are populated by
    the next patch "Add control port"
  - Commit message notes TX/RX paths are wired up by later patches
  - hif_exit has no caller at this commit, but ops->init is not called
    either, so nothing is allocated to leak; patch 5 wires up both
  - The trb_handler default case completes the TRB with -EINVAL
    instead of silently leaking the SKB
  - Documented the non-BD skb_headlen() == skb->len invariant
  - trb_open_priv overlay size is safe by construction
  - Removed the unused CLDMA4 support here at the source, instead of
    introducing it and deleting it again in Patch 5
  - Fixed an skb leak in the mtk_cldma_rxq_free() BD path: the parent
    skb was never freed after detaching frag_list
  - Replaced the free_idx/wr_idx boundary check added in v5 with a
    per-queue ring_lock following drivers/net/wwan/t7xx: the check
    raced with submit_tx and could stall TX forever; the lock also
    closes the submit/reclaim race the check was originally added for
  - Removed the write-only queues_cnt and tx_req.data_vm_addr fields
- Patch 5 (Add FSM thread):
  - Commit message: five FSM states (not seven); now describes the
    HS1/HS2/HS3 handshake and CLDMA bring-up/teardown in detail
  - Check ops->init() return in FSM_STATE_ON; NULL guard for srv in
    trb_srv_exit(); check mtk_cldma_dev_init() return in the listener
  - Kept the HS2 feature parser as is: the protocol requires dense
    ascending order and the ft_id check catches violations
  - Reject data_len == 0 before calling the query_rtft handler
  - Drop duplicate HS2 messages when hs_info->rt_data is already set
  - Added a notifier_lock mutex protecting FSM notifier list walks
  - Reordered HS2 handling: parse HS2, send HS3, and only then switch
    state, so FSM_STATE_READY implies HS3 was delivered
  - Unmask the HS channels on handshake error so the modem can retry
  - Moved wake_up_process() under evtq_lock with a fsm_handler NULL
    check; exit sets GATECLOSED and clears fsm_handler under the lock
  - mtk_fsm_hs_info_init() propagates event-registration failures and
    unwinds on error
  - The CLDMA ISR logs, clears and re-unmasks QUEUE_ERROR bits
  - CLDMA dev_exit unregisters the IRQ callback before mask +
    synchronize_irq, preventing ISR re-arm after teardown
  - Reordered FSM_STATE_OFF teardown: stop the TRB threads before
    freeing CLDMA, with the per-hif teardown moved into
    mtk_cldma_exit(); fixes a silently skipped teardown that broke the
    next reset cycle and leaked the CLDMA resources
  - mtk_trans_ctrl_exit() falls back to hif_exit when the
    FSM_STATE_OFF teardown never ran (guarded by trans->available)
  - mtk_pci_dev_start() propagates evt_submit/fsm_start errors
  - Restored TASK_INTERRUPTIBLE for the FSM event kthread: the v2 flip
    to TASK_UNINTERRUPTIBLE was not requested by review, and an idle
    D-state kthread blocks suspend and trips the hung-task detector
- Patch 6 (Add AT & MBIM WWAN ports):
  - Check and log mtk_port_enable_by_type() failure in the FSM state
    handler
  - A short WWAN write now returns -EIO; consume_skb() on success
  - The wwan_enable failure path calls mtk_port_ch_disable() to undo
    usr_cnt and partial resources
  - Publish the wwan_port pointer only after IS_ERR validation, so the
    RX path never sees an ERR_PTR
  - Pass wwan_port_caps (frag_len = negotiated tx_mtu, CCCI header
    headroom) to wwan_create_port(); a tx_mtu == 0 guard prevents
    userspace write() from spinning forever on frag_len == 0
  - Set PORT_S_WR/PORT_S_ENABLE before wwan_create_port() and clear
    both bits on failure
- Link to v6: https://patch.msgid.link/20260811-t9xx_driver_v1-v6-0-2c969fad57c6@compal.com

Changes in v6:
- Patch 3 (Add control DMA interface):
  - Move extern declarations for mtk_cldma_regs_m9xx and
    cldma_drv_ops_m9xx from Patch 5 to Patch 3 where the
    symbols are first defined, fixing sparse W=1 C=1 warnings
  - Fix RX done recycle path: skip buffer recycle in BD mode
    (nr_bds > 0) to prevent BDP flag mismatch causing HW to
    misinterpret data buffer address as BD chain pointer
    [sashiko]
  - Fix RX done recycle path: on dma_map_single failure,
    free skb and advance free_idx instead of goto out, to
    prevent ring stall [sashiko]
- Patch 4 (Add control port):
  - Replace wake_up_interruptible_all() with wake_up_all()
    for trb_wq and rx_wq in open/close/tx completion
    callbacks and common_close, fixing mismatch with
    wait_event_timeout (TASK_UNINTERRUPTIBLE) waiters in
    ch_enable/ch_disable that caused spurious timeouts
    [sashiko]
- Link to v5: https://patch.msgid.link/20260723-t9xx_driver_v1-v5-0-b27afb99ccbb@compal.com

Changes in v5:
- Patch 1 (Add PCIe core):
  - Remove LE32_TO_U32(cpu_to_le32()) no-op endianness
    roundtrip, return hw_bits directly
  - Use "%s" format string in pci_request_irq() to prevent
    format string injection
  - Return PCI_ERS_RESULT_DISCONNECT from AER error_detected,
    driver does not support AER recovery
- Patch 2 (Add control plane transaction layer):
  - Rewrite commit message to accurately describe skeleton
    content
- Patch 3 (Add control DMA interface):
  - Replace rmb() with dma_rmb() after HWO check in TX/RX
    done paths for correct DMA descriptor ordering
  - Add ring boundary check in TX done to prevent crossing
    into submit territory
  - Fix RX stall on reload failure: recycle old buffer with
    skb_trim() + dma_map_single() instead of losing descriptor
  - Convert devm_kcalloc to kcalloc for req_pool and
    bd_dsc_pool, matching existing kfree in teardown
  - Set bd_dsc->skb = NULL after dev_kfree_skb_any() in
    reload error path to prevent use-after-free
  - Use skb_headlen(skb) instead of skb->len for linear
    segment size in txbuf_set non-BD path
  - Use WRITE_ONCE() for wr_idx update in submit_tx
- Patch 4 (Add control port):
  - Fix direct mtk_port_release() calls to use kref_put()
    in mtk_port_free_or_backup and stale_list_grp_cleanup
  - Add ida_free() before kfree(s_list) in stale list
    cleanup to prevent IDA leak
  - Remove double-free of skb in mtk_port_internal_recv()
    drop_data path, caller owns the skb
  - Add port_cnt bounds check against data_len in
    mtk_port_status_update() to prevent OOB read
  - Replace wait_event_interruptible_timeout with
    wait_event_timeout in ch_enable/ch_disable to prevent
    busy-spin on pending signals
- Patch 5 (Add FSM thread):
  - Validate rtft data_len before passing to query_rtft
    callback to prevent OOB access
  - Zero-initialize HS3 skb data with memset() to prevent
    heap infoleak
  - Set hs_info->rt_data = NULL after dev_kfree_skb() in
    SAP/MD ctrl msg handlers to prevent use-after-free
  - Move wake_up_process() inside spinlock in evt_submit,
    clear fsm_handler under lock in exit to fix race
  - Replace kcalloc + radix_tree_gang_lookup with
    radix_tree_for_each_slot() in mtk_port_disable()
  - Add rt_data cleanup loop in FSM teardown path
- Patch 7 (Add maintainers entry):
  - Add Robert Yu and Jeff Chang as maintainers
- Link to v4: https://patch.msgid.link/20260709-t9xx_driver_v1-v4-0-a8c009d509c5@compal.com

Changes in v4:
- Patch 1 (Add PCIe core):
  - Add ACPI dependency to Kconfig (depends on PCI && ACPI)
    and remove #ifdef CONFIG_ACPI / #else guards from
    mtk_pci_fldr() and mtk_pci_pldr()
  - Remove unnecessary parentheses in (mdev)->dev across
    all dev_err/dev_warn calls
  - Remove devm_kfree() calls from mtk_pci_probe() error
    path and mtk_pci_remove(), devres handles cleanup
  - Remove mtk_dev_alloc()/mtk_dev_free() wrappers, use
    devm_kzalloc() directly in mtk_pci_probe()
  - Convert ext_evt callbacks from devm_kzalloc/devm_kfree
    to kzalloc/kfree for runtime-managed resources
- Patch 2 (Add control plane transaction layer):
  - Remove mtk_dev_alloc()/mtk_dev_free() declarations and
    implementations from mtk_dev.h and mtk_dev.c
  - Remove empty module_init/module_exit stubs, defer to
    Patch 4 where they have actual content
  - Remove devm_kfree() and unused ctrl_blk variable in
    mtk_ctrl_exit(), devres handles cleanup
- Patch 3 (Add control DMA interface):
  - Remove inline keyword from mtk_cldma_clr_bd_dsc() in
    .c file, let compiler decide inlining
  - Replace dev_warn() with dev_warn_ratelimited() for SKB
    alloc/map failure messages in data path
  - Replace open-coded do/while polling loop with
    read_poll_timeout() from <linux/iopoll.h> in
    mtk_cldma_stop_queue()
  - Remove ctrl_port_chl_mtu module_param, MODULE_PARM_DESC,
    and mtk_ctrl_queue_info_update() with associated macros
  - Remove unnecessary parentheses in (mdev)->dev across
    mtk_cldma.c and mtk_trans_ctrl.c
  - Convert CLDMA txq/rxq/cd from devm_kzalloc/devm_kfree
    to kzalloc/kfree for runtime-managed resources
  - Convert trb_srv/srv_que from devm_kzalloc/devm_kfree
    to kzalloc/kfree for runtime-managed resources
  - Remove devm_kfree() and unused variables from
    mtk_cldma_exit() and mtk_trans_ctrl_exit(), devres
    handles cleanup for device-lifetime resources
  - Remove unnecessary NULL checks before kfree() for
    bd_dsc_pool in txq_free/rxq_free
- Patch 4 (Add control port):
  - Add module_init/module_exit with mtk_port_io_init/exit
    content, moved from Patch 2
  - Remove unnecessary parentheses in (mdev)->dev in
    mtk_port.c
  - Remove devm_kfree() from mtk_port_mngr_init() error
    path and mtk_port_mngr_exit(), devres handles cleanup
  - Simplify mtk_trans_ctrl_init() error path: replace goto
    err_free_trans with direct return -ENOMEM
- Patch 5 (Add FSM thread):
  - Remove unnecessary parentheses in (mdev)->dev across
    mtk_fsm.c and mtk_ctrl_plane.c
  - Convert FSM notifiers from devm_kzalloc/devm_kfree to
    kzalloc/kfree for runtime-managed resources
  - Convert CLDMA drv_info from devm_kzalloc/devm_kfree to
    kzalloc/kfree for runtime-managed resources
  - Remove devm_kfree() from mtk_fsm_init()/mtk_fsm_exit()
    and mtk_ctrl_init()/mtk_ctrl_exit(), devres handles
    cleanup for device-lifetime resources
- Link to v3: https://patch.msgid.link/20260624-t9xx_driver_v1-v3-0-73ff03f60c48@compal.com

Changes in v3:
- Address sashiko AI code review comments and fix sparse warnings
- Patch 1 (Add PCIe core):
  - Move extern declaration of mtk_dev_cfg_0900 from mtk_pci.c to mtk_pci.h to fix sparse warning
  - Remove mtk_pci_bar_exit(): pcim_iounmap_region() was called with bitmask instead of BAR index, and pcim_iomap_regions() is devres-managed so manual unmap is redundant [sashiko]
  - Remove pci_disable_device() from probe error path and remove path: pcim_enable_device() registers devres cleanup, manual disable causes enable_cnt underflow [sashiko]
- Patch 3 (Add control DMA interface):
  - Add #include "mtk_cldma.h" in mtk_cldma_drv_m9xx.c to fix sparse undeclared symbol warnings for mtk_cldma_regs_m9xx and cldma_drv_ops_m9xx
  - Move extern declaration of mtk_ctrl_info_m9xx from mtk_trans_ctrl.c to mtk_trans_ctrl.h to fix sparse undeclared symbol warning
  - Replace kcalloc + radix_tree_gang_lookup with radix_tree_for_each_slot() in mtk_ctrl_remove_radix_tree() to eliminate allocation in teardown path [sashiko]
  - Add missing mtk_pci_dev_exit() call in mtk_pci_remove() to properly clean up FSM and trans_ctrl resources before device removal [sashiko]
- Patch 4 (Add control port):
  - Replace kcalloc + radix_tree_gang_lookup with radix_tree_for_each_slot() and single-entry gang_lookup in mtk_port_tbl_destroy() to eliminate allocation failure in teardown path [sashiko]
  - Add port = NULL after mtk_port_put_locked() in mtk_port_internal_open() error path to prevent returning un-refcounted pointer [sashiko]
  - Restore skb_unlink + trb_complete for non-EAGAIN errors in mtk_ctrl_trb_handler() TX path to prevent infinite retry loop [sashiko]
- Patch 5 (Add FSM thread):
  - Check mtk_pci_register_irq() return value in mtk_cldma_dev_init() and add err_destroy_wq error label to fix unreachable dead code [sashiko]
  - Propagate specific error codes (ENOMEM/EIO/EINVAL) from mtk_cldma_dev_init() error paths instead of generic -EIO [sashiko]
  - Free head SKB after detaching frag_list in mtk_cldma_rxq_free() scatter-gather mode to fix memory leak [sashiko]
- Patch 6 (Add AT & MBIM WWAN ports):
  - Remove unused write_lock mutex from struct mtk_port and its mutex_init call [sashiko]
- Link to v2: https://patch.msgid.link/20260610-t9xx_driver_v1-v2-0-c65addf23b3f@compal.com

Changes in v2:
- Split series into control plane (this v2) and data plane (follow-up)
- Patch 1 (Add PCIe core):
  - Rename BAR_NUM to MTK_PCI_BAR_NUM for driver prefix consistency
  - Replace magic numbers in mtk_pci_setup_atr() with named defines
  - Remove redundant ATR register comments, use blank line separators
  - Add kernel-doc comments to all non-static functions
  - Convert 4 MMIO wrapper functions to static inline in header [sashiko]
  - Remove unnecessary unlikely() from IRQ validation paths
  - Add irq_cnt == 0 and irq_id < 0 guards in mtk_pci_get_virq_id() [sashiko]
  - Initialize hw_bits at declaration for consistency
  - Merge same-type variable declarations into single lines
  - Add #else/#endif comments for CONFIG_ACPI blocks
  - Add newlines in mtk_pci_pldr() for readability
  - Move return into default case in mtk_pci_dev_reset()
  - Simplify mtk_mhccif_init() error path to use direct returns
  - Change -EFAULT to -ENOLINK for PCIe link check failure
  - Rename goto label "out" to "log_err" in mtk_pci_probe()
  - Wrap long lines to stay within 80 columns
  - Fix IRQ vector leak: add pci_free_irq_vectors() on error path [sashiko]
  - Fix mtk_pci_remove() ordering: free IRQ before cancel_work_sync [sashiko]
  - Fix mtk_pci_pldr() ACPI buffer leak: free first result before second call [sashiko]
  - Replace msleep(500) with MTK_PLDR_POWER_OFF_DELAY_MS define
  - Remove unused EXT_EVT_H2D_DRM_DISABLE_AP and related register define [sashiko]
  - Increase MTK_IRQ_NAME_LEN from 20 to 32 to fix W=1 format-truncation warning [sashiko]
- Patch 2 (Add control plane transaction layer):
  - Add kernel-doc comments to mtk_ctrl_init() and mtk_ctrl_exit()
  - Change mtk_ctrl_exit() return type from int to void
  - Set mdev->ctrl_blk to NULL after freeing in mtk_ctrl_exit() [sashiko]
  - Change ctrl_blk from void* to typed struct mtk_ctrl_blk* [sashiko]
  - Remove redundant "depends on MTK_T9XX" from MTK_T9XX_PCI Kconfig [sashiko]
  - Use mtk_dev_free() instead of devm_kfree() in mtk_pci_probe() error path [sashiko]
- Patch 3 (Add control DMA interface):
  - Add @ops kernel-doc parameter for mtk_ctrl_init()
  - Rename 'err' to 'ret' consistently throughout the patch
  - Reorder variable declarations to follow reverse Christmas tree style
  - Change mtk_cldma_txq_free() return type from int to void
  - Change mtk_cldma_rxq_free() return type from int to void
  - Change mtk_cldma_exit() return type from int to void
  - Remove unnecessary zero-initialization of ret in mtk_cldma_start_xfer()
  - Remove unnecessary zero-initialization of ret in mtk_cldma_tx()
  - Use direct return instead of goto out in mtk_cldma_submit_tx() error paths
  - Move software state before HWO flag in mtk_cldma_submit_tx()
  - Squash variable declarations in mtk_cldma_check_intr_status()
  - Remove unlikely() from validation paths in mtk_cldma_check_ch_cfg()
  - Clamp data_recv_len with min_t to prevent skb_over_panic in mtk_cldma_rx_skb_adjust() [sashiko]
  - Use READ_ONCE() for HWO flag polling in mtk_cldma_check_rx_req() [sashiko]
  - Fix mtk_cldma_rx_done_work() to always unmask interrupt on error path [sashiko]
  - Add DMA address guard in mtk_cldma_txq_free() teardown loop [sashiko]
  - Add IS_ERR() check for kthread_run() in mtk_ctrl_trb_srv_init() [sashiko]
  - Fix queue_info memory leak on validation failure in mtk_pcie_hif_init() [sashiko]
  - Handle non-EAGAIN errors in mtk_ctrl_trb_handler() TX path [sashiko]
  - Fix 'err' typo to 'ret' in mtk_cldma_txbuf_set() error message
  - Remove unused variable mdev in mtk_cldma_rx_check_again() [sashiko]
  - Remove unused variables trans and ctrl_blk in mtk_cldma_txq_free() and mtk_cldma_rxq_free() [sashiko]
- Patch 4 (Add control port):
  - Add @cfg kernel-doc parameter for mtk_ctrl_init()
  - Update mtk_ctrl_init() return description to cover additional error codes
  - Fix double list_del in mtk_port_stale_list_grp_cleanup() [sashiko]
  - Fix direct mtk_port_trb_free() call to use kref_put() in mtk_port_ch_enable() error path [sashiko]
  - Fix direct mtk_port_trb_free() call to use kref_put() in mtk_port_ch_disable() error path [sashiko]
  - Add mtk_port_tbl_destroy() in mtk_port_mngr_init() error path to prevent port memory leak [sashiko]
  - Change port_ops exit/reset/enable/disable callbacks from int to void
  - Move -EIO dispatch comment to where the code was introduced
- Patch 5 (Add FSM thread):
  - Add bounds check for rtft_entry in mtk_fsm_parse_hs2_msg() [sashiko]
  - Add skb length validation before accessing ctrl_msg_header in mtk_fsm_sap_ctrl_msg_handler() [sashiko]
  - Fix skb leak on CTRL_MSG_HS2 mismatch return in mtk_fsm_sap_ctrl_msg_handler() [sashiko]
  - Add skb length validation before accessing ctrl_msg_header in mtk_fsm_md_ctrl_msg_handler() [sashiko]
  - Replace devm_kzalloc/devm_kfree with kzalloc/kfree for FSM events [sashiko]
  - Fix mtk_fsm_evt_submit() to return -ETIMEDOUT on blocking event timeout [sashiko]
  - Change FSM kthread from TASK_INTERRUPTIBLE to TASK_UNINTERRUPTIBLE [sashiko]
  - Remove unused variable hw_id in mtk_cldma_dev_exit() [sashiko]
- Patch 6 (Add AT & MBIM WWAN ports):
  - Use imperative mode in commit message
  - Remove unnecessary zero-initialization of ret in mtk_port_copy_data_from()
  - Change copy_from_user() error code from -EFAULT to -EINVAL in mtk_port_copy_data_from()
  - Return -EINVAL for zero-length write in mtk_port_common_write()
  - Change mtk_port_wwan_exit/enable/disable() return type from int to void
  - Fix packet_size to account for CCCI header reservation in mtk_port_common_write() [sashiko]
  - Fix WWAN tx callbacks to consume skb and return 0 per wwan_port_ops contract [sashiko]
  - Fix wwan_create_port() error path: clear ERR_PTR to NULL and call mtk_port_ch_disable() [sashiko]
- Patch 7 (Add maintainers entry): new patch
- Link to v1: https://patch.msgid.link/20260529-t9xx_driver_v1-v1-0-bdbfe2c01e57@compal.com

To: Loic Poulain <loic.poulain@oss.qualcomm.com>
To: Sergey Ryazanov <ryazanov.s.a@gmail.com>
To: Johannes Berg <johannes@sipsolutions.net>
To: Andrew Lunn <andrew+netdev@lunn.ch>
To: "David S. Miller" <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Jack Wu <jackbb_wu@compal.com>
To: Robert Yu <robert_yu@compal.com>
To: Jeff Chang <Jeff_Chang@compal.com>
To: Wen-Zhi Huang <wen-zhi.huang@mediatek.com>
To: Shi-Wei Yeh <shi-wei.yeh@mediatek.com>
To: Minano Tseng <Minano.tseng@mediatek.com>
To: Matthias Brugger <matthias.bgg@gmail.com>
To: AngeloGioacchino Del Regno <angelogioacchino.delregno@collabora.com>
Cc: linux-kernel@vger.kernel.org
Cc: netdev@vger.kernel.org
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-mediatek@lists.infradead.org

---
Jack Wu (6):
      net: wwan: t9xx: Add PCIe core
      net: wwan: t9xx: Add control plane transaction layer
      net: wwan: t9xx: Add control DMA interface
      net: wwan: t9xx: Add control port
      net: wwan: t9xx: Add FSM thread
      net: wwan: t9xx: Add AT & MBIM WWAN ports

 MAINTAINERS                                 |   11 +
 drivers/net/wwan/Kconfig                    |   11 +
 drivers/net/wwan/Makefile                   |    1 +
 drivers/net/wwan/t9xx/Makefile              |   16 +
 drivers/net/wwan/t9xx/mtk_ctrl_plane.c      |  109 ++
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      |   79 ++
 drivers/net/wwan/t9xx/mtk_dev.h             |   46 +
 drivers/net/wwan/t9xx/mtk_fsm.c             | 1174 +++++++++++++++++
 drivers/net/wwan/t9xx/mtk_fsm.h             |  158 +++
 drivers/net/wwan/t9xx/mtk_port.c            |  793 +++++++++++
 drivers/net/wwan/t9xx/mtk_port.h            |  142 ++
 drivers/net/wwan/t9xx/mtk_port_io.c         |  613 +++++++++
 drivers/net/wwan/t9xx/mtk_port_io.h         |   36 +
 drivers/net/wwan/t9xx/mtk_utility.h         |   29 +
 drivers/net/wwan/t9xx/pcie/mtk_cldma.c      | 1893 +++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_cldma.h      |  169 +++
 drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c  |  561 ++++++++
 drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h  |  142 ++
 drivers/net/wwan/t9xx/pcie/mtk_pci.c        | 1122 ++++++++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_pci.h        |  143 ++
 drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h    |   35 +
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c |  661 ++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h |   93 ++
 23 files changed, 8037 insertions(+)
---
base-commit: eb3f4b7426cfd2b79d65b7d37155480b32259a11
change-id: 20260529-t9xx_driver_v1-1744f8af7739

Best regards,
--  
Jack Wu <jackbb_wu@compal.com>



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  2026-09-30  7:46 ` [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer Jack Wu via B4 Relay
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

Registers the T900 device driver with the kernel. Set up all
the fundamental configurations for the device: PCIe layer,
Modem Host Cross Core Interface (MHCCIF), modem common control
operations, build infrastructure and MAINTAINERS entry.

* PCIe layer code implements driver probe and removal, MSI-X
  interrupt initialization and de-initialization, and the way
  of resetting the device.
* MHCCIF provides the interrupt channels used for the boot
  handshake and for the host-to-device reset doorbell.
* Modem common control operations provide the basic read/write
  functions of the device's hardware registers,
  mask/unmask/get/clear functions of the device's interrupt
  registers and inquiry functions of the device's status.
* Add MAINTAINERS entry for the MediaTek T9XX 5G WWAN modem
  device driver.

Removal resets the modem.  There is no software reset that
returns the device from running firmware to a state the next
probe can boot, and the next driver load has to redo the
device handshake and status sync, so mtk_pci_remove()
evaluates MRST._RST on the device's own ACPI node.

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 MAINTAINERS                              |   11 +
 drivers/net/wwan/Kconfig                 |   11 +
 drivers/net/wwan/Makefile                |    1 +
 drivers/net/wwan/t9xx/Makefile           |    9 +
 drivers/net/wwan/t9xx/mtk_dev.h          |   42 ++
 drivers/net/wwan/t9xx/pcie/mtk_pci.c     | 1033 ++++++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_pci.h     |  143 +++++
 drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h |   34 +
 8 files changed, 1284 insertions(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index 461a3eed6129..8f13ef7440ce 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -16494,6 +16494,17 @@ L:	netdev@vger.kernel.org
 S:	Supported
 F:	drivers/net/wwan/t7xx/
 
+MEDIATEK T9XX 5G WWAN MODEM DRIVER
+M:	Jack Wu <jackbb_wu@compal.com>
+M:	Robert Yu <robert_yu@compal.com>
+M:	Jeff Chang <Jeff_Chang@compal.com>
+R:	Wen-Zhi Huang <wen-zhi.huang@mediatek.com>
+R:	Shi-Wei Yeh <shi-wei.yeh@mediatek.com>
+R:	Minano Tseng <Minano.tseng@mediatek.com>
+L:	netdev@vger.kernel.org
+S:	Supported
+F:	drivers/net/wwan/t9xx/
+
 MEDIATEK USB3 DRD IP DRIVER
 M:	Chunfeng Yun <chunfeng.yun@mediatek.com>
 L:	linux-usb@vger.kernel.org
diff --git a/drivers/net/wwan/Kconfig b/drivers/net/wwan/Kconfig
index 88df55d78d90..55c45af410ee 100644
--- a/drivers/net/wwan/Kconfig
+++ b/drivers/net/wwan/Kconfig
@@ -121,6 +121,17 @@ config MTK_T7XX
 
 	  If unsure, say N.
 
+config MTK_T9XX
+	tristate "MediaTek PCIe 5G WWAN modem T9xx device"
+	depends on PCI && ACPI
+	help
+	  Enables MediaTek PCIe based 5G WWAN modem (T9xx series) device.
+
+	  To compile this driver as a module, choose M here: the module will be
+	  called mtk_t9xx.
+
+	  If unsure, say N.
+
 endif # WWAN
 
 endmenu
diff --git a/drivers/net/wwan/Makefile b/drivers/net/wwan/Makefile
index 3960c0ae2445..7361eef4c472 100644
--- a/drivers/net/wwan/Makefile
+++ b/drivers/net/wwan/Makefile
@@ -14,3 +14,4 @@ obj-$(CONFIG_QCOM_BAM_DMUX) += qcom_bam_dmux.o
 obj-$(CONFIG_RPMSG_WWAN_CTRL) += rpmsg_wwan_ctrl.o
 obj-$(CONFIG_IOSM) += iosm/
 obj-$(CONFIG_MTK_T7XX) += t7xx/
+obj-$(CONFIG_MTK_T9XX) += t9xx/
diff --git a/drivers/net/wwan/t9xx/Makefile b/drivers/net/wwan/t9xx/Makefile
new file mode 100644
index 000000000000..bd6063c85c2b
--- /dev/null
+++ b/drivers/net/wwan/t9xx/Makefile
@@ -0,0 +1,9 @@
+# SPDX-License-Identifier: GPL-2.0-only
+
+ccflags-y += -I$(src)/pcie
+ccflags-y += -I$(src)
+
+obj-$(CONFIG_MTK_T9XX) += mtk_t9xx.o
+
+mtk_t9xx-y := \
+	pcie/mtk_pci.o
diff --git a/drivers/net/wwan/t9xx/mtk_dev.h b/drivers/net/wwan/t9xx/mtk_dev.h
new file mode 100644
index 000000000000..329e5e6695df
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_dev.h
@@ -0,0 +1,42 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_DEV_H__
+#define __MTK_DEV_H__
+
+#include <linux/dma-mapping.h>
+#include <linux/dmapool.h>
+#include <linux/kernel.h>
+#include <linux/module.h>
+#include <linux/sched.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+
+#define MTK_DEV_STR_LEN 16
+
+enum mtk_user_id {
+	MTK_USER_CTRL,
+	MTK_USER_MAX
+};
+
+enum mtk_dev_evt_h2d {
+	DEV_EVT_H2D_DEVICE_RESET	= BIT(2),
+};
+
+enum mtk_dev_evt_d2h {
+	DEV_EVT_D2H_BOOT_FLOW_SYNC	= BIT(4),
+	DEV_EVT_D2H_ASYNC_HS_NOTIFY_SAP = BIT(5),
+	DEV_EVT_D2H_ASYNC_HS_NOTIFY_MD	= BIT(6),
+};
+
+/* mtk_md_dev defines the structure of MTK modem device */
+struct mtk_md_dev {
+	struct device *dev;
+	void *hw_priv;
+	u32 hw_ver;
+	char dev_str[MTK_DEV_STR_LEN];
+};
+
+#endif /* __MTK_DEV_H__ */
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
new file mode 100644
index 000000000000..34ee823119fc
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
@@ -0,0 +1,1033 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#include <linux/acpi.h>
+#include <linux/aer.h>
+#include <linux/bitfield.h>
+#include <linux/device.h>
+#include <linux/dma-mapping.h>
+#include <linux/kernel.h>
+#include <linux/module.h>
+
+#include "mtk_dev.h"
+#include "mtk_pci.h"
+#include "mtk_pci_reg.h"
+
+#define MTK_MHCCIF_RC_BASE_ADDR		0x1000A000
+
+static const int mtk_pci_irq_tbl[MTK_IRQ_SRC_MAX] = {
+	[MTK_IRQ_SRC_MHCCIF] = 28,
+	[MTK_IRQ_SRC_CLDMA0] = 27,
+	[MTK_IRQ_SRC_CLDMA1] = 26,
+};
+
+#define MTK_PCI_TRANSPARENT_ATR_SIZE	(0x3F)
+#define MTK_PCI_MINIMUM_ATR_SIZE	(0x1000)
+#define ATR_SIZE_LO32_MASK		GENMASK_ULL(31, 0)
+#define ATR_SIZE_HI32_MASK		GENMASK_ULL(63, 32)
+#define ATR_SIZE_BIAS_FROM_LO32		2
+#define ATR_ADDR_ALIGN_MASK		0xFFFFF000
+#define ATR_EN				BIT(0)
+#define ATR_PARAM_OFFSET		16
+#define SET_HW_BITS(dest, chs, mhccif, dev)		\
+	({						\
+		if ((chs) & (dev))					\
+			(dest) |= FIELD_PREP(mhccif, 1);		\
+	})
+
+struct mtk_mhccif_cb {
+	struct list_head entry;
+	int (*evt_cb)(u32 status, void *data);
+	void *data;
+	u32 chs;
+};
+
+/**
+ * mtk_pci_setup_atr() - Configure a PCIe address translation rule
+ * @mdev: MTK MD device
+ * @cfg: ATR configuration parameters
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_setup_atr(struct mtk_md_dev *mdev, struct mtk_atr_cfg *cfg)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	u32 addr, val, size_h, size_l;
+	int atr_size, pos, offset;
+
+	if (cfg->transparent) {
+		/* No address conversion is performed */
+		atr_size = MTK_PCI_TRANSPARENT_ATR_SIZE;
+	} else {
+		if (cfg->size < MTK_PCI_MINIMUM_ATR_SIZE)
+			cfg->size = MTK_PCI_MINIMUM_ATR_SIZE;
+
+		if (cfg->src_addr & (cfg->size - 1)) {
+			dev_err(mdev->dev, "Invalid atr src addr is not aligned to size\n");
+			return -EFAULT;
+		}
+
+		if (cfg->trsl_addr & (cfg->size - 1)) {
+			dev_err(mdev->dev,
+				"Invalid atr trsl addr is not aligned to size, %llx, %llx\n",
+				cfg->trsl_addr, cfg->size - 1);
+			return -EFAULT;
+		}
+
+		size_l = FIELD_GET(ATR_SIZE_LO32_MASK, cfg->size);
+		size_h = FIELD_GET(ATR_SIZE_HI32_MASK, cfg->size);
+		pos = ffs(size_l);
+		if (pos) {
+			atr_size = pos - ATR_SIZE_BIAS_FROM_LO32;
+		} else {
+			pos = ffs(size_h);
+			atr_size = pos + 32 - ATR_SIZE_BIAS_FROM_LO32;
+		}
+	}
+
+	/* Calculate table offset */
+	offset = ATR_PORT_OFFSET * cfg->port + ATR_TABLE_OFFSET * cfg->table;
+
+	addr = REG_ATR_PCIE_WIN0_T0_SRC_ADDR_MSB + offset;
+	val = (u32)(cfg->src_addr >> 32);
+	mtk_pci_mac_write32(priv, addr, val);
+
+	addr = REG_ATR_PCIE_WIN0_T0_TRSL_ADDR_MSB + offset;
+	val = (u32)(cfg->trsl_addr >> 32);
+	mtk_pci_mac_write32(priv, addr, val);
+
+	addr = REG_ATR_PCIE_WIN0_T0_TRSL_ADDR_LSB + offset;
+	val = (u32)(cfg->trsl_addr & ATR_ADDR_ALIGN_MASK);
+	mtk_pci_mac_write32(priv, addr, val);
+
+	/* TRSL_PARAM */
+	addr = REG_ATR_PCIE_WIN0_T0_TRSL_PARAM + offset;
+	val = (cfg->trsl_param << ATR_PARAM_OFFSET) | cfg->trsl_id;
+	mtk_pci_mac_write32(priv, addr, val);
+
+	/* Enable ATR last, after translation target is fully programmed */
+	addr = REG_ATR_PCIE_WIN0_T0_SRC_ADDR_LSB + offset;
+	val = (u32)(cfg->src_addr & ATR_ADDR_ALIGN_MASK) | (atr_size << 1) | ATR_EN;
+	mtk_pci_mac_write32(priv, addr, val);
+
+	/* Ensure ATR is set */
+	mtk_pci_mac_read32(priv, addr);
+
+	return 0;
+}
+
+/**
+ * mtk_pci_atr_disable() - Disable all PCIe address translation rules
+ * @priv: MTK PCI private data
+ */
+void mtk_pci_atr_disable(struct mtk_pci_priv *priv)
+{
+	int port, tbl, offset;
+	u32 val;
+
+	/* Disable all ATR table for all ports */
+	for (port = ATR_SRC_PCI_WIN0; port <= ATR_SRC_AXIS_3; port++)
+		for (tbl = 0; tbl < ATR_TABLE_NUM_PER_ATR; tbl++) {
+			/* Calculate table offset */
+			offset = ATR_PORT_OFFSET * port + ATR_TABLE_OFFSET * tbl;
+			val = mtk_pci_mac_read32(priv, REG_ATR_PCIE_WIN0_T0_SRC_ADDR_LSB + offset);
+			val = val & (~BIT(0));
+			/* Disable table by SRC_ADDR_L */
+			mtk_pci_mac_write32(priv, REG_ATR_PCIE_WIN0_T0_SRC_ADDR_LSB + offset, val);
+		}
+}
+
+/**
+ * mtk_pci_get_dev_state() - Read the device state from the modem
+ * @mdev: MTK MD device
+ *
+ * Return: Device state value.
+ */
+u32 mtk_pci_get_dev_state(struct mtk_md_dev *mdev)
+{
+	return mtk_pci_mac_read32(mdev->hw_priv, REG_PCIE_DEBUG_DUMMY_7);
+}
+
+/**
+ * mtk_pci_ack_dev_state() - Acknowledge the device state to the modem
+ * @mdev: MTK MD device
+ * @state: State value to acknowledge
+ */
+void mtk_pci_ack_dev_state(struct mtk_md_dev *mdev, u32 state)
+{
+	mtk_pci_mac_write32(mdev->hw_priv, REG_PCIE_DEBUG_DUMMY_7, state);
+}
+
+/**
+ * mtk_pci_get_irq_id() - Map an IRQ source to its hardware IRQ ID
+ * @mdev: MTK MD device
+ * @irq_src: IRQ source enum
+ *
+ * Return: IRQ ID on success, -EINVAL on failure.
+ */
+int mtk_pci_get_irq_id(struct mtk_md_dev *mdev, enum mtk_irq_src irq_src)
+{
+	int irq_id = -EINVAL;
+
+	if (irq_src > MTK_IRQ_SRC_MIN && irq_src < MTK_IRQ_SRC_MAX) {
+		irq_id = mtk_pci_irq_tbl[irq_src];
+		if (irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX)
+			irq_id = -EINVAL;
+	}
+
+	return irq_id;
+}
+
+/**
+ * mtk_pci_get_virq_id() - Get the Linux virtual IRQ for a hardware IRQ ID
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ *
+ * Return: Virtual IRQ number on success, negative error code on failure.
+ */
+int mtk_pci_get_virq_id(struct mtk_md_dev *mdev, int irq_id)
+{
+	struct pci_dev *pdev = to_pci_dev(mdev->dev);
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	if (irq_id < 0 || irq_id >= priv->irq_cnt)
+		return -EINVAL;
+
+	return pci_irq_vector(pdev, irq_id);
+}
+
+/**
+ * mtk_pci_register_irq() - Register a callback for a hardware IRQ
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ * @irq_cb: Callback function
+ * @data: Private data passed to callback
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_register_irq(struct mtk_md_dev *mdev, int irq_id,
+			 int (*irq_cb)(int irq_id, void *data), void *data)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	if ((irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX) || !irq_cb)
+		return -EINVAL;
+
+	spin_lock(&priv->irq_cb_lock);
+	if (priv->irq_cb_list[irq_id]) {
+		spin_unlock(&priv->irq_cb_lock);
+		dev_err(mdev->dev,
+			"Unable to register irq, irq_id=%d, it's already been register by %ps.\n",
+			irq_id, priv->irq_cb_list[irq_id]);
+		return -EFAULT;
+	}
+	priv->irq_cb_data[irq_id] = data;
+	smp_wmb(); /* Ensure data is visible before callback */
+	WRITE_ONCE(priv->irq_cb_list[irq_id], irq_cb);
+	spin_unlock(&priv->irq_cb_lock);
+
+	return 0;
+}
+
+/**
+ * mtk_pci_unregister_irq() - Unregister a hardware IRQ callback
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_unregister_irq(struct mtk_md_dev *mdev, int irq_id)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	int virq_id;
+
+	if (irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX)
+		return -EINVAL;
+
+	if (!READ_ONCE(priv->irq_cb_list[irq_id])) {
+		dev_err(mdev->dev, "irq_id=%d has not been registered\n", irq_id);
+		return -EFAULT;
+	}
+
+	/* Stop the source and wait for in-flight handlers
+	 * before the callback or its data can disappear.
+	 */
+	mtk_pci_mask_irq(mdev, irq_id);
+	virq_id = mtk_pci_get_virq_id(mdev, irq_id);
+	if (virq_id >= 0)
+		synchronize_irq(virq_id);
+
+	spin_lock(&priv->irq_cb_lock);
+	WRITE_ONCE(priv->irq_cb_list[irq_id], NULL);
+	priv->irq_cb_data[irq_id] = NULL;
+	spin_unlock(&priv->irq_cb_lock);
+
+	return 0;
+}
+
+/**
+ * mtk_pci_mask_irq() - Mask (disable) a hardware IRQ
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_mask_irq(struct mtk_md_dev *mdev, int irq_id)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	if (irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX) {
+		dev_err(mdev->dev, "Failed to mask irq: input irq_id=%d\n", irq_id);
+		return -EINVAL;
+	}
+
+	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, BIT(irq_id));
+
+	return 0;
+}
+
+/**
+ * mtk_pci_unmask_irq() - Unmask (enable) a hardware IRQ
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_unmask_irq(struct mtk_md_dev *mdev, int irq_id)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	if (irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX) {
+		dev_err(mdev->dev, "Failed to unmask irq: input irq_id=%d\n", irq_id);
+		return -EINVAL;
+	}
+
+	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_SET_GRP0_0, BIT(irq_id));
+
+	return 0;
+}
+
+/**
+ * mtk_pci_clear_irq() - Clear (acknowledge) a hardware IRQ
+ * @mdev: MTK MD device
+ * @irq_id: Hardware IRQ ID
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_clear_irq(struct mtk_md_dev *mdev, int irq_id)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	if (irq_id < 0 || irq_id >= MTK_IRQ_CNT_MAX) {
+		dev_err(mdev->dev, "Failed to clear irq: input irq_id=%d\n", irq_id);
+		return -EINVAL;
+	}
+
+	mtk_pci_mac_write32(priv, REG_MSIX_ISTATUS_HOST_GRP0_0, BIT(irq_id));
+
+	return 0;
+}
+
+static u32 mtk_pci_ext_d2h_evt_hw_bits(u32 chs)
+{
+	u32 hw_bits = 0;
+
+	SET_HW_BITS(hw_bits, chs, MHCCIF_EP2RC_EVT_BOOT_FLOW_SYNC,
+		    DEV_EVT_D2H_BOOT_FLOW_SYNC);
+	SET_HW_BITS(hw_bits, chs, MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_SAP,
+		    DEV_EVT_D2H_ASYNC_HS_NOTIFY_SAP);
+	SET_HW_BITS(hw_bits, chs, MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_MD,
+		    DEV_EVT_D2H_ASYNC_HS_NOTIFY_MD);
+
+	return hw_bits;
+}
+
+static u32 mtk_pci_ext_d2h_evt_chs(u32 hw_bits)
+{
+	u32 chs = 0;
+
+	if (!hw_bits)
+		return chs;
+
+	chs = FIELD_PREP(DEV_EVT_D2H_BOOT_FLOW_SYNC,
+			 FIELD_GET(MHCCIF_EP2RC_EVT_BOOT_FLOW_SYNC, hw_bits)) |
+	      FIELD_PREP(DEV_EVT_D2H_ASYNC_HS_NOTIFY_SAP,
+			 FIELD_GET(MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_SAP, hw_bits)) |
+	      FIELD_PREP(DEV_EVT_D2H_ASYNC_HS_NOTIFY_MD,
+			 FIELD_GET(MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_MD, hw_bits));
+
+	return chs;
+}
+
+/**
+ * mtk_pci_register_ext_evt() - Register a callback for MHCCIF device events
+ * @mdev: MTK MD device
+ * @chs: Bitmask of event channels to register
+ * @evt_cb: Callback function
+ * @data: Private data passed to callback
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_register_ext_evt(struct mtk_md_dev *mdev, u32 chs,
+			     int (*evt_cb)(u32 status, void *data), void *data)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct mtk_mhccif_cb *cb;
+	int ret = 0;
+
+	if (!chs || !evt_cb)
+		return -EINVAL;
+
+	spin_lock_bh(&priv->mhccif_lock);
+	list_for_each_entry(cb, &priv->mhccif_cb_list, entry) {
+		if (cb->chs & chs) {
+			ret = -EFAULT;
+			dev_err(mdev->dev,
+				"Unable to register evt, intersection: chs=0x%08x&0x%08x cb=%ps\n",
+				chs, cb->chs, cb->evt_cb);
+			goto err_spin_unlock;
+		}
+	}
+	cb = kzalloc(sizeof(*cb), GFP_ATOMIC);
+	if (!cb) {
+		ret = -ENOMEM;
+		goto err_spin_unlock;
+	}
+	cb->evt_cb = evt_cb;
+	cb->data = data;
+	cb->chs = chs;
+	list_add_tail(&cb->entry, &priv->mhccif_cb_list);
+err_spin_unlock:
+	spin_unlock_bh(&priv->mhccif_lock);
+
+	return ret;
+}
+
+/**
+ * mtk_pci_unregister_ext_evt() - Unregister an MHCCIF device event callback
+ * @mdev: MTK MD device
+ * @chs: Bitmask of event channels to unregister
+ */
+void mtk_pci_unregister_ext_evt(struct mtk_md_dev *mdev, u32 chs)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct mtk_mhccif_cb *cb, *next;
+
+	if (!chs)
+		return;
+
+	spin_lock_bh(&priv->mhccif_lock);
+	list_for_each_entry_safe(cb, next, &priv->mhccif_cb_list, entry) {
+		if (cb->chs == chs) {
+			list_del(&cb->entry);
+			kfree(cb);
+			goto out;
+		}
+	}
+	dev_warn(mdev->dev,
+		 "Unable to unregister evt, no chs=0x%08x has been registered.\n", chs);
+out:
+	spin_unlock_bh(&priv->mhccif_lock);
+}
+
+/**
+ * mtk_pci_mask_ext_evt() - Mask (disable) MHCCIF device events
+ * @mdev: MTK MD device
+ * @chs: Bitmask of event channels to mask
+ */
+void mtk_pci_mask_ext_evt(struct mtk_md_dev *mdev, u32 chs)
+{
+	u32 hw_bits = mtk_pci_ext_d2h_evt_hw_bits(chs);
+
+	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			MHCCIF_EP2RC_SW_INT_EAP_MASK_SET, hw_bits);
+}
+
+/**
+ * mtk_pci_unmask_ext_evt() - Unmask (enable) MHCCIF device events
+ * @mdev: MTK MD device
+ * @chs: Bitmask of event channels to unmask
+ */
+void mtk_pci_unmask_ext_evt(struct mtk_md_dev *mdev, u32 chs)
+{
+	u32 hw_bits = mtk_pci_ext_d2h_evt_hw_bits(chs);
+
+	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			MHCCIF_EP2RC_SW_INT_EAP_MASK_CLR, hw_bits);
+}
+
+/**
+ * mtk_pci_clear_ext_evt() - Clear (acknowledge) MHCCIF device events
+ * @mdev: MTK MD device
+ * @chs: Bitmask of event channels to clear
+ */
+void mtk_pci_clear_ext_evt(struct mtk_md_dev *mdev, u32 chs)
+{
+	u32 hw_bits = mtk_pci_ext_d2h_evt_hw_bits(chs);
+
+	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			MHCCIF_EP2RC_SW_INT_ACK, hw_bits);
+}
+
+static u32 mtk_pci_ext_h2d_evt_hw_bits(u32 chs)
+{
+	u32 hw_bits = 0;
+
+	SET_HW_BITS(hw_bits, chs, MHCCIF_RC2EP_EVT_DEVICE_RESET,
+		    DEV_EVT_H2D_DEVICE_RESET);
+	return hw_bits;
+}
+
+/**
+ * mtk_pci_send_ext_evt() - Send an MHCCIF event to the modem
+ * @mdev: MTK MD device
+ * @ch: Event channel to trigger (must be a single bit)
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int mtk_pci_send_ext_evt(struct mtk_md_dev *mdev, u32 ch)
+{
+	u32 rc_base = MTK_MHCCIF_RC_BASE_ADDR;
+	u32 hw_bits;
+
+	/* Only allow one ch to be triggered at a time */
+	if (!is_power_of_2(ch)) {
+		dev_err(mdev->dev, "Unsupported ext evt ch=0x%08x\n", ch);
+		return -EINVAL;
+	}
+
+	hw_bits = mtk_pci_ext_h2d_evt_hw_bits(ch);
+	if (!hw_bits) {
+		dev_err(mdev->dev, "Unmapped ext evt ch=0x%08x\n", ch);
+		return -EINVAL;
+	}
+
+	mtk_pci_write32(mdev, rc_base + MHCCIF_RC2EP_SW_BSY, hw_bits);
+	mtk_pci_write32(mdev, rc_base + MHCCIF_RC2EP_SW_TCHNUM, ffs(hw_bits) - 1);
+	return 0;
+}
+
+static u32 mtk_pci_get_ext_evt_hw_status(struct mtk_md_dev *mdev)
+{
+	return mtk_pci_read32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			      MHCCIF_EP2RC_SW_INT_STS);
+}
+
+static void mtk_pci_ack_ext_evt_hw(struct mtk_md_dev *mdev, u32 hw_bits)
+{
+	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			MHCCIF_EP2RC_SW_INT_ACK, hw_bits);
+	/* Ensure the ack lands before level 1 is cleared */
+	mtk_pci_read32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+		       MHCCIF_EP2RC_SW_INT_STS);
+}
+
+/**
+ * mtk_pci_pldr() - Reset the modem via its ACPI MRST._RST method
+ * @mdev: MTK MD device
+ *
+ * Return:
+ * * 0       - Success, device was reset.
+ * * -ENODEV - No ACPI, no handle or no MRST._RST method.  Device
+ *             untouched.
+ * * -EIO    - MRST._RST failed.  Device untouched, still running;
+ *             a different reset may be attempted.
+ */
+int mtk_pci_pldr(struct mtk_md_dev *mdev)
+{
+	acpi_status acpi_ret;
+	acpi_handle handle;
+
+	if (acpi_disabled) {
+		dev_err(mdev->dev, "Unsupported, acpi function isn't enable\n");
+		return -ENODEV;
+	}
+
+	handle = ACPI_HANDLE(mdev->dev);
+	if (!handle) {
+		dev_err(mdev->dev, "Unsupported, acpi handle isn't found\n");
+		return -ENODEV;
+	}
+	if (!acpi_has_method(handle, "MRST._RST")) {
+		dev_err(mdev->dev, "Unsupported, pldr method isn't supported\n");
+		return -ENODEV;
+	}
+	acpi_ret = acpi_evaluate_object(handle, "MRST._RST", NULL, NULL);
+	if (ACPI_FAILURE(acpi_ret)) {
+		dev_err(mdev->dev, "Failed to execute MRST._RST method: %s\n",
+			acpi_format_exception(acpi_ret));
+		return -EIO;
+	}
+
+	return 0;
+}
+
+/**
+ * mtk_pci_link_check() - Check if the PCIe link to the modem is active
+ * @mdev: MTK MD device
+ *
+ * Return: true if the device is present, false otherwise.
+ */
+bool mtk_pci_link_check(struct mtk_md_dev *mdev)
+{
+	return pci_device_is_present(to_pci_dev(mdev->dev));
+}
+
+static void mtk_mhccif_isr_work(struct work_struct *work)
+{
+	struct mtk_pci_priv *priv =
+		container_of(work, struct mtk_pci_priv, mhccif_work);
+	struct mtk_md_dev *mdev = priv->irq_desc->mdev;
+	struct mtk_mhccif_cb *cb;
+	u32 stat, mask, chs;
+
+	stat = mtk_pci_get_ext_evt_hw_status(mdev);
+	mask = mtk_pci_read32(mdev, MTK_MHCCIF_RC_BASE_ADDR
+		+ MHCCIF_EP2RC_SW_INT_EAP_MASK);
+	if (unlikely(stat == U32_MAX && !(mtk_pci_link_check(mdev)))) {
+		/* When link failed, we don't need to unmask/clear. */
+		dev_err(mdev->dev, "Failed to check link in MHCCIF handler.\n");
+		return;
+	}
+
+	stat &= ~mask;
+	/* Acknowledge level 2 before dispatch: an event that re-asserts
+	 * while a callback runs must survive as a new status bit.
+	 */
+	if (stat)
+		mtk_pci_ack_ext_evt_hw(mdev, stat);
+
+	chs = mtk_pci_ext_d2h_evt_chs(stat);
+	/* Callbacks must not sleep or modify mhccif_cb_list */
+	spin_lock_bh(&priv->mhccif_lock);
+	list_for_each_entry(cb, &priv->mhccif_cb_list, entry) {
+		if (cb->chs & chs)
+			cb->evt_cb(cb->chs & chs, cb->data);
+	}
+	spin_unlock_bh(&priv->mhccif_lock);
+
+	mtk_pci_clear_irq(mdev, priv->mhccif_irq_id);
+	mtk_pci_unmask_irq(mdev, priv->mhccif_irq_id);
+}
+
+static const struct pci_device_id t9xx_pci_table[] = {
+	{ PCI_DEVICE(MTK_PCI_VENDOR_ID, 0x0900), MTK_PCI_CLASS, PCI_ANY_ID },
+	{ PCI_DEVICE(CEI_PCI_VENDOR_ID, 0x01CA), MTK_PCI_CLASS, PCI_ANY_ID },
+	{/* end: all zeroes */}
+};
+
+MODULE_DEVICE_TABLE(pci, t9xx_pci_table);
+
+static int mtk_pci_atr_init(struct mtk_md_dev *mdev)
+{
+	struct pci_dev *pdev = to_pci_dev(mdev->dev);
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct mtk_atr_cfg cfg;
+	int port, ret;
+
+	mtk_pci_atr_disable(priv);
+
+	/* Config ATR for RC to access device's register */
+	cfg.src_addr = pci_resource_start(pdev, MTK_BAR_2_3_IDX);
+	cfg.size = ATR_PCIE_REG_SIZE;
+	cfg.trsl_addr = ATR_PCIE_REG_TRSL_ADDR;
+	cfg.port = ATR_PCIE_REG_PORT;
+	cfg.table = ATR_PCIE_REG_TABLE_NUM;
+	cfg.trsl_id = ATR_PCIE_REG_TRSL_PORT;
+	cfg.trsl_param = 0x0;
+	cfg.transparent = 0x0;
+	ret = mtk_pci_setup_atr(mdev, &cfg);
+	if (ret)
+		return ret;
+
+	/* Config ATR for EP to access RC's memory */
+	for (port = ATR_SRC_AXIS_0; port <= ATR_SRC_AXIS_3; port++) {
+		cfg.src_addr = ATR_PCIE_DEV_DMA_SRC_ADDR;
+		cfg.size = ATR_PCIE_DEV_DMA_SIZE;
+		cfg.trsl_addr = ATR_PCIE_DEV_DMA_TRSL_ADDR;
+		cfg.port = port;
+		cfg.table = ATR_PCIE_DEV_DMA_TABLE_NUM;
+		cfg.trsl_id = ATR_DST_PCI_TRX;
+		cfg.trsl_param = 0x0;
+		/* Enable transparent translation */
+		cfg.transparent = ATR_PCIE_DEV_DMA_TRANSPARENT;
+		ret = mtk_pci_setup_atr(mdev, &cfg);
+		if (ret)
+			return ret;
+	}
+
+	return 0;
+}
+
+static int mtk_pci_bar_init(struct mtk_md_dev *mdev)
+{
+	struct pci_dev *pdev = to_pci_dev(mdev->dev);
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+
+	/* Fixed offsets used on these mappings (MSI-X registers up to
+	 * 0x3080 on BAR0/1, the ATR-biased MHCCIF window on BAR2/3) are
+	 * guaranteed by the only two devices this driver binds (0x0900,
+	 * 0x01CA), whose BAR sizes are fixed by hardware; an undersized
+	 * BAR here is impossible, so no pci_resource_len() checks.
+	 */
+	priv->mac_reg_base = pcim_iomap_region(pdev, MTK_BAR_0_1_IDX,
+					       mdev->dev_str);
+	if (IS_ERR(priv->mac_reg_base)) {
+		dev_err(mdev->dev, "Failed to map BAR0/1\n");
+		return PTR_ERR(priv->mac_reg_base);
+	}
+
+	priv->bar23_addr = pcim_iomap_region(pdev, MTK_BAR_2_3_IDX,
+					     mdev->dev_str);
+	if (IS_ERR(priv->bar23_addr)) {
+		dev_err(mdev->dev, "Failed to map BAR2/3\n");
+		return PTR_ERR(priv->bar23_addr);
+	}
+
+	/* We use MD view base address "0" to observe registers */
+	priv->ext_reg_base = priv->bar23_addr - ATR_PCIE_REG_TRSL_ADDR;
+
+	return 0;
+}
+
+static int mtk_mhccif_irq_cb(int irq_id, void *data)
+{
+	struct mtk_md_dev *mdev = data;
+	struct mtk_pci_priv *priv;
+
+	priv = mdev->hw_priv;
+	queue_work(system_highpri_wq, &priv->mhccif_work);
+
+	return 0;
+}
+
+static int mtk_mhccif_init(struct mtk_md_dev *mdev)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	int ret;
+
+	INIT_LIST_HEAD(&priv->mhccif_cb_list);
+	spin_lock_init(&priv->mhccif_lock);
+	INIT_WORK(&priv->mhccif_work, mtk_mhccif_isr_work);
+
+	/* The write below goes through the BAR 2/3 ATR window, so this
+	 * function must run after mtk_pci_atr_init().
+	 * Mask every channel; consumers unmask the channels they need.
+	 * Do NOT ack here: EP2RC status bits are latched one-shot
+	 * notifications (e.g. BOOT_FLOW_SYNC raised before the driver
+	 * loads) that the device never re-sends.  Masking blocks
+	 * delivery but preserves the bit for delivery on unmask; an
+	 * ack would erase it and the boot flow would hang silently.
+	 */
+	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
+			MHCCIF_EP2RC_SW_INT_EAP_MASK_SET, U32_MAX);
+
+	ret = mtk_pci_get_irq_id(mdev, MTK_IRQ_SRC_MHCCIF);
+	if (ret < 0) {
+		dev_err(mdev->dev, "Failed to get mhccif_irq_id. ret=%d\n", ret);
+		return ret;
+	}
+	priv->mhccif_irq_id = ret;
+
+	ret = mtk_pci_register_irq(mdev, priv->mhccif_irq_id, mtk_mhccif_irq_cb, mdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to register mhccif_irq callback\n");
+		return ret;
+	}
+
+	return 0;
+}
+
+static void mtk_mhccif_exit(struct mtk_md_dev *mdev)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct mtk_mhccif_cb *cb, *tmp;
+
+	mtk_pci_unregister_irq(mdev, priv->mhccif_irq_id);
+	cancel_work_sync(&priv->mhccif_work);
+	/* the work unmasks the source on its way out */
+	mtk_pci_mask_irq(mdev, priv->mhccif_irq_id);
+
+	list_for_each_entry_safe(cb, tmp, &priv->mhccif_cb_list, entry) {
+		list_del(&cb->entry);
+		kfree(cb);
+	}
+}
+
+static irqreturn_t mtk_pci_irq_handler(struct mtk_md_dev *mdev, u32 irq_state)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	int irq_id;
+
+	/* Check whether each set bit has a callback, if has, call it */
+	do {
+		int (*cb)(int irq_id, void *data);
+
+		irq_id = fls(irq_state) - 1;
+		irq_state &= ~BIT(irq_id);
+		cb = READ_ONCE(priv->irq_cb_list[irq_id]);
+		if (likely(cb)) {
+			smp_rmb(); /* Ensure data is read after callback */
+			cb(irq_id, priv->irq_cb_data[irq_id]);
+		} else {
+			dev_err_ratelimited(mdev->dev,
+					    "Unhandled irq_id=%d, no callback for it.\n",
+					    irq_id);
+			mtk_pci_clear_irq(mdev, irq_id);
+		}
+	} while (irq_state);
+
+	return IRQ_HANDLED;
+}
+
+static irqreturn_t mtk_pci_irq_msix(int irq, void *data)
+{
+	struct mtk_pci_irq_desc *irq_desc = data;
+	struct mtk_md_dev *mdev = irq_desc->mdev;
+	struct mtk_pci_priv *priv;
+	u32 irq_state, irq_enable;
+
+	priv = mdev->hw_priv;
+	irq_state = mtk_pci_mac_read32(priv, REG_MSIX_ISTATUS_HOST_GRP0_0);
+	irq_enable = mtk_pci_mac_read32(priv, REG_IMASK_HOST_MSIX_GRP0_0);
+	irq_state &= irq_enable;
+
+	if (unlikely(irq_state == U32_MAX && irq_enable == U32_MAX))
+		return IRQ_NONE; /* device gone */
+
+	irq_state &= irq_desc->msix_bits; /* scope to this vector */
+	if (unlikely(!irq_state))
+		return IRQ_NONE;
+
+	/* Mask the bit; the consumer unmasks it when it is done */
+	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, irq_state);
+
+	return mtk_pci_irq_handler(mdev, irq_state);
+}
+
+static int mtk_pci_request_irq_msix(struct mtk_md_dev *mdev)
+{
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct mtk_pci_irq_desc *irq_desc;
+	struct pci_dev *pdev;
+	int ret, i;
+
+	pdev = to_pci_dev(mdev->dev);
+	irq_desc = priv->irq_desc;
+
+	priv->irq_cnt = MTK_IRQ_CNT_MAX;
+
+	for (i = 0; i < MTK_IRQ_CNT_MAX; i++) {
+		irq_desc[i].mdev = mdev;
+		irq_desc[i].msix_bits = BIT(i);
+		snprintf(irq_desc[i].name, MTK_IRQ_NAME_LEN, "msix%d-%s", i, mdev->dev_str);
+		ret = pci_request_irq(pdev, i, mtk_pci_irq_msix, NULL,
+				      &irq_desc[i], "%s", irq_desc[i].name);
+		if (ret) {
+			dev_err(mdev->dev, "Failed to request %s: ret=%d\n",
+				irq_desc[i].name, ret);
+			for (i--; i >= 0; i--)
+				pci_free_irq(pdev, i, &irq_desc[i]);
+			priv->irq_cnt = 0;
+			return ret;
+		}
+	}
+
+	return 0;
+}
+
+static int mtk_pci_request_irq(struct mtk_md_dev *mdev)
+{
+	struct pci_dev *pdev = to_pci_dev(mdev->dev);
+	int ret;
+
+	ret = pci_alloc_irq_vectors(pdev, MTK_IRQ_CNT_MAX,
+				    MTK_IRQ_CNT_MAX, PCI_IRQ_MSIX);
+	if (ret < 0) {
+		dev_err(mdev->dev,
+			"Unable to alloc %d MSI-X vectors: ret=%d\n",
+			MTK_IRQ_CNT_MAX, ret);
+		return ret;
+	}
+
+	ret = mtk_pci_request_irq_msix(mdev);
+	if (ret)
+		pci_free_irq_vectors(pdev);
+
+	return ret;
+}
+
+static void mtk_pci_free_irq(struct mtk_md_dev *mdev)
+{
+	struct pci_dev *pdev = to_pci_dev(mdev->dev);
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	int i;
+
+	for (i = 0; i < priv->irq_cnt; i++)
+		pci_free_irq(pdev, i, &priv->irq_desc[i]);
+
+	pci_free_irq_vectors(pdev);
+}
+
+static int mtk_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
+{
+	struct device *dev = &pdev->dev;
+	struct mtk_pci_priv *priv;
+	struct mtk_md_dev *mdev;
+	int ret;
+
+	mdev = devm_kzalloc(dev, sizeof(*mdev), GFP_KERNEL);
+	if (!mdev) {
+		ret = -ENOMEM;
+		goto log_err;
+	}
+	mdev->dev = dev;
+
+	priv = devm_kzalloc(dev, sizeof(*priv), GFP_KERNEL);
+	if (!priv) {
+		ret = -ENOMEM;
+		goto log_err;
+	}
+
+	pci_set_drvdata(pdev, mdev);
+	priv->mdev = mdev;
+	mdev->hw_ver  = pdev->device;
+	mdev->hw_priv = priv;
+	mdev->dev     = dev;
+	snprintf(mdev->dev_str, MTK_DEV_STR_LEN, "%02x%02x%d",
+		 pdev->bus->number, PCI_SLOT(pdev->devfn), PCI_FUNC(pdev->devfn));
+	if (pdev->state_saved)
+		pci_restore_state(pdev);
+
+	ret = pci_enable_device(pdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to enable pci device.\n");
+		goto log_err;
+	}
+
+	ret = dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(64));
+	if (ret) {
+		dev_err(mdev->dev, "Failed to set DMA Mask and Coherent. (ret=%d)\n", ret);
+		goto disable_dev;
+	}
+
+	ret = mtk_pci_bar_init(mdev);
+	if (ret)
+		goto disable_dev;
+
+	ret = mtk_pci_atr_init(mdev);
+	if (ret)
+		goto disable_dev;
+
+	spin_lock_init(&priv->irq_cb_lock);
+
+	ret = mtk_mhccif_init(mdev);
+	if (ret)
+		goto disable_dev;
+
+	/* Mask every source before the handlers go in, so a source
+	 * left enabled by firmware or a previous bind cannot fire
+	 * into a half-initialised driver.
+	 */
+	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, U32_MAX);
+
+	ret = mtk_pci_request_irq(mdev);
+	if (ret)
+		goto free_mhccif;
+
+	pci_set_master(pdev);
+	mtk_pci_unmask_irq(mdev, priv->mhccif_irq_id);
+
+	if (!mtk_pci_link_check(mdev)) {
+		ret = -ENOLINK;
+		goto clear_master;
+	}
+
+	ret = pci_save_state(pdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to save PCI state: %d\n", ret);
+		goto clear_master;
+	}
+
+	priv->saved_state = pci_store_saved_state(pdev);
+	if (!priv->saved_state) {
+		ret = -ENOMEM;
+		goto clear_master;
+	}
+
+	return 0;
+
+clear_master:
+	pci_clear_master(pdev);
+	mtk_pci_free_irq(mdev);
+free_mhccif:
+	mtk_mhccif_exit(mdev);
+disable_dev:
+	pci_disable_device(pdev);
+log_err:
+	dev_err(dev, "Failed to probe device, ret=%d\n", ret);
+
+	return ret;
+}
+
+static void mtk_pci_remove(struct pci_dev *pdev)
+{
+	struct mtk_md_dev *mdev = pci_get_drvdata(pdev);
+	struct mtk_pci_priv *priv = mdev->hw_priv;
+	struct device *dev = &pdev->dev;
+	int ret;
+
+	/* Silence every source before tearing anything down. */
+	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, U32_MAX);
+
+	/* Unregisters the callback (masks and synchronises the vector),
+	 * then cancels the work.  With the callback gone the work cannot
+	 * be requeued.
+	 */
+	mtk_mhccif_exit(mdev);
+	mtk_pci_free_irq(mdev);
+	pci_clear_master(pdev);
+	pci_disable_device(pdev);
+
+	/* Reset last: the endpoint comes back at power-on defaults,
+	 * so no MMIO, config write or MSI-X teardown may follow it.
+	 */
+	ret = mtk_pci_pldr(mdev);
+	if (ret && mtk_pci_link_check(mdev)) {
+		dev_warn(dev, "PLDR failed (%d), trying MHCCIF reset\n", ret);
+		if (mtk_pci_send_ext_evt(mdev, DEV_EVT_H2D_DEVICE_RESET))
+			dev_err(dev, "MHCCIF reset failed\n");
+	}
+
+	pci_load_and_free_saved_state(pdev, &priv->saved_state);
+}
+
+static pci_ers_result_t mtk_pci_error_detected(struct pci_dev *pdev,
+					       pci_channel_state_t state)
+{
+	struct mtk_md_dev *mdev = pci_get_drvdata(pdev);
+
+	dev_err(mdev->dev, "AER detected: pci_channel_state_t=%u\n", state);
+
+	/* AER recovery not supported, disconnect the device */
+	return PCI_ERS_RESULT_DISCONNECT;
+}
+
+static const struct pci_error_handlers mtk_pci_err_handler = {
+	.error_detected = mtk_pci_error_detected,
+};
+
+static struct pci_driver mtk_pci_drv = {
+	.name = "mtk_pci_drv",
+	.id_table = t9xx_pci_table,
+	.probe = mtk_pci_probe,
+	.remove = mtk_pci_remove,
+	.err_handler = &mtk_pci_err_handler
+};
+
+module_pci_driver(mtk_pci_drv);
+
+MODULE_DESCRIPTION("MediaTek T9xx PCIe WWAN driver");
+MODULE_LICENSE("GPL");
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.h b/drivers/net/wwan/t9xx/pcie/mtk_pci.h
new file mode 100644
index 000000000000..a155846dbba3
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.h
@@ -0,0 +1,143 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_PCI_H__
+#define __MTK_PCI_H__
+
+#include <linux/pci.h>
+
+#include "../mtk_dev.h"
+
+enum mtk_irq_src {
+	MTK_IRQ_SRC_MIN,
+	MTK_IRQ_SRC_MHCCIF,
+	MTK_IRQ_SRC_CLDMA0,
+	MTK_IRQ_SRC_CLDMA1,
+	MTK_IRQ_SRC_MAX
+};
+
+enum mtk_atr_src_port {
+	ATR_SRC_PCI_WIN0 = 0,
+	ATR_SRC_PCI_WIN1,
+	ATR_SRC_AXIS_0,
+	ATR_SRC_AXIS_1,
+	ATR_SRC_AXIS_2,
+	ATR_SRC_AXIS_3,
+};
+
+enum mtk_atr_dst_port {
+	ATR_DST_PCI_TRX = 0,
+	ATR_DST_AXIM_0 = 4,
+};
+
+#define MTK_PCI_CLASS                 0x0D4000
+#define MTK_PCI_VENDOR_ID             0x14C3
+#define CEI_PCI_VENDOR_ID             0x03F0
+
+#define MTK_BAR_0_1_IDX                 0
+#define MTK_BAR_2_3_IDX                 2
+
+#define MTK_IRQ_CNT_MAX				32
+#define MTK_IRQ_NAME_LEN			32
+
+#define ATR_PORT_OFFSET				0x100
+#define ATR_TABLE_OFFSET			0x20
+#define ATR_TABLE_NUM_PER_ATR			8
+#define ATR_PCIE_REG_TRSL_ADDR			0x10000000
+#define ATR_PCIE_REG_SIZE			0x00400000
+#define ATR_PCIE_REG_PORT			ATR_SRC_PCI_WIN0
+#define ATR_PCIE_REG_TABLE_NUM			1
+#define ATR_PCIE_REG_TRSL_PORT			ATR_DST_AXIM_0
+#define ATR_PCIE_DEV_DMA_SRC_ADDR		0x00000000
+#define ATR_PCIE_DEV_DMA_TRANSPARENT		1
+#define ATR_PCIE_DEV_DMA_SIZE			0
+#define ATR_PCIE_DEV_DMA_TABLE_NUM		0
+#define ATR_PCIE_DEV_DMA_TRSL_ADDR		0x00000000
+
+struct mtk_pci_irq_desc {
+	struct mtk_md_dev *mdev;
+	u32 msix_bits;
+	char name[MTK_IRQ_NAME_LEN];
+};
+
+struct mtk_pci_priv {
+	struct mtk_md_dev *mdev;
+	void __iomem *bar23_addr;
+	void __iomem *mac_reg_base;
+	void __iomem *ext_reg_base;
+	int irq_cnt;
+	/* irq_cb_lock: protects irq_cb_list[] and irq_cb_data[] */
+	spinlock_t irq_cb_lock;
+	void *irq_cb_data[MTK_IRQ_CNT_MAX];
+
+	int (*irq_cb_list[MTK_IRQ_CNT_MAX])(int irq_id, void *data);
+	struct mtk_pci_irq_desc irq_desc[MTK_IRQ_CNT_MAX];
+	struct list_head mhccif_cb_list;
+	/* mhccif_lock: lock to protect mhccif_cb_list */
+	spinlock_t mhccif_lock;
+	struct work_struct mhccif_work;
+	int mhccif_irq_id;
+	struct pci_saved_state *saved_state;
+};
+
+struct mtk_atr_cfg {
+	u64 src_addr;
+	u64 trsl_addr;
+	u64 size;
+	u32 port;      /* Port number */
+	u32 table;     /* Table number (8 tables for each port) */
+	u32 trsl_id;
+	u32 trsl_param;
+	u32 transparent;
+};
+
+/* BAR 0/1 MMIO access */
+static inline u32 mtk_pci_mac_read32(struct mtk_pci_priv *priv, u64 addr)
+{
+	return ioread32(priv->mac_reg_base + addr);
+}
+
+static inline void mtk_pci_mac_write32(struct mtk_pci_priv *priv, u64 addr, u32 val)
+{
+	iowrite32(val, priv->mac_reg_base + addr);
+}
+
+/* BAR 2/3 MMIO access */
+static inline u32 mtk_pci_read32(struct mtk_md_dev *mdev, u64 addr)
+{
+	return ioread32(((struct mtk_pci_priv *)mdev->hw_priv)->ext_reg_base + addr);
+}
+
+static inline void mtk_pci_write32(struct mtk_md_dev *mdev, u64 addr, u32 val)
+{
+	iowrite32(val, ((struct mtk_pci_priv *)mdev->hw_priv)->ext_reg_base + addr);
+}
+
+/* Device operations */
+u32 mtk_pci_get_dev_state(struct mtk_md_dev *mdev);
+void mtk_pci_ack_dev_state(struct mtk_md_dev *mdev, u32 state);
+/* IRQ Related operations */
+int mtk_pci_get_irq_id(struct mtk_md_dev *mdev, enum mtk_irq_src irq_src);
+int mtk_pci_get_virq_id(struct mtk_md_dev *mdev, int irq_id);
+int mtk_pci_register_irq(struct mtk_md_dev *mdev, int irq_id,
+			 int (*irq_cb)(int irq_id, void *data), void *data);
+int mtk_pci_unregister_irq(struct mtk_md_dev *mdev, int irq_id);
+int mtk_pci_mask_irq(struct mtk_md_dev *mdev, int irq_id);
+int mtk_pci_unmask_irq(struct mtk_md_dev *mdev, int irq_id);
+int mtk_pci_clear_irq(struct mtk_md_dev *mdev, int irq_id);
+/* External event related */
+int mtk_pci_register_ext_evt(struct mtk_md_dev *mdev, u32 chs,
+			     int (*evt_cb)(u32 status, void *data), void *data);
+void mtk_pci_unregister_ext_evt(struct mtk_md_dev *mdev, u32 chs);
+void mtk_pci_mask_ext_evt(struct mtk_md_dev *mdev, u32 chs);
+void mtk_pci_unmask_ext_evt(struct mtk_md_dev *mdev, u32 chs);
+void mtk_pci_clear_ext_evt(struct mtk_md_dev *mdev, u32 chs);
+int mtk_pci_send_ext_evt(struct mtk_md_dev *mdev, u32 ch);
+int mtk_pci_pldr(struct mtk_md_dev *mdev);
+bool mtk_pci_link_check(struct mtk_md_dev *mdev);
+int mtk_pci_setup_atr(struct mtk_md_dev *mdev, struct mtk_atr_cfg *cfg);
+void mtk_pci_atr_disable(struct mtk_pci_priv *priv);
+
+#endif /* __MTK_PCI_H__ */
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h b/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h
new file mode 100644
index 000000000000..a97ad6630dcf
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h
@@ -0,0 +1,34 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_PCI_REG_H__
+#define __MTK_PCI_REG_H__
+
+#define REG_ATR_PCIE_WIN0_T0_SRC_ADDR_LSB	0x0600
+#define REG_ATR_PCIE_WIN0_T0_SRC_ADDR_MSB	0x0604
+#define REG_ATR_PCIE_WIN0_T0_TRSL_ADDR_LSB	0x0608
+#define REG_ATR_PCIE_WIN0_T0_TRSL_ADDR_MSB	0x060C
+#define REG_ATR_PCIE_WIN0_T0_TRSL_PARAM		0x0610
+#define REG_PCIE_DEBUG_DUMMY_7			0x0D1C
+#define REG_MSIX_ISTATUS_HOST_GRP0_0		0x0F00
+#define REG_IMASK_HOST_MSIX_SET_GRP0_0		0x3000
+#define REG_IMASK_HOST_MSIX_CLR_GRP0_0		0x3080
+#define REG_IMASK_HOST_MSIX_GRP0_0		0x3100
+
+/* mhccif registers */
+#define MHCCIF_RC2EP_SW_BSY			0x4
+#define MHCCIF_RC2EP_SW_TCHNUM			0xC
+#define MHCCIF_RC2EP_EVT_DEVICE_RESET		BIT(13)
+
+#define MHCCIF_EP2RC_SW_INT_STS			0x10
+#define MHCCIF_EP2RC_SW_INT_ACK			0x14
+#define MHCCIF_EP2RC_SW_INT_EAP_MASK		0x20
+#define MHCCIF_EP2RC_SW_INT_EAP_MASK_SET	0x30
+#define MHCCIF_EP2RC_SW_INT_EAP_MASK_CLR	0x40
+#define MHCCIF_EP2RC_EVT_BOOT_FLOW_SYNC		BIT(5)
+#define MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_SAP	BIT(15)
+#define MHCCIF_EP2RC_EVT_ASYNC_HS_NOTIFY_MD	BIT(16)
+
+#endif /* __MTK_PCI_REG_H__ */

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
  2026-09-30  7:46 ` [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  2026-09-30  7:46 ` [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface Jack Wu via B4 Relay
                   ` (3 subsequent siblings)
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

Introduce the control plane and transaction layer framework
for the T9XX WWAN driver. This patch adds the core data
structures (mtk_ctrl_blk, mtk_ctrl_trans) and initialization
entry points that subsequent patches build upon.

The control plane block (ctrl_blk) serves as the central
context for managing the transaction layer, port layer, and
FSM interactions. The actual DMA engine and TX/RX service
implementations are added in subsequent patches.

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 drivers/net/wwan/t9xx/Makefile              |  1 +
 drivers/net/wwan/t9xx/mtk_ctrl_plane.c      | 44 +++++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      | 22 +++++++++++++++
 drivers/net/wwan/t9xx/mtk_dev.h             |  3 ++
 drivers/net/wwan/t9xx/pcie/mtk_pci.c        |  1 +
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h | 21 ++++++++++++++
 6 files changed, 92 insertions(+)

diff --git a/drivers/net/wwan/t9xx/Makefile b/drivers/net/wwan/t9xx/Makefile
index bd6063c85c2b..0a07507f1ddf 100644
--- a/drivers/net/wwan/t9xx/Makefile
+++ b/drivers/net/wwan/t9xx/Makefile
@@ -6,4 +6,5 @@ ccflags-y += -I$(src)
 obj-$(CONFIG_MTK_T9XX) += mtk_t9xx.o
 
 mtk_t9xx-y := \
+	mtk_ctrl_plane.o \
 	pcie/mtk_pci.o
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
new file mode 100644
index 000000000000..fa2ab8c3e757
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
@@ -0,0 +1,44 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ * Copyright (c) 2022-2023, Intel Corporation.
+ */
+
+#include <linux/device.h>
+
+#include "mtk_ctrl_plane.h"
+
+/**
+ * mtk_ctrl_init() - Initialize the control plane block.
+ * @mdev: Pointer to the MTK modem device.
+ *
+ * Allocates and initializes the control plane block
+ * associated with @mdev.
+ *
+ * Return: 0 on success, -ENOMEM on allocation failure.
+ */
+int mtk_ctrl_init(struct mtk_md_dev *mdev)
+{
+	struct mtk_ctrl_blk *ctrl_blk;
+
+	ctrl_blk = devm_kzalloc(mdev->dev, sizeof(*ctrl_blk), GFP_KERNEL);
+	if (!ctrl_blk)
+		return -ENOMEM;
+
+	ctrl_blk->mdev = mdev;
+	mdev->ctrl_blk = ctrl_blk;
+
+	return 0;
+}
+
+/**
+ * mtk_ctrl_exit() - Clean up the control plane block.
+ * @mdev: Pointer to the MTK modem device.
+ *
+ * Clears the control plane block pointer. The allocation
+ * itself is managed by devres and freed on driver detach.
+ */
+void mtk_ctrl_exit(struct mtk_md_dev *mdev)
+{
+	mdev->ctrl_blk = NULL;
+}
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
new file mode 100644
index 000000000000..c141876ef95d
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
@@ -0,0 +1,22 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_CTRL_PLANE_H__
+#define __MTK_CTRL_PLANE_H__
+
+#include <linux/kref.h>
+#include <linux/skbuff.h>
+
+#include "mtk_dev.h"
+
+struct mtk_ctrl_blk {
+	struct mtk_md_dev *mdev;
+	struct mtk_ctrl_trans *trans;
+};
+
+int mtk_ctrl_init(struct mtk_md_dev *mdev);
+void mtk_ctrl_exit(struct mtk_md_dev *mdev);
+
+#endif /* __MTK_CTRL_PLANE_H__ */
diff --git a/drivers/net/wwan/t9xx/mtk_dev.h b/drivers/net/wwan/t9xx/mtk_dev.h
index 329e5e6695df..84fcd980110b 100644
--- a/drivers/net/wwan/t9xx/mtk_dev.h
+++ b/drivers/net/wwan/t9xx/mtk_dev.h
@@ -31,12 +31,15 @@ enum mtk_dev_evt_d2h {
 	DEV_EVT_D2H_ASYNC_HS_NOTIFY_MD	= BIT(6),
 };
 
+struct mtk_ctrl_blk;
+
 /* mtk_md_dev defines the structure of MTK modem device */
 struct mtk_md_dev {
 	struct device *dev;
 	void *hw_priv;
 	u32 hw_ver;
 	char dev_str[MTK_DEV_STR_LEN];
+	struct mtk_ctrl_blk *ctrl_blk;
 };
 
 #endif /* __MTK_DEV_H__ */
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
index 34ee823119fc..c2c2b824ff31 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_pci.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
@@ -12,6 +12,7 @@
 #include <linux/module.h>
 
 #include "mtk_dev.h"
+#include "mtk_trans_ctrl.h"
 #include "mtk_pci.h"
 #include "mtk_pci_reg.h"
 
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
new file mode 100644
index 000000000000..d6de4c43b529
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
@@ -0,0 +1,21 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_TRANS_CTRL_H__
+#define __MTK_TRANS_CTRL_H__
+
+#include <linux/kref.h>
+#include <linux/list.h>
+#include <linux/skbuff.h>
+#include <linux/types.h>
+
+#include "mtk_dev.h"
+
+struct mtk_ctrl_trans {
+	struct mtk_ctrl_blk *ctrl_blk;
+	struct mtk_md_dev *mdev;
+};
+
+#endif

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
  2026-09-30  7:46 ` [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core Jack Wu via B4 Relay
  2026-09-30  7:46 ` [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  2026-09-30  7:46 ` [PATCH v9 4/6] net: wwan: t9xx: Add control port Jack Wu via B4 Relay
                   ` (2 subsequent siblings)
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

Cross Layer Direct Memory Access(CLDMA) is the hardware
interface used by the control plane and designated to
translate data between the host and the device. It supports
8 hardware queues for the device AP and modem respectively.

CLDMA driver uses General Purpose Descriptor (GPD) to
describe transaction information that can be recognized by
CLDMA hardware. Once CLDMA hardware transaction is started,
it would fetch and parse GPD to transfer data correctly.
To facilitate the CLDMA transaction, a GPD ring for each
queue is used. Once the transaction is started, CLDMA
hardware will traverse the GPD ring to transfer data between
the host and the device until no GPD is available.

CLDMA TX flow:
Once a TX service receives the TX data from the port layer,
it uses APIs exported by the CLDMA driver to configure GPD
with the DMA address of TX data. After that, the service
triggers CLDMA to fetch the first available GPD to transfer
data.

CLDMA RX flow:
When there is RX data from the MD, CLDMA hardware asserts an
interrupt to notify the host to fetch data and dispatch it
to FSM (for handshake messages) or the port layer.
After CLDMA opening is finished, All RX GPDs are fulfilled
and ready to receive data from the device.

With this commit, probe registers the transport plane: the
queue info table for the AP and modem control channels is
set up and the transaction context is attached to the
control plane. Nothing is brought up yet - the FSM listener
that calls mtk_pcie_hif_init(), and with it the CLDMA
instances and the TRB service thread, arrives with the later
"Add FSM thread" patch; the userspace-visible port
interfaces arrive with "Add control port".

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 drivers/net/wwan/t9xx/Makefile              |    5 +-
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      |   50 +-
 drivers/net/wwan/t9xx/pcie/mtk_cldma.c      | 1520 +++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_cldma.h      |  165 +++
 drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c  |  561 ++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h  |  135 +++
 drivers/net/wwan/t9xx/pcie/mtk_pci.c        |   41 +
 drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h    |    1 +
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c |  615 +++++++++++
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h |   68 ++
 10 files changed, 3159 insertions(+), 2 deletions(-)

diff --git a/drivers/net/wwan/t9xx/Makefile b/drivers/net/wwan/t9xx/Makefile
index 0a07507f1ddf..74bef150e3a1 100644
--- a/drivers/net/wwan/t9xx/Makefile
+++ b/drivers/net/wwan/t9xx/Makefile
@@ -7,4 +7,7 @@ obj-$(CONFIG_MTK_T9XX) += mtk_t9xx.o
 
 mtk_t9xx-y := \
 	mtk_ctrl_plane.o \
-	pcie/mtk_pci.o
+	pcie/mtk_pci.o \
+	pcie/mtk_trans_ctrl.o \
+	pcie/mtk_cldma.o \
+	pcie/mtk_cldma_drv.o
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
index c141876ef95d..0dcc0ec2c45d 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
@@ -11,9 +11,57 @@
 
 #include "mtk_dev.h"
 
+#define Q_MTU_3_5K	(0xE00)
+#define Q_FRAG_3_5K	(0xE00)
+
+enum mtk_ccci_ch {
+	/* to sAP */
+	CCCI_SAP_CONTROL_RX = 0x1000,
+	CCCI_SAP_CONTROL_TX = 0x1001,
+	/* to MD */
+	CCCI_CONTROL_RX = 0x2000,
+	CCCI_CONTROL_TX = 0x2001,
+};
+
+enum mtk_trb_cmd_type {
+	TRB_CMD_MIN,
+	TRB_CMD_ENABLE,
+	TRB_CMD_TX,
+	TRB_CMD_DISABLE,
+	TRB_CMD_STOP,
+	TRB_CMD_RECOVER,
+	TRB_CMD_MAX,
+};
+
+enum mtk_hif_dev_ctrl_cmd {
+	HIF_CTRL_CMD_CHECK_TX_FULL,
+};
+
+struct trb_open_priv {
+	u8 log_rg_offset;
+	u32 tx_mtu;
+	u32 rx_mtu;
+	u32 tx_frag_size;
+	u32 rx_frag_size;
+	int (*rx_done)(struct sk_buff *skb, void *priv, bool force_recv);
+};
+
+struct trb {
+	u32 channel_id;
+	enum mtk_trb_cmd_type cmd;
+	int status;
+	struct kref kref;
+	void *priv;
+	int (*trb_complete)(struct sk_buff *skb);
+};
+
+union ctrl_hif_cmd_data {
+	u32 rx_ch;
+};
+
 struct mtk_ctrl_blk {
 	struct mtk_md_dev *mdev;
-	struct mtk_ctrl_trans *trans;
+	void *ctrl_hw_priv;
 };
 
 int mtk_ctrl_init(struct mtk_md_dev *mdev);
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
new file mode 100644
index 000000000000..0f281bfb38cc
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
@@ -0,0 +1,1520 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#include <linux/delay.h>
+#include <linux/device.h>
+#include <linux/dma-mapping.h>
+#include <linux/dmapool.h>
+#include <linux/err.h>
+#include <linux/interrupt.h>
+#include <linux/kdev_t.h>
+#include <linux/kernel.h>
+#include <linux/kthread.h>
+#include <linux/list.h>
+#include <linux/module.h>
+#include <linux/mutex.h>
+#include <linux/netdevice.h>
+#include <linux/sched.h>
+#include <linux/skbuff.h>
+#include <linux/slab.h>
+#include <linux/timer.h>
+#include <linux/wait.h>
+#include <linux/workqueue.h>
+#include "mtk_pci.h"
+#include "mtk_cldma.h"
+#include "mtk_cldma_drv.h"
+#include "mtk_dev.h"
+
+#define DMA_POOL_NAME_LEN	(64)
+#define WAIT_HWO_ROUND		(10)
+#define WAIT_HWO_TIME		(5)
+#define NO_BUDGET		(0)
+
+static const int mtk_cldma_hw_id_tbl[NR_CLDMA] = {
+	[CLDMA0] = CLDMA0_HW_ID,
+	[CLDMA1] = CLDMA1_HW_ID,
+};
+
+static void mtk_cldma_clr_bd_dsc(struct cldma_drv_info *drv_info,
+				 struct bd_dsc *bd_dsc_pool, int nr_bds)
+{
+	struct bd_dsc *bd_dsc;
+	int i;
+
+	for (i = 0; i < nr_bds; i++) {
+		bd_dsc = bd_dsc_pool + i;
+		dma_unmap_single(drv_info->mdev->dev, bd_dsc->data_dma_addr,
+				 bd_dsc->data_len, DMA_TO_DEVICE);
+		bd_dsc->data_dma_addr = 0;
+		bd_dsc->data_len = 0;
+		if (bd_dsc->bd->tx_bd.bd_flags & CLDMA_BD_FLAG_EOL) {
+			bd_dsc->bd->tx_bd.bd_flags &= ~CLDMA_BD_FLAG_EOL;
+			break;
+		}
+	}
+}
+
+static void mtk_cldma_tx_done_work(struct work_struct *work)
+{
+	struct txq *txq = container_of(work, struct txq, tx_done_work);
+	struct cldma_drv_info *drv_info;
+	struct mtk_ctrl_trans *trans;
+	struct mtk_md_dev *mdev;
+	struct sk_buff *skb;
+	struct trb_srv *srv;
+	struct tx_req *req;
+	unsigned int state;
+	bool was_starved;
+	struct trb *trb;
+	int i, hif_id;
+	u32 txqno;
+
+	drv_info = txq->drv_info;
+	hif_id = drv_info->hif_id;
+	txqno = txq->txqno;
+	mdev = drv_info->mdev;
+	trans = drv_info->cd->trans;
+
+again:
+	for (i = 0; i < txq->nr_gpds; i++) {
+		spin_lock(&txq->ring_lock);
+
+		req = txq->req_pool + txq->free_idx;
+
+		/* Ownership is decided by the slot alone: an empty slot has no
+		 * skb, and one the hardware still owns has HWO set. The
+		 * producer publishes both under ring_lock, so a slot can never
+		 * be observed half-filled and there is no need to bound the
+		 * walk by wr_idx.
+		 */
+		if (!req->skb || (req->gpd->tx_gpd.gpd_flags & CLDMA_GPD_FLAG_HWO)) {
+			spin_unlock(&txq->ring_lock);
+			break;
+		}
+
+		dma_rmb(); /* read descriptor fields after HWO check */
+
+		if (txq->nr_bds)
+			mtk_cldma_clr_bd_dsc(drv_info, req->bd_dsc_pool, txq->nr_bds);
+		else
+			dma_unmap_single(mdev->dev, req->data_dma_addr,
+					 req->data_len, DMA_TO_DEVICE);
+
+		skb = req->skb;
+		req->data_dma_addr = 0;
+		req->data_len = 0;
+		req->skb = NULL;
+
+		txq->free_idx = (txq->free_idx + 1) % txq->nr_gpds;
+		was_starved = atomic_fetch_inc(&txq->req_budget) == NO_BUDGET;
+
+		spin_unlock(&txq->ring_lock);
+
+		trb = (struct trb *)skb->cb;
+		trb->status = 0;
+		trb->trb_complete(skb);
+
+		/* The service may already have been torn down; completing the
+		 * TRB is still correct, there is just nobody left to wake.
+		 */
+		srv = trans->trb_srv[trans->srv_cfg[hif_id][txqno]];
+		if (was_starved && srv)
+			wake_up(&srv->trb_waitq);
+	}
+
+	state = mtk_cldma_check_intr_status(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
+	if (state) {
+		if (unlikely(state == LINK_ERROR_VAL))
+			goto out;
+
+		mtk_cldma_clr_intr_status(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
+
+		cond_resched();
+
+		goto again;
+	}
+
+out:
+	mtk_cldma_unmask_intr(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
+}
+
+/* Clamp every length the device reports before it reaches skb_put().
+ *
+ * In BD mode the bound is the per-BD data_allow_len, not frag_size or mtu:
+ * on the last BD of a packet data_allow_len is smaller than frag_size.  The
+ * running total is then clamped against req->mtu, so a device that inflates
+ * data_allow_len cannot grow the packet past the buffer.
+ *
+ * In non-BD mode the bound is req->mtu directly.  data_allow_len lives in a
+ * buffer the device can write, whereas req->mtu is host-only state and is by
+ * construction the value data_allow_len was programmed with, so a well-formed
+ * descriptor gives an identical result.
+ *
+ * Returns -EPROTO when the device over-reports so the caller drops the packet
+ * instead of delivering a malformed skb.
+ */
+static int mtk_cldma_rx_skb_adjust(struct mtk_md_dev *mdev, struct rxq *rxq,
+				   struct rx_req *req)
+{
+	u32 recv_len, allow_len, total_len = 0;
+	struct bd_dsc *bd_dsc;
+	int ret = 0;
+	int i;
+
+	for (i = 0; i < rxq->nr_bds; i++) {
+		bd_dsc = req->bd_dsc_pool + i;
+		if (bd_dsc->data_dma_addr) {
+			dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+					 req->frag_size, DMA_FROM_DEVICE);
+			bd_dsc->data_dma_addr = 0;
+		}
+		recv_len = le16_to_cpu(bd_dsc->bd->rx_bd.data_recv_len);
+		allow_len = le16_to_cpu(bd_dsc->bd->rx_bd.data_allow_len);
+		if (recv_len > allow_len) {
+			ret = -EPROTO;
+			recv_len = allow_len;
+		}
+		total_len += recv_len;
+		if (total_len > req->mtu) {
+			ret = -EPROTO;
+			recv_len -= min(total_len - req->mtu, recv_len);
+			total_len = req->mtu;
+		}
+		bd_dsc->skb->len = 0;
+		skb_reset_tail_pointer(bd_dsc->skb);
+		skb_put(bd_dsc->skb, recv_len);
+		if (req->skb != bd_dsc->skb) {
+			req->skb->len += bd_dsc->skb->len;
+			req->skb->data_len += bd_dsc->skb->len;
+		}
+		bd_dsc->bd->rx_bd.data_recv_len = 0;
+		bd_dsc->skb = NULL;
+	}
+	if (!rxq->nr_bds) {
+		if (req->data_dma_addr) {
+			dma_unmap_single(mdev->dev, req->data_dma_addr,
+					 req->mtu, DMA_FROM_DEVICE);
+			req->data_dma_addr = 0;
+		}
+		recv_len = le16_to_cpu(req->gpd->rx_gpd.data_recv_len);
+		if (recv_len > req->mtu) {
+			ret = -EPROTO;
+			recv_len = req->mtu;
+		}
+		req->skb->len = 0;
+		skb_reset_tail_pointer(req->skb);
+		skb_put(req->skb, recv_len);
+	}
+
+	req->gpd->rx_gpd.data_recv_len = 0;
+
+	return ret;
+}
+
+static int mtk_cldma_reload_rx_skb(struct mtk_md_dev *mdev, struct rxq *rxq,
+				   struct rx_req *req)
+{
+	struct sk_buff *tail = NULL;
+	struct bd_dsc *bd_dsc;
+	int nr_bds;
+	int i, ret;
+
+	nr_bds = rxq->nr_bds;
+
+	for (i = 0; i < nr_bds; i++) {
+		bd_dsc = req->bd_dsc_pool + i;
+		bd_dsc->skb = __dev_alloc_skb(req->frag_size, GFP_KERNEL);
+		if (!bd_dsc->skb) {
+			dev_warn_ratelimited(mdev->dev, "Failed to alloc SKB\n");
+			ret = -ENOMEM;
+			goto err_free_skb;
+		}
+		bd_dsc->skb->next = NULL;
+		bd_dsc->data_dma_addr = dma_map_single(mdev->dev, bd_dsc->skb->data,
+						       req->frag_size, DMA_FROM_DEVICE);
+		ret = dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr);
+		if (unlikely(ret)) {
+			dev_warn_ratelimited(mdev->dev, "Failed to map SKB data\n");
+			ret = -EFAULT;
+			goto err_free_skb;
+		}
+		bd_dsc->bd->rx_bd.data_buff_ptr_h =
+			cpu_to_le32((u64)(bd_dsc->data_dma_addr) >> 32);
+		bd_dsc->bd->rx_bd.data_buff_ptr_l =
+			cpu_to_le32(bd_dsc->data_dma_addr);
+		if (tail) {
+			tail->next = bd_dsc->skb;
+			tail = bd_dsc->skb;
+			continue;
+		}
+		if (!req->skb) {
+			req->skb = bd_dsc->skb;
+		} else {
+			skb_shinfo(req->skb)->frag_list = bd_dsc->skb;
+			tail = bd_dsc->skb;
+		}
+	}
+	if (!nr_bds) {
+		req->skb = __dev_alloc_skb(req->mtu, GFP_KERNEL);
+		if (!req->skb) {
+			ret = -ENOMEM;
+			goto err_free_skb;
+		}
+
+		req->data_dma_addr = dma_map_single(mdev->dev, req->skb->data,
+						    req->mtu, DMA_FROM_DEVICE);
+		ret = dma_mapping_error(mdev->dev, req->data_dma_addr);
+		if (unlikely(ret)) {
+			dev_warn_ratelimited(mdev->dev, "Failed to map SKB data\n");
+			ret = -EFAULT;
+			goto err_free_skb;
+		}
+		req->gpd->rx_gpd.data_buff_ptr_h = cpu_to_le32((u64)req->data_dma_addr >> 32);
+		req->gpd->rx_gpd.data_buff_ptr_l = cpu_to_le32(req->data_dma_addr);
+	}
+	return 0;
+
+err_free_skb:
+	if (nr_bds) {
+		if (req->skb)
+			skb_shinfo(req->skb)->frag_list = NULL;
+		for (i = 0; i < nr_bds; i++) {
+			bd_dsc = req->bd_dsc_pool + i;
+			if (!bd_dsc->skb)
+				break;
+			if (!dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr))
+				dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+						 req->frag_size, DMA_FROM_DEVICE);
+			bd_dsc->data_dma_addr = 0;
+			bd_dsc->skb->next = NULL;
+			dev_kfree_skb_any(bd_dsc->skb);
+			bd_dsc->skb = NULL;
+		}
+	} else {
+		req->data_dma_addr = 0;
+		if (req->skb)
+			dev_kfree_skb_any(req->skb);
+	}
+	req->skb = NULL;
+
+	return ret;
+}
+
+static int mtk_cldma_check_rx_req(struct cldma_drv_info *drv_info, struct rxq *rxq)
+{
+	struct rx_req *req = rxq->req_pool + rxq->free_idx;
+	u64 curr_addr;
+	int i;
+
+	curr_addr = mtk_cldma_get_rx_curr_addr(drv_info, rxq->rxqno);
+	if (unlikely(!curr_addr))
+		return -ENXIO;
+
+	if (req->gpd_dma_addr == curr_addr)
+		return -EAGAIN;
+	for (i = 0; i < WAIT_HWO_ROUND; i++) {
+		udelay(WAIT_HWO_TIME);
+		if (!(READ_ONCE(req->gpd->rx_gpd.gpd_flags) & CLDMA_GPD_FLAG_HWO))
+			break;
+	}
+	if (i == WAIT_HWO_ROUND) {
+		dev_err((drv_info->mdev)->dev, "Failed to check HWO=0\n");
+		return -EAGAIN;
+	}
+
+	return 0;
+}
+
+static bool mtk_cldma_rx_check_again(struct rxq *rxq)
+{
+	struct cldma_drv_info *drv_info;
+	bool need_check_again = false;
+	u32 state;
+	int rxqno;
+
+	drv_info = rxq->drv_info;
+	rxqno = rxq->rxqno;
+
+	do {
+		state = mtk_cldma_check_intr_status(drv_info, DIR_RX,
+						    rxqno, QUEUE_XFER_DONE);
+		if (state) {
+			if (unlikely(state == LINK_ERROR_VAL))
+				break;
+
+			mtk_cldma_clr_intr_status(drv_info, DIR_RX,
+						  rxqno, QUEUE_XFER_DONE);
+			cond_resched();
+			return true;
+		}
+	} while (need_check_again);
+
+	return false;
+}
+
+/* Point the device at the GPD it should fill next and start the queue.
+ * free_idx is owned by mtk_cldma_rx_done_work(), so this may only run
+ * from that worker, at a point where every consumed GPD is re-armed.
+ */
+static void mtk_cldma_rxq_restart(struct cldma_drv_info *drv_info, struct rxq *rxq)
+{
+	mtk_cldma_setup_start_addr(drv_info, DIR_RX, rxq->rxqno,
+				   rxq->req_pool[rxq->free_idx].gpd_dma_addr);
+	mtk_cldma_start_queue(drv_info, DIR_RX, rxq->rxqno);
+}
+
+static void mtk_cldma_rx_done_work(struct work_struct *work)
+{
+	struct rx_req *req = NULL, *pre_req = NULL;
+	struct rxq *rxq = container_of(work, struct rxq, rx_done_work);
+	struct cldma_drv_info *drv_info;
+	struct mtk_md_dev *mdev;
+	struct sk_buff *rx_skb;
+	int i, ret, idx, len_err;
+
+	drv_info = rxq->drv_info;
+	mdev = drv_info->mdev;
+
+again:
+	for (i = 0; i < rxq->nr_gpds; i++) {
+		req = rxq->req_pool + rxq->free_idx;
+		if (!req->skb) {
+			dev_err(mdev->dev,
+				"Failed to get valid req cldma%d rxq%d req%d\n",
+				drv_info->hw_id, rxq->rxqno, rxq->free_idx);
+			goto out;
+		}
+
+		if (req->gpd->rx_gpd.gpd_flags & CLDMA_GPD_FLAG_HWO)
+			break;
+
+		dma_rmb(); /* read descriptor fields after HWO check */
+
+		len_err = mtk_cldma_rx_skb_adjust(mdev, rxq, req);
+		rx_skb = req->skb;
+		req->skb = NULL;
+
+		ret = mtk_cldma_reload_rx_skb(mdev, rxq, req);
+		if (ret) {
+			/* Alloc failed — recycle old buffer, drop packet.
+			 * BD mode cannot recycle directly (BDP flag mismatch),
+			 * so accept the stall and let reset recovery handle it.
+			 */
+			if (rxq->nr_bds) {
+				dev_kfree_skb_any(rx_skb);
+				goto out;
+			}
+
+			skb_trim(rx_skb, 0);
+			req->skb = rx_skb;
+			req->data_dma_addr = dma_map_single(mdev->dev,
+							    rx_skb->data,
+							    req->mtu,
+							    DMA_FROM_DEVICE);
+			if (dma_mapping_error(mdev->dev, req->data_dma_addr)) {
+				req->data_dma_addr = 0;
+				/* Keep rx_skb in req->skb for clean stall.
+				 * HWO is not set — HW won't touch this slot.
+				 * Queue stalls until modem reset recovery.
+				 */
+				goto out;
+			} else {
+				req->gpd->rx_gpd.data_buff_ptr_h =
+					cpu_to_le32((u64)req->data_dma_addr >> 32);
+				req->gpd->rx_gpd.data_buff_ptr_l =
+					cpu_to_le32(req->data_dma_addr);
+			}
+		} else if (len_err) {
+			dev_err_ratelimited(mdev->dev,
+					    "Drop oversized packet on cldma%d rxq%d\n",
+					    drv_info->hw_id, rxq->rxqno);
+			dev_kfree_skb_any(rx_skb);
+		} else {
+			do {
+				ret = rxq->rx_done(rx_skb, rxq->arg,
+						   atomic_read(&rxq->need_exit) ? true : false);
+				if (ret == -EAGAIN)
+					usleep_range(1000, 2000);
+			} while (ret == -EAGAIN);
+		}
+
+		wmb(); /* ensure addr set done before HWO setup done  */
+
+		idx = rxq->free_idx == 0 ? rxq->nr_gpds - 1 : rxq->free_idx - 1;
+		pre_req = rxq->req_pool + idx;
+		pre_req->gpd->rx_gpd.gpd_flags |= CLDMA_GPD_FLAG_HWO;
+		rxq->free_idx = (rxq->free_idx + 1) % rxq->nr_gpds;
+	}
+
+	ret = mtk_cldma_check_rx_req(drv_info, rxq);
+	if (!ret)
+		goto again;
+
+	/* -ENXIO means the queue has no start address programmed, which is
+	 * exactly what a restart installs, so consume need_restart first.
+	 * Only the resume below is skipped in that case.
+	 */
+	if (!atomic_read(&rxq->need_exit)) {
+		if (atomic_xchg(&rxq->need_restart, 0))
+			mtk_cldma_rxq_restart(drv_info, rxq);
+		else if (ret != -ENXIO)
+			mtk_cldma_resume_queue(drv_info, DIR_RX, rxq->rxqno);
+	}
+
+	if (ret == -ENXIO)
+		goto out;
+
+	if (mtk_cldma_rx_check_again(rxq))
+		goto again;
+
+out:
+	mtk_cldma_unmask_intr(drv_info, DIR_RX, rxq->rxqno, QUEUE_XFER_DONE);
+	mtk_cldma_clear_ip_busy(drv_info);
+}
+
+static int mtk_cldma_alloc_tx_bd(struct cldma_drv_info *drv_info, struct txq *txq,
+				 struct tx_req *req)
+{
+	struct bd_dsc *bd_dsc, *last_bd_dsc = NULL;
+	int i;
+
+	req->bd_dsc_pool = kcalloc(txq->nr_bds, sizeof(*bd_dsc),
+				   GFP_KERNEL);
+	if (!req->bd_dsc_pool)
+		return -ENOMEM;
+
+	for (i = 0; i < txq->nr_bds; i++) {
+		bd_dsc = req->bd_dsc_pool + i;
+		bd_dsc->bd = dma_pool_zalloc(drv_info->bd_dma_pool, GFP_KERNEL,
+					     &bd_dsc->bd_dma_addr);
+		if (!bd_dsc->bd)
+			return -ENOMEM;
+		if (!last_bd_dsc) {
+			req->gpd->tx_gpd.data_buff_ptr_h =
+				cpu_to_le32((u64)(bd_dsc->bd_dma_addr) >> 32);
+			req->gpd->tx_gpd.data_buff_ptr_l =
+				cpu_to_le32(bd_dsc->bd_dma_addr);
+		} else {
+			last_bd_dsc->bd->tx_bd.next_bd_ptr_h =
+				cpu_to_le32((u64)(bd_dsc->bd_dma_addr) >> 32);
+			last_bd_dsc->bd->tx_bd.next_bd_ptr_l =
+				cpu_to_le32(bd_dsc->bd_dma_addr);
+		}
+		last_bd_dsc = bd_dsc;
+	}
+	return 0;
+}
+
+static struct txq *mtk_cldma_txq_alloc(struct cldma_drv_info *drv_info, struct queue_info *que)
+{
+	struct mtk_md_dev *mdev;
+	struct bd_dsc *bd_dsc;
+	struct tx_req *next;
+	struct tx_req *req;
+	u16 tx_frag_size;
+	struct txq *txq;
+	int i, j, ret;
+
+	mdev = drv_info->mdev;
+
+	txq = kzalloc_obj(*txq);
+	if (!txq)
+		return NULL;
+
+	txq->que = que;
+	txq->drv_info = drv_info;
+	txq->txqno = txq->que->txqno;
+	txq->nr_gpds = txq->que->tx_nr_gpds;
+	atomic_set(&txq->req_budget, txq->que->tx_nr_gpds);
+	spin_lock_init(&txq->ring_lock);
+	tx_frag_size = txq->que->tx_frag_size;
+	if (txq->que->tx_mtu > tx_frag_size && tx_frag_size)
+		txq->nr_bds = (txq->que->tx_mtu + tx_frag_size - 1) / tx_frag_size;
+
+	txq->req_pool = kcalloc(txq->nr_gpds, sizeof(*req), GFP_KERNEL);
+	if (!txq->req_pool)
+		goto err_free_txq;
+
+	for (i = 0; i < txq->nr_gpds; i++) {
+		req = txq->req_pool + i;
+		req->mtu = txq->que->tx_mtu;
+		req->frag_size = tx_frag_size;
+		req->gpd = dma_pool_zalloc(drv_info->gpd_dma_pool, GFP_KERNEL, &req->gpd_dma_addr);
+		if (!req->gpd)
+			goto err_free_req;
+		if (txq->nr_bds) {
+			ret = mtk_cldma_alloc_tx_bd(drv_info, txq, req);
+			if (ret)
+				goto err_free_req;
+			req->gpd->tx_gpd.gpd_flags |= CLDMA_GPD_FLAG_BDP;
+		}
+	}
+
+	for (i = 0; i < txq->nr_gpds; i++) {
+		req = txq->req_pool + i;
+		next = txq->req_pool + ((i + 1) % txq->nr_gpds);
+		req->gpd->tx_gpd.gpd_flags |= CLDMA_GPD_FLAG_IOC;
+		req->gpd->tx_gpd.next_gpd_ptr_h = cpu_to_le32((u64)(next->gpd_dma_addr) >> 32);
+		req->gpd->tx_gpd.next_gpd_ptr_l = cpu_to_le32(next->gpd_dma_addr);
+	}
+
+	INIT_WORK(&txq->tx_done_work, mtk_cldma_tx_done_work);
+
+	/* Publish the queue before unmasking its interrupts: an interrupt
+	 * that fires in between must find a valid txq, or the sources it
+	 * masks are never unmasked again.  The release pairs with the
+	 * acquire load in the ISR.
+	 */
+	smp_store_release(&drv_info->txq[txq->txqno], txq);
+	ret = mtk_cldma_stop_queue(drv_info, DIR_TX, txq->txqno);
+	if (ret) {
+		dev_warn(mdev->dev, "Failed to stop TX queue %d before alloc\n", txq->txqno);
+		WRITE_ONCE(drv_info->txq[txq->txqno], NULL);
+		goto err_free_req;
+	}
+	txq->tx_started = false;
+	mtk_cldma_setup_start_addr(drv_info, DIR_TX, txq->txqno,
+				   txq->req_pool[0].gpd_dma_addr);
+	mtk_cldma_unmask_intr(drv_info, DIR_TX, txq->txqno, QUEUE_ERROR);
+	mtk_cldma_unmask_intr(drv_info, DIR_TX, txq->txqno, QUEUE_XFER_DONE);
+
+	return txq;
+
+err_free_req:
+	for (i = 0; i < txq->nr_gpds; i++) {
+		req = txq->req_pool + i;
+		if (!req->gpd)
+			break;
+		if (req->bd_dsc_pool) {
+			for (j = 0; j < txq->nr_bds; j++) {
+				bd_dsc = req->bd_dsc_pool + j;
+				if (!bd_dsc->bd)
+					break;
+				dma_pool_free(drv_info->bd_dma_pool, bd_dsc->bd,
+					      bd_dsc->bd_dma_addr);
+			}
+			kfree(req->bd_dsc_pool);
+		}
+		dma_pool_free(drv_info->gpd_dma_pool, req->gpd, req->gpd_dma_addr);
+	}
+	kfree(txq->req_pool);
+err_free_txq:
+	kfree(txq);
+	return NULL;
+}
+
+/* mtk_cldma_drv_init() rewrites IP-global configuration (including the interrupt
+ * mask) and mtk_cldma_drv_reset() wipes every queue of the instance, so all
+ * queues still published in drv_info have to be re-armed, not only the one
+ * that triggered the reset.  A queue already unpublished (NULL slot) is being
+ * torn down and is deliberately left stopped.
+ */
+static void mtk_cldma_rearm_queues(struct cldma_drv_info *drv_info)
+{
+	struct txq *txq;
+	struct rxq *rxq;
+	int i;
+
+	mtk_cldma_drv_init(drv_info);
+
+	for (i = 0; i < HW_QUEUE_NUM; i++) {
+		txq = drv_info->txq[i];
+		if (txq) {
+			spin_lock(&txq->ring_lock);
+			mtk_cldma_setup_start_addr(drv_info, DIR_TX, i,
+						   txq->req_pool[txq->free_idx].gpd_dma_addr);
+			mtk_cldma_unmask_intr(drv_info, DIR_TX, i, QUEUE_ERROR);
+			mtk_cldma_unmask_intr(drv_info, DIR_TX, i, QUEUE_XFER_DONE);
+			if (READ_ONCE(txq->tx_started))
+				mtk_cldma_start_queue(drv_info, DIR_TX, i);
+			spin_unlock(&txq->ring_lock);
+		}
+
+		rxq = drv_info->rxq[i];
+		if (rxq) {
+			mtk_cldma_unmask_intr(drv_info, DIR_RX, i, QUEUE_ERROR);
+			mtk_cldma_unmask_intr(drv_info, DIR_RX, i, QUEUE_XFER_DONE);
+			/* free_idx belongs to rx_done_work, which may be halfway
+			 * through a packet: let it program the start address
+			 * once it reaches a consistent point.
+			 */
+			atomic_set(&rxq->need_restart, 1);
+			queue_work(drv_info->wq, &rxq->rx_done_work);
+		}
+	}
+}
+
+static void mtk_cldma_txq_free(struct cldma_drv_info *drv_info, u32 txqno)
+{
+	struct mtk_md_dev *mdev;
+	struct bd_dsc *bd_dsc;
+	struct tx_req *req;
+	struct txq *txq;
+	struct trb *trb;
+	int irq_id;
+	int i, j, ret;
+
+	mdev = drv_info->mdev;
+
+	txq = drv_info->txq[txqno];
+	/* Unpublish before teardown; pairs with the acquire loads of this
+	 * slot. No release is needed, there is no prior store to expose.
+	 */
+	WRITE_ONCE(drv_info->txq[txqno], NULL);
+	/* stop HW tx transaction; -ENODEV means the link is dead, so the
+	 * device cannot be walking the ring and the free may proceed.
+	 */
+	ret = mtk_cldma_stop_queue(drv_info, DIR_TX, txqno);
+	if (ret == -ETIMEDOUT) {
+		dev_err(mdev->dev, "TX queue %d stop timed out, resetting CLDMA%d\n",
+			txqno, drv_info->hw_id);
+		mtk_cldma_drv_reset(drv_info);
+		ret = mtk_cldma_stop_queue(drv_info, DIR_TX, txqno);
+		/* the reset took the whole instance down with this queue */
+		mtk_cldma_rearm_queues(drv_info);
+	}
+	txq->tx_started = false;
+
+	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
+	synchronize_irq(irq_id);
+	/* flush on-going work */
+	flush_work(&txq->tx_done_work);
+	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
+	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_ERROR);
+
+	if (ret == -ETIMEDOUT) {
+		/* Still a live DMA target: leak the ring rather than hand it
+		 * back to the allocator.
+		 */
+		dev_err(mdev->dev, "TX queue %d cannot be stopped, leaking its ring\n",
+			txqno);
+		return;
+	}
+
+	/* Free tx req resource. No ring_lock is taken here: txq was already
+	 * unpublished from drv_info->txq[] above, so no new producer can enter,
+	 * and synchronize_irq() plus flush_work() have retired the consumer.
+	 */
+	for (i = 0; i < txq->nr_gpds; i++) {
+		req = txq->req_pool + txq->free_idx;
+		if (req->skb && req->data_len) {
+			if (!txq->nr_bds)
+				dma_unmap_single(mdev->dev, req->data_dma_addr,
+						 req->data_len, DMA_TO_DEVICE);
+			for (j = 0; j < txq->nr_bds; j++) {
+				bd_dsc = req->bd_dsc_pool + j;
+				if (!bd_dsc->data_dma_addr)
+					continue;
+				dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+						 bd_dsc->data_len, DMA_TO_DEVICE);
+			}
+			trb = (struct trb *)req->skb->cb;
+			trb->status = -EPIPE;
+			trb->trb_complete(req->skb);
+		}
+		for (j = 0; j < txq->nr_bds; j++) {
+			bd_dsc = req->bd_dsc_pool + j;
+			dma_pool_free(drv_info->bd_dma_pool, bd_dsc->bd,
+				      bd_dsc->bd_dma_addr);
+		}
+		kfree(req->bd_dsc_pool);
+		dma_pool_free(drv_info->gpd_dma_pool, req->gpd, req->gpd_dma_addr);
+		txq->free_idx = (txq->free_idx + 1) % txq->nr_gpds;
+	}
+
+	kfree(txq->req_pool);
+	kfree(txq);
+}
+
+static int mtk_cldma_alloc_rx_bd(struct cldma_drv_info *drv_info, struct rx_req *req,
+				 int nr_bds)
+{
+	struct bd_dsc *bd_dsc, *last_bd_dsc = NULL;
+	struct sk_buff *tail = NULL;
+	struct mtk_md_dev *mdev;
+	u32 left_size;
+	int ret;
+	int i;
+
+	mdev = drv_info->mdev;
+	left_size = req->mtu;
+
+	req->bd_dsc_pool = kcalloc(nr_bds, sizeof(*bd_dsc),
+				   GFP_KERNEL);
+	if (!req->bd_dsc_pool)
+		return -ENOMEM;
+	for (i = 0; i < nr_bds; i++) {
+		bd_dsc = req->bd_dsc_pool + i;
+		bd_dsc->bd = dma_pool_zalloc(drv_info->bd_dma_pool, GFP_KERNEL,
+					     &bd_dsc->bd_dma_addr);
+		if (!bd_dsc->bd)
+			return -ENOMEM;
+
+		bd_dsc->skb = __dev_alloc_skb(req->frag_size, GFP_KERNEL);
+		if (!bd_dsc->skb)
+			return -ENOMEM;
+		bd_dsc->skb->next = NULL;
+		bd_dsc->data_dma_addr =
+			dma_map_single(mdev->dev, bd_dsc->skb->data,
+				       req->frag_size, DMA_FROM_DEVICE);
+		ret = dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr);
+		if (unlikely(ret))
+			return -ENOMEM;
+
+		bd_dsc->bd->rx_bd.data_buff_ptr_h =
+			cpu_to_le32((u64)(bd_dsc->data_dma_addr) >> 32);
+		bd_dsc->bd->rx_bd.data_buff_ptr_l =
+			cpu_to_le32(bd_dsc->data_dma_addr);
+		bd_dsc->bd->rx_bd.data_allow_len =
+			cpu_to_le16(min(req->frag_size, left_size));
+		left_size -= min(req->frag_size, left_size);
+		if (!last_bd_dsc) {
+			req->gpd->rx_gpd.data_buff_ptr_h =
+				cpu_to_le32((u64)(bd_dsc->bd_dma_addr) >> 32);
+			req->gpd->rx_gpd.data_buff_ptr_l =
+				cpu_to_le32(bd_dsc->bd_dma_addr);
+		} else {
+			last_bd_dsc->bd->rx_bd.next_bd_ptr_h =
+				cpu_to_le32((u64)(bd_dsc->bd_dma_addr) >> 32);
+			last_bd_dsc->bd->rx_bd.next_bd_ptr_l =
+				cpu_to_le32(bd_dsc->bd_dma_addr);
+		}
+		last_bd_dsc = bd_dsc;
+		if (tail) {
+			tail->next = bd_dsc->skb;
+			tail = bd_dsc->skb;
+			continue;
+		}
+		if (!req->skb) {
+			req->skb = bd_dsc->skb;
+		} else {
+			skb_shinfo(req->skb)->frag_list = bd_dsc->skb;
+			tail = bd_dsc->skb;
+		}
+	}
+	last_bd_dsc->bd->rx_bd.bd_flags |= CLDMA_BD_FLAG_EOL;
+	return 0;
+}
+
+static void mtk_cldma_rxq_alloc_cancel(struct cldma_drv_info *drv_info, struct rx_req *req,
+				       int nr_bds)
+{
+	struct mtk_md_dev *mdev;
+	struct bd_dsc *bd_dsc;
+	int i;
+
+	mdev = drv_info->mdev;
+
+	if (nr_bds) {
+		if (req->skb)
+			skb_shinfo(req->skb)->frag_list = NULL;
+		if (req->bd_dsc_pool) {
+			for (i = 0; i < nr_bds; i++) {
+				bd_dsc = req->bd_dsc_pool + i;
+				if (!bd_dsc->bd)
+					break;
+				if (bd_dsc->skb) {
+					if (!dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr))
+						dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+								 req->frag_size, DMA_FROM_DEVICE);
+					bd_dsc->data_dma_addr = 0;
+					bd_dsc->skb->next = NULL;
+					dev_kfree_skb_any(bd_dsc->skb);
+				}
+				dma_pool_free(drv_info->bd_dma_pool, bd_dsc->bd,
+					      bd_dsc->bd_dma_addr);
+			}
+			kfree(req->bd_dsc_pool);
+		}
+	} else {
+		if (req->skb) {
+			if (!dma_mapping_error(mdev->dev, req->data_dma_addr))
+				dma_unmap_single(mdev->dev, req->data_dma_addr,
+						 req->mtu, DMA_FROM_DEVICE);
+			req->data_dma_addr = 0;
+			dev_kfree_skb_any(req->skb);
+		}
+	}
+	dma_pool_free(drv_info->gpd_dma_pool, req->gpd, req->gpd_dma_addr);
+}
+
+static struct rxq *mtk_cldma_rxq_alloc(struct cldma_drv_info *drv_info, struct queue_info *que,
+				       struct sk_buff *skb)
+{
+	struct trb_open_priv *trb_open_priv = (struct trb_open_priv *)skb->data;
+	struct trb *trb = (struct trb *)skb->cb;
+	struct mtk_md_dev *mdev;
+	struct rx_req *next;
+	struct rx_req *req;
+	u16 rx_frag_size;
+	struct rxq *rxq;
+	int ret;
+	int i;
+
+	mdev = drv_info->mdev;
+
+	rxq = kzalloc_obj(*rxq);
+	if (!rxq)
+		return NULL;
+
+	rxq->que = que;
+	rxq->drv_info = drv_info;
+	rxq->rxqno = rxq->que->rxqno;
+	if (rxq->que->rx_nr_gpds < MIN_GPD_NUM) {
+		dev_err(mdev->dev,
+			"Failed to alloc cldma%d rxq%d due to gpd number < 2\n",
+			drv_info->hw_id, rxq->rxqno);
+		goto err_free_rxq;
+	}
+	rxq->nr_gpds = rxq->que->rx_nr_gpds;
+	rxq->arg = trb->priv;
+	rxq->rx_done = trb_open_priv->rx_done;
+	atomic_set(&rxq->need_exit, 0);
+	atomic_set(&rxq->need_restart, 0);
+	rx_frag_size = rxq->que->rx_frag_size;
+	if (rxq->que->rx_mtu > rx_frag_size && rx_frag_size)
+		rxq->nr_bds = (rxq->que->rx_mtu + rx_frag_size - 1) / rx_frag_size;
+
+	rxq->req_pool = kcalloc(rxq->nr_gpds, sizeof(*req), GFP_KERNEL);
+	if (!rxq->req_pool)
+		goto err_free_rxq;
+
+	/* setup rx request */
+	for (i = 0; i < rxq->nr_gpds; i++) {
+		req = rxq->req_pool + i;
+		req->mtu = rxq->que->rx_mtu;
+		req->frag_size = rx_frag_size;
+		req->gpd = dma_pool_zalloc(drv_info->gpd_dma_pool, GFP_KERNEL, &req->gpd_dma_addr);
+		if (!req->gpd)
+			goto err_free_req;
+		if (rxq->nr_bds) {
+			ret = mtk_cldma_alloc_rx_bd(drv_info, req, rxq->nr_bds);
+			if (ret)
+				goto err_free_req;
+			req->gpd->rx_gpd.gpd_flags |= CLDMA_GPD_FLAG_BDP;
+		} else {
+			req->skb = __dev_alloc_skb(req->mtu, GFP_KERNEL);
+			if (!req->skb)
+				goto err_free_req;
+			req->data_dma_addr = dma_map_single(mdev->dev, req->skb->data,
+							    req->mtu, DMA_FROM_DEVICE);
+			ret = dma_mapping_error(mdev->dev, req->data_dma_addr);
+			if (unlikely(ret))
+				goto err_free_req;
+		}
+	}
+
+	for (i = 0; i < rxq->nr_gpds; i++) {
+		req = rxq->req_pool + i;
+		next = rxq->req_pool + ((i + 1) % rxq->nr_gpds);
+		req->gpd->rx_gpd.gpd_flags |= CLDMA_GPD_FLAG_IOC;
+		req->gpd->rx_gpd.data_allow_len = cpu_to_le16(req->mtu);
+		req->gpd->rx_gpd.next_gpd_ptr_h = cpu_to_le32((u64)(next->gpd_dma_addr) >> 32);
+		req->gpd->rx_gpd.next_gpd_ptr_l = cpu_to_le32(next->gpd_dma_addr);
+		if (!rxq->nr_bds) {
+			req->gpd->rx_gpd.data_buff_ptr_h =
+				cpu_to_le32((u64)(req->data_dma_addr) >> 32);
+			req->gpd->rx_gpd.data_buff_ptr_l = cpu_to_le32(req->data_dma_addr);
+		}
+		if (i != rxq->nr_gpds - 1)
+			req->gpd->rx_gpd.gpd_flags |= CLDMA_GPD_FLAG_HWO;
+	}
+
+	INIT_WORK(&rxq->rx_done_work, mtk_cldma_rx_done_work);
+
+	/* Pairs with the acquire load in the ISR, as for txq above. */
+	smp_store_release(&drv_info->rxq[rxq->rxqno], rxq);
+	ret = mtk_cldma_stop_queue(drv_info, DIR_RX, rxq->rxqno);
+	if (ret) {
+		dev_warn(mdev->dev, "Failed to stop RX queue %d before alloc\n", rxq->rxqno);
+		WRITE_ONCE(drv_info->rxq[rxq->rxqno], NULL);
+		goto err_free_req;
+	}
+	mtk_cldma_setup_start_addr(drv_info, DIR_RX,
+				   rxq->rxqno, rxq->req_pool[0].gpd_dma_addr);
+	mtk_cldma_start_queue(drv_info, DIR_RX, rxq->rxqno);
+	mtk_cldma_unmask_intr(drv_info, DIR_RX, rxq->rxqno, QUEUE_ERROR);
+	mtk_cldma_unmask_intr(drv_info, DIR_RX, rxq->rxqno, QUEUE_XFER_DONE);
+
+	return rxq;
+
+err_free_req:
+	for (i = 0; i < rxq->nr_gpds; i++) {
+		req = rxq->req_pool + i;
+		if (!req->gpd)
+			break;
+		mtk_cldma_rxq_alloc_cancel(drv_info, req, rxq->nr_bds);
+	}
+
+	kfree(rxq->req_pool);
+err_free_rxq:
+	kfree(rxq);
+	return NULL;
+}
+
+static void mtk_cldma_rxq_free(struct cldma_drv_info *drv_info, u32 rxqno)
+{
+	struct mtk_md_dev *mdev;
+	struct bd_dsc *bd_dsc;
+	struct rx_req *req;
+	struct rxq *rxq;
+	int irq_id;
+	int i, j, ret;
+
+	mdev = drv_info->mdev;
+
+	rxq = drv_info->rxq[rxqno];
+	/* Unpublish before teardown; pairs with the acquire loads of this
+	 * slot. No release is needed, there is no prior store to expose.
+	 */
+	WRITE_ONCE(drv_info->rxq[rxqno], NULL);
+
+	/* stop HW rx transaction; -ENODEV means the link is dead, so the
+	 * device cannot be walking the ring and the free may proceed.
+	 */
+	atomic_set(&rxq->need_exit, 1);
+	ret = mtk_cldma_stop_queue(drv_info, DIR_RX, rxqno);
+	if (ret == -ETIMEDOUT) {
+		dev_err(mdev->dev, "RX queue %d stop timed out, resetting CLDMA%d\n",
+			rxqno, drv_info->hw_id);
+		mtk_cldma_drv_reset(drv_info);
+		ret = mtk_cldma_stop_queue(drv_info, DIR_RX, rxqno);
+		/* the reset took the whole instance down with this queue */
+		mtk_cldma_rearm_queues(drv_info);
+	}
+
+	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
+	synchronize_irq(irq_id);
+	/* flush on-going work */
+	flush_work(&rxq->rx_done_work);
+	/* mask L2 RX interrupt again to avoid race condition causing use-after-free issue */
+	mtk_cldma_mask_intr(drv_info, DIR_RX, rxqno, QUEUE_XFER_DONE);
+	mtk_cldma_mask_intr(drv_info, DIR_RX, rxqno, QUEUE_ERROR);
+
+	if (ret == -ETIMEDOUT) {
+		/* Still a live DMA target: leak the ring rather than hand it
+		 * back to the allocator.
+		 */
+		dev_err(mdev->dev, "RX queue %d cannot be stopped, leaking its ring\n",
+			rxqno);
+		return;
+	}
+
+	/* free rx req resource */
+	for (i = 0; i < rxq->nr_gpds; i++) {
+		req = rxq->req_pool + rxq->free_idx;
+		if (!(req->gpd->rx_gpd.gpd_flags & CLDMA_GPD_FLAG_HWO) &&
+		    le16_to_cpu(req->gpd->rx_gpd.data_recv_len)) {
+			if (!mtk_cldma_rx_skb_adjust(mdev, rxq, req)) {
+				rxq->rx_done(req->skb, rxq->arg, true);
+				req->skb = NULL;
+			}
+		}
+		if (req->skb) {
+			if (!rxq->nr_bds) {
+				if (req->data_dma_addr)
+					dma_unmap_single(mdev->dev, req->data_dma_addr,
+							 req->mtu, DMA_FROM_DEVICE);
+				dev_kfree_skb_any(req->skb);
+			} else if (req->bd_dsc_pool[0].skb) {
+				/* The head skb aliases bd_dsc_pool[0].skb and the
+				 * frag_list chain aliases the other BD skbs: break
+				 * the aliases and let the BD walk below free every
+				 * skb exactly once, as rxq_alloc_cancel() does.
+				 */
+				skb_shinfo(req->skb)->frag_list = NULL;
+			} else {
+				/* The BD skbs are already detached, the head owns
+				 * the whole chain through its frag_list.
+				 */
+				dev_kfree_skb_any(req->skb);
+			}
+			req->skb = NULL;
+		}
+		for (j = 0; j < rxq->nr_bds; j++) {
+			bd_dsc = req->bd_dsc_pool + j;
+			if (bd_dsc->skb) {
+				if (bd_dsc->data_dma_addr)
+					dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+							 req->frag_size, DMA_FROM_DEVICE);
+				bd_dsc->skb->next = NULL;
+				dev_kfree_skb_any(bd_dsc->skb);
+			}
+			dma_pool_free(drv_info->bd_dma_pool,
+				      bd_dsc->bd, bd_dsc->bd_dma_addr);
+		}
+		kfree(req->bd_dsc_pool);
+		dma_pool_free(drv_info->gpd_dma_pool, req->gpd, req->gpd_dma_addr);
+		rxq->free_idx = (rxq->free_idx + 1) % rxq->nr_gpds;
+	}
+
+	kfree(rxq->req_pool);
+	kfree(rxq);
+}
+
+/* The device reported a zero TX start address: it lost its queue state. */
+static int mtk_cldma_hw_recovery(struct cldma_drv_info *drv_info, u32 qno)
+{
+	u64 val;
+
+	dev_err(drv_info->mdev->dev,
+		"CLDMA%d lost its queue state, re-initializing\n",
+		drv_info->hw_id);
+
+	mtk_cldma_rearm_queues(drv_info);
+
+	val = mtk_cldma_get_tx_start_addr(drv_info, qno);
+	if (!val || val == U64_MAX)
+		return -EIO;
+
+	return 0;
+}
+
+static int mtk_cldma_start_xfer(struct cldma_drv_info *drv_info, u32 qno)
+{
+	struct txq *txq;
+	u64 val;
+	int ret;
+
+	txq = drv_info->txq[qno];
+
+	val = mtk_cldma_get_tx_start_addr(drv_info, qno);
+	if (unlikely(val == U64_MAX))
+		return -EIO;
+
+	if (unlikely(!val)) {
+		ret = mtk_cldma_hw_recovery(drv_info, qno);
+		if (ret)
+			return ret;
+	}
+
+	/* Hold ring_lock across programming and kicking the queue so the
+	 * start address given to the device is the one free_idx still
+	 * names; tx_done_work advances free_idx under the same lock.
+	 */
+	spin_lock(&txq->ring_lock);
+	if (unlikely(!READ_ONCE(txq->tx_started))) {
+		mtk_cldma_setup_start_addr(drv_info, DIR_TX, qno,
+					   txq->req_pool[txq->free_idx].gpd_dma_addr);
+		mtk_cldma_start_queue(drv_info, DIR_TX, qno);
+		WRITE_ONCE(txq->tx_started, true);
+	} else {
+		mtk_cldma_resume_queue(drv_info, DIR_TX, qno);
+	}
+	spin_unlock(&txq->ring_lock);
+
+	return 0;
+}
+
+int mtk_cldma_init(struct mtk_ctrl_trans *trans)
+{
+	struct cldma_dev *cd;
+
+	cd = kzalloc_obj(*cd);
+	if (!cd)
+		return -ENOMEM;
+
+	cd->trans = trans;
+	trans->dev = cd;
+
+	return 0;
+}
+
+void mtk_cldma_exit(struct mtk_ctrl_trans *trans)
+{
+	if (!trans->dev)
+		return;
+
+	kfree(trans->dev);
+	trans->dev = NULL;
+}
+
+static int mtk_cldma_open(struct cldma_dev *cd, struct sk_buff *skb)
+{
+	struct trb_open_priv *trb_open_priv = (struct trb_open_priv *)skb->data;
+	struct trb *trb = (struct trb *)skb->cb;
+	struct cldma_drv_info *drv_info;
+	struct queue_info *que;
+	struct txq *txq;
+	struct rxq *rxq;
+	int ret = 0;
+
+	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
+	if (unlikely(!que)) {
+		trb->status = -EINVAL;
+		trb->trb_complete(skb);
+		return -EINVAL;
+	}
+	drv_info = cd->cldma_drv_info[que->hif_id];
+	if (!drv_info) {
+		ret = -EIO;
+		goto out;
+	}
+
+	if (que->tx_mtu == 0 || que->rx_mtu == 0) {
+		dev_err((cd->trans->mdev)->dev,
+			"Failed to enable cldma%d txq%d rxq%d due to wrong mtu\n",
+			drv_info->hw_id, que->txqno, que->rxqno);
+		ret = -EINVAL;
+		goto out;
+	}
+
+	trb_open_priv->tx_mtu = que->tx_mtu;
+	trb_open_priv->rx_mtu = que->rx_mtu;
+	trb_open_priv->tx_frag_size = que->tx_frag_size;
+	trb_open_priv->rx_frag_size = que->rx_frag_size;
+
+	if (drv_info->txq[que->txqno] || drv_info->rxq[que->rxqno]) {
+		ret = -EBUSY;
+		goto out;
+	}
+
+	txq = mtk_cldma_txq_alloc(drv_info, que);
+	if (!txq) {
+		ret = -ENOMEM;
+		goto out;
+	}
+
+	rxq = mtk_cldma_rxq_alloc(drv_info, que, skb);
+	if (!rxq) {
+		ret = -ENOMEM;
+		mtk_cldma_txq_free(drv_info, txq->txqno);
+		goto out;
+	}
+
+out:
+	if (ret)
+		cd->trans->usr_cnt[que->hif_id][que->txqno]--;
+
+	trb->status = ret;
+	trb->trb_complete(skb);
+
+	return ret;
+}
+
+/* Complete every submitted-but-unkicked request with an error so no TRB
+ * is stranded in the ring when the doorbell cannot be rung.
+ */
+static void mtk_cldma_txq_flush(struct cldma_drv_info *drv_info,
+				struct txq *txq, int err)
+{
+	struct mtk_ctrl_trans *trans = drv_info->cd->trans;
+	int hif_id = drv_info->hif_id;
+	u32 txqno = txq->txqno;
+	struct sk_buff *skb;
+	struct trb_srv *srv;
+	struct tx_req *req;
+	bool was_starved;
+	struct trb *trb;
+	int i;
+
+	for (i = 0; i < txq->nr_gpds; i++) {
+		spin_lock(&txq->ring_lock);
+
+		req = txq->req_pool + txq->free_idx;
+		if (!req->skb) {
+			spin_unlock(&txq->ring_lock);
+			break;
+		}
+
+		req->gpd->tx_gpd.gpd_flags &= ~CLDMA_GPD_FLAG_HWO;
+
+		if (txq->nr_bds)
+			mtk_cldma_clr_bd_dsc(drv_info, req->bd_dsc_pool, txq->nr_bds);
+		else
+			dma_unmap_single(drv_info->mdev->dev, req->data_dma_addr,
+					 req->data_len, DMA_TO_DEVICE);
+
+		skb = req->skb;
+		req->data_dma_addr = 0;
+		req->data_len = 0;
+		req->skb = NULL;
+
+		txq->free_idx = (txq->free_idx + 1) % txq->nr_gpds;
+		was_starved = atomic_fetch_inc(&txq->req_budget) == NO_BUDGET;
+
+		spin_unlock(&txq->ring_lock);
+
+		trb = (struct trb *)skb->cb;
+		trb->status = err;
+		trb->trb_complete(skb);
+
+		/* The service may already have been torn down; completing the
+		 * TRB is still correct, there is just nobody left to wake.
+		 */
+		srv = trans->trb_srv[trans->srv_cfg[hif_id][txqno]];
+		if (was_starved && srv)
+			wake_up(&srv->trb_waitq);
+	}
+}
+
+static int mtk_cldma_tx(struct cldma_dev *cd, struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct cldma_drv_info *drv_info;
+	struct mtk_md_dev *mdev;
+	struct queue_info *que;
+	struct txq *txq;
+	int ret;
+
+	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
+	if (unlikely(!que))
+		return -EPIPE;
+	drv_info = cd->cldma_drv_info[que->hif_id];
+	if (unlikely(!drv_info))
+		return -EPIPE;
+	txq = drv_info->txq[que->txqno];
+	if (unlikely(!txq))
+		return -EPIPE;
+
+	mdev = drv_info->mdev;
+
+	ret = mtk_cldma_start_xfer(drv_info, que->txqno);
+	if (unlikely(ret)) {
+		dev_err(mdev->dev, "Failed to trigger cldma tx\n");
+		mtk_cldma_txq_flush(drv_info, txq, ret);
+	}
+
+	return ret;
+}
+
+static int mtk_cldma_close(struct cldma_dev *cd, struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct cldma_drv_info *drv_info;
+	struct queue_info *que;
+
+	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
+	if (unlikely(!que)) {
+		trb->status = -EPIPE;
+		trb->trb_complete(skb);
+		return -EPIPE;
+	}
+	drv_info = cd->cldma_drv_info[que->hif_id];
+	if (unlikely(!drv_info)) {
+		trb->status = -EPIPE;
+		trb->trb_complete(skb);
+		return -EPIPE;
+	}
+
+	if (drv_info->txq[que->txqno])
+		mtk_cldma_txq_free(drv_info, que->txqno);
+	if (drv_info->rxq[que->rxqno])
+		mtk_cldma_rxq_free(drv_info, que->rxqno);
+
+	trb->status = 0;
+	trb->trb_complete(skb);
+
+	return 0;
+}
+
+static int mtk_cldma_txbuf_set(struct cldma_drv_info *drv_info, struct sk_buff *skb,
+			       struct tx_req *req, int nr_bds)
+{
+	struct sk_buff *curr_skb, *next_skb;
+	struct mtk_md_dev *mdev;
+	struct bd_dsc *bd_dsc;
+	int ret;
+	int i;
+
+	mdev = drv_info->mdev;
+
+	if (nr_bds) {
+		bd_dsc = req->bd_dsc_pool;
+		curr_skb = skb;
+		for (i = 0; i < nr_bds && curr_skb; i++) {
+			bd_dsc = req->bd_dsc_pool + i;
+			if (req->bd_dsc_pool == bd_dsc) {
+				bd_dsc->data_len = skb->len - skb->data_len;
+				next_skb = skb_shinfo(skb)->frag_list;
+			} else {
+				bd_dsc->data_len = curr_skb->len;
+				next_skb = curr_skb->next;
+			}
+			bd_dsc->data_dma_addr = dma_map_single(mdev->dev, curr_skb->data,
+							       bd_dsc->data_len, DMA_TO_DEVICE);
+			ret = dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr);
+			if (unlikely(ret))
+				goto err_unmap_buffer;
+
+			bd_dsc->bd->tx_bd.data_buff_ptr_h =
+				cpu_to_le32((u64)(bd_dsc->data_dma_addr) >> 32);
+			bd_dsc->bd->tx_bd.data_buff_ptr_l = cpu_to_le32(bd_dsc->data_dma_addr);
+			bd_dsc->bd->tx_bd.data_buffer_len = cpu_to_le16(bd_dsc->data_len);
+			curr_skb = next_skb;
+		}
+		bd_dsc->bd->tx_bd.bd_flags = CLDMA_BD_FLAG_EOL;
+	} else {
+		/* Non-BD mode maps only the linear area; a nonlinear SKB
+		 * here would be silently truncated to skb_headlen().
+		 */
+		if (WARN_ON_ONCE(skb_is_nonlinear(skb)))
+			return -EINVAL;
+
+		req->data_dma_addr = dma_map_single(mdev->dev, skb->data,
+						    skb_headlen(skb), DMA_TO_DEVICE);
+		ret = dma_mapping_error(mdev->dev, req->data_dma_addr);
+		if (unlikely(ret)) {
+			req->data_dma_addr = 0;
+			goto err_exit;
+		}
+
+		req->gpd->tx_gpd.data_buff_ptr_h = cpu_to_le32((u64)(req->data_dma_addr) >> 32);
+		req->gpd->tx_gpd.data_buff_ptr_l = cpu_to_le32(req->data_dma_addr);
+	}
+
+	return 0;
+
+err_unmap_buffer:
+	for (i = 0; i < nr_bds; i++) {
+		bd_dsc = req->bd_dsc_pool + i;
+		if (dma_mapping_error(mdev->dev, bd_dsc->data_dma_addr)) {
+			bd_dsc->data_dma_addr = 0;
+			break;
+		}
+		dma_unmap_single(mdev->dev, bd_dsc->data_dma_addr,
+				 bd_dsc->data_len, DMA_TO_DEVICE);
+		bd_dsc->data_dma_addr = 0;
+	}
+err_exit:
+	dev_err_ratelimited(mdev->dev, "Failed to map dma! error:%d\n", ret);
+	return -ENOMEM;
+}
+
+int mtk_cldma_submit_tx(void *dev, struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct cldma_drv_info *drv_info;
+	struct cldma_dev *cd = dev;
+	struct queue_info *que;
+	struct tx_req *req;
+	struct txq *txq;
+	int ret;
+
+	/* the CLDMA device is unpublished before it is freed, so a submitter
+	 * that raced the teardown lands here with a NULL dev
+	 */
+	if (unlikely(!cd))
+		return -EINVAL;
+
+	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
+	if (unlikely(!que))
+		return -EINVAL;
+	drv_info = cd->cldma_drv_info[que->hif_id];
+	if (unlikely(!drv_info))
+		return -EINVAL;
+
+	txq = drv_info->txq[que->txqno];
+	if (unlikely(!txq))
+		return -EINVAL;
+
+	spin_lock(&txq->ring_lock);
+
+	if (!atomic_read(&txq->req_budget)) {
+		spin_unlock(&txq->ring_lock);
+		return -EAGAIN;
+	}
+
+	req = txq->req_pool + txq->wr_idx;
+	req->gpd->tx_gpd.debug_id = 0x01;
+	ret = mtk_cldma_txbuf_set(drv_info, skb, req, txq->nr_bds);
+	if (ret) {
+		spin_unlock(&txq->ring_lock);
+		return ret;
+	}
+
+	req->gpd->tx_gpd.data_buff_len = cpu_to_le16(skb->len);
+
+	req->data_len = skb->len;
+	req->skb = skb;
+
+	dma_wmb(); /* ensure req and data msg set done before HWO setup */
+
+	req->gpd->tx_gpd.gpd_flags |= CLDMA_GPD_FLAG_HWO;
+
+	txq->wr_idx = (txq->wr_idx + 1) % txq->nr_gpds;
+	atomic_dec(&txq->req_budget);
+
+	spin_unlock(&txq->ring_lock);
+
+	return 0;
+}
+
+int mtk_cldma_get_tx_budget(void *dev, enum mtk_hif_id hif_id, u32 qno)
+{
+	struct cldma_drv_info *drv_info;
+	struct cldma_dev *cd = dev;
+	struct txq *txq;
+
+	if (unlikely(hif_id >= NR_CLDMA || qno >= HW_QUE_NUM || !cd))
+		return -EINVAL;
+
+	drv_info = cd->cldma_drv_info[hif_id];
+	if (!drv_info)
+		return -EINVAL;
+	txq = drv_info->txq[qno];
+	if (!txq)
+		return -EINVAL;
+	return atomic_read(&txq->req_budget);
+}
+
+static int (*trb_act_tbl[TRB_CMD_MAX])(struct cldma_dev *cd, struct sk_buff *skb) = {
+	[TRB_CMD_ENABLE] = mtk_cldma_open,
+	[TRB_CMD_TX] = mtk_cldma_tx,
+	[TRB_CMD_DISABLE] = mtk_cldma_close,
+};
+
+int mtk_cldma_trb_process(void *dev, struct sk_buff *skb)
+{
+	struct cldma_dev *cd;
+	struct trb *trb;
+
+	if (!dev || !skb)
+		return -EINVAL;
+
+	cd = (struct cldma_dev *)dev;
+	trb = (struct trb *)skb->cb;
+
+	if (!(trb->cmd > TRB_CMD_MIN && trb->cmd < TRB_CMD_STOP))
+		return -EINVAL;
+
+	return trb_act_tbl[trb->cmd](cd, skb);
+}
+
+int mtk_cldma_check_ch_cfg(void *dev, struct queue_info *que)
+{
+	struct cldma_drv_info *drv_info;
+	struct cldma_dev *cd = dev;
+	struct mtk_md_dev *mdev;
+	struct txq *txq;
+	struct rxq *rxq;
+
+	mdev = cd->trans->mdev;
+	drv_info = cd->cldma_drv_info[que->hif_id];
+
+	if (!drv_info) {
+		dev_err(mdev->dev, "CLDMA%d has not been initialized\n",
+			mtk_cldma_hw_id_tbl[que->hif_id]);
+		return -EINVAL;
+	}
+
+	txq = drv_info->txq[que->txqno];
+	rxq = drv_info->rxq[que->rxqno];
+	if (!txq || !rxq) {
+		dev_err(mdev->dev,
+			"CLDMA%d txq%d rxq%d has not been enabled\n",
+			mtk_cldma_hw_id_tbl[que->hif_id], que->txqno, que->rxqno);
+		return -EINVAL;
+	}
+
+	if (que->tx_mtu != txq->que->tx_mtu || que->rx_mtu != rxq->que->rx_mtu) {
+		dev_err(mdev->dev,
+			"Channel:%08x tx_mtu:%08x rx_mtu:%08x do not match ch cfg\n",
+			que->tx_chl, que->tx_mtu, que->rx_mtu);
+		return -EINVAL;
+	}
+
+	return 0;
+}
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.h b/drivers/net/wwan/t9xx/pcie/mtk_cldma.h
new file mode 100644
index 000000000000..fa6d79f8b5df
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.h
@@ -0,0 +1,165 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_CLDMA_H__
+#define __MTK_CLDMA_H__
+
+#include <linux/dma-mapping.h>
+#include <linux/dmapool.h>
+#include <linux/interrupt.h>
+#include <linux/list.h>
+#include <linux/spinlock.h>
+#include <linux/types.h>
+
+#include "mtk_ctrl_plane.h"
+#include "mtk_trans_ctrl.h"
+
+struct mtk_fsm_param;
+
+#define TXQ(N)					(N)
+#define RXQ(N)					(N)
+
+#define CLDMA_GPD_FLAG_HWO			BIT(0)
+#define CLDMA_GPD_FLAG_BDP			BIT(1)
+#define CLDMA_GPD_FLAG_IOC			BIT(7)
+#define CLDMA_BD_FLAG_EOL			BIT(0)
+
+union gpd {
+	struct {
+		u8 gpd_flags;
+		u8 non_used1;
+		__le16 data_allow_len;
+		__le32 next_gpd_ptr_h;
+		__le32 next_gpd_ptr_l;
+		__le32 data_buff_ptr_h;
+		__le32 data_buff_ptr_l;
+		__le16 data_recv_len;
+		u8 non_used2;
+		u8 debug_id;
+	} rx_gpd;
+
+	struct {
+		u8 gpd_flags;
+		u8 non_used1;
+		u8 non_used2;
+		u8 debug_id;
+		__le32 next_gpd_ptr_h;
+		__le32 next_gpd_ptr_l;
+		__le32 data_buff_ptr_h;
+		__le32 data_buff_ptr_l;
+		__le16 data_buff_len;
+		__le16 non_used3;
+	} tx_gpd;
+} __packed;
+
+union bd {
+	struct {
+		u8 bd_flags;
+		u8 non_used1;
+		__le16 data_allow_len;
+		__le32 next_bd_ptr_h;
+		__le32 next_bd_ptr_l;
+		__le32 data_buff_ptr_h;
+		__le32 data_buff_ptr_l;
+		__le16 data_recv_len;
+		__le16 non_used2;
+	} rx_bd;
+
+	struct {
+		u8 bd_flags;
+		u8 non_used1;
+		__le16 non_used2;
+		__le32 next_bd_ptr_h;
+		__le32 next_bd_ptr_l;
+		__le32 data_buff_ptr_h;
+		__le32 data_buff_ptr_l;
+		__le16 data_buffer_len;
+		u8 extension_len;
+		u8 non_used3;
+	} tx_bd;
+} __packed;
+
+struct bd_dsc {
+	union bd *bd;
+	struct sk_buff *skb;
+	dma_addr_t bd_dma_addr;
+	dma_addr_t data_dma_addr;
+	size_t data_len;
+};
+
+struct rx_req {
+	union gpd *gpd;
+	u32 mtu;
+	struct sk_buff *skb;
+	size_t data_len;
+	dma_addr_t gpd_dma_addr;
+	dma_addr_t data_dma_addr;
+	u32 frag_size;
+	struct bd_dsc *bd_dsc_pool;
+};
+
+struct rxq {
+	struct cldma_drv_info *drv_info;
+	u32 rxqno;
+	struct queue_info *que;
+	struct work_struct rx_done_work;
+	struct rx_req *req_pool;
+	u32 nr_gpds;
+	u32 free_idx;
+	void *arg;
+	int (*rx_done)(struct sk_buff *skb, void *priv, bool force_recv);
+	u32 nr_bds;
+	atomic_t need_exit;
+	/* set when the queue must be programmed and started again after an IP
+	 * reset; consumed by mtk_cldma_rx_done_work(), the owner of free_idx
+	 */
+	atomic_t need_restart;
+};
+
+struct tx_req {
+	union gpd *gpd;
+	u32 mtu;
+	size_t data_len;
+	dma_addr_t data_dma_addr;
+	dma_addr_t gpd_dma_addr;
+	struct sk_buff *skb;
+	int (*trb_complete)(struct sk_buff *skb);
+	u32 frag_size;
+	struct bd_dsc *bd_dsc_pool;
+};
+
+struct txq {
+	struct cldma_drv_info *drv_info;
+	u32 txqno;
+	struct queue_info *que;
+	struct work_struct tx_done_work;
+	struct tx_req *req_pool;
+	u32 nr_gpds;
+	atomic_t req_budget;
+	/* ring_lock: serializes the producer (TRB service thread) against the
+	 * consumer (tx_done_work) over req_budget, wr_idx, free_idx and the
+	 * tx_req/GPD slots they index.
+	 */
+	spinlock_t ring_lock;
+	u32 wr_idx;
+	u32 free_idx;
+	bool tx_started;
+	u32 nr_bds;
+};
+
+struct cldma_dev {
+	struct cldma_drv_info *cldma_drv_info[NR_CLDMA];
+	struct mtk_ctrl_trans *trans;
+};
+
+int mtk_cldma_init(struct mtk_ctrl_trans *trans);
+void mtk_cldma_exit(struct mtk_ctrl_trans *trans);
+int mtk_cldma_submit_tx(void *dev, struct sk_buff *skb);
+int mtk_cldma_get_tx_budget(void *dev, enum mtk_hif_id hif_id, u32 qno);
+int mtk_cldma_trb_process(void *dev, struct sk_buff *skb);
+void mtk_cldma_fsm_state_listener(struct mtk_fsm_param *param, struct mtk_ctrl_trans *trans);
+int mtk_cldma_check_ch_cfg(void *dev, struct queue_info *que);
+
+#endif
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c
new file mode 100644
index 000000000000..8ab153da2083
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c
@@ -0,0 +1,561 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2023, MediaTek Inc.
+ */
+
+#include <linux/delay.h>
+#include <linux/device.h>
+#include <linux/iopoll.h>
+#include <linux/dma-mapping.h>
+#include <linux/dmapool.h>
+#include <linux/err.h>
+#include <linux/interrupt.h>
+#include <linux/kdev_t.h>
+#include <linux/kernel.h>
+#include <linux/kthread.h>
+#include <linux/list.h>
+#include <linux/module.h>
+#include <linux/mutex.h>
+#include <linux/netdevice.h>
+#include <linux/sched.h>
+#include <linux/skbuff.h>
+#include <linux/slab.h>
+#include <linux/timer.h>
+#include <linux/wait.h>
+#include <linux/workqueue.h>
+
+#include "mtk_cldma_drv.h"
+#include "mtk_dev.h"
+#include "mtk_pci.h"
+#include "mtk_pci_reg.h"
+
+#define WAIT_QUEUE_STOP		(70)
+
+#define CLDMA0_BASE_ADDR				(0x1021C000)
+#define CLDMA1_BASE_ADDR				(0x1021E000)
+
+#define CLDMA_RX_SKB_POOL_MAX_SIZE			(64)
+#define CLDMA_RX_SKB_RELOAD_THRESHOLD			(16)
+
+/* L2TISAR0 */
+#define TQ_ERR_INT_OFFSET				(16)
+#define TQ_ERR_INT_BITMASK				(0x00FF0000)
+#define TQ_ACTIVE_START_ERR_INT_OFFSET			(24)
+#define TQ_ACTIVE_START_ERR_INT_BITMASK			(0xFF000000)
+
+/* L2RISAR0 */
+#define RQ_ERR_INT_OFFSET				(16)
+#define RQ_ERR_INT_BITMASK				(0x00FF0000)
+#define RQ_ACTIVE_START_ERR_INT_OFFSET			(24)
+#define RQ_ACTIVE_START_ERR_INT_BITMASK			(0xFF000000)
+
+/* CLDMA IN(Tx) */
+#define REG_CLDMA_UL_START_ADDRL_0			(0x0004)
+#define REG_CLDMA_UL_START_ADDRH_0			(0x0008)
+#define REG_CLDMA_UL_CURRENT_ADDRL_0			(0x0044)
+#define REG_CLDMA_UL_CURRENT_ADDRH_0			(0x0048)
+#define REG_CLDMA_UL_STATUS				(0x0084)
+#define REG_CLDMA_UL_START_CMD				(0x0088)
+#define REG_CLDMA_UL_RESUME_CMD				(0x008C)
+#define REG_CLDMA_UL_STOP_CMD				(0x0090)
+#define REG_CLDMA_UL_ERROR				(0x0094)
+#define REG_CLDMA_UL_CFG				(0x0098)
+#define REG_CLDMA_UL_DUMMY_0				(0x009C)
+
+/* CLDMA OUT(Rx) */
+#define REG_CLDMA_SO_ERROR				(0x0400 + 0x0100)
+#define REG_CLDMA_SO_START_CMD				(0x0400 + 0x01BC)
+#define REG_CLDMA_SO_RESUME_CMD				(0x0400 + 0x01C0)
+#define REG_CLDMA_SO_STOP_CMD				(0x0400 + 0x01C4)
+#define REG_CLDMA_SO_DUMMY_0				(0x0400 + 0x0108)
+#define REG_CLDMA_SO_CFG				(0x0400 + 0x0004)
+#define REG_CLDMA_SO_START_ADDRL_0			(0x0400 + 0x0078)
+#define REG_CLDMA_SO_START_ADDRH_0			(0x0400 + 0x007C)
+#define REG_CLDMA_SO_CUR_ADDRL_0			(0x0400 + 0x00B8)
+#define REG_CLDMA_SO_CUR_ADDRH_0			(0x0400 + 0x00BC)
+#define REG_CLDMA_SO_STATUS				(0x0400 + 0x00F8)
+#define REG_CLDMA_DEBUG_ID_EN				(0x0400 + 0x00FC)
+#define REG_CLDMA_SO_LAST_UPDATE_ADDRL_0		(0x0400 + 0x01C8)
+#define REG_CLDMA_SO_LAST_UPDATE_ADDRH_0		(0x0400 + 0x01CC)
+
+/* CLDMA MISC */
+#define REG_CLDMA_L2TISAR0				(0x0800 + 0x0010)
+#define REG_CLDMA_L2TISAR1				(0x0800 + 0x0014)
+#define REG_CLDMA_L2TIMR0				(0x0800 + 0x0018)
+#define REG_CLDMA_L2TIMR1				(0x0800 + 0x001C)
+#define REG_CLDMA_L2TIMCR0				(0x0800 + 0x0020)
+#define REG_CLDMA_L2TIMCR1				(0x0800 + 0x0024)
+#define REG_CLDMA_L2TIMSR0				(0x0800 + 0x0028)
+#define REG_CLDMA_L2TIMSR1				(0x0800 + 0x002C)
+#define REG_CLDMA_L3TISAR0				(0x0800 + 0x0030)
+#define REG_CLDMA_L3TISAR1				(0x0800 + 0x0034)
+#define REG_CLDMA_L2RISAR0				(0x0800 + 0x0050)
+#define REG_CLDMA_L2RISAR1				(0x0800 + 0x0054)
+#define REG_CLDMA_L3RISAR0				(0x0800 + 0x0070)
+#define REG_CLDMA_L3RISAR1				(0x0800 + 0x0074)
+#define REG_CLDMA_IP_BUSY				(0x0800 + 0x00B4)
+#define REG_CLDMA_L3TISAR2				(0x0800 + 0x00C0)
+
+#define REG_CLDMA_L2RIMR0				(0x0800 + 0x00E8)
+#define REG_CLDMA_L2RIMR1				(0x0800 + 0x00EC)
+#define REG_CLDMA_L2RIMCR0				(0x0800 + 0x00F0)
+#define REG_CLDMA_L2RIMCR1				(0x0800 + 0x00F4)
+#define REG_CLDMA_L2RIMSR0				(0x0800 + 0x00F8)
+#define REG_CLDMA_L2RIMSR1				(0x0800 + 0x00FC)
+
+#define REG_CLDMA_INT_EAP_USIP_MASK			(0x0800 + 0x011C)
+
+#define REG_CLDMA_IP_BUSY_TO_PCIE_MASK			(0x0800 + 0x0194)
+#define REG_CLDMA_IP_BUSY_TO_PCIE_MASK_SET		(0x0800 + 0x0198)
+#define REG_CLDMA_IP_BUSY_TO_PCIE_MASK_CLR		(0x0800 + 0x019C)
+
+#define REG_CLDMA_IP_BUSY_TO_AP_MASK			(0x0800 + 0x0200)
+#define REG_CLDMA_IP_BUSY_TO_AP_MASK_SET		(0x0800 + 0x0204)
+#define REG_CLDMA_IP_BUSY_TO_AP_MASK_CLR		(0x0800 + 0x0208)
+#define REG_CLDMA_IP_BUSY_TO_MD_MASK_SET		(0x0800 + 0x0210)
+#define REG_CLDMA_RX_WORK_TO_REG_MASK_SET		(0x0800 + 0x021C)
+
+/* CLDMA RESET */
+#define REG_INFRA_RST0_SET				(0x120)
+#define REG_INFRA_RST0_CLR				(0x124)
+#define REG_CLDMA0_RST_SET_BIT				(8)
+#define REG_CLDMA0_RST_CLR_BIT				(8)
+
+struct cldma_hw_regs mtk_cldma_regs_m9xx = {
+	.cldma0_base_addr = CLDMA0_BASE_ADDR,
+	.cldma1_base_addr = CLDMA1_BASE_ADDR,
+	.cldma_rx_skb_pool_max_size = CLDMA_RX_SKB_POOL_MAX_SIZE,
+	.cldma_rx_skb_reload_threshold = CLDMA_RX_SKB_RELOAD_THRESHOLD,
+	.tq_err_int_offset = TQ_ERR_INT_OFFSET,
+	.tq_err_int_bitmask = TQ_ERR_INT_BITMASK,
+	.tq_active_start_err_int_offset = TQ_ACTIVE_START_ERR_INT_OFFSET,
+	.tq_active_start_err_int_bitmask = TQ_ACTIVE_START_ERR_INT_BITMASK,
+	.rq_err_int_offset = RQ_ERR_INT_OFFSET,
+	.rq_err_int_bitmask = RQ_ERR_INT_BITMASK,
+	.rq_active_start_err_int_offset = RQ_ACTIVE_START_ERR_INT_OFFSET,
+	.rq_active_start_err_int_bitmask = RQ_ACTIVE_START_ERR_INT_BITMASK,
+	.reg_cldma_ul_start_addrl_0 = REG_CLDMA_UL_START_ADDRL_0,
+	.reg_cldma_ul_start_addrh_0 = REG_CLDMA_UL_START_ADDRH_0,
+	.reg_cldma_ul_current_addrl_0 = REG_CLDMA_UL_CURRENT_ADDRL_0,
+	.reg_cldma_ul_current_addrh_0 = REG_CLDMA_UL_CURRENT_ADDRH_0,
+	.reg_cldma_ul_status = REG_CLDMA_UL_STATUS,
+	.reg_cldma_ul_start_cmd = REG_CLDMA_UL_START_CMD,
+	.reg_cldma_ul_resume_cmd = REG_CLDMA_UL_RESUME_CMD,
+	.reg_cldma_ul_stop_cmd = REG_CLDMA_UL_STOP_CMD,
+	.reg_cldma_ul_error = REG_CLDMA_UL_ERROR,
+	.reg_cldma_ul_cfg = REG_CLDMA_UL_CFG,
+	.reg_cldma_ul_dummy_0 = REG_CLDMA_UL_DUMMY_0,
+	.reg_cldma_so_error = REG_CLDMA_SO_ERROR,
+	.reg_cldma_so_start_cmd = REG_CLDMA_SO_START_CMD,
+	.reg_cldma_so_resume_cmd = REG_CLDMA_SO_RESUME_CMD,
+	.reg_cldma_so_stop_cmd = REG_CLDMA_SO_STOP_CMD,
+	.reg_cldma_so_dummy_0 = REG_CLDMA_SO_DUMMY_0,
+	.reg_cldma_so_cfg = REG_CLDMA_SO_CFG,
+	.reg_cldma_so_start_addrl_0 = REG_CLDMA_SO_START_ADDRL_0,
+	.reg_cldma_so_start_addrh_0 = REG_CLDMA_SO_START_ADDRH_0,
+	.reg_cldma_so_current_addrl_0 = REG_CLDMA_SO_CUR_ADDRL_0,
+	.reg_cldma_so_current_addrh_0 = REG_CLDMA_SO_CUR_ADDRH_0,
+	.reg_cldma_so_status = REG_CLDMA_SO_STATUS,
+	.reg_cldma_debug_id_en = REG_CLDMA_DEBUG_ID_EN,
+	.reg_cldma_so_last_update_addrl_0 = REG_CLDMA_SO_LAST_UPDATE_ADDRL_0,
+	.reg_cldma_so_last_update_addrh_0 = REG_CLDMA_SO_LAST_UPDATE_ADDRH_0,
+	.reg_cldma_l2tisar0 = REG_CLDMA_L2TISAR0,
+	.reg_cldma_l2tisar1 = REG_CLDMA_L2TISAR1,
+	.reg_cldma_l2timr0 = REG_CLDMA_L2TIMR0,
+	.reg_cldma_l2timr1 = REG_CLDMA_L2TIMR1,
+	.reg_cldma_l2timcr0 = REG_CLDMA_L2TIMCR0,
+	.reg_cldma_l2timcr1 = REG_CLDMA_L2TIMCR1,
+	.reg_cldma_l2timsr0 = REG_CLDMA_L2TIMSR0,
+	.reg_cldma_l2timsr1 = REG_CLDMA_L2TIMSR1,
+	.reg_cldma_l3tisar0 = REG_CLDMA_L3TISAR0,
+	.reg_cldma_l3tisar1 = REG_CLDMA_L3TISAR1,
+	.reg_cldma_l3tisar2 = REG_CLDMA_L3TISAR2,
+	.reg_cldma_l2risar0 = REG_CLDMA_L2RISAR0,
+	.reg_cldma_l2risar1 = REG_CLDMA_L2RISAR1,
+	.reg_cldma_l2rimr0 = REG_CLDMA_L2RIMR0,
+	.reg_cldma_l2rimr1 = REG_CLDMA_L2RIMR1,
+	.reg_cldma_l2rimcr0 = REG_CLDMA_L2RIMCR0,
+	.reg_cldma_l2rimcr1 = REG_CLDMA_L2RIMCR1,
+	.reg_cldma_l2rimsr0 = REG_CLDMA_L2RIMSR0,
+	.reg_cldma_l2rimsr1 = REG_CLDMA_L2RIMSR1,
+	.reg_cldma_l3risar0 = REG_CLDMA_L3RISAR0,
+	.reg_cldma_l3risar1 = REG_CLDMA_L3RISAR1,
+	.reg_cldma_ip_busy = REG_CLDMA_IP_BUSY,
+	.reg_cldma_int_mask = REG_CLDMA_INT_EAP_USIP_MASK,
+	.reg_cldma_ip_busy_to_pcie_mask = REG_CLDMA_IP_BUSY_TO_PCIE_MASK,
+	.reg_cldma_ip_busy_to_pcie_mask_set = REG_CLDMA_IP_BUSY_TO_PCIE_MASK_SET,
+	.reg_cldma_ip_busy_to_pcie_mask_clr = REG_CLDMA_IP_BUSY_TO_PCIE_MASK_CLR,
+	.reg_cldma_ip_busy_to_ap_mask = REG_CLDMA_IP_BUSY_TO_AP_MASK,
+	.reg_cldma_ip_busy_to_ap_mask_set = REG_CLDMA_IP_BUSY_TO_AP_MASK_SET,
+	.reg_cldma_ip_busy_to_ap_mask_clr = REG_CLDMA_IP_BUSY_TO_AP_MASK_CLR,
+	.reg_cldma_ip_busy_to_md_mask_set = REG_CLDMA_IP_BUSY_TO_MD_MASK_SET,
+	.reg_cldma_rx_work_to_reg_mask_set = REG_CLDMA_RX_WORK_TO_REG_MASK_SET,
+	.reg_infra_rst0_set = REG_INFRA_RST0_SET,
+	.reg_infra_rst0_clr = REG_INFRA_RST0_CLR,
+};
+
+void mtk_cldma_drv_reset(struct cldma_drv_info *drv_info)
+{
+	struct cldma_hw_regs *hw_regs;
+	struct mtk_md_dev *mdev;
+	u32 val;
+
+	mdev = drv_info->mdev;
+	hw_regs = drv_info->hw_regs;
+
+	val = mtk_pci_read32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_set);
+	val |= 1 << (REG_CLDMA0_RST_SET_BIT + drv_info->hw_id);
+	mtk_pci_write32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_set, val);
+	udelay(1);
+	val = mtk_pci_read32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_clr);
+	val |= 1 << (REG_CLDMA0_RST_CLR_BIT + drv_info->hw_id);
+	mtk_pci_write32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_clr, val);
+}
+
+void mtk_cldma_drv_init(struct cldma_drv_info *drv_info)
+{
+	struct cldma_hw_regs *hw_regs;
+	struct mtk_md_dev *mdev;
+	int base;
+	u32 val;
+
+	mdev = drv_info->mdev;
+	base = drv_info->base_addr;
+	hw_regs = drv_info->hw_regs;
+
+	/* set CLDMA to 64 bit mode GPD */
+	val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_ul_cfg);
+	val = (val & (~(0x7 << 5))) | ((0x4) << 5);
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_ul_cfg, val);
+
+	val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_so_cfg);
+	val = (val & (~(0x7 << 10))) | ((0x4) << 10) | (1 << 2);
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_so_cfg, val);
+
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_rx_work_to_reg_mask_set, ALLQ);
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_ip_busy_to_pcie_mask_set,
+			ALLQ << 16);
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_ip_busy_to_pcie_mask_clr,
+			ALLQ << 24);
+
+	/* enable interrupt to PCIe */
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_int_mask, 0);
+
+	/* disable illegal memory check */
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_ul_dummy_0, 1);
+	mtk_pci_write32(mdev, base + hw_regs->reg_cldma_so_dummy_0, 1);
+}
+
+void mtk_cldma_setup_start_addr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+				u32 qno, dma_addr_t addr)
+{
+	struct cldma_hw_regs *hw_regs;
+	unsigned int addr_l;
+	unsigned int addr_h;
+	int base;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX) {
+		addr_l = base + hw_regs->reg_cldma_ul_start_addrl_0 + qno * HW_QUEUE_NUM;
+		addr_h = base + hw_regs->reg_cldma_ul_start_addrh_0 + qno * HW_QUEUE_NUM;
+	} else {
+		addr_l = base + hw_regs->reg_cldma_so_start_addrl_0 + qno * HW_QUEUE_NUM;
+		addr_h = base + hw_regs->reg_cldma_so_start_addrh_0 + qno * HW_QUEUE_NUM;
+	}
+
+	mtk_pci_write32(drv_info->mdev, addr_l, (u32)addr);
+	mtk_pci_write32(drv_info->mdev, addr_h, (u32)((u64)addr >> 32));
+}
+
+void mtk_cldma_mask_intr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			 u32 qno, enum mtk_intr_type type)
+{
+	struct cldma_hw_regs *hw_regs;
+	int base;
+	u32 addr;
+	u32 val;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_l2timsr0;
+	else
+		addr = base + hw_regs->reg_cldma_l2rimsr0;
+
+	if (qno == ALLQ)
+		val = qno << type;
+	else
+		val = BIT(qno) << type;
+
+	mtk_pci_write32(drv_info->mdev, addr, val);
+}
+
+void mtk_cldma_unmask_intr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			   u32 qno, enum mtk_intr_type type)
+{
+	struct cldma_hw_regs *hw_regs;
+	int base;
+	u32 addr;
+	u32 val;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_l2timcr0;
+	else
+		addr = base + hw_regs->reg_cldma_l2rimcr0;
+
+	if (qno == ALLQ)
+		val = qno << type;
+	else
+		val = BIT(qno) << type;
+
+	mtk_pci_write32(drv_info->mdev, addr, val);
+}
+
+void mtk_cldma_clr_intr_status(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			       u32 qno, enum mtk_intr_type type)
+{
+	struct cldma_hw_regs *hw_regs;
+	struct mtk_md_dev *mdev;
+	int base;
+	u32 addr;
+	u32 val;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+	mdev = drv_info->mdev;
+
+	if (type == QUEUE_ERROR) {
+		if (dir == DIR_TX) {
+			val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l3tisar0);
+			mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l3tisar0, val);
+			val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l3tisar1);
+			mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l3tisar1, val);
+			val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l3tisar2);
+			mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l3tisar2, val);
+		} else {
+			val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l3risar0);
+			mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l3risar0, val);
+			val = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l3risar1);
+			mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l3risar1, val);
+		}
+	}
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_l2tisar0;
+	else
+		addr = base + hw_regs->reg_cldma_l2risar0;
+
+	if (qno == ALLQ)
+		val = qno << type;
+	else
+		val = BIT(qno) << type;
+
+	mtk_pci_write32(mdev, addr, val);
+	val = mtk_pci_read32(mdev, addr);
+}
+
+u32 mtk_cldma_check_intr_status(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+				u32 qno, enum mtk_intr_type type)
+{
+	struct cldma_hw_regs *hw_regs;
+	u32 addr, val, sta;
+	int base;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_l2tisar0;
+	else
+		addr = base + hw_regs->reg_cldma_l2risar0;
+
+	val = mtk_pci_read32(drv_info->mdev, addr);
+	if (val == LINK_ERROR_VAL)
+		sta = val;
+	else if (qno == ALLQ)
+		sta = (val >> type) & 0xFF;
+	else
+		sta = (val >> type) & BIT(qno);
+
+	return sta;
+}
+
+void mtk_cldma_start_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno)
+{
+	struct cldma_hw_regs *hw_regs;
+	u32 val = BIT(qno);
+	int base;
+	u32 addr;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_ul_start_cmd;
+	else
+		addr = base + hw_regs->reg_cldma_so_start_cmd;
+
+	mtk_pci_write32(drv_info->mdev, addr, val);
+}
+
+void mtk_cldma_resume_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno)
+{
+	struct cldma_hw_regs *hw_regs;
+	u32 val = BIT(qno);
+	int base;
+	u32 addr;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_ul_resume_cmd;
+	else
+		addr = base + hw_regs->reg_cldma_so_resume_cmd;
+
+	mtk_pci_write32(drv_info->mdev, addr, val);
+}
+
+static u32 mtk_cldma_queue_status(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno)
+{
+	struct cldma_hw_regs *hw_regs;
+	int base;
+	u32 addr;
+	u32 val;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_ul_status;
+	else
+		addr = base + hw_regs->reg_cldma_so_status;
+
+	val = mtk_pci_read32(drv_info->mdev, addr);
+
+	if (qno == ALLQ || val == LINK_ERROR_VAL)
+		return val;
+
+	return val & BIT(qno);
+}
+
+int mtk_cldma_stop_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno)
+{
+	u32 val = (qno == ALLQ) ? qno : BIT(qno);
+	struct cldma_hw_regs *hw_regs;
+	unsigned int active;
+	int base;
+	u32 addr;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+
+	if (dir == DIR_TX)
+		addr = base + hw_regs->reg_cldma_ul_stop_cmd;
+	else
+		addr = base + hw_regs->reg_cldma_so_stop_cmd;
+
+	mtk_pci_write32(drv_info->mdev, addr, val);
+
+	if (read_poll_timeout(mtk_cldma_queue_status, active,
+			      active == LINK_ERROR_VAL || !active,
+			      WAIT_QUEUE_STOP, WAIT_QUEUE_STOP * 10, false,
+			      drv_info, dir, qno))
+		return -ETIMEDOUT;
+
+	/* An all-ones read means the link is dead, not that the queue
+	 * stopped; let the caller tell the two apart.
+	 */
+	if (active == LINK_ERROR_VAL)
+		return -ENODEV;
+
+	return 0;
+}
+
+void mtk_cldma_clear_ip_busy(struct cldma_drv_info *drv_info)
+{
+	mtk_pci_write32(drv_info->mdev, drv_info->base_addr +
+			drv_info->hw_regs->reg_cldma_ip_busy, 0x01);
+}
+
+void mtk_cldma_get_intr_status(struct cldma_drv_info *drv_info, u32 *tx_sta, u32 *rx_sta)
+{
+	struct cldma_hw_regs *hw_regs;
+	struct mtk_md_dev *mdev;
+	u32 tx_mask, rx_mask;
+	int base;
+
+	mdev = drv_info->mdev;
+	base = drv_info->base_addr;
+	hw_regs = drv_info->hw_regs;
+
+	*tx_sta = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l2tisar0);
+	tx_mask = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l2timr0);
+	*rx_sta = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l2risar0);
+	rx_mask = mtk_pci_read32(mdev, base + hw_regs->reg_cldma_l2rimr0);
+
+	*tx_sta = (*tx_sta) & (~tx_mask);
+	*rx_sta = (*rx_sta) & (~rx_mask);
+
+	if (*tx_sta) {
+		/* TX XFER_DONE and QUEUE_ERROR mask */
+		mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l2timsr0, *tx_sta);
+		/* TX XFER_DONE clear */
+		mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l2tisar0,
+				(*tx_sta) & (0xFF << QUEUE_XFER_DONE));
+	}
+
+	if (*rx_sta) {
+		/* RX XFER_DONE and QUEUE_ERROR mask */
+		mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l2rimsr0, *rx_sta);
+		/* RX XFER_DONE clear */
+		mtk_pci_write32(mdev, base + hw_regs->reg_cldma_l2risar0,
+				(*rx_sta) & (0xFF << QUEUE_XFER_DONE));
+	}
+}
+
+u64 mtk_cldma_get_tx_start_addr(struct cldma_drv_info *drv_info, u32 qno)
+{
+	struct cldma_hw_regs *hw_regs = drv_info->hw_regs;
+	struct mtk_md_dev *mdev = drv_info->mdev;
+	int base = drv_info->base_addr;
+	u32 addr_h, addr_l;
+
+	addr_l = mtk_pci_read32(mdev,
+				base + hw_regs->reg_cldma_ul_start_addrl_0 + qno * HW_QUEUE_NUM);
+	addr_h = mtk_pci_read32(mdev,
+				base + hw_regs->reg_cldma_ul_start_addrh_0 + qno * HW_QUEUE_NUM);
+
+	return ((u64)addr_h << 32) | addr_l;
+}
+
+u64 mtk_cldma_get_rx_curr_addr(struct cldma_drv_info *drv_info, u32 qno)
+{
+	struct cldma_hw_regs *hw_regs;
+	u32 curr_addr_h, curr_addr_l;
+	struct mtk_md_dev *mdev;
+	u64 curr_addr;
+	int base;
+	u64 addr;
+
+	hw_regs = drv_info->hw_regs;
+	base = drv_info->base_addr;
+	mdev = drv_info->mdev;
+
+	addr = base + hw_regs->reg_cldma_so_current_addrh_0 +
+	       (u64)qno * HW_QUEUE_NUM;
+	curr_addr_h = mtk_pci_read32(mdev, addr);
+	addr = base + hw_regs->reg_cldma_so_current_addrl_0 +
+	       (u64)qno * HW_QUEUE_NUM;
+	curr_addr_l = mtk_pci_read32(mdev, addr);
+	curr_addr = ((u64)curr_addr_h << 32) | curr_addr_l;
+	if (curr_addr_h == LINK_ERROR_VAL && curr_addr_l == LINK_ERROR_VAL)
+		curr_addr = 0;
+	return curr_addr;
+}
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h
new file mode 100644
index 000000000000..d7abf4b48895
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h
@@ -0,0 +1,135 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2023, MediaTek Inc.
+ */
+
+#ifndef __MTK_CLDMA_DRV_H__
+#define __MTK_CLDMA_DRV_H__
+
+#define HW_QUEUE_NUM		(8)
+#define ALLQ			(0xFF)
+#define LINK_ERROR_VAL		(0xFFFFFFFF)
+#define CLDMA0_HW_ID		(0)
+#define CLDMA1_HW_ID		(1)
+
+struct cldma_hw_regs {
+	u8 cldma_rx_skb_pool_max_size;
+	u8 cldma_rx_skb_reload_threshold;
+	u8 tq_err_int_offset;
+	u8 tq_active_start_err_int_offset;
+	u8 rq_err_int_offset;
+	u8 rq_active_start_err_int_offset;
+	u16 reg_cldma_so_cfg;
+	u16 reg_cldma_so_start_addrl_0;
+	u16 reg_cldma_so_start_addrh_0;
+	u16 reg_cldma_so_current_addrl_0;
+	u16 reg_cldma_so_current_addrh_0;
+	u16 reg_cldma_so_status;
+	u16 reg_cldma_debug_id_en;
+	u16 reg_cldma_so_last_update_addrl_0;
+	u16 reg_cldma_so_last_update_addrh_0;
+	u16 reg_cldma_l2rimr0;
+	u16 reg_cldma_l2rimr1;
+	u16 reg_cldma_l2rimcr0;
+	u16 reg_cldma_l2rimcr1;
+	u16 reg_cldma_l2rimsr0;
+	u16 reg_cldma_l2rimsr1;
+	u16 reg_cldma_int_mask;
+	u16 reg_cldma_ip_busy_to_pcie_mask;
+	u16 reg_cldma_ip_busy_to_pcie_mask_set;
+	u16 reg_cldma_ip_busy_to_pcie_mask_clr;
+	u16 reg_cldma_ip_busy_to_ap_mask;
+	u16 reg_cldma_ip_busy_to_ap_mask_set;
+	u16 reg_cldma_ip_busy_to_ap_mask_clr;
+	u16 reg_cldma_ip_busy_to_md_mask_set;
+	u16 reg_cldma_rx_work_to_reg_mask_set;
+	u16 reg_infra_rst0_set;
+	u16 reg_infra_rst0_clr;
+	u32 tq_err_int_bitmask;
+	u32 tq_active_start_err_int_bitmask;
+	u32 rq_err_int_bitmask;
+	u32 cldma0_base_addr;
+	u32 cldma1_base_addr;
+	u32 rq_active_start_err_int_bitmask;
+	u32 reg_cldma_ul_start_addrl_0;
+	u32 reg_cldma_ul_start_addrh_0;
+	u32 reg_cldma_ul_current_addrl_0;
+	u32 reg_cldma_ul_current_addrh_0;
+	u32 reg_cldma_ul_status;
+	u32 reg_cldma_ul_start_cmd;
+	u32 reg_cldma_ul_resume_cmd;
+	u32 reg_cldma_ul_stop_cmd;
+	u32 reg_cldma_ul_error;
+	u32 reg_cldma_ul_cfg;
+	u32 reg_cldma_ul_dummy_0;
+	u32 reg_cldma_so_error;
+	u32 reg_cldma_so_start_cmd;
+	u32 reg_cldma_so_resume_cmd;
+	u32 reg_cldma_so_stop_cmd;
+	u32 reg_cldma_so_dummy_0;
+	u32 reg_cldma_l2tisar0;
+	u32 reg_cldma_l2tisar1;
+	u32 reg_cldma_l2timr0;
+	u32 reg_cldma_l2timr1;
+	u32 reg_cldma_l2timcr0;
+	u32 reg_cldma_l2timcr1;
+	u32 reg_cldma_l2timsr0;
+	u32 reg_cldma_l2timsr1;
+	u32 reg_cldma_l2risar0;
+	u32 reg_cldma_l2risar1;
+	u32 reg_cldma_l3tisar0;
+	u32 reg_cldma_l3tisar1;
+	u32 reg_cldma_l3tisar2;
+	u32 reg_cldma_l3risar0;
+	u32 reg_cldma_l3risar1;
+	u32 reg_cldma_ip_busy;
+};
+
+enum mtk_intr_type {
+	QUEUE_XFER_DONE = 0,
+	QUEUE_ERROR = 16,
+};
+
+enum mtk_tx_rx {
+	DIR_TX,
+	DIR_RX,
+};
+
+struct cldma_drv_info {
+	int hif_id;
+	int hw_id;
+	int base_addr;
+	int pci_ext_irq_id;
+	struct mtk_md_dev *mdev;
+	struct cldma_dev *cd;
+	struct txq *txq[HW_QUEUE_NUM];
+	struct rxq *rxq[HW_QUEUE_NUM];
+	struct dma_pool *gpd_dma_pool;
+	struct dma_pool *bd_dma_pool;
+	struct workqueue_struct *wq;
+	struct cldma_hw_regs *hw_regs;
+};
+
+extern struct cldma_hw_regs mtk_cldma_regs_m9xx;
+
+void mtk_cldma_drv_init(struct cldma_drv_info *drv_info);
+void mtk_cldma_drv_reset(struct cldma_drv_info *drv_info);
+void mtk_cldma_setup_start_addr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+				u32 qno, dma_addr_t addr);
+void mtk_cldma_mask_intr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			 u32 qno, enum mtk_intr_type type);
+void mtk_cldma_unmask_intr(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			   u32 qno, enum mtk_intr_type type);
+void mtk_cldma_clr_intr_status(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+			       u32 qno, enum mtk_intr_type type);
+u32 mtk_cldma_check_intr_status(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir,
+				u32 qno, enum mtk_intr_type type);
+void mtk_cldma_start_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno);
+void mtk_cldma_resume_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno);
+int mtk_cldma_stop_queue(struct cldma_drv_info *drv_info, enum mtk_tx_rx dir, u32 qno);
+void mtk_cldma_clear_ip_busy(struct cldma_drv_info *drv_info);
+void mtk_cldma_get_intr_status(struct cldma_drv_info *drv_info, u32 *tx_sta, u32 *rx_sta);
+u64 mtk_cldma_get_tx_start_addr(struct cldma_drv_info *drv_info, u32 qno);
+u64 mtk_cldma_get_rx_curr_addr(struct cldma_drv_info *drv_info, u32 qno);
+
+#endif
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
index c2c2b824ff31..6c31a5d80693 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_pci.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
@@ -872,6 +872,29 @@ static void mtk_pci_free_irq(struct mtk_md_dev *mdev)
 	pci_free_irq_vectors(pdev);
 }
 
+static int mtk_pci_dev_init(struct mtk_md_dev *mdev)
+{
+	int ret;
+
+	ret = mtk_trans_ctrl_init(mdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to initialize control plane: %d\n", ret);
+		return ret;
+	}
+
+	return 0;
+}
+
+static void mtk_pci_dev_exit(struct mtk_md_dev *mdev)
+{
+	mtk_trans_ctrl_exit(mdev);
+}
+
+static int mtk_pci_dev_start(struct mtk_md_dev *mdev)
+{
+	return 0;
+}
+
 static int mtk_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 {
 	struct device *dev = &pdev->dev;
@@ -938,6 +961,12 @@ static int mtk_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	if (ret)
 		goto free_mhccif;
 
+	ret = mtk_pci_dev_init(mdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to init dev.\n");
+		goto free_irq;
+	}
+
 	pci_set_master(pdev);
 	mtk_pci_unmask_irq(mdev, priv->mhccif_irq_id);
 
@@ -958,10 +987,20 @@ static int mtk_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 		goto clear_master;
 	}
 
+	ret = mtk_pci_dev_start(mdev);
+	if (ret) {
+		dev_err(mdev->dev, "Failed to start dev.\n");
+		goto free_saved_state;
+	}
+
 	return 0;
 
+free_saved_state:
+	pci_load_and_free_saved_state(pdev, &priv->saved_state);
 clear_master:
 	pci_clear_master(pdev);
+	mtk_pci_dev_exit(mdev);
+free_irq:
 	mtk_pci_free_irq(mdev);
 free_mhccif:
 	mtk_mhccif_exit(mdev);
@@ -980,6 +1019,8 @@ static void mtk_pci_remove(struct pci_dev *pdev)
 	struct device *dev = &pdev->dev;
 	int ret;
 
+	mtk_pci_dev_exit(mdev);
+
 	/* Silence every source before tearing anything down. */
 	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, U32_MAX);
 
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h b/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h
index a97ad6630dcf..f65a8065d945 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci_reg.h
@@ -16,6 +16,7 @@
 #define REG_IMASK_HOST_MSIX_SET_GRP0_0		0x3000
 #define REG_IMASK_HOST_MSIX_CLR_GRP0_0		0x3080
 #define REG_IMASK_HOST_MSIX_GRP0_0		0x3100
+#define REG_DEV_INFRA_BASE			0x10001000
 
 /* mhccif registers */
 #define MHCCIF_RC2EP_SW_BSY			0x4
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
new file mode 100644
index 000000000000..531940b06e2a
--- /dev/null
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
@@ -0,0 +1,615 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#include <linux/device.h>
+#include <linux/freezer.h>
+#include <linux/hashtable.h>
+#include <linux/kthread.h>
+#include <linux/list.h>
+#include <linux/nospec.h>
+#include <linux/sched.h>
+#include <linux/wait.h>
+
+#include "mtk_cldma.h"
+#include "mtk_ctrl_plane.h"
+#include "mtk_dev.h"
+#include "mtk_pci.h"
+#include "mtk_trans_ctrl.h"
+
+#define QUEUE_CHL_MASK	0xFFFF
+#define TRB_SRV_NUM	(1)
+
+static const int mtk_srv_cfg[NR_CLDMA][HW_QUE_NUM] = {
+	{0},
+	{0},
+};
+
+/* the number of RX GPDs should be at least two */
+static const struct queue_info mtk_queue_info[] = {
+	{CCCI_CONTROL_TX, CCCI_CONTROL_RX, CLDMA1, TXQ(0), RXQ(0),
+	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
+	{CCCI_SAP_CONTROL_TX, CCCI_SAP_CONTROL_RX, CLDMA0, TXQ(0), RXQ(0),
+	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
+};
+
+static bool mtk_queue_list_is_full(struct mtk_ctrl_trans *trans, struct queue_info *que)
+{
+	return skb_queue_len_lockless(&trans->trans_list[que->hif_id].skb_list[que->txqno]) >=
+	       SKB_LIST_MAX_LEN;
+}
+
+static bool mtk_ctrl_chs_is_busy_or_empty(struct trb_srv *srv)
+{
+	struct srv_que *srv_que;
+	int i;
+
+	for (i = 0; i < NR_CLDMA; i++) {
+		list_for_each_entry(srv_que, &srv->srv_q_list[i], list) {
+			struct sk_buff *skb;
+			struct trb *trb;
+
+			skb = skb_peek(&srv->trans->trans_list[i].skb_list[srv_que->qno]);
+			if (!skb)
+				continue;
+
+			/* ENABLE and DISABLE are software-only and are queued at
+			 * the head, so gating them on TX budget would make a queue
+			 * that cannot drain impossible to close.
+			 */
+			trb = (struct trb *)skb->cb;
+			if (trb->cmd != TRB_CMD_TX ||
+			    mtk_cldma_get_tx_budget(srv->trans->dev, i, srv_que->qno))
+				return false;
+		}
+	}
+
+	return true;
+}
+
+static void mtk_ctrl_ch_flush(struct sk_buff_head *skb_list)
+{
+	struct sk_buff *skb;
+	struct trb *trb;
+
+	while (!skb_queue_empty(skb_list)) {
+		skb = skb_dequeue(skb_list);
+		trb = (struct trb *)skb->cb;
+		trb->status = -EIO;
+		trb->trb_complete(skb);
+	}
+}
+
+static void mtk_ctrl_chs_flush(struct trb_srv *srv)
+{
+	struct srv_que *srv_que;
+	int i;
+
+	for (i = 0; i < NR_CLDMA; i++)
+		list_for_each_entry(srv_que, &srv->srv_q_list[i], list)
+			mtk_ctrl_ch_flush(&srv->trans->trans_list[i].skb_list[srv_que->qno]);
+}
+
+static int mtk_ch_status_check(struct mtk_ctrl_trans *trans, struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct trb_open_priv *trb_open_priv;
+	struct queue_info *que;
+	int ret = 0;
+
+	que = radix_tree_lookup(&trans->queue_tbl, trb->channel_id & QUEUE_CHL_MASK);
+
+	switch (trb->cmd) {
+	case TRB_CMD_ENABLE:
+		trb_open_priv = (struct trb_open_priv *)skb->data;
+		trb_open_priv->log_rg_offset = que->log_rg_offset;
+		trans->usr_cnt[que->hif_id][que->txqno]++;
+		if (trans->usr_cnt[que->hif_id][que->txqno] == 1)
+			break;
+		trb_open_priv->tx_mtu = que->tx_mtu;
+		trb_open_priv->rx_mtu = que->rx_mtu;
+		trb_open_priv->tx_frag_size = que->tx_frag_size;
+		trb_open_priv->rx_frag_size = que->rx_frag_size;
+		if (mtk_cldma_check_ch_cfg(trans->dev, que)) {
+			trb->status = -EINVAL;
+			ret = -EINVAL;
+		} else {
+			trb->status = -EBUSY;
+			ret = -EBUSY;
+		}
+		trb->trb_complete(skb);
+		break;
+	case TRB_CMD_DISABLE:
+		if (trans->usr_cnt[que->hif_id][que->txqno] > 0) {
+			trans->usr_cnt[que->hif_id][que->txqno]--;
+			if (!trans->usr_cnt[que->hif_id][que->txqno])
+				break;
+		}
+		trb->status = -EBUSY;
+		trb->trb_complete(skb);
+		ret = -EBUSY;
+		break;
+	default:
+		dev_err((trans->mdev)->dev, "Invalid trb command(%d)\n", trb->cmd);
+		ret = -EINVAL;
+		break;
+	}
+	return ret;
+}
+
+/* Single consumer per srv_que — only this kthread dequeues from skb_list.
+ * The list lock is held around every list read (peek, is_last, peek_next,
+ * unlink) so a producer inserting a DISABLE at the head cannot race the
+ * traversal, but it is dropped across submit and dispatch: those paths
+ * may allocate with GFP_KERNEL and thus sleep. Dropping the lock there
+ * is safe because no other consumer can steal the peeked skb.
+ */
+static void mtk_ctrl_trb_handler(struct trb_srv *srv, struct trans_list *trans_list, u32 qno)
+{
+	struct sk_buff_head *skb_list = &trans_list->skb_list[qno];
+	struct mtk_ctrl_trans *trans = srv->trans;
+	struct sk_buff *skb, *skb_next;
+	struct trb *trb, *trb_next;
+	unsigned long flags;
+	bool kick = false;
+	int loop = 0;
+	int err;
+
+	do {
+		spin_lock_irqsave(&skb_list->lock, flags);
+		skb = skb_peek(skb_list);
+		if (!skb) {
+			spin_unlock_irqrestore(&skb_list->lock, flags);
+			break;
+		}
+		trb = (struct trb *)skb->cb;
+
+		switch (trb->cmd) {
+		case TRB_CMD_ENABLE:
+		case TRB_CMD_DISABLE:
+			__skb_unlink(skb, skb_list);
+			spin_unlock_irqrestore(&skb_list->lock, flags);
+			err = mtk_ch_status_check(trans, skb);
+			if (!err) {
+				kick = true;
+				if (trb->cmd == TRB_CMD_DISABLE)
+					mtk_ctrl_ch_flush(skb_list);
+			}
+			break;
+		case TRB_CMD_TX:
+			spin_unlock_irqrestore(&skb_list->lock, flags);
+			err = mtk_cldma_submit_tx(trans->dev, skb);
+			if (err) {
+				if (trans_list->tx_burst_cnt[qno]) {
+					kick = true;
+					break;
+				}
+				if (err == -EAGAIN)
+					return;
+
+				skb_unlink(skb, skb_list);
+				trb->status = err;
+				trb->trb_complete(skb);
+				break;
+			}
+
+			trans_list->tx_burst_cnt[qno]++;
+			spin_lock_irqsave(&skb_list->lock, flags);
+			if (trans_list->tx_burst_cnt[qno] >= TX_BURST_MAX_CNT ||
+			    skb_queue_is_last(skb_list, skb)) {
+				kick = true;
+			} else {
+				skb_next = skb_peek_next(skb, skb_list);
+				trb_next = (struct trb *)skb_next->cb;
+				if (trb_next->cmd != TRB_CMD_TX)
+					kick = true;
+			}
+
+			__skb_unlink(skb, skb_list);
+			spin_unlock_irqrestore(&skb_list->lock, flags);
+			break;
+		default:
+			__skb_unlink(skb, skb_list);
+			spin_unlock_irqrestore(&skb_list->lock, flags);
+			trb->status = -EINVAL;
+			trb->trb_complete(skb);
+			break;
+		}
+
+		if (kick) {
+			err = mtk_cldma_trb_process(trans->dev, skb);
+			if (err)
+				dev_err_ratelimited((trans->mdev)->dev,
+						    "Failed to process trb on queue %u: %d\n",
+						    qno, err);
+			trans_list->tx_burst_cnt[qno] = 0;
+			kick = false;
+		}
+
+		loop++;
+	} while (loop < TRB_NUM_PER_ROUND);
+}
+
+static void mtk_ctrl_trb_process(struct trb_srv *srv)
+{
+	struct mtk_ctrl_trans *trans = srv->trans;
+	struct srv_que *srv_que;
+	int i;
+
+	for (i = 0; i < NR_CLDMA; i++)
+		list_for_each_entry(srv_que, &srv->srv_q_list[i], list)
+			mtk_ctrl_trb_handler(srv, &trans->trans_list[i], srv_que->qno);
+}
+
+static int mtk_ctrl_trb_thread(void *args)
+{
+	struct trb_srv *srv = args;
+
+	for (;;) {
+		wait_event_interruptible(srv->trb_waitq,
+					 !mtk_ctrl_chs_is_busy_or_empty(srv) ||
+					 kthread_should_stop() || kthread_should_park());
+		if (kthread_should_stop())
+			break;
+
+		if (kthread_should_park())
+			kthread_parkme();
+
+		do {
+			mtk_ctrl_trb_process(srv);
+			cond_resched();
+		} while (!mtk_ctrl_chs_is_busy_or_empty(srv) && !kthread_should_stop() &&
+			 !kthread_should_park());
+	}
+	mtk_ctrl_chs_flush(srv);
+	return 0;
+}
+
+static int mtk_ctrl_trb_srv_init(struct mtk_ctrl_trans *trans)
+{
+	struct srv_que *srv_que;
+	struct trb_srv *srv;
+	int i, j;
+	int ret;
+
+	for (i = 0; i < trans->trb_srv_num; i++) {
+		srv = kzalloc_obj(*srv);
+		if (!srv) {
+			ret = -ENOMEM;
+			goto err_free_srv;
+		}
+
+		srv->trans = trans;
+		srv->srv_id = i;
+		trans->trb_srv[i] = srv;
+
+		init_waitqueue_head(&srv->trb_waitq);
+		for (j = 0; j < NR_CLDMA; j++)
+			INIT_LIST_HEAD(&srv->srv_q_list[j]);
+	}
+
+	for (i = 0; i < NR_CLDMA; i++)
+		for (j = 0; j < HW_QUE_NUM; j++) {
+			if (trans->srv_cfg[i][j] < 0 ||
+			    trans->srv_cfg[i][j] >= trans->trb_srv_num)
+				trans->srv_cfg[i][j] = 0;
+			srv_que = kzalloc_obj(*srv_que);
+			if (!srv_que) {
+				ret = -ENOMEM;
+				goto err_free_srv_que;
+			}
+			srv_que->hif_id = i;
+			srv_que->qno = j;
+			list_add_tail(&srv_que->list,
+				      &trans->trb_srv[trans->srv_cfg[i][j]]->srv_q_list[i]);
+		}
+
+	for (i = 0; i < trans->trb_srv_num; i++) {
+		trans->trb_srv[i]->trb_thread = kthread_run(mtk_ctrl_trb_thread, trans->trb_srv[i],
+							    "mtk_trb_srv%d_%s", i,
+							    trans->mdev->dev_str);
+		if (IS_ERR(trans->trb_srv[i]->trb_thread)) {
+			ret = PTR_ERR(trans->trb_srv[i]->trb_thread);
+			trans->trb_srv[i]->trb_thread = NULL;
+			goto err_stop_kthread;
+		}
+	}
+
+	return 0;
+err_stop_kthread:
+	while (--i >= 0)
+		kthread_stop(trans->trb_srv[i]->trb_thread);
+err_free_srv_que:
+	for (i = 0; i < trans->trb_srv_num; i++) {
+		for (j = 0; j < NR_CLDMA; j++) {
+			struct srv_que *next_srv_que;
+
+			list_for_each_entry_safe(srv_que, next_srv_que,
+						 &trans->trb_srv[i]->srv_q_list[j], list) {
+				list_del(&srv_que->list);
+				kfree(srv_que);
+			}
+		}
+	}
+err_free_srv:
+	for (i = 0; i < trans->trb_srv_num; i++) {
+		if (!trans->trb_srv[i])
+			break;
+		kfree(trans->trb_srv[i]);
+		trans->trb_srv[i] = NULL;
+	}
+
+	return ret;
+}
+
+static void mtk_ctrl_trb_srv_exit(struct mtk_ctrl_trans *trans)
+{
+	struct srv_que *srv_que, *next_srv_que;
+	struct trb_srv *srv;
+	int i, j;
+
+	for (i = 0; i < trans->trb_srv_num; i++) {
+		srv = trans->trb_srv[i];
+		if (!srv)
+			continue;
+		kthread_stop(srv->trb_thread);
+		for (j = 0; j < NR_CLDMA; j++) {
+			list_for_each_entry_safe(srv_que, next_srv_que,
+						 &trans->trb_srv[i]->srv_q_list[j], list) {
+				list_del(&srv_que->list);
+				kfree(srv_que);
+			}
+		}
+		kfree(srv);
+		trans->trb_srv[i] = NULL;
+	}
+}
+
+static void mtk_ctrl_remove_radix_tree(struct mtk_ctrl_trans *trans)
+{
+	struct radix_tree_iter iter;
+	struct queue_info *queue;
+	void __rcu **slot;
+
+	radix_tree_for_each_slot(slot, &trans->queue_tbl, &iter, 0) {
+		queue = radix_tree_deref_slot(slot);
+		if (!queue)
+			continue;
+		radix_tree_delete(&trans->queue_tbl, iter.index);
+		kfree(queue);
+	}
+}
+
+int mtk_pcie_hif_init(struct mtk_md_dev *mdev)
+{
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+	struct queue_info *queue, *queue_info;
+	struct mtk_ctrl_trans *trans;
+	int i, j;
+	int ret;
+
+	trans = ctrl_blk->ctrl_hw_priv;
+	trans->ctrl_blk = ctrl_blk;
+	queue_info = trans->queue_info;
+
+	INIT_RADIX_TREE(&trans->queue_tbl, GFP_KERNEL);
+	for (i = 0; i < trans->queue_info_num; i++) {
+		queue = kmemdup(queue_info + i, sizeof(*queue), GFP_KERNEL);
+		if (!queue) {
+			ret = -ENOMEM;
+			goto err_free_radix_tree;
+		}
+		if (queue->txqno >= HW_QUE_NUM || queue->rxqno >= HW_QUE_NUM ||
+		    queue->hif_id >= NR_CLDMA) {
+			dev_err(mdev->dev, "Failed to get correct queue info %x\n",
+				queue->rx_chl);
+			kfree(queue);
+			ret = -EINVAL;
+			goto err_free_radix_tree;
+		}
+		ret = radix_tree_insert(&trans->queue_tbl, queue->rx_chl & QUEUE_CHL_MASK, queue);
+		if (ret) {
+			dev_err(mdev->dev, "Insert %x fail, ret: %d", queue->rx_chl, ret);
+			kfree(queue);
+			goto err_free_radix_tree;
+		}
+	}
+
+	for (i = 0; i < NR_CLDMA; i++) {
+		for (j = 0; j < HW_QUE_NUM; j++) {
+			skb_queue_head_init(&trans->trans_list[i].skb_list[j]);
+			trans->trans_list[i].tx_burst_cnt[j] = 0;
+			/* usr_cnt tracks the queues rebuilt by mtk_cldma_init()
+			 * below, so it must be reset with them. Otherwise a
+			 * count left over from a torn-down cycle makes the
+			 * channel permanently unopenable.
+			 */
+			trans->usr_cnt[i][j] = 0;
+		}
+	}
+	ret = mtk_cldma_init(trans);
+	if (ret)
+		goto err_free_radix_tree;
+
+	ret = mtk_ctrl_trb_srv_init(trans);
+	if (ret)
+		goto err_cldma_exit;
+
+	atomic_set(&trans->available, 1);
+
+	return 0;
+
+err_cldma_exit:
+	mtk_cldma_exit(trans);
+err_free_radix_tree:
+	mtk_ctrl_remove_radix_tree(trans);
+
+	return ret;
+}
+
+int mtk_pcie_hif_exit(struct mtk_md_dev *mdev)
+{
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+	struct mtk_ctrl_trans *trans;
+
+	trans = ctrl_blk->ctrl_hw_priv;
+
+	mutex_lock(&trans->submit_lock);
+	atomic_set(&trans->available, 0);
+	mutex_unlock(&trans->submit_lock);
+
+	/* Join the service threads before freeing what they dereference: they
+	 * read trans->dev without a lock, so clearing it cannot stop a
+	 * consumer that has already loaded the pointer.
+	 */
+	mtk_ctrl_trb_srv_exit(trans);
+	mtk_cldma_exit(trans);
+
+	/* Late submitters may still hold the lock and walk the tree. */
+	mutex_lock(&trans->submit_lock);
+	mtk_ctrl_remove_radix_tree(trans);
+	mutex_unlock(&trans->submit_lock);
+
+	return 0;
+}
+
+int mtk_pcie_hif_submit_skb(struct mtk_md_dev *mdev, struct sk_buff *skb, bool force_send)
+{
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+	struct mtk_ctrl_trans *trans;
+	struct queue_info *que;
+	struct trb *trb;
+	int ret;
+
+	trans = ctrl_blk->ctrl_hw_priv;
+	trb = (struct trb *)skb->cb;
+
+	if (trb->cmd == TRB_CMD_STOP || trb->cmd == TRB_CMD_RECOVER) {
+		trb->trb_complete(skb);
+		return 0;
+	}
+
+	mutex_lock(&trans->submit_lock);
+
+	if (!atomic_read(&trans->available)) {
+		ret = -EIO;
+		goto unlock;
+	}
+
+	que = radix_tree_lookup(&trans->queue_tbl, trb->channel_id & QUEUE_CHL_MASK);
+	if (!que) {
+		dev_warn(mdev->dev, "lookup que fail, ch_id: %x\n",
+			 trb->channel_id);
+		ret = -EINVAL;
+		goto unlock;
+	}
+
+	if (mtk_queue_list_is_full(trans, que) && !force_send) {
+		ret = -EAGAIN;
+		goto unlock;
+	}
+
+	if (trb->cmd == TRB_CMD_DISABLE) {
+		struct sk_buff *entry = NULL;
+		struct sk_buff_head *list;
+		struct sk_buff *iter;
+		unsigned long flags;
+
+		/* A disable may overtake queued data, so teardown does not
+		 * wait for a TX backlog, but it must never overtake a pending
+		 * ENABLE for the same queue: the two do not commute, and a
+		 * DISABLE consumed before its ENABLE closes nothing while the
+		 * ENABLE then arms rings nobody owns.
+		 */
+		list = &trans->trans_list[que->hif_id].skb_list[que->txqno];
+		spin_lock_irqsave(&list->lock, flags);
+		skb_queue_walk(list, iter) {
+			if (((struct trb *)iter->cb)->cmd != TRB_CMD_ENABLE) {
+				entry = iter;
+				break;
+			}
+		}
+		if (entry)
+			__skb_queue_before(list, entry, skb);
+		else
+			__skb_queue_tail(list, skb);
+		spin_unlock_irqrestore(&list->lock, flags);
+	} else {
+		skb_queue_tail(&trans->trans_list[que->hif_id].skb_list[que->txqno], skb);
+	}
+
+	wake_up(&trans->trb_srv[trans->srv_cfg[que->hif_id][que->txqno]]->trb_waitq);
+	ret = 0;
+
+unlock:
+	mutex_unlock(&trans->submit_lock);
+	return ret;
+}
+
+int mtk_pcie_hif_cmd_func(struct mtk_md_dev *mdev, int cmd, void *data)
+{
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+	struct mtk_ctrl_trans *trans;
+	struct queue_info *que;
+	int ret;
+
+	switch (cmd) {
+	case HIF_CTRL_CMD_CHECK_TX_FULL:
+		trans = ctrl_blk->ctrl_hw_priv;
+		mutex_lock(&trans->submit_lock);
+		if (!atomic_read(&trans->available)) {
+			ret = -EIO;
+			break;
+		}
+		que = radix_tree_lookup(&trans->queue_tbl,
+					((union ctrl_hif_cmd_data *)data)->rx_ch & QUEUE_CHL_MASK);
+		if (!que) {
+			dev_warn(mdev->dev, "Failed to find que to check tx full\n");
+			ret = -EINVAL;
+			break;
+		}
+		ret = mtk_queue_list_is_full(trans, que);
+		break;
+	default:
+		return -EINVAL;
+	}
+	mutex_unlock(&trans->submit_lock);
+
+	return ret;
+}
+
+int mtk_trans_ctrl_init(struct mtk_md_dev *mdev)
+{
+	struct mtk_ctrl_trans *trans;
+	struct mtk_ctrl_blk *ctrl_blk;
+	int err;
+
+	trans = devm_kzalloc(mdev->dev, sizeof(*trans), GFP_KERNEL);
+	if (!trans)
+		return -ENOMEM;
+	trans->mdev = mdev;
+	mutex_init(&trans->submit_lock);
+	atomic_set(&trans->available, 0);
+
+	memcpy(trans->srv_cfg, mtk_srv_cfg, sizeof(mtk_srv_cfg));
+	trans->queue_info = (struct queue_info *)mtk_queue_info;
+	trans->queue_info_num = ARRAY_SIZE(mtk_queue_info);
+	trans->trb_srv_num = TRB_SRV_NUM;
+
+	err = mtk_ctrl_init(mdev);
+	if (err)
+		return err;
+
+	ctrl_blk = mdev->ctrl_blk;
+	ctrl_blk->ctrl_hw_priv = trans;
+
+	return 0;
+}
+
+int mtk_trans_ctrl_exit(struct mtk_md_dev *mdev)
+{
+	mtk_ctrl_exit(mdev);
+
+	return 0;
+}
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
index d6de4c43b529..f9144c60edf3 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
@@ -8,14 +8,82 @@
 
 #include <linux/kref.h>
 #include <linux/list.h>
+#include <linux/mutex.h>
 #include <linux/skbuff.h>
 #include <linux/types.h>
 
 #include "mtk_dev.h"
 
+#define TRB_SRV_MAX_NUM			(1)
+#define HW_QUE_NUM			(8)
+#define TX_GPD_NUM			(16)
+#define RX_GPD_NUM			(TX_GPD_NUM)
+#define MIN_GPD_NUM			(2)
+#define SKB_LIST_MAX_LEN		(16)
+#define TRB_NUM_PER_ROUND		(TX_GPD_NUM)
+#define TX_BURST_MAX_CNT		(TX_GPD_NUM / 4 + 1)
+
+enum mtk_hif_id {
+	CLDMA0,
+	CLDMA1,
+	NR_CLDMA
+};
+
+struct queue_info {
+	u32 tx_chl;
+	u32 rx_chl;
+	enum mtk_hif_id hif_id;
+	u32 txqno;
+	u32 rxqno;
+	u32 tx_mtu;
+	u32 rx_mtu;
+	u32 tx_nr_gpds;
+	u32 rx_nr_gpds;
+	u32 tx_frag_size;
+	u32 rx_frag_size;
+	u8 log_rg_offset;
+};
+
+struct trans_list {
+	struct sk_buff_head skb_list[HW_QUE_NUM];
+	u8 tx_burst_cnt[HW_QUE_NUM];
+};
+
 struct mtk_ctrl_trans {
 	struct mtk_ctrl_blk *ctrl_blk;
+	struct trb_srv *trb_srv[TRB_SRV_MAX_NUM];
+	struct trans_list trans_list[NR_CLDMA];
+	void *dev;
+	struct radix_tree_root queue_tbl;
 	struct mtk_md_dev *mdev;
+	int usr_cnt[NR_CLDMA][HW_QUE_NUM];
+	struct mutex submit_lock; /* protects available flag and submit/exit paths */
+	atomic_t available;
+	int srv_cfg[NR_CLDMA][HW_QUE_NUM];
+	struct queue_info *queue_info;
+	int queue_info_num;
+	int trb_srv_num;
+};
+
+struct srv_que {
+	u32 hif_id;
+	u32 qno;
+	struct list_head list;
 };
 
+struct trb_srv {
+	u32 srv_id;
+	struct list_head srv_q_list[NR_CLDMA];
+	struct mtk_ctrl_trans *trans;
+	wait_queue_head_t trb_waitq;
+	struct task_struct *trb_thread;
+};
+
+int mtk_pcie_hif_init(struct mtk_md_dev *mdev);
+int mtk_pcie_hif_exit(struct mtk_md_dev *mdev);
+int mtk_pcie_hif_submit_skb(struct mtk_md_dev *mdev, struct sk_buff *skb, bool force_send);
+int mtk_pcie_hif_cmd_func(struct mtk_md_dev *mdev, int cmd, void *data);
+int mtk_trans_ctrl_init(struct mtk_md_dev *mdev);
+int mtk_trans_ctrl_exit(struct mtk_md_dev *mdev);
+
 #endif

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 4/6] net: wwan: t9xx: Add control port
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
                   ` (2 preceding siblings ...)
  2026-09-30  7:46 ` [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  2026-09-30  7:46 ` [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread Jack Wu via B4 Relay
  2026-09-30  7:46 ` [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports Jack Wu via B4 Relay
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

The control port consists of port I/O and port manager.
Port I/O provides a common operation as defined by "struct port_ops",
and the operation is managed by the "port manager". It provides
interfaces to internal users, the implemented internal interfaces are
open, close, write and recv_register.

The port manager defines and implements port management interfaces and
structures. It is responsible for port creation, destroying, and managing
port states. It sends data from port I/O to CLDMA via TRB ( Transaction
Request Block ), and dispatches received data from CLDMA to port I/O.

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 drivers/net/wwan/t9xx/Makefile              |   2 +
 drivers/net/wwan/t9xx/mtk_ctrl_plane.c      |  18 +-
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      |  10 +-
 drivers/net/wwan/t9xx/mtk_port.c            | 693 ++++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/mtk_port.h            | 131 ++++++
 drivers/net/wwan/t9xx/mtk_port_io.c         | 258 +++++++++++
 drivers/net/wwan/t9xx/mtk_port_io.h         |  35 ++
 drivers/net/wwan/t9xx/pcie/mtk_pci.c        |  25 +-
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c |  22 +-
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h |   1 +
 10 files changed, 1187 insertions(+), 8 deletions(-)

diff --git a/drivers/net/wwan/t9xx/Makefile b/drivers/net/wwan/t9xx/Makefile
index 74bef150e3a1..d44266009362 100644
--- a/drivers/net/wwan/t9xx/Makefile
+++ b/drivers/net/wwan/t9xx/Makefile
@@ -7,6 +7,8 @@ obj-$(CONFIG_MTK_T9XX) += mtk_t9xx.o
 
 mtk_t9xx-y := \
 	mtk_ctrl_plane.o \
+	mtk_port.o \
+	mtk_port_io.o \
 	pcie/mtk_pci.o \
 	pcie/mtk_trans_ctrl.o \
 	pcie/mtk_cldma.o \
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
index fa2ab8c3e757..10f0d3267b3b 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
@@ -7,19 +7,22 @@
 #include <linux/device.h>
 
 #include "mtk_ctrl_plane.h"
+#include "mtk_port.h"
 
 /**
  * mtk_ctrl_init() - Initialize the control plane block.
  * @mdev: Pointer to the MTK modem device.
+ * @port_layer_cfg: Port layer configuration.
  *
  * Allocates and initializes the control plane block
  * associated with @mdev.
  *
- * Return: 0 on success, -ENOMEM on allocation failure.
+ * Return: 0 on success, negative error code on failure.
  */
-int mtk_ctrl_init(struct mtk_md_dev *mdev)
+int mtk_ctrl_init(struct mtk_md_dev *mdev, struct mtk_port_layer_cfg *port_layer_cfg)
 {
 	struct mtk_ctrl_blk *ctrl_blk;
+	int err;
 
 	ctrl_blk = devm_kzalloc(mdev->dev, sizeof(*ctrl_blk), GFP_KERNEL);
 	if (!ctrl_blk)
@@ -28,7 +31,15 @@ int mtk_ctrl_init(struct mtk_md_dev *mdev)
 	ctrl_blk->mdev = mdev;
 	mdev->ctrl_blk = ctrl_blk;
 
+	err = mtk_port_mngr_init(ctrl_blk, port_layer_cfg->port_cfg,
+				 port_layer_cfg->port_cnt);
+	if (err)
+		goto err_free_mem;
+
 	return 0;
+
+err_free_mem:
+	return err;
 }
 
 /**
@@ -40,5 +51,8 @@ int mtk_ctrl_init(struct mtk_md_dev *mdev)
  */
 void mtk_ctrl_exit(struct mtk_md_dev *mdev)
 {
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+
+	mtk_port_mngr_exit(ctrl_blk);
 	mdev->ctrl_blk = NULL;
 }
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
index 0dcc0ec2c45d..933ba6193b70 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
@@ -11,8 +11,8 @@
 
 #include "mtk_dev.h"
 
-#define Q_MTU_3_5K	(0xE00)
-#define Q_FRAG_3_5K	(0xE00)
+#define Q_MTU_3_5K			(0xE00)
+#define Q_FRAG_3_5K			(0xE00)
 
 enum mtk_ccci_ch {
 	/* to sAP */
@@ -59,12 +59,16 @@ union ctrl_hif_cmd_data {
 	u32 rx_ch;
 };
 
+struct mtk_port_layer_cfg;
+struct mtk_port_mngr;
+
 struct mtk_ctrl_blk {
 	struct mtk_md_dev *mdev;
+	struct mtk_port_mngr *port_mngr;
 	void *ctrl_hw_priv;
 };
 
-int mtk_ctrl_init(struct mtk_md_dev *mdev);
+int mtk_ctrl_init(struct mtk_md_dev *mdev, struct mtk_port_layer_cfg *port_layer_cfg);
 void mtk_ctrl_exit(struct mtk_md_dev *mdev);
 
 #endif /* __MTK_CTRL_PLANE_H__ */
diff --git a/drivers/net/wwan/t9xx/mtk_port.c b/drivers/net/wwan/t9xx/mtk_port.c
new file mode 100644
index 000000000000..a8ff06ab4de4
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_port.c
@@ -0,0 +1,693 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#include <linux/bitfield.h>
+#include <linux/device.h>
+#include <linux/err.h>
+#include <linux/kernel.h>
+#include <linux/netdevice.h>
+#include <linux/slab.h>
+#include <linux/wait.h>
+
+#include "mtk_port.h"
+#include "mtk_port_io.h"
+#include "mtk_trans_ctrl.h"
+
+#define MTK_DFLT_TRB_TIMEOUT		(5 * HZ)
+#define MTK_DFLT_TRB_STATUS		(0x1)
+#define MTK_TRB_HEADER_ADDED		(0xADDED)
+#define MTK_CHECK_RX_SEQ_MASK		(0x7fff)
+
+#define MTK_PORT_ENUM_VER		(0)
+#define MTK_PORT_ENUM_HEAD_PATTERN	(0x5a5a5a5a)
+#define MTK_PORT_ENUM_TAIL_PATTERN	(0xa5a5a5a5)
+
+#define MTK_PORT_SEARCH_FROM_RADIX_TREE(p, s) ({\
+	struct mtk_port *_p;			\
+	_p = rcu_dereference_raw(*(s));		\
+	if (!_p)				\
+		continue;			\
+	p = _p;					\
+})
+
+#define MTK_PORT_INTERNAL_NODE_CHECK(p, s, i) ({\
+	if (radix_tree_is_internal_node(p)) {	\
+		s = radix_tree_iter_retry(&(i));\
+		continue;			\
+	}					\
+})
+
+struct mtk_port_info {
+	__le16 channel;
+	__le16 reserved;
+} __packed;
+
+struct mtk_port_enum_msg {
+	__le32 head_pattern;
+	__le16 port_cnt;
+	__le16 version;
+	__le32 tail_pattern;
+	u8 data[];
+} __packed;
+
+/* mutex lock for the port refcount */
+DEFINE_MUTEX(port_mngr_grp_mtx);
+
+/* This function working always under mutex lock port_mngr_grp_mtx */
+void mtk_port_release(struct kref *port_kref)
+{
+	struct mtk_port *port;
+
+	port = container_of(port_kref, struct mtk_port, kref);
+	ports_ops[port->info.type]->exit(port);
+	kfree_rcu(port, rcu);
+}
+
+static int mtk_port_tbl_add(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
+{
+	int ret;
+
+	mutex_lock(&port_mngr->port_tbl_mtx);
+	ret = radix_tree_insert(&port_mngr->port_tbl[MTK_PORT_TBL_TYPE(port->info.rx_ch)],
+				port->info.rx_ch & 0xFFF, port);
+	if (!ret)
+		port_mngr->port_cnt++;
+	mutex_unlock(&port_mngr->port_tbl_mtx);
+
+	if (ret)
+		dev_err(port_mngr->ctrl_blk->mdev->dev,
+			"port(%s) add to port_tbl failed, return %d\n",
+			port->info.name, ret);
+
+	return ret;
+}
+
+static void mtk_port_tbl_del(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
+{
+	mutex_lock(&port_mngr->port_tbl_mtx);
+	radix_tree_delete(&port_mngr->port_tbl[MTK_PORT_TBL_TYPE(port->info.rx_ch)],
+			  port->info.rx_ch & 0xFFF);
+	port_mngr->port_cnt--;
+	mutex_unlock(&port_mngr->port_tbl_mtx);
+}
+
+static struct mtk_port *mtk_port_alloc_and_add(struct mtk_port_mngr *port_mngr,
+					       struct mtk_port_cfg *dflt_info)
+{
+	struct mtk_port *port;
+	int ret;
+
+	port = kzalloc_obj(*port, GFP_KERNEL);
+	if (!port) {
+		ret = -ENOMEM;
+		goto err_alloc_port;
+	}
+	memcpy(&port->info, dflt_info, sizeof(*dflt_info));
+
+	ret = mtk_port_tbl_add(port_mngr, port);
+	if (ret < 0) {
+		dev_err(port_mngr->ctrl_blk->mdev->dev,
+			"Failed to add port(%s) to port tbl\n", dflt_info->name);
+		goto err_free_port;
+	}
+
+	port->port_mngr = port_mngr;
+	ret = ports_ops[port->info.type]->init(port);
+	if (ret < 0) {
+		mtk_port_tbl_del(port_mngr, port);
+		goto err_free_port;
+	}
+
+	return port;
+
+err_free_port:
+	kfree(port);
+err_alloc_port:
+	return ERR_PTR(ret);
+}
+
+static void mtk_port_free(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
+{
+	mutex_lock(&port_mngr_grp_mtx);
+	mtk_port_tbl_del(port_mngr, port);
+	kref_put(&port->kref, mtk_port_release);
+	mutex_unlock(&port_mngr_grp_mtx);
+}
+
+static struct mtk_port *mtk_port_search_by_id(struct mtk_port_mngr *port_mngr, int rx_ch)
+{
+	int tbl_type = MTK_PORT_TBL_TYPE(rx_ch);
+
+	if (tbl_type < PORT_TBL_SAP || tbl_type >= PORT_TBL_MAX)
+		return NULL;
+
+	return radix_tree_lookup(&port_mngr->port_tbl[tbl_type], MTK_CH_ID(rx_ch));
+}
+
+struct mtk_port *mtk_port_search_by_name(struct mtk_port_mngr *port_mngr, char *name)
+{
+	int tbl_type = PORT_TBL_SAP;
+	struct radix_tree_iter iter;
+	struct mtk_port *port;
+	void __rcu **slot;
+
+	do {
+		radix_tree_for_each_slot(slot, &port_mngr->port_tbl[tbl_type], &iter, 0) {
+			MTK_PORT_SEARCH_FROM_RADIX_TREE(port, slot);
+			MTK_PORT_INTERNAL_NODE_CHECK(port, slot, iter);
+			if (!strncmp(port->info.name, name, MTK_DFLT_PORT_NAME_LEN))
+				return port;
+		}
+		tbl_type++;
+	} while (tbl_type < PORT_TBL_MAX);
+
+	return NULL;
+}
+
+static int mtk_port_tbl_create(struct mtk_port_mngr *port_mngr, struct mtk_port_cfg *cfg,
+			       const int port_cnt)
+{
+	struct mtk_port_cfg *dflt_port;
+	struct mtk_port *port;
+	int i;
+
+	INIT_RADIX_TREE(&port_mngr->port_tbl[PORT_TBL_SAP], GFP_KERNEL);
+	INIT_RADIX_TREE(&port_mngr->port_tbl[PORT_TBL_MD], GFP_KERNEL);
+
+	/* copy ports from static port cfg table */
+	for (i = 0; i < port_cnt; i++) {
+		dflt_port = cfg + i;
+		if (!mtk_port_search_by_id(port_mngr, dflt_port->rx_ch)) {
+			port = mtk_port_alloc_and_add(port_mngr, dflt_port);
+			if (IS_ERR(port))
+				return PTR_ERR(port);
+		}
+	}
+
+	return 0;
+}
+
+static void mtk_port_tbl_destroy(struct mtk_port_mngr *port_mngr)
+{
+	struct radix_tree_iter iter;
+	struct mtk_port *port;
+	void __rcu **slot;
+	int tbl_type;
+
+	tbl_type = PORT_TBL_SAP;
+	do {
+		radix_tree_for_each_slot(slot, &port_mngr->port_tbl[tbl_type], &iter, 0) {
+			MTK_PORT_SEARCH_FROM_RADIX_TREE(port, slot);
+			MTK_PORT_INTERNAL_NODE_CHECK(port, slot, iter);
+			ports_ops[port->info.type]->disable(port);
+		}
+
+		while (radix_tree_gang_lookup(&port_mngr->port_tbl[tbl_type],
+					      (void **)&port, 0, 1))
+			mtk_port_free(port_mngr, port);
+	} while (++tbl_type < PORT_TBL_MAX);
+}
+
+void mtk_port_trb_init(struct mtk_port *port, struct trb *trb, enum mtk_trb_cmd_type cmd,
+		       int (*trb_complete)(struct sk_buff *skb))
+{
+	kref_init(&trb->kref);
+	trb->channel_id = port->info.rx_ch;
+	trb->status = MTK_DFLT_TRB_STATUS;
+	trb->priv = port;
+	trb->cmd = cmd;
+	trb->trb_complete = trb_complete;
+}
+
+void mtk_port_trb_free(struct kref *trb_kref)
+{
+	struct trb *trb = container_of(trb_kref, struct trb, kref);
+	struct sk_buff *skb, *frag_skb, *next_skb;
+
+	skb = container_of((char *)trb, struct sk_buff, cb[0]);
+	/* Free frag_list for scatter gather TX */
+	if (trb->cmd == TRB_CMD_TX && skb_has_frag_list(skb)) {
+		frag_skb = skb_shinfo(skb)->frag_list;
+		while (frag_skb) {
+			next_skb = frag_skb->next;
+			frag_skb->next = NULL;
+			dev_kfree_skb_any(frag_skb);
+			frag_skb = next_skb;
+		}
+		skb_shinfo(skb)->frag_list = NULL;
+		skb->data_len = 0;
+	}
+	dev_kfree_skb_any(skb);
+}
+
+static int mtk_port_open_trb_complete(struct sk_buff *skb)
+{
+	struct trb_open_priv *trb_open_priv = (struct trb_open_priv *)skb->data;
+	struct trb *trb = (struct trb *)skb->cb;
+	struct mtk_port *port = trb->priv;
+
+	if (!trb->status) {
+		port->tx_mtu = trb_open_priv->tx_mtu;
+		port->rx_mtu = trb_open_priv->rx_mtu;
+		port->tx_frag_size = trb_open_priv->tx_frag_size;
+		port->rx_frag_size = trb_open_priv->rx_frag_size;
+		port->tx_mtu -= MTK_CCCI_H_ELEN;
+		port->rx_mtu -= MTK_CCCI_H_ELEN;
+	}
+
+	wake_up_all(&port->trb_wq);
+
+	kref_put(&trb->kref, mtk_port_trb_free);
+	return 0;
+}
+
+static int mtk_port_close_trb_complete(struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct mtk_port *port = trb->priv;
+
+	wake_up_all(&port->trb_wq);
+	wake_up_all(&port->rx_wq);
+	kref_put(&trb->kref, mtk_port_trb_free);
+
+	return 0;
+}
+
+static int mtk_port_tx_complete(struct sk_buff *skb)
+{
+	struct trb *trb = (struct trb *)skb->cb;
+	struct mtk_port *port = trb->priv;
+
+	if (trb->status < 0)
+		dev_warn(port->port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to send data: status:%d, port:%s\n",
+			 trb->status, port->info.name);
+
+	wake_up_all(&port->trb_wq);
+	kref_put(&trb->kref, mtk_port_trb_free);
+
+	return 0;
+}
+
+int mtk_port_status_check(struct mtk_port *port)
+{
+	if (!test_bit(PORT_S_ENABLE, &port->status))
+		return -ENODEV;
+
+	if (!test_bit(PORT_S_OPEN, &port->status) || test_bit(PORT_S_FLUSH, &port->status) ||
+	    !test_bit(PORT_S_WR, &port->status))
+		return -EBADF;
+
+	return 0;
+}
+
+int mtk_port_send_data(struct mtk_port *port, void *data, bool blocking, bool force_send)
+{
+	struct mtk_port_mngr *port_mngr;
+	struct sk_buff *skb = data;
+	struct trb *trb;
+	int ret, len;
+
+	port_mngr = port->port_mngr;
+
+	trb = (struct trb *)skb->cb;
+	mtk_port_trb_init(port, trb, TRB_CMD_TX, mtk_port_tx_complete);
+	len = skb->len;
+	kref_get(&trb->kref); /* kref count 1->2 */
+
+	/* add ccci header */
+	mtk_port_add_header(skb);
+	ret = mtk_port_status_check(port);
+	if (!ret)
+		ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb,
+					      force_send);
+
+	if (ret < 0) {
+		kref_put(&trb->kref, mtk_port_trb_free); /* kref count 2->1 */
+		kref_put(&trb->kref, mtk_port_trb_free); /* kref count 1->0 */
+		port->tx_seq--;
+		goto out;
+	}
+
+	if (!blocking) {
+		kref_put(&trb->kref, mtk_port_trb_free);
+		ret = len;
+		goto out;
+	}
+start_wait:
+
+	/* wait trb done, and no timeout in tx blocking mode */
+	ret = wait_event_interruptible_timeout(port->trb_wq,
+					       trb->status <= 0 ||
+					       test_bit(PORT_S_FLUSH, &port->status) ||
+					       !test_bit(PORT_S_WR, &port->status),
+					       MTK_DFLT_TRB_TIMEOUT);
+	if (!ret) {
+		goto start_wait;
+	} else if (ret == -ERESTARTSYS) {
+		ret = -EINTR;
+	} else if (ret > 0) {
+		if (test_bit(PORT_S_FLUSH, &port->status))
+			ret = len;
+		else
+			ret = (!trb->status) ? len : trb->status;
+	}
+	kref_put(&trb->kref, mtk_port_trb_free);
+
+out:
+	return ret;
+}
+
+static int mtk_port_check_rx_seq(struct mtk_port *port, struct mtk_ccci_header *ccci_h)
+{
+	u16 seq_num, assert_bit, channel;
+	struct mtk_md_dev *mdev;
+
+	seq_num = FIELD_GET(MTK_HDR_FLD_SEQ, le32_to_cpu(ccci_h->status));
+	assert_bit = FIELD_GET(MTK_HDR_FLD_AST, le32_to_cpu(ccci_h->status));
+	if (assert_bit && port->rx_seq &&
+	    ((seq_num - port->rx_seq) & MTK_CHECK_RX_SEQ_MASK) != 1) {
+		mdev = port->port_mngr->ctrl_blk->mdev;
+		channel = FIELD_GET(MTK_HDR_FLD_CHN, le32_to_cpu(ccci_h->status));
+		dev_warn(mdev->dev,
+			 "<ch: %04x> seq num out-of-order %d->%d, len(%u)\n",
+			 channel, seq_num, port->rx_seq,
+			 le32_to_cpu(ccci_h->packet_len));
+
+		port->rx_seq = seq_num;
+		return -EPROTO;
+	}
+
+	return 0;
+}
+
+static int mtk_port_rx_dispatch_frag_skb(struct mtk_port *port, struct sk_buff *skb)
+{
+	struct sk_buff *frag_skb, *frag_next;
+	int ret;
+
+	frag_skb = skb_shinfo(skb)->frag_list;
+	skb->len -= skb->data_len;
+	skb->data_len = 0;
+	skb_shinfo(skb)->frag_list = NULL;
+
+	ret = ports_ops[port->info.type]->recv(port, skb);
+	if (ret < 0) {
+		skb_shinfo(skb)->frag_list = frag_skb;
+		return ret;
+	}
+
+	while (frag_skb) {
+		frag_next = frag_skb->next;
+		if (!frag_skb->len) {
+			frag_skb->next = NULL;
+			dev_kfree_skb_any(frag_skb);
+			frag_skb = frag_next;
+			continue;
+		}
+		frag_skb->next = NULL;
+		ret = ports_ops[port->info.type]->recv(port, frag_skb);
+		if (ret < 0) {
+			frag_skb->next = frag_next;
+			while (frag_skb) {
+				frag_next = frag_skb->next;
+				frag_skb->next = NULL;
+				dev_kfree_skb_any(frag_skb);
+				frag_skb = frag_next;
+			}
+			return -EIO;
+		}
+		frag_skb = frag_next;
+	}
+
+	return 0;
+}
+
+static int mtk_port_rx_dispatch(struct sk_buff *skb, void *priv, bool force_recv)
+{
+	struct mtk_port_mngr *port_mngr;
+	struct mtk_ccci_header *ccci_h;
+	struct mtk_port *port = priv;
+	int ret = -EPROTO;
+	u16 channel;
+
+	if (!skb || !priv) {
+		pr_err("Invalid input value in rx dispatch\n");
+		return -EINVAL;
+	}
+
+	port_mngr = port->port_mngr;
+
+	ccci_h = mtk_port_strip_header(skb);
+	if (unlikely(!ccci_h)) {
+		dev_warn(port_mngr->ctrl_blk->mdev->dev,
+			 "Unsupported: skb length(%d) is less than ccci header\n",
+			 skb->len);
+		goto drop_data;
+	}
+
+	channel = FIELD_GET(MTK_HDR_FLD_CHN, le32_to_cpu(ccci_h->status));
+	port = mtk_port_search_by_id(port_mngr, channel);
+	if (unlikely(!port)) {
+		dev_warn(port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to find port by channel:%d\n", channel);
+		goto drop_data;
+	}
+
+	ret = mtk_port_check_rx_seq(port, ccci_h);
+	if (unlikely(ret))
+		goto drop_data;
+
+	port->rx_seq = FIELD_GET(MTK_HDR_FLD_SEQ, le32_to_cpu(ccci_h->status));
+	skb_pull(skb, sizeof(*ccci_h));
+
+	/* Support scatter gather transmission */
+	if (port->rx_mtu > port->rx_frag_size) {
+		ret = mtk_port_rx_dispatch_frag_skb(port, skb);
+		/* -EIO means partial data dispatch complete, does not goto drop flow */
+		if (ret < 0 && ret != -EIO)
+			goto drop_frag_skb;
+	} else {
+		ret = ports_ops[port->info.type]->recv(port, skb);
+		if (ret < 0)
+			goto drop_data;
+	}
+
+	return ret;
+
+drop_frag_skb:
+	{
+		struct sk_buff *frag_skb, *tmp;
+
+		frag_skb = skb_shinfo(skb)->frag_list;
+		while (frag_skb) {
+			tmp = frag_skb->next;
+			frag_skb->next = NULL;
+			dev_kfree_skb_any(frag_skb);
+			frag_skb = tmp;
+		}
+		skb_shinfo(skb)->frag_list = NULL;
+	}
+drop_data:
+	dev_kfree_skb_any(skb);
+	return ret;
+}
+
+int mtk_port_add_header(struct sk_buff *skb)
+{
+	struct mtk_ccci_header *ccci_h;
+	struct mtk_port *port;
+	struct trb *trb;
+
+	trb = (struct trb *)skb->cb;
+	if (trb->status == MTK_TRB_HEADER_ADDED)
+		return 0;
+
+	port = trb->priv;
+	if (!port)
+		return -EINVAL;
+
+	ccci_h = skb_push(skb, sizeof(*ccci_h));
+
+	ccci_h->packet_header = cpu_to_le32(0);
+	ccci_h->packet_len = cpu_to_le32(skb->len);
+	ccci_h->ex_msg = cpu_to_le32(0);
+	ccci_h->status = cpu_to_le32(FIELD_PREP(MTK_HDR_FLD_CHN, port->info.tx_ch) |
+				     FIELD_PREP(MTK_HDR_FLD_SEQ, port->tx_seq++) |
+				     FIELD_PREP(MTK_HDR_FLD_AST, 1));
+
+	trb->status = MTK_TRB_HEADER_ADDED;
+
+	return 0;
+}
+
+struct mtk_ccci_header *mtk_port_strip_header(struct sk_buff *skb)
+{
+	struct mtk_ccci_header *ccci_h;
+
+	if (skb->len < sizeof(*ccci_h)) {
+		pr_err("Invalid input value\n");
+		return NULL;
+	}
+
+	ccci_h = (struct mtk_ccci_header *)skb->data;
+
+	return ccci_h;
+}
+
+int mtk_port_status_update(struct mtk_md_dev *mdev, void *data, u32 data_len)
+{
+	struct mtk_port_enum_msg *msg = data;
+	struct mtk_port_info *port_info;
+	struct mtk_port_mngr *port_mngr;
+	struct mtk_ctrl_blk *ctrl_blk;
+	struct mtk_port *port;
+	int port_id;
+	u16 ch_id;
+
+	if (unlikely(!mdev || !msg))
+		return -EINVAL;
+
+	ctrl_blk = mdev->ctrl_blk;
+	port_mngr = ctrl_blk->port_mngr;
+	if (le16_to_cpu(msg->version) != MTK_PORT_ENUM_VER ||
+	    le32_to_cpu(msg->head_pattern) != MTK_PORT_ENUM_HEAD_PATTERN ||
+	    le32_to_cpu(msg->tail_pattern) != MTK_PORT_ENUM_TAIL_PATTERN)
+		return -EPROTO;
+
+	if (data_len < sizeof(*msg) +
+	    le16_to_cpu(msg->port_cnt) * sizeof(*port_info))
+		return -EPROTO;
+
+	for (port_id = 0; port_id < le16_to_cpu(msg->port_cnt); port_id++) {
+		port_info = (struct mtk_port_info *)(msg->data +
+						   (sizeof(*port_info) * port_id));
+		ch_id = FIELD_GET(MTK_INFO_FLD_CHID, le16_to_cpu(port_info->channel));
+		port = mtk_port_search_by_id(port_mngr, ch_id);
+		if (!port)
+			continue;
+		port->enable = FIELD_GET(MTK_INFO_FLD_EN, le16_to_cpu(port_info->channel));
+	}
+
+	return 0;
+}
+
+int mtk_port_ch_enable(struct mtk_port *port)
+{
+	struct mtk_port_mngr *port_mngr = port->port_mngr;
+	struct trb_open_priv *trb_open_priv;
+	struct sk_buff *skb;
+	struct trb *trb;
+	int ret;
+
+	skb = __dev_alloc_skb(Q_MTU_3_5K, GFP_KERNEL);
+	if (!skb)
+		return -ENOMEM;
+
+	trb_open_priv = (struct trb_open_priv *)skb->data;
+	trb_open_priv->rx_done = mtk_port_rx_dispatch;
+
+	skb_put(skb, sizeof(struct trb_open_priv));
+	trb = (struct trb *)skb->cb;
+	mtk_port_trb_init(port, trb, TRB_CMD_ENABLE, mtk_port_open_trb_complete);
+	kref_get(&trb->kref);
+
+	ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb, true);
+	if (ret) {
+		dev_err(port_mngr->ctrl_blk->mdev->dev,
+			"Failed to submit trb for port(%s), ret=%d\n",
+			port->info.name, ret);
+		kref_put(&trb->kref, mtk_port_trb_free);
+		kref_put(&trb->kref, mtk_port_trb_free);
+		return ret;
+	}
+
+	ret = wait_event_timeout(port->trb_wq, trb->status <= 0,
+				 MTK_DFLT_TRB_TIMEOUT);
+	if (!ret)
+		ret = -ETIMEDOUT;
+	else
+		ret = trb->status;
+
+	kref_put(&trb->kref, mtk_port_trb_free);
+
+	return ret;
+}
+
+int mtk_port_ch_disable(struct mtk_port *port)
+{
+	struct mtk_port_mngr *port_mngr = port->port_mngr;
+	struct sk_buff *skb;
+	struct trb *trb;
+	int ret;
+
+	skb = __dev_alloc_skb(Q_MTU_3_5K, GFP_KERNEL);
+	if (!skb)
+		return -ENOMEM;
+
+	trb = (struct trb *)skb->cb;
+	mtk_port_trb_init(port, trb, TRB_CMD_DISABLE, mtk_port_close_trb_complete);
+	kref_get(&trb->kref);
+
+	ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb, true);
+	if (ret) {
+		dev_warn(port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to submit trb for port(%s), ret=%d\n",
+			 port->info.name, ret);
+		kref_put(&trb->kref, mtk_port_trb_free);
+		kref_put(&trb->kref, mtk_port_trb_free);
+		return ret;
+	}
+
+	ret = wait_event_timeout(port->trb_wq, trb->status <= 0,
+				 MTK_DFLT_TRB_TIMEOUT);
+	if (!ret)
+		ret = -ETIMEDOUT;
+	else
+		ret = trb->status;
+
+	kref_put(&trb->kref, mtk_port_trb_free);
+
+	return ret;
+}
+
+int mtk_port_mngr_init(struct mtk_ctrl_blk *ctrl_blk, struct mtk_port_cfg *port_cfg, int port_cnt)
+{
+	struct mtk_port_mngr *port_mngr;
+	int ret = -ENOMEM;
+
+	port_mngr = devm_kzalloc(ctrl_blk->mdev->dev, sizeof(*port_mngr), GFP_KERNEL);
+	if (unlikely(!port_mngr)) {
+		dev_err((ctrl_blk->mdev)->dev, "Failed to alloc memory for port_mngr\n");
+		goto err_out;
+	}
+
+	port_mngr->ctrl_blk = ctrl_blk;
+	mutex_init(&port_mngr->port_tbl_mtx);
+
+	ret = mtk_port_tbl_create(port_mngr, port_cfg, port_cnt);
+	if (unlikely(ret)) {
+		dev_err((ctrl_blk->mdev)->dev, "Failed to create port_tbl\n");
+		goto err_free_port_mngr;
+	}
+
+	ctrl_blk->port_mngr = port_mngr;
+
+	return ret;
+
+err_free_port_mngr:
+	mtk_port_tbl_destroy(port_mngr);
+err_out:
+	return ret;
+}
+
+void mtk_port_mngr_exit(struct mtk_ctrl_blk *ctrl_blk)
+{
+	struct mtk_port_mngr *port_mngr = ctrl_blk->port_mngr;
+
+	mtk_port_tbl_destroy(port_mngr);
+
+	ctrl_blk->port_mngr = NULL;
+}
diff --git a/drivers/net/wwan/t9xx/mtk_port.h b/drivers/net/wwan/t9xx/mtk_port.h
new file mode 100644
index 000000000000..9b8f63ce3f3c
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_port.h
@@ -0,0 +1,131 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_PORT_H__
+#define __MTK_PORT_H__
+
+#include <linux/bits.h>
+#include <linux/device.h>
+#include <linux/radix-tree.h>
+#include <linux/skbuff.h>
+#include <linux/types.h>
+
+#include "mtk_ctrl_plane.h"
+#include "mtk_dev.h"
+
+#define MTK_PEER_ID_MASK			(0xF000)
+#define MTK_PEER_ID_SHIFT			(12)
+#define MTK_PEER_ID(ch)				(((ch) & MTK_PEER_ID_MASK) >> MTK_PEER_ID_SHIFT)
+#define MTK_CH_ID_MASK				(0x0FFF)
+#define MTK_CH_ID(ch)				((ch) & MTK_CH_ID_MASK)
+#define MTK_DFLT_PORT_NAME_LEN			(20)
+
+/* Mapping MTK_PEER_ID and mtk_port_tbl index */
+#define MTK_PORT_TBL_TYPE(ch)			(MTK_PEER_ID(ch) - 1)
+
+/* ccci header length + reserved space that is used in exception flow */
+#define MTK_CCCI_H_ELEN		(128)
+
+#define MTK_HDR_FLD_AST		((u32)BIT(31))
+#define MTK_HDR_FLD_SEQ		GENMASK(30, 16)
+#define MTK_HDR_FLD_CHN		GENMASK(15, 0)
+
+#define MTK_INFO_FLD_EN		((u16)BIT(15))
+#define MTK_INFO_FLD_CHID	GENMASK(14, 0)
+
+enum mtk_port_status {
+	PORT_S_ENABLE,
+	PORT_S_OPEN,
+	PORT_S_WR,
+	PORT_S_FLUSH,
+};
+
+enum mtk_port_flag {
+	PORT_F_DFLT = 0,
+	PORT_F_BLOCKING = BIT(1),
+};
+
+enum mtk_port_tbl {
+	PORT_TBL_SAP,
+	PORT_TBL_MD,
+	PORT_TBL_MAX
+};
+
+enum mtk_port_type {
+	PORT_TYPE_INTERNAL,
+	PORT_TYPE_MAX
+};
+
+struct mtk_internal_port {
+	void *arg;
+	int (*recv_cb)(void *arg, struct sk_buff *skb);
+};
+
+struct mtk_port_cfg {
+	enum mtk_ccci_ch tx_ch;
+	enum mtk_ccci_ch rx_ch;
+	enum mtk_port_type type;
+	char name[MTK_DFLT_PORT_NAME_LEN];
+	unsigned char flags;
+};
+
+struct mtk_port {
+	struct mtk_port_cfg info;
+	struct kref kref;
+	struct rcu_head rcu;
+	bool enable;
+	unsigned long status;
+	unsigned short tx_seq;
+	unsigned short rx_seq;
+	unsigned int tx_mtu;
+	unsigned int rx_mtu;
+	u32 tx_frag_size;
+	u32 rx_frag_size;
+	struct sk_buff_head rx_skb_list;
+	unsigned int rx_data_len;
+	unsigned int rx_buf_size;
+	wait_queue_head_t trb_wq;
+	wait_queue_head_t rx_wq;
+	struct mtk_port_mngr *port_mngr;
+	struct mtk_internal_port i_priv;
+};
+
+struct mtk_port_mngr {
+	struct mtk_ctrl_blk *ctrl_blk;
+	struct radix_tree_root port_tbl[PORT_TBL_MAX];
+	/* protects port_tbl[] and port_cnt against concurrent mutators */
+	struct mutex port_tbl_mtx;
+	unsigned int port_cnt;
+};
+
+struct mtk_ccci_header {
+	__le32 packet_header;
+	__le32 packet_len;
+	__le32 status;
+	__le32 ex_msg;
+};
+
+struct mtk_port_layer_cfg {
+	struct mtk_port_cfg *port_cfg;
+	int port_cnt;
+};
+
+extern const struct port_ops *ports_ops[PORT_TYPE_MAX];
+
+void mtk_port_release(struct kref *port_kref);
+void mtk_port_trb_free(struct kref *trb_kref);
+struct mtk_port *mtk_port_search_by_name(struct mtk_port_mngr *port_mngr, char *name);
+int mtk_port_add_header(struct sk_buff *skb);
+struct mtk_ccci_header *mtk_port_strip_header(struct sk_buff *skb);
+int mtk_port_status_check(struct mtk_port *port);
+int mtk_port_send_data(struct mtk_port *port, void *data, bool blocking, bool force_send);
+int mtk_port_status_update(struct mtk_md_dev *mdev, void *data, u32 data_len);
+int mtk_port_ch_enable(struct mtk_port *port);
+int mtk_port_ch_disable(struct mtk_port *port);
+int mtk_port_mngr_init(struct mtk_ctrl_blk *ctrl_blk, struct mtk_port_cfg *port_cfg, int port_cnt);
+void mtk_port_mngr_exit(struct mtk_ctrl_blk *ctrl_blk);
+void mtk_port_trb_init(struct mtk_port *port, struct trb *trb, enum mtk_trb_cmd_type cmd,
+		       int (*trb_complete)(struct sk_buff *skb));
+#endif /* __MTK_PORT_H__ */
diff --git a/drivers/net/wwan/t9xx/mtk_port_io.c b/drivers/net/wwan/t9xx/mtk_port_io.c
new file mode 100644
index 000000000000..6cff0704b7fc
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_port_io.c
@@ -0,0 +1,258 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+#include <linux/netdevice.h>
+
+#include "mtk_port_io.h"
+
+static int mtk_port_get_locked(struct mtk_port *port)
+{
+	int ret = 0;
+
+	mutex_lock(&port_mngr_grp_mtx);
+	if (!port) {
+		mutex_unlock(&port_mngr_grp_mtx);
+		pr_err("Port does not exist\n");
+		return -ENODEV;
+	}
+	kref_get(&port->kref);
+	mutex_unlock(&port_mngr_grp_mtx);
+
+	return ret;
+}
+
+static void mtk_port_put_locked(struct mtk_port *port)
+{
+	mutex_lock(&port_mngr_grp_mtx);
+	kref_put(&port->kref, mtk_port_release);
+	mutex_unlock(&port_mngr_grp_mtx);
+}
+
+static void mtk_port_struct_init(struct mtk_port *port)
+{
+	port->tx_seq = 0;
+	port->rx_seq = -1;
+	clear_bit(PORT_S_ENABLE, &port->status);
+	kref_init(&port->kref);
+	skb_queue_head_init(&port->rx_skb_list);
+	port->rx_buf_size = MTK_RX_BUF_SIZE;
+	init_waitqueue_head(&port->trb_wq);
+	init_waitqueue_head(&port->rx_wq);
+}
+
+static int mtk_port_internal_init(struct mtk_port *port)
+{
+	mtk_port_struct_init(port);
+	port->enable = false;
+
+	return 0;
+}
+
+static void mtk_port_internal_exit(struct mtk_port *port)
+{
+	if (test_bit(PORT_S_ENABLE, &port->status))
+		ports_ops[port->info.type]->disable(port);
+}
+
+static void mtk_port_internal_enable(struct mtk_port *port)
+{
+	int ret;
+
+	if (test_bit(PORT_S_ENABLE, &port->status))
+		return;
+
+	ret = mtk_port_ch_enable(port);
+	if (ret && ret != -EBUSY) {
+		/* On -ETIMEDOUT the ENABLE trb may still be queued; queue a
+		 * DISABLE behind it so the channel does not end up armed with
+		 * no software owner.
+		 */
+		mtk_port_ch_disable(port);
+		return;
+	}
+
+	set_bit(PORT_S_WR, &port->status);
+	set_bit(PORT_S_ENABLE, &port->status);
+}
+
+static void mtk_port_internal_disable(struct mtk_port *port)
+{
+	int ret;
+
+	if (!test_and_clear_bit(PORT_S_ENABLE, &port->status))
+		return;
+
+	/* PORT_S_WR is part of the blocking TX wait condition, and the waiter
+	 * loops back on timeout, so without this wake it only notices at the
+	 * next trb timeout rather than now.
+	 */
+	clear_bit(PORT_S_WR, &port->status);
+	wake_up_all(&port->trb_wq);
+
+	/* Failure must not abort teardown or keep PORT_S_ENABLE set:
+	 * usr_cnt and the queues are rebuilt together at the next
+	 * FSM_STATE_ON, and a port left with the bit set could never
+	 * be re-enabled (mtk_port_internal_enable() early-returns).
+	 */
+	ret = mtk_port_ch_disable(port);
+	if (ret)
+		dev_warn(port->port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to disable channel for port(%s): %d\n",
+			 port->info.name, ret);
+}
+
+static int mtk_port_internal_recv(struct mtk_port *port, struct sk_buff *skb)
+{
+	struct mtk_internal_port *priv;
+	int ret = -ENXIO;
+
+	if (!test_bit(PORT_S_OPEN, &port->status))
+		goto drop_data;
+
+	priv = &port->i_priv;
+	if (!priv->recv_cb || !priv->arg)
+		goto drop_data;
+
+	ret = priv->recv_cb(priv->arg, skb);
+	return ret;
+
+drop_data:
+	return ret;
+}
+
+static int mtk_port_common_open(struct mtk_port *port)
+{
+	int ret = 0;
+
+	if (!test_bit(PORT_S_ENABLE, &port->status))
+		return -ENODEV;
+
+	if (test_bit(PORT_S_OPEN, &port->status))
+		return -EBUSY;
+
+	skb_queue_purge(&port->rx_skb_list);
+	set_bit(PORT_S_OPEN, &port->status);
+	clear_bit(PORT_S_FLUSH, &port->status);
+
+	return ret;
+}
+
+static void mtk_port_common_close(struct mtk_port *port)
+{
+	clear_bit(PORT_S_OPEN, &port->status);
+
+	skb_queue_purge(&port->rx_skb_list);
+	port->rx_data_len = 0;
+
+	set_bit(PORT_S_FLUSH, &port->status);
+	wake_up_all(&port->trb_wq);
+	wake_up_all(&port->rx_wq);
+}
+
+void *mtk_port_internal_open(struct mtk_md_dev *mdev, char *name, int flag)
+{
+	struct mtk_port_mngr *port_mngr;
+	struct mtk_ctrl_blk *ctrl_blk;
+	struct mtk_port *port;
+	int ret;
+
+	ctrl_blk = mdev->ctrl_blk;
+	port_mngr = ctrl_blk->port_mngr;
+
+	port = mtk_port_search_by_name(port_mngr, name);
+	if (port && port->info.type != PORT_TYPE_INTERNAL) {
+		port = NULL;
+		goto out;
+	}
+
+	ret = mtk_port_get_locked(port);
+	if (ret)
+		goto out;
+
+	ret = mtk_port_common_open(port);
+	if (ret) {
+		mtk_port_put_locked(port);
+		port = NULL;
+		goto out;
+	}
+
+	if (flag & O_NONBLOCK)
+		port->info.flags &= ~PORT_F_BLOCKING;
+	else
+		port->info.flags |= PORT_F_BLOCKING;
+out:
+	return port;
+}
+
+int mtk_port_internal_close(void *i_port)
+{
+	struct mtk_port *port = i_port;
+	int ret = 0;
+
+	if (!port) {
+		ret = -EINVAL;
+		goto end;
+	}
+
+	if (!test_bit(PORT_S_OPEN, &port->status)) {
+		pr_err("Port(%s) has been closed\n", port->info.name);
+		ret = -EBADF;
+		goto end;
+	}
+
+	mtk_port_common_close(port);
+	mtk_port_put_locked(port);
+end:
+	return ret;
+}
+
+int mtk_port_internal_write(void *i_port, struct sk_buff *skb)
+{
+	struct mtk_port *port = i_port;
+	bool blocking;
+
+	if (!port || !skb) {
+		if (skb)
+			dev_kfree_skb_any(skb);
+		pr_err_ratelimited("Internal write: invalid input\n");
+		return -EINVAL;
+	}
+
+	blocking = !!(port->info.flags & PORT_F_BLOCKING);
+
+	return mtk_port_send_data(port, skb, blocking, blocking);
+}
+
+void mtk_port_internal_recv_register(void *i_port,
+				     int (*cb)(void *priv, struct sk_buff *skb),
+				     void *arg)
+{
+	struct mtk_port *port = i_port;
+	struct mtk_internal_port *priv;
+
+	priv = &port->i_priv;
+	priv->arg = arg;
+	priv->recv_cb = cb;
+}
+
+int mtk_port_io_init(void)
+{
+	return 0;
+}
+
+void mtk_port_io_exit(void)
+{
+}
+
+static const struct port_ops port_internal_ops = {
+	.init = mtk_port_internal_init,
+	.exit = mtk_port_internal_exit,
+	.enable = mtk_port_internal_enable,
+	.disable = mtk_port_internal_disable,
+	.recv = mtk_port_internal_recv,
+};
+
+const struct port_ops *ports_ops[PORT_TYPE_MAX] = {
+	&port_internal_ops,
+};
diff --git a/drivers/net/wwan/t9xx/mtk_port_io.h b/drivers/net/wwan/t9xx/mtk_port_io.h
new file mode 100644
index 000000000000..5a4e36c075a5
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_port_io.h
@@ -0,0 +1,35 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_PORT_IO_H__
+#define __MTK_PORT_IO_H__
+
+#include <linux/skbuff.h>
+
+#include "mtk_port.h"
+
+#define MTK_RX_BUF_SIZE			(1024 * 1024)
+
+extern struct mutex port_mngr_grp_mtx;
+
+struct port_ops {
+	int (*init)(struct mtk_port *port);
+	void (*exit)(struct mtk_port *port);
+	void (*enable)(struct mtk_port *port);
+	void (*disable)(struct mtk_port *port);
+	int (*recv)(struct mtk_port *port, struct sk_buff *skb);
+};
+
+void *mtk_port_internal_open(struct mtk_md_dev *mdev, char *name, int flag);
+int mtk_port_internal_close(void *i_port);
+int mtk_port_internal_write(void *i_port, struct sk_buff *skb);
+void mtk_port_internal_recv_register(void *i_port,
+				     int (*cb)(void *priv, struct sk_buff *skb),
+				     void *arg);
+
+int mtk_port_io_init(void);
+void mtk_port_io_exit(void);
+
+#endif /* __MTK_PORT_IO_H__ */
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
index 6c31a5d80693..84b7fb3c82fd 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_pci.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
@@ -15,6 +15,8 @@
 #include "mtk_trans_ctrl.h"
 #include "mtk_pci.h"
 #include "mtk_pci_reg.h"
+#include "mtk_port.h"
+#include "mtk_port_io.h"
 
 #define MTK_MHCCIF_RC_BASE_ADDR		0x1000A000
 
@@ -1069,7 +1071,28 @@ static struct pci_driver mtk_pci_drv = {
 	.err_handler = &mtk_pci_err_handler
 };
 
-module_pci_driver(mtk_pci_drv);
+static int __init mtk_drv_init(void)
+{
+	int ret;
+
+	ret = mtk_port_io_init();
+	if (ret)
+		return ret;
+
+	ret = pci_register_driver(&mtk_pci_drv);
+	if (ret)
+		mtk_port_io_exit();
+
+	return ret;
+}
+module_init(mtk_drv_init);
+
+static void __exit mtk_drv_exit(void)
+{
+	pci_unregister_driver(&mtk_pci_drv);
+	mtk_port_io_exit();
+}
+module_exit(mtk_drv_exit);
 
 MODULE_DESCRIPTION("MediaTek T9xx PCIe WWAN driver");
 MODULE_LICENSE("GPL");
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
index 531940b06e2a..d485bd7e3bf9 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
@@ -16,6 +16,7 @@
 #include "mtk_ctrl_plane.h"
 #include "mtk_dev.h"
 #include "mtk_pci.h"
+#include "mtk_port.h"
 #include "mtk_trans_ctrl.h"
 
 #define QUEUE_CHL_MASK	0xFFFF
@@ -34,6 +35,18 @@ static const struct queue_info mtk_queue_info[] = {
 	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
 };
 
+static const struct mtk_port_cfg mtk_port_cfg_tbl[] = {
+	{CCCI_CONTROL_TX, CCCI_CONTROL_RX, PORT_TYPE_INTERNAL, "MDCTRL",
+		PORT_F_DFLT},
+	{CCCI_SAP_CONTROL_TX, CCCI_SAP_CONTROL_RX, PORT_TYPE_INTERNAL, "SAPCTRL",
+		PORT_F_DFLT},
+};
+
+static struct mtk_port_layer_cfg mtk_port_layer_cfg = {
+	.port_cfg = (struct mtk_port_cfg *)mtk_port_cfg_tbl,
+	.port_cnt = ARRAY_SIZE(mtk_port_cfg_tbl),
+};
+
 static bool mtk_queue_list_is_full(struct mtk_ctrl_trans *trans, struct queue_info *que)
 {
 	return skb_queue_len_lockless(&trans->trans_list[que->hif_id].skb_list[que->txqno]) >=
@@ -164,6 +177,7 @@ static void mtk_ctrl_trb_handler(struct trb_srv *srv, struct trans_list *trans_l
 			break;
 		}
 		trb = (struct trb *)skb->cb;
+		kref_get(&trb->kref);
 
 		switch (trb->cmd) {
 		case TRB_CMD_ENABLE:
@@ -185,8 +199,10 @@ static void mtk_ctrl_trb_handler(struct trb_srv *srv, struct trans_list *trans_l
 					kick = true;
 					break;
 				}
-				if (err == -EAGAIN)
+				if (err == -EAGAIN) {
+					kref_put(&trb->kref, mtk_port_trb_free);
 					return;
+				}
 
 				skb_unlink(skb, skb_list);
 				trb->status = err;
@@ -227,6 +243,8 @@ static void mtk_ctrl_trb_handler(struct trb_srv *srv, struct trans_list *trans_l
 			kick = false;
 		}
 
+		kref_put(&trb->kref, mtk_port_trb_free);
+
 		loop++;
 	} while (loop < TRB_NUM_PER_ROUND);
 }
@@ -597,7 +615,7 @@ int mtk_trans_ctrl_init(struct mtk_md_dev *mdev)
 	trans->queue_info_num = ARRAY_SIZE(mtk_queue_info);
 	trans->trb_srv_num = TRB_SRV_NUM;
 
-	err = mtk_ctrl_init(mdev);
+	err = mtk_ctrl_init(mdev, &mtk_port_layer_cfg);
 	if (err)
 		return err;
 
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
index f9144c60edf3..1228f40f2d39 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
@@ -13,6 +13,7 @@
 #include <linux/types.h>
 
 #include "mtk_dev.h"
+#include "mtk_port.h"
 
 #define TRB_SRV_MAX_NUM			(1)
 #define HW_QUE_NUM			(8)

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
                   ` (3 preceding siblings ...)
  2026-09-30  7:46 ` [PATCH v9 4/6] net: wwan: t9xx: Add control port Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  2026-09-30  7:46 ` [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports Jack Wu via B4 Relay
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

The FSM (Finite-state Machine) thread is responsible for
synchronizing the actions of different modules. The
asynchronous events from the device or the OS will trigger
a state transition.

The FSM thread will append it to the event queue when an
event arrives. It handles the events sequentially. After
processing the event, the FSM thread notifies other modules
before and after the state transition.

Five FSM states are defined. They can transition from one
state to another, and self-transition in some states.

Also wire the HS1/HS2/HS3 runtime feature handshake into
the FSM: the host opens a control port on device readiness,
sends HS1 with the supported feature set, parses the HS2
reply, and confirms with HS3. MD and SAP each run this
sequence independently; the FSM advances to READY only
after both complete.

Every state transition also emits a KOBJ_CHANGE uevent
carrying "<kobj name>:event_id=1, info=state=<state>,
fsm_flag=0x<flag>". It is informational only: no driver
behaviour depends on it and userspace is not required to
consume it.

CLDMA hardware bring-up and teardown are driven by FSM
state transitions: mtk_cldma_dev_init() on FSM_STATE_BOOTUP
allocates DMA pools, creates a workqueue and registers the
MSI-X IRQ, and mtk_cldma_dev_exit() tears them down on
FSM_STATE_OFF.

Registering that IRQ is what first makes the completion
path built by the earlier patches run: mtk_cldma_isr()
dispatches tx_done_work and rx_done_work, and queues
err_work on a QUEUE_ERROR interrupt. err_work owns the
recovery policy - it stops the offending queue, completes
every outstanding TX request with -EPIPE rather than
retrying it, and hands a stopped RX queue back to its
worker to be restarted.

A queue that will not stop within the timeout is still a
live DMA target, so its ring, and the DMA pools its
descriptors live in, are deliberately leaked instead of
being returned to the allocator. Teardown is ordered to
match: mtk_pcie_hif_exit() joins the trb service threads
before mtk_cldma_exit() frees what they dereference.

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 drivers/net/wwan/t9xx/Makefile              |    1 +
 drivers/net/wwan/t9xx/mtk_ctrl_plane.c      |   51 ++
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      |    1 +
 drivers/net/wwan/t9xx/mtk_dev.h             |    1 +
 drivers/net/wwan/t9xx/mtk_fsm.c             | 1174 +++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/mtk_fsm.h             |  158 ++++
 drivers/net/wwan/t9xx/mtk_port.c            |   66 ++
 drivers/net/wwan/t9xx/mtk_port.h            |    2 +
 drivers/net/wwan/t9xx/mtk_utility.h         |   29 +
 drivers/net/wwan/t9xx/pcie/mtk_cldma.c      |  393 ++++++++-
 drivers/net/wwan/t9xx/pcie/mtk_cldma.h      |    4 +
 drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h  |    7 +
 drivers/net/wwan/t9xx/pcie/mtk_pci.c        |   30 +-
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c |   30 +-
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h |    3 +
 15 files changed, 1932 insertions(+), 18 deletions(-)

diff --git a/drivers/net/wwan/t9xx/Makefile b/drivers/net/wwan/t9xx/Makefile
index d44266009362..fcfddf2e5147 100644
--- a/drivers/net/wwan/t9xx/Makefile
+++ b/drivers/net/wwan/t9xx/Makefile
@@ -9,6 +9,7 @@ mtk_t9xx-y := \
 	mtk_ctrl_plane.o \
 	mtk_port.o \
 	mtk_port_io.o \
+	mtk_fsm.o \
 	pcie/mtk_pci.o \
 	pcie/mtk_trans_ctrl.o \
 	pcie/mtk_cldma.o \
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
index 10f0d3267b3b..4ea51b4d6a7c 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
@@ -5,9 +5,50 @@
  */
 
 #include <linux/device.h>
+#include <linux/freezer.h>
+#include <linux/kthread.h>
+#include <linux/list.h>
+#include <linux/pm_runtime.h>
+#include <linux/sched.h>
+#include <linux/wait.h>
 
 #include "mtk_ctrl_plane.h"
 #include "mtk_port.h"
+#include "mtk_trans_ctrl.h"
+
+static void mtk_ctrl_trans_fsm_state_handler(struct mtk_fsm_param *param,
+					     struct mtk_ctrl_blk *ctrl_blk)
+{
+	struct mtk_md_dev *mdev = ctrl_blk->mdev;
+	int ret;
+
+	switch (param->to) {
+	case FSM_STATE_OFF:
+		mtk_pcie_hif_exit(mdev);
+		mtk_pcie_hif_fsm_indication(mdev, param);
+		break;
+	case FSM_STATE_ON:
+		ret = mtk_pcie_hif_init(mdev);
+		if (ret) {
+			dev_err(mdev->dev, "Failed to init HIF: %d\n", ret);
+			mtk_fsm_hif_err_record(mdev, MTK_FSM_HIF_USER_CTRL, ret);
+			break;
+		}
+		fallthrough;
+	default:
+		mtk_pcie_hif_fsm_indication(mdev, param);
+		break;
+	}
+}
+
+static void mtk_ctrl_fsm_state_listener(struct mtk_fsm_param *param, void *data)
+{
+	struct mtk_ctrl_blk *ctrl_blk = data;
+
+	mtk_port_mngr_fsm_state_handler(param, ctrl_blk->port_mngr);
+	mtk_ctrl_trans_fsm_state_handler(param, ctrl_blk);
+	mtk_port_mngr_fsm_state_handler_late(param, ctrl_blk->port_mngr);
+}
 
 /**
  * mtk_ctrl_init() - Initialize the control plane block.
@@ -36,8 +77,17 @@ int mtk_ctrl_init(struct mtk_md_dev *mdev, struct mtk_port_layer_cfg *port_layer
 	if (err)
 		goto err_free_mem;
 
+	err = mtk_fsm_notifier_register(mdev, MTK_USER_CTRL, mtk_ctrl_fsm_state_listener,
+					ctrl_blk, FSM_PRIO_1, false);
+	if (err) {
+		dev_err(mdev->dev, "Fail to register fsm notification(ret = %d)\n", err);
+		goto err_port_exit;
+	}
+
 	return 0;
 
+err_port_exit:
+	mtk_port_mngr_exit(ctrl_blk);
 err_free_mem:
 	return err;
 }
@@ -53,6 +103,7 @@ void mtk_ctrl_exit(struct mtk_md_dev *mdev)
 {
 	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
 
+	mtk_fsm_notifier_unregister(mdev, MTK_USER_CTRL);
 	mtk_port_mngr_exit(ctrl_blk);
 	mdev->ctrl_blk = NULL;
 }
diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
index 933ba6193b70..1f13e25502ad 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
@@ -10,6 +10,7 @@
 #include <linux/skbuff.h>
 
 #include "mtk_dev.h"
+#include "mtk_fsm.h"
 
 #define Q_MTU_3_5K			(0xE00)
 #define Q_FRAG_3_5K			(0xE00)
diff --git a/drivers/net/wwan/t9xx/mtk_dev.h b/drivers/net/wwan/t9xx/mtk_dev.h
index 84fcd980110b..658fd5b8431e 100644
--- a/drivers/net/wwan/t9xx/mtk_dev.h
+++ b/drivers/net/wwan/t9xx/mtk_dev.h
@@ -40,6 +40,7 @@ struct mtk_md_dev {
 	u32 hw_ver;
 	char dev_str[MTK_DEV_STR_LEN];
 	struct mtk_ctrl_blk *ctrl_blk;
+	struct mtk_md_fsm *fsm;
 };
 
 #endif /* __MTK_DEV_H__ */
diff --git a/drivers/net/wwan/t9xx/mtk_fsm.c b/drivers/net/wwan/t9xx/mtk_fsm.c
new file mode 100644
index 000000000000..94617933e771
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_fsm.c
@@ -0,0 +1,1174 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#include <linux/bitfield.h>
+#include <linux/device.h>
+#include <linux/kref.h>
+#include <linux/kthread.h>
+#include <linux/list.h>
+#include <linux/pci.h>
+#include <linux/sched/signal.h>
+#include <linux/sched/task.h>
+#include <linux/skbuff.h>
+#include <linux/wait.h>
+
+#include "mtk_fsm.h"
+#include "mtk_pci.h"
+#include "mtk_port.h"
+#include "mtk_port_io.h"
+#include "mtk_utility.h"
+
+#define EVT_TF_GATECLOSED (1)
+#define MTK_FSM_INFO_LEN	(64)
+
+#define FSM_HS_START_MASK	(FSM_F_SAP_HS_START | FSM_F_MD_HS_START)
+#define FSM_HS2_DONE_MASK	(FSM_F_SAP_HS2_DONE | FSM_F_MD_HS2_DONE)
+
+#define RTFT_DATA_SIZE		(3 * 1024)
+#define EVT_HANDLER_TIMEOUT	(HZ * 30)
+#define BLOCKING_EVT_TIMEOUT	(2 * EVT_HANDLER_TIMEOUT)
+
+#define REGION_BITMASK		0xF
+#define DEVICE_CFG_SHIFT	24
+#define DEVICE_CFG_REGION_MASK	0x3
+
+enum device_stage {
+	DEV_STAGE_IDLE = 4,
+	DEV_STAGE_MAX
+};
+
+enum device_cfg {
+	DEV_CFG_MD_ONLY = 1,
+};
+
+enum runtime_feature_support_type {
+	RTFT_TYPE_NOT_EXIST = 0,
+	RTFT_TYPE_NOT_SUPPORT = 1,
+	RTFT_TYPE_MUST_SUPPORT = 2,
+	RTFT_TYPE_OPTIONAL_SUPPORT = 3,
+	RTFT_TYPE_SUPPORT_BACKWARD_COMPAT = 4,
+};
+
+enum query_runtime_feature_id {
+	QUERY_RTFT_ID_MD_PORT_ENUM = 0,
+	QUERY_RTFT_ID_SAP_PORT_ENUM = 1,
+	QUERY_RTFT_ID_MD_PORT_CFG = 2,
+};
+
+enum ctrl_msg_id {
+	CTRL_MSG_HS1 = 0,
+	CTRL_MSG_HS2 = 1,
+	CTRL_MSG_HS3 = 2,
+};
+
+struct ctrl_msg_header {
+	__le32 id;
+	__le32 ex_msg;
+	__le32 data_len;
+	u8 reserved[];
+} __packed;
+
+struct runtime_feature_entry {
+	u8 feature_id;
+	struct runtime_feature_info support_info;
+	u8 reserved[2];
+	__le32 data_len;
+	u8 data[];
+} __packed;
+
+struct feature_query {
+	__le32 head_pattern;
+	struct runtime_feature_info ft_set[FEATURE_CNT];
+	__le32 tail_pattern;
+};
+
+static int mtk_fsm_send_hs1_msg(struct fsm_hs_info *hs_info)
+{
+	struct ctrl_msg_header *ctrl_msg_h;
+	struct feature_query *ft_query;
+	struct sk_buff *skb;
+	int ret, msg_size;
+
+	msg_size = sizeof(*ctrl_msg_h) + sizeof(*ft_query);
+	skb = __dev_alloc_skb(msg_size, GFP_KERNEL);
+	if (!skb)
+		return -ENOMEM;
+
+	skb_put(skb, msg_size);
+	ctrl_msg_h = (struct ctrl_msg_header *)skb->data;
+	ctrl_msg_h->id = cpu_to_le32(CTRL_MSG_HS1);
+	ctrl_msg_h->ex_msg = 0;
+	ctrl_msg_h->data_len = cpu_to_le32(sizeof(*ft_query));
+
+	ft_query = (struct feature_query *)(skb->data + sizeof(*ctrl_msg_h));
+	ft_query->head_pattern = cpu_to_le32(FEATURE_QUERY_PATTERN);
+	memcpy(ft_query->ft_set, hs_info->query_ft_set, sizeof(hs_info->query_ft_set));
+	ft_query->tail_pattern = cpu_to_le32(FEATURE_QUERY_PATTERN);
+
+	/* send handshake1 message to device */
+	ret = mtk_port_internal_write(hs_info->ctrl_port, skb);
+	if (ret <= 0)
+		return ret;
+
+	return 0;
+}
+
+static int mtk_fsm_feature_set_match(enum runtime_feature_support_type *cur_ft_spt,
+				     struct runtime_feature_info rtft_info_st,
+				     struct runtime_feature_info rtft_info_cfg)
+{
+	int ret = 0;
+
+	switch (FIELD_GET(FEATURE_TYPE, rtft_info_st.feature)) {
+	case RTFT_TYPE_NOT_EXIST:
+		fallthrough;
+	case RTFT_TYPE_NOT_SUPPORT:
+		/* The device refusing a feature the host declared mandatory
+		 * must fail the handshake, symmetrically with the
+		 * MUST_SUPPORT case below.
+		 */
+		if (FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_MUST_SUPPORT)
+			ret = -EPROTO;
+		else
+			*cur_ft_spt = RTFT_TYPE_NOT_EXIST;
+		break;
+	case RTFT_TYPE_MUST_SUPPORT:
+		if (FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_EXIST ||
+		    FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_SUPPORT)
+			ret = -EPROTO;
+		else
+			*cur_ft_spt = RTFT_TYPE_MUST_SUPPORT;
+		break;
+	case RTFT_TYPE_OPTIONAL_SUPPORT:
+		if (FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_EXIST ||
+		    FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_SUPPORT) {
+			*cur_ft_spt = RTFT_TYPE_NOT_SUPPORT;
+		} else {
+			if (FIELD_GET(FEATURE_VER, rtft_info_st.feature) ==
+			    FIELD_GET(FEATURE_VER, rtft_info_cfg.feature))
+				*cur_ft_spt = RTFT_TYPE_MUST_SUPPORT;
+			else
+				*cur_ft_spt = RTFT_TYPE_NOT_SUPPORT;
+		}
+		break;
+	case RTFT_TYPE_SUPPORT_BACKWARD_COMPAT:
+		if (FIELD_GET(FEATURE_VER, rtft_info_st.feature) >=
+		    FIELD_GET(FEATURE_VER, rtft_info_cfg.feature))
+			*cur_ft_spt = RTFT_TYPE_MUST_SUPPORT;
+		else
+			*cur_ft_spt = RTFT_TYPE_NOT_EXIST;
+		break;
+	default:
+		ret = -EPROTO;
+	}
+
+	return ret;
+}
+
+static int (*query_rtft_action[FEATURE_CNT])(struct mtk_md_dev *mdev,
+					     void *rt_data, u32 data_len) = {
+	[QUERY_RTFT_ID_MD_PORT_ENUM] = mtk_port_status_update,
+	[QUERY_RTFT_ID_SAP_PORT_ENUM] = mtk_port_status_update,
+};
+
+/* The runtime data skb is owned by the caller (the FSM kthread, which got
+ * it attached to the STARTUP event); its pointer and length are read once
+ * from the skb itself, never from a shared field the rx path could rewrite.
+ */
+static int mtk_fsm_parse_hs2_msg(struct mtk_md_fsm *fsm, struct fsm_hs_info *hs_info,
+				 struct sk_buff *skb)
+{
+	enum runtime_feature_support_type cur_ft_spt;
+	struct runtime_feature_entry *rtft_entry;
+	unsigned int ft_id, offset, data_len;
+	unsigned int rt_data_len;
+	char *rt_data;
+	int ret = 0;
+
+	if (!skb)
+		return -EINVAL;
+
+	rt_data = skb->data;
+	rt_data_len = skb->len;
+
+	offset = sizeof(struct feature_query);
+	for (ft_id = 0; ft_id < FEATURE_CNT; ft_id++) {
+		if (offset + sizeof(*rtft_entry) > rt_data_len)
+			break;
+
+		rtft_entry = (struct runtime_feature_entry *)(rt_data + offset);
+		ret = mtk_fsm_feature_set_match(&cur_ft_spt,
+						rtft_entry->support_info,
+						hs_info->query_ft_set[ft_id]);
+		if (ret < 0)
+			break;
+
+		data_len = le32_to_cpu(rtft_entry->data_len);
+		if (data_len > rt_data_len - offset - sizeof(*rtft_entry))
+			break;
+
+		if (cur_ft_spt == RTFT_TYPE_MUST_SUPPORT && query_rtft_action[ft_id]) {
+			if (!data_len) {
+				dev_err(fsm->mdev->dev,
+					"RTFT feature %u: zero data_len\n", ft_id);
+				ret = -EPROTO;
+				break;
+			}
+			ret = query_rtft_action[ft_id](fsm->mdev,
+						       rtft_entry->data,
+						       data_len);
+		}
+		if (ret < 0)
+			break;
+
+		offset += sizeof(*rtft_entry) + data_len;
+	}
+
+	if (ft_id != FEATURE_CNT) {
+		dev_err((fsm->mdev)->dev, "Unable to handle mistake hs2 msg, ft_id=%d\n", ft_id);
+		ret = -EPROTO;
+	}
+
+	return ret;
+}
+
+static int mtk_fsm_append_rtft_entries(struct mtk_md_dev *mdev, void *feature_data,
+				       unsigned int *len, struct fsm_hs_info *hs_info,
+				       struct sk_buff *skb)
+{
+	struct runtime_feature_entry *rtft_entry;
+	int ft_id, ret = 0, rtdata_len = 0;
+	struct feature_query *ft_query;
+
+	if (!skb || skb->len < sizeof(*ft_query)) {
+		ret = -EPROTO;
+		goto hs_err;
+	}
+
+	ft_query = (struct feature_query *)skb->data;
+	if (le32_to_cpu(ft_query->head_pattern) != FEATURE_QUERY_PATTERN ||
+	    le32_to_cpu(ft_query->tail_pattern) != FEATURE_QUERY_PATTERN) {
+		ret = -EPROTO;
+		goto hs_err;
+	}
+
+	/* parse runtime feature query and fill runtime feature entry */
+	rtft_entry = feature_data;
+	for (ft_id = 0; ft_id < FEATURE_CNT && rtdata_len < RTFT_DATA_SIZE; ft_id++) {
+		rtft_entry->feature_id = ft_id;
+		rtft_entry->data_len = 0;
+
+		switch (FIELD_GET(FEATURE_TYPE, ft_query->ft_set[ft_id].feature)) {
+		case RTFT_TYPE_NOT_EXIST:
+			fallthrough;
+		case RTFT_TYPE_NOT_SUPPORT:
+			fallthrough;
+		case RTFT_TYPE_MUST_SUPPORT:
+			rtft_entry->support_info = ft_query->ft_set[ft_id];
+			break;
+		case RTFT_TYPE_OPTIONAL_SUPPORT:
+			fallthrough;
+		case RTFT_TYPE_SUPPORT_BACKWARD_COMPAT:
+			rtft_entry->support_info.feature = FEATURE_TYPE_NOT;
+			rtft_entry->support_info.feature |= FEATURE_VER_0;
+			break;
+		}
+
+		rtdata_len += sizeof(*rtft_entry) + le32_to_cpu(rtft_entry->data_len);
+		rtft_entry = (struct runtime_feature_entry *)(feature_data + rtdata_len);
+	}
+	*len = rtdata_len;
+	return 0;
+
+hs_err:
+	*len = 0;
+	return ret;
+}
+
+static int mtk_fsm_send_hs3_msg(struct fsm_hs_info *hs_info, struct sk_buff *rt_skb)
+{
+	struct mtk_md_fsm *fsm = container_of(hs_info, struct mtk_md_fsm, hs_info[hs_info->id]);
+	unsigned int data_len, msg_size = 0;
+	struct ctrl_msg_header *ctrl_msg_h;
+	struct sk_buff *skb;
+	int ret;
+
+	skb = __dev_alloc_skb(RTFT_DATA_SIZE, GFP_KERNEL);
+	if (!skb)
+		return -ENOMEM;
+	memset(skb->data, 0, RTFT_DATA_SIZE);
+
+	msg_size += sizeof(*ctrl_msg_h);
+	ctrl_msg_h = (struct ctrl_msg_header *)skb->data;
+	ctrl_msg_h->id = cpu_to_le32(CTRL_MSG_HS3);
+	ctrl_msg_h->ex_msg = 0;
+	ret = mtk_fsm_append_rtft_entries(fsm->mdev,
+					  skb->data + sizeof(*ctrl_msg_h),
+					  &data_len, hs_info, rt_skb);
+	if (ret) {
+		dev_kfree_skb(skb);
+		return ret;
+	}
+
+	ctrl_msg_h->data_len = cpu_to_le32(data_len);
+	msg_size += data_len;
+	skb_put(skb, msg_size);
+	ret = mtk_port_internal_write(hs_info->ctrl_port, skb);
+	if (ret <= 0)
+		return ret;
+
+	return 0;
+}
+
+/* Both handlers own the skb on every path and always return 0: the caller
+ * (mtk_port_rx_dispatch) frees the skb itself on a negative return, so
+ * returning an error after dev_kfree_skb() would free it twice.
+ *
+ * On success the skb is attached to the submitted event (event->data) and
+ * from then on belongs to the FSM kthread; hs_info->rt_data is kept only
+ * as the duplicate-HS2 guard and is cleared by the kthread once the event
+ * has been consumed.  That store and the accesses below are the only
+ * concurrent users of the field, so they are marked: a stale non-NULL read
+ * at worst drops one retransmitted HS2, which the device repeats.
+ */
+static int mtk_fsm_sap_ctrl_msg_handler(void *__fsm, struct sk_buff *skb)
+{
+	struct ctrl_msg_header *ctrl_msg_h;
+	struct mtk_md_fsm *fsm = __fsm;
+	struct fsm_hs_info *hs_info;
+	int ret;
+
+	if (skb->len < sizeof(*ctrl_msg_h)) {
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	ctrl_msg_h = (struct ctrl_msg_header *)skb->data;
+	skb_pull(skb, sizeof(*ctrl_msg_h));
+
+	hs_info = &fsm->hs_info[HS_ID_SAP];
+	if (le32_to_cpu(ctrl_msg_h->id) != CTRL_MSG_HS2) {
+		dev_err(fsm->mdev->dev, "Invalid SAP ctrl msg id\n");
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	if (READ_ONCE(hs_info->rt_data)) {
+		dev_warn(fsm->mdev->dev, "Duplicate SAP HS2, dropping\n");
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	WRITE_ONCE(hs_info->rt_data, skb);
+	ret = mtk_fsm_evt_submit(fsm->mdev, FSM_EVT_STARTUP,
+				 hs_info->fsm_flag_hs2, skb, skb->len, 0);
+	if (ret == FSM_EVT_RET_FAIL) {
+		WRITE_ONCE(hs_info->rt_data, NULL);
+		dev_kfree_skb(skb);
+	}
+
+	return 0;
+}
+
+static int mtk_fsm_md_ctrl_msg_handler(void *__fsm, struct sk_buff *skb)
+{
+	struct ctrl_msg_header *ctrl_msg_h;
+	struct mtk_md_fsm *fsm = __fsm;
+	struct fsm_hs_info *hs_info;
+	int ret;
+
+	if (skb->len < sizeof(*ctrl_msg_h)) {
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	ctrl_msg_h = (struct ctrl_msg_header *)skb->data;
+	hs_info = &fsm->hs_info[HS_ID_MD];
+	if (le32_to_cpu(ctrl_msg_h->id) != CTRL_MSG_HS2) {
+		dev_err(fsm->mdev->dev, "Invalid ctrl msg id\n");
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	if (READ_ONCE(hs_info->rt_data)) {
+		/* The guarded skb belongs to a still-queued event; only the
+		 * new skb may be freed here.
+		 */
+		dev_warn(fsm->mdev->dev, "Duplicate MD HS2, dropping\n");
+		dev_kfree_skb(skb);
+		return 0;
+	}
+
+	skb_pull(skb, sizeof(*ctrl_msg_h));
+	WRITE_ONCE(hs_info->rt_data, skb);
+	ret = mtk_fsm_evt_submit(fsm->mdev, FSM_EVT_STARTUP,
+				 hs_info->fsm_flag_hs2, skb, skb->len, 0);
+	if (ret == FSM_EVT_RET_FAIL) {
+		WRITE_ONCE(hs_info->rt_data, NULL);
+		dev_kfree_skb(skb);
+	}
+
+	return 0;
+}
+
+static int (*ctrl_msg_handler[HS_ID_MAX])(void *__fsm, struct sk_buff *skb) = {
+	[HS_ID_MD] = mtk_fsm_md_ctrl_msg_handler,
+	[HS_ID_SAP] = mtk_fsm_sap_ctrl_msg_handler,
+};
+
+static int mtk_fsm_idle_evt_handler(struct mtk_md_dev *mdev,
+				    u32 dev_state, struct mtk_md_fsm *fsm)
+{
+	u32 dev_cfg = dev_state >> DEVICE_CFG_SHIFT & DEVICE_CFG_REGION_MASK;
+	int hs_id;
+
+	if (dev_cfg == DEV_CFG_MD_ONLY)
+		fsm->hs_done_flag = FSM_F_MD_HS_START | FSM_F_MD_HS2_DONE;
+	else
+		fsm->hs_done_flag = FSM_HS_START_MASK | FSM_HS2_DONE_MASK;
+
+	/* On failure keep the handshake channels masked and report it, so
+	 * the device's next boot-flow notification can retrigger us.
+	 */
+	if (mtk_fsm_evt_submit(mdev, FSM_EVT_STARTUP, FSM_F_DFLT,
+			       NULL, 0, 0) == FSM_EVT_RET_FAIL) {
+		dev_err(mdev->dev, "Failed to submit STARTUP evt, waiting for retry\n");
+		return -ENOMEM;
+	}
+
+	for (hs_id = 0; hs_id < HS_ID_MAX; hs_id++)
+		mtk_pci_unmask_ext_evt(mdev, fsm->hs_info[hs_id].mhccif_ch);
+
+	return 0;
+}
+
+static int mtk_fsm_early_bootup_handler(u32 status, void *__fsm)
+{
+	struct mtk_md_fsm *fsm = __fsm;
+	struct mtk_md_dev *mdev;
+	u32 dev_state, dev_stage;
+
+	mdev = fsm->mdev;
+	mtk_pci_mask_ext_evt(mdev, status);
+	mtk_pci_clear_ext_evt(mdev, status);
+
+	dev_state = mtk_pci_get_dev_state(mdev);
+	dev_stage = dev_state & REGION_BITMASK;
+	if (dev_stage >= DEV_STAGE_MAX) {
+		dev_err(mdev->dev, "Invalid dev state 0x%x\n", dev_state);
+		goto exit;
+	}
+
+	if (dev_state == fsm->last_dev_state)
+		goto exit;
+
+	/* Only latch the state once it has been acted on; a failed submit
+	 * leaves last_dev_state unchanged so the repeated notification is
+	 * not filtered out.
+	 */
+	if (dev_stage == DEV_STAGE_IDLE &&
+	    mtk_fsm_idle_evt_handler(mdev, dev_state, fsm))
+		goto exit;
+
+	fsm->last_dev_state = dev_state;
+
+exit:
+	/* Re-arm the channel masked on entry: the device notifies once
+	 * per boot stage, so leaving it masked drops every later stage.
+	 * A device still booting when the driver binds would never
+	 * deliver its DEV_STAGE_IDLE notification, and a STARTUP submit
+	 * that failed above would never see a second one to retry on.
+	 */
+	mtk_pci_unmask_ext_evt(mdev, status);
+	return 0;
+}
+
+static int mtk_fsm_ctrl_ch_start(struct mtk_md_fsm *fsm, struct fsm_hs_info *hs_info, int flag)
+{
+	if (!hs_info->ctrl_port) {
+		hs_info->ctrl_port = mtk_port_internal_open(fsm->mdev, hs_info->port_name, flag);
+		if (!hs_info->ctrl_port) {
+			dev_err(fsm->mdev->dev, "Failed to open ctrl port(%s)\n",
+				hs_info->port_name);
+			return -ENODEV;
+		}
+
+		mtk_port_internal_recv_register(hs_info->ctrl_port,
+						ctrl_msg_handler[hs_info->id], fsm);
+	}
+
+	return 0;
+}
+
+static void mtk_fsm_ctrl_ch_stop(struct mtk_md_fsm *fsm)
+{
+	struct fsm_hs_info *hs_info;
+	int hs_id;
+
+	for (hs_id = 0; hs_id < HS_ID_MAX; hs_id++) {
+		hs_info = &fsm->hs_info[hs_id];
+		if (hs_info->ctrl_port) {
+			mtk_port_internal_close(hs_info->ctrl_port);
+			hs_info->ctrl_port = NULL;
+		}
+	}
+}
+
+static void mtk_fsm_switch_state(struct mtk_md_fsm *fsm,
+				 enum mtk_fsm_state to_state, struct mtk_fsm_evt *event)
+{
+	char fsm_info[MTK_FSM_INFO_LEN];
+	struct mtk_fsm_notifier *nt;
+	struct mtk_fsm_param param;
+
+	param.from = fsm->state;
+	param.to = to_state;
+	param.evt_id = event ? event->id : FSM_EVT_MAX;
+	param.fsm_flag = event ? event->fsm_flag : FSM_F_DFLT;
+
+	mutex_lock(&fsm->notifier_lock);
+	list_for_each_entry(nt, &fsm->pre_notifiers, entry)
+		nt->cb(&param, nt->data);
+
+	fsm->state = to_state;
+	fsm->fsm_flag |= event ? event->fsm_flag : FSM_F_DFLT;
+
+	snprintf(fsm_info, MTK_FSM_INFO_LEN,
+		 "state=%d, fsm_flag=0x%x", to_state, fsm->fsm_flag);
+	mtk_uevent_notify(fsm->mdev->dev, MTK_UEVENT_FSM, fsm_info);
+
+	list_for_each_entry(nt, &fsm->post_notifiers, entry)
+		nt->cb(&param, nt->data);
+	mutex_unlock(&fsm->notifier_lock);
+}
+
+static int mtk_fsm_startup_act(struct mtk_md_fsm *fsm, struct mtk_fsm_evt *event)
+{
+	enum mtk_fsm_state to_state = FSM_STATE_BOOTUP;
+	struct mtk_md_dev *mdev = fsm->mdev;
+	struct fsm_hs_info *hs_info;
+	struct sk_buff *skb = NULL;
+	int ret = 0;
+
+	if (event->fsm_flag & FSM_HS2_DONE_MASK) {
+		/* The HS2 event owns its runtime-data skb. */
+		skb = event->data;
+		hs_info = &fsm->hs_info[(event->fsm_flag & FSM_F_MD_HS2_DONE) ?
+					HS_ID_MD : HS_ID_SAP];
+	} else {
+		/* HS1 events carry the hs_info; the bare STARTUP carries NULL. */
+		hs_info = event->data;
+	}
+
+	if (fsm->state != FSM_STATE_ON && fsm->state != FSM_STATE_BOOTUP) {
+		ret = -EPROTO;
+		goto free_skb;
+	}
+
+	if (!(event->fsm_flag & (FSM_HS_START_MASK | FSM_HS2_DONE_MASK))) {
+		/* Bare STARTUP from the boot-flow notification: only the
+		 * state switch.
+		 */
+		if (fsm->state != FSM_STATE_BOOTUP)
+			mtk_fsm_switch_state(fsm, to_state, event);
+		return 0;
+	}
+
+	if (event->fsm_flag & FSM_HS_START_MASK) {
+		/* An HS1 that raced ahead of the bare STARTUP must both
+		 * switch ON to BOOTUP and run the handshake below; returning
+		 * after the switch would consume the only trigger.
+		 */
+		mtk_fsm_switch_state(fsm, to_state, event);
+
+		ret = mtk_fsm_ctrl_ch_start(fsm, hs_info, O_NONBLOCK);
+		if (!ret)
+			ret = mtk_fsm_send_hs1_msg(hs_info);
+		if (ret)
+			goto hs_err;
+	} else if (event->fsm_flag & FSM_HS2_DONE_MASK) {
+		ret = mtk_fsm_parse_hs2_msg(fsm, hs_info, skb);
+		if (!ret)
+			ret = mtk_fsm_send_hs3_msg(hs_info, skb);
+		dev_kfree_skb(skb);
+		skb = NULL;
+		/* re-open the duplicate guard */
+		WRITE_ONCE(hs_info->rt_data, NULL);
+		if (ret)
+			goto hs_err;
+		mtk_fsm_switch_state(fsm, to_state, event);
+	}
+
+	if (((fsm->fsm_flag | event->fsm_flag) & fsm->hs_done_flag) == fsm->hs_done_flag) {
+		int user;
+
+		for (user = 0; user < MTK_FSM_HIF_USER_MAX; user++) {
+			if (fsm->hif_err[user]) {
+				dev_err(mdev->dev,
+					"Refusing READY: HIF user %d init failed with %d\n",
+					user, fsm->hif_err[user]);
+				return fsm->hif_err[user];
+			}
+		}
+
+		to_state = FSM_STATE_READY;
+		mtk_fsm_switch_state(fsm, to_state, NULL);
+	}
+
+	return 0;
+
+free_skb:
+	if (skb) {
+		dev_kfree_skb(skb);
+		WRITE_ONCE(hs_info->rt_data, NULL);
+	}
+hs_err:
+	for (int hs_id = 0; hs_id < HS_ID_MAX; hs_id++)
+		mtk_pci_unmask_ext_evt(mdev, fsm->hs_info[hs_id].mhccif_ch);
+	dev_err(mdev->dev, "Failed to hs with device %d:0x%x, ret=%d",
+		fsm->state, fsm->fsm_flag, ret);
+	return ret;
+}
+
+static void mtk_fsm_evt_release(struct kref *kref)
+{
+	struct mtk_fsm_evt *event = container_of(kref, struct mtk_fsm_evt, kref);
+
+	kfree(event);
+}
+
+static void mtk_fsm_evt_put(struct mtk_fsm_evt *event)
+{
+	kref_put(&event->kref, mtk_fsm_evt_release);
+}
+
+static void mtk_fsm_evt_finish(struct mtk_md_fsm *fsm,
+			       struct mtk_fsm_evt *event, int retval)
+{
+	if (event->mode & EVT_MODE_BLOCKING) {
+		event->status = retval;
+		wake_up(&fsm->evt_waitq);
+	}
+	mtk_fsm_evt_put(event);
+}
+
+static void mtk_fsm_evt_cleanup(struct mtk_md_fsm *fsm, struct list_head *evtq)
+{
+	struct mtk_fsm_evt *event, *tmp;
+
+	list_for_each_entry_safe(event, tmp, evtq, entry) {
+		list_del(&event->entry);
+		mtk_fsm_evt_finish(fsm, event, FSM_EVT_RET_FAIL);
+	}
+}
+
+static int mtk_fsm_enter_off_state(struct mtk_md_fsm *fsm, struct mtk_fsm_evt *event)
+{
+	struct mtk_md_dev *mdev = fsm->mdev;
+	int hs_id;
+
+	if (fsm->state == FSM_STATE_OFF || fsm->state == FSM_STATE_INVALID)
+		return -EPROTO;
+
+	mtk_pci_mask_ext_evt(mdev, DEV_EVT_D2H_BOOT_FLOW_SYNC);
+	for (hs_id = 0; hs_id < HS_ID_MAX; hs_id++)
+		mtk_pci_mask_ext_evt(mdev, fsm->hs_info[hs_id].mhccif_ch);
+
+	mtk_fsm_ctrl_ch_stop(fsm);
+	mtk_fsm_switch_state(fsm, FSM_STATE_OFF, event);
+
+	return 0;
+}
+
+static int mtk_fsm_dev_rm_act(struct mtk_md_fsm *fsm, struct mtk_fsm_evt *event)
+{
+	unsigned long flags;
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	set_bit(EVT_TF_GATECLOSED, &fsm->t_flag);
+	mtk_fsm_evt_cleanup(fsm, &fsm->evtq);
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	return mtk_fsm_enter_off_state(fsm, event);
+}
+
+static int mtk_fsm_hs1_handler(u32 status, void *__hs_info)
+{
+	struct fsm_hs_info *hs_info = __hs_info;
+	struct mtk_md_dev *mdev;
+	struct mtk_md_fsm *fsm;
+
+	fsm = container_of(hs_info, struct mtk_md_fsm, hs_info[hs_info->id]);
+	mdev = fsm->mdev;
+	/* Only consume the notification once the event is queued; on a
+	 * failed submit the channel stays unmasked and uncleared so the
+	 * device's retry still reaches us.
+	 */
+	if (mtk_fsm_evt_submit(mdev, FSM_EVT_STARTUP, hs_info->fsm_flag_hs1,
+			       hs_info, sizeof(*hs_info), 0) == FSM_EVT_RET_FAIL) {
+		dev_err(mdev->dev, "Failed to submit HS1 evt(hs%d), waiting for retry\n",
+			hs_info->id);
+		return -ENOMEM;
+	}
+	mtk_pci_mask_ext_evt(mdev, hs_info->mhccif_ch);
+	mtk_pci_clear_ext_evt(mdev, hs_info->mhccif_ch);
+
+	return 0;
+}
+
+static void mtk_fsm_hs_info_init_by_hsid(struct mtk_md_fsm *fsm, int hs_id)
+{
+	struct fsm_hs_info *hs_info;
+
+	if (hs_id < 0 || hs_id >= HS_ID_MAX) {
+		dev_warn((fsm->mdev)->dev, "hs_id = %d, invalid.\n", hs_id);
+		return;
+	}
+
+	hs_info = &fsm->hs_info[hs_id];
+	hs_info->id = hs_id;
+	hs_info->ctrl_port = NULL;
+	hs_info->rt_data = NULL;
+	switch (hs_id) {
+	case HS_ID_MD:
+		snprintf(hs_info->port_name, PORT_NAME_LEN, "MDCTRL");
+		hs_info->mhccif_ch = DEV_EVT_D2H_ASYNC_HS_NOTIFY_MD;
+		hs_info->fsm_flag_hs1 = FSM_F_MD_HS_START;
+		hs_info->fsm_flag_hs2 = FSM_F_MD_HS2_DONE;
+		hs_info->query_ft_set[QUERY_RTFT_ID_MD_PORT_ENUM].feature =
+			FIELD_PREP(FEATURE_TYPE, RTFT_TYPE_MUST_SUPPORT);
+		hs_info->query_ft_set[QUERY_RTFT_ID_MD_PORT_ENUM].feature |=
+			FIELD_PREP(FEATURE_VER, 0);
+		hs_info->query_ft_set[QUERY_RTFT_ID_MD_PORT_CFG].feature =
+			FIELD_PREP(FEATURE_TYPE, RTFT_TYPE_NOT_SUPPORT);
+		break;
+	case HS_ID_SAP:
+		snprintf(hs_info->port_name, PORT_NAME_LEN, "SAPCTRL");
+		hs_info->mhccif_ch = DEV_EVT_D2H_ASYNC_HS_NOTIFY_SAP;
+		hs_info->fsm_flag_hs1 = FSM_F_SAP_HS_START;
+		hs_info->fsm_flag_hs2 = FSM_F_SAP_HS2_DONE;
+		hs_info->query_ft_set[QUERY_RTFT_ID_SAP_PORT_ENUM].feature =
+			FIELD_PREP(FEATURE_TYPE, RTFT_TYPE_MUST_SUPPORT);
+		hs_info->query_ft_set[QUERY_RTFT_ID_SAP_PORT_ENUM].feature |=
+			FIELD_PREP(FEATURE_VER, 0);
+		break;
+	}
+}
+
+static int mtk_fsm_hs_info_init(struct mtk_md_fsm *fsm)
+{
+	struct mtk_md_dev *mdev = fsm->mdev;
+	struct fsm_hs_info *hs_info;
+	int hs_id, ret;
+
+	for (hs_id = 0; hs_id < HS_ID_MAX; hs_id++) {
+		mtk_fsm_hs_info_init_by_hsid(fsm, hs_id);
+		hs_info = &fsm->hs_info[hs_id];
+		ret = mtk_pci_register_ext_evt(mdev, hs_info->mhccif_ch,
+					       mtk_fsm_hs1_handler, hs_info);
+		if (ret)
+			goto err_unregister;
+	}
+
+	return 0;
+
+err_unregister:
+	/* All or nothing: the caller must not have to know how far we got. */
+	while (--hs_id >= 0)
+		mtk_pci_unregister_ext_evt(mdev, fsm->hs_info[hs_id].mhccif_ch);
+
+	return ret;
+}
+
+static void mtk_fsm_hs_info_exit(struct mtk_md_fsm *fsm)
+{
+	struct mtk_md_dev *mdev = fsm->mdev;
+	struct fsm_hs_info *hs_info;
+	int hs_id;
+
+	for (hs_id = 0; hs_id < HS_ID_MAX; hs_id++) {
+		hs_info = &fsm->hs_info[hs_id];
+		mtk_pci_unregister_ext_evt(mdev, hs_info->mhccif_ch);
+	}
+}
+
+static int mtk_fsm_dev_add_act(struct mtk_md_fsm *fsm, struct mtk_fsm_evt *event)
+{
+	if (fsm->state != FSM_STATE_OFF && fsm->state != FSM_STATE_INVALID)
+		return -EPROTO;
+
+	/* a fresh device lifecycle starts with a clean HIF error record */
+	memset(fsm->hif_err, 0, sizeof(fsm->hif_err));
+	mtk_fsm_switch_state(fsm, FSM_STATE_ON, event);
+	mtk_pci_unmask_ext_evt(fsm->mdev, DEV_EVT_D2H_BOOT_FLOW_SYNC);
+
+	return 0;
+}
+
+static int (*evts_act_tbl[FSM_EVT_MAX])(struct mtk_md_fsm *__fsm, struct mtk_fsm_evt *event) = {
+	[FSM_EVT_STARTUP] = mtk_fsm_startup_act,
+	[FSM_EVT_DEV_RM] = mtk_fsm_dev_rm_act,
+	[FSM_EVT_DEV_ADD] = mtk_fsm_dev_add_act,
+};
+
+int mtk_fsm_start(struct mtk_md_dev *mdev)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+
+	if (!fsm)
+		return -EINVAL;
+
+	if (!fsm->fsm_handler)
+		return -EFAULT;
+
+	wake_up_process(fsm->fsm_handler);
+	return 0;
+}
+
+/* Record a transport-plane init failure so the STARTUP path refuses the
+ * BOOTUP-to-READY promotion: the notifier callbacks are void and cannot
+ * reject a transition, so without this the FSM would advertise readiness
+ * on top of a data path that was never created.  Each producer owns one
+ * slot and stores its latest result there, so a retried init that
+ * succeeds clears the failure it recorded before; DEV_ADD clears all.
+ */
+void mtk_fsm_hif_err_record(struct mtk_md_dev *mdev, enum mtk_fsm_hif_user user, int err)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+
+	if (fsm)
+		fsm->hif_err[user] = err;
+}
+
+static void mkt_fsm_notifier_cleanup(struct mtk_md_dev *mdev, struct list_head *ntq)
+{
+	struct mtk_fsm_notifier *nt, *tmp;
+
+	list_for_each_entry_safe(nt, tmp, ntq, entry) {
+		list_del(&nt->entry);
+		dev_warn(mdev->dev, "Having to free notifier(%d) by FSM!\n", nt->id);
+		kfree(nt);
+	}
+}
+
+static void mtk_fsm_notifier_insert(struct mtk_fsm_notifier *notifier, struct list_head *head)
+{
+	struct mtk_fsm_notifier *nt;
+
+	list_for_each_entry(nt, head, entry) {
+		if (notifier->prio > nt->prio) {
+			list_add(&notifier->entry, nt->entry.prev);
+			return;
+		}
+	}
+	list_add_tail(&notifier->entry, head);
+}
+
+int mtk_fsm_notifier_register(struct mtk_md_dev *mdev, enum mtk_user_id id,
+			      void (*cb)(struct mtk_fsm_param *, void *data),
+			      void *data, enum mtk_fsm_prio prio, bool is_pre)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+	struct mtk_fsm_notifier *notifier;
+
+	if (!fsm)
+		return -EINVAL;
+
+	if (id >= MTK_USER_MAX || !cb || prio >= FSM_PRIO_MAX)
+		return -EINVAL;
+
+	notifier = kzalloc_obj(*notifier);
+	if (!notifier)
+		return -ENOMEM;
+
+	INIT_LIST_HEAD(&notifier->entry);
+	notifier->id = id;
+	notifier->cb = cb;
+	notifier->data = data;
+	notifier->prio = prio;
+
+	mutex_lock(&fsm->notifier_lock);
+	if (is_pre)
+		mtk_fsm_notifier_insert(notifier, &fsm->pre_notifiers);
+	else
+		mtk_fsm_notifier_insert(notifier, &fsm->post_notifiers);
+	mutex_unlock(&fsm->notifier_lock);
+
+	return 0;
+}
+
+int mtk_fsm_notifier_unregister(struct mtk_md_dev *mdev, enum mtk_user_id id)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+	struct mtk_fsm_notifier *nt, *tmp;
+
+	if (!fsm)
+		return -EINVAL;
+
+	mutex_lock(&fsm->notifier_lock);
+	list_for_each_entry_safe(nt, tmp, &fsm->pre_notifiers, entry) {
+		if (nt->id == id) {
+			list_del(&nt->entry);
+			kfree(nt);
+			break;
+		}
+	}
+	list_for_each_entry_safe(nt, tmp, &fsm->post_notifiers, entry) {
+		if (nt->id == id) {
+			list_del(&nt->entry);
+			kfree(nt);
+			break;
+		}
+	}
+	mutex_unlock(&fsm->notifier_lock);
+	return 0;
+}
+
+int mtk_fsm_evt_submit(struct mtk_md_dev *mdev,
+		       enum mtk_fsm_evt_id id, enum mtk_fsm_flag flag,
+		       void *data, unsigned int len, unsigned char mode)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+	struct mtk_fsm_evt *event;
+	unsigned long flags;
+	int ret = 0;
+
+	if (!fsm || id >= FSM_EVT_MAX) {
+		dev_err(mdev->dev, "Invalid param!\n");
+		return FSM_EVT_RET_FAIL;
+	}
+
+	if (test_bit(EVT_TF_GATECLOSED, &fsm->t_flag)) {
+		dev_err(mdev->dev, "Failed to submit evt, fsm has been removed!\n");
+		return FSM_EVT_RET_FAIL;
+	}
+
+	event = kzalloc(sizeof(*event),
+			(in_hardirq() || in_softirq() || irqs_disabled()) ?
+			GFP_ATOMIC : GFP_KERNEL);
+	if (!event)
+		return FSM_EVT_RET_FAIL;
+
+	kref_init(&event->kref);
+	event->mdev = mdev;
+	event->id = id;
+	event->fsm_flag = flag;
+	event->status = FSM_EVT_RET_ONGOING;
+	event->data = data;
+	event->len = len;
+	event->mode = mode;
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	if (test_bit(EVT_TF_GATECLOSED, &fsm->t_flag)) {
+		spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+		mtk_fsm_evt_put(event);
+		dev_err(mdev->dev, "Failed to add event, fsm dev has been removed!\n");
+		return FSM_EVT_RET_FAIL;
+	}
+
+	kref_get(&event->kref);
+	if (mode & EVT_MODE_TOHEAD)
+		list_add(&event->entry, &fsm->evtq);
+	else
+		list_add_tail(&event->entry, &fsm->evtq);
+	if (fsm->fsm_handler)
+		wake_up_process(fsm->fsm_handler);
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	if (mode & EVT_MODE_BLOCKING) {
+		ret = wait_event_timeout(fsm->evt_waitq,
+					 (event->status != 0), BLOCKING_EVT_TIMEOUT);
+		if (!ret && event->status != FSM_EVT_RET_DONE) {
+			dev_err(mdev->dev, "Handling fsm blocking event timeout!\n");
+			ret = -ETIMEDOUT;
+		} else {
+			ret = event->status;
+		}
+	}
+	mtk_fsm_evt_put(event);
+
+	return ret;
+}
+
+static int mtk_fsm_evt_handler(void *__fsm)
+{
+	struct mtk_md_fsm *fsm = __fsm;
+	struct mtk_fsm_evt *event;
+	unsigned long flags;
+	int ret;
+
+wake_up:
+	set_current_state(TASK_INTERRUPTIBLE);
+	while (!kthread_should_stop() && !list_empty(&fsm->evtq)) {
+		set_current_state(TASK_RUNNING);
+		spin_lock_irqsave(&fsm->evtq_lock, flags);
+		event = list_first_entry(&fsm->evtq, struct mtk_fsm_evt, entry);
+		list_del(&event->entry);
+		spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+		if (event->id < FSM_EVT_MAX) {
+			ret = evts_act_tbl[event->id](fsm, event);
+			if (ret) {
+				dev_err((fsm->mdev)->dev,
+					"Failed to handle evt, fsm state = %d, ret = %d\n",
+					fsm->state, ret);
+				mtk_fsm_evt_finish(fsm, event, FSM_EVT_RET_FAIL);
+			} else {
+				mtk_fsm_evt_finish(fsm, event, FSM_EVT_RET_DONE);
+			}
+		} else {
+			mtk_fsm_evt_finish(fsm, event, FSM_EVT_RET_DONE);
+		}
+	}
+
+	if (kthread_should_stop()) {
+		set_current_state(TASK_RUNNING);
+		return 0;
+	}
+
+	schedule();
+	goto wake_up;
+}
+
+int mtk_fsm_init(struct mtk_md_dev *mdev)
+{
+	struct mtk_md_fsm *fsm;
+	int ret;
+
+	fsm = devm_kzalloc(mdev->dev, sizeof(*fsm), GFP_KERNEL);
+	if (!fsm)
+		return -ENOMEM;
+
+	fsm->fsm_handler = kthread_create(mtk_fsm_evt_handler, fsm, "fsm_evt_thread%d_%s",
+					  mdev->hw_ver, mdev->dev_str);
+	if (IS_ERR(fsm->fsm_handler))
+		return PTR_ERR(fsm->fsm_handler);
+
+	/* Keep our own reference so that kthread_stop() in the teardown
+	 * paths stays safe even if the thread has already exited (e.g.
+	 * it was killed by an oops in a listener it called): without it
+	 * the task_struct may be freed and kthread_stop() would hit a
+	 * use-after-free through its internal get_task_struct().
+	 */
+	get_task_struct(fsm->fsm_handler);
+
+	fsm->mdev = mdev;
+	fsm->state = FSM_STATE_INVALID;
+	fsm->fsm_flag = FSM_F_DFLT;
+
+	INIT_LIST_HEAD(&fsm->evtq);
+	spin_lock_init(&fsm->evtq_lock);
+	init_waitqueue_head(&fsm->evt_waitq);
+
+	INIT_LIST_HEAD(&fsm->pre_notifiers);
+	INIT_LIST_HEAD(&fsm->post_notifiers);
+	mutex_init(&fsm->notifier_lock);
+
+	ret = mtk_pci_register_ext_evt(mdev, DEV_EVT_D2H_BOOT_FLOW_SYNC,
+				       mtk_fsm_early_bootup_handler, fsm);
+	if (ret)
+		goto err_stop_thread;
+
+	ret = mtk_fsm_hs_info_init(fsm);
+	if (ret)
+		goto err_unregister_evt;
+
+	mdev->fsm = fsm;
+	return 0;
+
+err_unregister_evt:
+	mtk_pci_unregister_ext_evt(mdev, DEV_EVT_D2H_BOOT_FLOW_SYNC);
+err_stop_thread:
+	/* mdev->fsm is never published on failure, so mtk_fsm_exit() cannot
+	 * stop the parked kthread for us; it must be stopped here.
+	 */
+	kthread_stop(fsm->fsm_handler);
+	put_task_struct(fsm->fsm_handler);
+	return ret;
+}
+
+/* Close the event gate and join the handler thread, then release the
+ * control ports while the port table they reference is still alive.
+ * Called from the removal path before the transport plane is dismantled;
+ * every step is idempotent, so mtk_fsm_exit() may repeat them safely.
+ */
+int mtk_fsm_stop(struct mtk_md_dev *mdev)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+	struct task_struct *handler;
+	unsigned long flags;
+
+	if (!fsm)
+		return -EINVAL;
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	set_bit(EVT_TF_GATECLOSED, &fsm->t_flag);
+	handler = fsm->fsm_handler;
+	fsm->fsm_handler = NULL;
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	if (handler) {
+		kthread_stop(handler);
+		put_task_struct(handler);
+	}
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	mtk_fsm_evt_cleanup(fsm, &fsm->evtq);
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	mtk_fsm_ctrl_ch_stop(fsm);
+
+	return 0;
+}
+
+int mtk_fsm_exit(struct mtk_md_dev *mdev)
+{
+	struct mtk_md_fsm *fsm = mdev->fsm;
+	struct task_struct *handler;
+	unsigned long flags;
+
+	if (!fsm)
+		return -EINVAL;
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	set_bit(EVT_TF_GATECLOSED, &fsm->t_flag);
+	handler = fsm->fsm_handler;
+	fsm->fsm_handler = NULL;
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	if (handler) {
+		kthread_stop(handler);
+		put_task_struct(handler);
+	}
+
+	spin_lock_irqsave(&fsm->evtq_lock, flags);
+	if (WARN_ON(!list_empty(&fsm->evtq)))
+		mtk_fsm_evt_cleanup(fsm, &fsm->evtq);
+	spin_unlock_irqrestore(&fsm->evtq_lock, flags);
+
+	/* Release the port references and recv callbacks even when the
+	 * DEV_RM event never ran (its enter-off path is the only other
+	 * caller); mtk_fsm_ctrl_ch_stop() is idempotent.
+	 */
+	mtk_fsm_ctrl_ch_stop(fsm);
+
+	for (int i = 0; i < HS_ID_MAX; i++) {
+		if (fsm->hs_info[i].rt_data) {
+			dev_kfree_skb(fsm->hs_info[i].rt_data);
+			fsm->hs_info[i].rt_data = NULL;
+		}
+	}
+
+	mutex_lock(&fsm->notifier_lock);
+	mkt_fsm_notifier_cleanup(mdev, &fsm->pre_notifiers);
+	mkt_fsm_notifier_cleanup(mdev, &fsm->post_notifiers);
+	mutex_unlock(&fsm->notifier_lock);
+
+	mtk_pci_unregister_ext_evt(mdev, DEV_EVT_D2H_BOOT_FLOW_SYNC);
+	mtk_fsm_hs_info_exit(fsm);
+	mdev->fsm = NULL;
+
+	return 0;
+}
diff --git a/drivers/net/wwan/t9xx/mtk_fsm.h b/drivers/net/wwan/t9xx/mtk_fsm.h
new file mode 100644
index 000000000000..1284a6188a8e
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_fsm.h
@@ -0,0 +1,158 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_FSM_H__
+#define __MTK_FSM_H__
+
+#include <linux/mutex.h>
+
+#include "mtk_dev.h"
+
+#define FEATURE_CNT		(64)
+#define FEATURE_QUERY_PATTERN	(0x49434343)
+
+#define FEATURE_TYPE		GENMASK(3, 0)
+#define FEATURE_VER		GENMASK(7, 4)
+
+#define FEATURE_TYPE_NOT	FIELD_PREP(FEATURE_TYPE, RTFT_TYPE_NOT_SUPPORT)
+#define FEATURE_VER_0		FIELD_PREP(FEATURE_VER, 0)
+
+#define EVT_MODE_BLOCKING	(0x01)
+#define EVT_MODE_TOHEAD		(0x02)
+
+#define FSM_EVT_RET_FAIL	(-1)
+#define FSM_EVT_RET_ONGOING	(0)
+#define FSM_EVT_RET_DONE	(1)
+
+enum mtk_fsm_flag {
+	FSM_F_DFLT = 0,
+	FSM_F_SAP_HS_START	= BIT(0),
+	FSM_F_SAP_HS2_DONE	= BIT(1),
+	FSM_F_MD_HS_START	= BIT(2),
+	FSM_F_MD_HS2_DONE	= BIT(3),
+};
+
+enum mtk_fsm_state {
+	FSM_STATE_INVALID = 0,
+	FSM_STATE_OFF,
+	FSM_STATE_ON,
+	FSM_STATE_BOOTUP,
+	FSM_STATE_READY,
+};
+
+enum mtk_fsm_evt_id {
+	FSM_EVT_STARTUP = 0,
+	FSM_EVT_DEV_RM,
+	FSM_EVT_DEV_ADD,
+	FSM_EVT_MAX
+};
+
+enum mtk_fsm_prio {
+	FSM_PRIO_1 = 1,
+	FSM_PRIO_MAX
+};
+
+/* One slot per transport-plane producer, so that a retried init clears
+ * only its own record.
+ */
+enum mtk_fsm_hif_user {
+	MTK_FSM_HIF_USER_CTRL = 0,
+	MTK_FSM_HIF_USER_CLDMA0,
+	MTK_FSM_HIF_USER_CLDMA1,
+	MTK_FSM_HIF_USER_MAX
+};
+
+struct mtk_fsm_param {
+	enum mtk_fsm_state from;
+	enum mtk_fsm_state to;
+	enum mtk_fsm_evt_id evt_id;
+	enum mtk_fsm_flag fsm_flag;
+};
+
+#define PORT_NAME_LEN 20
+
+enum handshake_info_id {
+	HS_ID_MD = 0,
+	HS_ID_SAP,
+	HS_ID_MAX
+};
+
+struct runtime_feature_info {
+	u8 feature;
+};
+
+struct fsm_hs_info {
+	unsigned char id;
+	void *ctrl_port;
+	char port_name[PORT_NAME_LEN];
+	unsigned int mhccif_ch;
+	unsigned int fsm_flag_hs1;
+	unsigned int fsm_flag_hs2;
+	/* the feature that the device should support */
+	struct runtime_feature_info query_ft_set[FEATURE_CNT];
+	/* Duplicate-HS2 guard only: the in-flight runtime-data skb itself
+	 * is owned by the STARTUP event it was submitted with.
+	 */
+	void *rt_data;
+};
+
+struct mtk_md_fsm {
+	struct mtk_md_dev *mdev;
+	struct task_struct *fsm_handler;
+	struct fsm_hs_info hs_info[HS_ID_MAX];
+	unsigned int hs_done_flag;
+	unsigned long t_flag;
+	u32 last_dev_state;
+	/* last transport-plane init result per producer; any non-zero entry
+	 * refuses READY
+	 */
+	int hif_err[MTK_FSM_HIF_USER_MAX];
+	enum mtk_fsm_state state;
+	unsigned int fsm_flag;
+	struct list_head evtq;
+	/* protect evtq */
+	spinlock_t evtq_lock;
+	/* waitq for fsm blocking submit */
+	wait_queue_head_t evt_waitq;
+	struct list_head pre_notifiers;
+	struct list_head post_notifiers;
+	/* protects pre_notifiers and post_notifiers lists */
+	struct mutex notifier_lock;
+};
+
+struct mtk_fsm_evt {
+	struct list_head entry;
+	struct kref kref;
+	struct mtk_md_dev *mdev;
+	enum mtk_fsm_evt_id id;
+	unsigned int fsm_flag;
+	int status;
+	unsigned char mode;
+	unsigned int len;
+	void *data;
+};
+
+struct mtk_fsm_notifier {
+	struct list_head entry;
+	enum mtk_user_id id;
+	void (*cb)(struct mtk_fsm_param *param, void *data);
+	void *data;
+	enum mtk_fsm_prio prio;
+};
+
+int mtk_fsm_init(struct mtk_md_dev *mdev);
+int mtk_fsm_exit(struct mtk_md_dev *mdev);
+int mtk_fsm_start(struct mtk_md_dev *mdev);
+int mtk_fsm_stop(struct mtk_md_dev *mdev);
+void mtk_fsm_hif_err_record(struct mtk_md_dev *mdev, enum mtk_fsm_hif_user user, int err);
+int mtk_fsm_notifier_register(struct mtk_md_dev *mdev, enum mtk_user_id id,
+			      void (*cb)(struct mtk_fsm_param *, void *data),
+			      void *data, enum mtk_fsm_prio prio, bool is_pre);
+int mtk_fsm_notifier_unregister(struct mtk_md_dev *mdev, enum mtk_user_id id);
+int mtk_fsm_evt_submit(struct mtk_md_dev *mdev,
+		       enum mtk_fsm_evt_id id, enum mtk_fsm_flag flag,
+		       void *data, unsigned int len, unsigned char mode);
+
+#endif /* __MTK_FSM_H__ */
diff --git a/drivers/net/wwan/t9xx/mtk_port.c b/drivers/net/wwan/t9xx/mtk_port.c
index a8ff06ab4de4..d7c50519e1fd 100644
--- a/drivers/net/wwan/t9xx/mtk_port.c
+++ b/drivers/net/wwan/t9xx/mtk_port.c
@@ -550,6 +550,9 @@ int mtk_port_status_update(struct mtk_md_dev *mdev, void *data, u32 data_len)
 	if (unlikely(!mdev || !msg))
 		return -EINVAL;
 
+	if (data_len < sizeof(*msg))
+		return -EPROTO;
+
 	ctrl_blk = mdev->ctrl_blk;
 	port_mngr = ctrl_blk->port_mngr;
 	if (le16_to_cpu(msg->version) != MTK_PORT_ENUM_VER ||
@@ -653,6 +656,69 @@ int mtk_port_ch_disable(struct mtk_port *port)
 	return ret;
 }
 
+static void mtk_port_disable(struct mtk_port_mngr *port_mngr)
+{
+	struct radix_tree_iter iter;
+	struct mtk_port *port;
+	void __rcu **slot;
+	int tbl_type;
+
+	tbl_type = PORT_TBL_SAP;
+	do {
+		radix_tree_for_each_slot(slot, &port_mngr->port_tbl[tbl_type],
+					 &iter, 0) {
+			MTK_PORT_SEARCH_FROM_RADIX_TREE(port, slot);
+			MTK_PORT_INTERNAL_NODE_CHECK(port, slot, iter);
+			ports_ops[port->info.type]->disable(port);
+		}
+	} while (++tbl_type < PORT_TBL_MAX);
+}
+
+void mtk_port_mngr_fsm_state_handler(struct mtk_fsm_param *fsm_param, void *arg)
+{
+	struct mtk_port_mngr *port_mngr;
+
+	if (!fsm_param || !arg)
+		return;
+
+	port_mngr = arg;
+
+	switch (fsm_param->to) {
+	case FSM_STATE_OFF:
+		mtk_port_disable(port_mngr);
+		break;
+	default:
+		break;
+	}
+}
+
+void mtk_port_mngr_fsm_state_handler_late(struct mtk_fsm_param *fsm_param, void *arg)
+{
+	struct mtk_port_mngr *port_mngr;
+	struct mtk_port *port;
+
+	if (!fsm_param || !arg)
+		return;
+
+	port_mngr = arg;
+
+	switch (fsm_param->to) {
+	case FSM_STATE_BOOTUP:
+		if (fsm_param->fsm_flag & FSM_F_MD_HS_START) {
+			port = mtk_port_search_by_id(port_mngr, CCCI_CONTROL_RX);
+			if (port)
+				ports_ops[port->info.type]->enable(port);
+		} else if (fsm_param->fsm_flag & FSM_F_SAP_HS_START) {
+			port = mtk_port_search_by_id(port_mngr, CCCI_SAP_CONTROL_RX);
+			if (port)
+				ports_ops[port->info.type]->enable(port);
+		}
+		break;
+	default:
+		break;
+	}
+}
+
 int mtk_port_mngr_init(struct mtk_ctrl_blk *ctrl_blk, struct mtk_port_cfg *port_cfg, int port_cnt)
 {
 	struct mtk_port_mngr *port_mngr;
diff --git a/drivers/net/wwan/t9xx/mtk_port.h b/drivers/net/wwan/t9xx/mtk_port.h
index 9b8f63ce3f3c..2fc0e450ab6b 100644
--- a/drivers/net/wwan/t9xx/mtk_port.h
+++ b/drivers/net/wwan/t9xx/mtk_port.h
@@ -124,6 +124,8 @@ int mtk_port_send_data(struct mtk_port *port, void *data, bool blocking, bool fo
 int mtk_port_status_update(struct mtk_md_dev *mdev, void *data, u32 data_len);
 int mtk_port_ch_enable(struct mtk_port *port);
 int mtk_port_ch_disable(struct mtk_port *port);
+void mtk_port_mngr_fsm_state_handler(struct mtk_fsm_param *fsm_param, void *arg);
+void mtk_port_mngr_fsm_state_handler_late(struct mtk_fsm_param *fsm_param, void *arg);
 int mtk_port_mngr_init(struct mtk_ctrl_blk *ctrl_blk, struct mtk_port_cfg *port_cfg, int port_cnt);
 void mtk_port_mngr_exit(struct mtk_ctrl_blk *ctrl_blk);
 void mtk_port_trb_init(struct mtk_port *port, struct trb *trb, enum mtk_trb_cmd_type cmd,
diff --git a/drivers/net/wwan/t9xx/mtk_utility.h b/drivers/net/wwan/t9xx/mtk_utility.h
new file mode 100644
index 000000000000..a7b5f3f4d54e
--- /dev/null
+++ b/drivers/net/wwan/t9xx/mtk_utility.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0-only
+ *
+ * Copyright (c) 2022, MediaTek Inc.
+ */
+
+#ifndef __MTK_UTILITY_H__
+#define __MTK_UTILITY_H__
+
+#include <linux/device.h>
+#include "mtk_dev.h"
+
+#define MTK_UEVENT_INFO_LEN 128
+
+/* MTK uevent */
+enum mtk_uevent_id {
+	MTK_UEVENT_FSM = 1,
+};
+
+static inline void mtk_uevent_notify(struct device *dev, enum mtk_uevent_id id, const char *info)
+{
+	char buf[MTK_UEVENT_INFO_LEN];
+	char *ext[2] = {NULL, NULL};
+
+	snprintf(buf, MTK_UEVENT_INFO_LEN, "%s:event_id=%d, info=%s",
+		 dev->kobj.name, id, info);
+	ext[0] = buf;
+	kobject_uevent_env(&dev->kobj, KOBJ_CHANGE, ext);
+}
+#endif /* __MTK_UTILITY_H__ */
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
index 0f281bfb38cc..080ae29d3a88 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
@@ -32,11 +32,184 @@
 #define WAIT_HWO_TIME		(5)
 #define NO_BUDGET		(0)
 
+static void mtk_cldma_err_work(struct work_struct *work);
+
+static int mtk_cldma_isr(int irq_id, void *param)
+{
+	struct cldma_drv_info *drv_info = param;
+	u32 tx_err, rx_err;
+	struct mtk_md_dev *mdev;
+	u32 tx_done, rx_done;
+	u32 tx_sta, rx_sta;
+	struct txq *txq;
+	struct rxq *rxq;
+	int i;
+
+	mdev = drv_info->mdev;
+	mtk_cldma_get_intr_status(drv_info, &tx_sta, &rx_sta);
+	tx_done = (tx_sta >> QUEUE_XFER_DONE) & 0xFF;
+	rx_done = (rx_sta >> QUEUE_XFER_DONE) & 0xFF;
+	tx_err = (tx_sta >> QUEUE_ERROR) & 0xFF;
+	rx_err = (rx_sta >> QUEUE_ERROR) & 0xFF;
+
+	if (tx_err || rx_err) {
+		dev_err_ratelimited(mdev->dev, "CLDMA%d queue error: TX 0x%x RX 0x%x\n",
+				    drv_info->hif_id, tx_err, rx_err);
+		for (i = 0; i < HW_QUEUE_NUM; i++) {
+			if (tx_err & BIT(i)) {
+				mtk_cldma_clr_intr_status(drv_info, DIR_TX,
+							  i, QUEUE_ERROR);
+				mtk_cldma_unmask_intr(drv_info, DIR_TX,
+						      i, QUEUE_ERROR);
+			}
+			if (rx_err & BIT(i)) {
+				mtk_cldma_clr_intr_status(drv_info, DIR_RX,
+							  i, QUEUE_ERROR);
+				mtk_cldma_unmask_intr(drv_info, DIR_RX,
+						      i, QUEUE_ERROR);
+			}
+		}
+		atomic_or(tx_err, &drv_info->tx_err_qs);
+		atomic_or(rx_err, &drv_info->rx_err_qs);
+		queue_work(drv_info->wq, &drv_info->err_work);
+	}
+
+	if (tx_done) {
+		for (i = 0; i < HW_QUEUE_NUM; i++) {
+			/* pairs with smp_store_release() in txq_alloc */
+			txq = smp_load_acquire(&drv_info->txq[i]);
+			if (!(tx_done & BIT(i)) || !txq)
+				continue;
+			queue_work(drv_info->wq, &txq->tx_done_work);
+		}
+	}
+	if (rx_done) {
+		for (i = 0; i < HW_QUEUE_NUM; i++) {
+			/* pairs with smp_store_release() in rxq_alloc */
+			rxq = smp_load_acquire(&drv_info->rxq[i]);
+			if (!(rx_done & BIT(i)) || !rxq)
+				continue;
+			queue_work(drv_info->wq, &rxq->rx_done_work);
+		}
+	}
+
+	mtk_pci_clear_irq(mdev, drv_info->pci_ext_irq_id);
+	mtk_pci_unmask_irq(mdev, drv_info->pci_ext_irq_id);
+
+	return IRQ_HANDLED;
+}
+
 static const int mtk_cldma_hw_id_tbl[NR_CLDMA] = {
 	[CLDMA0] = CLDMA0_HW_ID,
 	[CLDMA1] = CLDMA1_HW_ID,
 };
 
+static int mtk_cldma_dev_init(struct cldma_dev *cd, int hif_id)
+{
+	char gpd_pool_name[DMA_POOL_NAME_LEN];
+	char bd_pool_name[DMA_POOL_NAME_LEN];
+	struct cldma_drv_info *drv_info;
+	struct cldma_hw_regs *hw_regs;
+	struct mtk_md_dev *mdev;
+	unsigned int flag;
+	int hw_id, ret;
+
+	if (!cd || hif_id >= NR_CLDMA)
+		return -EINVAL;
+
+	if (cd->cldma_drv_info[hif_id])
+		return 0;
+
+	hw_id = mtk_cldma_hw_id_tbl[hif_id];
+	mdev = cd->trans->mdev;
+	drv_info = kzalloc_obj(*drv_info);
+	if (!drv_info)
+		return -ENOMEM;
+
+	drv_info->cd = cd;
+	drv_info->mdev = mdev;
+	drv_info->hif_id = hif_id;
+	drv_info->hw_id = hw_id;
+	drv_info->hw_regs = &mtk_cldma_regs_m9xx;
+
+	hw_regs = drv_info->hw_regs;
+	snprintf(gpd_pool_name, DMA_POOL_NAME_LEN, "cldma%d_gpd_pool_%s",
+		 hw_id, mdev->dev_str);
+	snprintf(bd_pool_name, DMA_POOL_NAME_LEN, "cldma%d_bd_pool_%s",
+		 hw_id, mdev->dev_str);
+	drv_info->gpd_dma_pool = dma_pool_create(gpd_pool_name, mdev->dev,
+						 sizeof(union gpd), 16, 0);
+	if (!drv_info->gpd_dma_pool) {
+		dev_err(mdev->dev, "Failed to alloc gpd dma pool for cldma%d\n", hw_id);
+		ret = -ENOMEM;
+		goto err_free_drv_info;
+	}
+	drv_info->bd_dma_pool = dma_pool_create(bd_pool_name, mdev->dev,
+						sizeof(union bd), 16, 0);
+	if (!drv_info->bd_dma_pool) {
+		dev_err(mdev->dev, "Failed to alloc bd dma pool for cldma%d\n", hw_id);
+		ret = -ENOMEM;
+		goto err_destroy_gpd_pool;
+	}
+
+	switch (hif_id) {
+	case CLDMA0:
+		drv_info->pci_ext_irq_id = mtk_pci_get_irq_id(mdev, MTK_IRQ_SRC_CLDMA0);
+		drv_info->base_addr = hw_regs->cldma0_base_addr;
+		break;
+	case CLDMA1:
+		drv_info->pci_ext_irq_id = mtk_pci_get_irq_id(mdev, MTK_IRQ_SRC_CLDMA1);
+		drv_info->base_addr = hw_regs->cldma1_base_addr;
+		break;
+	default:
+		ret = -EINVAL;
+		goto err_destroy_dma_pool;
+	}
+
+	flag = WQ_UNBOUND | WQ_MEM_RECLAIM | WQ_HIGHPRI;
+	drv_info->wq = alloc_workqueue("cldma%d_workq_%s", flag, 0, hw_id, mdev->dev_str);
+	if (!drv_info->wq) {
+		dev_err(mdev->dev, "Failed to alloc work queue for cldma%d\n", hw_id);
+		ret = -ENOMEM;
+		goto err_destroy_dma_pool;
+	}
+
+	INIT_WORK(&drv_info->err_work, mtk_cldma_err_work);
+	atomic_set(&drv_info->tx_err_qs, 0);
+	atomic_set(&drv_info->rx_err_qs, 0);
+
+	mtk_cldma_drv_init(drv_info);
+
+	/* mask/clear PCI CLDMA L1 interrupt */
+	mtk_pci_mask_irq(mdev, drv_info->pci_ext_irq_id);
+	mtk_pci_clear_irq(mdev, drv_info->pci_ext_irq_id);
+
+	/* register CLDMA interrupt handler */
+	ret = mtk_pci_register_irq(mdev, drv_info->pci_ext_irq_id, mtk_cldma_isr, drv_info);
+	if (ret)
+		goto err_destroy_wq;
+
+	/* unmask PCI CLDMA L1 interrupt */
+	mtk_pci_unmask_irq(mdev, drv_info->pci_ext_irq_id);
+
+	/* Publish only once drv_info is fully built; pairs with the acquire
+	 * loads in the transfer paths below.
+	 */
+	smp_store_release(&cd->cldma_drv_info[hif_id], drv_info);
+	return 0;
+
+err_destroy_wq:
+	destroy_workqueue(drv_info->wq);
+err_destroy_dma_pool:
+	dma_pool_destroy(drv_info->bd_dma_pool);
+err_destroy_gpd_pool:
+	dma_pool_destroy(drv_info->gpd_dma_pool);
+err_free_drv_info:
+	kfree(drv_info);
+
+	return ret;
+}
+
 static void mtk_cldma_clr_bd_dsc(struct cldma_drv_info *drv_info,
 				 struct bd_dsc *bd_dsc_pool, int nr_bds)
 {
@@ -529,6 +702,7 @@ static struct txq *mtk_cldma_txq_alloc(struct cldma_drv_info *drv_info, struct q
 	txq->nr_gpds = txq->que->tx_nr_gpds;
 	atomic_set(&txq->req_budget, txq->que->tx_nr_gpds);
 	spin_lock_init(&txq->ring_lock);
+	txq->is_stopping = false;
 	tx_frag_size = txq->que->tx_frag_size;
 	if (txq->que->tx_mtu > tx_frag_size && tx_frag_size)
 		txq->nr_bds = (txq->que->tx_mtu + tx_frag_size - 1) / tx_frag_size;
@@ -679,8 +853,11 @@ static void mtk_cldma_txq_free(struct cldma_drv_info *drv_info, u32 txqno)
 
 	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
 	synchronize_irq(irq_id);
-	/* flush on-going work */
+	/* flush on-going work; the error worker may have loaded this txq
+	 * before it was unpublished above, so it has to be retired too
+	 */
 	flush_work(&txq->tx_done_work);
+	flush_work(&drv_info->err_work);
 	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
 	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_ERROR);
 
@@ -690,6 +867,7 @@ static void mtk_cldma_txq_free(struct cldma_drv_info *drv_info, u32 txqno)
 		 */
 		dev_err(mdev->dev, "TX queue %d cannot be stopped, leaking its ring\n",
 			txqno);
+		drv_info->ring_leaked = true;
 		return;
 	}
 
@@ -987,8 +1165,11 @@ static void mtk_cldma_rxq_free(struct cldma_drv_info *drv_info, u32 rxqno)
 
 	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
 	synchronize_irq(irq_id);
-	/* flush on-going work */
+	/* flush on-going work; the error worker may still be about to stop
+	 * this queue number, which a later allocation could already reuse
+	 */
 	flush_work(&rxq->rx_done_work);
+	flush_work(&drv_info->err_work);
 	/* mask L2 RX interrupt again to avoid race condition causing use-after-free issue */
 	mtk_cldma_mask_intr(drv_info, DIR_RX, rxqno, QUEUE_XFER_DONE);
 	mtk_cldma_mask_intr(drv_info, DIR_RX, rxqno, QUEUE_ERROR);
@@ -999,6 +1180,7 @@ static void mtk_cldma_rxq_free(struct cldma_drv_info *drv_info, u32 rxqno)
 		 */
 		dev_err(mdev->dev, "RX queue %d cannot be stopped, leaking its ring\n",
 			rxqno);
+		drv_info->ring_leaked = true;
 		return;
 	}
 
@@ -1072,6 +1254,67 @@ static int mtk_cldma_hw_recovery(struct cldma_drv_info *drv_info, u32 qno)
 	return 0;
 }
 
+static int mtk_cldma_dev_exit(struct cldma_dev *cd, int hif_id)
+{
+	struct cldma_drv_info *drv_info;
+	struct mtk_md_dev *mdev;
+	int virq_id;
+	int i;
+
+	if (!cd || hif_id >= NR_CLDMA)
+		return -EINVAL;
+
+	if (!cd->cldma_drv_info[hif_id])
+		return 0;
+
+	/* free cldma descriptor */
+	drv_info = cd->cldma_drv_info[hif_id];
+	mdev = cd->trans->mdev;
+	virq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
+	/* mask first so no new interrupt can fire, then wait out any
+	 * in-flight handler before tearing the registration down
+	 */
+	mtk_pci_mask_irq(mdev, drv_info->pci_ext_irq_id);
+	synchronize_irq(virq_id);
+	mtk_pci_unregister_irq(mdev, drv_info->pci_ext_irq_id);
+	for (i = 0; i < HW_QUEUE_NUM; i++) {
+		if (drv_info->txq[i])
+			mtk_cldma_txq_free(drv_info, drv_info->txq[i]->txqno);
+		if (drv_info->rxq[i])
+			mtk_cldma_rxq_free(drv_info, drv_info->rxq[i]->rxqno);
+	}
+
+	flush_workqueue(drv_info->wq);
+	destroy_workqueue(drv_info->wq);
+
+	/* quiesce the IP before releasing descriptor memory: disable its
+	 * interrupt output and reset it, so it cannot touch the rings again
+	 */
+	mtk_pci_write32(mdev, drv_info->base_addr + drv_info->hw_regs->reg_cldma_int_mask,
+			LINK_ERROR_VAL);
+	mtk_cldma_drv_reset(drv_info);
+
+	if (drv_info->ring_leaked) {
+		/* A ring is leaked and its descriptors live in these pools:
+		 * the device may still master DMA into them, so handing them
+		 * back to the allocator would open a use-after-free window.
+		 */
+		dev_err(mdev->dev, "CLDMA%d rings leaked, leaking DMA pools too\n",
+			drv_info->hw_id);
+	} else {
+		dma_pool_destroy(drv_info->bd_dma_pool);
+		dma_pool_destroy(drv_info->gpd_dma_pool);
+	}
+
+	/* Unpublish before teardown; pairs with the acquire loads of this
+	 * slot. No release is needed, there is no prior store to expose.
+	 */
+	WRITE_ONCE(cd->cldma_drv_info[hif_id], NULL);
+	kfree(drv_info);
+
+	return 0;
+}
+
 static int mtk_cldma_start_xfer(struct cldma_drv_info *drv_info, u32 qno)
 {
 	struct txq *txq;
@@ -1095,6 +1338,11 @@ static int mtk_cldma_start_xfer(struct cldma_drv_info *drv_info, u32 qno)
 	 * names; tx_done_work advances free_idx under the same lock.
 	 */
 	spin_lock(&txq->ring_lock);
+	if (unlikely(txq->is_stopping)) {
+		spin_unlock(&txq->ring_lock);
+		return -EPIPE;
+	}
+
 	if (unlikely(!READ_ONCE(txq->tx_started))) {
 		mtk_cldma_setup_start_addr(drv_info, DIR_TX, qno,
 					   txq->req_pool[txq->free_idx].gpd_dma_addr);
@@ -1124,11 +1372,22 @@ int mtk_cldma_init(struct mtk_ctrl_trans *trans)
 
 void mtk_cldma_exit(struct mtk_ctrl_trans *trans)
 {
-	if (!trans->dev)
+	struct cldma_dev *cd = trans->dev;
+	int i;
+
+	if (!cd)
 		return;
 
-	kfree(trans->dev);
+	/* Latch-and-clear up front: the caller holds trans->submit_lock, so
+	 * publishing the NULL here makes any later submit path bail out in
+	 * mtk_cldma_get_tx_budget() instead of walking freed queues.
+	 */
 	trans->dev = NULL;
+
+	for (i = 0; i < NR_CLDMA; i++)
+		mtk_cldma_dev_exit(cd, i);
+
+	kfree(cd);
 }
 
 static int mtk_cldma_open(struct cldma_dev *cd, struct sk_buff *skb)
@@ -1147,7 +1406,8 @@ static int mtk_cldma_open(struct cldma_dev *cd, struct sk_buff *skb)
 		trb->trb_complete(skb);
 		return -EINVAL;
 	}
-	drv_info = cd->cldma_drv_info[que->hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[que->hif_id]);
 	if (!drv_info) {
 		ret = -EIO;
 		goto out;
@@ -1250,6 +1510,82 @@ static void mtk_cldma_txq_flush(struct cldma_drv_info *drv_info,
 	}
 }
 
+/* Handle QUEUE_ERROR interrupts out of atomic context: stop the errored
+ * queues so the device stops walking their rings, and complete pending TX
+ * requests with an error so the failure is reported upward instead of
+ * being silently logged.
+ */
+static void mtk_cldma_err_work(struct work_struct *work)
+{
+	struct cldma_drv_info *drv_info = container_of(work, struct cldma_drv_info, err_work);
+	u32 tx_err, rx_err;
+	struct txq *txq;
+	struct rxq *rxq;
+	int i, ret;
+
+	tx_err = atomic_xchg(&drv_info->tx_err_qs, 0);
+	rx_err = atomic_xchg(&drv_info->rx_err_qs, 0);
+
+	for (i = 0; i < HW_QUEUE_NUM; i++) {
+		if (tx_err & BIT(i)) {
+			/* pairs with smp_store_release() in txq_alloc */
+			txq = smp_load_acquire(&drv_info->txq[i]);
+			if (!txq)
+				continue;
+			/* Keep the producer out for as long as this queue is
+			 * being stopped, so start_xfer cannot ring the
+			 * doorbell between the stop and the flush.
+			 */
+			spin_lock(&txq->ring_lock);
+			txq->is_stopping = true;
+			spin_unlock(&txq->ring_lock);
+
+			ret = mtk_cldma_stop_queue(drv_info, DIR_TX, i);
+			if (ret) {
+				/* the device may still be walking the ring:
+				 * unmapping its buffers here would leave it
+				 * writing into unmapped memory. is_stopping is
+				 * left set so nothing submits to it again.
+				 */
+				dev_err(drv_info->mdev->dev,
+					"TX queue %d stop failed (%d), keeping its requests\n",
+					i, ret);
+				continue;
+			}
+
+			spin_lock(&txq->ring_lock);
+			WRITE_ONCE(txq->tx_started, false);
+			spin_unlock(&txq->ring_lock);
+			mtk_cldma_txq_flush(drv_info, txq, -EPIPE);
+
+			spin_lock(&txq->ring_lock);
+			txq->is_stopping = false;
+			spin_unlock(&txq->ring_lock);
+		}
+		if (rx_err & BIT(i)) {
+			/* pairs with smp_store_release() in rxq_alloc */
+			rxq = smp_load_acquire(&drv_info->rxq[i]);
+			if (!rxq)
+				continue;
+			ret = mtk_cldma_stop_queue(drv_info, DIR_RX, i);
+			if (ret) {
+				dev_err(drv_info->mdev->dev,
+					"RX queue %d stop failed (%d), leaving it stopped\n",
+					i, ret);
+				continue;
+			}
+
+			/* Hand the queue back to the worker that owns free_idx
+			 * instead of programming it here: a stopped queue
+			 * raises no further interrupt, so nothing else would
+			 * ever re-arm it.
+			 */
+			atomic_set(&rxq->need_restart, 1);
+			queue_work(drv_info->wq, &rxq->rx_done_work);
+		}
+	}
+}
+
 static int mtk_cldma_tx(struct cldma_dev *cd, struct sk_buff *skb)
 {
 	struct trb *trb = (struct trb *)skb->cb;
@@ -1262,7 +1598,8 @@ static int mtk_cldma_tx(struct cldma_dev *cd, struct sk_buff *skb)
 	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
 	if (unlikely(!que))
 		return -EPIPE;
-	drv_info = cd->cldma_drv_info[que->hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[que->hif_id]);
 	if (unlikely(!drv_info))
 		return -EPIPE;
 	txq = drv_info->txq[que->txqno];
@@ -1292,7 +1629,8 @@ static int mtk_cldma_close(struct cldma_dev *cd, struct sk_buff *skb)
 		trb->trb_complete(skb);
 		return -EPIPE;
 	}
-	drv_info = cd->cldma_drv_info[que->hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[que->hif_id]);
 	if (unlikely(!drv_info)) {
 		trb->status = -EPIPE;
 		trb->trb_complete(skb);
@@ -1402,7 +1740,8 @@ int mtk_cldma_submit_tx(void *dev, struct sk_buff *skb)
 	que = radix_tree_lookup(&cd->trans->queue_tbl, trb->channel_id & 0xFFFF);
 	if (unlikely(!que))
 		return -EINVAL;
-	drv_info = cd->cldma_drv_info[que->hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[que->hif_id]);
 	if (unlikely(!drv_info))
 		return -EINVAL;
 
@@ -1451,7 +1790,8 @@ int mtk_cldma_get_tx_budget(void *dev, enum mtk_hif_id hif_id, u32 qno)
 	if (unlikely(hif_id >= NR_CLDMA || qno >= HW_QUE_NUM || !cd))
 		return -EINVAL;
 
-	drv_info = cd->cldma_drv_info[hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[hif_id]);
 	if (!drv_info)
 		return -EINVAL;
 	txq = drv_info->txq[qno];
@@ -1483,6 +1823,38 @@ int mtk_cldma_trb_process(void *dev, struct sk_buff *skb)
 	return trb_act_tbl[trb->cmd](cd, skb);
 }
 
+void mtk_cldma_fsm_state_listener(struct mtk_fsm_param *param, struct mtk_ctrl_trans *trans)
+{
+	struct cldma_dev *cd = trans->dev;
+	enum mtk_fsm_hif_user user;
+	enum mtk_hif_id hif_id;
+	int ret;
+
+	switch (param->to) {
+	case FSM_STATE_BOOTUP:
+		if (param->fsm_flag & FSM_F_SAP_HS_START) {
+			hif_id = CLDMA0;
+			user = MTK_FSM_HIF_USER_CLDMA0;
+		} else if (param->fsm_flag & FSM_F_MD_HS_START) {
+			hif_id = CLDMA1;
+			user = MTK_FSM_HIF_USER_CLDMA1;
+		} else {
+			break;
+		}
+
+		ret = mtk_cldma_dev_init(cd, hif_id);
+		if (ret)
+			dev_err(trans->mdev->dev, "Failed to init CLDMA%d: %d\n", hif_id, ret);
+		/* Record the success too: HS1 is retried, so a later good init
+		 * has to clear what the failed attempt left behind.
+		 */
+		mtk_fsm_hif_err_record(trans->mdev, user, ret);
+		break;
+	default:
+		break;
+	}
+}
+
 int mtk_cldma_check_ch_cfg(void *dev, struct queue_info *que)
 {
 	struct cldma_drv_info *drv_info;
@@ -1492,7 +1864,8 @@ int mtk_cldma_check_ch_cfg(void *dev, struct queue_info *que)
 	struct rxq *rxq;
 
 	mdev = cd->trans->mdev;
-	drv_info = cd->cldma_drv_info[que->hif_id];
+	/* pairs with smp_store_release() in mtk_cldma_dev_init */
+	drv_info = smp_load_acquire(&cd->cldma_drv_info[que->hif_id]);
 
 	if (!drv_info) {
 		dev_err(mdev->dev, "CLDMA%d has not been initialized\n",
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.h b/drivers/net/wwan/t9xx/pcie/mtk_cldma.h
index fa6d79f8b5df..bc898f2f52ce 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_cldma.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.h
@@ -146,6 +146,10 @@ struct txq {
 	u32 wr_idx;
 	u32 free_idx;
 	bool tx_started;
+	/* Set by err_work while it stops and flushes this queue; read under
+	 * ring_lock by the producer so no request is submitted in between.
+	 */
+	bool is_stopping;
 	u32 nr_bds;
 };
 
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h
index d7abf4b48895..1917bd70b51d 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.h
@@ -108,6 +108,13 @@ struct cldma_drv_info {
 	struct dma_pool *bd_dma_pool;
 	struct workqueue_struct *wq;
 	struct cldma_hw_regs *hw_regs;
+	struct work_struct err_work;
+	atomic_t tx_err_qs;
+	atomic_t rx_err_qs;
+	/* a ring could not be quiesced and was leaked on free; skip
+	 * destroying the DMA pools its descriptors still live in
+	 */
+	bool ring_leaked;
 };
 
 extern struct cldma_hw_regs mtk_cldma_regs_m9xx;
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
index 84b7fb3c82fd..000a481398a4 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_pci.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
@@ -878,23 +878,47 @@ static int mtk_pci_dev_init(struct mtk_md_dev *mdev)
 {
 	int ret;
 
-	ret = mtk_trans_ctrl_init(mdev);
+	ret = mtk_fsm_init(mdev);
 	if (ret) {
-		dev_err(mdev->dev, "Failed to initialize control plane: %d\n", ret);
+		dev_err(mdev->dev, "Failed to initialize FSM: %d\n", ret);
 		return ret;
 	}
 
+	ret = mtk_trans_ctrl_init(mdev);
+	if (ret)
+		goto free_fsm;
+
 	return 0;
+free_fsm:
+	mtk_fsm_exit(mdev);
+	return ret;
 }
 
 static void mtk_pci_dev_exit(struct mtk_md_dev *mdev)
 {
+	int ret;
+
+	ret = mtk_fsm_evt_submit(mdev, FSM_EVT_DEV_RM, 0, NULL, 0,
+				 EVT_MODE_BLOCKING | EVT_MODE_TOHEAD);
+	if (ret < 0 || ret == FSM_EVT_RET_FAIL)
+		dev_err(mdev->dev, "FSM DEV_RM failed: %d, forcing cleanup\n", ret);
+	/* Close the event gate and park the FSM thread before the transport
+	 * plane it drives is torn down.
+	 */
+	mtk_fsm_stop(mdev);
 	mtk_trans_ctrl_exit(mdev);
+	mtk_fsm_exit(mdev);
 }
 
 static int mtk_pci_dev_start(struct mtk_md_dev *mdev)
 {
-	return 0;
+	int ret;
+
+	ret = mtk_fsm_evt_submit(mdev, FSM_EVT_DEV_ADD, 0, NULL, 0, 0);
+	if (ret == FSM_EVT_RET_FAIL)
+		return -ENOMEM;
+
+	return mtk_fsm_start(mdev);
 }
 
 static int mtk_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
index d485bd7e3bf9..658189dace01 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
@@ -473,19 +473,23 @@ int mtk_pcie_hif_exit(struct mtk_md_dev *mdev)
 
 	trans = ctrl_blk->ctrl_hw_priv;
 
+	/* Hold submit_lock across the whole teardown: late submitters either
+	 * see available and finish before the teardown starts, or block here
+	 * and then bail out on !available. Also makes this exit idempotent,
+	 * so both the FSM listener and device removal may call it.
+	 */
 	mutex_lock(&trans->submit_lock);
+	if (!atomic_read(&trans->available)) {
+		mutex_unlock(&trans->submit_lock);
+		return 0;
+	}
 	atomic_set(&trans->available, 0);
-	mutex_unlock(&trans->submit_lock);
-
 	/* Join the service threads before freeing what they dereference: they
 	 * read trans->dev without a lock, so clearing it cannot stop a
 	 * consumer that has already loaded the pointer.
 	 */
 	mtk_ctrl_trb_srv_exit(trans);
 	mtk_cldma_exit(trans);
-
-	/* Late submitters may still hold the lock and walk the tree. */
-	mutex_lock(&trans->submit_lock);
 	mtk_ctrl_remove_radix_tree(trans);
 	mutex_unlock(&trans->submit_lock);
 
@@ -565,6 +569,15 @@ int mtk_pcie_hif_submit_skb(struct mtk_md_dev *mdev, struct sk_buff *skb, bool f
 	return ret;
 }
 
+void mtk_pcie_hif_fsm_indication(struct mtk_md_dev *mdev, struct mtk_fsm_param *param)
+{
+	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
+	struct mtk_ctrl_trans *trans;
+
+	trans = ctrl_blk->ctrl_hw_priv;
+	mtk_cldma_fsm_state_listener(param, trans);
+}
+
 int mtk_pcie_hif_cmd_func(struct mtk_md_dev *mdev, int cmd, void *data)
 {
 	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
@@ -627,6 +640,13 @@ int mtk_trans_ctrl_init(struct mtk_md_dev *mdev)
 
 int mtk_trans_ctrl_exit(struct mtk_md_dev *mdev)
 {
+	/* FSM_STATE_OFF normally tears the HIF down. If that never ran, the
+	 * trb kthreads and the CLDMA irq callback would outlive the devm
+	 * allocations they point at, so do it here. mtk_pcie_hif_exit() is
+	 * idempotent under trans->submit_lock, so calling it again is safe.
+	 */
+	mtk_pcie_hif_exit(mdev);
+
 	mtk_ctrl_exit(mdev);
 
 	return 0;
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
index 1228f40f2d39..b281f49ea4f0 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.h
@@ -80,9 +80,12 @@ struct trb_srv {
 	struct task_struct *trb_thread;
 };
 
+struct mtk_fsm_param;
+
 int mtk_pcie_hif_init(struct mtk_md_dev *mdev);
 int mtk_pcie_hif_exit(struct mtk_md_dev *mdev);
 int mtk_pcie_hif_submit_skb(struct mtk_md_dev *mdev, struct sk_buff *skb, bool force_send);
+void mtk_pcie_hif_fsm_indication(struct mtk_md_dev *mdev, struct mtk_fsm_param *param);
 int mtk_pcie_hif_cmd_func(struct mtk_md_dev *mdev, int cmd, void *data);
 int mtk_trans_ctrl_init(struct mtk_md_dev *mdev);
 int mtk_trans_ctrl_exit(struct mtk_md_dev *mdev);

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports
  2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
                   ` (4 preceding siblings ...)
  2026-09-30  7:46 ` [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread Jack Wu via B4 Relay
@ 2026-09-30  7:46 ` Jack Wu via B4 Relay
  2026-10-04  9:12   ` netdev-bot+sashiko
  5 siblings, 1 reply; 13+ messages in thread
From: Jack Wu via B4 Relay @ 2026-09-30  7:46 UTC (permalink / raw)
  To: Loic Poulain, Sergey Ryazanov, Johannes Berg, Andrew Lunn,
	David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Jack Wu, Wen-Zhi Huang, Shi-Wei Yeh, Minano Tseng,
	Matthias Brugger, AngeloGioacchino Del Regno, Simon Horman,
	Jonathan Corbet, Shuah Khan, Robert Yu, Jeff Chang
  Cc: linux-kernel, netdev, linux-arm-kernel, linux-mediatek, linux-doc

From: Jack Wu <jackbb_wu@compal.com>

Add AT & MBIM ports to the port infrastructure.

mtk_port_wwan_init() only prepares the port object; wwan_create_port()
is called from mtk_port_wwan_enable(). enable() runs on FSM_STATE_READY
through the new mtk_port_enable_by_type(PORT_TBL_MD) hook, which enables
every MD-table port the device marked enabled, internal ports included.
CCCI_UART2 and CCCI_MBIM get CLDMA1 TXQ(5)/RXQ(5) and TXQ(2)/RXQ(2) in
mtk_queue_info[].

The implemented WWAN port operations are start, stop, tx and
tx_blocking. The tx path copies with skb_copy_bits(), handling the
non-linear skb the WWAN core supplies for writes larger than one
fragment.

TX back-pressure is reported through wwan_port_txon()/txoff() rather
than a .tx_poll operation: a queue-full submit failure pauses the port
and a per-port tx_complete hook resumes it from TRB completion, so
poll() sleeps on the WWAN core's own waitqueue - whose lifetime is
pinned by the open file - instead of a driver waitqueue that
wwan_remove_port() can outlive.

Signed-off-by: Jack Wu <jackbb_wu@compal.com>
---
 drivers/net/wwan/t9xx/mtk_ctrl_plane.h      |   4 +
 drivers/net/wwan/t9xx/mtk_port.c            |  34 +++
 drivers/net/wwan/t9xx/mtk_port.h            |   9 +
 drivers/net/wwan/t9xx/mtk_port_io.c         | 355 ++++++++++++++++++++++++++++
 drivers/net/wwan/t9xx/mtk_port_io.h         |   1 +
 drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c |   8 +
 6 files changed, 411 insertions(+)

diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
index 1f13e25502ad..2b031160eaa9 100644
--- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
+++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
@@ -22,6 +22,10 @@ enum mtk_ccci_ch {
 	/* to MD */
 	CCCI_CONTROL_RX = 0x2000,
 	CCCI_CONTROL_TX = 0x2001,
+	CCCI_UART2_RX = 0x200A,
+	CCCI_UART2_TX = 0x200C,
+	CCCI_MBIM_RX = 0x20D0,
+	CCCI_MBIM_TX = 0x20D1,
 };
 
 enum mtk_trb_cmd_type {
diff --git a/drivers/net/wwan/t9xx/mtk_port.c b/drivers/net/wwan/t9xx/mtk_port.c
index d7c50519e1fd..90ac551887a8 100644
--- a/drivers/net/wwan/t9xx/mtk_port.c
+++ b/drivers/net/wwan/t9xx/mtk_port.c
@@ -285,6 +285,10 @@ static int mtk_port_tx_complete(struct sk_buff *skb)
 			 "Failed to send data: status:%d, port:%s\n",
 			 trb->status, port->info.name);
 
+	/* Runs in the trb_srv kthread, so the hook may sleep on a mutex. */
+	if (ports_ops[port->info.type]->tx_complete)
+		ports_ops[port->info.type]->tx_complete(port);
+
 	wake_up_all(&port->trb_wq);
 	kref_put(&trb->kref, mtk_port_trb_free);
 
@@ -656,6 +660,29 @@ int mtk_port_ch_disable(struct mtk_port *port)
 	return ret;
 }
 
+static int mtk_port_enable_by_type(struct mtk_port_mngr *port_mngr, int tbl_type)
+{
+	struct mtk_port **ports;
+	int ret, idx;
+
+	if (tbl_type < 0 || tbl_type >= PORT_TBL_MAX)
+		return -EINVAL;
+
+	ports = kcalloc(port_mngr->port_cnt, sizeof(struct mtk_port *), GFP_KERNEL);
+	if (!ports)
+		return -ENOMEM;
+
+	ret = radix_tree_gang_lookup(&port_mngr->port_tbl[tbl_type],
+				     (void **)ports, 0, port_mngr->port_cnt);
+	for (idx = 0; idx < ret; idx++) {
+		if (ports[idx]->enable)
+			ports_ops[ports[idx]->info.type]->enable(ports[idx]);
+	}
+
+	kfree(ports);
+	return 0;
+}
+
 static void mtk_port_disable(struct mtk_port_mngr *port_mngr)
 {
 	struct radix_tree_iter iter;
@@ -677,6 +704,7 @@ static void mtk_port_disable(struct mtk_port_mngr *port_mngr)
 void mtk_port_mngr_fsm_state_handler(struct mtk_fsm_param *fsm_param, void *arg)
 {
 	struct mtk_port_mngr *port_mngr;
+	int ret;
 
 	if (!fsm_param || !arg)
 		return;
@@ -687,6 +715,12 @@ void mtk_port_mngr_fsm_state_handler(struct mtk_fsm_param *fsm_param, void *arg)
 	case FSM_STATE_OFF:
 		mtk_port_disable(port_mngr);
 		break;
+	case FSM_STATE_READY:
+		ret = mtk_port_enable_by_type(port_mngr, PORT_TBL_MD);
+		if (ret)
+			dev_err(port_mngr->ctrl_blk->mdev->dev,
+				"Failed to enable MD ports: %d\n", ret);
+		break;
 	default:
 		break;
 	}
diff --git a/drivers/net/wwan/t9xx/mtk_port.h b/drivers/net/wwan/t9xx/mtk_port.h
index 2fc0e450ab6b..acfd52792dc7 100644
--- a/drivers/net/wwan/t9xx/mtk_port.h
+++ b/drivers/net/wwan/t9xx/mtk_port.h
@@ -55,6 +55,7 @@ enum mtk_port_tbl {
 
 enum mtk_port_type {
 	PORT_TYPE_INTERNAL,
+	PORT_TYPE_WWAN,
 	PORT_TYPE_MAX
 };
 
@@ -63,6 +64,13 @@ struct mtk_internal_port {
 	int (*recv_cb)(void *arg, struct sk_buff *skb);
 };
 
+struct mtk_wwan_port {
+	/* w_lock protects wwan_port when recv data and disable port at the same time */
+	struct mutex w_lock;
+	int w_type;
+	void *w_port;
+};
+
 struct mtk_port_cfg {
 	enum mtk_ccci_ch tx_ch;
 	enum mtk_ccci_ch rx_ch;
@@ -90,6 +98,7 @@ struct mtk_port {
 	wait_queue_head_t rx_wq;
 	struct mtk_port_mngr *port_mngr;
 	struct mtk_internal_port i_priv;
+	struct mtk_wwan_port w_priv;
 };
 
 struct mtk_port_mngr {
diff --git a/drivers/net/wwan/t9xx/mtk_port_io.c b/drivers/net/wwan/t9xx/mtk_port_io.c
index 6cff0704b7fc..0d8f63e2dbda 100644
--- a/drivers/net/wwan/t9xx/mtk_port_io.c
+++ b/drivers/net/wwan/t9xx/mtk_port_io.c
@@ -3,8 +3,13 @@
  * Copyright (c) 2022, MediaTek Inc.
  */
 #include <linux/netdevice.h>
+#include <linux/poll.h>
+#include <linux/slab.h>
+#include <linux/wait.h>
+#include <linux/wwan.h>
 
 #include "mtk_port_io.h"
+#include "mtk_trans_ctrl.h"
 
 static int mtk_port_get_locked(struct mtk_port *port)
 {
@@ -41,6 +46,73 @@ static void mtk_port_struct_init(struct mtk_port *port)
 	init_waitqueue_head(&port->rx_wq);
 }
 
+/* Splits the source skb into CCCI packets and submits them.  The source may
+ * be non-linear: the WWAN core hands the tx ops a head skb whose linear area
+ * is one fragment (caps.frag_len bytes) with the rest chained on frag_list,
+ * so every read goes through skb_copy_bits(), which walks that chain and is
+ * bounded by src->len by construction.
+ *
+ * Every packet is built before any is submitted, because the core reports a
+ * partial write as a whole-write failure and userspace then re-sends the
+ * message, duplicating the prefix already on the wire.  So no packet is
+ * submitted unless all of them were built, and only the first submit may
+ * fail with -EAGAIN; the rest carry force_send, leaving only failures that
+ * mean the channel itself is going away.  -EINTR from a blocking wait
+ * arrives after the skb was submitted, so the packet counts as accepted.
+ */
+static int mtk_port_common_write(struct mtk_port *port, struct sk_buff *src, bool blocking)
+{
+	u32 packet_size, left_cnt = src->len, cur_pos;
+	struct sk_buff_head list;
+	bool force_send = false;
+	struct sk_buff *skb;
+	int ret;
+
+	ret = mtk_port_status_check(port);
+	if (ret)
+		return ret;
+
+	__skb_queue_head_init(&list);
+
+	while (left_cnt) {
+		skb = __dev_alloc_skb(port->tx_mtu, GFP_KERNEL);
+		if (!skb) {
+			ret = -ENOMEM;
+			goto err_purge;
+		}
+
+		skb_reserve(skb, sizeof(struct mtk_ccci_header));
+
+		packet_size = min_t(u32, left_cnt,
+				    port->tx_mtu - sizeof(struct mtk_ccci_header));
+		cur_pos = src->len - left_cnt;
+		ret = skb_copy_bits(src, cur_pos, skb_put(skb, packet_size), packet_size);
+		if (ret) {
+			dev_err(port->port_mngr->ctrl_blk->mdev->dev,
+				"Failed to copy data for port(%s)\n", port->info.name);
+			dev_kfree_skb_any(skb);
+			goto err_purge;
+		}
+
+		__skb_queue_tail(&list, skb);
+		left_cnt -= packet_size;
+	}
+
+	while ((skb = __skb_dequeue(&list))) {
+		ret = mtk_port_send_data(port, skb, blocking, force_send);
+		if (ret < 0 && ret != -EINTR)
+			goto err_purge;
+
+		force_send = true;
+	}
+
+	return 0;
+
+err_purge:
+	__skb_queue_purge(&list);
+	return ret;
+}
+
 static int mtk_port_internal_init(struct mtk_port *port)
 {
 	mtk_port_struct_init(port);
@@ -253,6 +325,289 @@ static const struct port_ops port_internal_ops = {
 	.recv = mtk_port_internal_recv,
 };
 
+static int mtk_port_wwan_open(struct wwan_port *w_port)
+{
+	struct mtk_port *port;
+	int ret;
+
+	port = wwan_port_get_drvdata(w_port);
+	ret = mtk_port_get_locked(port);
+	if (ret)
+		return ret;
+
+	ret = mtk_port_common_open(port);
+	if (ret) {
+		mtk_port_put_locked(port);
+		return ret;
+	}
+
+	/* WWAN_PORT_TX_OFF persists on the wwan_port across close/open
+	 * cycles; start every session writable.
+	 */
+	wwan_port_txon(w_port);
+
+	return 0;
+}
+
+static void mtk_port_wwan_close(struct wwan_port *w_port)
+{
+	struct mtk_port *port = wwan_port_get_drvdata(w_port);
+
+	/* Clear PORT_S_OPEN under the same lock mtk_port_wwan_recv() holds, so
+	 * a receive that saw the port open has finished its wwan_port_rx()
+	 * before stop() returns and the core purges its rx queue.
+	 */
+	mutex_lock(&port->w_priv.w_lock);
+	mtk_port_common_close(port);
+	mutex_unlock(&port->w_priv.w_lock);
+
+	mtk_port_put_locked(port);
+}
+
+/* Pause TX after a queue-full submit failure.  The queue may have drained
+ * between that failure and the txoff below - the tx_complete that emptied
+ * it ran before this txoff and nothing later would ever resume TX - so
+ * recheck under the same lock and undo the txoff if the queue is no longer
+ * full.  An error from the recheck also resumes TX: reporting it as "not
+ * writable" would leave the poller asleep with nothing left to wake it,
+ * while the next write returns the real errno.
+ */
+static void mtk_port_wwan_tx_pause(struct mtk_port *port)
+{
+	union ctrl_hif_cmd_data hif_cmd;
+	struct mtk_ctrl_blk *ctrl_blk;
+	int ret;
+
+	ctrl_blk = port->port_mngr->ctrl_blk;
+
+	mutex_lock(&port->w_priv.w_lock);
+	if (!port->w_priv.w_port)
+		goto unlock;
+
+	wwan_port_txoff(port->w_priv.w_port);
+
+	hif_cmd.rx_ch = port->info.rx_ch;
+	ret = mtk_pcie_hif_cmd_func(ctrl_blk->mdev, HIF_CTRL_CMD_CHECK_TX_FULL,
+				    &hif_cmd);
+	if (ret <= 0)
+		wwan_port_txon(port->w_priv.w_port);
+unlock:
+	mutex_unlock(&port->w_priv.w_lock);
+}
+
+/* Called from the trb_srv kthread when a TX trb for this port completes.
+ * w_lock serializes it against mtk_port_wwan_tx_pause(), closing the
+ * txoff-after-drain window described there.
+ */
+static void mtk_port_wwan_tx_complete(struct mtk_port *port)
+{
+	mutex_lock(&port->w_priv.w_lock);
+	if (port->w_priv.w_port)
+		wwan_port_txon(port->w_priv.w_port);
+	mutex_unlock(&port->w_priv.w_lock);
+}
+
+static int mtk_port_wwan_tx(struct wwan_port *w_port, struct sk_buff *skb, bool blocking)
+{
+	struct mtk_port *port = wwan_port_get_drvdata(w_port);
+	int ret;
+
+	if (unlikely(!skb->len)) {
+		consume_skb(skb);
+		return 0;
+	}
+
+	ret = mtk_port_common_write(port, skb, blocking);
+	if (ret < 0) {
+		if (ret == -EAGAIN)
+			mtk_port_wwan_tx_pause(port);
+		return ret;
+	}
+
+	consume_skb(skb);
+	return 0;
+}
+
+static int mtk_port_wwan_write(struct wwan_port *w_port, struct sk_buff *skb)
+{
+	return mtk_port_wwan_tx(w_port, skb, false);
+}
+
+static int mtk_port_wwan_write_blocking(struct wwan_port *w_port, struct sk_buff *skb)
+{
+	return mtk_port_wwan_tx(w_port, skb, true);
+}
+
+/* No .tx_poll: it would poll_wait() on port->trb_wq, which lives in a
+ * mtk_port that wwan_remove_port() lets go of while the file stays open,
+ * and wake_up_pollfree() is not available to modules.  TX back-pressure
+ * reaches poll() through wwan_port_txon()/txoff() on the core's own
+ * waitqueue instead, whose lifetime is pinned by the open file.
+ */
+static const struct wwan_port_ops wwan_ops = {
+	.start = mtk_port_wwan_open,
+	.stop = mtk_port_wwan_close,
+	.tx = mtk_port_wwan_write,
+	.tx_blocking = mtk_port_wwan_write_blocking,
+};
+
+static int mtk_port_wwan_init(struct mtk_port *port)
+{
+	mtk_port_struct_init(port);
+	port->enable = false;
+
+	mutex_init(&port->w_priv.w_lock);
+
+	switch (port->info.rx_ch) {
+	case CCCI_MBIM_RX:
+		port->w_priv.w_type = WWAN_PORT_MBIM;
+		break;
+	case CCCI_UART2_RX:
+		port->w_priv.w_type = WWAN_PORT_AT;
+		break;
+	default:
+		port->w_priv.w_type = WWAN_PORT_UNKNOWN;
+		break;
+	}
+
+	return 0;
+}
+
+static void mtk_port_wwan_exit(struct mtk_port *port)
+{
+	if (test_bit(PORT_S_ENABLE, &port->status))
+		ports_ops[port->info.type]->disable(port);
+}
+
+static void mtk_port_wwan_enable(struct mtk_port *port)
+{
+	struct mtk_port_mngr *port_mngr;
+	struct wwan_port_caps caps;
+	struct wwan_port *wp;
+	int ret;
+
+	port_mngr = port->port_mngr;
+
+	if (test_bit(PORT_S_ENABLE, &port->status))
+		return;
+
+	ret = mtk_port_ch_enable(port);
+	if (ret && ret != -EBUSY) {
+		/* On -ETIMEDOUT the enable's outcome is not yet known: the
+		 * ENABLE trb may still be queued.  The DISABLE is queued
+		 * behind it, so the channel cannot stay armed unowned.
+		 */
+		mtk_port_ch_disable(port);
+		return;
+	}
+
+	/* tx_mtu is only valid once the channel open trb has completed. A zero
+	 * frag_len would make wwan_port_fops_write() loop forever.
+	 */
+	if (!port->tx_mtu) {
+		dev_err(port_mngr->ctrl_blk->mdev->dev,
+			"Invalid tx_mtu for port(%s)\n", port->info.name);
+		mtk_port_ch_disable(port);
+		return;
+	}
+
+	/* The core allocates frag_len + headroom_len and skb_put()s frag_len,
+	 * so frag_len is the payload budget: subtract the CCCI header to make
+	 * one core fragment exactly one CCCI packet.
+	 */
+	caps.frag_len = port->tx_mtu - sizeof(struct mtk_ccci_header);
+	caps.headroom_len = sizeof(struct mtk_ccci_header);
+
+	/* These bits must be set before wwan_create_port(): the device node
+	 * becomes openable inside it and mtk_port_common_open() rejects a
+	 * port without PORT_S_ENABLE.  w_port cannot be published first - it is
+	 * this call's return value - so an RX frame arriving in between is
+	 * dropped with -ENXIO by design.
+	 */
+	set_bit(PORT_S_WR, &port->status);
+	set_bit(PORT_S_ENABLE, &port->status);
+
+	wp = wwan_create_port(port_mngr->ctrl_blk->mdev->dev,
+			      port->w_priv.w_type,
+			      &wwan_ops, &caps, port);
+	if (IS_ERR(wp)) {
+		dev_warn(port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to create wwan port for (%s)\n", port->info.name);
+		clear_bit(PORT_S_ENABLE, &port->status);
+		clear_bit(PORT_S_WR, &port->status);
+		mtk_port_ch_disable(port);
+		return;
+	}
+
+	mutex_lock(&port->w_priv.w_lock);
+	port->w_priv.w_port = wp;
+	mutex_unlock(&port->w_priv.w_lock);
+}
+
+static void mtk_port_wwan_disable(struct mtk_port *port)
+{
+	struct wwan_port *w_port;
+	int ret;
+
+	if (!test_and_clear_bit(PORT_S_ENABLE, &port->status))
+		return;
+
+	/* PORT_S_WR is part of the blocking TX wait condition, and the waiter
+	 * loops back on timeout, so without this wake it only notices at the
+	 * next trb timeout rather than now.
+	 */
+	clear_bit(PORT_S_WR, &port->status);
+	wake_up_all(&port->trb_wq);
+
+	/* w_lock must be dropped before wwan_remove_port(): that takes the
+	 * core's ops_lock and forces stop(), which is close() taking w_lock.
+	 */
+	mutex_lock(&port->w_priv.w_lock);
+	w_port = port->w_priv.w_port;
+	port->w_priv.w_port = NULL;
+	mutex_unlock(&port->w_priv.w_lock);
+
+	ret = mtk_port_ch_disable(port);
+	if (ret)
+		dev_warn(port->port_mngr->ctrl_blk->mdev->dev,
+			 "Failed to disable channel for port(%s): %d\n",
+			 port->info.name, ret);
+
+	wwan_remove_port(w_port);
+}
+
+static int mtk_port_wwan_recv(struct mtk_port *port, struct sk_buff *skb)
+{
+	/* Drop frames when nobody has the device open: wwan_port_rx() queues
+	 * without bound and only a reader drains the queue, so accepting
+	 * unsolicited traffic here would grow the rxq indefinitely.  Both
+	 * conditions are read under w_lock, which mtk_port_wwan_close() also
+	 * takes, so a frame accepted here cannot land on a queue the core is
+	 * about to purge.
+	 */
+	mutex_lock(&port->w_priv.w_lock);
+	if (!test_bit(PORT_S_OPEN, &port->status) || !port->w_priv.w_port) {
+		mutex_unlock(&port->w_priv.w_lock);
+		dev_dbg_ratelimited(port->port_mngr->ctrl_blk->mdev->dev,
+				    "Drop RX for unopened port(%s)\n", port->info.name);
+		return -ENXIO;
+	}
+
+	wwan_port_rx(port->w_priv.w_port, skb);
+	mutex_unlock(&port->w_priv.w_lock);
+	return 0;
+}
+
+static const struct port_ops port_wwan_ops = {
+	.init = mtk_port_wwan_init,
+	.exit = mtk_port_wwan_exit,
+	.enable = mtk_port_wwan_enable,
+	.disable = mtk_port_wwan_disable,
+	.recv = mtk_port_wwan_recv,
+	.tx_complete = mtk_port_wwan_tx_complete,
+};
+
 const struct port_ops *ports_ops[PORT_TYPE_MAX] = {
 	&port_internal_ops,
+	&port_wwan_ops,
 };
diff --git a/drivers/net/wwan/t9xx/mtk_port_io.h b/drivers/net/wwan/t9xx/mtk_port_io.h
index 5a4e36c075a5..88aeb28d536b 100644
--- a/drivers/net/wwan/t9xx/mtk_port_io.h
+++ b/drivers/net/wwan/t9xx/mtk_port_io.h
@@ -20,6 +20,7 @@ struct port_ops {
 	void (*enable)(struct mtk_port *port);
 	void (*disable)(struct mtk_port *port);
 	int (*recv)(struct mtk_port *port, struct sk_buff *skb);
+	void (*tx_complete)(struct mtk_port *port);
 };
 
 void *mtk_port_internal_open(struct mtk_md_dev *mdev, char *name, int flag);
diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
index 658189dace01..a1f1f86838f7 100644
--- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
+++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
@@ -29,6 +29,10 @@ static const int mtk_srv_cfg[NR_CLDMA][HW_QUE_NUM] = {
 
 /* the number of RX GPDs should be at least two */
 static const struct queue_info mtk_queue_info[] = {
+	{CCCI_UART2_TX, CCCI_UART2_RX, CLDMA1, TXQ(5), RXQ(5),
+	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
+	{CCCI_MBIM_TX, CCCI_MBIM_RX, CLDMA1, TXQ(2), RXQ(2),
+	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
 	{CCCI_CONTROL_TX, CCCI_CONTROL_RX, CLDMA1, TXQ(0), RXQ(0),
 	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
 	{CCCI_SAP_CONTROL_TX, CCCI_SAP_CONTROL_RX, CLDMA0, TXQ(0), RXQ(0),
@@ -36,6 +40,10 @@ static const struct queue_info mtk_queue_info[] = {
 };
 
 static const struct mtk_port_cfg mtk_port_cfg_tbl[] = {
+	{CCCI_UART2_TX, CCCI_UART2_RX, PORT_TYPE_WWAN, "AT",
+		PORT_F_DFLT},
+	{CCCI_MBIM_TX, CCCI_MBIM_RX, PORT_TYPE_WWAN, "MBIM",
+		PORT_F_DFLT},
 	{CCCI_CONTROL_TX, CCCI_CONTROL_RX, PORT_TYPE_INTERNAL, "MDCTRL",
 		PORT_F_DFLT},
 	{CCCI_SAP_CONTROL_TX, CCCI_SAP_CONTROL_RX, PORT_TYPE_INTERNAL, "SAPCTRL",

-- 
2.34.1



^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core
  2026-09-30  7:46 ` [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found 11 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 5 · Low: 5

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] The teardown order in mtk_pci_remove() lets MSI-X handlers run
  during callback unpublication and after cancel_work_sync().
- [Medium] The commit message says 'There is no software reset that
  returns the device from running firmware to a state the next probe can…
- [Medium] mtk_pci_remove() resets the endpoint after
  pci_disable_device(), using ACPI MRST._RST or an asynchronous MHCCIF…
- [Medium] mtk_pci_atr_init() programs the endpoint's inbound WIN0 ATR
  source address with `cfg.src_addr = pci_resource_start(pdev,…
- [Medium] mtk_pci_bar_init() maps BAR0/1 and BAR2/3 with
  pcim_iomap_region(), which calls pci_iomap(pdev, bar, 0) and so maps…
- [Medium] mtk_pci_drv has no .shutdown callback.
- [Low] mtk_pci.h defines `CEI_PCI_VENDOR_ID 0x03F0`, and t9xx_pci_table
  binds `PCI_DEVICE(CEI_PCI_VENDOR_ID, 0x01CA)` with it.
- [Low] The comment in mtk_pci_bar_init() says the fixed BAR0/1 offsets
  reach only 'MSI-X registers up to 0x3080' and uses that to justify…
- [Low] mtk_pci.h defines its own `MTK_PCI_VENDOR_ID 0x14C3`, which
  duplicates PCI_VENDOR_ID_MEDIATEK (include/linux/pci_ids.h:2142, value…
- [Low] Kconfig lets MTK_T9XX be built with only `depends on PCI &&
  ACPI`, so CONFIG_PCI_MSI=n is allowed.
- [Low] Baseline assessment: false positive (retained because
  requires_validation=true).

--

net: wwan: t9xx: Add PCIe core

This commit adds the PCIe core for the MediaTek T9xx 5G WWAN modem driver.
It covers probe and remove, MSI-X setup, ATR programming and the MHCCIF
event interface. On removal the driver resets the modem through the ACPI
MRST._RST method on the device's node.

> diff --git a/drivers/net/wwan/Kconfig b/drivers/net/wwan/Kconfig
> index 88df55d78d902..55c45af410eee 100644
> --- a/drivers/net/wwan/Kconfig
> +++ b/drivers/net/wwan/Kconfig
> @@ -121,6 +121,17 @@ config MTK_T7XX
>  
>  	  If unsure, say N.
>  
> +config MTK_T9XX
> +	tristate "MediaTek PCIe 5G WWAN modem T9xx device"
> +	depends on PCI && ACPI

[Severity: Low]
Should this also depend on PCI_MSI?

mtk_pci_request_irq() asks for exactly 32 MSI-X vectors and has no INTx
fallback:

	ret = pci_alloc_irq_vectors(pdev, MTK_IRQ_CNT_MAX,
				    MTK_IRQ_CNT_MAX, PCI_IRQ_MSIX);

With CONFIG_PCI_MSI=n, the pci_alloc_irq_vectors_affinity() stub in
include/linux/pci.h returns -ENOSPC for any request other than a single
INTx vector. Probe of both supported devices would then always fail in
that configuration.

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
> new file mode 100644
> index 0000000000000..34ee823119fce
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
> @@ -0,0 +1,1033 @@

[ ... ]

> +int mtk_pci_unregister_irq(struct mtk_md_dev *mdev, int irq_id)
> +{
> +	struct mtk_pci_priv *priv = mdev->hw_priv;
> +	int virq_id;

[ ... ]

> +	/* Stop the source and wait for in-flight handlers
> +	 * before the callback or its data can disappear.
> +	 */
> +	mtk_pci_mask_irq(mdev, irq_id);
> +	virq_id = mtk_pci_get_virq_id(mdev, irq_id);
> +	if (virq_id >= 0)
> +		synchronize_irq(virq_id);

[Severity: Low]
This may be fine. The mask in mtk_pci_mask_irq() is a posted iowrite32(),
and it is not read back before synchronize_irq().

mtk_pci_irq_msix() seems to close that window by itself. It reads
REG_IMASK_HOST_MSIX_GRP0_0 and applies irq_state &= irq_enable before it
dispatches, so a handler that starts after synchronize_irq() should see
the mask.

The one remaining way for the source to be re-enabled is the unmask in
mtk_mhccif_isr_work(). That is covered in the comment on mtk_pci_remove()
below.

[ ... ]

> +static const struct pci_device_id t9xx_pci_table[] = {
> +	{ PCI_DEVICE(MTK_PCI_VENDOR_ID, 0x0900), MTK_PCI_CLASS, PCI_ANY_ID },

[Severity: Low]
This isn't a bug, but MTK_PCI_VENDOR_ID (0x14C3) in mtk_pci.h duplicates
PCI_VENDOR_ID_MEDIATEK from include/linux/pci_ids.h. Could the table use
PCI_VENDOR_ID_MEDIATEK, as t7xx does?

> +	{ PCI_DEVICE(CEI_PCI_VENDOR_ID, 0x01CA), MTK_PCI_CLASS, PCI_ANY_ID },

[Severity: Low]
Is CEI_PCI_VENDOR_ID the right name for 0x03F0? Other in-tree drivers
treat that vendor ID as HP. For example, t7xx_pci.c has:

	{ PCI_DEVICE(0x03f0, 0x09c8) }, // HP DRMR-H01

and mhi pci_generic.c labels 0x03f0 as "T99W175 (sdx55), HP variant".

The commit message also mentions only the T900 device. It does not say
that the driver binds this second vendor/device ID, which becomes visible
through MODULE_DEVICE_TABLE and module autoloading.

> +	{/* end: all zeroes */}
> +};

[ ... ]

> +static int mtk_pci_atr_init(struct mtk_md_dev *mdev)
> +{
> +	struct pci_dev *pdev = to_pci_dev(mdev->dev);
> +	struct mtk_pci_priv *priv = mdev->hw_priv;
> +	struct mtk_atr_cfg cfg;
> +	int port, ret;
> +
> +	mtk_pci_atr_disable(priv);
> +
> +	/* Config ATR for RC to access device's register */
> +	cfg.src_addr = pci_resource_start(pdev, MTK_BAR_2_3_IDX);

[Severity: Medium]
Should this be pci_bus_address(pdev, MTK_BAR_2_3_IDX)?

pci_resource_start() returns the CPU physical address of the BAR.
mtk_pci_setup_atr() writes cfg->src_addr into
REG_ATR_PCIE_WIN0_T0_SRC_ADDR_{MSB,LSB}, and the endpoint compares that
against the PCI bus address of incoming memory TLPs.

Some hosts have a bridge window with a non-zero CPU-to-bus offset (see
pcibios_resource_to_bus()). On those hosts, would the window fail to
match BAR2 accesses? Every access through ext_reg_base (MHCCIF mask, ack
and reset) would then hit an untranslated device address.

The other possible outcome is that the 4 MiB alignment check in
mtk_pci_setup_atr() fails with -EFAULT, even though the bus address is
aligned.

t7xx has the same pattern in t7xx_pcie_mac.c, but this patch adds new
code that repeats it. It is unchanged at the end of the series. This
also depends on how the endpoint matches ATR windows, which the code
alone cannot confirm.

> +	cfg.size = ATR_PCIE_REG_SIZE;

[ ... ]

> +static int mtk_pci_bar_init(struct mtk_md_dev *mdev)
> +{
> +	struct pci_dev *pdev = to_pci_dev(mdev->dev);
> +	struct mtk_pci_priv *priv = mdev->hw_priv;
> +
> +	/* Fixed offsets used on these mappings (MSI-X registers up to
> +	 * 0x3080 on BAR0/1, the ATR-biased MHCCIF window on BAR2/3) are

[Severity: Low]
Is 0x3080 the right upper bound? mtk_pci_reg.h defines
REG_IMASK_HOST_MSIX_GRP0_0 as 0x3100, and mtk_pci_irq_msix() reads it on
every interrupt:

	irq_enable = mtk_pci_mac_read32(priv, REG_IMASK_HOST_MSIX_GRP0_0);

That makes the highest BAR0 offset accessed 0x3103. The comment is
unchanged at the end of the series.

> +	 * guaranteed by the only two devices this driver binds (0x0900,
> +	 * 0x01CA), whose BAR sizes are fixed by hardware; an undersized
> +	 * BAR here is impossible, so no pci_resource_len() checks.
> +	 */
> +	priv->mac_reg_base = pcim_iomap_region(pdev, MTK_BAR_0_1_IDX,
> +					       mdev->dev_str);

[ ... ]

> +	/* We use MD view base address "0" to observe registers */
> +	priv->ext_reg_base = priv->bar23_addr - ATR_PCIE_REG_TRSL_ADDR;

[Severity: Medium]
Can a device with smaller BARs cause out-of-bounds MMIO here?

pcim_iomap_region() maps only the BAR length the device reports, and
nothing checks pci_resource_len().

The vendor and device IDs come from the device's own config space. A
buggy or spoofed device (a Thunderbolt/USB4 enclosure, an FPGA, an
emulated device) could present 0x14C3:0x0900 or 0x03F0:0x01CA with a
4 KiB BAR2 at a 4 MiB-aligned address, and that passes the ATR alignment
check.

The driver then uses fixed offsets. For example, mtk_mhccif_init() does:

	mtk_pci_write32(mdev, MTK_MHCCIF_RC_BASE_ADDR +
			MHCCIF_EP2RC_SW_INT_EAP_MASK_SET, U32_MAX);

This writes U32_MAX to bar23_addr + 0xA030. On BAR0, probe writes
REG_IMASK_HOST_MSIX_CLR_GRP0_0 (0x3080) and the IRQ handler reads
REG_IMASK_HOST_MSIX_GRP0_0 (0x3100).

Could both BAR lengths be checked before they are used?

> +
> +	return 0;
> +}

[ ... ]

> +static void mtk_pci_remove(struct pci_dev *pdev)
> +{
> +	struct mtk_md_dev *mdev = pci_get_drvdata(pdev);
> +	struct mtk_pci_priv *priv = mdev->hw_priv;
> +	struct device *dev = &pdev->dev;
> +	int ret;
> +
> +	/* Silence every source before tearing anything down. */
> +	mtk_pci_mac_write32(priv, REG_IMASK_HOST_MSIX_CLR_GRP0_0, U32_MAX);
> +
> +	/* Unregisters the callback (masks and synchronises the vector),
> +	 * then cancels the work.  With the callback gone the work cannot
> +	 * be requeued.
> +	 */
> +	mtk_mhccif_exit(mdev);
> +	mtk_pci_free_irq(mdev);

[Severity: High]
Can the MHCCIF vector fire again while mtk_mhccif_exit() runs, or after
it returns?

mtk_mhccif_isr_work() always ends by unmasking the source:

	mtk_pci_clear_irq(mdev, priv->mhccif_irq_id);
	mtk_pci_unmask_irq(mdev, priv->mhccif_irq_id);

If the work is running when remove starts, this unmask can land after the
mask-all write above and after synchronize_irq() in
mtk_pci_unregister_irq(). The comment in mtk_mhccif_exit() ("the work
unmasks the source on its way out") seems to say the same.

A new vector 28 interrupt then runs
mtk_pci_irq_msix()->mtk_pci_irq_handler() while the handlers are still
requested.

mtk_pci_irq_handler() loads the callback and its data separately, and it
does not take irq_cb_lock:

	cb = READ_ONCE(priv->irq_cb_list[irq_id]);
	if (likely(cb)) {
		smp_rmb(); /* Ensure data is read after callback */
		cb(irq_id, priv->irq_cb_data[irq_id]);

Suppose it reads a non-NULL cb, then mtk_pci_unregister_irq() clears both
entries, and then the handler reads NULL data. In that case
mtk_mhccif_irq_cb(28, NULL) dereferences mdev->hw_priv on a NULL mdev in
hardirq context.

Now suppose the handler gets valid data but is delayed until
cancel_work_sync() has returned. mtk_mhccif_irq_cb() then does:

	queue_work(system_highpri_wq, &priv->mhccif_work);

free_irq() waits only for the handler, not for the work. After remove
returns, devres frees priv and mdev and unmaps both BARs.

The queued work would then run container_of() on freed memory, take the
freed mhccif_lock, and do MMIO through the unmapped ext_reg_base and
mac_reg_base.

The probe error path calls mtk_pci_free_irq() before mtk_mhccif_exit().
Would using that order in remove avoid this?

The ordering is unchanged at the end of the series, where
mtk_pci_dev_exit() is added in front of it.

> +	pci_clear_master(pdev);
> +	pci_disable_device(pdev);
> +
> +	/* Reset last: the endpoint comes back at power-on defaults,
> +	 * so no MMIO, config write or MSI-X teardown may follow it.
> +	 */
> +	ret = mtk_pci_pldr(mdev);
> +	if (ret && mtk_pci_link_check(mdev)) {
> +		dev_warn(dev, "PLDR failed (%d), trying MHCCIF reset\n", ret);
> +		if (mtk_pci_send_ext_evt(mdev, DEV_EVT_H2D_DEVICE_RESET))
> +			dev_err(dev, "MHCCIF reset failed\n");
> +	}

[Severity: Medium]
Does the commit message match this code? It says:

  There is no software reset that returns the device from running
  firmware to a state the next probe can boot, ... so mtk_pci_remove()
  evaluates MRST._RST on the device's own ACPI node.

This code also has a fallback. When mtk_pci_pldr() fails and the link is
still up, remove rings the MHCCIF DEV_EVT_H2D_DEVICE_RESET doorbell,
which is a software reset. mtk_pci_pldr() fails with -ENODEV on any
platform that has no ACPI handle or no MRST._RST method.

If the doorbell cannot restore a bootable state, isn't rebind broken on
every platform without MRST._RST? On those platforms every unbind would
also log "pldr method isn't supported" and the warning.

If the doorbell can restore it, is the commit message's reason for doing
the ACPI reset in remove accurate?

Probe also never checks that MRST._RST exists. For comparison, t7xx uses
its device reset doorbell only together with FASTBOOT_DL_NOTIFY, and it
always goes through MRST._RST for PLDR. This is unchanged at the end of
the series.

> +
> +	pci_load_and_free_saved_state(pdev, &priv->saved_state);
> +}

[Severity: Medium]
Does pci_load_and_free_saved_state() restore anything in hardware here?

According to the comment above, the reset brings the endpoint back at
power-on defaults, so its BARs and command register are cleared.
pci_load_and_free_saved_state() only reloads pdev->saved_config_space and
writes nothing to the device. t7xx calls pci_save_state() before the ACPI
reset and pci_restore_state() after it.

After remove returns, pci_device_remove() sets pci_dev->driver = NULL and
calls pm_runtime_put_sync(). Userspace may have enabled runtime PM
(power/control=auto, which TLP, powertop or udev rules often set on
laptops). In that case pci_pm_runtime_suspend() takes this path:

	if (!pci_dev->driver) {
		pci_save_state(pci_dev);
		return 0;
	}

That overwrites the snapshot with the reset device's config space. The
next mtk_pci_probe() then does:

	if (pdev->state_saved)
		pci_restore_state(pdev);

This writes the power-on-default BAR and command values back, so BAR MMIO
in probe would fail.

A different driver bound afterwards (for example vfio-pci through
driver_override) would see BARs the device does not decode, because only
this driver's probe restores them.

With the MHCCIF fallback, the reset is only a doorbell with no delay and
no readiness wait. Could the device reset in the middle of the next probe
and wipe the ATR and MSI-X mask setup?

This is unchanged at the end of the series.

[ ... ]

> +static struct pci_driver mtk_pci_drv = {
> +	.name = "mtk_pci_drv",
> +	.id_table = t9xx_pci_table,
> +	.probe = mtk_pci_probe,
> +	.remove = mtk_pci_remove,
> +	.err_handler = &mtk_pci_err_handler
> +};

[Severity: Medium]
Is a .shutdown callback needed here?

On reboot or kexec, remove is not called, so the modem stays in running
firmware. The commit message says no software reset can recover from that
state.

On kexec the PCI core only clears bus mastering. The next kernel's
mtk_pci_probe() reprograms the ATR tables with AXIS ports 0-3 as
transparent pass-through. It then calls pci_set_master() without
resetting the endpoint or checking mtk_pci_get_dev_state(). In the final
series this happens after mtk_pci_dev_init() and before any handshake.

Without IOMMU translation, could a DMA engine left armed by the previous
kernel (such as CLDMA with stale descriptor addresses) resume transfers
into host memory the new kernel owns?

The probe error paths also never reset the device. t7xx provides
t7xx_pci_shutdown() for this.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer
  2026-09-30  7:46 ` [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] At this commit, mtk_ctrl_init() and mtk_ctrl_exit() in
  drivers/net/wwan/t9xx/mtk_ctrl_plane.c have no callers.
- [Low] struct mtk_ctrl_blk (mtk_ctrl_plane.h) has a 'struct
  mtk_ctrl_trans *trans' member, and struct mtk_ctrl_trans…
- [Low] mtk_ctrl_init() allocates ctrl_blk with devm_kzalloc(mdev->dev,
  ...).

--

net: wwan: t9xx: Add control plane transaction layer

This commit adds the core control plane structures (struct mtk_ctrl_blk
and struct mtk_ctrl_trans) and the mtk_ctrl_init()/mtk_ctrl_exit() entry
points for the T9XX WWAN driver. The DMA engine and TX/RX service
implementations are left to later patches.

> diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
> new file mode 100644
> index 0000000000000..fa2ab8c3e7577
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
> @@ -0,0 +1,44 @@

[ ... ]

> +int mtk_ctrl_init(struct mtk_md_dev *mdev)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk;
> +
> +	ctrl_blk = devm_kzalloc(mdev->dev, sizeof(*ctrl_blk), GFP_KERNEL);
> +	if (!ctrl_blk)
> +		return -ENOMEM;
> +
> +	ctrl_blk->mdev = mdev;
> +	mdev->ctrl_blk = ctrl_blk;
> +
> +	return 0;
> +}

[Severity: Low]

This isn't a bug, but at this point in the series nothing calls
mtk_ctrl_init() or mtk_ctrl_exit(). That means:

- mdev->ctrl_blk in struct mtk_md_dev is never filled in.
- Nothing creates a struct mtk_ctrl_trans.
- Nothing in pcie/mtk_pci.c uses the newly added include:

    #include "mtk_trans_ctrl.h"

The subject says "Add control plane transaction layer", but this patch
only adds structures and helpers that nothing can reach yet.

The next commit, "net: wwan: t9xx: Add control DMA interface", adds the
callers:

    mtk_pci_probe()->mtk_pci_dev_init()->mtk_trans_ctrl_init()->mtk_ctrl_init()
    mtk_pci_remove()->mtk_pci_dev_exit()->mtk_trans_ctrl_exit()->mtk_ctrl_exit()

Could these helpers and the include move into the patch that adds their
callers? Otherwise, could the subject say that this patch only adds the
scaffolding?

> +
> +/**
> + * mtk_ctrl_exit() - Clean up the control plane block.
> + * @mdev: Pointer to the MTK modem device.
> + *
> + * Clears the control plane block pointer. The allocation
> + * itself is managed by devres and freed on driver detach.
> + */
> +void mtk_ctrl_exit(struct mtk_md_dev *mdev)
> +{
> +	mdev->ctrl_blk = NULL;
> +}

[Severity: Low]

This isn't a bug, but mtk_ctrl_init() allocates ctrl_blk with
devm_kzalloc(), and mtk_ctrl_exit() only clears mdev->ctrl_blk. Suppose
init and exit ran more than once while the device stayed bound. Each
mtk_ctrl_init() call would add another devres allocation and overwrite
mdev->ctrl_blk without freeing the old block.

That can't happen at the end of the series. The only path to
mtk_ctrl_init() is mtk_pci_probe()->mtk_pci_dev_init()->
mtk_trans_ctrl_init(). mtk_ctrl_exit() is reached only from
mtk_pci_remove() or the probe error path. Devres then frees the block at
unbind, as the kerneldoc says.

By the end of the series, mtk_ctrl_exit() also calls
mtk_fsm_notifier_unregister() and mtk_port_mngr_exit(). So the teardown
hook does end up in place.

Would it help to add a comment saying mtk_ctrl_init() may only run once
per bind? That would keep a future reset path from calling it again.

> diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.h b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
> new file mode 100644
> index 0000000000000..c141876ef95d3
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.h
> @@ -0,0 +1,22 @@

[ ... ]

> +struct mtk_ctrl_blk {
> +	struct mtk_md_dev *mdev;
> +	struct mtk_ctrl_trans *trans;
> +};

[Severity: Low]

This isn't a bug, but nothing assigns ctrl_blk->trans in this patch.
The back-pointers in the new struct mtk_ctrl_trans in
pcie/mtk_trans_ctrl.h are never set either:

    struct mtk_ctrl_trans {
    	struct mtk_ctrl_blk *ctrl_blk;
    	struct mtk_md_dev *mdev;
    };

mtk_ctrl_plane.h is a generic header, but it names the pcie-only struct
mtk_ctrl_trans. It does not include or forward-declare the header that
defines it. The code still builds, because C allows a member that points
to an undeclared struct tag.

The next commit, "net: wwan: t9xx: Add control DMA interface", replaces
this member with void *ctrl_hw_priv. Later in the series,
mtk_trans_ctrl_init() sets trans->mdev and mtk_pcie_hif_init() sets
trans->ctrl_blk.

Could the ctrl_hw_priv member be added here directly? That would avoid
adding a trans member that is removed one patch later.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface
  2026-09-30  7:46 ` [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found 12 potential
issue(s) to consider.

Critical: 0 · High: 3 · Medium: 8 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] Use-after-free of struct trb_srv on HIF teardown.
- [High] Race between mtk_cldma_rx_done_work() and mtk_cldma_rxq_free()
  that restarts DMA into freed memory.
- [High] net: wwan: t9xx: use-after-free when deleting radix tree entries
  during iteration
- [Medium] In BD mode, mtk_cldma_rx_skb_adjust() limits each BD's
  recv_len only by data_allow_len, which it reads back from the device-…
- [Medium] mtk_pcie_hif_submit_skb() inserts a DISABLE before the first
  non-ENABLE entry in the queue list.
- [Medium] In non-BD mode, when mtk_cldma_reload_rx_skb() fails and the
  recycle dma_map_single() also fails, mtk_cldma_rx_done_work() jumps to…
- [Medium] In mtk_ch_status_check(), a TRB_CMD_ENABLE increments
  usr_cnt[hif_id][txqno] unconditionally.
- [Medium] Pending TRBs in skb_list are flushed (completed with -EIO)
  only at the end of mtk_ctrl_trb_thread(), via mtk_ctrl_chs_flush().
- [Medium] mtk_cldma_drv_reset() drives the infracfg RST0_SET and
  RST0_CLR registers with read-modify-write: 'val = read(SET); val |=…
- [Medium] When mtk_cldma_start_xfer() fails, mtk_cldma_tx() calls
  mtk_cldma_txq_flush() without stopping the TX queue first.
- [Medium] mtk_cldma_rearm_queues() restarts each TX queue at
  req_pool[free_idx] without reclaiming completed descriptors first.
- [Low] mtk_ctrl_trb_handler() keeps using the skb after
  mtk_cldma_submit_tx() has published it to the hardware ring.

--

net: wwan: t9xx: Add control DMA interface

This commit adds the CLDMA (Cross Layer DMA) driver for the t9xx control
plane. The driver manages GPD/BD rings for TX and RX, and the patch also
adds a transport control layer with a TRB service kthread. Probe now
registers the transport plane and the queue info table. The FSM listener
that brings CLDMA up arrives later in the series.

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
> new file mode 100644
> index 0000000000000..0f281bfb38cc7
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
> @@ -0,0 +1,1520 @@

[ ... ]

> +static int mtk_cldma_rx_skb_adjust(struct mtk_md_dev *mdev, struct rxq *rxq,
> +				   struct rx_req *req)
> +{

[ ... ]

> +		recv_len = le16_to_cpu(bd_dsc->bd->rx_bd.data_recv_len);
> +		allow_len = le16_to_cpu(bd_dsc->bd->rx_bd.data_allow_len);
> +		if (recv_len > allow_len) {
> +			ret = -EPROTO;
> +			recv_len = allow_len;
> +		}
> +		total_len += recv_len;
> +		if (total_len > req->mtu) {
> +			ret = -EPROTO;
> +			recv_len -= min(total_len - req->mtu, recv_len);
> +			total_len = req->mtu;
> +		}
> +		bd_dsc->skb->len = 0;
> +		skb_reset_tail_pointer(bd_dsc->skb);
> +		skb_put(bd_dsc->skb, recv_len);

[Severity: Medium]
Can a device overrun a single BD skb here?

The only per-BD bound is data_allow_len, and that value is read back from
the dma_pool BD, which the device can write. The BD skbs themselves are
only frag_size bytes:

mtk_cldma_alloc_rx_bd()
    bd_dsc->skb = __dev_alloc_skb(req->frag_size, GFP_KERNEL);
    ...
    bd_dsc->bd->rx_bd.data_allow_len =
        cpu_to_le16(min(req->frag_size, left_size));

mtk_cldma_reload_rx_skb() uses the same size.

Suppose the device writes data_allow_len and data_recv_len up to mtu into
BD 0. Then recv_len passes both checks, and skb_put() runs past skb->end
into skb_over_panic(). The comment above the function says "a device that
inflates data_allow_len cannot grow the packet past the buffer". That is
true of the running total, but not of each BD.

Should recv_len also be clamped to req->frag_size?

BD mode is not configured today, because Q_MTU_3_5K == Q_FRAG_3_5K keeps
nr_bds at 0. This path becomes reachable as soon as a queue uses
mtu > frag_size.

[ ... ]

> +		ret = mtk_cldma_reload_rx_skb(mdev, rxq, req);
> +		if (ret) {

[ ... ]

> +			skb_trim(rx_skb, 0);
> +			req->skb = rx_skb;
> +			req->data_dma_addr = dma_map_single(mdev->dev,
> +							    rx_skb->data,
> +							    req->mtu,
> +							    DMA_FROM_DEVICE);
> +			if (dma_mapping_error(mdev->dev, req->data_dma_addr)) {
> +				req->data_dma_addr = 0;
> +				/* Keep rx_skb in req->skb for clean stall.
> +				 * HWO is not set — HW won't touch this slot.
> +				 * Queue stalls until modem reset recovery.
> +				 */
> +				goto out;

[Severity: Medium]
What brings this slot back after the goto out?

At this point req->skb is rx_skb trimmed to 0, data_dma_addr is 0, HWO is
clear and data_recv_len is already 0. free_idx has not advanced, the
previous GPD has not been re-armed, and the queue has not been resumed.
Nothing marks the slot dead or schedules a retry.

On the next run of mtk_cldma_rx_done_work(), the slot passes both the
!req->skb check and the HWO check as a completed descriptor.
mtk_cldma_rx_skb_adjust() then does skb_put(skb, 0). If the reload
succeeds this time, a zero-length skb goes up through rxq->rx_done().

If no further completion arrives at all, doesn't RX stay stalled even
after memory pressure clears? Nothing appears to implement the "clean
stall until modem reset recovery" in the comment.

[ ... ]

> +	if (!atomic_read(&rxq->need_exit)) {
> +		if (atomic_xchg(&rxq->need_restart, 0))
> +			mtk_cldma_rxq_restart(drv_info, rxq);
> +		else if (ret != -ENXIO)
> +			mtk_cldma_resume_queue(drv_info, DIR_RX, rxq->rxqno);
> +	}

[Severity: High]
Is the need_exit check atomic with respect to mtk_cldma_rxq_free()?
Consider this ordering:

rx_done_work                        mtk_cldma_rxq_free()
atomic_read(need_exit) == 0
<preempted>
                                    atomic_set(&rxq->need_exit, 1);
                                    mtk_cldma_stop_queue() returns 0
                                    synchronize_irq()
                                    flush_work(&rxq->rx_done_work)
mtk_cldma_resume_queue() or
mtk_cldma_rxq_restart()
return
                                    ret == 0, unmap and free RX skbs,
                                    BDs and GPDs

mtk_cldma_rxq_free() does not stop the queue again after flush_work(). Can
the engine then be running on GPDs with HWO set while their buffers and
descriptors are being unmapped and freed?

[ ... ]

> +static void mtk_cldma_rearm_queues(struct cldma_drv_info *drv_info)
> +{

[ ... ]

> +		txq = drv_info->txq[i];
> +		if (txq) {
> +			spin_lock(&txq->ring_lock);
> +			mtk_cldma_setup_start_addr(drv_info, DIR_TX, i,
> +						   txq->req_pool[txq->free_idx].gpd_dma_addr);
> +			mtk_cldma_unmask_intr(drv_info, DIR_TX, i, QUEUE_ERROR);
> +			mtk_cldma_unmask_intr(drv_info, DIR_TX, i, QUEUE_XFER_DONE);
> +			if (READ_ONCE(txq->tx_started))
> +				mtk_cldma_start_queue(drv_info, DIR_TX, i);
> +			spin_unlock(&txq->ring_lock);
> +		}

[Severity: Medium]
What happens if free_idx names a GPD the hardware has already completed
(HWO=0) but mtk_cldma_tx_done_work() has not reclaimed yet?

In that case the engine would restart at that descriptor, stop there, and
never reach the HWO=1 descriptors behind it. tx_started stays true, so
later kicks from mtk_cldma_start_xfer() only resume:

mtk_cldma_start_xfer()
    } else {
        mtk_cldma_resume_queue(drv_info, DIR_TX, qno);
    }

That never moves the hardware cursor. Could the pending slots between the
old free_idx and wr_idx then stay unprocessed, draining req_budget to zero
and stalling TX? This depends on how the hardware handles RESUME at an
HWO=0 descriptor.

[ ... ]

> +static int mtk_cldma_tx(struct cldma_dev *cd, struct sk_buff *skb)
> +{

[ ... ]

> +	ret = mtk_cldma_start_xfer(drv_info, que->txqno);
> +	if (unlikely(ret)) {
> +		dev_err(mdev->dev, "Failed to trigger cldma tx\n");
> +		mtk_cldma_txq_flush(drv_info, txq, ret);
> +	}

[Severity: Medium]
Is it safe to flush here without stopping the TX queue first?

mtk_cldma_start_xfer() can fail after mtk_cldma_hw_recovery() ->
mtk_cldma_rearm_queues() has already issued UL_START. It can also fail on
a transient all-ones read. mtk_cldma_txq_flush() then does three things
on descriptors the engine may still own:

  - clears CLDMA_GPD_FLAG_HWO in place
  - dma_unmap_single()s their buffers
  - advances free_idx

tx_started also stays true. The next successful kick therefore only
resumes from the hardware's stale current-GPD pointer, which now names an
HWO=0 slot behind wr_idx.

The err_work added later in the series stops the queue before flushing,
with the comment "the device may still be walking the ring: unmapping its
buffers here would leave it writing into unmapped memory".

Should this path also stop the queue and clear tx_started before calling
mtk_cldma_txq_flush()?

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c
> new file mode 100644
> index 0000000000000..8ab153da2083e
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma_drv.c
> @@ -0,0 +1,561 @@

[ ... ]

> +void mtk_cldma_drv_reset(struct cldma_drv_info *drv_info)
> +{
> +	struct cldma_hw_regs *hw_regs;
> +	struct mtk_md_dev *mdev;
> +	u32 val;
> +
> +	mdev = drv_info->mdev;
> +	hw_regs = drv_info->hw_regs;
> +
> +	val = mtk_pci_read32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_set);
> +	val |= 1 << (REG_CLDMA0_RST_SET_BIT + drv_info->hw_id);
> +	mtk_pci_write32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_set, val);
> +	udelay(1);
> +	val = mtk_pci_read32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_clr);
> +	val |= 1 << (REG_CLDMA0_RST_CLR_BIT + drv_info->hw_id);
> +	mtk_pci_write32(mdev, REG_DEV_INFRA_BASE + hw_regs->reg_infra_rst0_clr, val);
> +}

[Severity: Medium]
Are RST0_SET and RST0_CLR write-1-to-set and write-1-to-clear registers?
On MediaTek set/clear reset banks every 1 written is normally a command,
and the in-tree MediaTek reset code writes only BIT(id).

If these registers read back the current reset status, the
read-modify-write on RST0_CLR would release every other block in bank 0
that is currently held in reset. Would writing just the CLDMA bit be
safer?

Separately, udelay(1) follows a posted write with no read-back. Is the
reset assert width actually guaranteed?

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
> new file mode 100644
> index 0000000000000..531940b06e2a9
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
> @@ -0,0 +1,615 @@

[ ... ]

> +	switch (trb->cmd) {
> +	case TRB_CMD_ENABLE:
> +		trb_open_priv = (struct trb_open_priv *)skb->data;
> +		trb_open_priv->log_rg_offset = que->log_rg_offset;
> +		trans->usr_cnt[que->hif_id][que->txqno]++;
> +		if (trans->usr_cnt[que->hif_id][que->txqno] == 1)
> +			break;
> +		trb_open_priv->tx_mtu = que->tx_mtu;
> +		trb_open_priv->rx_mtu = que->rx_mtu;
> +		trb_open_priv->tx_frag_size = que->tx_frag_size;
> +		trb_open_priv->rx_frag_size = que->rx_frag_size;
> +		if (mtk_cldma_check_ch_cfg(trans->dev, que)) {
> +			trb->status = -EINVAL;
> +			ret = -EINVAL;
> +		} else {

[Severity: Medium]
Should usr_cnt be rolled back when the ENABLE is rejected with -EINVAL?

mtk_cldma_open() drops the count on its own failure paths:

mtk_cldma_open()
    if (ret)
        cd->trans->usr_cnt[que->hif_id][que->txqno]--;

This path leaves the increment in place. One leaked increment keeps
usr_cnt at 1 or more. The real owner's DISABLE would then never reach
mtk_cldma_close(), and later ENABLEs would keep failing until
mtk_pcie_hif_init() resets the count.

[ ... ]

> +		case TRB_CMD_TX:
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			err = mtk_cldma_submit_tx(trans->dev, skb);

[ ... ]

> +			trans_list->tx_burst_cnt[qno]++;
> +			spin_lock_irqsave(&skb_list->lock, flags);
> +			if (trans_list->tx_burst_cnt[qno] >= TX_BURST_MAX_CNT ||
> +			    skb_queue_is_last(skb_list, skb)) {
> +				kick = true;
> +			} else {
> +				skb_next = skb_peek_next(skb, skb_list);
> +				trb_next = (struct trb *)skb_next->cb;
> +				if (trb_next->cmd != TRB_CMD_TX)
> +					kick = true;
> +			}
> +
> +			__skb_unlink(skb, skb_list);
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			break;

[Severity: Low]
Once mtk_cldma_submit_tx() has published the skb to the ring,
mtk_cldma_tx_done_work() can complete it through trb_complete(). The
handler then still uses the skb for skb_queue_is_last(), __skb_unlink()
and mtk_cldma_trb_process(trans->dev, skb). What keeps it alive?

At this commit the path is not reachable, because submit_tx always fails
and trb_complete does not free anything. The later patch "net: wwan: t9xx:
Add control port" adds kref_get(&trb->kref) in mtk_ctrl_trb_handler()
under skb_list->lock, which covers this.

Would it be clearer to take that reference here, where the code is
introduced?

[ ... ]

> +static int mtk_ctrl_trb_thread(void *args)
> +{

[ ... ]

> +	mtk_ctrl_chs_flush(srv);
> +	return 0;
> +}

[Severity: Medium]
What happens to queued TRBs if kthread_stop() arrives before this thread
has run for the first time?

kthread() skips threadfn entirely when KTHREAD_SHOULD_STOP is already set,
so mtk_ctrl_chs_flush() would never run. mtk_pcie_hif_init() sets
available to 1 right after kthread_run(), so skbs can be queued before
the thread is scheduled.

mtk_ctrl_trb_srv_exit() does not flush the lists itself, and the next
mtk_pcie_hif_init() calls skb_queue_head_init() over the orphaned
entries. Are those skbs then leaked and their waiters never completed?

Would flushing in mtk_ctrl_trb_srv_exit() after kthread_stop() cover this?

[ ... ]

> +int mtk_pcie_hif_exit(struct mtk_md_dev *mdev)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
> +	struct mtk_ctrl_trans *trans;
> +
> +	trans = ctrl_blk->ctrl_hw_priv;
> +
> +	mutex_lock(&trans->submit_lock);
> +	atomic_set(&trans->available, 0);
> +	mutex_unlock(&trans->submit_lock);
> +
> +	/* Join the service threads before freeing what they dereference: they
> +	 * read trans->dev without a lock, so clearing it cannot stop a
> +	 * consumer that has already loaded the pointer.
> +	 */
> +	mtk_ctrl_trb_srv_exit(trans);
> +	mtk_cldma_exit(trans);

[Severity: High]
Can this free struct trb_srv while the CLDMA completion paths still
dereference it?

mtk_ctrl_trb_srv_exit() calls kthread_stop() and kfree(srv), and only
afterwards sets trans->trb_srv[i] = NULL. The CLDMA IRQ and TX queues are
torn down after that point (at the end of the series in
mtk_cldma_exit()->mtk_cldma_dev_exit()). Until then,
mtk_cldma_tx_done_work() and mtk_cldma_txq_flush() can still run:

mtk_cldma_tx_done_work()
    srv = trans->trb_srv[trans->srv_cfg[hif_id][txqno]];
    if (was_starved && srv)
        wake_up(&srv->trb_waitq);

A worker can load srv before the NULL store, or between kfree() and the
NULL store. wake_up() would then take the waitqueue lock inside the freed
trb_srv.

The comment "The service may already have been torn down" treats the NULL
check as sufficient, but nothing orders that check against the free.

Would it be safer to quiesce the CLDMA IRQ and workers before freeing the
service?

[ ... ]

> +	if (trb->cmd == TRB_CMD_DISABLE) {
> +		struct sk_buff *entry = NULL;
> +		struct sk_buff_head *list;
> +		struct sk_buff *iter;
> +		unsigned long flags;
> +
> +		/* A disable may overtake queued data, so teardown does not
> +		 * wait for a TX backlog, but it must never overtake a pending
> +		 * ENABLE for the same queue: the two do not commute, and a
> +		 * DISABLE consumed before its ENABLE closes nothing while the
> +		 * ENABLE then arms rings nobody owns.
> +		 */
> +		list = &trans->trans_list[que->hif_id].skb_list[que->txqno];
> +		spin_lock_irqsave(&list->lock, flags);
> +		skb_queue_walk(list, iter) {
> +			if (((struct trb *)iter->cb)->cmd != TRB_CMD_ENABLE) {
> +				entry = iter;
> +				break;
> +			}
> +		}
> +		if (entry)
> +			__skb_queue_before(list, entry, skb);
> +		else
> +			__skb_queue_tail(list, skb);

[Severity: Medium]
Does this walk hold up the promise in the comment when an ENABLE sits
behind a TX entry?

The walk stops at the first non-ENABLE entry. With a list of TX, ENABLE,
a new DISABLE is inserted at the front, which gives DISABLE, TX, ENABLE.

The DISABLE is then handled while usr_cnt is 0. It completes with -EBUSY
in mtk_ch_status_check() without closing anything. The later ENABLE then
arms the rings after the close has already been reported.

Should the DISABLE go after the last pending ENABLE for this queue
instead?

[ ... ]

> @@ -0,0 +1,615 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/*
> + * Copyright (c) 2022, MediaTek Inc.
> + */
> +
> +#include <linux/device.h>
> +#include <linux/freezer.h>
> +#include <linux/hashtable.h>
> +#include <linux/kthread.h>
> +#include <linux/list.h>
> +#include <linux/nospec.h>
> +#include <linux/sched.h>
> +#include <linux/wait.h>
> +
> +#include "mtk_cldma.h"
> +#include "mtk_ctrl_plane.h"
> +#include "mtk_dev.h"
> +#include "mtk_pci.h"
> +#include "mtk_trans_ctrl.h"
> +
> +#define QUEUE_CHL_MASK	0xFFFF
> +#define TRB_SRV_NUM	(1)
> +
> +static const int mtk_srv_cfg[NR_CLDMA][HW_QUE_NUM] = {
> +	{0},
> +	{0},
> +};
> +
> +/* the number of RX GPDs should be at least two */
> +static const struct queue_info mtk_queue_info[] = {
> +	{CCCI_CONTROL_TX, CCCI_CONTROL_RX, CLDMA1, TXQ(0), RXQ(0),
> +	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
> +	{CCCI_SAP_CONTROL_TX, CCCI_SAP_CONTROL_RX, CLDMA0, TXQ(0), RXQ(0),
> +	 Q_MTU_3_5K, Q_MTU_3_5K, TX_GPD_NUM, RX_GPD_NUM, Q_FRAG_3_5K, Q_FRAG_3_5K, 0},
> +};
> +
> +static bool mtk_queue_list_is_full(struct mtk_ctrl_trans *trans, struct queue_info *que)
> +{
> +	return skb_queue_len_lockless(&trans->trans_list[que->hif_id].skb_list[que->txqno]) >=
> +	       SKB_LIST_MAX_LEN;
> +}
> +
> +static bool mtk_ctrl_chs_is_busy_or_empty(struct trb_srv *srv)
> +{
> +	struct srv_que *srv_que;
> +	int i;
> +
> +	for (i = 0; i < NR_CLDMA; i++) {
> +		list_for_each_entry(srv_que, &srv->srv_q_list[i], list) {
> +			struct sk_buff *skb;
> +			struct trb *trb;
> +
> +			skb = skb_peek(&srv->trans->trans_list[i].skb_list[srv_que->qno]);
> +			if (!skb)
> +				continue;
> +
> +			/* ENABLE and DISABLE are software-only and are queued at
> +			 * the head, so gating them on TX budget would make a queue
> +			 * that cannot drain impossible to close.
> +			 */
> +			trb = (struct trb *)skb->cb;
> +			if (trb->cmd != TRB_CMD_TX ||
> +			    mtk_cldma_get_tx_budget(srv->trans->dev, i, srv_que->qno))
> +				return false;
> +		}
> +	}
> +
> +	return true;
> +}
> +
> +static void mtk_ctrl_ch_flush(struct sk_buff_head *skb_list)
> +{
> +	struct sk_buff *skb;
> +	struct trb *trb;
> +
> +	while (!skb_queue_empty(skb_list)) {
> +		skb = skb_dequeue(skb_list);
> +		trb = (struct trb *)skb->cb;
> +		trb->status = -EIO;
> +		trb->trb_complete(skb);
> +	}
> +}
> +
> +static void mtk_ctrl_chs_flush(struct trb_srv *srv)
> +{
> +	struct srv_que *srv_que;
> +	int i;
> +
> +	for (i = 0; i < NR_CLDMA; i++)
> +		list_for_each_entry(srv_que, &srv->srv_q_list[i], list)
> +			mtk_ctrl_ch_flush(&srv->trans->trans_list[i].skb_list[srv_que->qno]);
> +}
> +
> +static int mtk_ch_status_check(struct mtk_ctrl_trans *trans, struct sk_buff *skb)
> +{
> +	struct trb *trb = (struct trb *)skb->cb;
> +	struct trb_open_priv *trb_open_priv;
> +	struct queue_info *que;
> +	int ret = 0;
> +
> +	que = radix_tree_lookup(&trans->queue_tbl, trb->channel_id & QUEUE_CHL_MASK);
> +
> +	switch (trb->cmd) {
> +	case TRB_CMD_ENABLE:
> +		trb_open_priv = (struct trb_open_priv *)skb->data;
> +		trb_open_priv->log_rg_offset = que->log_rg_offset;
> +		trans->usr_cnt[que->hif_id][que->txqno]++;
> +		if (trans->usr_cnt[que->hif_id][que->txqno] == 1)
> +			break;
> +		trb_open_priv->tx_mtu = que->tx_mtu;
> +		trb_open_priv->rx_mtu = que->rx_mtu;
> +		trb_open_priv->tx_frag_size = que->tx_frag_size;
> +		trb_open_priv->rx_frag_size = que->rx_frag_size;
> +		if (mtk_cldma_check_ch_cfg(trans->dev, que)) {
> +			trb->status = -EINVAL;
> +			ret = -EINVAL;
> +		} else {
> +			trb->status = -EBUSY;
> +			ret = -EBUSY;
> +		}
> +		trb->trb_complete(skb);
> +		break;
> +	case TRB_CMD_DISABLE:
> +		if (trans->usr_cnt[que->hif_id][que->txqno] > 0) {
> +			trans->usr_cnt[que->hif_id][que->txqno]--;
> +			if (!trans->usr_cnt[que->hif_id][que->txqno])
> +				break;
> +		}
> +		trb->status = -EBUSY;
> +		trb->trb_complete(skb);
> +		ret = -EBUSY;
> +		break;
> +	default:
> +		dev_err((trans->mdev)->dev, "Invalid trb command(%d)\n", trb->cmd);
> +		ret = -EINVAL;
> +		break;
> +	}
> +	return ret;
> +}
> +
> +/* Single consumer per srv_que — only this kthread dequeues from skb_list.
> + * The list lock is held around every list read (peek, is_last, peek_next,
> + * unlink) so a producer inserting a DISABLE at the head cannot race the
> + * traversal, but it is dropped across submit and dispatch: those paths
> + * may allocate with GFP_KERNEL and thus sleep. Dropping the lock there
> + * is safe because no other consumer can steal the peeked skb.
> + */
> +static void mtk_ctrl_trb_handler(struct trb_srv *srv, struct trans_list *trans_list, u32 qno)
> +{
> +	struct sk_buff_head *skb_list = &trans_list->skb_list[qno];
> +	struct mtk_ctrl_trans *trans = srv->trans;
> +	struct sk_buff *skb, *skb_next;
> +	struct trb *trb, *trb_next;
> +	unsigned long flags;
> +	bool kick = false;
> +	int loop = 0;
> +	int err;
> +
> +	do {
> +		spin_lock_irqsave(&skb_list->lock, flags);
> +		skb = skb_peek(skb_list);
> +		if (!skb) {
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			break;
> +		}
> +		trb = (struct trb *)skb->cb;
> +
> +		switch (trb->cmd) {
> +		case TRB_CMD_ENABLE:
> +		case TRB_CMD_DISABLE:
> +			__skb_unlink(skb, skb_list);
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			err = mtk_ch_status_check(trans, skb);
> +			if (!err) {
> +				kick = true;
> +				if (trb->cmd == TRB_CMD_DISABLE)
> +					mtk_ctrl_ch_flush(skb_list);
> +			}
> +			break;
> +		case TRB_CMD_TX:
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			err = mtk_cldma_submit_tx(trans->dev, skb);
> +			if (err) {
> +				if (trans_list->tx_burst_cnt[qno]) {
> +					kick = true;
> +					break;
> +				}
> +				if (err == -EAGAIN)
> +					return;
> +
> +				skb_unlink(skb, skb_list);
> +				trb->status = err;
> +				trb->trb_complete(skb);
> +				break;
> +			}
> +
> +			trans_list->tx_burst_cnt[qno]++;
> +			spin_lock_irqsave(&skb_list->lock, flags);
> +			if (trans_list->tx_burst_cnt[qno] >= TX_BURST_MAX_CNT ||
> +			    skb_queue_is_last(skb_list, skb)) {
> +				kick = true;
> +			} else {
> +				skb_next = skb_peek_next(skb, skb_list);
> +				trb_next = (struct trb *)skb_next->cb;
> +				if (trb_next->cmd != TRB_CMD_TX)
> +					kick = true;
> +			}
> +
> +			__skb_unlink(skb, skb_list);
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			break;
> +		default:
> +			__skb_unlink(skb, skb_list);
> +			spin_unlock_irqrestore(&skb_list->lock, flags);
> +			trb->status = -EINVAL;
> +			trb->trb_complete(skb);
> +			break;
> +		}
> +
> +		if (kick) {
> +			err = mtk_cldma_trb_process(trans->dev, skb);
> +			if (err)
> +				dev_err_ratelimited((trans->mdev)->dev,
> +						    "Failed to process trb on queue %u: %d\n",
> +						    qno, err);
> +			trans_list->tx_burst_cnt[qno] = 0;
> +			kick = false;
> +		}
> +
> +		loop++;
> +	} while (loop < TRB_NUM_PER_ROUND);
> +}
> +
> +static void mtk_ctrl_trb_process(struct trb_srv *srv)
> +{
> +	struct mtk_ctrl_trans *trans = srv->trans;
> +	struct srv_que *srv_que;
> +	int i;
> +
> +	for (i = 0; i < NR_CLDMA; i++)
> +		list_for_each_entry(srv_que, &srv->srv_q_list[i], list)
> +			mtk_ctrl_trb_handler(srv, &trans->trans_list[i], srv_que->qno);
> +}
> +
> +static int mtk_ctrl_trb_thread(void *args)
> +{
> +	struct trb_srv *srv = args;
> +
> +	for (;;) {
> +		wait_event_interruptible(srv->trb_waitq,
> +					 !mtk_ctrl_chs_is_busy_or_empty(srv) ||
> +					 kthread_should_stop() || kthread_should_park());
> +		if (kthread_should_stop())
> +			break;
> +
> +		if (kthread_should_park())
> +			kthread_parkme();
> +
> +		do {
> +			mtk_ctrl_trb_process(srv);
> +			cond_resched();
> +		} while (!mtk_ctrl_chs_is_busy_or_empty(srv) && !kthread_should_stop() &&
> +			 !kthread_should_park());
> +	}
> +	mtk_ctrl_chs_flush(srv);
> +	return 0;
> +}
> +
> +static int mtk_ctrl_trb_srv_init(struct mtk_ctrl_trans *trans)
> +{
> +	struct srv_que *srv_que;
> +	struct trb_srv *srv;
> +	int i, j;
> +	int ret;
> +
> +	for (i = 0; i < trans->trb_srv_num; i++) {
> +		srv = kzalloc_obj(*srv);
> +		if (!srv) {
> +			ret = -ENOMEM;
> +			goto err_free_srv;
> +		}
> +
> +		srv->trans = trans;
> +		srv->srv_id = i;
> +		trans->trb_srv[i] = srv;
> +
> +		init_waitqueue_head(&srv->trb_waitq);
> +		for (j = 0; j < NR_CLDMA; j++)
> +			INIT_LIST_HEAD(&srv->srv_q_list[j]);
> +	}
> +
> +	for (i = 0; i < NR_CLDMA; i++)
> +		for (j = 0; j < HW_QUE_NUM; j++) {
> +			if (trans->srv_cfg[i][j] < 0 ||
> +			    trans->srv_cfg[i][j] >= trans->trb_srv_num)
> +				trans->srv_cfg[i][j] = 0;
> +			srv_que = kzalloc_obj(*srv_que);
> +			if (!srv_que) {
> +				ret = -ENOMEM;
> +				goto err_free_srv_que;
> +			}
> +			srv_que->hif_id = i;
> +			srv_que->qno = j;
> +			list_add_tail(&srv_que->list,
> +				      &trans->trb_srv[trans->srv_cfg[i][j]]->srv_q_list[i]);
> +		}
> +
> +	for (i = 0; i < trans->trb_srv_num; i++) {
> +		trans->trb_srv[i]->trb_thread = kthread_run(mtk_ctrl_trb_thread, trans->trb_srv[i],
> +							    "mtk_trb_srv%d_%s", i,
> +							    trans->mdev->dev_str);
> +		if (IS_ERR(trans->trb_srv[i]->trb_thread)) {
> +			ret = PTR_ERR(trans->trb_srv[i]->trb_thread);
> +			trans->trb_srv[i]->trb_thread = NULL;
> +			goto err_stop_kthread;
> +		}
> +	}
> +
> +	return 0;
> +err_stop_kthread:
> +	while (--i >= 0)
> +		kthread_stop(trans->trb_srv[i]->trb_thread);
> +err_free_srv_que:
> +	for (i = 0; i < trans->trb_srv_num; i++) {
> +		for (j = 0; j < NR_CLDMA; j++) {
> +			struct srv_que *next_srv_que;
> +
> +			list_for_each_entry_safe(srv_que, next_srv_que,
> +						 &trans->trb_srv[i]->srv_q_list[j], list) {
> +				list_del(&srv_que->list);
> +				kfree(srv_que);
> +			}
> +		}
> +	}
> +err_free_srv:
> +	for (i = 0; i < trans->trb_srv_num; i++) {
> +		if (!trans->trb_srv[i])
> +			break;
> +		kfree(trans->trb_srv[i]);
> +		trans->trb_srv[i] = NULL;
> +	}
> +
> +	return ret;
> +}
> +
> +static void mtk_ctrl_trb_srv_exit(struct mtk_ctrl_trans *trans)
> +{
> +	struct srv_que *srv_que, *next_srv_que;
> +	struct trb_srv *srv;
> +	int i, j;
> +
> +	for (i = 0; i < trans->trb_srv_num; i++) {
> +		srv = trans->trb_srv[i];
> +		if (!srv)
> +			continue;
> +		kthread_stop(srv->trb_thread);
> +		for (j = 0; j < NR_CLDMA; j++) {
> +			list_for_each_entry_safe(srv_que, next_srv_que,
> +						 &trans->trb_srv[i]->srv_q_list[j], list) {
> +				list_del(&srv_que->list);
> +				kfree(srv_que);
> +			}
> +		}
> +		kfree(srv);
> +		trans->trb_srv[i] = NULL;
> +	}
> +}
> +
> +static void mtk_ctrl_remove_radix_tree(struct mtk_ctrl_trans *trans)
> +{
> +	struct radix_tree_iter iter;
> +	struct queue_info *queue;
> +	void __rcu **slot;
> +
> +	radix_tree_for_each_slot(slot, &trans->queue_tbl, &iter, 0) {
> +		queue = radix_tree_deref_slot(slot);
> +		if (!queue)
> +			continue;
> +		radix_tree_delete(&trans->queue_tbl, iter.index);
> +		kfree(queue);
> +	}
> +}
> +
> +int mtk_pcie_hif_init(struct mtk_md_dev *mdev)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
> +	struct queue_info *queue, *queue_info;
> +	struct mtk_ctrl_trans *trans;
> +	int i, j;
> +	int ret;
> +
> +	trans = ctrl_blk->ctrl_hw_priv;
> +	trans->ctrl_blk = ctrl_blk;
> +	queue_info = trans->queue_info;
> +
> +	INIT_RADIX_TREE(&trans->queue_tbl, GFP_KERNEL);
> +	for (i = 0; i < trans->queue_info_num; i++) {
> +		queue = kmemdup(queue_info + i, sizeof(*queue), GFP_KERNEL);
> +		if (!queue) {
> +			ret = -ENOMEM;
> +			goto err_free_radix_tree;
> +		}
> +		if (queue->txqno >= HW_QUE_NUM || queue->rxqno >= HW_QUE_NUM ||
> +		    queue->hif_id >= NR_CLDMA) {
> +			dev_err(mdev->dev, "Failed to get correct queue info %x\n",
> +				queue->rx_chl);
> +			kfree(queue);
> +			ret = -EINVAL;
> +			goto err_free_radix_tree;
> +		}
> +		ret = radix_tree_insert(&trans->queue_tbl, queue->rx_chl & QUEUE_CHL_MASK, queue);
> +		if (ret) {
> +			dev_err(mdev->dev, "Insert %x fail, ret: %d", queue->rx_chl, ret);
> +			kfree(queue);
> +			goto err_free_radix_tree;
> +		}
> +	}
> +
> +	for (i = 0; i < NR_CLDMA; i++) {
> +		for (j = 0; j < HW_QUE_NUM; j++) {
> +			skb_queue_head_init(&trans->trans_list[i].skb_list[j]);
> +			trans->trans_list[i].tx_burst_cnt[j] = 0;
> +			/* usr_cnt tracks the queues rebuilt by mtk_cldma_init()
> +			 * below, so it must be reset with them. Otherwise a
> +			 * count left over from a torn-down cycle makes the
> +			 * channel permanently unopenable.
> +			 */
> +			trans->usr_cnt[i][j] = 0;
> +		}
> +	}
> +	ret = mtk_cldma_init(trans);
> +	if (ret)
> +		goto err_free_radix_tree;
> +
> +	ret = mtk_ctrl_trb_srv_init(trans);
> +	if (ret)
> +		goto err_cldma_exit;
> +
> +	atomic_set(&trans->available, 1);
> +
> +	return 0;
> +
> +err_cldma_exit:
> +	mtk_cldma_exit(trans);
> +err_free_radix_tree:
> +	mtk_ctrl_remove_radix_tree(trans);
> +
> +	return ret;
> +}
> +
> +int mtk_pcie_hif_exit(struct mtk_md_dev *mdev)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
> +	struct mtk_ctrl_trans *trans;
> +
> +	trans = ctrl_blk->ctrl_hw_priv;
> +
> +	mutex_lock(&trans->submit_lock);
> +	atomic_set(&trans->available, 0);
> +	mutex_unlock(&trans->submit_lock);
> +
> +	/* Join the service threads before freeing what they dereference: they
> +	 * read trans->dev without a lock, so clearing it cannot stop a
> +	 * consumer that has already loaded the pointer.
> +	 */
> +	mtk_ctrl_trb_srv_exit(trans);
> +	mtk_cldma_exit(trans);
> +
> +	/* Late submitters may still hold the lock and walk the tree. */
> +	mutex_lock(&trans->submit_lock);
> +	mtk_ctrl_remove_radix_tree(trans);
> +	mutex_unlock(&trans->submit_lock);
> +
> +	return 0;
> +}
> +
> +int mtk_pcie_hif_submit_skb(struct mtk_md_dev *mdev, struct sk_buff *skb, bool force_send)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
> +	struct mtk_ctrl_trans *trans;
> +	struct queue_info *que;
> +	struct trb *trb;
> +	int ret;
> +
> +	trans = ctrl_blk->ctrl_hw_priv;
> +	trb = (struct trb *)skb->cb;
> +
> +	if (trb->cmd == TRB_CMD_STOP || trb->cmd == TRB_CMD_RECOVER) {
> +		trb->trb_complete(skb);
> +		return 0;
> +	}
> +
> +	mutex_lock(&trans->submit_lock);
> +
> +	if (!atomic_read(&trans->available)) {
> +		ret = -EIO;
> +		goto unlock;
> +	}
> +
> +	que = radix_tree_lookup(&trans->queue_tbl, trb->channel_id & QUEUE_CHL_MASK);
> +	if (!que) {
> +		dev_warn(mdev->dev, "lookup que fail, ch_id: %x\n",
> +			 trb->channel_id);
> +		ret = -EINVAL;
> +		goto unlock;
> +	}
> +
> +	if (mtk_queue_list_is_full(trans, que) && !force_send) {
> +		ret = -EAGAIN;
> +		goto unlock;
> +	}
> +
> +	if (trb->cmd == TRB_CMD_DISABLE) {
> +		struct sk_buff *entry = NULL;
> +		struct sk_buff_head *list;
> +		struct sk_buff *iter;
> +		unsigned long flags;
> +
> +		/* A disable may overtake queued data, so teardown does not
> +		 * wait for a TX backlog, but it must never overtake a pending
> +		 * ENABLE for the same queue: the two do not commute, and a
> +		 * DISABLE consumed before its ENABLE closes nothing while the
> +		 * ENABLE then arms rings nobody owns.
> +		 */
> +		list = &trans->trans_list[que->hif_id].skb_list[que->txqno];
> +		spin_lock_irqsave(&list->lock, flags);
> +		skb_queue_walk(list, iter) {
> +			if (((struct trb *)iter->cb)->cmd != TRB_CMD_ENABLE) {
> +				entry = iter;
> +				break;
> +			}
> +		}
> +		if (entry)
> +			__skb_queue_before(list, entry, skb);
> +		else
> +			__skb_queue_tail(list, skb);
> +		spin_unlock_irqrestore(&list->lock, flags);
> +	} else {
> +		skb_queue_tail(&trans->trans_list[que->hif_id].skb_list[que->txqno], skb);
> +	}
> +
> +	wake_up(&trans->trb_srv[trans->srv_cfg[que->hif_id][que->txqno]]->trb_waitq);
> +	ret = 0;
> +
> +unlock:
> +	mutex_unlock(&trans->submit_lock);
> +	return ret;
> +}
> +
> +int mtk_pcie_hif_cmd_func(struct mtk_md_dev *mdev, int cmd, void *data)
> +{
> +	struct mtk_ctrl_blk *ctrl_blk = mdev->ctrl_blk;
> +	struct mtk_ctrl_trans *trans;
> +	struct queue_info *que;
> +	int ret;
> +
> +	switch (cmd) {
> +	case HIF_CTRL_CMD_CHECK_TX_FULL:
> +		trans = ctrl_blk->ctrl_hw_priv;
> +		mutex_lock(&trans->submit_lock);
> +		if (!atomic_read(&trans->available)) {
> +			ret = -EIO;
> +			break;
> +		}
> +		que = radix_tree_lookup(&trans->queue_tbl,
> +					((union ctrl_hif_cmd_data *)data)->rx_ch & QUEUE_CHL_MASK);
> +		if (!que) {
> +			dev_warn(mdev->dev, "Failed to find que to check tx full\n");
> +			ret = -EINVAL;
> +			break;
> +		}
> +		ret = mtk_queue_list_is_full(trans, que);
> +		break;
> +	default:
> +		return -EINVAL;
> +	}
> +	mutex_unlock(&trans->submit_lock);
> +
> +	return ret;
> +}
> +
> +int mtk_trans_ctrl_init(struct mtk_md_dev *mdev)
> +{
> +	struct mtk_ctrl_trans *trans;
> +	struct mtk_ctrl_blk *ctrl_blk;
> +	int err;
> +
> +	trans = devm_kzalloc(mdev->dev, sizeof(*trans), GFP_KERNEL);
> +	if (!trans)
> +		return -ENOMEM;
> +	trans->mdev = mdev;
> +	mutex_init(&trans->submit_lock);
> +	atomic_set(&trans->available, 0);
> +
> +	memcpy(trans->srv_cfg, mtk_srv_cfg, sizeof(mtk_srv_cfg));
> +	trans->queue_info = (struct queue_info *)mtk_queue_info;
> +	trans->queue_info_num = ARRAY_SIZE(mtk_queue_info);
> +	trans->trb_srv_num = TRB_SRV_NUM;
> +
> +	err = mtk_ctrl_init(mdev);
> +	if (err)
> +		return err;
> +
> +	ctrl_blk = mdev->ctrl_blk;
> +	ctrl_blk->ctrl_hw_priv = trans;
> +
> +	return 0;
> +}
> +
> +int mtk_trans_ctrl_exit(struct mtk_md_dev *mdev)
> +{
> +	mtk_ctrl_exit(mdev);
> +
> +	return 0;
> +}

[Severity: High]
Is it safe to call radix_tree_delete() from inside
radix_tree_for_each_slot() in mtk_ctrl_remove_radix_tree()?

radix_tree_delete() frees the node holding the current slot once its
last entry is gone, and it may also shrink the tree. The iterator
then advances by stepping slot within that same node, so the next
radix_tree_deref_slot(slot) reads from it. The node is released
through call_rcu(), but this walk holds no rcu_read_lock(), so
nothing keeps the node alive.

With the two channels in mtk_queue_info[] this is the normal case,
not a corner case. 0x1000 and 0x2000 land in separate leaf nodes of
a three level tree, so deleting the first entry already empties and
frees its node before the iterator moves on.

The walk runs from mtk_pcie_hif_exit() and from the error unwind in
mtk_pcie_hif_init(). Once the FSM patch later in the series calls
those, every teardown and every failed init goes through it.

radix_tree_iter_delete() exists for deleting at the iterator
position, and it moves the iterator past a freed node. Would that
work here? Or, since this is new code, would an xarray with
xa_for_each() and xa_erase() be simpler?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 4/6] net: wwan: t9xx: Add control port
  2026-09-30  7:46 ` [PATCH v9 4/6] net: wwan: t9xx: Add control port Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found 20 potential
issue(s) to consider.

Critical: 0 · High: 2 · Medium: 1 · Low: 17

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] net: core: skb_push() called without verifying headroom causes
  kernel panic
- [High] net: rcu: lockless radix tree lookup causes UAF in port search
- [Medium] mtk_port_ch_enable() treats `trb->status <= 0` as meaning the
  ENABLE completion has finished.
- [Low] The commit message leaves out two changes to existing code.
- [Low] Concern: at this commit nothing calls ->enable(),
  mtk_port_internal_open/write/recv_register or mtk_port_status_update,…
- [Low] The patch adds non-static globals with no driver prefix:
  `DEFINE_MUTEX(port_mngr_grp_mtx)` in mtk_port.c and `const struct…
- [Low] Two small quality problems.
- [Low] Concern: mtk_cldma_rx_done_work() retries rx_done() with the same
  skb when it returns -EAGAIN.
- [Low] The port layer handles an ENABLE that completes with -EBUSY
  inconsistently.
- [Low] Concern: the comment in mtk_port_internal_disable() refers to
  FSM_STATE_ON and to usr_cnt being rebuilt there, but neither exists at…
- [Low] mtk_port_send_data() can return the in-flight sentinel
  MTK_TRB_HEADER_ADDED (0xADDED = 712173) as if it were a byte count.
- [Low] Concern: rx_seq and tx_seq are initialised only once, in
  mtk_port_struct_init(), so the first packet after a modem reset or…
- [Low] Concern: mtk_port_status_update() reads the 12-byte enumeration
  header before checking data_len >= sizeof(*msg).
- [Low] Concern: in the scatter-gather RX path, -EIO means both 'partial
  dispatch' and a raw head-skb recv() failure.
- [Low] Concern: TRBs hold `trb->priv = port` without a port reference.
- [Low] Concern: mtk_port_internal_open() looks up the port by name
  without holding a lock, and only then takes kref_get() under…
- [Low] Concern: mtk_port_common_open()/mtk_port_internal_close() do a
  non-atomic test_bit+set_bit/clear_bit on PORT_S_OPEN, so concurrent…
- [Low] Concern: recv_cb/arg are published with plain stores and read
  without ordering, and close does not synchronise with in-flight…
- [Low] Concern: port->tx_seq++ and tx_seq-- are not protected, so
  concurrent writers could duplicate or reorder sequence numbers.
- [Low] Concern: the port layer's SG predicate (rx_mtu - 128 >
  rx_frag_size) differs from CLDMA's BD-mode predicate (rx_mtu >…

--

net: wwan: t9xx: Add control port

This commit adds a port manager and a port I/O layer to the t9xx WWAN
driver. Two internal control ports (MDCTRL, SAPCTRL) submit TRBs to CLDMA
and receive dispatched data. The port manager is hooked into
mtk_ctrl_init()/mtk_ctrl_exit(), and module_pci_driver() is replaced with
explicit module init/exit functions.

> diff --git a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
> index fa2ab8c3e7577..10f0d3267b3b1 100644
> --- a/drivers/net/wwan/t9xx/mtk_ctrl_plane.c
> +++ b/drivers/net/wwan/t9xx/mtk_ctrl_plane.c

[ ... ]

> @@ -28,7 +31,15 @@ int mtk_ctrl_init(struct mtk_md_dev *mdev)
>  	ctrl_blk->mdev = mdev;
>  	mdev->ctrl_blk = ctrl_blk;
>  
> +	err = mtk_port_mngr_init(ctrl_blk, port_layer_cfg->port_cfg,
> +				 port_layer_cfg->port_cnt);
> +	if (err)
> +		goto err_free_mem;
> +
>  	return 0;
> +
> +err_free_mem:
> +	return err;
>  }

[Severity: Low]
This isn't a bug, but the err_free_mem label frees nothing, because
ctrl_blk comes from devm_kzalloc(). Also, mdev->ctrl_blk still points at
ctrl_blk after the failure. Could this just return err, or clear
mdev->ctrl_blk and use a label name that matches what it does?

Similarly, mtk_trans_ctrl.c casts away const:

	.port_cfg = (struct mtk_port_cfg *)mtk_port_cfg_tbl,

mtk_port_mngr_init() and mtk_port_tbl_create() only read the table (via
memcpy() in mtk_port_alloc_and_add()). Could they take a
const struct mtk_port_cfg * so the cast isn't needed?

> diff --git a/drivers/net/wwan/t9xx/mtk_port.c b/drivers/net/wwan/t9xx/mtk_port.c
> new file mode 100644
> index 0000000000000..a8ff06ab4de4e
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/mtk_port.c
> @@ -0,0 +1,693 @@

[ ... ]

> +/* mutex lock for the port refcount */
> +DEFINE_MUTEX(port_mngr_grp_mtx);

[Severity: Low]
This isn't a bug, but port_mngr_grp_mtx, ports_ops[] (in mtk_port_io.c)
and struct port_ops are all global and have no driver prefix. With
CONFIG_MTK_T9XX=y they become kernel-wide symbols. The single mutex is
also shared by every device instance.

Could they get an mtk_ prefix, or be made static where possible?

[ ... ]

> +start_wait:
> +
> +	/* wait trb done, and no timeout in tx blocking mode */
> +	ret = wait_event_interruptible_timeout(port->trb_wq,
> +					       trb->status <= 0 ||
> +					       test_bit(PORT_S_FLUSH, &port->status) ||
> +					       !test_bit(PORT_S_WR, &port->status),
> +					       MTK_DFLT_TRB_TIMEOUT);
> +	if (!ret) {
> +		goto start_wait;
> +	} else if (ret == -ERESTARTSYS) {
> +		ret = -EINTR;
> +	} else if (ret > 0) {
> +		if (test_bit(PORT_S_FLUSH, &port->status))
> +			ret = len;
> +		else
> +			ret = (!trb->status) ? len : trb->status;

[Severity: Low]
Can this return MTK_TRB_HEADER_ADDED as a byte count?

mtk_port_add_header() sets trb->status = MTK_TRB_HEADER_ADDED (0xADDED)
before the skb is submitted. The wait also ends when PORT_S_WR is
cleared, and mtk_port_internal_disable() clears it before the DISABLE
TRB is even queued:

	clear_bit(PORT_S_WR, &port->status);
	wake_up_all(&port->trb_wq);

In that case PORT_S_FLUSH is clear and the TX TRB has not completed yet.
The code then returns trb->status, 712173, as if that many bytes had
been written.

mtk_ctrl_ch_flush() later completes the TX TRB with -EIO. The data is
dropped, but the blocking write has already reported success.

Later in the series this can happen for a blocking WWAN write during
mtk_port_wwan_disable(). Should the !PORT_S_WR case return an error
instead?

[ ... ]

> +int mtk_port_ch_enable(struct mtk_port *port)
> +{

[ ... ]

> +	ret = wait_event_timeout(port->trb_wq, trb->status <= 0,
> +				 MTK_DFLT_TRB_TIMEOUT);
> +	if (!ret)
> +		ret = -ETIMEDOUT;
> +	else
> +		ret = trb->status;

[Severity: Medium]
Can mtk_port_ch_enable() return 0 before mtk_port_open_trb_complete()
has filled in the port geometry?

The transport stores the status before it calls the completion callback.
For example, in mtk_cldma_open():

	trb->status = ret;
	trb->trb_complete(skb);

The MTU and fragment sizes are only written inside the callback:

	if (!trb->status) {
		port->tx_mtu = trb_open_priv->tx_mtu;
		...

wait_event_timeout() checks trb->status <= 0 before it sleeps. A waiter
that checks between those two points, or while the trb_srv thread is
preempted inside the callback, returns 0 with port->tx_mtu still 0.
There is also no acquire/release pairing, so on weakly ordered CPUs the
MTU stores may not be visible yet.

Later in the series, mtk_port_wwan_enable() checks port->tx_mtu right
after this returns:

	if (!port->tx_mtu) {
		dev_err(port_mngr->ctrl_blk->mdev->dev,
			"Invalid tx_mtu for port(%s)\n", port->info.name);

It then disables the channel, so the AT or MBIM port would not be
created for the rest of that device's lifetime.

Would it be better to signal completion from
mtk_port_open_trb_complete() after its last store? A struct completion,
or a separate done flag using smp_store_release()/smp_load_acquire(),
would do that.

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/mtk_port_io.c b/drivers/net/wwan/t9xx/mtk_port_io.c
> new file mode 100644
> index 0000000000000..6cff0704b7fc8
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/mtk_port_io.c
> @@ -0,0 +1,258 @@

[ ... ]

> +static void mtk_port_internal_enable(struct mtk_port *port)
> +{
> +	int ret;
> +
> +	if (test_bit(PORT_S_ENABLE, &port->status))
> +		return;
> +
> +	ret = mtk_port_ch_enable(port);
> +	if (ret && ret != -EBUSY) {

[ ... ]

> +	set_bit(PORT_S_WR, &port->status);
> +	set_bit(PORT_S_ENABLE, &port->status);
> +}

[Severity: Low]
If the ENABLE TRB completes with -EBUSY, does the port end up enabled
with zero MTU and fragment sizes?

Two paths fill trb_open_priv->tx_mtu, rx_mtu, tx_frag_size and
rx_frag_size and then complete with -EBUSY:

  - mtk_ch_status_check(), when usr_cnt is already non-zero
  - mtk_cldma_open(), when the txq/rxq already exist

mtk_port_open_trb_complete() copies those values only when the status
is 0:

	if (!trb->status) {
		port->tx_mtu = trb_open_priv->tx_mtu;

This function treats -EBUSY as success, so it sets PORT_S_WR and
PORT_S_ENABLE. If this port never had a successful ENABLE, tx_mtu,
rx_mtu and the frag sizes stay 0.

The mtk_cldma_open() -EBUSY path also rolls back usr_cnt:

	if (ret)
		cd->trans->usr_cnt[que->hif_id][que->txqno]--;

The port still believes it owns the channel, so its later DISABLE would
also complete with -EBUSY.

Should the geometry also be copied on -EBUSY, or should -EBUSY not
count as a successful enable?

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_pci.c b/drivers/net/wwan/t9xx/pcie/mtk_pci.c
> index 6c31a5d80693d..84b7fb3c82fd0 100644
> --- a/drivers/net/wwan/t9xx/pcie/mtk_pci.c
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_pci.c

[ ... ]

> @@ -1069,7 +1071,28 @@ static struct pci_driver mtk_pci_drv = {
>  	.err_handler = &mtk_pci_err_handler
>  };
>  
> -module_pci_driver(mtk_pci_drv);
> +static int __init mtk_drv_init(void)
> +{
> +	int ret;
> +
> +	ret = mtk_port_io_init();
> +	if (ret)
> +		return ret;
> +
> +	ret = pci_register_driver(&mtk_pci_drv);
> +	if (ret)
> +		mtk_port_io_exit();
> +
> +	return ret;
> +}
> +module_init(mtk_drv_init);

[Severity: Low]
The commit message leaves out two changes to existing code.

First, mtk_drv_init()/mtk_drv_exit() replace module_pci_driver() only to
call mtk_port_io_init() and mtk_port_io_exit(), and both are empty:

	int mtk_port_io_init(void)
	{
		return 0;
	}

They are still empty at the end of the series. Could module_pci_driver()
stay and the stubs be dropped?

Second, mtk_ctrl_trb_handler() in mtk_trans_ctrl.c now takes and drops a
reference around every TRB, including on the -EAGAIN early return:

	kref_get(&trb->kref);
	...
	kref_put(&trb->kref, mtk_port_trb_free);

This makes the transport layer depend on the port layer's release
function. Could the commit message describe this change to TRB
ownership?

[ ... ]

> @@ -0,0 +1,693 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/*
> + * Copyright (c) 2022, MediaTek Inc.
> + */
> +
> +#include <linux/bitfield.h>
> +#include <linux/device.h>
> +#include <linux/err.h>
> +#include <linux/kernel.h>
> +#include <linux/netdevice.h>
> +#include <linux/slab.h>
> +#include <linux/wait.h>
> +
> +#include "mtk_port.h"
> +#include "mtk_port_io.h"
> +#include "mtk_trans_ctrl.h"
> +
> +#define MTK_DFLT_TRB_TIMEOUT		(5 * HZ)
> +#define MTK_DFLT_TRB_STATUS		(0x1)
> +#define MTK_TRB_HEADER_ADDED		(0xADDED)
> +#define MTK_CHECK_RX_SEQ_MASK		(0x7fff)
> +
> +#define MTK_PORT_ENUM_VER		(0)
> +#define MTK_PORT_ENUM_HEAD_PATTERN	(0x5a5a5a5a)
> +#define MTK_PORT_ENUM_TAIL_PATTERN	(0xa5a5a5a5)
> +
> +#define MTK_PORT_SEARCH_FROM_RADIX_TREE(p, s) ({\
> +	struct mtk_port *_p;			\
> +	_p = rcu_dereference_raw(*(s));		\
> +	if (!_p)				\
> +		continue;			\
> +	p = _p;					\
> +})
> +
> +#define MTK_PORT_INTERNAL_NODE_CHECK(p, s, i) ({\
> +	if (radix_tree_is_internal_node(p)) {	\
> +		s = radix_tree_iter_retry(&(i));\
> +		continue;			\
> +	}					\
> +})
> +
> +struct mtk_port_info {
> +	__le16 channel;
> +	__le16 reserved;
> +} __packed;
> +
> +struct mtk_port_enum_msg {
> +	__le32 head_pattern;
> +	__le16 port_cnt;
> +	__le16 version;
> +	__le32 tail_pattern;
> +	u8 data[];
> +} __packed;
> +
> +/* mutex lock for the port refcount */
> +DEFINE_MUTEX(port_mngr_grp_mtx);
> +
> +/* This function working always under mutex lock port_mngr_grp_mtx */
> +void mtk_port_release(struct kref *port_kref)
> +{
> +	struct mtk_port *port;
> +
> +	port = container_of(port_kref, struct mtk_port, kref);
> +	ports_ops[port->info.type]->exit(port);
> +	kfree_rcu(port, rcu);
> +}
> +
> +static int mtk_port_tbl_add(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
> +{
> +	int ret;
> +
> +	mutex_lock(&port_mngr->port_tbl_mtx);
> +	ret = radix_tree_insert(&port_mngr->port_tbl[MTK_PORT_TBL_TYPE(port->info.rx_ch)],
> +				port->info.rx_ch & 0xFFF, port);
> +	if (!ret)
> +		port_mngr->port_cnt++;
> +	mutex_unlock(&port_mngr->port_tbl_mtx);
> +
> +	if (ret)
> +		dev_err(port_mngr->ctrl_blk->mdev->dev,
> +			"port(%s) add to port_tbl failed, return %d\n",
> +			port->info.name, ret);
> +
> +	return ret;
> +}
> +
> +static void mtk_port_tbl_del(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
> +{
> +	mutex_lock(&port_mngr->port_tbl_mtx);
> +	radix_tree_delete(&port_mngr->port_tbl[MTK_PORT_TBL_TYPE(port->info.rx_ch)],
> +			  port->info.rx_ch & 0xFFF);
> +	port_mngr->port_cnt--;
> +	mutex_unlock(&port_mngr->port_tbl_mtx);
> +}
> +
> +static struct mtk_port *mtk_port_alloc_and_add(struct mtk_port_mngr *port_mngr,
> +					       struct mtk_port_cfg *dflt_info)
> +{
> +	struct mtk_port *port;
> +	int ret;
> +
> +	port = kzalloc_obj(*port, GFP_KERNEL);
> +	if (!port) {
> +		ret = -ENOMEM;
> +		goto err_alloc_port;
> +	}
> +	memcpy(&port->info, dflt_info, sizeof(*dflt_info));
> +
> +	ret = mtk_port_tbl_add(port_mngr, port);
> +	if (ret < 0) {
> +		dev_err(port_mngr->ctrl_blk->mdev->dev,
> +			"Failed to add port(%s) to port tbl\n", dflt_info->name);
> +		goto err_free_port;
> +	}
> +
> +	port->port_mngr = port_mngr;
> +	ret = ports_ops[port->info.type]->init(port);
> +	if (ret < 0) {
> +		mtk_port_tbl_del(port_mngr, port);
> +		goto err_free_port;
> +	}
> +
> +	return port;
> +
> +err_free_port:
> +	kfree(port);
> +err_alloc_port:
> +	return ERR_PTR(ret);
> +}
> +
> +static void mtk_port_free(struct mtk_port_mngr *port_mngr, struct mtk_port *port)
> +{
> +	mutex_lock(&port_mngr_grp_mtx);
> +	mtk_port_tbl_del(port_mngr, port);
> +	kref_put(&port->kref, mtk_port_release);
> +	mutex_unlock(&port_mngr_grp_mtx);
> +}
> +
> +static struct mtk_port *mtk_port_search_by_id(struct mtk_port_mngr *port_mngr, int rx_ch)

[Severity: High]
Does anything keep the port returned here alive until
mtk_port_get_locked() takes its reference?

mtk_port_search_by_name() walks port_tbl[] without rcu_read_lock(),
port_tbl_mtx or port_mngr_grp_mtx, and returns a bare pointer.
mtk_port_get_locked() only takes port_mngr_grp_mtx afterwards. Its
!port check tests the pointer value, not whether the port is still in
the table.

On the mtk_ctrl_exit() path, mtk_port_tbl_destroy() calls
mtk_port_free(). That takes port_mngr_grp_mtx, deletes the port from
the radix tree and drops the last reference, so mtk_port_release()
queues the port with kfree_rcu(). The reader is not inside an RCU
read-side critical section, so the grace period does not wait for it.

If the free runs between the lookup and the mutex_lock() in
mtk_port_get_locked(), then kref_get() and mtk_port_common_open() use
freed memory. The tree walk itself can also see nodes that a
concurrent radix_tree_delete() is freeing.

mtk_port_search_by_id() has the same problem. That includes the RX
path in mtk_port_rx_dispatch(), which uses the port without taking any
reference.

Could the lookup and kref_get() be done inside one port_mngr_grp_mtx
section? Another option is rcu_read_lock() with
kref_get_unless_zero(). Also, the rcu_dereference_raw() in
MTK_PORT_SEARCH_FROM_RADIX_TREE hides this from lockdep.

> +{
> +	int tbl_type = MTK_PORT_TBL_TYPE(rx_ch);
> +
> +	if (tbl_type < PORT_TBL_SAP || tbl_type >= PORT_TBL_MAX)
> +		return NULL;
> +
> +	return radix_tree_lookup(&port_mngr->port_tbl[tbl_type], MTK_CH_ID(rx_ch));
> +}
> +
> +struct mtk_port *mtk_port_search_by_name(struct mtk_port_mngr *port_mngr, char *name)
> +{
> +	int tbl_type = PORT_TBL_SAP;
> +	struct radix_tree_iter iter;
> +	struct mtk_port *port;
> +	void __rcu **slot;
> +
> +	do {
> +		radix_tree_for_each_slot(slot, &port_mngr->port_tbl[tbl_type], &iter, 0) {
> +			MTK_PORT_SEARCH_FROM_RADIX_TREE(port, slot);
> +			MTK_PORT_INTERNAL_NODE_CHECK(port, slot, iter);
> +			if (!strncmp(port->info.name, name, MTK_DFLT_PORT_NAME_LEN))
> +				return port;
> +		}
> +		tbl_type++;
> +	} while (tbl_type < PORT_TBL_MAX);
> +
> +	return NULL;
> +}
> +
> +static int mtk_port_tbl_create(struct mtk_port_mngr *port_mngr, struct mtk_port_cfg *cfg,
> +			       const int port_cnt)
> +{
> +	struct mtk_port_cfg *dflt_port;
> +	struct mtk_port *port;
> +	int i;
> +
> +	INIT_RADIX_TREE(&port_mngr->port_tbl[PORT_TBL_SAP], GFP_KERNEL);
> +	INIT_RADIX_TREE(&port_mngr->port_tbl[PORT_TBL_MD], GFP_KERNEL);
> +
> +	/* copy ports from static port cfg table */
> +	for (i = 0; i < port_cnt; i++) {
> +		dflt_port = cfg + i;
> +		if (!mtk_port_search_by_id(port_mngr, dflt_port->rx_ch)) {
> +			port = mtk_port_alloc_and_add(port_mngr, dflt_port);
> +			if (IS_ERR(port))
> +				return PTR_ERR(port);
> +		}
> +	}
> +
> +	return 0;
> +}
> +
> +static void mtk_port_tbl_destroy(struct mtk_port_mngr *port_mngr)
> +{
> +	struct radix_tree_iter iter;
> +	struct mtk_port *port;
> +	void __rcu **slot;
> +	int tbl_type;
> +
> +	tbl_type = PORT_TBL_SAP;
> +	do {
> +		radix_tree_for_each_slot(slot, &port_mngr->port_tbl[tbl_type], &iter, 0) {
> +			MTK_PORT_SEARCH_FROM_RADIX_TREE(port, slot);
> +			MTK_PORT_INTERNAL_NODE_CHECK(port, slot, iter);
> +			ports_ops[port->info.type]->disable(port);
> +		}
> +
> +		while (radix_tree_gang_lookup(&port_mngr->port_tbl[tbl_type],
> +					      (void **)&port, 0, 1))
> +			mtk_port_free(port_mngr, port);
> +	} while (++tbl_type < PORT_TBL_MAX);
> +}
> +
> +void mtk_port_trb_init(struct mtk_port *port, struct trb *trb, enum mtk_trb_cmd_type cmd,
> +		       int (*trb_complete)(struct sk_buff *skb))
> +{
> +	kref_init(&trb->kref);
> +	trb->channel_id = port->info.rx_ch;
> +	trb->status = MTK_DFLT_TRB_STATUS;
> +	trb->priv = port;
> +	trb->cmd = cmd;
> +	trb->trb_complete = trb_complete;
> +}
> +
> +void mtk_port_trb_free(struct kref *trb_kref)
> +{
> +	struct trb *trb = container_of(trb_kref, struct trb, kref);
> +	struct sk_buff *skb, *frag_skb, *next_skb;
> +
> +	skb = container_of((char *)trb, struct sk_buff, cb[0]);
> +	/* Free frag_list for scatter gather TX */
> +	if (trb->cmd == TRB_CMD_TX && skb_has_frag_list(skb)) {
> +		frag_skb = skb_shinfo(skb)->frag_list;
> +		while (frag_skb) {
> +			next_skb = frag_skb->next;
> +			frag_skb->next = NULL;
> +			dev_kfree_skb_any(frag_skb);
> +			frag_skb = next_skb;
> +		}
> +		skb_shinfo(skb)->frag_list = NULL;
> +		skb->data_len = 0;
> +	}
> +	dev_kfree_skb_any(skb);
> +}
> +
> +static int mtk_port_open_trb_complete(struct sk_buff *skb)
> +{
> +	struct trb_open_priv *trb_open_priv = (struct trb_open_priv *)skb->data;
> +	struct trb *trb = (struct trb *)skb->cb;
> +	struct mtk_port *port = trb->priv;
> +
> +	if (!trb->status) {
> +		port->tx_mtu = trb_open_priv->tx_mtu;
> +		port->rx_mtu = trb_open_priv->rx_mtu;
> +		port->tx_frag_size = trb_open_priv->tx_frag_size;
> +		port->rx_frag_size = trb_open_priv->rx_frag_size;
> +		port->tx_mtu -= MTK_CCCI_H_ELEN;
> +		port->rx_mtu -= MTK_CCCI_H_ELEN;
> +	}
> +
> +	wake_up_all(&port->trb_wq);
> +
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +	return 0;
> +}
> +
> +static int mtk_port_close_trb_complete(struct sk_buff *skb)
> +{
> +	struct trb *trb = (struct trb *)skb->cb;
> +	struct mtk_port *port = trb->priv;
> +
> +	wake_up_all(&port->trb_wq);
> +	wake_up_all(&port->rx_wq);
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +
> +	return 0;
> +}
> +
> +static int mtk_port_tx_complete(struct sk_buff *skb)
> +{
> +	struct trb *trb = (struct trb *)skb->cb;
> +	struct mtk_port *port = trb->priv;
> +
> +	if (trb->status < 0)
> +		dev_warn(port->port_mngr->ctrl_blk->mdev->dev,
> +			 "Failed to send data: status:%d, port:%s\n",
> +			 trb->status, port->info.name);
> +
> +	wake_up_all(&port->trb_wq);
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +
> +	return 0;
> +}
> +
> +int mtk_port_status_check(struct mtk_port *port)
> +{
> +	if (!test_bit(PORT_S_ENABLE, &port->status))
> +		return -ENODEV;
> +
> +	if (!test_bit(PORT_S_OPEN, &port->status) || test_bit(PORT_S_FLUSH, &port->status) ||
> +	    !test_bit(PORT_S_WR, &port->status))
> +		return -EBADF;
> +
> +	return 0;
> +}
> +
> +int mtk_port_send_data(struct mtk_port *port, void *data, bool blocking, bool force_send)
> +{
> +	struct mtk_port_mngr *port_mngr;
> +	struct sk_buff *skb = data;
> +	struct trb *trb;
> +	int ret, len;
> +
> +	port_mngr = port->port_mngr;
> +
> +	trb = (struct trb *)skb->cb;
> +	mtk_port_trb_init(port, trb, TRB_CMD_TX, mtk_port_tx_complete);
> +	len = skb->len;
> +	kref_get(&trb->kref); /* kref count 1->2 */
> +
> +	/* add ccci header */
> +	mtk_port_add_header(skb);
> +	ret = mtk_port_status_check(port);
> +	if (!ret)
> +		ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb,
> +					      force_send);
> +
> +	if (ret < 0) {
> +		kref_put(&trb->kref, mtk_port_trb_free); /* kref count 2->1 */
> +		kref_put(&trb->kref, mtk_port_trb_free); /* kref count 1->0 */
> +		port->tx_seq--;
> +		goto out;
> +	}
> +
> +	if (!blocking) {
> +		kref_put(&trb->kref, mtk_port_trb_free);
> +		ret = len;
> +		goto out;
> +	}
> +start_wait:
> +
> +	/* wait trb done, and no timeout in tx blocking mode */
> +	ret = wait_event_interruptible_timeout(port->trb_wq,
> +					       trb->status <= 0 ||
> +					       test_bit(PORT_S_FLUSH, &port->status) ||
> +					       !test_bit(PORT_S_WR, &port->status),
> +					       MTK_DFLT_TRB_TIMEOUT);
> +	if (!ret) {
> +		goto start_wait;
> +	} else if (ret == -ERESTARTSYS) {
> +		ret = -EINTR;
> +	} else if (ret > 0) {
> +		if (test_bit(PORT_S_FLUSH, &port->status))
> +			ret = len;
> +		else
> +			ret = (!trb->status) ? len : trb->status;
> +	}
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +
> +out:
> +	return ret;
> +}
> +
> +static int mtk_port_check_rx_seq(struct mtk_port *port, struct mtk_ccci_header *ccci_h)
> +{
> +	u16 seq_num, assert_bit, channel;
> +	struct mtk_md_dev *mdev;
> +
> +	seq_num = FIELD_GET(MTK_HDR_FLD_SEQ, le32_to_cpu(ccci_h->status));
> +	assert_bit = FIELD_GET(MTK_HDR_FLD_AST, le32_to_cpu(ccci_h->status));
> +	if (assert_bit && port->rx_seq &&
> +	    ((seq_num - port->rx_seq) & MTK_CHECK_RX_SEQ_MASK) != 1) {
> +		mdev = port->port_mngr->ctrl_blk->mdev;
> +		channel = FIELD_GET(MTK_HDR_FLD_CHN, le32_to_cpu(ccci_h->status));
> +		dev_warn(mdev->dev,
> +			 "<ch: %04x> seq num out-of-order %d->%d, len(%u)\n",
> +			 channel, seq_num, port->rx_seq,
> +			 le32_to_cpu(ccci_h->packet_len));
> +
> +		port->rx_seq = seq_num;
> +		return -EPROTO;
> +	}
> +
> +	return 0;
> +}
> +
> +static int mtk_port_rx_dispatch_frag_skb(struct mtk_port *port, struct sk_buff *skb)
> +{
> +	struct sk_buff *frag_skb, *frag_next;
> +	int ret;
> +
> +	frag_skb = skb_shinfo(skb)->frag_list;
> +	skb->len -= skb->data_len;
> +	skb->data_len = 0;
> +	skb_shinfo(skb)->frag_list = NULL;
> +
> +	ret = ports_ops[port->info.type]->recv(port, skb);
> +	if (ret < 0) {
> +		skb_shinfo(skb)->frag_list = frag_skb;
> +		return ret;
> +	}
> +
> +	while (frag_skb) {
> +		frag_next = frag_skb->next;
> +		if (!frag_skb->len) {
> +			frag_skb->next = NULL;
> +			dev_kfree_skb_any(frag_skb);
> +			frag_skb = frag_next;
> +			continue;
> +		}
> +		frag_skb->next = NULL;
> +		ret = ports_ops[port->info.type]->recv(port, frag_skb);
> +		if (ret < 0) {
> +			frag_skb->next = frag_next;
> +			while (frag_skb) {
> +				frag_next = frag_skb->next;
> +				frag_skb->next = NULL;
> +				dev_kfree_skb_any(frag_skb);
> +				frag_skb = frag_next;
> +			}
> +			return -EIO;
> +		}
> +		frag_skb = frag_next;
> +	}
> +
> +	return 0;
> +}
> +
> +static int mtk_port_rx_dispatch(struct sk_buff *skb, void *priv, bool force_recv)
> +{
> +	struct mtk_port_mngr *port_mngr;
> +	struct mtk_ccci_header *ccci_h;
> +	struct mtk_port *port = priv;
> +	int ret = -EPROTO;
> +	u16 channel;
> +
> +	if (!skb || !priv) {
> +		pr_err("Invalid input value in rx dispatch\n");
> +		return -EINVAL;
> +	}
> +
> +	port_mngr = port->port_mngr;
> +
> +	ccci_h = mtk_port_strip_header(skb);
> +	if (unlikely(!ccci_h)) {
> +		dev_warn(port_mngr->ctrl_blk->mdev->dev,
> +			 "Unsupported: skb length(%d) is less than ccci header\n",
> +			 skb->len);
> +		goto drop_data;
> +	}
> +
> +	channel = FIELD_GET(MTK_HDR_FLD_CHN, le32_to_cpu(ccci_h->status));
> +	port = mtk_port_search_by_id(port_mngr, channel);
> +	if (unlikely(!port)) {
> +		dev_warn(port_mngr->ctrl_blk->mdev->dev,
> +			 "Failed to find port by channel:%d\n", channel);
> +		goto drop_data;
> +	}
> +
> +	ret = mtk_port_check_rx_seq(port, ccci_h);
> +	if (unlikely(ret))
> +		goto drop_data;
> +
> +	port->rx_seq = FIELD_GET(MTK_HDR_FLD_SEQ, le32_to_cpu(ccci_h->status));
> +	skb_pull(skb, sizeof(*ccci_h));
> +
> +	/* Support scatter gather transmission */
> +	if (port->rx_mtu > port->rx_frag_size) {
> +		ret = mtk_port_rx_dispatch_frag_skb(port, skb);
> +		/* -EIO means partial data dispatch complete, does not goto drop flow */
> +		if (ret < 0 && ret != -EIO)
> +			goto drop_frag_skb;
> +	} else {
> +		ret = ports_ops[port->info.type]->recv(port, skb);
> +		if (ret < 0)
> +			goto drop_data;
> +	}
> +
> +	return ret;
> +
> +drop_frag_skb:
> +	{
> +		struct sk_buff *frag_skb, *tmp;
> +
> +		frag_skb = skb_shinfo(skb)->frag_list;
> +		while (frag_skb) {
> +			tmp = frag_skb->next;
> +			frag_skb->next = NULL;
> +			dev_kfree_skb_any(frag_skb);
> +			frag_skb = tmp;
> +		}
> +		skb_shinfo(skb)->frag_list = NULL;
> +	}
> +drop_data:
> +	dev_kfree_skb_any(skb);
> +	return ret;
> +}
> +
> +int mtk_port_add_header(struct sk_buff *skb)
> +{
> +	struct mtk_ccci_header *ccci_h;
> +	struct mtk_port *port;
> +	struct trb *trb;
> +
> +	trb = (struct trb *)skb->cb;
> +	if (trb->status == MTK_TRB_HEADER_ADDED)
> +		return 0;
> +
> +	port = trb->priv;
> +	if (!port)
> +		return -EINVAL;
> +
> +	ccci_h = skb_push(skb, sizeof(*ccci_h));
> +
> +	ccci_h->packet_header = cpu_to_le32(0);
> +	ccci_h->packet_len = cpu_to_le32(skb->len);
> +	ccci_h->ex_msg = cpu_to_le32(0);
> +	ccci_h->status = cpu_to_le32(FIELD_PREP(MTK_HDR_FLD_CHN, port->info.tx_ch) |
> +				     FIELD_PREP(MTK_HDR_FLD_SEQ, port->tx_seq++) |
> +				     FIELD_PREP(MTK_HDR_FLD_AST, 1));
> +
> +	trb->status = MTK_TRB_HEADER_ADDED;
> +
> +	return 0;
> +}
> +
> +struct mtk_ccci_header *mtk_port_strip_header(struct sk_buff *skb)
> +{
> +	struct mtk_ccci_header *ccci_h;
> +
> +	if (skb->len < sizeof(*ccci_h)) {
> +		pr_err("Invalid input value\n");
> +		return NULL;
> +	}
> +
> +	ccci_h = (struct mtk_ccci_header *)skb->data;
> +
> +	return ccci_h;
> +}
> +
> +int mtk_port_status_update(struct mtk_md_dev *mdev, void *data, u32 data_len)
> +{
> +	struct mtk_port_enum_msg *msg = data;
> +	struct mtk_port_info *port_info;
> +	struct mtk_port_mngr *port_mngr;
> +	struct mtk_ctrl_blk *ctrl_blk;
> +	struct mtk_port *port;
> +	int port_id;
> +	u16 ch_id;
> +
> +	if (unlikely(!mdev || !msg))
> +		return -EINVAL;
> +
> +	ctrl_blk = mdev->ctrl_blk;
> +	port_mngr = ctrl_blk->port_mngr;
> +	if (le16_to_cpu(msg->version) != MTK_PORT_ENUM_VER ||
> +	    le32_to_cpu(msg->head_pattern) != MTK_PORT_ENUM_HEAD_PATTERN ||
> +	    le32_to_cpu(msg->tail_pattern) != MTK_PORT_ENUM_TAIL_PATTERN)
> +		return -EPROTO;
> +
> +	if (data_len < sizeof(*msg) +
> +	    le16_to_cpu(msg->port_cnt) * sizeof(*port_info))
> +		return -EPROTO;
> +
> +	for (port_id = 0; port_id < le16_to_cpu(msg->port_cnt); port_id++) {
> +		port_info = (struct mtk_port_info *)(msg->data +
> +						   (sizeof(*port_info) * port_id));
> +		ch_id = FIELD_GET(MTK_INFO_FLD_CHID, le16_to_cpu(port_info->channel));
> +		port = mtk_port_search_by_id(port_mngr, ch_id);
> +		if (!port)
> +			continue;
> +		port->enable = FIELD_GET(MTK_INFO_FLD_EN, le16_to_cpu(port_info->channel));
> +	}
> +
> +	return 0;
> +}
> +
> +int mtk_port_ch_enable(struct mtk_port *port)
> +{
> +	struct mtk_port_mngr *port_mngr = port->port_mngr;
> +	struct trb_open_priv *trb_open_priv;
> +	struct sk_buff *skb;
> +	struct trb *trb;
> +	int ret;
> +
> +	skb = __dev_alloc_skb(Q_MTU_3_5K, GFP_KERNEL);
> +	if (!skb)
> +		return -ENOMEM;
> +
> +	trb_open_priv = (struct trb_open_priv *)skb->data;
> +	trb_open_priv->rx_done = mtk_port_rx_dispatch;
> +
> +	skb_put(skb, sizeof(struct trb_open_priv));
> +	trb = (struct trb *)skb->cb;
> +	mtk_port_trb_init(port, trb, TRB_CMD_ENABLE, mtk_port_open_trb_complete);
> +	kref_get(&trb->kref);
> +
> +	ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb, true);
> +	if (ret) {
> +		dev_err(port_mngr->ctrl_blk->mdev->dev,
> +			"Failed to submit trb for port(%s), ret=%d\n",
> +			port->info.name, ret);
> +		kref_put(&trb->kref, mtk_port_trb_free);
> +		kref_put(&trb->kref, mtk_port_trb_free);
> +		return ret;
> +	}
> +
> +	ret = wait_event_timeout(port->trb_wq, trb->status <= 0,
> +				 MTK_DFLT_TRB_TIMEOUT);
> +	if (!ret)
> +		ret = -ETIMEDOUT;
> +	else
> +		ret = trb->status;
> +
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +
> +	return ret;
> +}
> +
> +int mtk_port_ch_disable(struct mtk_port *port)
> +{
> +	struct mtk_port_mngr *port_mngr = port->port_mngr;
> +	struct sk_buff *skb;
> +	struct trb *trb;
> +	int ret;
> +
> +	skb = __dev_alloc_skb(Q_MTU_3_5K, GFP_KERNEL);
> +	if (!skb)
> +		return -ENOMEM;
> +
> +	trb = (struct trb *)skb->cb;
> +	mtk_port_trb_init(port, trb, TRB_CMD_DISABLE, mtk_port_close_trb_complete);
> +	kref_get(&trb->kref);
> +
> +	ret = mtk_pcie_hif_submit_skb(port_mngr->ctrl_blk->mdev, skb, true);
> +	if (ret) {
> +		dev_warn(port_mngr->ctrl_blk->mdev->dev,
> +			 "Failed to submit trb for port(%s), ret=%d\n",
> +			 port->info.name, ret);
> +		kref_put(&trb->kref, mtk_port_trb_free);
> +		kref_put(&trb->kref, mtk_port_trb_free);
> +		return ret;
> +	}
> +
> +	ret = wait_event_timeout(port->trb_wq, trb->status <= 0,
> +				 MTK_DFLT_TRB_TIMEOUT);
> +	if (!ret)
> +		ret = -ETIMEDOUT;
> +	else
> +		ret = trb->status;
> +
> +	kref_put(&trb->kref, mtk_port_trb_free);
> +
> +	return ret;
> +}
> +
> +int mtk_port_mngr_init(struct mtk_ctrl_blk *ctrl_blk, struct mtk_port_cfg *port_cfg, int port_cnt)
> +{
> +	struct mtk_port_mngr *port_mngr;
> +	int ret = -ENOMEM;
> +
> +	port_mngr = devm_kzalloc(ctrl_blk->mdev->dev, sizeof(*port_mngr), GFP_KERNEL);
> +	if (unlikely(!port_mngr)) {
> +		dev_err((ctrl_blk->mdev)->dev, "Failed to alloc memory for port_mngr\n");
> +		goto err_out;
> +	}
> +
> +	port_mngr->ctrl_blk = ctrl_blk;
> +	mutex_init(&port_mngr->port_tbl_mtx);
> +
> +	ret = mtk_port_tbl_create(port_mngr, port_cfg, port_cnt);
> +	if (unlikely(ret)) {
> +		dev_err((ctrl_blk->mdev)->dev, "Failed to create port_tbl\n");
> +		goto err_free_port_mngr;
> +	}
> +
> +	ctrl_blk->port_mngr = port_mngr;
> +
> +	return ret;
> +
> +err_free_port_mngr:
> +	mtk_port_tbl_destroy(port_mngr);
> +err_out:
> +	return ret;
> +}
> +
> +void mtk_port_mngr_exit(struct mtk_ctrl_blk *ctrl_blk)
> +{
> +	struct mtk_port_mngr *port_mngr = ctrl_blk->port_mngr;
> +
> +	mtk_port_tbl_destroy(port_mngr);
> +
> +	ctrl_blk->port_mngr = NULL;
> +}

[Severity: High]
Can mtk_port_add_header() hit skb_under_panic() here?

The skb comes straight from the caller of mtk_port_internal_write().
Nothing on the way to this skb_push() checks skb_headroom(skb) against
sizeof(struct mtk_ccci_header).

An internal user might allocate its skb with alloc_skb() and no
skb_reserve(), or pass in an skb whose headroom is already used up.
In that case skb_push() calls BUG() instead of failing the write.

If the skb is cloned, the header is also written into data that is
shared with the clone.

Should this call skb_cow_head(skb, sizeof(*ccci_h)) and return an
error on failure? If so, mtk_port_send_data() would need to stop
ignoring the return value of mtk_port_add_header(). The port->tx_seq--
in its error path would also have to be skipped when no sequence
number was consumed.

Or, if the API requires callers to reserve the headroom, could that be
documented and checked here?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread
  2026-09-30  7:46 ` [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found 15 potential
issue(s) to consider.

Critical: 0 · High: 4 · Medium: 7 · Low: 4

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] Recovery can unmap TX buffers the device still owns.
- [High] Use-after-free of rxq in mtk_cldma_rxq_free().
- [High] Use-after-free on trb_srv during HIF teardown.
- [High] The RX DMA engine can be restarted after mtk_cldma_rxq_free()
  has stopped it, and the buffers are then freed under live DMA.
- [Medium] An RX error can be dropped for good in mtk_cldma_err_work().
- [Medium] mtk_mhccif_isr_work() acks MHCCIF_EP2RC_SW_INT_ACK before it
  dispatches callbacks, and it ignores their return values.
- [Medium] The IRQ teardown in mtk_cldma_dev_exit() relies on 'mask first
  so no new interrupt can fire, then wait out any in-flight handler'.
- [Medium] In mtk_cldma_txq_free() and mtk_cldma_rxq_free(),
  synchronize_irq() and flush_work(&drv_info->err_work) run before the…
- [Medium] mtk_cldma_err_work() calls mtk_cldma_stop_queue(DIR_RX) with
  no serialization against rxq->rx_done_work.
- [Medium] An incompatible OPTIONAL_SUPPORT reply satisfies the handshake
  even though the host marked the feature mandatory.
- [Medium] HS1 masking race.
- [Low] The justification for leaking the DMA pools contradicts itself.
- [Low] The new comment in mtk_cldma_exit() says 'the caller holds
  trans->submit_lock'.
- [Low] In mtk_cldma_dev_exit(), WRITE_ONCE(cd->cldma_drv_info[hif_id],
  NULL) carries the comment 'Unpublish before teardown; pairs with the…
- [Low] In mtk_fsm_feature_set_match(), the
  RTFT_TYPE_SUPPORT_BACKWARD_COMPAT case only compares versions and…

--

net: wwan: t9xx: Add FSM thread

This adds an FSM kthread. It serializes device and OS events into state
transitions and notifies registered listeners before and after each one.
The MD and SAP HS1/HS2/HS3 runtime feature handshake is wired into the
FSM, and CLDMA bring-up and teardown are driven from FSM state
transitions. That includes the IRQ registration and a new QUEUE_ERROR
recovery worker.

> diff --git a/drivers/net/wwan/t9xx/mtk_fsm.c b/drivers/net/wwan/t9xx/mtk_fsm.c
> new file mode 100644
> index 0000000000000..94617933e7713
> --- /dev/null
> +++ b/drivers/net/wwan/t9xx/mtk_fsm.c
> @@ -0,0 +1,1174 @@

[ ... ]

> +static int mtk_fsm_feature_set_match(enum runtime_feature_support_type *cur_ft_spt,
> +				     struct runtime_feature_info rtft_info_st,
> +				     struct runtime_feature_info rtft_info_cfg)
> +{

[ ... ]

> +	case RTFT_TYPE_OPTIONAL_SUPPORT:
> +		if (FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_EXIST ||
> +		    FIELD_GET(FEATURE_TYPE, rtft_info_cfg.feature) == RTFT_TYPE_NOT_SUPPORT) {
> +			*cur_ft_spt = RTFT_TYPE_NOT_SUPPORT;
> +		} else {
> +			if (FIELD_GET(FEATURE_VER, rtft_info_st.feature) ==
> +			    FIELD_GET(FEATURE_VER, rtft_info_cfg.feature))
> +				*cur_ft_spt = RTFT_TYPE_MUST_SUPPORT;
> +			else
> +				*cur_ft_spt = RTFT_TYPE_NOT_SUPPORT;
> +		}
> +		break;

[Severity: Medium]
Should a version mismatch here fail the handshake when the host has
marked the feature mandatory?

mtk_fsm_hs_info_init_by_hsid() configures MD_PORT_ENUM and SAP_PORT_ENUM
as RTFT_TYPE_MUST_SUPPORT, version 0. Suppose the device replies
OPTIONAL_SUPPORT with a different version. This code sets cur_ft_spt to
RTFT_TYPE_NOT_SUPPORT and returns 0. BACKWARD_COMPAT with a lower version
gives NOT_EXIST in the same way.

mtk_fsm_parse_hs2_msg() then skips the port enumeration action:

    if (cur_ft_spt == RTFT_TYPE_MUST_SUPPORT && query_rtft_action[ft_id]) {

HS3 still goes out, and the FSM can reach READY without the mandatory
enumeration ever running. That seems to contradict the comment in the
NOT_SUPPORT case above: "The device refusing a feature the host declared
mandatory must fail the handshake".

> +	case RTFT_TYPE_SUPPORT_BACKWARD_COMPAT:
> +		if (FIELD_GET(FEATURE_VER, rtft_info_st.feature) >=
> +		    FIELD_GET(FEATURE_VER, rtft_info_cfg.feature))
> +			*cur_ft_spt = RTFT_TYPE_MUST_SUPPORT;
> +		else
> +			*cur_ft_spt = RTFT_TYPE_NOT_EXIST;
> +		break;

[Severity: Low]
This case only compares versions and never looks at the host's
configured type. For a feature the host never offered, query_ft_set[] is
zero (NOT_EXIST, version 0). The >= check then always passes and
cur_ft_spt becomes MUST_SUPPORT.

mtk_fsm_parse_hs2_msg() will then run query_rtft_action[ft_id] on data
supplied by the device. For example, during the MD handshake the device
could trigger the SAP_PORT_ENUM action (mtk_port_status_update()).

Should this case reject or downgrade when the host config is NOT_EXIST
or NOT_SUPPORT, as the MUST_SUPPORT and OPTIONAL_SUPPORT cases do?

[ ... ]

> +static int mtk_fsm_idle_evt_handler(struct mtk_md_dev *mdev,
> +				    u32 dev_state, struct mtk_md_fsm *fsm)
> +{

[ ... ]

> +	/* On failure keep the handshake channels masked and report it, so
> +	 * the device's next boot-flow notification can retrigger us.
> +	 */
> +	if (mtk_fsm_evt_submit(mdev, FSM_EVT_STARTUP, FSM_F_DFLT,
> +			       NULL, 0, 0) == FSM_EVT_RET_FAIL) {
> +		dev_err(mdev->dev, "Failed to submit STARTUP evt, waiting for retry\n");
> +		return -ENOMEM;
> +	}

[Severity: Medium]
Can this retry ever arrive?

mtk_mhccif_isr_work() acks MHCCIF_EP2RC_SW_INT_ACK before it dispatches
the callbacks, and it ignores their return values. mtk_mhccif_init()
documents EP2RC notifications as one-shot: the device never re-sends
them.

The callback runs under spin_lock_bh, so mtk_fsm_evt_submit() allocates
with GFP_ATOMIC here and can fail under memory pressure. When it fails,
this function returns before it unmasks the HS channels. In addition,
mtk_fsm_early_bootup_handler() leaves last_dev_state unlatched so that
the repeated notification is not filtered out.

mtk_fsm_hs1_handler() relies on the same assumption:

    /* Only consume the notification once the event is queued; on a
     * failed submit the channel stays unmasked and uncleared so the
     * device's retry still reaches us.
     */

The status has already been acked, and DEV_STAGE_IDLE is the last stage,
so no second notification seems to come. The FSM would then stay in ON
with the HS channels masked, and the device would never reach READY.

The callbacks also ack a second time through mtk_pci_clear_ext_evt().
That contradicts the dispatcher's ack-first ordering, and it can erase
an event that was re-asserted in the meantime.

[ ... ]

> +static int mtk_fsm_hs1_handler(u32 status, void *__hs_info)
> +{

[ ... ]

> +	if (mtk_fsm_evt_submit(mdev, FSM_EVT_STARTUP, hs_info->fsm_flag_hs1,
> +			       hs_info, sizeof(*hs_info), 0) == FSM_EVT_RET_FAIL) {
> +		dev_err(mdev->dev, "Failed to submit HS1 evt(hs%d), waiting for retry\n",
> +			hs_info->id);
> +		return -ENOMEM;
> +	}
> +	mtk_pci_mask_ext_evt(mdev, hs_info->mhccif_ch);
> +	mtk_pci_clear_ext_evt(mdev, hs_info->mhccif_ch);

[Severity: Medium]
Can the mask here undo the FSM thread's re-arm?

mtk_fsm_evt_submit() queues the event and wakes the FSM thread before
this callback masks and clears the channel. Meanwhile the FSM thread on
another CPU can process the STARTUP event and fail. Examples are the
-EPROTO state check, a mtk_fsm_ctrl_ch_start() failure, or a failed HS1
send. On failure it unmasks the HS channels in mtk_fsm_startup_act() so
that a retry can come in:

    hs_err:
        for (int hs_id = 0; hs_id < HS_ID_MAX; hs_id++)
            mtk_pci_unmask_ext_evt(mdev, fsm->hs_info[hs_id].mhccif_ch);

This callback then masks the channel again, and later HS notifications
stay masked.

The FSM side does not take mhccif_lock, so the lock held here does not
order the two paths. On PREEMPT_RT the spin_lock_bh section can also be
preempted, which makes the window wider.

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
> index 0f281bfb38cc7..080ae29d3a887 100644
> --- a/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_cldma.c
> @@ -32,11 +32,184 @@

[ ... ]

> +static int mtk_cldma_isr(int irq_id, void *param)
> +{

[ ... ]

> +	mtk_pci_clear_irq(mdev, drv_info->pci_ext_irq_id);
> +	mtk_pci_unmask_irq(mdev, drv_info->pci_ext_irq_id);
> +
> +	return IRQ_HANDLED;
> +}

[Severity: Medium]
Does this unconditional unmask defeat the mask in mtk_cldma_dev_exit()?

mtk_cldma_dev_exit() relies on this sequence:

    mtk_pci_mask_irq(mdev, drv_info->pci_ext_irq_id);
    synchronize_irq(virq_id);
    mtk_pci_unregister_irq(mdev, drv_info->pci_ext_irq_id);

Its comment says "mask first so no new interrupt can fire". A handler
that is already running would unmask L1 again here before
synchronize_irq() returns.

mtk_pci_unregister_irq() uses the same mask plus synchronize_irq()
sequence and then clears irq_cb_list[] and irq_cb_data[].
mtk_pci_irq_handler() reads cb and data as two separate loads, without
irq_cb_lock:

    cb = READ_ONCE(priv->irq_cb_list[irq_id]);
    if (likely(cb)) {
        smp_rmb();
        cb(irq_id, priv->irq_cb_data[irq_id]);

A new interrupt could then call mtk_cldma_isr() with NULL data. It could
also run against a drv_info whose workqueue is being destroyed and which
is freed later.

In the queue-timeout path, mtk_cldma_rearm_queues() and
mtk_cldma_drv_init() can also rewrite int_mask and unmask L2 sources
after the handler has been removed.

mtk_pci_irq_msix() masks L1 before dispatch, so an interrupt that finds
no handler leaves L1 masked rather than armed indefinitely.

[ ... ]

> @@ -679,8 +853,11 @@ static void mtk_cldma_txq_free(struct cldma_drv_info *drv_info, u32 txqno)
>  
>  	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
>  	synchronize_irq(irq_id);
> -	/* flush on-going work */
> +	/* flush on-going work; the error worker may have loaded this txq
> +	 * before it was unpublished above, so it has to be retired too
> +	 */
>  	flush_work(&txq->tx_done_work);
> +	flush_work(&drv_info->err_work);
>  	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_XFER_DONE);
>  	mtk_cldma_mask_intr(drv_info, DIR_TX, txqno, QUEUE_ERROR);

[Severity: Medium]
Is err_work flushed too early here? synchronize_irq() and
flush_work(&drv_info->err_work) both run before QUEUE_ERROR is masked
for this queue.

If a QUEUE_ERROR interrupt lands in between, mtk_cldma_isr() sets the
tx_err_qs bit, unmasks QUEUE_ERROR again and queues err_work after the
flush has finished.

A later OPEN could reallocate the same queue number before the stale
err_work runs. err_work would then stop the new queue and complete its
pending TX requests with -EPIPE.

mtk_cldma_rxq_free() uses the same ordering. On RX, a stale bit would
stop and restart a healthy new queue, which is the case the new comment
in rxq_free says the flush prevents.

> @@ -987,8 +1165,11 @@ static void mtk_cldma_rxq_free(struct cldma_drv_info *drv_info, u32 rxqno)
>  
>  	irq_id = mtk_pci_get_virq_id(mdev, drv_info->pci_ext_irq_id);
>  	synchronize_irq(irq_id);
> -	/* flush on-going work */
> +	/* flush on-going work; the error worker may still be about to stop
> +	 * this queue number, which a later allocation could already reuse
> +	 */
>  	flush_work(&rxq->rx_done_work);
> +	flush_work(&drv_info->err_work);

[Severity: High]
Is this flush ordering enough to keep rxq alive?

mtk_cldma_err_work() can load rxq with
smp_load_acquire(&drv_info->rxq[i]) before rxq_free() unpublishes the
slot. It then spends time in mtk_cldma_stop_queue(DIR_RX), and finally
does:

    atomic_set(&rxq->need_restart, 1);
    queue_work(drv_info->wq, &rxq->rx_done_work);

If that queue_work() runs after flush_work(&rxq->rx_done_work) has
returned, flush_work(&drv_info->err_work) waits only for err_work
itself. rxq_free() then goes on to kfree(rxq) while rx_done_work is
still pending, and mtk_cldma_rx_done_work() later runs on freed memory.

need_exit does not help, because the work_struct itself is inside the
freed rxq.

Should err_work be flushed before rx_done_work? This path can be reached
from both mtk_cldma_close() and mtk_cldma_dev_exit().

[ ... ]

> @@ -1072,6 +1254,67 @@ static int mtk_cldma_hw_recovery(struct cldma_drv_info *drv_info, u32 qno)

[ ... ]

> +static int mtk_cldma_dev_exit(struct cldma_dev *cd, int hif_id)
> +{

[ ... ]

> +	/* quiesce the IP before releasing descriptor memory: disable its
> +	 * interrupt output and reset it, so it cannot touch the rings again
> +	 */
> +	mtk_pci_write32(mdev, drv_info->base_addr + drv_info->hw_regs->reg_cldma_int_mask,
> +			LINK_ERROR_VAL);
> +	mtk_cldma_drv_reset(drv_info);
> +
> +	if (drv_info->ring_leaked) {
> +		/* A ring is leaked and its descriptors live in these pools:
> +		 * the device may still master DMA into them, so handing them
> +		 * back to the allocator would open a use-after-free window.
> +		 */
> +		dev_err(mdev->dev, "CLDMA%d rings leaked, leaking DMA pools too\n",
> +			drv_info->hw_id);
> +	} else {
> +		dma_pool_destroy(drv_info->bd_dma_pool);
> +		dma_pool_destroy(drv_info->gpd_dma_pool);
> +	}

[Severity: Low]
Which of these two comments is correct? The reset above is said to
quiesce the IP "so it cannot touch the rings again". This branch then
skips dma_pool_destroy() because "the device may still master DMA into
them".

ring_leaked is only set after txq_free() or rxq_free() has already reset
the instance and the stop still timed out. So the code already assumes
that a reset is not enough in that case.

Either the pool leak is unnecessary, or the "cannot touch the rings
again" claim is wrong. The leak policy in the commit message does not
say which guarantee it relies on.

> +
> +	/* Unpublish before teardown; pairs with the acquire loads of this
> +	 * slot. No release is needed, there is no prior store to expose.
> +	 */
> +	WRITE_ONCE(cd->cldma_drv_info[hif_id], NULL);
> +	kfree(drv_info);

[Severity: Low]
Does this comment match the code? It says "Unpublish before teardown".
The store actually runs after the IRQ unregister, the queue frees,
destroy_workqueue(), the reset and dma_pool_destroy().

The comment looks copied from txq_free() and rxq_free(), where the order
really is the other way round. The protection today comes from
mtk_cldma_exit() clearing trans->dev first and from the trb threads
being joined.

[ ... ]

> @@ -1095,6 +1338,11 @@ static int mtk_cldma_start_xfer(struct cldma_drv_info *drv_info, u32 qno)
>  	 * names; tx_done_work advances free_idx under the same lock.
>  	 */
>  	spin_lock(&txq->ring_lock);
> +	if (unlikely(txq->is_stopping)) {
> +		spin_unlock(&txq->ring_lock);
> +		return -EPIPE;
> +	}

[Severity: High]
Can this -EPIPE return lead to unmapping TX buffers that the device
still owns?

mtk_cldma_tx() treats any error from mtk_cldma_start_xfer() as a reason
to flush the ring:

    ret = mtk_cldma_start_xfer(drv_info, que->txqno);
    if (unlikely(ret)) {
        dev_err(mdev->dev, "Failed to trigger cldma tx\n");
        mtk_cldma_txq_flush(drv_info, txq, ret);
    }

mtk_cldma_txq_flush() clears CLDMA_GPD_FLAG_HWO, calls
dma_unmap_single() and completes every pending skb.

mtk_cldma_err_work() sets is_stopping before it polls
mtk_cldma_stop_queue(). If the stop fails, it leaves is_stopping set on
purpose and skips the flush, because the device may still be walking the
ring. A TRB_CMD_TX that arrives during the stop poll, or any time after
a failed stop, would flush exactly the requests err_work is trying to
keep.

mtk_cldma_submit_tx() also never checks is_stopping, although the new
comment in struct txq says the producer reads it under ring_lock. So new
HWO descriptors keep being published into the ring.

Nothing clears is_stopping after a failed stop. Wouldn't every later TX
on that queue take the flush path?

[ ... ]

> @@ -1124,11 +1372,22 @@ int mtk_cldma_init(struct mtk_ctrl_trans *trans)

[ ... ]

> +	/* Latch-and-clear up front: the caller holds trans->submit_lock, so
> +	 * publishing the NULL here makes any later submit path bail out in
> +	 * mtk_cldma_get_tx_budget() instead of walking freed queues.
> +	 */
>  	trans->dev = NULL;

[Severity: Low]
Is the locking precondition in this comment accurate? The err_cldma_exit
path in mtk_pcie_hif_init() is taken when mtk_ctrl_trb_srv_init() fails,
and it calls mtk_cldma_exit() without holding submit_lock.

That is harmless today because trans->available is still 0 on that
path. Should the comment cover this case as well?

[ ... ]

> @@ -1250,6 +1510,82 @@ static void mtk_cldma_txq_flush(struct cldma_drv_info *drv_info,

[ ... ]

> +	tx_err = atomic_xchg(&drv_info->tx_err_qs, 0);
> +	rx_err = atomic_xchg(&drv_info->rx_err_qs, 0);
> +
> +	for (i = 0; i < HW_QUEUE_NUM; i++) {
> +		if (tx_err & BIT(i)) {
> +			/* pairs with smp_store_release() in txq_alloc */
> +			txq = smp_load_acquire(&drv_info->txq[i]);
> +			if (!txq)
> +				continue;

[ ... ]

> +			ret = mtk_cldma_stop_queue(drv_info, DIR_TX, i);
> +			if (ret) {
> +				/* the device may still be walking the ring:
> +				 * unmapping its buffers here would leave it
> +				 * writing into unmapped memory. is_stopping is
> +				 * left set so nothing submits to it again.
> +				 */
> +				dev_err(drv_info->mdev->dev,
> +					"TX queue %d stop failed (%d), keeping its requests\n",
> +					i, ret);
> +				continue;
> +			}

[Severity: Medium]
Does this continue, and the one after the !txq check above, also skip
RX error handling for the same queue index?

Both masks are consumed up front with atomic_xchg(). When the TX branch
hits either continue, the "if (rx_err & BIT(i))" branch for the same i
never runs.

The ISR has already cleared and unmasked the RX QUEUE_ERROR source. As
the comment in the RX branch notes, a stopped RX queue raises no further
interrupt. Wouldn't that RX queue then stay stalled until it is closed
and reopened?

[ ... ]

> +		if (rx_err & BIT(i)) {
> +			/* pairs with smp_store_release() in rxq_alloc */
> +			rxq = smp_load_acquire(&drv_info->rxq[i]);
> +			if (!rxq)
> +				continue;
> +			ret = mtk_cldma_stop_queue(drv_info, DIR_RX, i);

[Severity: Medium]
Is this stop serialized with rxq->rx_done_work? drv_info->wq is
allocated with WQ_UNBOUND, so err_work and rx_done_work can run at the
same time on different CPUs.

Consider an rx_done_work near its end. It can read need_restart == 0 and
call mtk_cldma_resume_queue(DIR_RX) right after this stop. err_work then
sets need_restart and queues rx_done_work again. That call runs
mtk_cldma_rxq_restart() and reprograms the start address on a queue that
may already be running again.

The comment below says the queue is handed back to its worker instead
of being programmed here. The stop command is also queue programming,
though, and it is not serialized with the owner's resume or restart.

> +			if (ret) {

[ ... ]

> +			atomic_set(&rxq->need_restart, 1);
> +			queue_work(drv_info->wq, &rxq->rx_done_work);

[Severity: High]
Can the RX engine be restarted after mtk_cldma_rxq_free() has stopped
it?

The end of mtk_cldma_rx_done_work() checks need_exit and then acts on
it, with no lock:

    if (!atomic_read(&rxq->need_exit)) {
        if (atomic_xchg(&rxq->need_restart, 0))
            mtk_cldma_rxq_restart(drv_info, rxq);
        else if (ret != -ENXIO)
            mtk_cldma_resume_queue(drv_info, DIR_RX, rxq->rxqno);
    }

rxq_free() sets need_exit and calls mtk_cldma_stop_queue(DIR_RX), then
synchronize_irq() and flush_work(). A worker that read need_exit == 0
before that can still call mtk_cldma_rxq_restart() after the stop has
completed. flush_work() only waits for the restart to finish.

The stop returned 0, so ring_leaked stays false. rxq_free() then unmaps
and frees the RX skbs and returns the GPDs, while the SO engine may be
running on HWO descriptors.

Setting need_restart here on every RX QUEUE_ERROR makes this
interleaving concrete. The added flush_work(&drv_info->err_work) in
rxq_free() does not stop the hardware again.

[ ... ]

> diff --git a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
> index d485bd7e3bf9c..658189dace01b 100644
> --- a/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
> +++ b/drivers/net/wwan/t9xx/pcie/mtk_trans_ctrl.c
> @@ -473,19 +473,23 @@ int mtk_pcie_hif_exit(struct mtk_md_dev *mdev)

[ ... ]

>  	/* Join the service threads before freeing what they dereference: they
>  	 * read trans->dev without a lock, so clearing it cannot stop a
>  	 * consumer that has already loaded the pointer.
>  	 */
>  	mtk_ctrl_trb_srv_exit(trans);
>  	mtk_cldma_exit(trans);

[Severity: High]
Does this ordering leave a window where trb_srv is used after it has
been freed?

mtk_ctrl_trb_srv_exit() calls kthread_stop() and kfree(srv), and only
then sets trans->trb_srv[i] = NULL. The CLDMA IRQ is masked, and
drv_info->wq drained, only later in
mtk_cldma_exit()->mtk_cldma_dev_exit().

This patch registers mtk_cldma_isr(), so mtk_cldma_tx_done_work() and
mtk_cldma_err_work()->mtk_cldma_txq_flush() can run inside that window.
Both do this without a lock:

    srv = trans->trb_srv[trans->srv_cfg[hif_id][txqno]];
    if (was_starved && srv)
        wake_up(&srv->trb_waitq);

srv can already be freed at that point. It is even still non-NULL
between the kfree() and the NULL store.

The commit message says "mtk_pcie_hif_exit() joins the trb service
threads before mtk_cldma_exit() frees what they dereference". However,
the IRQ and work producers that also dereference trb_srv[] are not
quiesced before srv is freed.

A queue can still be open at this point. That happens, for example, when
the DEV_RM submit fails and removal forces cleanup, or when
mtk_port_ch_disable() times out.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports
  2026-09-30  7:46 ` [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports Jack Wu via B4 Relay
@ 2026-10-04  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-04  9:12 UTC (permalink / raw)
  To: jackbb_wu
  Cc: loic.poulain, ryazanov.s.a, johannes, andrew+netdev, davem,
	edumazet, kuba, pabeni, wen-zhi.huang, shi-wei.yeh, Minano.tseng,
	matthias.bgg, angelogioacchino.delregno, horms, corbet, skhan,
	robert_yu, Jeff_Chang, linux-kernel, netdev, linux-arm-kernel,
	linux-mediatek, linux-doc

Thank you for your contribution! Sashiko AI review found 8 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 7 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] ABBA deadlock between port->w_priv.w_lock and
  trans->submit_lock.
- [Medium] A blocking write() to the new AT or MBIM WWAN port can fail
  with -EAGAIN even though the fd is blocking.
- [Medium] A blocking multi-packet write can report the whole write as
  failed after part of it has been sent.
- [Medium] mtk_port_wwan_enable() sets PORT_S_ENABLE and calls
  wwan_create_port(), which registers an openable char device, and only…
- [Medium] The new WWAN enable path reads port->tx_mtu straight after
  mtk_port_ch_enable() returns.
- [Medium] One write() to the AT or MBIM device can queue any number of
  CCCI packets on the CLDMA TX skb_list, far past SKB_LIST_MAX_LEN (16).
- [Medium] mtk_port_common_write() calls skb_copy_bits(src, cur_pos, ...)
  on the same non-linear source head for each 3440-byte packet.
- [Medium] The CCCI sequence counters for the new AT and MBIM ports
  (port->tx_seq, port->rx_seq) are reset only in mtk_port_struct_init(),…

--

net: wwan: t9xx: Add AT & MBIM WWAN ports

This commit adds AT and MBIM WWAN ports to the t9xx port infrastructure.
The ports are created on FSM_STATE_READY through a new
mtk_port_enable_by_type() hook. TX back-pressure is reported with
wwan_port_txon()/txoff() and a new per-port tx_complete hook.

> diff --git a/drivers/net/wwan/t9xx/mtk_port_io.c b/drivers/net/wwan/t9xx/mtk_port_io.c
> index 6cff0704b7fc8..0d8f63e2dbdaa 100644
> --- a/drivers/net/wwan/t9xx/mtk_port_io.c
> +++ b/drivers/net/wwan/t9xx/mtk_port_io.c

[ ... ]

> @@ -41,6 +46,73 @@ static void mtk_port_struct_init(struct mtk_port *port)

[ ... ]

> +static int mtk_port_common_write(struct mtk_port *port, struct sk_buff *src, bool blocking)
> +{
> +	u32 packet_size, left_cnt = src->len, cur_pos;
> +	struct sk_buff_head list;
> +	bool force_send = false;
> +	struct sk_buff *skb;
> +	int ret;
> +
> +	ret = mtk_port_status_check(port);
> +	if (ret)
> +		return ret;
> +
> +	__skb_queue_head_init(&list);
> +
> +	while (left_cnt) {
> +		skb = __dev_alloc_skb(port->tx_mtu, GFP_KERNEL);
> +		if (!skb) {
> +			ret = -ENOMEM;
> +			goto err_purge;
> +		}
> +
> +		skb_reserve(skb, sizeof(struct mtk_ccci_header));
> +
> +		packet_size = min_t(u32, left_cnt,
> +				    port->tx_mtu - sizeof(struct mtk_ccci_header));
> +		cur_pos = src->len - left_cnt;
> +		ret = skb_copy_bits(src, cur_pos, skb_put(skb, packet_size), packet_size);

[Severity: Medium]
Does this make large writes quadratic in the number of fragments?

On every call, skb_copy_bits() starts at the head of src's frag_list
(skb_walk_frags()) and adds up the lengths of the earlier fragments before
it reaches cur_pos. Here it is called once per 3440-byte packet on the same
source head. A 1 GiB write is about 312k fragments, which works out to
about 5e10 list-node visits.

The loop doesn't check for signals and has no size cap. It also runs under
the core's ops_lock in wwan_port_op_tx(), so other users of the port and
wwan_remove_port() are blocked while it runs.

> +		if (ret) {
> +			dev_err(port->port_mngr->ctrl_blk->mdev->dev,
> +				"Failed to copy data for port(%s)\n", port->info.name);
> +			dev_kfree_skb_any(skb);
> +			goto err_purge;
> +		}
> +
> +		__skb_queue_tail(&list, skb);
> +		left_cnt -= packet_size;
> +	}
> +
> +	while ((skb = __skb_dequeue(&list))) {
> +		ret = mtk_port_send_data(port, skb, blocking, force_send);
> +		if (ret < 0 && ret != -EINTR)
> +			goto err_purge;

[Severity: Medium]
Can a later packet fail here for some reason other than the channel going
away?

In the blocking case, packet 1 has already completed on the wire before
packet 2 is submitted. Suppose packet 2 then fails DMA mapping in
mtk_cldma_txbuf_set():

    dev_err_ratelimited(mdev->dev, "Failed to map dma! error:%d\n", ret);
    return -ENOMEM;

mtk_ctrl_trb_handler() completes the trb with that status, and
mtk_port_send_data() returns it:

    ret = (!trb->status) ? len : trb->status;

This loop then returns -ENOMEM for the whole write. wwan_port_fops_write()
can't report a partial count, so userspace sees a failed write. If it
resends, the prefix that was already sent is duplicated.

The comment above the function says only failures that mean the channel
itself is going away can happen after the first submit. Doesn't a
transient DMA mapping failure, with the channel still enabled afterwards,
contradict that?

> +
> +		force_send = true;

[Severity: Medium]
Is there anything that limits how many packets one write() can queue here?

Every packet after the first is submitted with force_send set. The only
depth check in mtk_pcie_hif_submit_skb() is skipped in that case:

    if (mtk_queue_list_is_full(trans, que) && !force_send) {

Take an O_NONBLOCK write close to MAX_RW_COUNT. The whole second copy built
above stays on the CLDMA TX skb_list, far past SKB_LIST_MAX_LEN, until the
modem drains it. At peak the kernel holds about twice the write size (the
core's skbs plus these copies), and none of it is charged to a memcg.

A blocking write can end up the same way. After a signal, each later
wait_event_interruptible_timeout() in mtk_port_send_data() returns
-ERESTARTSYS at once. The resulting -EINTR is treated as accepted, so the
remaining packets are force-queued without waiting.

> +	}
> +
> +	return 0;
> +
> +err_purge:
> +	__skb_queue_purge(&list);
> +	return ret;
> +}

[ ... ]

> @@ -253,6 +325,289 @@ static const struct port_ops port_internal_ops = {

[ ... ]

> +static void mtk_port_wwan_tx_pause(struct mtk_port *port)
> +{
> +	union ctrl_hif_cmd_data hif_cmd;
> +	struct mtk_ctrl_blk *ctrl_blk;
> +	int ret;
> +
> +	ctrl_blk = port->port_mngr->ctrl_blk;
> +
> +	mutex_lock(&port->w_priv.w_lock);
> +	if (!port->w_priv.w_port)
> +		goto unlock;
> +
> +	wwan_port_txoff(port->w_priv.w_port);
> +
> +	hif_cmd.rx_ch = port->info.rx_ch;
> +	ret = mtk_pcie_hif_cmd_func(ctrl_blk->mdev, HIF_CTRL_CMD_CHECK_TX_FULL,
> +				    &hif_cmd);

[Severity: High]
Can this deadlock against mtk_pcie_hif_exit()?

This path takes w_lock, then takes submit_lock inside
mtk_pcie_hif_cmd_func(). mtk_pcie_hif_exit() holds submit_lock for the
whole teardown, and during that teardown it completes this port's trbs:

mtk_pcie_hif_exit()
  mutex_lock(&trans->submit_lock)
  mtk_ctrl_trb_srv_exit()
    kthread_stop() -> mtk_ctrl_chs_flush() -> mtk_ctrl_ch_flush()
      trb->trb_complete() -> mtk_port_tx_complete()
        mtk_port_wwan_tx_complete()
          mutex_lock(&port->w_priv.w_lock)
  mtk_cldma_exit()
    mtk_cldma_txq_free()
      flush_work(&txq->tx_done_work), then trb_complete(-EPIPE)
    mtk_cldma_rxq_free()
      rxq->rx_done() -> ... -> mtk_port_wwan_recv()
        mutex_lock(&port->w_priv.w_lock)

That is submit_lock -> w_lock, the reverse of the order used here.

On removal, mtk_pci_dev_exit() carries on with "forcing cleanup" when
FSM_EVT_DEV_RM fails, for example when the event allocation fails. That
means mtk_trans_ctrl_exit() -> mtk_pcie_hif_exit() can run while the WWAN
ports are still registered and open.

Consider a writer that got -EAGAIN and is in this function. It holds
w_lock and the core's ops_lock, and blocks on submit_lock. Meanwhile
hif_exit waits in kthread_stop() or flush_work() for a completion that is
itself blocked on w_lock.

The comment added in mtk_port_tx_complete() says:

    /* Runs in the trb_srv kthread, so the hook may sleep on a mutex. */

However, the hook also runs under submit_lock during HIF teardown.

In the normal FSM_STATE_OFF path the ports are disabled before the trans
handler runs, so that path should not deadlock. It can still record the
submit_lock -> w_lock edge whenever a WWAN channel disable fails, which
would give a lockdep circular dependency report.

> +	if (ret <= 0)
> +		wwan_port_txon(port->w_priv.w_port);
> +unlock:
> +	mutex_unlock(&port->w_priv.w_lock);
> +}

[ ... ]

> +static int mtk_port_wwan_tx(struct wwan_port *w_port, struct sk_buff *skb, bool blocking)
> +{
> +	struct mtk_port *port = wwan_port_get_drvdata(w_port);
> +	int ret;
> +
> +	if (unlikely(!skb->len)) {
> +		consume_skb(skb);
> +		return 0;
> +	}
> +
> +	ret = mtk_port_common_write(port, skb, blocking);
> +	if (ret < 0) {
> +		if (ret == -EAGAIN)
> +			mtk_port_wwan_tx_pause(port);
> +		return ret;
> +	}

[Severity: Medium]
Can a blocking write() get -EAGAIN here?

Without O_NONBLOCK, the core calls ops->tx_blocking, which reaches this
function with blocking=true. However, mtk_port_common_write() always
submits the first packet with force_send=false:

    bool force_send = false;

If the queue's skb_list already holds SKB_LIST_MAX_LEN entries,
mtk_pcie_hif_submit_skb() returns -EAGAIN at once, and tx_blocking returns
it unchanged. wwan_port_fops_write() passes it straight back to the
blocking caller:

    ret = wwan_port_op_tx(port, head, !!(filp->f_flags & O_NONBLOCK));

The wait on WWAN_PORT_TX_OFF in wwan_wait_tx() only runs at the start of
the next write.

The queue can already be full when a blocking write starts. An earlier
non-blocking write (on another fd, or on the same fd before fcntl) or an
interrupted blocking write can leave it that way, because every packet
after the first skips the limit.

Should the blocking path wait for space and retry instead?

[ ... ]

> +static void mtk_port_wwan_enable(struct mtk_port *port)
> +{
> +	struct mtk_port_mngr *port_mngr;
> +	struct wwan_port_caps caps;
> +	struct wwan_port *wp;
> +	int ret;
> +
> +	port_mngr = port->port_mngr;
> +
> +	if (test_bit(PORT_S_ENABLE, &port->status))
> +		return;
> +
> +	ret = mtk_port_ch_enable(port);

[Severity: Medium]
Should tx_seq and rx_seq be reset for each modem session?

They are only reset in mtk_port_struct_init(), which runs once from
mtk_port_wwan_init(). Neither mtk_port_wwan_enable() (run on every
FSM_STATE_READY) nor mtk_port_wwan_disable() resets them.

After an OFF -> READY cycle without a driver reload, rx_seq still holds the
previous session's value. Suppose the modem restarts its sequence numbers
at 0. Its first frame with the assert bit set would then fail this check
in mtk_port_check_rx_seq():

    if (assert_bit && port->rx_seq &&
        ((seq_num - port->rx_seq) & MTK_CHECK_RX_SEQ_MASK) != 1) {

and be dropped with -EPROTO. In the other direction, mtk_port_add_header()
stamps host packets with the stale tx_seq and AST=1.

This depends on the modem firmware resetting its CCCI sequence numbers
across the cycle. The initial rx_seq = -1 suggests it does, but the host
code can't confirm it. The internal port enable path has the same gap.

> +	if (ret && ret != -EBUSY) {

[ ... ]

> +	/* tx_mtu is only valid once the channel open trb has completed. A zero
> +	 * frag_len would make wwan_port_fops_write() loop forever.
> +	 */
> +	if (!port->tx_mtu) {
> +		dev_err(port_mngr->ctrl_blk->mdev->dev,
> +			"Invalid tx_mtu for port(%s)\n", port->info.name);
> +		mtk_port_ch_disable(port);
> +		return;
> +	}

[Severity: Medium]
Is tx_mtu guaranteed to be visible at this point?

mtk_port_ch_enable() only waits for trb->status <= 0. In mtk_cldma_open(),
the trb_srv kthread publishes the status before it runs the completion
callback:

    trb->status = ret;
    trb->trb_complete(skb);

It is mtk_port_open_trb_complete() that writes port->tx_mtu:

    if (!trb->status) {
        port->tx_mtu = trb_open_priv->tx_mtu;

There is no barrier or separate completion between the two stores. A waiter
that checks the condition while the kthread is between them can see status
0 with tx_mtu still 0. That can happen on the first check after being
preempted after submit, or on a timeout recheck. On weakly ordered CPUs the
two stores can also become visible out of order.

On the first enable, this would log "Invalid tx_mtu", disable a channel
that opened successfully, and never create the AT or MBIM device for that
session.

> +
> +	/* The core allocates frag_len + headroom_len and skb_put()s frag_len,
> +	 * so frag_len is the payload budget: subtract the CCCI header to make
> +	 * one core fragment exactly one CCCI packet.
> +	 */
> +	caps.frag_len = port->tx_mtu - sizeof(struct mtk_ccci_header);
> +	caps.headroom_len = sizeof(struct mtk_ccci_header);
> +
> +	/* These bits must be set before wwan_create_port(): the device node
> +	 * becomes openable inside it and mtk_port_common_open() rejects a
> +	 * port without PORT_S_ENABLE.  w_port cannot be published first - it is
> +	 * this call's return value - so an RX frame arriving in between is
> +	 * dropped with -ENXIO by design.
> +	 */
> +	set_bit(PORT_S_WR, &port->status);
> +	set_bit(PORT_S_ENABLE, &port->status);
> +
> +	wp = wwan_create_port(port_mngr->ctrl_blk->mdev->dev,
> +			      port->w_priv.w_type,
> +			      &wwan_ops, &caps, port);
> +	if (IS_ERR(wp)) {

[ ... ]

> +	mutex_lock(&port->w_priv.w_lock);
> +	port->w_priv.w_port = wp;
> +	mutex_unlock(&port->w_priv.w_lock);
> +}

[Severity: Medium]
Can a reply to a legitimate command be dropped in this window?

Once wwan_create_port() returns, the char device can be opened, and
mtk_port_wwan_open() only requires PORT_S_ENABLE. Userspace (udev or
ModemManager reacting to the new node) can open the port and send a
command, since the TX path doesn't use w_port.

If the modem's reply arrives before the assignment above,
mtk_port_wwan_recv() sees PORT_S_OPEN set while w_port is still NULL:

    if (!test_bit(PORT_S_OPEN, &port->status) || !port->w_priv.w_port) {

and drops the reply with -ENXIO. The comment before wwan_create_port()
covers unsolicited RX arriving in between. It does not cover a reply to a
command sent after a successful open.

The start() callback is passed the wwan_port. Could it publish w_port
itself?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930-t9xx_driver_v1-v9-0-3248b9256956%40compal.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-10-04  9:12 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  7:46 [PATCH v9 0/6] net: wwan: t9xx: Add MediaTek T9XX WWAN driver Jack Wu via B4 Relay
2026-09-30  7:46 ` [PATCH v9 1/6] net: wwan: t9xx: Add PCIe core Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko
2026-09-30  7:46 ` [PATCH v9 2/6] net: wwan: t9xx: Add control plane transaction layer Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko
2026-09-30  7:46 ` [PATCH v9 3/6] net: wwan: t9xx: Add control DMA interface Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko
2026-09-30  7:46 ` [PATCH v9 4/6] net: wwan: t9xx: Add control port Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko
2026-09-30  7:46 ` [PATCH v9 5/6] net: wwan: t9xx: Add FSM thread Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko
2026-09-30  7:46 ` [PATCH v9 6/6] net: wwan: t9xx: Add AT & MBIM WWAN ports Jack Wu via B4 Relay
2026-10-04  9:12   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®