* [PATCH v15 net-next 0/2] octeontx2-pf: mqprio bandwidth offload for NIX TX schedulers
@ 2026-09-11 10:55 Ratheesh Kannoth
2026-09-11 10:55 ` [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
2026-09-11 10:55 ` [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
0 siblings, 2 replies; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-09-11 10:55 UTC (permalink / raw)
To: bpf, linux-kernel, netdev
Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk,
john.fastabend, kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth
This series adds hardware offload for channel-mode mqprio with
TC_MQPRIO_SHAPER_BW_RATE on Marvell octeontx2 PF devices. Each
non-QoS transmit queue is shaped by programming MDQ CIR/PIR on the NIX TX
scheduler. When bandwidth offload is active, the driver allocates one
SMQ per queue, parents every MDQ under TL4[0], and maps each traffic
class min/max rate to the queue(s) in that class.
The NIX TX scheduler hierarchy cannot be reprogrammed live today, so
mqprio add, replace, delete, and failed-replace rollback rebuild it by
bouncing the netdev through ndo_stop()/ndo_open(). That intentionally
drops in-flight traffic on each change. Before ndo_stop(), quiesce the
transmit path with dev_deactivate() so xmit cannot race queue teardown.
After a successful bounce on a running interface, reactivate TX queues
with dev_activate(). Cache the active rates and restore MDQ shapers from
otx2_mqprio_up() during ndo_open(); fail open if restoration fails.
Track mqprio configuration in mq_offload_snap snapshots (TC layout and
rates). On tc qdisc replace, stage the new configuration while keeping
the previous snapshot for rollback: failed setup restores the old
snapshot via netdev restart when the interface is running, successful
graft is recorded through TC_ROOT_GRAFT, and teardown of the replaced
qdisc instance commits the staged snapshot without tearing down the live
offload.
Patch 1 converts PF/VF and representor flag access to atomic bitops.
Patch 2 depends on it for safe OTX2_FLAG_INTF_DOWN and OTX2_FLAG_PORT_UP
updates on asynchronous mbox paths and during the mqprio netdev bounce.
The driver rejects offload unless the interface is running and the device
advertises CIR+PIR support. Per-TC rates are rejected when a traffic class
spans more than one queue. Concurrent PFC, XDP, SDP rep, or HTB use is
blocked, and ethtool channel count changes are blocked while mqprio
bandwidth offload is active.
Ratheesh Kannoth (2):
octeontx2: use atomic bitops for PF/VF and rep flags
octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
.../net/ethernet/marvell/octeontx2/af/rvu_nix.c | 6 +-
.../ethernet/marvell/octeontx2/nic/cn10k_ipsec.c | 8 +-
.../ethernet/marvell/octeontx2/nic/otx2_common.c | 154 +++-
.../ethernet/marvell/octeontx2/nic/otx2_common.h | 100 ++-
.../ethernet/marvell/octeontx2/nic/otx2_dcbnl.c | 6 +
.../ethernet/marvell/octeontx2/nic/otx2_devlink.c | 2 +-
.../ethernet/marvell/octeontx2/nic/otx2_ethtool.c | 29 +-
.../ethernet/marvell/octeontx2/nic/otx2_flows.c | 34 +-
.../net/ethernet/marvell/octeontx2/nic/otx2_pf.c | 94 +--
.../net/ethernet/marvell/octeontx2/nic/otx2_tc.c | 776 ++++++++++++++++++++-
.../net/ethernet/marvell/octeontx2/nic/otx2_txrx.c | 16 +-
.../net/ethernet/marvell/octeontx2/nic/otx2_vf.c | 10 +-
.../net/ethernet/marvell/octeontx2/nic/otx2_xsk.c | 4 +-
drivers/net/ethernet/marvell/octeontx2/nic/qos.c | 11 +
.../net/ethernet/marvell/octeontx2/nic/qos_sq.c | 4 +-
drivers/net/ethernet/marvell/octeontx2/nic/rep.c | 32 +-
drivers/net/ethernet/marvell/octeontx2/nic/rep.h | 3 +-
17 files changed, 1138 insertions(+), 151 deletions(-)
---
v14 -> v15: Addressed sashiko comments.
- Split atomic PF/VF and representor flag access into a preparatory patch
so mqprio netdev-restart and mbox paths can update OTX2_FLAG_INTF_DOWN
and OTX2_FLAG_PORT_UP without data races on the shared flags word.
- Clear mqprio software state when hardware shaper teardown fails, warn,
and still bounce the netdev on delete so offload does not remain stuck
active after a mailbox error.
https://lore.kernel.org/netdev/20260904031553.3196916-1-rkannoth@marvell.com/
v13 -> v14: Addressed sashiko comments.
- Quiesce TX with dev_deactivate() before ndo_stop() and dev_activate()
after ndo_open() in otx2_mqprio_restart_netdev() to avoid xmit racing
queue teardown.
- Use atomic set_bit()/clear_bit() for OTX2_FLAG_INTF_DOWN and
OTX2_FLAG_PORT_UP updates on netdev-restart and mbox paths.
- Block concurrent mqprio bandwidth offload and HTB shaping.
- Fail ndo_open() if otx2_mqprio_up() cannot restore MDQ shapers.
- Rebuild the TX scheduler via netdev restart in otx2_mqprio_restore_old()
when rolling back a failed replace on a running interface.
- Return an error from otx2_mqprio_down() if clearing hardware shapers
fails instead of clearing software state anyway.
https://lore.kernel.org/netdev/20260904031553.3196916-1-rkannoth@marvell.com/
v12 -> v13: Addressed sashiko comments.
https://sashiko.dev/#/patchset/20260903023324.3078284-1-rkannoth%40marvell.com
v11 -> v12: Addressed sashiko comments.
https://sashiko.dev/#/patchset/20260902015500.2985371-1-rkannoth%40marvell.com
v10 -> v11: Addressed sashiko comments.
https://sashiko.dev/#/patchset/20260831131014.2639581-1-rkannoth%40marvell.com
v9 -> v10: Addressed sashiko/jacub comments.
https://sashiko.dev/#/message/20260817032747.1765883-1-rkannoth%40marvell.com
v8 -> v9: Addressed Sashiko comments
https://lore.kernel.org/netdev/aoJ6FhtWue0FHDQV@rkannoth-OptiPlex-7090/
v7 -> v8: Addressed Sashiko comments
https://sashiko.dev/#/patchset/20260811085050.3212280-1-rkannoth%40marvell.com
v6 -> v7: Addressed Sashiko comments
https://sashiko.dev/#/message/20260810034738.1786029-1-rkannoth%40marvell.com
v5 -> v6: Addressed Sashiko comments
https://lore.kernel.org/netdev/20260806095434.1144397-1-rkannoth@marvell.com/
v4 -> v5: Addressed sashiko comments
https://sashiko.dev/#/patchset/20260803042724.3380209-1-rkannoth%40marvell.com
v3 -> v4: Addressed sashiko comments
https://lore.kernel.org/netdev/20260729105139.2302908-1-rkannoth@marvell.com/
v2 -> v3: Addressed sashiko comments
https://lore.kernel.org/netdev/amnYX866mYx02cBe@rkannoth-OptiPlex-7090/T/#m67310cbec48b21c7720858ab3a1ea083a0f8dc10
v1 -> v2: Addressed sashiko comments
https://lore.kernel.org/netdev/20260724075010.2665758-1-rkannoth@marvell.com/
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags
2026-09-11 10:55 [PATCH v15 net-next 0/2] octeontx2-pf: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
@ 2026-09-11 10:55 ` Ratheesh Kannoth
2026-09-17 11:31 ` Paolo Abeni
2026-09-11 10:55 ` [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
1 sibling, 1 reply; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-09-11 10:55 UTC (permalink / raw)
To: bpf, linux-kernel, netdev
Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk,
john.fastabend, kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth
Replace non-atomic u64 flag read-modify-write with unsigned long
bitmaps and set_bit/clear_bit/test_bit access across the NIC driver.
Add otx2_set_flag(), otx2_clear_flag() and otx2_test_flag() helpers
for struct otx2_nic, use bitops directly on rep_dev->flags, and sync
representor state to the PF mailbox context via otx2_sync_flags_from_rep().
Define representor VF initialization as OTX2_REP_VF_INITIALIZED (bit 21)
in the shared flag namespace.
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
---
.../marvell/octeontx2/nic/cn10k_ipsec.c | 8 +-
.../marvell/octeontx2/nic/otx2_common.c | 8 +-
.../marvell/octeontx2/nic/otx2_common.h | 73 +++++++++++------
.../marvell/octeontx2/nic/otx2_devlink.c | 2 +-
.../marvell/octeontx2/nic/otx2_ethtool.c | 21 +++--
.../marvell/octeontx2/nic/otx2_flows.c | 34 ++++----
.../ethernet/marvell/octeontx2/nic/otx2_pf.c | 78 +++++++++----------
.../ethernet/marvell/octeontx2/nic/otx2_tc.c | 30 +++----
.../marvell/octeontx2/nic/otx2_txrx.c | 16 ++--
.../ethernet/marvell/octeontx2/nic/otx2_vf.c | 10 +--
.../ethernet/marvell/octeontx2/nic/otx2_xsk.c | 4 +-
.../ethernet/marvell/octeontx2/nic/qos_sq.c | 4 +-
.../net/ethernet/marvell/octeontx2/nic/rep.c | 32 ++++----
.../net/ethernet/marvell/octeontx2/nic/rep.h | 3 +-
14 files changed, 174 insertions(+), 149 deletions(-)
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c b/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
index 77543d472345..50ec4542c418 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
@@ -334,7 +334,7 @@ static int cn10k_outb_cpt_init(struct net_device *netdev)
CN10K_CPT_LF_NQX(0));
/* Set ipsec offload enabled for this device */
- pf->flags |= OTX2_FLAG_IPSEC_OFFLOAD_ENABLED;
+ otx2_set_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED);
cn10k_cpt_device_set_available(pf);
return 0;
@@ -356,7 +356,7 @@ static int cn10k_outb_cpt_clean(struct otx2_nic *pf)
}
/* Set ipsec offload disabled for this device */
- pf->flags &= ~OTX2_FLAG_IPSEC_OFFLOAD_ENABLED;
+ otx2_clear_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED);
/* Disable CPTLF Instruction Queue (IQ) */
cn10k_outb_cptlf_iq_disable(pf);
@@ -820,7 +820,7 @@ void cn10k_ipsec_clean(struct otx2_nic *pf)
if (!is_dev_support_ipsec_offload(pf->pdev))
return;
- if (!(pf->flags & OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
+ if (!otx2_test_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
return;
if (pf->ipsec.sa_workq) {
@@ -945,7 +945,7 @@ bool cn10k_ipsec_transmit(struct otx2_nic *pf, struct netdev_queue *txq,
u16 dlen;
/* Check for IPSEC offload enabled */
- if (!(pf->flags & OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
+ if (!otx2_test_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
goto drop;
sp = skb_sec_path(skb);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
index 175992188c18..b421cb75e44b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
@@ -220,10 +220,10 @@ int otx2_set_mac_address(struct net_device *netdev, void *p)
eth_hw_addr_set(netdev, addr->sa_data);
/* update dmac field in vlan offload rule */
if (netif_running(netdev) &&
- pfvf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+ otx2_test_flag(pfvf, OTX2_FLAG_RX_VLAN_SUPPORT))
otx2_install_rxvlan_offload_flow(pfvf);
/* update dmac address in ntuple and DMAC filter list */
- if (pfvf->flags & OTX2_FLAG_DMACFLTR_SUPPORT)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_DMACFLTR_SUPPORT))
otx2_dmacflt_update_pfmac_flow(pfvf);
} else {
return -EPERM;
@@ -275,8 +275,8 @@ int otx2_config_pause_frm(struct otx2_nic *pfvf)
goto unlock;
}
- req->rx_pause = !!(pfvf->flags & OTX2_FLAG_RX_PAUSE_ENABLED);
- req->tx_pause = !!(pfvf->flags & OTX2_FLAG_TX_PAUSE_ENABLED);
+ req->rx_pause = otx2_test_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
+ req->tx_pause = otx2_test_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
req->set = 1;
err = otx2_sync_mbox_msg(&pfvf->mbox);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
index eecee612b7b2..7e09c1444a6d 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
@@ -491,28 +491,29 @@ struct otx2_nic {
u16 tx_max_pktlen;
u16 rbsize; /* Receive buffer size */
-#define OTX2_FLAG_RX_TSTAMP_ENABLED BIT_ULL(0)
-#define OTX2_FLAG_TX_TSTAMP_ENABLED BIT_ULL(1)
-#define OTX2_FLAG_INTF_DOWN BIT_ULL(2)
-#define OTX2_FLAG_MCAM_ENTRIES_ALLOC BIT_ULL(3)
-#define OTX2_FLAG_NTUPLE_SUPPORT BIT_ULL(4)
-#define OTX2_FLAG_UCAST_FLTR_SUPPORT BIT_ULL(5)
-#define OTX2_FLAG_RX_VLAN_SUPPORT BIT_ULL(6)
-#define OTX2_FLAG_VF_VLAN_SUPPORT BIT_ULL(7)
-#define OTX2_FLAG_PF_SHUTDOWN BIT_ULL(8)
-#define OTX2_FLAG_RX_PAUSE_ENABLED BIT_ULL(9)
-#define OTX2_FLAG_TX_PAUSE_ENABLED BIT_ULL(10)
-#define OTX2_FLAG_TC_FLOWER_SUPPORT BIT_ULL(11)
-#define OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED BIT_ULL(12)
-#define OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED BIT_ULL(13)
-#define OTX2_FLAG_DMACFLTR_SUPPORT BIT_ULL(14)
-#define OTX2_FLAG_PTP_ONESTEP_SYNC BIT_ULL(15)
-#define OTX2_FLAG_ADPTV_INT_COAL_ENABLED BIT_ULL(16)
-#define OTX2_FLAG_TC_MARK_ENABLED BIT_ULL(17)
-#define OTX2_FLAG_REP_MODE_ENABLED BIT_ULL(18)
-#define OTX2_FLAG_PORT_UP BIT_ULL(19)
-#define OTX2_FLAG_IPSEC_OFFLOAD_ENABLED BIT_ULL(20)
- u64 flags;
+#define OTX2_FLAG_RX_TSTAMP_ENABLED 0
+#define OTX2_FLAG_TX_TSTAMP_ENABLED 1
+#define OTX2_FLAG_INTF_DOWN 2
+#define OTX2_FLAG_MCAM_ENTRIES_ALLOC 3
+#define OTX2_FLAG_NTUPLE_SUPPORT 4
+#define OTX2_FLAG_UCAST_FLTR_SUPPORT 5
+#define OTX2_FLAG_RX_VLAN_SUPPORT 6
+#define OTX2_FLAG_VF_VLAN_SUPPORT 7
+#define OTX2_FLAG_PF_SHUTDOWN 8
+#define OTX2_FLAG_RX_PAUSE_ENABLED 9
+#define OTX2_FLAG_TX_PAUSE_ENABLED 10
+#define OTX2_FLAG_TC_FLOWER_SUPPORT 11
+#define OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED 12
+#define OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED 13
+#define OTX2_FLAG_DMACFLTR_SUPPORT 14
+#define OTX2_FLAG_PTP_ONESTEP_SYNC 15
+#define OTX2_FLAG_ADPTV_INT_COAL_ENABLED 16
+#define OTX2_FLAG_TC_MARK_ENABLED 17
+#define OTX2_FLAG_REP_MODE_ENABLED 18
+#define OTX2_FLAG_PORT_UP 19
+#define OTX2_FLAG_IPSEC_OFFLOAD_ENABLED 20
+#define OTX2_REP_VF_INITIALIZED 21
+ unsigned long flags;
u64 *cq_op_addr;
struct bpf_prog *xdp_prog;
@@ -594,6 +595,34 @@ struct otx2_nic {
unsigned long *af_xdp_zc_qidx;
};
+static inline void otx2_set_flag(struct otx2_nic *nic, unsigned int flag)
+{
+ set_bit(flag, &nic->flags);
+}
+
+static inline void otx2_clear_flag(struct otx2_nic *nic, unsigned int flag)
+{
+ clear_bit(flag, &nic->flags);
+}
+
+static inline bool otx2_test_flag(struct otx2_nic *nic, unsigned int flag)
+{
+ return test_bit(flag, &nic->flags);
+}
+
+static inline void otx2_sync_flags_from_rep(struct otx2_nic *dst,
+ unsigned long *src_flags)
+{
+ unsigned int flag;
+
+ for (flag = 0; flag <= OTX2_REP_VF_INITIALIZED; flag++) {
+ if (test_bit(flag, src_flags))
+ set_bit(flag, &dst->flags);
+ else
+ clear_bit(flag, &dst->flags);
+ }
+}
+
static inline bool is_otx2_lbkvf(struct pci_dev *pdev)
{
return (pdev->device == PCI_DEVID_OCTEONTX2_RVU_AFVF) ||
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
index 4a5ce0e67dda..863a5ced9a26 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
@@ -104,7 +104,7 @@ static int otx2_dl_ucast_flt_cnt_validate(struct devlink *devlink, u32 id,
struct otx2_nic *pfvf = otx2_dl->pfvf;
/* Check for UNICAST filter support*/
- if (!(pfvf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT)) {
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_UCAST_FLTR_SUPPORT)) {
NL_SET_ERR_MSG_MOD(extack,
"Unicast filter not enabled");
return -EINVAL;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
index 9bee1b91eeaa..9a05aa5d902d 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
@@ -354,14 +354,14 @@ static int otx2_set_pauseparam(struct net_device *netdev,
return -EOPNOTSUPP;
if (pause->rx_pause)
- pfvf->flags |= OTX2_FLAG_RX_PAUSE_ENABLED;
+ otx2_set_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
else
- pfvf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
+ otx2_clear_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
if (pause->tx_pause)
- pfvf->flags |= OTX2_FLAG_TX_PAUSE_ENABLED;
+ otx2_set_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
else
- pfvf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+ otx2_clear_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
return otx2_config_pause_frm(pfvf);
}
@@ -470,8 +470,7 @@ static int otx2_get_coalesce(struct net_device *netdev,
cmd->rx_max_coalesced_frames = hw->cq_ecount_wait;
cmd->tx_coalesce_usecs = hw->cq_time_wait;
cmd->tx_max_coalesced_frames = hw->cq_ecount_wait;
- if ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) ==
- OTX2_FLAG_ADPTV_INT_COAL_ENABLED) {
+ if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
cmd->use_adaptive_rx_coalesce = 1;
cmd->use_adaptive_tx_coalesce = 1;
} else {
@@ -502,15 +501,14 @@ static int otx2_set_coalesce(struct net_device *netdev,
}
/* Check and update coalesce status */
- if ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) ==
- OTX2_FLAG_ADPTV_INT_COAL_ENABLED) {
+ if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
priv_coalesce_status = 1;
if (!ec->use_adaptive_rx_coalesce)
- pfvf->flags &= ~OTX2_FLAG_ADPTV_INT_COAL_ENABLED;
+ otx2_clear_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED);
} else {
priv_coalesce_status = 0;
if (ec->use_adaptive_rx_coalesce)
- pfvf->flags |= OTX2_FLAG_ADPTV_INT_COAL_ENABLED;
+ otx2_set_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED);
}
/* 'cq_time_wait' is 8bit and is in multiple of 100ns,
@@ -556,8 +554,7 @@ static int otx2_set_coalesce(struct net_device *netdev,
* 'on' to 'off'.
*/
if (priv_coalesce_status &&
- ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) !=
- OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
+ (!otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED))) {
hw->cq_time_wait = CQ_TIMER_THRESH_DEFAULT;
hw->cq_ecount_wait = CQ_CQE_THRESH_DEFAULT;
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
index 99d78fc5a2c4..b8ff49f0f6e3 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
@@ -270,9 +270,9 @@ int otx2_alloc_mcam_entries(struct otx2_nic *pfvf, u16 count)
flow_cfg->max_flows = allocated;
if (allocated) {
- pfvf->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
- pfvf->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
- pfvf->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+ otx2_set_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
+ otx2_set_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT);
+ otx2_set_flag(pfvf, OTX2_FLAG_TC_FLOWER_SUPPORT);
}
if (allocated != count)
@@ -376,7 +376,7 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
flow_cfg->unicast_offset = vf_vlan_max_flows;
flow_cfg->rx_vlan_offset = flow_cfg->unicast_offset +
flow_cfg->ucast_flt_cnt;
- pfvf->flags |= OTX2_FLAG_UCAST_FLTR_SUPPORT;
+ otx2_set_flag(pfvf, OTX2_FLAG_UCAST_FLTR_SUPPORT);
/* Check if NPC_DMAC field is supported
* by the mkex profile before setting VLAN support flag.
@@ -401,11 +401,11 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
}
if (frsp->enable) {
- pfvf->flags |= OTX2_FLAG_RX_VLAN_SUPPORT;
- pfvf->flags |= OTX2_FLAG_VF_VLAN_SUPPORT;
+ otx2_set_flag(pfvf, OTX2_FLAG_RX_VLAN_SUPPORT);
+ otx2_set_flag(pfvf, OTX2_FLAG_VF_VLAN_SUPPORT);
}
- pfvf->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
+ otx2_set_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
mutex_unlock(&pfvf->mbox.lock);
/* Allocate entries for Ntuple filters */
@@ -415,7 +415,7 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
return 0;
}
- pfvf->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+ otx2_set_flag(pfvf, OTX2_FLAG_TC_FLOWER_SUPPORT);
refcount_set(&flow_cfg->mark_flows, 1);
return 0;
@@ -479,7 +479,7 @@ int otx2_mcam_flow_init(struct otx2_nic *pf)
return err;
/* Check if MCAM entries are allocate or not */
- if (!(pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT))
+ if (!otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
return 0;
pf->mac_table = devm_kzalloc(pf->dev, sizeof(struct otx2_mac_table)
@@ -501,7 +501,7 @@ int otx2_mcam_flow_init(struct otx2_nic *pf)
if (!pf->flow_cfg->bmap_to_dmacindex)
return -ENOMEM;
- pf->flags |= OTX2_FLAG_DMACFLTR_SUPPORT;
+ otx2_set_flag(pf, OTX2_FLAG_DMACFLTR_SUPPORT);
return 0;
}
@@ -521,7 +521,7 @@ static int otx2_do_add_macfilter(struct otx2_nic *pf, const u8 *mac)
struct npc_install_flow_req *req;
int err, i;
- if (!(pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT))
+ if (!otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
return -ENOMEM;
/* dont have free mcam entries or uc list is greater than alloted */
@@ -1167,7 +1167,7 @@ static int otx2_is_flow_rule_dmacfilter(struct otx2_nic *pfvf,
u64 ring_cookie = fsp->ring_cookie;
u32 flow_type;
- if (!(pfvf->flags & OTX2_FLAG_DMACFLTR_SUPPORT))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_DMACFLTR_SUPPORT))
return false;
flow_type = fsp->flow_type & ~(FLOW_EXT | FLOW_MAC_EXT | FLOW_RSS);
@@ -1364,7 +1364,7 @@ int otx2_add_flow(struct otx2_nic *pfvf, struct ethtool_rxnfc *nfc)
}
ring = ethtool_get_flow_spec_ring(fsp->ring_cookie);
- if (!(pfvf->flags & OTX2_FLAG_NTUPLE_SUPPORT))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT))
return -ENOMEM;
/* Number of queues on a VF can be greater or less than
@@ -1596,7 +1596,7 @@ int otx2_destroy_ntuple_flows(struct otx2_nic *pfvf)
struct otx2_flow *iter, *tmp;
int err;
- if (!(pfvf->flags & OTX2_FLAG_NTUPLE_SUPPORT))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT))
return 0;
if (!flow_cfg->max_flows)
@@ -1629,7 +1629,7 @@ int otx2_destroy_mcam_flows(struct otx2_nic *pfvf)
struct otx2_flow *iter, *tmp;
int err;
- if (!(pfvf->flags & OTX2_FLAG_MCAM_ENTRIES_ALLOC))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC))
return 0;
/* remove all flows */
@@ -1658,7 +1658,7 @@ int otx2_destroy_mcam_flows(struct otx2_nic *pfvf)
return err;
}
- pfvf->flags &= ~OTX2_FLAG_MCAM_ENTRIES_ALLOC;
+ otx2_clear_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
flow_cfg->max_flows = 0;
mutex_unlock(&pfvf->mbox.lock);
@@ -1721,7 +1721,7 @@ int otx2_enable_rxvlan(struct otx2_nic *pf, bool enable)
int err;
/* Dont have enough mcam entries */
- if (!(pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT))
+ if (!otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
return -ENOMEM;
if (enable) {
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
index c0e2100de1d9..32582b6347ea 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
@@ -879,7 +879,7 @@ static void otx2_handle_link_event(struct otx2_nic *pf)
struct cgx_link_user_info *linfo = &pf->linfo;
struct net_device *netdev = pf->netdev;
- if (pf->flags & OTX2_FLAG_PORT_UP)
+ if (otx2_test_flag(pf, OTX2_FLAG_PORT_UP))
return;
pr_info("%s NIC Link is %s %d Mbps %s duplex\n", netdev->name,
@@ -907,11 +907,11 @@ static int otx2_mbox_up_handler_rep_event_up_notify(struct otx2_nic *pf,
if (info->event == RVU_EVENT_PORT_STATE) {
if (info->evt_data.port_state) {
- pf->flags |= OTX2_FLAG_PORT_UP;
+ otx2_set_flag(pf, OTX2_FLAG_PORT_UP);
netif_carrier_on(netdev);
netif_tx_start_all_queues(netdev);
} else {
- pf->flags &= ~OTX2_FLAG_PORT_UP;
+ otx2_clear_flag(pf, OTX2_FLAG_PORT_UP);
netif_tx_stop_all_queues(netdev);
netif_carrier_off(netdev);
}
@@ -953,7 +953,7 @@ int otx2_mbox_up_handler_cgx_link_event(struct otx2_nic *pf,
}
/* interface has not been fully configured yet */
- if (pf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
return 0;
otx2_handle_link_event(pf);
@@ -1828,7 +1828,7 @@ void otx2_free_hw_resources(struct otx2_nic *pf)
free_req = otx2_mbox_alloc_msg_nix_lf_free(mbox);
if (free_req) {
free_req->flags = NIX_LF_DISABLE_FLOWS | NIX_LF_DONT_FREE_DFT_IDXS;
- if (!(pf->flags & OTX2_FLAG_PF_SHUTDOWN))
+ if (!otx2_test_flag(pf, OTX2_FLAG_PF_SHUTDOWN))
free_req->flags |= NIX_LF_DONT_FREE_TX_VTAG;
if (otx2_sync_mbox_msg(mbox))
dev_err(pf->dev, "%s failed to free nixlf\n", __func__);
@@ -2135,21 +2135,21 @@ int otx2_open(struct net_device *netdev)
}
otx2_write64(pf, NIX_LF_RAS_ENA_W1S, NIX_LF_RAS_MASK);
- if (pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
otx2_enable_rxvlan(pf, true);
/* When reinitializing enable time stamping if it is enabled before */
- if (pf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED) {
- pf->flags &= ~OTX2_FLAG_TX_TSTAMP_ENABLED;
+ if (otx2_test_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED)) {
+ otx2_clear_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED);
otx2_config_hw_tx_tstamp(pf, true);
}
- if (pf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED) {
- pf->flags &= ~OTX2_FLAG_RX_TSTAMP_ENABLED;
+ if (otx2_test_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED)) {
+ otx2_clear_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED);
otx2_config_hw_rx_tstamp(pf, true);
}
- pf->flags &= ~OTX2_FLAG_INTF_DOWN;
- pf->flags &= ~OTX2_FLAG_PORT_UP;
+ otx2_clear_flag(pf, OTX2_FLAG_INTF_DOWN);
+ otx2_clear_flag(pf, OTX2_FLAG_PORT_UP);
/* 'intf_down' may be checked on any cpu */
smp_wmb();
@@ -2161,7 +2161,7 @@ int otx2_open(struct net_device *netdev)
otx2_handle_link_event(pf);
/* Install DMAC Filters */
- if (pf->flags & OTX2_FLAG_DMACFLTR_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_DMACFLTR_SUPPORT))
otx2_dmacflt_reinstall_flows(pf);
otx2_tc_apply_ingress_police_rules(pf);
@@ -2186,7 +2186,7 @@ int otx2_open(struct net_device *netdev)
err_tx_stop_queues:
netif_tx_stop_all_queues(netdev);
netif_carrier_off(netdev);
- pf->flags |= OTX2_FLAG_INTF_DOWN;
+ otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
/* free NIXLF POISON irq */
vec = pci_irq_vector(pf->pdev,
pf->hw.nix_msixoff + NIX_LF_POISON_VEC);
@@ -2220,13 +2220,13 @@ int otx2_stop(struct net_device *netdev)
int qidx, vec, wrk;
/* If the DOWN flag is set resources are already freed */
- if (pf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
return 0;
netif_carrier_off(netdev);
netif_tx_stop_all_queues(netdev);
- pf->flags |= OTX2_FLAG_INTF_DOWN;
+ otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
/* 'intf_down' may be checked on any cpu */
smp_wmb();
@@ -2457,7 +2457,7 @@ static int otx2_config_hw_rx_tstamp(struct otx2_nic *pfvf, bool enable)
struct msg_req *req;
int err;
- if (pfvf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED && enable)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED) && enable)
return 0;
mutex_lock(&pfvf->mbox.lock);
@@ -2478,9 +2478,9 @@ static int otx2_config_hw_rx_tstamp(struct otx2_nic *pfvf, bool enable)
mutex_unlock(&pfvf->mbox.lock);
if (enable)
- pfvf->flags |= OTX2_FLAG_RX_TSTAMP_ENABLED;
+ otx2_set_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED);
else
- pfvf->flags &= ~OTX2_FLAG_RX_TSTAMP_ENABLED;
+ otx2_clear_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED);
return 0;
}
@@ -2489,7 +2489,7 @@ static int otx2_config_hw_tx_tstamp(struct otx2_nic *pfvf, bool enable)
struct msg_req *req;
int err;
- if (pfvf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED && enable)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED) && enable)
return 0;
mutex_lock(&pfvf->mbox.lock);
@@ -2510,9 +2510,9 @@ static int otx2_config_hw_tx_tstamp(struct otx2_nic *pfvf, bool enable)
mutex_unlock(&pfvf->mbox.lock);
if (enable)
- pfvf->flags |= OTX2_FLAG_TX_TSTAMP_ENABLED;
+ otx2_set_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
else
- pfvf->flags &= ~OTX2_FLAG_TX_TSTAMP_ENABLED;
+ otx2_clear_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
return 0;
}
@@ -2537,8 +2537,8 @@ int otx2_config_hwtstamp_set(struct net_device *netdev,
switch (config->tx_type) {
case HWTSTAMP_TX_OFF:
- if (pfvf->flags & OTX2_FLAG_PTP_ONESTEP_SYNC)
- pfvf->flags &= ~OTX2_FLAG_PTP_ONESTEP_SYNC;
+ if (otx2_test_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC))
+ otx2_clear_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC);
cancel_delayed_work(&pfvf->ptp->synctstamp_work);
otx2_config_hw_tx_tstamp(pfvf, false);
@@ -2549,7 +2549,7 @@ int otx2_config_hwtstamp_set(struct net_device *netdev,
"One-step time stamping is not supported");
return -ERANGE;
}
- pfvf->flags |= OTX2_FLAG_PTP_ONESTEP_SYNC;
+ otx2_set_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC);
schedule_delayed_work(&pfvf->ptp->synctstamp_work,
msecs_to_jiffies(500));
fallthrough;
@@ -2835,7 +2835,7 @@ static int otx2_set_vf_vlan(struct net_device *netdev, int vf, u16 vlan, u8 qos,
if (proto != htons(ETH_P_8021Q))
return -EPROTONOSUPPORT;
- if (!(pf->flags & OTX2_FLAG_VF_VLAN_SUPPORT))
+ if (!otx2_test_flag(pf, OTX2_FLAG_VF_VLAN_SUPPORT))
return -EOPNOTSUPP;
return otx2_do_set_vf_vlan(pf, vf, vlan, qos, proto);
@@ -3086,7 +3086,7 @@ int otx2_realloc_msix_vectors(struct otx2_nic *pf)
* interrupt range (QINT, CINT, GINT, ERR and POISON vectors).
*/
num_vec = hw->nix_msixoff;
- if (pf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+ if (otx2_test_flag(pf, OTX2_FLAG_REP_MODE_ENABLED))
num_vec += NIX_LF_CINT_VEC_START + hw->max_queues;
else
num_vec += NIX_LF_POISON_VEC + 1;
@@ -3273,7 +3273,7 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
pf->pdev = pdev;
pf->dev = dev;
pf->total_vfs = pci_sriov_get_totalvfs(pdev);
- pf->flags |= OTX2_FLAG_INTF_DOWN;
+ otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
hw = &pf->hw;
hw->pdev = pdev;
@@ -3328,23 +3328,23 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
if (err)
goto err_del_mcam_entries;
- if (pf->flags & OTX2_FLAG_NTUPLE_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_NTUPLE_SUPPORT))
netdev->hw_features |= NETIF_F_NTUPLE;
- if (pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
netdev->priv_flags |= IFF_UNICAST_FLT;
/* Support TSO on tag interface */
netdev->vlan_features |= netdev->features;
netdev->hw_features |= NETIF_F_HW_VLAN_CTAG_TX |
NETIF_F_HW_VLAN_STAG_TX;
- if (pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
netdev->hw_features |= NETIF_F_HW_VLAN_CTAG_RX |
NETIF_F_HW_VLAN_STAG_RX;
netdev->features |= netdev->hw_features;
/* HW supports tc offload but mutually exclusive with n-tuple filters */
- if (pf->flags & OTX2_FLAG_TC_FLOWER_SUPPORT)
+ if (otx2_test_flag(pf, OTX2_FLAG_TC_FLOWER_SUPPORT))
netdev->hw_features |= NETIF_F_HW_TC;
netdev->hw_features |= NETIF_F_LOOPBACK | NETIF_F_RXALL;
@@ -3595,18 +3595,18 @@ static void otx2_remove(struct pci_dev *pdev)
pf = netdev_priv(netdev);
- pf->flags |= OTX2_FLAG_PF_SHUTDOWN;
+ otx2_set_flag(pf, OTX2_FLAG_PF_SHUTDOWN);
- if (pf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED)
+ if (otx2_test_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED))
otx2_config_hw_tx_tstamp(pf, false);
- if (pf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED)
+ if (otx2_test_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED))
otx2_config_hw_rx_tstamp(pf, false);
/* Disable 802.3x pause frames */
- if (pf->flags & OTX2_FLAG_RX_PAUSE_ENABLED ||
- (pf->flags & OTX2_FLAG_TX_PAUSE_ENABLED)) {
- pf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
- pf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+ if (otx2_test_flag(pf, OTX2_FLAG_RX_PAUSE_ENABLED) ||
+ otx2_test_flag(pf, OTX2_FLAG_TX_PAUSE_ENABLED)) {
+ otx2_clear_flag(pf, OTX2_FLAG_RX_PAUSE_ENABLED);
+ otx2_clear_flag(pf, OTX2_FLAG_TX_PAUSE_ENABLED);
otx2_config_pause_frm(pf);
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
index 039fd47ebf52..ddb46b580c3b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
@@ -159,7 +159,7 @@ static int otx2_tc_validate_flow(struct otx2_nic *nic,
struct flow_action *actions,
struct netlink_ext_ack *extack)
{
- if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+ if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
return -EINVAL;
}
@@ -223,7 +223,7 @@ static int otx2_tc_egress_matchall_install(struct otx2_nic *nic,
if (err)
return err;
- if (nic->flags & OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED) {
+ if (otx2_test_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED)) {
NL_SET_ERR_MSG_MOD(extack,
"Only one Egress MATCHALL ratelimiter can be offloaded");
return -ENOMEM;
@@ -244,7 +244,7 @@ static int otx2_tc_egress_matchall_install(struct otx2_nic *nic,
otx2_convert_rate(entry->police.rate_bytes_ps));
if (err)
return err;
- nic->flags |= OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED;
+ otx2_set_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED);
break;
default:
NL_SET_ERR_MSG_MOD(extack,
@@ -261,13 +261,13 @@ static int otx2_tc_egress_matchall_delete(struct otx2_nic *nic,
struct netlink_ext_ack *extack = cls->common.extack;
int err;
- if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+ if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
return -EINVAL;
}
err = otx2_set_matchall_egress_rate(nic, 0, 0);
- nic->flags &= ~OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED;
+ otx2_clear_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED);
return err;
}
@@ -505,7 +505,7 @@ static int otx2_tc_parse_actions(struct otx2_nic *nic,
mark = act->mark;
req->match_id = mark & OTX2_RX_MATCH_ID_MASK;
req->op = NIX_RX_ACTION_DEFAULT;
- nic->flags |= OTX2_FLAG_TC_MARK_ENABLED;
+ otx2_set_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
refcount_inc(&nic->flow_cfg->mark_flows);
break;
@@ -942,7 +942,7 @@ static void otx2_destroy_tc_flow_list(struct otx2_nic *pfvf)
struct otx2_flow_config *flow_cfg = pfvf->flow_cfg;
struct otx2_tc_flow *iter, *tmp;
- if (!(pfvf->flags & OTX2_FLAG_MCAM_ENTRIES_ALLOC))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC))
return;
list_for_each_entry_safe(iter, tmp, &flow_cfg->flow_list_tc, list) {
@@ -1195,12 +1195,12 @@ static int otx2_tc_del_flow(struct otx2_nic *nic,
/* Disable TC MARK flag if they are no rules with skbedit mark action */
if (flow_node->req.match_id)
if (!refcount_dec_and_test(&flow_cfg->mark_flows))
- nic->flags &= ~OTX2_FLAG_TC_MARK_ENABLED;
+ otx2_clear_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
if (flow_node->is_act_police) {
__clear_bit(flow_node->rq, &nic->rq_bmap);
- if (nic->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN))
goto free_mcam_flow;
mutex_lock(&nic->mbox.lock);
@@ -1246,10 +1246,10 @@ static int otx2_tc_add_flow(struct otx2_nic *nic,
struct npc_install_flow_req *req, dummy;
int rc, err, entry;
- if (!(nic->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
+ if (!otx2_test_flag(nic, OTX2_FLAG_TC_FLOWER_SUPPORT))
return -ENOMEM;
- if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+ if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
return -EINVAL;
}
@@ -1444,7 +1444,7 @@ static int otx2_tc_ingress_matchall_install(struct otx2_nic *nic,
if (err)
return err;
- if (nic->flags & OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED) {
+ if (otx2_test_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED)) {
NL_SET_ERR_MSG_MOD(extack,
"Only one ingress MATCHALL ratelimitter can be offloaded");
return -ENOMEM;
@@ -1469,7 +1469,7 @@ static int otx2_tc_ingress_matchall_install(struct otx2_nic *nic,
err = cn10k_set_matchall_ipolicer_rate(nic, entry->police.burst, rate);
if (err)
return err;
- nic->flags |= OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED;
+ otx2_set_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED);
break;
default:
NL_SET_ERR_MSG_MOD(extack,
@@ -1486,13 +1486,13 @@ static int otx2_tc_ingress_matchall_delete(struct otx2_nic *nic,
struct netlink_ext_ack *extack = cls->common.extack;
int err;
- if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+ if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
return -EINVAL;
}
err = cn10k_free_matchall_ipolicer(nic);
- nic->flags &= ~OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED;
+ otx2_clear_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED);
return err;
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
index 8d2d607bc92f..f65ba44db60b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
@@ -171,7 +171,7 @@ static void otx2_set_rxtstamp(struct otx2_nic *pfvf,
u64 timestamp, tsns;
int err;
- if (!(pfvf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED))
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED))
return;
timestamp = pfvf->ptp->convert_rx_ptp_tstmp(*(u64 *)data);
@@ -374,13 +374,13 @@ static void otx2_rcv_pkt_handler(struct otx2_nic *pfvf,
}
otx2_set_rxhash(pfvf, cqe, skb);
- if (!(pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)) {
+ if (!otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED)) {
skb_record_rx_queue(skb, cq->cq_idx);
if (pfvf->netdev->features & NETIF_F_RXCSUM)
skb->ip_summed = CHECKSUM_UNNECESSARY;
}
- if (pfvf->flags & OTX2_FLAG_TC_MARK_ENABLED)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_TC_MARK_ENABLED))
skb->mark = parse->match_id;
skb_mark_for_recycle(skb);
@@ -513,7 +513,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
((u64)cq->cq_idx << 32) | processed_cqe);
#if IS_ENABLED(CONFIG_RVU_ESWITCH)
- if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
ndev = pfvf->reps[qidx]->netdev;
else
#endif
@@ -526,7 +526,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
if (qidx >= pfvf->hw.tx_queues)
qidx -= pfvf->hw.xdp_queues;
- if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
qidx = 0;
txq = netdev_get_tx_queue(ndev, qidx);
netdev_tx_completed_queue(txq, tx_pkts, tx_bytes);
@@ -599,11 +599,11 @@ int otx2_napi_handler(struct napi_struct *napi, int budget)
if (workdone < budget && napi_complete_done(napi, workdone)) {
/* If interface is going down, don't re-enable IRQ */
- if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
return workdone;
/* Adjust irq coalese using net_dim */
- if (pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED))
otx2_adjust_adaptive_coalese(pfvf, cq_poll);
if (likely(cq))
@@ -1137,7 +1137,7 @@ static void otx2_set_txtstamp(struct otx2_nic *pfvf, struct sk_buff *skb,
if (unlikely(!skb_shinfo(skb)->gso_size &&
(skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP))) {
- if (unlikely(pfvf->flags & OTX2_FLAG_PTP_ONESTEP_SYNC &&
+ if (unlikely(otx2_test_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC) &&
otx2_ptp_is_sync(skb, &ptp_offset, &udp_csum_crt))) {
origin_tstamp = (struct ptpv2_tstamp *)
((u8 *)skb->data + ptp_offset +
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
index f7765e19d78a..5f7915231ca3 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
@@ -610,7 +610,7 @@ static int otx2vf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
vf->dev = dev;
vf->iommu_domain = iommu_get_domain_for_dev(dev);
- vf->flags |= OTX2_FLAG_INTF_DOWN;
+ otx2_set_flag(vf, OTX2_FLAG_INTF_DOWN);
hw = &vf->hw;
hw->pdev = vf->pdev;
hw->rx_queues = qcount;
@@ -824,10 +824,10 @@ static void otx2vf_remove(struct pci_dev *pdev)
vf = netdev_priv(netdev);
/* Disable 802.3x pause frames */
- if (vf->flags & OTX2_FLAG_RX_PAUSE_ENABLED ||
- (vf->flags & OTX2_FLAG_TX_PAUSE_ENABLED)) {
- vf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
- vf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+ if (otx2_test_flag(vf, OTX2_FLAG_RX_PAUSE_ENABLED) ||
+ otx2_test_flag(vf, OTX2_FLAG_TX_PAUSE_ENABLED)) {
+ otx2_clear_flag(vf, OTX2_FLAG_RX_PAUSE_ENABLED);
+ otx2_clear_flag(vf, OTX2_FLAG_TX_PAUSE_ENABLED);
otx2_config_pause_frm(vf);
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
index 0e8a6a6486c4..7808588a0234 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
@@ -96,7 +96,7 @@ static void otx2_clean_up_rq(struct otx2_nic *pfvf, int qidx)
u64 iova;
/* If the DOWN flag is set SQs are already freed */
- if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
return;
cq = &qset->cq[qidx];
@@ -172,7 +172,7 @@ int otx2_xsk_wakeup(struct net_device *dev, u32 queue_id, u32 flags)
struct otx2_cq_poll *cq_poll = NULL;
struct otx2_qset *qset = &pf->qset;
- if (pf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
return -ENETDOWN;
if (queue_id >= pf->hw.rx_queues || queue_id >= pf->hw.tx_queues)
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
index 2872adabc830..5f09e2960144 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
@@ -238,7 +238,7 @@ int otx2_qos_enable_sq(struct otx2_nic *pfvf, int qidx)
struct otx2_hw *hw = &pfvf->hw;
int pool_id, sq_idx, err;
- if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
return -EPERM;
sq_idx = hw->non_qos_queues + qidx;
@@ -288,7 +288,7 @@ void otx2_qos_disable_sq(struct otx2_nic *pfvf, int qidx)
sq_idx = hw->non_qos_queues + qidx;
/* If the DOWN flag is set SQs are already freed */
- if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+ if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
return;
sq = &pfvf->qset.sq[sq_idx];
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
index 0f5d5642d3f7..8a8c0088fd20 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
@@ -93,9 +93,9 @@ static int rvu_rep_mcam_flow_init(struct rep_dev *rep)
rep->flow_cfg->max_flows = allocated;
if (allocated) {
- rep->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
- rep->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
- rep->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+ set_bit(OTX2_FLAG_MCAM_ENTRIES_ALLOC, &rep->flags);
+ set_bit(OTX2_FLAG_NTUPLE_SUPPORT, &rep->flags);
+ set_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags);
}
INIT_LIST_HEAD(&rep->flow_cfg->flow_list);
@@ -109,14 +109,14 @@ static int rvu_rep_setup_tc_cb(enum tc_setup_type type,
struct rep_dev *rep = cb_priv;
struct otx2_nic *priv = rep->mdev;
- if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+ if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
return -EINVAL;
- if (!(rep->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
+ if (!test_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags))
rvu_rep_mcam_flow_init(rep);
priv->netdev = rep->netdev;
- priv->flags = rep->flags;
+ otx2_sync_flags_from_rep(priv, &rep->flags);
priv->pcifunc = rep->pcifunc;
priv->flow_cfg = rep->flow_cfg;
@@ -303,9 +303,9 @@ static void rvu_rep_state_evt_handler(struct otx2_nic *priv,
rep_id = rvu_rep_get_repid(priv, info->pcifunc);
rep = priv->reps[rep_id];
if (info->evt_data.vf_state)
- rep->flags |= RVU_REP_VF_INITIALIZED;
+ set_bit(OTX2_REP_VF_INITIALIZED, &rep->flags);
else
- rep->flags &= ~RVU_REP_VF_INITIALIZED;
+ clear_bit(OTX2_REP_VF_INITIALIZED, &rep->flags);
}
int rvu_event_up_notify(struct otx2_nic *pf, struct rep_event *info)
@@ -382,7 +382,7 @@ static void rvu_rep_get_stats64(struct net_device *dev,
{
struct rep_dev *rep = netdev_priv(dev);
- if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+ if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
return;
stats->rx_packets = rep->stats.rx_frames;
@@ -453,7 +453,7 @@ static int rvu_rep_open(struct net_device *dev)
struct otx2_nic *priv = rep->mdev;
struct rep_event evt = {0};
- if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+ if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
return 0;
netif_carrier_on(dev);
@@ -472,7 +472,7 @@ static int rvu_rep_stop(struct net_device *dev)
struct otx2_nic *priv = rep->mdev;
struct rep_event evt = {0};
- if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+ if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
return 0;
netif_carrier_off(dev);
@@ -547,7 +547,7 @@ static int rvu_rep_napi_init(struct otx2_nic *priv,
otx2_write64(priv, NIX_LF_CINTX_INT(qidx), BIT_ULL(0));
otx2_write64(priv, NIX_LF_CINTX_ENA_W1S(qidx), BIT_ULL(0));
}
- priv->flags &= ~OTX2_FLAG_INTF_DOWN;
+ otx2_clear_flag(priv, OTX2_FLAG_INTF_DOWN);
return 0;
err_free_cints:
@@ -632,7 +632,7 @@ void rvu_rep_destroy(struct otx2_nic *priv)
int rep_id;
rvu_eswitch_config(priv, false);
- priv->flags |= OTX2_FLAG_INTF_DOWN;
+ otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
rvu_rep_free_cq_rsrc(priv);
for (rep_id = 0; rep_id < priv->rep_cnt; rep_id++) {
rep = priv->reps[rep_id];
@@ -801,8 +801,8 @@ static int rvu_rep_probe(struct pci_dev *pdev, const struct pci_device_id *id)
pci_set_drvdata(pdev, priv);
priv->pdev = pdev;
priv->dev = dev;
- priv->flags |= OTX2_FLAG_INTF_DOWN;
- priv->flags |= OTX2_FLAG_REP_MODE_ENABLED;
+ otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
+ otx2_set_flag(priv, OTX2_FLAG_REP_MODE_ENABLED);
hw = &priv->hw;
hw->pdev = pdev;
@@ -845,7 +845,7 @@ static void rvu_rep_remove(struct pci_dev *pdev)
struct otx2_nic *priv = pci_get_drvdata(pdev);
otx2_unregister_dl(priv);
- if (!(priv->flags & OTX2_FLAG_INTF_DOWN))
+ if (!otx2_test_flag(priv, OTX2_FLAG_INTF_DOWN))
rvu_rep_destroy(priv);
otx2_detach_resources(&priv->mbox);
if (priv->hw.lmt_info)
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.h b/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
index 5bc9e2c7d800..45707c434d89 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
@@ -37,8 +37,7 @@ struct rep_dev {
struct delayed_work stats_wrk;
struct devlink_port dl_port;
struct otx2_flow_config *flow_cfg;
-#define RVU_REP_VF_INITIALIZED BIT_ULL(0)
- u64 flags;
+ unsigned long flags;
u16 rep_id;
u16 pcifunc;
u8 mac[ETH_ALEN];
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
2026-09-11 10:55 [PATCH v15 net-next 0/2] octeontx2-pf: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-09-11 10:55 ` [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
@ 2026-09-11 10:55 ` Ratheesh Kannoth
2026-09-17 11:31 ` Paolo Abeni
1 sibling, 1 reply; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-09-11 10:55 UTC (permalink / raw)
To: bpf, linux-kernel, netdev
Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk,
john.fastabend, kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth
Add TC_SETUP_QDISC_MQPRIO offload for channel-mode mqprio with
TC_MQPRIO_SHAPER_BW_RATE. Program per-queue MDQ CIR/PIR through the
NIX TX scheduler mailbox for each non-QoS transmit queue. When offload
is active, allocate one SMQ per such queue, parent every MDQ under
TL4[0], and map each traffic-class min/max rate to the queue(s) in that
class.
The NIX TX scheduler hierarchy cannot be reprogrammed live today, so
mqprio add, replace, delete, and failed-replace rollback rebuild it by
bouncing the netdev through ndo_stop()/ndo_open(). That intentionally
drops in-flight traffic on each change. Before ndo_stop(), quiesce the
transmit path with dev_deactivate() so xmit cannot race queue teardown.
After a successful bounce on a running interface, reactivate TX queues
with dev_activate(). Cache the active rates and restore MDQ shapers from
otx2_mqprio_up() during ndo_open(); fail open if restoration fails.
Track mqprio configuration in mq_offload_snap snapshots (TC layout and
rates). On tc qdisc replace, stage the new configuration while keeping
the previous snapshot for rollback: failed setup restores the old
snapshot via netdev restart when the interface is running, successful
graft is recorded through TC_ROOT_GRAFT, and teardown of the replaced
qdisc instance commits the staged snapshot without tearing down the
live offload.
Reject offload unless the interface is running and the device advertises
CIR+PIR support. Reject per-TC rates for traffic classes mapped to more
than one queue. Block concurrent use with PFC, XDP, SDP rep, or HTB (and
block HTB while mqprio offload is active). Block ethtool channel count
changes while mqprio bandwidth offload is active.
Use atomic bit operations when updating OTX2_FLAG_PORT_UP and
OTX2_FLAG_INTF_DOWN on asynchronous and netdev-restart paths.
Add a ratelimited AF debug message when validating TX scheduler queue
ownership to aid mqprio hierarchy setup failures.
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
---
.../ethernet/marvell/octeontx2/af/rvu_nix.c | 6 +-
.../marvell/octeontx2/nic/otx2_common.c | 146 +++-
.../marvell/octeontx2/nic/otx2_common.h | 27 +
.../marvell/octeontx2/nic/otx2_dcbnl.c | 6 +
.../marvell/octeontx2/nic/otx2_ethtool.c | 8 +
.../ethernet/marvell/octeontx2/nic/otx2_pf.c | 16 +
.../ethernet/marvell/octeontx2/nic/otx2_tc.c | 746 ++++++++++++++++++
.../net/ethernet/marvell/octeontx2/nic/qos.c | 11 +
8 files changed, 964 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
index d80d2c00bd84..2d1ab8c761c4 100644
--- a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
+++ b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
@@ -331,8 +331,12 @@ static bool is_valid_txschq(struct rvu *rvu, int blkaddr,
return true;
}
- if (map_func != pcifunc)
+ if (map_func != pcifunc) {
+ dev_err_ratelimited(rvu->dev,
+ "pcifunc %x map pcifunc %x not equal, lvl=%u schq=%u\n",
+ pcifunc, map_func, lvl, schq);
return false;
+ }
return true;
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
index b421cb75e44b..5bad2466da0c 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
@@ -615,6 +615,142 @@ void otx2_get_mac_from_af(struct net_device *netdev)
}
EXPORT_SYMBOL(otx2_get_mac_from_af);
+static int
+otx2_nix_tmq_reg_write(struct otx2_nic *pfvf, int cnt,
+ u64 reg_addr[MAX_REGS_PER_MBOX_MSG],
+ u64 reg_val[MAX_REGS_PER_MBOX_MSG])
+{
+ struct mbox *mbox = &pfvf->mbox;
+ struct nix_txschq_config *req;
+ int i, err;
+
+ mutex_lock(&mbox->lock);
+ req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+ if (!req) {
+ mutex_unlock(&mbox->lock);
+ return -ENOMEM;
+ }
+
+ req->lvl = NIX_TXSCH_LVL_MDQ;
+ req->num_regs = cnt;
+
+ for (i = 0; i < cnt; i++) {
+ req->reg[i] = reg_addr[i];
+ req->regval[i] = reg_val[i];
+ }
+
+ err = otx2_sync_mbox_msg(mbox);
+ mutex_unlock(&mbox->lock);
+
+ return err;
+}
+
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf)
+{
+ u64 reg_addr[MAX_REGS_PER_MBOX_MSG];
+ u64 reg_val[MAX_REGS_PER_MBOX_MSG];
+ int err, smq, i, cnt = 0;
+
+ for (i = 0; i < pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]; i++) {
+ smq = pfvf->hw.txschq_list[NIX_TXSCH_LVL_SMQ][i];
+
+ reg_addr[cnt] = NIX_AF_MDQX_PIR(smq);
+ reg_val[cnt] = 0;
+ cnt++;
+
+ reg_addr[cnt] = NIX_AF_MDQX_CIR(smq);
+ reg_val[cnt] = 0;
+ cnt++;
+
+ if (cnt < MAX_REGS_PER_MBOX_MSG - 1)
+ continue;
+
+ err = otx2_nix_tmq_reg_write(pfvf, cnt,
+ reg_addr, reg_val);
+ if (err)
+ goto fail;
+ cnt = 0;
+ }
+
+ if (cnt) {
+ err = otx2_nix_tmq_reg_write(pfvf, cnt,
+ reg_addr, reg_val);
+ if (err)
+ goto fail;
+ }
+
+ return 0;
+fail:
+ return err;
+}
+
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf,
+ int txq, u64 minrate, u64 maxrate)
+{
+ struct mbox *mbox = &pfvf->mbox;
+ struct nix_txschq_config *req;
+ int err, smq, n = 0;
+ u64 reg_addr[2];
+ u64 reg_val[2];
+ u64 rate;
+
+ if (!maxrate && !minrate) {
+ smq = otx2_get_smq_idx(pfvf, txq);
+ reg_addr[0] = NIX_AF_MDQX_PIR(smq);
+ reg_val[0] = 0;
+ reg_addr[1] = NIX_AF_MDQX_CIR(smq);
+ reg_val[1] = 0;
+ return otx2_nix_tmq_reg_write(pfvf, 2, reg_addr, reg_val);
+ }
+
+ smq = otx2_get_smq_idx(pfvf, txq);
+
+ mutex_lock(&mbox->lock);
+ req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+ if (!req) {
+ mutex_unlock(&mbox->lock);
+ return -ENOMEM;
+ }
+
+ req->lvl = NIX_TXSCH_LVL_MDQ;
+
+ /* MQPRIO exposes only min/max rate, not burst. Pass burst 0 so
+ * otx2_get_egress_burst_cfg() programmes the largest burst the NIX
+ * encoding supports (CN10K_MAX_BURST_SIZE on CN10K). This differs
+ * from the 65536 byte default used in the HTB path, which is a
+ * kernel-side default when no explicit burst is configured, not a
+ * hardware cap.
+ *
+ * mqprio setup restarts the netdev (otx2_mqprio_restart_netdev),
+ * which resets MDQ shapers to zero. Program both PIR and CIR on
+ * every update so omitted rates are applied explicitly rather than
+ * relying on stale hardware state.
+ */
+ req->reg[n] = NIX_AF_MDQX_PIR(smq);
+ if (maxrate) {
+ rate = otx2_convert_rate(maxrate);
+ req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 0);
+ } else {
+ req->regval[n] = 0;
+ }
+ n++;
+
+ /* CIR+PIR support is required and checked at mqprio setup. */
+ req->reg[n] = NIX_AF_MDQX_CIR(smq);
+ if (minrate) {
+ rate = otx2_convert_rate(minrate);
+ req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 0);
+ } else {
+ req->regval[n] = 0;
+ }
+ n++;
+ req->num_regs = n;
+
+ err = otx2_sync_mbox_msg(mbox);
+ mutex_unlock(&mbox->lock);
+ return err;
+}
+
int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for_pfc)
{
u16 (*schq_list)[MAX_TXSCHQ_PER_FUNC];
@@ -651,7 +787,11 @@ int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for
(u64)hw->smq_link_type);
req->num_regs++;
/* MDQ config */
- parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+ if (pfvf->mqprio.rate_limit)
+ parent = schq_list[NIX_TXSCH_LVL_TL4][0];
+ else
+ parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+
req->reg[1] = NIX_AF_MDQX_PARENT(schq);
req->regval[1] = parent << 16;
req->num_regs++;
@@ -779,6 +919,9 @@ int otx2_txsch_alloc(struct otx2_nic *pfvf)
req->schq[NIX_TXSCH_LVL_TL4] = chan_cnt;
}
+ if (pfvf->mqprio.rate_limit)
+ req->schq[NIX_TXSCH_LVL_SMQ] = pfvf->hw.non_qos_queues;
+
rc = otx2_sync_mbox_msg(&pfvf->mbox);
if (rc)
return rc;
@@ -844,6 +987,7 @@ void otx2_txschq_stop(struct otx2_nic *pfvf)
/* Clear the txschq list */
for (lvl = 0; lvl < NIX_TXSCH_LVL_CNT; lvl++) {
+ pfvf->hw.txschq_cnt[lvl] = 0;
for (schq = 0; schq < MAX_TXSCHQ_PER_FUNC; schq++)
pfvf->hw.txschq_list[lvl][schq] = 0;
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
index 7e09c1444a6d..6e7132f27b62 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
@@ -17,6 +17,7 @@
#include <linux/soc/marvell/silicons.h>
#include <linux/soc/marvell/octeontx2/asm.h>
#include <net/macsec.h>
+#include <uapi/linux/pkt_sched.h>
#include <net/pkt_cls.h>
#include <net/devlink.h>
#include <linux/time64.h>
@@ -483,6 +484,23 @@ struct pf_irq_data {
int mdevs;
};
+struct mq_offload_snap {
+ u64 min_rate[TC_QOPT_MAX_QUEUE];
+ u64 max_rate[TC_QOPT_MAX_QUEUE];
+ __u8 num_tc;
+ __u16 count[TC_QOPT_MAX_QUEUE];
+ __u16 offset[TC_QOPT_MAX_QUEUE];
+};
+
+struct otx2_mqprio {
+ u32 flags;
+ u64 *min_rate;
+ u64 *max_rate;
+ bool rate_limit;
+ bool replace_setup_done;
+ bool replace_graft_done;
+};
+
struct otx2_nic {
void __iomem *reg_base;
struct net_device *netdev;
@@ -516,6 +534,10 @@ struct otx2_nic {
unsigned long flags;
u64 *cq_op_addr;
+ struct otx2_mqprio mqprio;
+ struct mq_offload_snap *cur_mq_snap;
+ struct mq_offload_snap *old_mq_snap;
+
struct bpf_prog *xdp_prog;
struct otx2_qset qset;
struct otx2_hw hw;
@@ -1275,6 +1297,11 @@ dma_addr_t otx2_dma_map_skb_frag(struct otx2_nic *pfvf,
struct sk_buff *skb, int seg, int *len);
void otx2_dma_unmap_skb_frags(struct otx2_nic *pfvf, struct sg_list *sg);
int otx2_read_free_sqe(struct otx2_nic *pfvf, u16 qidx);
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf, int txq,
+ u64 minrate, u64 maxrate);
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf);
+int otx2_mqprio_down(struct otx2_nic *pfvf);
+int otx2_mqprio_up(struct otx2_nic *pfvf);
void otx2_queue_vf_work(struct mbox *mw, struct workqueue_struct *mbox_wq,
int first, int mdevs, u64 intr);
int otx2_del_mcam_flow_entry(struct otx2_nic *nic, u16 entry,
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
index 91d346d114af..b7bd08129fb6 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
@@ -413,6 +413,12 @@ static int otx2_dcbnl_ieee_setpfc(struct net_device *dev, struct ieee_pfc *pfc)
u8 old_pfc_en;
int err;
+ if (pfvf->mqprio.rate_limit && pfc->pfc_en) {
+ netdev_err(dev,
+ "PFC: cannot enable while mqprio bandwidth offload is active\n");
+ return -EOPNOTSUPP;
+ }
+
old_pfc_en = pfvf->pfc_en;
pfvf->pfc_en = pfc->pfc_en;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
index 9a05aa5d902d..e63ead3b1113 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
@@ -287,6 +287,14 @@ static int otx2_set_channels(struct net_device *dev,
return -EINVAL;
}
+ if (pfvf->mqprio.rate_limit &&
+ (channel->tx_count != pfvf->hw.tx_queues ||
+ channel->rx_count != pfvf->hw.rx_queues)) {
+ netdev_info(dev,
+ "Not permitted to change channel count while MQ prio is active\n");
+ return -EINVAL;
+ }
+
if (if_up)
dev->netdev_ops->ndo_stop(dev);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
index 32582b6347ea..1e17f8d49fd4 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
@@ -2007,6 +2007,14 @@ int otx2_open(struct net_device *netdev)
if (err)
goto err_free_mem;
+ err = otx2_mqprio_up(pf);
+ if (err) {
+ netdev_err(pf->netdev,
+ "mqprio: failed to restore shapers during open: %d\n",
+ err);
+ goto err_free_hw;
+ }
+
/* Register NAPI handler */
for (qidx = 0; qidx < pf->hw.cint_cnt; qidx++) {
cq_poll = &qset->napi[qidx];
@@ -2205,6 +2213,7 @@ int otx2_open(struct net_device *netdev)
free_irq(vec, pf);
err_disable_napi:
otx2_disable_napi(pf);
+err_free_hw:
otx2_free_hw_resources(pf);
err_free_mem:
otx2_free_queue_mem(qset);
@@ -2280,6 +2289,7 @@ int otx2_stop(struct net_device *netdev)
for (qidx = 0; qidx < netdev->num_tx_queues; qidx++)
netdev_tx_reset_queue(netdev_get_tx_queue(netdev, qidx));
+ synchronize_net();
otx2_free_queue_mem(qset);
/* Do not clear RQ/SQ ringsize settings */
memset_startat(qset, 0, sqe_cnt);
@@ -2923,6 +2933,12 @@ static int otx2_xdp_setup(struct otx2_nic *pf, struct bpf_prog *prog)
bool if_up = netif_running(pf->netdev);
struct bpf_prog *old_prog;
+ if (prog && pf->mqprio.rate_limit) {
+ netdev_err(dev,
+ "XDP: cannot attach while mqprio bandwidth offload is active\n");
+ return -EOPNOTSUPP;
+ }
+
if (prog && dev->mtu > MAX_XDP_MTU) {
netdev_warn(dev, "Jumbo frames not yet supported with XDP\n");
return -EOPNOTSUPP;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
index ddb46b580c3b..70d99ae3defb 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
@@ -16,6 +16,8 @@
#include <net/tc_act/tc_mirred.h>
#include <net/tc_act/tc_vlan.h>
#include <net/ipv6.h>
+#include <net/pkt_sched.h>
+#include <net/sch_generic.h>
#include "cn10k.h"
#include "otx2_common.h"
@@ -31,6 +33,10 @@
#define MCAST_INVALID_GRP (-1U)
#define RATE_MANTISSA_BITS 8
+/* Min per-queue egress shaping rate the NIX TLX encoder supports (2 Mbps). */
+#define OTX2_MQPRIO_MIN_RATE_BYTES_PS 250000ULL
+/* Max egress shaping rate the NIX TLX encoder supports (130816 Mbps). */
+#define OTX2_MQPRIO_MAX_RATE_BYTES_PS ((MAX_BURST_SIZE * 1000000ULL) / 8ULL)
static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
u32 *burst_exp, u32 *burst_mantissa)
@@ -61,6 +67,9 @@ static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
*burst_mantissa = tmp / (1ULL << (*burst_exp - 7));
}
} else {
+ /* burst 0: largest encodable burst (CN10K_MAX_BURST_SIZE on
+ * CN10K), not a minimal burst.
+ */
*burst_exp = MAX_BURST_EXPONENT;
*burst_mantissa = max_mantissa;
}
@@ -1600,14 +1609,750 @@ static int otx2_setup_tc_block(struct net_device *netdev,
nic, nic, ingress);
}
+/* Free the per-queue min/max rate caches. */
+static void otx2_mqprio_free_cache(struct otx2_nic *pfvf)
+{
+ devm_kfree(pfvf->dev, pfvf->mqprio.min_rate);
+ devm_kfree(pfvf->dev, pfvf->mqprio.max_rate);
+ pfvf->mqprio.min_rate = NULL;
+ pfvf->mqprio.max_rate = NULL;
+ pfvf->mqprio.flags = 0;
+}
+
+static int otx2_mqprio_alloc_cache(struct otx2_nic *pfvf, bool replacing)
+{
+ u16 num_txq = pfvf->hw.non_qos_queues;
+
+ if (replacing && pfvf->mqprio.min_rate && pfvf->mqprio.max_rate) {
+ memset(pfvf->mqprio.min_rate, 0,
+ num_txq * sizeof(*pfvf->mqprio.min_rate));
+ memset(pfvf->mqprio.max_rate, 0,
+ num_txq * sizeof(*pfvf->mqprio.max_rate));
+ pfvf->mqprio.flags = 0;
+ return 0;
+ }
+
+ otx2_mqprio_free_cache(pfvf);
+
+ pfvf->mqprio.min_rate = devm_kcalloc(pfvf->dev, num_txq,
+ sizeof(*pfvf->mqprio.min_rate),
+ GFP_KERNEL);
+ pfvf->mqprio.max_rate = devm_kcalloc(pfvf->dev, num_txq,
+ sizeof(*pfvf->mqprio.max_rate),
+ GFP_KERNEL);
+ if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate) {
+ otx2_mqprio_free_cache(pfvf);
+ return -ENOMEM;
+ }
+
+ return 0;
+}
+
+static void otx2_mqprio_snap_free(struct otx2_nic *pfvf,
+ struct mq_offload_snap **snap)
+{
+ if (!*snap)
+ return;
+
+ devm_kfree(pfvf->dev, *snap);
+ *snap = NULL;
+}
+
+static int otx2_mqprio_snap_copy(struct otx2_nic *pfvf,
+ struct mq_offload_snap **dst,
+ const struct tc_mqprio_qopt_offload *mqprio)
+{
+ const struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ struct mq_offload_snap *snap;
+ int tc;
+
+ if (!*dst) {
+ snap = devm_kzalloc(pfvf->dev, sizeof(*snap), GFP_KERNEL);
+ if (!snap)
+ return -ENOMEM;
+ *dst = snap;
+ } else {
+ snap = *dst;
+ }
+
+ snap->num_tc = qopt->num_tc;
+ for (tc = 0; tc < TC_QOPT_MAX_QUEUE; tc++) {
+ snap->count[tc] = qopt->count[tc];
+ snap->offset[tc] = qopt->offset[tc];
+ snap->min_rate[tc] = 0;
+ snap->max_rate[tc] = 0;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ snap->min_rate[tc] = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ snap->max_rate[tc] = mqprio->max_rate[tc];
+ }
+
+ return 0;
+}
+
+static int otx2_mqprio_stage_cur(struct otx2_nic *pfvf,
+ const struct tc_mqprio_qopt_offload *mqprio)
+{
+ return otx2_mqprio_snap_copy(pfvf, &pfvf->cur_mq_snap, mqprio);
+}
+
+static void otx2_mqprio_snap_commit(struct otx2_nic *pfvf)
+{
+ otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+ pfvf->old_mq_snap = pfvf->cur_mq_snap;
+ pfvf->cur_mq_snap = NULL;
+}
+
+static void otx2_mqprio_clear_replace_state(struct otx2_nic *pfvf)
+{
+ pfvf->mqprio.replace_setup_done = false;
+ pfvf->mqprio.replace_graft_done = false;
+}
+
+static bool otx2_mqprio_mdq_allocated(struct otx2_nic *pfvf)
+{
+ return pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_MDQ] != 0;
+}
+
+static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit);
+
+static int otx2_mqprio_restore_old(struct otx2_nic *pfvf)
+{
+ struct mq_offload_snap *snap = pfvf->old_mq_snap;
+ struct net_device *netdev = pfvf->netdev;
+ u16 num_txq = pfvf->hw.non_qos_queues;
+ int tc, txq, err;
+
+ if (!snap)
+ return 0;
+
+ err = otx2_mqprio_alloc_cache(pfvf, false);
+ if (err)
+ return err;
+
+ memset(pfvf->mqprio.min_rate, 0, num_txq * sizeof(*pfvf->mqprio.min_rate));
+ memset(pfvf->mqprio.max_rate, 0, num_txq * sizeof(*pfvf->mqprio.max_rate));
+ pfvf->mqprio.flags = 0;
+
+ for (tc = 0; tc < snap->num_tc; tc++) {
+ u64 min_rate = snap->min_rate[tc];
+ u64 max_rate = snap->max_rate[tc];
+
+ if (min_rate)
+ pfvf->mqprio.flags |= TC_MQPRIO_F_MIN_RATE;
+ if (max_rate)
+ pfvf->mqprio.flags |= TC_MQPRIO_F_MAX_RATE;
+
+ for (txq = snap->offset[tc];
+ txq < snap->offset[tc] + snap->count[tc]; txq++) {
+ pfvf->mqprio.min_rate[txq] = min_rate;
+ pfvf->mqprio.max_rate[txq] = max_rate;
+ }
+ }
+
+ netdev_set_num_tc(netdev, snap->num_tc);
+ for (tc = 0; tc < snap->num_tc; tc++)
+ netdev_set_tc_queue(netdev, tc, snap->count[tc],
+ snap->offset[tc]);
+
+ if (otx2_mqprio_mdq_allocated(pfvf)) {
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+ if (err)
+ return err;
+ }
+
+ /* Rebuild the TX scheduler via netdev restart when running; otx2_mqprio_up()
+ * alone is insufficient after a failed replace that already bounced the
+ * interface. If open failed, TX schedulers were freed; defer shaper restore
+ * to the next successful ndo_open() via otx2_mqprio_up().
+ */
+ pfvf->mqprio.rate_limit = true;
+
+ if (netif_running(netdev)) {
+ err = otx2_mqprio_restart_netdev(netdev, true);
+ if (err)
+ return err;
+ } else if (pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]) {
+ err = otx2_mqprio_up(pfvf);
+ if (err)
+ return err;
+ }
+
+ otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+
+ return 0;
+}
+
+static void otx2_mqprio_snap_destroy(struct otx2_nic *pfvf)
+{
+ otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+ otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+}
+
+static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
+{
+ struct net_device *netdev = pfvf->netdev;
+
+ pfvf->mqprio.rate_limit = false;
+ otx2_mqprio_clear_replace_state(pfvf);
+ netdev_set_num_tc(netdev, 0);
+ otx2_mqprio_free_cache(pfvf);
+}
+
+/* Tear down mqprio bandwidth offload: clear per-queue shapers,
+ * mqprio_rate_limit, netdev TC mappings, and the cached rates. Called on
+ * explicit mqprio teardown (tc qdisc del) and error cleanup, not on
+ * routine netdev stop/open cycles where the offload stays active.
+ */
+int otx2_mqprio_down(struct otx2_nic *pfvf)
+{
+ int err = 0;
+
+ if (!pfvf->mqprio.rate_limit)
+ return 0;
+
+ if (netif_running(pfvf->netdev) &&
+ otx2_mqprio_mdq_allocated(pfvf))
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+
+ if (err) {
+ netdev_warn(pfvf->netdev,
+ "mqprio: failed to clear hardware shapers: %d; clearing software state\n",
+ err);
+ otx2_mqprio_clear_sw(pfvf);
+ return err;
+ }
+
+ otx2_mqprio_clear_sw(pfvf);
+
+ return 0;
+}
+
+int otx2_mqprio_up(struct otx2_nic *pfvf)
+{
+ struct net_device *netdev = pfvf->netdev;
+ int txq, err;
+
+ if (!pfvf->mqprio.rate_limit)
+ return 0;
+
+ if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate)
+ return 0;
+
+ for (txq = 0; txq < pfvf->hw.non_qos_queues; txq++) {
+ u64 min_rate = 0, max_rate = 0;
+
+ if (pfvf->mqprio.flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = pfvf->mqprio.min_rate[txq];
+ if (pfvf->mqprio.flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = pfvf->mqprio.max_rate[txq];
+
+ if (!min_rate && !max_rate)
+ continue;
+
+ err = otx2_nix_tm_set_queue_shaper(pfvf, txq, min_rate,
+ max_rate);
+ if (err) {
+ netdev_err(netdev,
+ "mqprio: failed to restore shaper for txq %d: %d\n",
+ txq, err);
+ return err;
+ }
+ }
+
+ return 0;
+}
+
+/* Restart the netdev to reprogram the TX scheduler hierarchy for mqprio
+ * bandwidth offload. Both mqprio add and delete (when offload was active)
+ * take this path via ndo_stop()/ndo_open() so VF-specific open logic (e.g.
+ * LBK carrier on) runs correctly.
+ *
+ * Intentional behaviour: this full stop/open cycle drops in-flight traffic
+ * (carrier off, IRQ/NAPI teardown, queue drain). The NIX TX scheduler must
+ * be reallocated (e.g. one SMQ per non-QoS queue) and cannot be reprogrammed
+ * live today, so a netdev bounce is required on every mqprio add, replace,
+ * delete, and rollback. Users see a brief connectivity blip; this is not a
+ * bug to "fix" without implementing the live-reprogramming path noted below.
+ * If open fails, the interface is left administratively down without calling
+ * ndo_stop() again on resources already torn down by the open error path.
+ *
+ * Quiesce the transmit path like __dev_close_many() before ndo_stop() so
+ * otx2_xmit() cannot race otx2_free_queue_mem().
+ */
+static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit)
+{
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ const struct net_device_ops *ops = netdev->netdev_ops;
+ bool running = netif_running(netdev);
+ int err;
+
+ /* TODO: Explore live TX scheduler reprogramming to avoid a full
+ * ndo_stop()/ndo_open() bounce on every mqprio change.
+ */
+ netdev_info(netdev,
+ "mqprio: restarting interface to reprogram TX scheduler; in-flight traffic will be dropped\n");
+
+ if (running) {
+ clear_bit(__LINK_STATE_START, &netdev->state);
+ smp_mb__after_atomic(); /* Commit netif_running(). */
+ }
+ dev_deactivate(netdev, true);
+
+ err = ops->ndo_stop(netdev);
+ if (err) {
+ if (running) {
+ set_bit(__LINK_STATE_START, &netdev->state);
+ dev_activate(netdev);
+ }
+ return err;
+ }
+
+ /* Set before ndo_open() so otx2_txsch_alloc() widens SMQ allocation.
+ * On teardown, drop mqprio software state so ndo_open() does not
+ * re-apply bandwidth limits via otx2_mqprio_up() after the kernel
+ * removed the qdisc.
+ */
+ if (rate_limit)
+ pfvf->mqprio.rate_limit = true;
+ else
+ otx2_mqprio_clear_sw(pfvf);
+
+ err = ops->ndo_open(netdev);
+ if (!err && running) {
+ set_bit(__LINK_STATE_START, &netdev->state);
+ dev_activate(netdev);
+ } else if (err) {
+ netdev_err(netdev,
+ "Failed to restart device after mqprio change: %d\n",
+ err);
+ /* ndo_open() already freed the TX schedulers on failure while
+ * netif_running() may still be true; drop mqprio software state
+ * only instead of sending shaper clears to freed queues.
+ */
+ otx2_mqprio_clear_sw(pfvf);
+ /* ndo_open() rolls back on failure; mark the interface down so
+ * netif_close() does not invoke ndo_stop() on freed NAPI/queue
+ * state. Caller holds RTNL; dev_close() would deadlock.
+ */
+ otx2_set_flag(pfvf, OTX2_FLAG_INTF_DOWN);
+ /* visible to otx2_stop() on other cpus */
+ smp_wmb();
+ netif_close(netdev);
+ }
+
+ return err;
+}
+
+static int otx2_mqprio_validate_tc_rate(struct net_device *netdev,
+ struct netlink_ext_ack *extack,
+ u64 rate, u32 qcount, int tc,
+ const char *name)
+{
+ if (!rate)
+ return 0;
+
+ if (qcount <= 1)
+ return 0;
+
+ /* TODO: mqprio min_rate/max_rate are per traffic class, but bandwidth
+ * offload shapes on per-queue MDQ nodes parented under a single TL4.
+ * Without per-TC TL4 shapers the driver cannot honor TC-level limits
+ * for a traffic class that spans multiple queues without either
+ * dividing the rate across queues (uAPI mismatch) or exceeding the TC
+ * cap when every member queue is active. Reject until per-TC TL4
+ * shaping can be implemented without allocating additional TL4 nodes
+ * beyond the existing hierarchy.
+ */
+ netdev_err(netdev,
+ "mqprio: %s rate for tc %d not supported with %u queues\n",
+ name, tc, qcount);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: %s rate for tc %d not supported with %u queues",
+ name, tc, qcount);
+ return -EOPNOTSUPP;
+}
+
+static int otx2_mqprio_validate_txqs(struct net_device *netdev,
+ struct netlink_ext_ack *extack,
+ struct tc_mqprio_qopt *qopt)
+{
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ u16 num_txq = pfvf->hw.non_qos_queues;
+ int tc, txq;
+
+ if (qopt->num_tc > num_txq) {
+ netdev_err(netdev, "Number of TCs (%u) exceeds hw queues %u\n",
+ qopt->num_tc, num_txq);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "Number of TCs (%u) exceeds hw queues %u",
+ qopt->num_tc, num_txq);
+ return -EINVAL;
+ }
+
+ if (num_txq > MAX_TXSCHQ_PER_FUNC) {
+ netdev_err(netdev,
+ "Number of queues (%u) exceeds max scheduler queues %u\n",
+ num_txq, MAX_TXSCHQ_PER_FUNC);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "Number of queues (%u) exceeds max scheduler queues %u",
+ num_txq, MAX_TXSCHQ_PER_FUNC);
+ return -EINVAL;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u32 qcount = qopt->count[tc];
+
+ for (txq = qopt->offset[tc];
+ txq < qopt->offset[tc] + qcount; txq++) {
+ if (txq >= num_txq) {
+ netdev_err(netdev,
+ "mqprio: txq %d exceeds offload queue count %u\n",
+ txq, num_txq);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: txq %d exceeds offload queue count %u",
+ txq, num_txq);
+ return -EINVAL;
+ }
+ }
+ }
+
+ return 0;
+}
+
+static bool otx2_mqprio_rate_valid(u64 rate_bytes_ps)
+{
+ u64 mbps;
+
+ if (!rate_bytes_ps)
+ return true;
+
+ if (rate_bytes_ps < OTX2_MQPRIO_MIN_RATE_BYTES_PS)
+ return false;
+
+ if (rate_bytes_ps > OTX2_MQPRIO_MAX_RATE_BYTES_PS)
+ return false;
+
+ if (rate_bytes_ps > div_u64(U64_MAX, 8))
+ return false;
+
+ mbps = otx2_convert_rate(rate_bytes_ps);
+ return ilog2(mbps / 2) <= MAX_RATE_EXPONENT;
+}
+
+static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
+ struct tc_mqprio_qopt_offload *mqprio)
+{
+ struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ bool had_mqprio = pfvf->mqprio.rate_limit;
+ struct net_device *netdev = pfvf->netdev;
+ bool if_up = netif_running(netdev);
+ int err;
+
+ qopt->hw = 0;
+
+ /* tc qdisc replace runs setup on the new mqprio before destroying the
+ * old one. replace_setup_done and TC_ROOT_GRAFT distinguish stale
+ * old-instance teardown from graft failure after setup.
+ */
+ if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
+ err = 0;
+ if (pfvf->mqprio.replace_graft_done)
+ otx2_mqprio_snap_commit(pfvf);
+ else
+ err = otx2_mqprio_restore_old(pfvf);
+ otx2_mqprio_clear_replace_state(pfvf);
+ return err;
+ }
+
+ /* Skip the netdev restart when mqprio offload was not active. */
+ if (!had_mqprio)
+ return 0;
+
+ if (if_up) {
+ int down_err, err;
+
+ down_err = otx2_mqprio_down(pfvf);
+ err = otx2_mqprio_restart_netdev(netdev, false);
+ if (err)
+ return err;
+ return down_err;
+ }
+
+ /* ndo_stop() already freed the TX scheduler TL nodes; drop software
+ * state only.
+ */
+ otx2_mqprio_clear_sw(pfvf);
+ return 0;
+}
+
+static int otx2_setup_tc_mqprio(struct net_device *netdev,
+ struct tc_mqprio_qopt_offload *mqprio)
+{
+ struct netlink_ext_ack *extack = mqprio->extack;
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ bool replacing = pfvf->mqprio.rate_limit;
+ bool if_up = netif_running(netdev);
+ int tc, txq, err, i;
+
+ if (!qopt->hw)
+ return otx2_teardown_tc_mqprio(pfvf, mqprio);
+
+ if (!if_up) {
+ netdev_err(netdev, "mqprio: setup requires interface UP\n");
+ NL_SET_ERR_MSG_MOD(extack, "mqprio: setup requires interface UP");
+ return -EOPNOTSUPP;
+ }
+
+ if (mqprio->shaper != TC_MQPRIO_SHAPER_BW_RATE) {
+ netdev_err(netdev, "Unsupported mqprio shaper %#x\n", mqprio->shaper);
+ NL_SET_ERR_MSG_FMT_MOD(extack, "Unsupported mqprio shaper %#x",
+ mqprio->shaper);
+ return -EOPNOTSUPP;
+ }
+
+ if (!test_bit(QOS_CIR_PIR_SUPPORT, &pfvf->hw.cap_flag)) {
+ netdev_err(netdev,
+ "mqprio: bandwidth offload requires CIR+PIR support\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: bandwidth offload requires CIR+PIR support");
+ return -EOPNOTSUPP;
+ }
+
+ if (is_otx2_sdp_rep(pfvf->pdev)) {
+ netdev_err(netdev, "mqprio: bandwidth offload not supported on SDP rep\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: bandwidth offload not supported on SDP rep");
+ return -EOPNOTSUPP;
+ }
+
+ if (pfvf->pfc_en) {
+ netdev_err(netdev,
+ "mqprio: cannot enable offload while PFC is enabled\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: cannot enable offload while PFC is enabled");
+ return -EOPNOTSUPP;
+ }
+
+ if (pfvf->xdp_prog) {
+ netdev_err(netdev,
+ "mqprio: cannot enable offload while XDP is active\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: cannot enable offload while XDP is active");
+ return -EOPNOTSUPP;
+ }
+
+ if (!list_empty(&pfvf->qos.qos_tree)) {
+ netdev_err(netdev,
+ "mqprio: cannot enable offload while HTB is active\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: cannot enable offload while HTB is active");
+ return -EOPNOTSUPP;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u64 min_rate = 0, max_rate = 0;
+ u32 qcount = qopt->count[tc];
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = mqprio->max_rate[tc];
+
+ if (min_rate && max_rate && min_rate > max_rate) {
+ netdev_err(netdev,
+ "min_rate %llu exceeds max_rate %llu for tc %d\n",
+ min_rate, max_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "min_rate %llu exceeds max_rate %llu for tc %d",
+ min_rate, max_rate, tc);
+ return -EINVAL;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE) {
+ err = otx2_mqprio_validate_tc_rate(netdev, extack, min_rate,
+ qcount, tc, "min");
+ if (err)
+ return err;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE) {
+ err = otx2_mqprio_validate_tc_rate(netdev, extack, max_rate,
+ qcount, tc, "max");
+ if (err)
+ return err;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE &&
+ !otx2_mqprio_rate_valid(min_rate)) {
+ netdev_err(netdev,
+ "mqprio: min_rate %llu for tc %d is outside hardware limits\n",
+ min_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: min_rate %llu for tc %d is outside hardware limits",
+ min_rate, tc);
+ return -EINVAL;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE &&
+ !otx2_mqprio_rate_valid(max_rate)) {
+ netdev_err(netdev,
+ "mqprio: max_rate %llu for tc %d is outside hardware limits\n",
+ max_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: max_rate %llu for tc %d is outside hardware limits",
+ max_rate, tc);
+ return -EINVAL;
+ }
+ }
+
+ err = otx2_mqprio_validate_txqs(netdev, extack, qopt);
+ if (err)
+ return err;
+
+ err = otx2_mqprio_stage_cur(pfvf, mqprio);
+ if (err)
+ return err;
+
+ err = otx2_mqprio_restart_netdev(pfvf->netdev, true);
+ if (err)
+ goto cleanup;
+
+ err = otx2_mqprio_alloc_cache(pfvf, replacing);
+ if (err)
+ goto cleanup;
+
+ /* otx2_mqprio_up() may have restored the previous configuration during
+ * the restart above. Clear every MDQ shaper before applying the new
+ * mapping so queues dropped from the TC layout do not keep stale
+ * limits in hardware.
+ */
+ if (otx2_mqprio_mdq_allocated(pfvf)) {
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+ if (err)
+ goto cleanup;
+ }
+
+ pfvf->mqprio.flags = mqprio->flags;
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u64 min_rate = 0, max_rate = 0;
+ u32 qcount = qopt->count[tc];
+
+ /* Rates omitted from tc mqprio are passed as zero and both MDQ
+ * shaper registers are programmed; see
+ * otx2_nix_tm_set_queue_shaper(). Multi-queue TCs with rates
+ * are rejected above.
+ */
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = mqprio->max_rate[tc];
+
+ for (txq = qopt->offset[tc];
+ txq < qopt->offset[tc] + qcount; txq++) {
+ netdev_dbg(netdev,
+ "mqprio: tc %d txq %d min_rate %llu max_rate %llu\n",
+ tc, txq, min_rate, max_rate);
+
+ pfvf->mqprio.min_rate[txq] = min_rate;
+ pfvf->mqprio.max_rate[txq] = max_rate;
+
+ err = otx2_nix_tm_set_queue_shaper(pfvf, txq,
+ min_rate, max_rate);
+ if (err)
+ goto cleanup;
+ }
+ }
+
+ netdev_set_num_tc(netdev, pfvf->cur_mq_snap->num_tc);
+ for (i = 0; i < pfvf->cur_mq_snap->num_tc; i++)
+ netdev_set_tc_queue(netdev, i, pfvf->cur_mq_snap->count[i],
+ qopt->offset[i]);
+
+ qopt->hw = TC_MQPRIO_HW_OFFLOAD_TCS;
+
+ if (replacing) {
+ pfvf->mqprio.replace_setup_done = true;
+ pfvf->mqprio.replace_graft_done = false;
+ } else {
+ otx2_mqprio_snap_commit(pfvf);
+ }
+
+ return 0;
+
+cleanup:
+ qopt->hw = 0;
+ if (replacing) {
+ int restore_err = otx2_mqprio_restore_old(pfvf);
+
+ otx2_mqprio_clear_replace_state(pfvf);
+ if (restore_err) {
+ netdev_err(netdev,
+ "mqprio: replace failed and prior configuration rollback failed: %d\n",
+ restore_err);
+ if (extack)
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: replace failed and prior configuration rollback failed: %d",
+ restore_err);
+ } else {
+ netdev_err(netdev,
+ "mqprio: replace failed; prior configuration restored\n");
+ if (extack)
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: replace failed; prior configuration restored");
+ }
+ return err ? err : -EIO;
+ }
+ otx2_teardown_tc_mqprio(pfvf, mqprio);
+ return err;
+}
+
+static int otx2_setup_tc_root(struct otx2_nic *pfvf,
+ struct tc_root_qopt_offload *root)
+{
+ switch (root->command) {
+ case TC_ROOT_GRAFT:
+ if (pfvf->mqprio.replace_setup_done)
+ pfvf->mqprio.replace_graft_done = true;
+ return 0;
+ default:
+ return -EOPNOTSUPP;
+ }
+}
+
+static int otx2_setup_tc_query_caps(void *type_data)
+{
+ struct tc_query_caps_base *base = type_data;
+ struct tc_mqprio_caps *caps;
+
+ if (base->type != TC_SETUP_QDISC_MQPRIO)
+ return -EOPNOTSUPP;
+
+ caps = base->caps;
+ caps->validate_queue_counts = true;
+
+ return 0;
+}
+
int otx2_setup_tc(struct net_device *netdev, enum tc_setup_type type,
void *type_data)
{
switch (type) {
+ case TC_QUERY_CAPS:
+ return otx2_setup_tc_query_caps(type_data);
case TC_SETUP_BLOCK:
return otx2_setup_tc_block(netdev, type_data);
case TC_SETUP_QDISC_HTB:
return otx2_setup_tc_htb(netdev, type_data);
+ case TC_SETUP_QDISC_MQPRIO:
+ return otx2_setup_tc_mqprio(netdev, type_data);
+ case TC_SETUP_ROOT_QDISC:
+ return otx2_setup_tc_root(netdev_priv(netdev), type_data);
default:
return -EOPNOTSUPP;
}
@@ -1632,6 +2377,7 @@ EXPORT_SYMBOL(otx2_init_tc);
void otx2_shutdown_tc(struct otx2_nic *nic)
{
otx2_destroy_tc_flow_list(nic);
+ otx2_mqprio_snap_destroy(nic);
}
EXPORT_SYMBOL(otx2_shutdown_tc);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
index f160b1618efa..9ef55a6db50b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
@@ -118,6 +118,9 @@ static void otx2_config_sched_shaping(struct otx2_nic *pfvf,
/* configure PIR */
maxrate = (node->rate > node->ceil) ? node->rate : node->ceil;
+ /* 65536 is the kernel-side default burst when HTB does not supply an
+ * explicit value, not the NIX hardware maximum (CN10K_MAX_BURST_SIZE).
+ */
cfg->regval[*num_regs] =
otx2_get_txschq_rate_regval(pfvf, maxrate, 65536);
(*num_regs)++;
@@ -1088,6 +1091,14 @@ static int otx2_qos_root_add(struct otx2_nic *pfvf, u16 htb_maj_id, u16 htb_defc
"TC_HTB_CREATE: handle=0x%x defcls=0x%x\n",
htb_maj_id, htb_defcls);
+ if (pfvf->mqprio.rate_limit) {
+ netdev_err(pfvf->netdev,
+ "HTB: cannot enable while mqprio bandwidth offload is active\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "HTB: cannot enable while mqprio bandwidth offload is active");
+ return -EOPNOTSUPP;
+ }
+
root = otx2_qos_alloc_root(pfvf);
if (IS_ERR(root)) {
err = PTR_ERR(root);
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags
2026-09-11 10:55 ` [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
@ 2026-09-17 11:31 ` Paolo Abeni
0 siblings, 0 replies; 5+ messages in thread
From: Paolo Abeni @ 2026-09-17 11:31 UTC (permalink / raw)
To: rkannoth
Cc: bpf, linux-kernel, netdev, andrew+netdev, ast, daniel, davem,
edumazet, hawk, john.fastabend, kuba, pabeni, sdf, sgoutham
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
Full review at: https://netdev-ai.bots.linux.dev/sashiko/
---
octeontx2: use atomic bitops for PF/VF and rep flags
The patch converts struct otx2_nic::flags and struct rep_dev::flags from
u64 bitmasks to unsigned long bitmaps accessed with set_bit/clear_bit/
test_bit, adds otx2_set_flag()/otx2_clear_flag()/otx2_test_flag() helpers,
and replaces the representor to PF flag publication with a new
otx2_sync_flags_from_rep() helper. RVU_REP_VF_INITIALIZED is folded into
the shared flag namespace as OTX2_REP_VF_INITIALIZED (bit 21).
A few questions below, mostly around the new sync helper and around
pre-existing issues in the representor paths that the patch touches.
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> index eecee612b7b2c..7e09c1444a6d5 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> @@ -491,28 +491,29 @@ struct otx2_nic {
> u16 tx_max_pktlen;
> u16 rbsize; /* Receive buffer size */
>
> -#define OTX2_FLAG_RX_TSTAMP_ENABLED BIT_ULL(0)
[ ... ]
> +#define OTX2_FLAG_TC_MARK_ENABLED 17
> +#define OTX2_FLAG_REP_MODE_ENABLED 18
> +#define OTX2_FLAG_PORT_UP 19
> +#define OTX2_FLAG_IPSEC_OFFLOAD_ENABLED 20
> +#define OTX2_REP_VF_INITIALIZED 21
> + unsigned long flags;
> u64 *cq_op_addr;
>
> struct bpf_prog *xdp_prog;
> @@ -594,6 +595,34 @@ struct otx2_nic {
> unsigned long *af_xdp_zc_qidx;
> };
>
> +static inline void otx2_set_flag(struct otx2_nic *nic, unsigned int flag)
> +{
> + set_bit(flag, &nic->flags);
> +}
> +
> +static inline void otx2_clear_flag(struct otx2_nic *nic, unsigned int flag)
> +{
> + clear_bit(flag, &nic->flags);
> +}
> +
> +static inline bool otx2_test_flag(struct otx2_nic *nic, unsigned int flag)
> +{
> + return test_bit(flag, &nic->flags);
> +}
> +
> +static inline void otx2_sync_flags_from_rep(struct otx2_nic *dst,
> + unsigned long *src_flags)
> +{
> + unsigned int flag;
> +
> + for (flag = 0; flag <= OTX2_REP_VF_INITIALIZED; flag++) {
> + if (test_bit(flag, src_flags))
> + set_bit(flag, &dst->flags);
> + else
> + clear_bit(flag, &dst->flags);
> + }
> +}
> +
[Severity: Critical]
This isn't a bug introduced by this patch, since the previous code did
priv->flags = rep->flags, but the new helper codifies the behaviour, so this
looks like the natural place to fix it.
Which bits can rep->flags ever hold? The only writers are
rvu_rep_mcam_flow_init():
if (allocated) {
set_bit(OTX2_FLAG_MCAM_ENTRIES_ALLOC, &rep->flags);
set_bit(OTX2_FLAG_NTUPLE_SUPPORT, &rep->flags);
set_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags);
}
and rvu_rep_state_evt_handler(), which sets or clears
OTX2_REP_VF_INITIALIZED.
Given that, does the else branch above erase PF owned state from the shared
eswitch context (priv is rep->mdev) on the very first tc filter add on any
representor, namely OTX2_FLAG_REP_MODE_ENABLED, OTX2_FLAG_INTF_DOWN,
OTX2_FLAG_PORT_UP and OTX2_FLAG_TC_MARK_ENABLED?
OTX2_FLAG_REP_MODE_ENABLED is only ever set in rvu_rep_probe(), so is there
anything that restores it for the remaining life of the PCI device?
With that bit gone, otx2_tx_napi_handler() takes the non representor path:
if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
ndev = pfvf->reps[qidx]->netdev;
else
ndev = pfvf->netdev;
...
if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
qidx = 0;
txq = netdev_get_tx_queue(ndev, qidx);
netdev_tx_completed_queue(txq, tx_pkts, tx_bytes);
but rvu_rep_setup_tc_cb() has just stored priv->netdev = rep->netdev, and
rvu_rep_create() allocates every representor with alloc_etherdev(), i.e.
num_tx_queues is 1. Can netdev_get_tx_queue(ndev, qidx) with qidx up to
rep_cnt - 1 then index past dev->_tx[0], so that
netdev_tx_completed_queue() and netif_tx_wake_queue() write outside the
netdev_queue array?
Two further effects of the same sweep: otx2_rcv_pkt_handler() stops doing
skb->mark = parse->match_id, because OTX2_FLAG_TC_MARK_ENABLED is only ever
produced in priv by otx2_tc_parse_actions() and is never mirrored back into
rep->flags, and clearing OTX2_FLAG_INTF_DOWN removes the guard in
otx2_napi_handler() that keeps NIX_LF_CINTX_ENA_W1S from re-arming CINTs
that rvu_rep_free_cq_rsrc() has just disabled.
Would copying only the representor owned bits (MCAM_ENTRIES_ALLOC,
NTUPLE_SUPPORT, TC_FLOWER_SUPPORT, REP_VF_INITIALIZED) and leaving the PF
owned bits untouched be a better fit here?
Separately, the loop bound hard-codes the current highest flag, so a bit 22
added later silently stops being synced. Could a BUILD_BUG_ON or an
explicit mask be used instead?
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> index 039fd47ebf522..ddb46b580c3b6 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
[ ... ]
> @@ -1195,12 +1195,12 @@ static int otx2_tc_del_flow(struct otx2_nic *nic,
> /* Disable TC MARK flag if they are no rules with skbedit mark action */
> if (flow_node->req.match_id)
> if (!refcount_dec_and_test(&flow_cfg->mark_flows))
> - nic->flags &= ~OTX2_FLAG_TC_MARK_ENABLED;
> + otx2_clear_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
[Severity: Medium]
This is a pre-existing issue and not introduced by this patch, but the
patch rewrites the statement, so a note here.
Is the sense of the test inverted? mark_flows is initialised with
refcount_set(&flow_cfg->mark_flows, 1) in otx2_mcam_entry_init() and is
incremented once per rule in otx2_tc_parse_actions():
otx2_set_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
refcount_inc(&nic->flow_cfg->mark_flows);
With two mark rules installed the counter is 3. Deleting one leaves 2, so
refcount_dec_and_test() returns false and OTX2_FLAG_TC_MARK_ENABLED is
cleared while a mark rule is still programmed in the MCAM.
otx2_rcv_pkt_handler() then stops executing skb->mark = parse->match_id for
the surviving rule.
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
> index 8d2d607bc92f9..f65ba44db60b6 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
[ ... ]
> @@ -513,7 +513,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
> ((u64)cq->cq_idx << 32) | processed_cqe);
>
> #if IS_ENABLED(CONFIG_RVU_ESWITCH)
> - if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
> + if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
> ndev = pfvf->reps[qidx]->netdev;
> else
> #endif
> @@ -526,7 +526,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
>
> if (qidx >= pfvf->hw.tx_queues)
> qidx -= pfvf->hw.xdp_queues;
> - if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
> + if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
> qidx = 0;
> txq = netdev_get_tx_queue(ndev, qidx);
> netdev_tx_completed_queue(txq, tx_pkts, tx_bytes);
[Severity: Medium]
This function reads OTX2_FLAG_REP_MODE_ENABLED twice, once to pick ndev and
once to clamp qidx. The publication of these flags used to be a single word
store, priv->flags = rep->flags, and is now 22 independent set_bit() and
clear_bit() calls in otx2_sync_flags_from_rep(), run from process context in
rvu_rep_setup_tc_cb() with no lock or RCU against the NAPI readers that use
the same otx2_nic (rvu_rep_napi_init() sets cq_poll->dev = priv).
Can a poll running concurrently with the sync observe a combination that
neither the old nor the new state contains, for example ndev =
pfvf->reps[qidx]->netdev from the first test and an unclamped qidx from the
second?
The unconditional clear_bit() in the else branch also discards bits another
context set moments earlier, for instance rvu_rep_napi_init() clearing
INTF_DOWN, otx2_mbox_up_handler_rep_event_up_notify() setting PORT_UP, or
otx2_tc_parse_actions() setting TC_MARK_ENABLED, with no way for those
writers to notice the loss. For a change whose stated purpose is to make
flag access race free, would restricting the sync to the representor owned
bits keep the single publication step intact?
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
> index 0e8a6a6486c4a..7808588a0234a 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
[ ... ]
> @@ -172,7 +172,7 @@ int otx2_xsk_wakeup(struct net_device *dev, u32 queue_id, u32 flags)
> struct otx2_cq_poll *cq_poll = NULL;
> struct otx2_qset *qset = &pf->qset;
>
> - if (pf->flags & OTX2_FLAG_INTF_DOWN)
> + if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
> return -ENETDOWN;
>
> if (queue_id >= pf->hw.rx_queues || queue_id >= pf->hw.tx_queues)
[Severity: High]
This is a pre-existing issue and not introduced by this patch, but does
converting the gate to test_bit() close the window here?
OTX2_FLAG_INTF_DOWN stays an ad-hoc lockless gate, and nothing protects
qset->napi between this test and its use further down the same function:
cq_poll = &qset->napi[queue_id];
if (!cq_poll)
return -EINVAL;
if (!napi_if_scheduled_mark_missed(&cq_poll->napi)) {
ndo_xsk_wakeup is reached from xsk_sendmsg() and xsk_poll() without RTNL, so
can this sequence happen?
CPU0 otx2_xsk_wakeup()
otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN) == false, continues
CPU1 otx2_stop()
otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
smp_wmb();
... synchronize_irq()/napi_synchronize() ...
otx2_free_queue_mem(qset)
kfree(qset->napi);
qset->napi = NULL;
CPU0 then computes &qset->napi[queue_id], and the !cq_poll test only catches
queue_id 0 when the pointer is NULL, so does napi_if_scheduled_mark_missed()
dereference freed or near NULL memory, followed by an MMIO write derived
from cq_poll->cint_idx? synchronize_irq() and napi_synchronize() do not
wait for a process context wakeup caller that already passed the flag test.
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> index 0f5d5642d3f7a..8a8c0088fd204 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> @@ -93,9 +93,9 @@ static int rvu_rep_mcam_flow_init(struct rep_dev *rep)
> rep->flow_cfg->max_flows = allocated;
>
> if (allocated) {
> - rep->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
> - rep->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
> - rep->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
> + set_bit(OTX2_FLAG_MCAM_ENTRIES_ALLOC, &rep->flags);
> + set_bit(OTX2_FLAG_NTUPLE_SUPPORT, &rep->flags);
> + set_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags);
> }
>
> INIT_LIST_HEAD(&rep->flow_cfg->flow_list);
> @@ -109,14 +109,14 @@ static int rvu_rep_setup_tc_cb(enum tc_setup_type type,
> struct rep_dev *rep = cb_priv;
> struct otx2_nic *priv = rep->mdev;
>
> - if (!(rep->flags & RVU_REP_VF_INITIALIZED))
> + if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
> return -EINVAL;
>
> - if (!(rep->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
> + if (!test_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags))
> rvu_rep_mcam_flow_init(rep);
[Severity: Medium]
This is a pre-existing issue and not introduced by this patch. The return
value of rvu_rep_mcam_flow_init() is discarded, and that function
unconditionally re-assigns both rep->flow_cfg and rep->flow_cfg->flow_ent:
rep->flow_cfg = kzalloc_objs(struct otx2_flow_config, 1);
...
rep->flow_cfg->flow_ent = kcalloc(count, sizeof(u16), GFP_KERNEL);
It only sets the support bits when allocated is non zero, so if the AF
mailbox allocation yields zero entries, or the flow_ent kcalloc fails and
-ENOMEM is returned, does OTX2_FLAG_TC_FLOWER_SUPPORT stay clear so that
every subsequent tc request re-enters the function and leaks the previous
otx2_flow_config plus its flow_ent array?
Also, rvu_rep_destroy() only does kfree(rep->flow_cfg) and never frees
flow_cfg->flow_ent, so does the u16 entry array leak on every representor
teardown even on the success path?
> priv->netdev = rep->netdev;
> - priv->flags = rep->flags;
> + otx2_sync_flags_from_rep(priv, &rep->flags);
> priv->pcifunc = rep->pcifunc;
> priv->flow_cfg = rep->flow_cfg;
[Severity: High]
This is a pre-existing issue and not introduced by this patch, but the new
sync helper is part of the same publication step.
Representor owned objects are published into priv here, and priv is the
devm allocated otx2_nic that lives for the whole PCI device lifetime.
rvu_rep_destroy() later frees them:
unregister_netdev(rep->netdev);
rvu_rep_devlink_port_unregister(rep);
free_netdev(rep->netdev);
kfree(rep->flow_cfg);
without clearing priv->netdev or priv->flow_cfg. After
devlink dev eswitch set <dev> mode legacy followed by mode switchdev,
rvu_rep_napi_init() clears INTF_DOWN and re-enables NAPI, and since this tc
callback also cleared OTX2_FLAG_REP_MODE_ENABLED, otx2_rcv_pkt_handler()
reads pfvf->netdev->features and otx2_tx_napi_handler() uses
ndev = pfvf->netdev for netdev_get_tx_queue(), netdev_tx_completed_queue()
and netif_tx_wake_queue().
Can that dereference the freed representor netdev, i.e. a use after free
read and write from the NAPI datapath? The stale priv->flow_cfg is
similarly still advertised as valid by the MCAM/NTUPLE/TC_FLOWER bits the
sync copied into priv->flags.
[ ... ]
> @@ -382,7 +382,7 @@ static void rvu_rep_get_stats64(struct net_device *dev,
> {
> struct rep_dev *rep = netdev_priv(dev);
>
> - if (!(rep->flags & RVU_REP_VF_INITIALIZED))
> + if (!test_bit(OTX2_REP_VF_INITIALIZED, &rep->flags))
> return;
>
> stats->rx_packets = rep->stats.rx_frames;
[Severity: High]
This is a pre-existing issue and not introduced by this patch. The tail of
this function arms a delayed work on every statistics read:
schedule_delayed_work(&rep->stats_wrk, msecs_to_jiffies(100));
and the callback rvu_rep_get_stats() recovers rep with
container_of(del_work, struct rep_dev, stats_wrk) and then dereferences
rep->mdev and takes priv->mbox.lock.
struct rep_dev is the private area of an alloc_etherdev() device, and
neither rvu_rep_destroy() nor the post registration error path in
rvu_rep_create() calls cancel_delayed_work_sync(&rep->stats_wrk) before
free_netdev(rep->netdev). unregister_netdev() does not cancel driver owned
delayed work.
If ip -s link show on a representor is followed within 100 ms by
devlink dev eswitch set <dev> mode legacy, can the pending work run against
the freed netdev allocation?
[ ... ]
> @@ -632,7 +632,7 @@ void rvu_rep_destroy(struct otx2_nic *priv)
> int rep_id;
>
> rvu_eswitch_config(priv, false);
> - priv->flags |= OTX2_FLAG_INTF_DOWN;
> + otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
> rvu_rep_free_cq_rsrc(priv);
> for (rep_id = 0; rep_id < priv->rep_cnt; rep_id++) {
> rep = priv->reps[rep_id];
[Severity: High]
This is a pre-existing issue and not introduced by this patch, but it is in
the function being modified. A little further down this loop:
free_netdev(rep->netdev);
kfree(rep->flow_cfg);
rep is netdev_priv() of a device allocated with alloc_etherdev() in
rvu_rep_create(), so free_netdev() releases the allocation that contains
rep. Does the following statement then read rep->flow_cfg out of freed
memory and hand whatever it finds to kfree()?
Saving flow_cfg in a local before free_netdev() would avoid the stale read,
and flow_cfg->flow_ent allocated by rvu_rep_mcam_flow_init() still needs a
matching kfree().
[Severity: Medium]
Also pre-existing and not introduced by this patch, on the creation side of
this pair: rvu_rep_create() returns directly if the representor array
allocation fails, right after resources were acquired:
err = rvu_rep_rsrc_init(priv);
if (err)
return -ENOMEM;
priv->reps = kzalloc_objs(struct rep_dev *, rep_cnt);
if (!priv->reps)
return -ENOMEM;
Every later failure jumps to the exit label which calls
rvu_rep_rsrc_free(). Does this second return leak the queue memory from
otx2_alloc_queue_mem() and the NIX/NPA contexts set up by
otx2_init_hw_resources()?
[ ... ]
--
This is an AI-generated review.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
2026-09-11 10:55 ` [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
@ 2026-09-17 11:31 ` Paolo Abeni
0 siblings, 0 replies; 5+ messages in thread
From: Paolo Abeni @ 2026-09-17 11:31 UTC (permalink / raw)
To: rkannoth
Cc: bpf, linux-kernel, netdev, andrew+netdev, ast, daniel, davem,
edumazet, hawk, john.fastabend, kuba, pabeni, sdf, sgoutham
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
Full review at: https://netdev-ai.bots.linux.dev/sashiko/
---
octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
This adds TC_SETUP_QDISC_MQPRIO channel-mode offload with
TC_MQPRIO_SHAPER_BW_RATE, programming per-queue MDQ CIR/PIR through the
NIX TX scheduler mailbox. Since the scheduler hierarchy cannot be
reprogrammed live, every mqprio add, replace, delete and rollback bounces
the netdev through ndo_stop()/ndo_open(), with the TC layout and rates
tracked in mq_offload_snap snapshots for rollback.
A few questions on the commit message first.
> Use atomic bit operations when updating OTX2_FLAG_PORT_UP and
> OTX2_FLAG_INTF_DOWN on asynchronous and netdev-restart paths.
>
> Add a ratelimited AF debug message when validating TX scheduler queue
> ownership to aid mqprio hierarchy setup failures.
[Severity: Low]
The first paragraph appears to describe the immediately preceding commit
("octeontx2: use atomic bitops for PF/VF and rep flags"), which is what
converted the flags field to bit numbers and introduced otx2_set_flag().
The only new flag write added here is the
otx2_set_flag(pfvf, OTX2_FLAG_INTF_DOWN) call in
otx2_mqprio_restart_netdev(), which simply inherits that conversion.
Should this paragraph be dropped?
For the second paragraph, the AF change adds dev_err_ratelimited() in
is_valid_txschq(), which is an error-level print rather than a debug
message, and it is reachable from VF mailbox requests. Could the wording
match the log level used, or should the print be pr_debug/dev_dbg level?
> diff --git a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
> index d80d2c00bd843..2d1ab8c761c40 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
> @@ -331,8 +331,12 @@ static bool is_valid_txschq(struct rvu *rvu, int blkaddr,
> return true;
> }
>
> - if (map_func != pcifunc)
> + if (map_func != pcifunc) {
> + dev_err_ratelimited(rvu->dev,
> + "pcifunc %x map pcifunc %x not equal, lvl=%u schq=%u\n",
> + pcifunc, map_func, lvl, schq);
> return false;
> + }
>
> return true;
> }
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> index 7e09c1444a6d5..6e7132f27b621 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> @@ -483,6 +484,23 @@ struct pf_irq_data {
> int mdevs;
> };
>
> +struct mq_offload_snap {
> + u64 min_rate[TC_QOPT_MAX_QUEUE];
> + u64 max_rate[TC_QOPT_MAX_QUEUE];
> + __u8 num_tc;
> + __u16 count[TC_QOPT_MAX_QUEUE];
> + __u16 offset[TC_QOPT_MAX_QUEUE];
> +};
[Severity: Medium]
Should the snapshot also record prio_tc_map[]?
sch_mqprio installs three pieces of netdev state for an offloaded mqprio,
and the priority map is applied from the new qopt after the driver
callback returns success:
net/sched/sch_mqprio.c:mqprio_init() {
...
/* Always use supplied priority mappings */
for (i = 0; i < TC_BITMASK + 1; i++)
netdev_set_prio_tc_map(dev, i, qopt->prio_tc_map[i]);
...
}
otx2_mqprio_restore_old() only replays num_tc and the per-TC
count/offset pair:
netdev_set_num_tc(netdev, snap->num_tc);
for (tc = 0; tc < snap->num_tc; tc++)
netdev_set_tc_queue(netdev, tc, snap->count[tc],
snap->offset[tc]);
After a rollback of a replacement whose setup succeeded but which was
never grafted (for example qdisc_create() failing at err_out4 with
TCQ_F_MQROOT and then calling ops->destroy()), dev->prio_tc_map[] still
describes the abandoned layout while num_tc/tc_to_txq describe the old
one, so skb_tx_hash() steers traffic to queues whose MDQ shapers belong
to a different configuration.
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> index 32582b6347ea4..1e17f8d49fd44 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> @@ -2007,6 +2007,14 @@ int otx2_open(struct net_device *netdev)
> if (err)
> goto err_free_mem;
>
> + err = otx2_mqprio_up(pf);
> + if (err) {
> + netdev_err(pf->netdev,
> + "mqprio: failed to restore shapers during open: %d\n",
> + err);
> + goto err_free_hw;
> + }
> +
> /* Register NAPI handler */
> for (qidx = 0; qidx < pf->hw.cint_cnt; qidx++) {
> cq_poll = &qset->napi[qidx];
> @@ -2205,6 +2213,7 @@ int otx2_open(struct net_device *netdev)
> free_irq(vec, pf);
> err_disable_napi:
> otx2_disable_napi(pf);
> +err_free_hw:
> otx2_free_hw_resources(pf);
> err_free_mem:
> otx2_free_queue_mem(qset);
[Severity: Medium]
This isn't a bug introduced by this patch, but does this new error label
leak pf->refill_wrk?
pf->refill_wrk is allocated (and its delayed works initialised) inside
otx2_init_hw_resources() -> otx2_config_nix_queues():
pfvf->refill_wrk = devm_kcalloc(pfvf->dev, pfvf->qset.cq_cnt,
sizeof(struct refill_work), GFP_KERNEL);
The only place it is cancelled and freed is otx2_stop():
for (wrk = 0; wrk < pf->qset.cq_cnt; wrk++)
cancel_delayed_work_sync(&pf->refill_wrk[wrk].pool_refill_work);
devm_kfree(pf->dev, pf->refill_wrk);
None of the otx2_open() error labels, including the new err_free_hw,
cancel the works or free the array, so each failed open leaks one array
until driver detach. This unbalanced lifecycle is shared by all the
pre-existing labels, but this patch adds a new one and makes ndo_open()
failure routinely reachable from userspace through
otx2_mqprio_restart_netdev() on tc qdisc add/replace/del.
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> index ddb46b580c3b6..70d99ae3defb5 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> @@ -1600,14 +1609,750 @@ static int otx2_setup_tc_block(struct net_device *netdev,
> nic, nic, ingress);
> }
>
> +/* Free the per-queue min/max rate caches. */
> +static void otx2_mqprio_free_cache(struct otx2_nic *pfvf)
> +{
> + devm_kfree(pfvf->dev, pfvf->mqprio.min_rate);
> + devm_kfree(pfvf->dev, pfvf->mqprio.max_rate);
> + pfvf->mqprio.min_rate = NULL;
> + pfvf->mqprio.max_rate = NULL;
> + pfvf->mqprio.flags = 0;
> +}
> +
> +static int otx2_mqprio_alloc_cache(struct otx2_nic *pfvf, bool replacing)
> +{
> + u16 num_txq = pfvf->hw.non_qos_queues;
> +
> + if (replacing && pfvf->mqprio.min_rate && pfvf->mqprio.max_rate) {
> + memset(pfvf->mqprio.min_rate, 0,
> + num_txq * sizeof(*pfvf->mqprio.min_rate));
> + memset(pfvf->mqprio.max_rate, 0,
> + num_txq * sizeof(*pfvf->mqprio.max_rate));
> + pfvf->mqprio.flags = 0;
> + return 0;
> + }
> +
> + otx2_mqprio_free_cache(pfvf);
> +
> + pfvf->mqprio.min_rate = devm_kcalloc(pfvf->dev, num_txq,
> + sizeof(*pfvf->mqprio.min_rate),
> + GFP_KERNEL);
> + pfvf->mqprio.max_rate = devm_kcalloc(pfvf->dev, num_txq,
> + sizeof(*pfvf->mqprio.max_rate),
> + GFP_KERNEL);
> + if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate) {
> + otx2_mqprio_free_cache(pfvf);
> + return -ENOMEM;
> + }
> +
> + return 0;
> +}
[Severity: Medium]
Can a failed replace lose the active configuration permanently here?
otx2_mqprio_free_cache() releases the live caches before the two
devm_kcalloc() calls, so on -ENOMEM both min_rate and max_rate are NULL.
otx2_mqprio_restore_old() bails out before rebuilding anything:
err = otx2_mqprio_alloc_cache(pfvf, false);
if (err)
return err;
At that point the old qdisc is still installed and mqprio.rate_limit is
still true, but the rates are gone. The next open cannot recover them
either, because otx2_mqprio_up() treats missing caches as success:
if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate)
return 0;
Would allocating into temporary pointers and only swapping them in on
success avoid destroying the live state?
[ ... ]
> +static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit);
> +
> +static int otx2_mqprio_restore_old(struct otx2_nic *pfvf)
> +{
> + struct mq_offload_snap *snap = pfvf->old_mq_snap;
> + struct net_device *netdev = pfvf->netdev;
> + u16 num_txq = pfvf->hw.non_qos_queues;
> + int tc, txq, err;
> +
> + if (!snap)
> + return 0;
> +
> + err = otx2_mqprio_alloc_cache(pfvf, false);
> + if (err)
> + return err;
[ ... ]
> + netdev_set_num_tc(netdev, snap->num_tc);
> + for (tc = 0; tc < snap->num_tc; tc++)
> + netdev_set_tc_queue(netdev, tc, snap->count[tc],
> + snap->offset[tc]);
> +
> + if (otx2_mqprio_mdq_allocated(pfvf)) {
> + err = otx2_nix_tm_clear_queue_shaper(pfvf);
> + if (err)
> + return err;
> + }
[Severity: Medium]
Does this early return leave hardware and software state out of sync?
By the time this runs, the software cache has already been rewritten from
old_mq_snap, and the failed replace has already programmed some MDQ
CIR/PIR registers with the new rates in the
otx2_nix_tm_set_queue_shaper() loop of otx2_setup_tc_mqprio(). Returning
here skips both
pfvf->mqprio.rate_limit = true;
and the netif_running() restart, so the interface keeps running with
per-queue shapers holding a partially applied configuration that the
driver's bookkeeping no longer describes, until some unrelated stop/open
cycle happens. Should a failure here force the netdev restart so the AF
resets shaping on freshly allocated MDQs?
> +
> + /* Rebuild the TX scheduler via netdev restart when running; otx2_mqprio_up()
> + * alone is insufficient after a failed replace that already bounced the
> + * interface. If open failed, TX schedulers were freed; defer shaper restore
> + * to the next successful ndo_open() via otx2_mqprio_up().
> + */
> + pfvf->mqprio.rate_limit = true;
> +
> + if (netif_running(netdev)) {
> + err = otx2_mqprio_restart_netdev(netdev, true);
> + if (err)
> + return err;
> + } else if (pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]) {
> + err = otx2_mqprio_up(pfvf);
> + if (err)
> + return err;
> + }
> +
> + otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
> +
> + return 0;
> +}
[ ... ]
> +static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
> +{
> + struct net_device *netdev = pfvf->netdev;
> +
> + pfvf->mqprio.rate_limit = false;
> + otx2_mqprio_clear_replace_state(pfvf);
> + netdev_set_num_tc(netdev, 0);
> + otx2_mqprio_free_cache(pfvf);
> +}
[ ... ]
> +int otx2_mqprio_up(struct otx2_nic *pfvf)
> +{
> + struct net_device *netdev = pfvf->netdev;
> + int txq, err;
> +
> + if (!pfvf->mqprio.rate_limit)
> + return 0;
> +
> + if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate)
> + return 0;
[ ... ]
> +static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit)
> +{
> + struct otx2_nic *pfvf = netdev_priv(netdev);
> + const struct net_device_ops *ops = netdev->netdev_ops;
> + bool running = netif_running(netdev);
> + int err;
> +
> + /* TODO: Explore live TX scheduler reprogramming to avoid a full
> + * ndo_stop()/ndo_open() bounce on every mqprio change.
> + */
> + netdev_info(netdev,
> + "mqprio: restarting interface to reprogram TX scheduler; in-flight traffic will be dropped\n");
> +
> + if (running) {
> + clear_bit(__LINK_STATE_START, &netdev->state);
> + smp_mb__after_atomic(); /* Commit netif_running(). */
> + }
> + dev_deactivate(netdev, true);
> +
> + err = ops->ndo_stop(netdev);
> + if (err) {
> + if (running) {
> + set_bit(__LINK_STATE_START, &netdev->state);
> + dev_activate(netdev);
> + }
> + return err;
> + }
[Severity: High]
Is it safe for a driver to drive dev_deactivate()/dev_activate() from
inside an ndo_setup_tc(TC_SETUP_QDISC_MQPRIO) callback?
On the teardown/replace path this callback runs from
mqprio_destroy() -> mqprio_disable_offload(), which the core invokes from
the middle of qdisc_graft():
net/sched/sch_api.c:qdisc_graft() {
if (dev->flags & IFF_UP)
dev_deactivate(dev, false);
qdisc_offload_graft_root(dev, new, old, extack);
...
notify_and_destroy(net, skb, n, classid, old, new, extack);
if (new && new->ops->attach)
new->ops->attach(new);
}
if (dev->flags & IFF_UP)
dev_activate(dev);
}
dev_deactivate() does not clear __LINK_STATE_START, so
otx2_teardown_tc_mqprio() sees netif_running() == true and calls
otx2_mqprio_restart_netdev(netdev, false), whose
set_bit(__LINK_STATE_START, &netdev->state);
dev_activate(netdev);
clears __QDISC_STATE_DEACTIVATED on the outgoing per-queue child qdiscs
and republishes them into dev_queue->qdisc, since new->ops->attach() has
not run yet. mq_attach()/mqprio_attach() then dev_graft_qdisc() and
qdisc_put() exactly those qdiscs without any further
some_qdisc_is_busy() wait, so a concurrent softirq __qdisc_run() can be
running inside a qdisc that is being reset and freed.
Reproducer would be "tc qdisc replace dev X root mq" on a running
interface carrying traffic with the offload active. The same reentrancy
also purges an unrelated live qdisc on the add path, where
dev_deactivate(netdev, true) is called from within mqprio_init().
Can the offload be restructured so the driver never touches the core's
activate/deactivate state machine from a qdisc callback?
[ ... ]
> +static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
> + struct tc_mqprio_qopt_offload *mqprio)
> +{
> + struct tc_mqprio_qopt *qopt = &mqprio->qopt;
> + bool had_mqprio = pfvf->mqprio.rate_limit;
> + struct net_device *netdev = pfvf->netdev;
> + bool if_up = netif_running(netdev);
> + int err;
> +
> + qopt->hw = 0;
> +
> + /* tc qdisc replace runs setup on the new mqprio before destroying the
> + * old one. replace_setup_done and TC_ROOT_GRAFT distinguish stale
> + * old-instance teardown from graft failure after setup.
> + */
> + if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
> + err = 0;
> + if (pfvf->mqprio.replace_graft_done)
> + otx2_mqprio_snap_commit(pfvf);
> + else
> + err = otx2_mqprio_restore_old(pfvf);
> + otx2_mqprio_clear_replace_state(pfvf);
> + return err;
> + }
> +
> + /* Skip the netdev restart when mqprio offload was not active. */
> + if (!had_mqprio)
> + return 0;
> +
> + if (if_up) {
> + int down_err, err;
> +
> + down_err = otx2_mqprio_down(pfvf);
> + err = otx2_mqprio_restart_netdev(netdev, false);
> + if (err)
> + return err;
> + return down_err;
> + }
[Severity: Medium]
Does replacing an offloaded mqprio with a non-offloaded configuration
wipe the TC layout of the qdisc that is already live?
For "tc qdisc replace dev X root mqprio ... hw 0" (or "... root mq"),
mqprio_init() installs the new layout without ever calling
ndo_setup_tc():
net/sched/sch_mqprio.c:mqprio_init() {
...
} else {
netdev_set_num_tc(dev, qopt->num_tc);
for (i = 0; i < qopt->num_tc; i++)
netdev_set_tc_queue(dev, i, qopt->count[i], qopt->offset[i]);
}
...
}
So replace_setup_done stays false, and the later destruction of the old
offloaded instance reaches this function and takes the full-teardown
branch above, whose otx2_mqprio_clear_sw() runs
netdev_set_num_tc(netdev, 0) on the qdisc the core has just grafted.
The result is dev->num_tc == 0 while a software mqprio expects N traffic
classes, so skb_tx_hash() no longer honours the prio/TC/txq ranges and
mqprio_dump() reports 0 TCs. The netdev is also bounced in that case.
[ ... ]
> +cleanup:
> + qopt->hw = 0;
> + if (replacing) {
> + int restore_err = otx2_mqprio_restore_old(pfvf);
> +
> + otx2_mqprio_clear_replace_state(pfvf);
> + if (restore_err) {
[ ... ]
> + return err ? err : -EIO;
> + }
[Severity: Medium]
Does the core immediately undo the netdev part of this rollback?
Because otx2_setup_tc_mqprio() returns an error here,
mqprio_enable_offload() propagates it before setting priv->hw_offload,
so when qdisc_create() unwinds via ops->destroy() the core takes the
non-offload branch:
net/sched/sch_mqprio.c:mqprio_destroy() {
...
if (priv->hw_offload && dev->netdev_ops->ndo_setup_tc)
mqprio_disable_offload(sch);
else
netdev_set_num_tc(dev, 0);
...
}
That leaves the old mqprio qdisc installed and the driver offload active
(mqprio.rate_limit true, MDQ CIR/PIR programmed from old_mq_snap) while
dev->num_tc is 0 and tc_to_txq[] is cleared, so traffic spreads over the
whole queue range and tc qdisc show reports 0 TCs for an offloaded
qdisc. This is reachable from any failure after the bounce, for example
an -ENOMEM from otx2_mqprio_alloc_cache() or a mailbox error from
otx2_nix_tm_set_queue_shaper().
[ ... ]
--
This is an AI-generated review.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-17 11:32 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-11 10:55 [PATCH v15 net-next 0/2] octeontx2-pf: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-09-11 10:55 ` [PATCH v15 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
2026-09-17 11:31 ` Paolo Abeni
2026-09-11 10:55 ` [PATCH v15 net-next 2/2] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-09-17 11:31 ` Paolo Abeni
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®