mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections
@ 2026-10-09 20:46 Daniel Zahka
  2026-10-09 20:46 ` [PATCH net-next v2 1/7] psp: support rx rekey operation Daniel Zahka
                   ` (6 more replies)
  0 siblings, 7 replies; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

The PSP architecture spec states the need for rekeying connections in
the event of a device key rotation. After a device key rotation has
occurred, a new rx derived key needs to be generated and sent out to the
other end of the connection sometime before the next device key rotation
occurs.

Because PSP connections involve two different keys at each endpoint,
one for decrypting ingress traffic, and one for encrypting egress
traffic, there are two types of rekeying events that need to be
supported. From the perspective of one endpoint of the connection:

1. rx rekey: we need to allocate a new spi + decryption key pair to
   provide to our peer.

2. tx rekey: our peer has provided us with a new spi + encryption key
   pair which we should use for encrypting traffic immediately.

In the case of rx rekeying, there is a period where it makes sense to
accept packets authenticated from either the previous or current spi. To
deal with that, we allow a psp_assoc to remember the last spi that was
valid on a socket due to a rekey. If authentication state does not match
the most recent assoc, the stored state from the previous assoc will be
tried.

In the case of tx rekeying, as soon as we install the new tx key, we
have no use for the previous one, and it can be disposed of immediately.
The only catch is in the case where hw uses a key handle in tx
descriptor state, as opposed to inlining the key directly. In this
case, psp core needs to be sure that any of these unaccounted for
references to key state are gone by the time it tries to sync a deleted
key to hw.

To deal with this race condition, we introduce a different path to key
deletion for psp_assocs removed from the socket during a tx rekey. psp
core will use bql byte counters filled out by the driver to determine a
conservative grace period where key handles can be disposed of.

Lastly, some test cases for rekeying are included that go through key
rotations and rekeying.

There are some packetdrill tests that are queued for upstreaming [1].

[1]: https://github.com/danieldzahka/packetdrill/commits/psp-rekey/

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>

Changes in v2:
- psp: support rx rekey operation
  - copy old rx state by value instead of pointer to prev
  - place non-datapath fields at end of struct psp_assoc
  - disallow rx rekey when socket is not in full psp state, or
    dev/version doesn't match
  - require the new rx spi to have the opposite phase bit from the
    previous one
- psp: support tx rekey operation
  - place new assoc on pas->assocs list instead of psd->active_assocs
  - disallow tx rekey when socket is not in full psp state
  - document rekeying in psp.rst
  - reject tx rekey on SADB devices until deferred tx key deletion
    lands
- psp: defer tx key deletions for SADB drivers
  - replaces "psp: add driver api for deferred tx key deletion"
  - use bql byte counters instead of driver callback for grace periods
  - only defer tx key deletion when socket outlives psp_assoc
  - document driver requirements in psp.rst
  - allow tx rekey on SADB devices
- psp: add core tracked stat for outstanding tx keys
  - replaces "psp: add core tracked stats for deferred key deletion"
  - drop grace-periods stat
  - rename tx-key-cnt to tx-key-count, only report it for SADB drivers
  - warn about outstanding tx keys in psp_dev_unregister()
- selftests: drv-net: psp: factor out psp connection setup
  - split psp_responder conn_setup_psp() into rx_assoc() and tx_assoc()
  - drop Reviewed-by from Willem due to psp_responder changes
- selftests: drv-net: psp: add rekey tests
  - add tests for rekeys rejected due to psp state and version
    mismatch
  - add test for rx rekey rejected without a device key rotation
  - add rx-assoc, tx-assoc, and key-rotate rpcs to psp_responder
- selftests: drv-net: psp: add a tx rekey drain test for SADB drivers
  - new patch
- mlx5: psp: implement deferred tx key deletion
  - dropped, bql based grace periods need no driver callback
- Link to v1: https://lore.kernel.org/r/20260204-psp-v1-0-5f034e2dfa36@gmail.com

---
Daniel Zahka (7):
      psp: support rx rekey operation
      psp: support tx rekey operation
      psp: defer tx key deletions for SADB drivers
      psp: add core tracked stat for outstanding tx keys
      selftests: drv-net: psp: factor out psp connection setup
      selftests: drv-net: psp: add rekey tests
      selftests: drv-net: psp: add a tx rekey drain test for SADB drivers

 Documentation/netlink/specs/psp.yaml               |   8 +
 Documentation/networking/psp.rst                   |  64 +++-
 include/net/psp/functions.h                        |  10 +-
 include/net/psp/types.h                            |  34 +++
 include/uapi/linux/psp.h                           |   1 +
 net/psp/Kconfig                                    |   1 +
 net/psp/Makefile                                   |   2 +-
 net/psp/psp.h                                      |   8 +-
 net/psp/psp_deferred_del.c                         | 231 +++++++++++++++
 net/psp/psp_main.c                                 |  25 +-
 net/psp/psp_nl.c                                   |   4 +
 net/psp/psp_sock.c                                 | 113 ++++++-
 tools/testing/selftests/drivers/net/psp.py         | 325 ++++++++++++++++++++-
 .../testing/selftests/drivers/net/psp_responder.c  | 198 +++++++++++--
 14 files changed, 970 insertions(+), 54 deletions(-)
---
base-commit: d8674294aefef02266c4d47ad10131f1bffbe534
change-id: 20260202-psp-3c8e2f65c5c4

Best regards,
-- 
Daniel Zahka <daniel.zahka@gmail.com>


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 1/7] psp: support rx rekey operation
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 2/7] psp: support tx " Daniel Zahka
                   ` (5 subsequent siblings)
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

Support updating the rx PSP state used by a socket. This is most
useful to do after a device key rotation has occurred and the key
being used by the peer will become stale after the next device key
rotation.

Create a new psp_assoc for the socket to use that contains the same tx
state, but with new rx state. Copy the previous rx state to use as a
fallback while the peer is in the process of switching to the new key.

skbs matching the current rx spi and generation or the previous rx spi
and generation will be accepted. It might be reasonable to stop
accepting skbs on the previous state after seeing skbs with the new
state and after waiting a grace period, but that is not implemented
for now.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
v2:
- copy old rx state by value instead of pointer to prev
- place non-datapath fields at end of struct psp_assoc
- disallow rx rekey when socket is not in full psp state, or
  dev/version doesn't match.
- require the new rx spi to have the opposite phase bit from the previous
  one
---
 include/net/psp/functions.h | 10 ++++++----
 include/net/psp/types.h     | 14 ++++++++++++++
 net/psp/psp.h               |  3 ++-
 net/psp/psp_sock.c          | 46 +++++++++++++++++++++++++++++++++++++++++----
 4 files changed, 64 insertions(+), 9 deletions(-)

diff --git a/include/net/psp/functions.h b/include/net/psp/functions.h
index b23c30898389..fb80e15e4368 100644
--- a/include/net/psp/functions.h
+++ b/include/net/psp/functions.h
@@ -77,10 +77,12 @@ psp_is_allowed_nondata(struct sk_buff *skb, struct psp_assoc *pas)
 static inline bool
 psp_pse_matches_pas(struct psp_skb_ext *pse, struct psp_assoc *pas)
 {
-	return pse && pas->rx.spi == pse->spi &&
-	       pas->generation == pse->generation &&
-	       pas->version == pse->version &&
-	       pas->dev_id == pse->dev_id;
+	return pse && pas->version == pse->version &&
+	       pas->dev_id == pse->dev_id &&
+	       ((pas->rx.spi == pse->spi &&
+		 pas->generation == pse->generation) ||
+		(pas->prev_spi && pas->prev_spi == pse->spi &&
+		 pas->prev_generation == pse->generation));
 }
 
 static inline enum skb_drop_reason
diff --git a/include/net/psp/types.h b/include/net/psp/types.h
index b8905efbd604..52c78cc72f1f 100644
--- a/include/net/psp/types.h
+++ b/include/net/psp/types.h
@@ -3,6 +3,7 @@
 #ifndef __NET_PSP_H
 #define __NET_PSP_H
 
+#include <linux/bits.h>
 #include <linux/mutex.h>
 #include <linux/refcount.h>
 #include <net/net_trackers.h>
@@ -152,6 +153,15 @@ struct psp_key_parsed {
 	u8 key[PSP_MAX_KEY];
 };
 
+/**
+ * enum psp_assoc_flags - flags of struct psp_assoc
+ * @PSP_ASSOC_SKIP_TX_KEY_DEL: Do not delete Tx key from psp_dev. It was
+ *	copied to a newer psp_assoc during an Rx rekey.
+ */
+enum psp_assoc_flags {
+	PSP_ASSOC_SKIP_TX_KEY_DEL	= BIT(0),
+};
+
 struct psp_assoc {
 	struct psp_dev *psd;
 
@@ -159,6 +169,10 @@ struct psp_assoc {
 	u8 generation;
 	u8 version;
 	u8 peer_tx;
+	u8 prev_generation;
+	u8 flags; /* Slow path, protected by psd->lock */
+
+	__be32 prev_spi;
 
 	u32 upgrade_seq;
 
diff --git a/net/psp/psp.h b/net/psp/psp.h
index b123c2427905..92d7b92acaed 100644
--- a/net/psp/psp.h
+++ b/net/psp/psp.h
@@ -63,7 +63,8 @@ static inline bool psp_dev_has_sadb(struct psp_dev *psd)
 static inline bool psp_assoc_needs_tx_key_del(struct psp_assoc *pas)
 {
 	lockdep_assert_held(&pas->psd->lock);
-	return psp_dev_has_sadb(pas->psd) && pas->tx.spi;
+	return psp_dev_has_sadb(pas->psd) && pas->tx.spi &&
+	       !(pas->flags & PSP_ASSOC_SKIP_TX_KEY_DEL);
 }
 
 #endif /* __PSP_PSP_H */
diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
index a6b1c42dd626..b5887171c84e 100644
--- a/net/psp/psp_sock.c
+++ b/net/psp/psp_sock.c
@@ -149,10 +149,42 @@ void psp_sk_assoc_free(struct sock *sk)
 	psp_assoc_put(pas);
 }
 
+static int psp_sock_rx_rekey(struct psp_assoc *pas, struct psp_assoc *prev,
+			     struct netlink_ext_ack *extack)
+{
+	if (pas->psd != prev->psd) {
+		NL_SET_ERR_MSG(extack, "PSP device mismatch with existing state");
+		return -EINVAL;
+	}
+	if (pas->version != prev->version) {
+		NL_SET_ERR_MSG(extack, "PSP version mismatch with existing state");
+		return -EINVAL;
+	}
+	if (!prev->tx.spi || !prev->peer_tx) {
+		NL_SET_ERR_MSG(extack, "Socket PSP state is not fully established");
+		return -EBUSY;
+	}
+	if (!((pas->rx.spi ^ prev->rx.spi) & cpu_to_be32(PSP_SPI_KEY_PHASE))) {
+		NL_SET_ERR_MSG(extack, "New and prev SPI have same phase bit");
+		return -EINVAL;
+	}
+
+	pas->peer_tx = 1;
+	pas->prev_spi = prev->rx.spi;
+	pas->prev_generation = prev->generation;
+
+	memcpy(&pas->tx, &prev->tx, sizeof(pas->tx));
+	memcpy(pas->drv_data, prev->drv_data, pas->psd->caps->assoc_drv_spc);
+	prev->flags |= PSP_ASSOC_SKIP_TX_KEY_DEL;
+
+	return 0;
+}
+
 int psp_sock_assoc_set_rx(struct sock *sk, struct psp_assoc *pas,
 			  struct psp_key_parsed *key,
 			  struct netlink_ext_ack *extack)
 {
+	struct psp_assoc *prev;
 	int err;
 
 	memcpy(&pas->rx, key, sizeof(*key));
@@ -165,10 +197,11 @@ int psp_sock_assoc_set_rx(struct sock *sk, struct psp_assoc *pas,
 		goto exit_unlock;
 	}
 
-	if (psp_sk_assoc(sk)) {
-		NL_SET_ERR_MSG(extack, "Socket already has PSP state");
-		err = -EBUSY;
-		goto exit_unlock;
+	prev = psp_sk_assoc(sk);
+	if (prev) {
+		err = psp_sock_rx_rekey(pas, prev, extack);
+		if (err)
+			goto exit_unlock;
 	} else if (sk_has_decrypt_user(sk)) {
 		NL_SET_ERR_MSG(extack, "Socket has incompatible state");
 		err = -EINVAL;
@@ -177,6 +210,7 @@ int psp_sock_assoc_set_rx(struct sock *sk, struct psp_assoc *pas,
 
 	refcount_inc(&pas->refcnt);
 	rcu_assign_pointer(sk->psp_assoc, pas);
+	psp_assoc_put(prev);
 	err = 0;
 
 exit_unlock:
@@ -304,6 +338,10 @@ void psp_assocs_key_rotated(struct psp_dev *psd)
 		pas->generation |= ~PSP_GEN_VALID_MASK;
 		psd->stats.stales++;
 	}
+
+	list_for_each_entry(pas, &psd->active_assocs, assocs_list)
+		pas->prev_generation |= ~PSP_GEN_VALID_MASK;
+
 	list_splice_init(&psd->prev_assocs, &psd->stale_assocs);
 	list_splice_init(&psd->active_assocs, &psd->prev_assocs);
 

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 2/7] psp: support tx rekey operation
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
  2026-10-09 20:46 ` [PATCH net-next v2 1/7] psp: support rx rekey operation Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers Daniel Zahka
                   ` (4 subsequent siblings)
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

The tx rekey operation creates a new psp_assoc with the same rx state,
but with new tx state. The new association is spliced into the socket
with RCU to be used immediately in the tx path.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
v2:
- place new assoc on pas->assocs list instead of psd->active assocs
  list
- reject tx rekey on SADB devices until deferred tx key deletion
  lands
---
 Documentation/networking/psp.rst | 26 +++++++++++++++++--
 net/psp/psp_sock.c               | 56 +++++++++++++++++++++++++++++++++++-----
 2 files changed, 74 insertions(+), 8 deletions(-)

diff --git a/Documentation/networking/psp.rst b/Documentation/networking/psp.rst
index 0f9b6b73f244..5c4e4215d906 100644
--- a/Documentation/networking/psp.rst
+++ b/Documentation/networking/psp.rst
@@ -140,14 +140,36 @@ The PSP assoc state of a socket is not reset when the connection is
 torn down. ``connect()`` on a socket that has PSP assoc state will
 return ``-EINVAL``.
 
+Rekeying
+--------
+
+A connection which has completed the exchange described above and is
+in the "PSP Full" state can be rekeyed in place without being torn
+down. The Tx and Rx directions are rekeyed separately using the same
+netlink calls as in connection setup.
+
+``rx-assoc`` allocates a new Rx key and SPI, which should be passed
+to the peer exactly as during the initial exchange. The SPI which was
+in use before the rekey remains acceptable on the socket until the
+next Rx rekey, so packets the peer sent before it learned the new SPI
+are still received.
+
+``tx-assoc`` installs a new Tx key, which takes effect immediately.
+The peer will accept data encrypted from the old and new SPI.
+
+A rekey is rejected with ``-EINVAL`` if it would change the PSP device
+or the PSP version of the association. It is rejected with ``-EBUSY``
+if the socket has not completed its initial exchange in both
+directions.
+
 Rotation notifications
 ----------------------
 
 The rotations of device key happen asynchronously and are usually
 performed by management daemons, not under application control.
 The PSP netlink family will generate a notification whenever keys
-are rotated. The applications are expected to re-establish connections
-before keys are rotated again.
+are rotated. Applications are expected to rekey their connections,
+as described above, before device keys are rotated again.
 
 Kernel implementation
 =====================
diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
index b5887171c84e..060149f3d72a 100644
--- a/net/psp/psp_sock.c
+++ b/net/psp/psp_sock.c
@@ -283,6 +283,52 @@ psp_sock_set_tx_key(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
 	return err;
 }
 
+static int
+psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
+		  struct psp_key_parsed *key, struct netlink_ext_ack *extack)
+{
+	struct psp_assoc *new;
+	int err;
+
+	if (psp_dev_has_sadb(psd)) {
+		NL_SET_ERR_MSG(extack, "Tx rekey not supported on this device");
+		return -EOPNOTSUPP;
+	}
+	if (!pas->peer_tx) {
+		NL_SET_ERR_MSG(extack, "Socket PSP state is not fully established");
+		return -EBUSY;
+	}
+
+	new = kzalloc_flex(*new, drv_data, psd->caps->assoc_drv_spc,
+			   GFP_KERNEL_ACCOUNT);
+	if (!new)
+		return -ENOMEM;
+
+	new->psd             = pas->psd;
+	new->dev_id          = pas->dev_id;
+	new->generation      = pas->generation;
+	new->version         = pas->version;
+	new->peer_tx         = 1;
+	new->prev_spi        = pas->prev_spi;
+	new->prev_generation = pas->prev_generation;
+	refcount_set(&new->refcnt, 1);
+	memcpy(&new->rx, &pas->rx, sizeof(new->rx));
+
+	err = psp_assoc_set_tx(psd, new, key, extack);
+	if (err) {
+		kfree(new);
+		return err;
+	}
+
+	psp_dev_get(new->psd);
+	list_add(&new->assocs_list, &pas->assocs_list);
+
+	rcu_assign_pointer(sk->psp_assoc, new);
+	psp_assoc_put(pas);
+
+	return 0;
+}
+
 int psp_sock_assoc_set_tx(struct sock *sk, struct psp_dev *psd,
 			  u32 version, struct psp_key_parsed *key,
 			  struct netlink_ext_ack *extack)
@@ -315,13 +361,11 @@ int psp_sock_assoc_set_tx(struct sock *sk, struct psp_dev *psd,
 		err = -EINVAL;
 		goto exit_unlock;
 	}
-	if (pas->tx.spi) {
-		NL_SET_ERR_MSG(extack, "Tx key already set");
-		err = -EBUSY;
-		goto exit_unlock;
-	}
+	if (pas->tx.spi)
+		err = psp_sock_tx_rekey(sk, psd, pas, key, extack);
+	else
+		err = psp_sock_set_tx_key(sk, psd, pas, key, extack);
 
-	err = psp_sock_set_tx_key(sk, psd, pas, key, extack);
 exit_unlock:
 	release_sock(sk);
 	return err;

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
  2026-10-09 20:46 ` [PATCH net-next v2 1/7] psp: support rx rekey operation Daniel Zahka
  2026-10-09 20:46 ` [PATCH net-next v2 2/7] psp: support tx " Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys Daniel Zahka
                   ` (3 subsequent siblings)
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

Devices that store tx keys in an SADB can hold references to key
handles in tx descriptor state. PSP core needs to make sure all of
these tx descriptors have been completed before calling
psp_dev_tx_key_del(). Before rekeying, this was ensured by skbs
holding references to skb->sk, and sk->psp_assoc being cleared only in
the socket destructor.

Rekeying introduces a code path where a psp_assoc is detached from its
socket and may be freed while the socket is still alive. This means
skbs may exist in the driver ring that were rendered into descriptors
using that old psp_assoc. We cannot delete the key from the SADB until
we know these descriptors have been completed.

After a psp_assoc has all references drained (refcount and rcu), PSP
core records the BQL queued count for active TX queue and starts a
descriptor grace period. It then polls the queued and completed byte
counts until all data outstanding at the start of the grace period has
completed. After that point, any keys queued for deletion before the
start of the grace period can be deleted from the SADB.

Make CONFIG_INET_PSP depend on CONFIG_BQL.

Drivers which install Tx keys must meet the requirements in
Documentation/networking/psp.rst.

Regarding lock ordering: the order of taking psp_dev::lock before
net_device::lock was established in commit 1e6e0cec0dab ("net/mlx5e:
psp: Make PSP steering config dynamic")

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
v2:
- move grace period mechanism to use bql instead of driver callback.
- allow tx rekey on SADB devices
---
 Documentation/networking/psp.rst |  38 +++++++
 include/net/psp/types.h          |  18 +++
 net/psp/Kconfig                  |   1 +
 net/psp/Makefile                 |   2 +-
 net/psp/psp.h                    |   5 +
 net/psp/psp_deferred_del.c       | 231 +++++++++++++++++++++++++++++++++++++++
 net/psp/psp_main.c               |  22 +++-
 net/psp/psp_sock.c               |  13 ++-
 8 files changed, 321 insertions(+), 9 deletions(-)

diff --git a/Documentation/networking/psp.rst b/Documentation/networking/psp.rst
index 5c4e4215d906..769c7b21045c 100644
--- a/Documentation/networking/psp.rst
+++ b/Documentation/networking/psp.rst
@@ -202,6 +202,44 @@ Drivers must use ``psp_skb_get_assoc_rcu()`` to check if PSP Tx offload
 was requested for given skb. On Rx drivers should allocate and populate
 the ``SKB_EXT_PSP`` skb extension, and set the skb->decrypted bit to 1.
 
+Deferred Tx key deletion
+~~~~~~~~~~~~~~~~~~~~~~~~
+
+Drivers which implement ``tx_key_add`` and ``tx_key_del`` may hold key
+references in their Tx rings, so PSP core waits until descriptors
+containing these references are completed before removing the key from
+the underlying device. PSP core uses accounting based on Byte Queue Limits
+(BQL) to wait for a suitable grace period before asking the driver to
+remove the key.
+
+Drivers which implement the Tx key callbacks must meet the following
+requirements:
+
+* The driver must be "ops locked"
+  (see Documentation/networking/netdevices.rst).
+* A PSP skb accepted by ``ndo_start_xmit()`` must remain socket-owned until
+  every Tx descriptor referencing its PSP key has been completed or otherwise
+  made incapable of using the key. The driver must not orphan the skb before
+  that point. Otherwise, the socket destructor may free the Tx key while key
+  references remain in the Tx ring.
+* Every PSP packet accepted by ``ndo_start_xmit()`` must be included in BQL
+  queued bytecount before the function returns.
+* Completion bytecounts must exactly match enqueue bytecounts and preserve
+  queue order. A later packet must not be reported as completed while an earlier
+  packet can still reference its key.
+* Before calling ``dql_reset()`` on a queue, the driver must ensure that all
+  outstanding descriptors containing key references have been completed or
+  discarded. If called under the netdev instance lock, then descriptors can
+  remain outstanding after the reset as long as they are completed or
+  discarded before the instance lock is released.
+* When decreasing ``real_num_tx_queues``, the driver must ensure that every
+  outstanding descriptor on the queues at or above the new Tx queue count
+  has been completed or otherwise made incapable of using its PSP key
+  before the netdev instance lock is released.
+* The BQL state of a queue below ``real_num_tx_queues`` must reflect all
+  bytes completed or be reset after the queue has been quiesced. Failure to
+  do so will stall the grace period algorithm.
+
 Kernel implementation notes
 ---------------------------
 
diff --git a/include/net/psp/types.h b/include/net/psp/types.h
index 52c78cc72f1f..8ffc566cec2a 100644
--- a/include/net/psp/types.h
+++ b/include/net/psp/types.h
@@ -6,6 +6,7 @@
 #include <linux/bits.h>
 #include <linux/mutex.h>
 #include <linux/refcount.h>
+#include <linux/workqueue.h>
 #include <net/net_trackers.h>
 
 struct netlink_ext_ack;
@@ -80,6 +81,11 @@ struct psp_assoc_dev {
  * @prev_assocs:	associations which use old (but still usable)
  *			device key
  * @stale_assocs:	associations which use a rotated out key
+ * @tx_del:		deferred TX key deletion state
+ * @tx_del.active:	TX keys to be deleted after the current grace period
+ * @tx_del.next:	TX keys queued for deletion during current grace period
+ * @tx_del.work:	delayed work item for periodic grace period checking
+ * @tx_del.txq_state:	TX queue state for deferred deletion grace periods
  *
  * @stats:	statistics maintained by the core
  * @stats.rotations:	See stats attr key-rotations
@@ -109,6 +115,13 @@ struct psp_dev {
 	struct list_head prev_assocs;
 	struct list_head stale_assocs;
 
+	struct {
+		struct list_head active;
+		struct list_head next;
+		struct delayed_work work;
+		struct psp_txq_state *txq_state;
+	} tx_del;
+
 	struct {
 		unsigned long rotations;
 		unsigned long stales;
@@ -157,9 +170,14 @@ struct psp_key_parsed {
  * enum psp_assoc_flags - flags of struct psp_assoc
  * @PSP_ASSOC_SKIP_TX_KEY_DEL: Do not delete Tx key from psp_dev. It was
  *	copied to a newer psp_assoc during an Rx rekey.
+ * @PSP_ASSOC_DEFER_TX_KEY_DEL: Send Tx key to psp_dev's deferred deletion
+ *	queue. This psp_assoc was detached from its socket during a Tx rekey
+ *	and requires device descriptor state referencing this key to be
+ *	drained.
  */
 enum psp_assoc_flags {
 	PSP_ASSOC_SKIP_TX_KEY_DEL	= BIT(0),
+	PSP_ASSOC_DEFER_TX_KEY_DEL	= BIT(1),
 };
 
 struct psp_assoc {
diff --git a/net/psp/Kconfig b/net/psp/Kconfig
index 84d6b0f25460..36c12493784f 100644
--- a/net/psp/Kconfig
+++ b/net/psp/Kconfig
@@ -5,6 +5,7 @@
 config INET_PSP
 	bool "PSP Security Protocol support"
 	depends on INET
+	depends on BQL
 	select SKB_DECRYPTED
 	select SKB_EXTENSIONS
 	select SOCK_VALIDATE_XMIT
diff --git a/net/psp/Makefile b/net/psp/Makefile
index eb5ff3c5bfb2..40decd8fe050 100644
--- a/net/psp/Makefile
+++ b/net/psp/Makefile
@@ -2,4 +2,4 @@
 
 obj-$(CONFIG_INET_PSP) += psp.o
 
-psp-y := psp_main.o psp_nl.o psp_sock.o psp-nl-gen.o
+psp-y := psp_main.o psp_nl.o psp_sock.o psp_deferred_del.o psp-nl-gen.o
diff --git a/net/psp/psp.h b/net/psp/psp.h
index 92d7b92acaed..0cfc82c6c711 100644
--- a/net/psp/psp.h
+++ b/net/psp/psp.h
@@ -18,6 +18,11 @@ int psp_dev_check_access(struct psp_dev *psd, struct net *net, bool admin);
 bool psp_has_assoc_dev_in_ns(struct psp_dev *psd, struct net *net);
 int psp_attach_netdev_notifier(void);
 
+int psp_deferred_del_init(struct psp_dev *psd);
+void psp_deferred_del_stop(struct psp_dev *psd);
+void psp_deferred_del_uninit(struct psp_dev *psd, bool unpublished);
+void psp_deferred_del_queue(struct psp_dev *psd, struct psp_assoc *pas);
+
 void psp_nl_notify_dev(struct psp_dev *psd, u32 cmd);
 void psp_nl_notify_disassoc(struct psp_dev *psd, struct net *net);
 
diff --git a/net/psp/psp_deferred_del.c b/net/psp/psp_deferred_del.c
new file mode 100644
index 000000000000..cfbac72577e9
--- /dev/null
+++ b/net/psp/psp_deferred_del.c
@@ -0,0 +1,231 @@
+// SPDX-License-Identifier: GPL-2.0-only
+
+#include <linux/bitmap.h>
+#include <linux/jiffies.h>
+#include <linux/list.h>
+#include <linux/netdevice.h>
+#include <linux/slab.h>
+#include <net/netdev_lock.h>
+#include <net/psp.h>
+
+#include "psp.h"
+
+#define PSP_TX_GRACE_POLL_MS	1000
+
+struct psp_txq_state {
+	unsigned int num_tx_queues;
+	unsigned long *drained_queues;
+	unsigned int to_complete[];
+};
+
+static void psp_txq_state_free(struct psp_txq_state *state)
+{
+	if (!state)
+		return;
+
+	bitmap_free(state->drained_queues);
+	kvfree(state);
+}
+
+static struct psp_txq_state *
+psp_txq_state_alloc(unsigned int max_tx_queues)
+{
+	struct psp_txq_state *state;
+
+	state = kvzalloc_flex(*state, to_complete, max_tx_queues);
+	if (!state)
+		return NULL;
+
+	state->drained_queues = bitmap_zalloc(max_tx_queues, GFP_KERNEL);
+	if (!state->drained_queues) {
+		psp_txq_state_free(state);
+		return NULL;
+	}
+
+	return state;
+}
+
+static void psp_tx_grace_start(struct psp_dev *psd)
+{
+	struct psp_txq_state *state = psd->tx_del.txq_state;
+	struct net_device *dev = psd->main_netdev;
+	unsigned int i;
+
+	netdev_assert_locked(dev);
+
+	/* Any instances of ndo_start_xmit() that observed these keys through
+	 * psp_skb_get_assoc_rcu() must already be accounted in BQL's queued
+	 * byte count. ndo_start_xmit() runs with BHs disabled, which forms an
+	 * implicit RCU read-side critical section, and we are now at least one
+	 * grace period after any of these psp_assocs have been cleared from
+	 * sk->psp_assoc.
+	 */
+	state->num_tx_queues = dev->real_num_tx_queues;
+	bitmap_zero(state->drained_queues, state->num_tx_queues);
+	for (i = 0; i < state->num_tx_queues; i++) {
+		struct netdev_queue *txq = netdev_get_tx_queue(dev, i);
+
+		state->to_complete[i] = READ_ONCE(txq->dql.num_queued);
+	}
+}
+
+static bool psp_tx_grace_done(struct psp_dev *psd)
+{
+	struct psp_txq_state *state = psd->tx_del.txq_state;
+	struct net_device *dev = psd->main_netdev;
+	bool all_completed = true;
+	unsigned int i;
+
+	netdev_assert_locked(dev);
+
+	if (dev->real_num_tx_queues < state->num_tx_queues)
+		state->num_tx_queues = dev->real_num_tx_queues;
+
+	for (i = 0; i < state->num_tx_queues; i++) {
+		struct netdev_queue *txq = netdev_get_tx_queue(dev, i);
+		unsigned int completed;
+		unsigned int queued;
+
+		if (test_bit(i, state->drained_queues))
+			continue;
+
+		/* Drivers must meet the requirements documented in
+		 * Documentation/networking/psp.rst to guarantee forward
+		 * progress and avoid ending the grace period prematurely.
+		 *
+		 * We snapshotted num_queued in psp_tx_grace_start() and need
+		 * to determine whether num_completed has since reached or
+		 * passed that mark. For a given queue, we can consider first
+		 * whether or not dql_reset() was called at any point between
+		 * when we read num_queued until we read num_completed and
+		 * num_queued here and then separate the space of possible
+		 * states based on that.
+		 *
+		 * Not reset: num_queued and num_completed only increase, and
+		 * dql keeps num_queued - num_completed <= INT_MAX at all
+		 * times. We can form three exhaustive sub cases:
+		 *
+		 * (int)(to_complete - completed) <= 0: completion reached or
+		 * passed the mark. At the time to_complete was read,
+		 * completed was in [to_complete - INT_MAX, to_complete].
+		 *
+		 * (int)(to_complete - queued) > 0: to_complete and queued are
+		 * both snapshots of num_queued, so queued has advanced more
+		 * than INT_MAX + 1 past to_complete. num_completed lags
+		 * queued by at most INT_MAX, so num_completed must have
+		 * crossed to_complete.
+		 *
+		 * Otherwise the mark may still be in flight, so poll again.
+		 *
+		 * Reset: By the driver conditions laid out in
+		 * Documentation/networking/psp.rst, if a dql_reset()
+		 * occurred, and we are holding the instance lock, then we
+		 * know the actual state of the queue is drained. The check
+		 * below is then conservative because there is no risk of
+		 * falsely declaring the queue drained. As long as dql makes
+		 * forward progress the check should eventually report
+		 * drained.
+		 */
+		completed = READ_ONCE(txq->dql.num_completed);
+		queued = READ_ONCE(txq->dql.num_queued);
+
+		if ((int)(state->to_complete[i] - completed) <= 0 ||
+		    (int)(state->to_complete[i] - queued) > 0)
+			__set_bit(i, state->drained_queues);
+		else
+			all_completed = false;
+	}
+
+	return all_completed;
+}
+
+static bool psp_tx_grace_active(struct psp_dev *psd)
+{
+	return !list_empty(&psd->tx_del.active);
+}
+
+static void psp_deferred_del_work(struct work_struct *work)
+{
+	struct delayed_work *dwork = to_delayed_work(work);
+	struct psp_assoc *pas, *tmp;
+	bool need_reschedule;
+	struct psp_dev *psd;
+	LIST_HEAD(to_free);
+
+	psd = container_of(dwork, struct psp_dev, tx_del.work);
+
+	mutex_lock(&psd->lock);
+	netdev_lock(psd->main_netdev);
+
+	if (psp_tx_grace_active(psd) && psp_tx_grace_done(psd))
+		list_splice_init(&psd->tx_del.active, &to_free);
+
+	if (!psp_tx_grace_active(psd) && !list_empty(&psd->tx_del.next)) {
+		psp_tx_grace_start(psd);
+		list_splice_init(&psd->tx_del.next, &psd->tx_del.active);
+	}
+
+	netdev_unlock(psd->main_netdev);
+
+	list_for_each_entry(pas, &to_free, assocs_list)
+		psp_dev_tx_key_del(psd, pas);
+
+	need_reschedule = psp_tx_grace_active(psd);
+	mutex_unlock(&psd->lock);
+
+	if (need_reschedule)
+		mod_delayed_work(system_percpu_wq, &psd->tx_del.work,
+				 msecs_to_jiffies(PSP_TX_GRACE_POLL_MS));
+
+	list_for_each_entry_safe(pas, tmp, &to_free, assocs_list) {
+		list_del(&pas->assocs_list);
+		psp_dev_put(psd);
+		kfree(pas);
+	}
+}
+
+int psp_deferred_del_init(struct psp_dev *psd)
+{
+	INIT_LIST_HEAD(&psd->tx_del.active);
+	INIT_LIST_HEAD(&psd->tx_del.next);
+	INIT_DELAYED_WORK(&psd->tx_del.work, psp_deferred_del_work);
+
+	if (!psd->ops->tx_key_del)
+		return 0;
+
+	psd->tx_del.txq_state =
+		psp_txq_state_alloc(psd->main_netdev->num_tx_queues);
+
+	return psd->tx_del.txq_state ? 0 : -ENOMEM;
+}
+
+void psp_deferred_del_stop(struct psp_dev *psd)
+{
+	disable_delayed_work_sync(&psd->tx_del.work);
+}
+
+void psp_deferred_del_uninit(struct psp_dev *psd, bool unpublished)
+{
+	struct psp_assoc *pas, *next;
+
+	lockdep_assert(unpublished || lockdep_is_held(&psd->lock));
+
+	psp_txq_state_free(psd->tx_del.txq_state);
+	psd->tx_del.txq_state = NULL;
+
+	list_splice_init(&psd->tx_del.active, &psd->tx_del.next);
+	list_for_each_entry_safe(pas, next, &psd->tx_del.next, assocs_list) {
+		list_del(&pas->assocs_list);
+		psp_dev_tx_key_del(psd, pas);
+		psp_dev_put(psd);
+		kfree(pas);
+	}
+}
+
+void psp_deferred_del_queue(struct psp_dev *psd, struct psp_assoc *pas)
+{
+	lockdep_assert_held(&psd->lock);
+
+	list_move_tail(&pas->assocs_list, &psd->tx_del.next);
+	schedule_delayed_work(&psd->tx_del.work, 0);
+}
diff --git a/net/psp/psp_main.c b/net/psp/psp_main.c
index 273b010d2355..e118b8019e7e 100644
--- a/net/psp/psp_main.c
+++ b/net/psp/psp_main.c
@@ -3,8 +3,10 @@
 #include <linux/bitfield.h>
 #include <linux/list.h>
 #include <linux/netdevice.h>
+#include <linux/slab.h>
 #include <linux/xarray.h>
 #include <net/net_namespace.h>
+#include <net/netdev_lock.h>
 #include <net/psp.h>
 #include <net/udp.h>
 
@@ -70,7 +72,8 @@ psp_dev_create(struct net_device *netdev,
 		    !psd_ops->rx_spi_alloc ||
 		    !psd_ops->get_stats ||
 		    (!psd_ops->tx_key_add != !psd_ops->tx_key_del) ||
-		    (psd_caps->assoc_drv_spc && !psd_ops->tx_key_add)))
+		    (psd_caps->assoc_drv_spc && !psd_ops->tx_key_add) ||
+		    (psd_ops->tx_key_add && !netdev_need_ops_lock(netdev))))
 		return ERR_PTR(-EINVAL);
 
 	psd = kzalloc_obj(*psd);
@@ -89,13 +92,16 @@ psp_dev_create(struct net_device *netdev,
 	INIT_LIST_HEAD(&psd->stale_assocs);
 	refcount_set(&psd->refcnt, 1);
 
+	err = psp_deferred_del_init(psd);
+	if (err)
+		goto err_free_psd;
+
 	mutex_lock(&psp_devs_lock);
 	err = xa_alloc_cyclic(&psp_devs, &psd->id, psd, xa_limit_16b,
 			      &last_id, GFP_KERNEL);
 	if (err) {
 		mutex_unlock(&psp_devs_lock);
-		kfree(psd);
-		return ERR_PTR(err);
+		goto err_uninit_deferred_del;
 	}
 	mutex_lock(&psd->lock);
 	mutex_unlock(&psp_devs_lock);
@@ -111,6 +117,12 @@ psp_dev_create(struct net_device *netdev,
 	mutex_unlock(&psd->lock);
 
 	return psd;
+
+err_uninit_deferred_del:
+	psp_deferred_del_uninit(psd, true);
+err_free_psd:
+	kfree(psd);
+	return ERR_PTR(err);
 }
 EXPORT_SYMBOL(psp_dev_create);
 
@@ -133,6 +145,8 @@ void psp_dev_unregister(struct psp_dev *psd)
 	struct psp_assoc_dev *entry, *entry_tmp;
 	struct psp_assoc *pas, *next;
 
+	psp_deferred_del_stop(psd);
+
 	mutex_lock(&psp_devs_lock);
 	mutex_lock(&psd->lock);
 
@@ -145,6 +159,8 @@ void psp_dev_unregister(struct psp_dev *psd)
 	xa_store(&psp_devs, psd->id, NULL, GFP_KERNEL);
 	mutex_unlock(&psp_devs_lock);
 
+	psp_deferred_del_uninit(psd, false);
+
 	list_splice_init(&psd->active_assocs, &psd->prev_assocs);
 	list_splice_init(&psd->prev_assocs, &psd->stale_assocs);
 	list_for_each_entry_safe(pas, next, &psd->stale_assocs, assocs_list) {
diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
index 060149f3d72a..d21f7afd830a 100644
--- a/net/psp/psp_sock.c
+++ b/net/psp/psp_sock.c
@@ -114,8 +114,14 @@ static void psp_assoc_free(struct work_struct *work)
 
 	mutex_lock(&psd->lock);
 	if (psp_dev_is_registered(psd)) {
-		if (psp_assoc_needs_tx_key_del(pas))
+		if (psp_assoc_needs_tx_key_del(pas)) {
+			if (pas->flags & PSP_ASSOC_DEFER_TX_KEY_DEL) {
+				psp_deferred_del_queue(psd, pas);
+				mutex_unlock(&psd->lock);
+				return;
+			}
 			psp_dev_tx_key_del(psd, pas);
+		}
 		list_del(&pas->assocs_list);
 	}
 	mutex_unlock(&psd->lock);
@@ -290,10 +296,6 @@ psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
 	struct psp_assoc *new;
 	int err;
 
-	if (psp_dev_has_sadb(psd)) {
-		NL_SET_ERR_MSG(extack, "Tx rekey not supported on this device");
-		return -EOPNOTSUPP;
-	}
 	if (!pas->peer_tx) {
 		NL_SET_ERR_MSG(extack, "Socket PSP state is not fully established");
 		return -EBUSY;
@@ -324,6 +326,7 @@ psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
 	list_add(&new->assocs_list, &pas->assocs_list);
 
 	rcu_assign_pointer(sk->psp_assoc, new);
+	pas->flags |= PSP_ASSOC_DEFER_TX_KEY_DEL;
 	psp_assoc_put(pas);
 
 	return 0;

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
                   ` (2 preceding siblings ...)
  2026-10-09 20:46 ` [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup Daniel Zahka
                   ` (2 subsequent siblings)
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

It is useful to keep track of tx_key_cnt: the number of outstanding tx
keys from the hw's perspective. Userspace can monitor this to
determine if there is a key leak (leak in the sense that some device
resource like an SADB entry has been leaked, not that the crypto key
has been exposed).

It also may be of interest to userspace if a device is approaching its
SADB capacity.

Only report this stat to userspace if the driver uses an SADB.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
 Documentation/netlink/specs/psp.yaml | 8 ++++++++
 include/net/psp/types.h              | 2 ++
 include/uapi/linux/psp.h             | 1 +
 net/psp/psp_main.c                   | 3 +++
 net/psp/psp_nl.c                     | 4 ++++
 net/psp/psp_sock.c                   | 6 +++++-
 6 files changed, 23 insertions(+), 1 deletion(-)

diff --git a/Documentation/netlink/specs/psp.yaml b/Documentation/netlink/specs/psp.yaml
index f3266763c325..d14c11d09891 100644
--- a/Documentation/netlink/specs/psp.yaml
+++ b/Documentation/netlink/specs/psp.yaml
@@ -190,6 +190,13 @@ attribute-sets:
         doc: |
           Number of PSP packets for transmission with errors.
           Device statistic (from the PSP spec).
+      -
+        name: tx-key-count
+        type: uint
+        doc: |
+          Current number of Tx keys installed on the device.
+          Only visible if the driver stores keys on device.
+          Kernel statistic.
 
 operations:
   list:
@@ -316,6 +323,7 @@ operations:
             - tx-packets
             - tx-bytes
             - tx-error
+            - tx-key-count
         pre: psp-device-get-locked
         post: psp-device-unlock
       dump:
diff --git a/include/net/psp/types.h b/include/net/psp/types.h
index 8ffc566cec2a..970801cb6502 100644
--- a/include/net/psp/types.h
+++ b/include/net/psp/types.h
@@ -90,6 +90,7 @@ struct psp_assoc_dev {
  * @stats:	statistics maintained by the core
  * @stats.rotations:	See stats attr key-rotations
  * @stats.stales:	See stats attr stale-events
+ * @stats.tx_key_cnt:	See stats attr tx-key-count
  *
  * @rcu:	RCU head for freeing the structure
  */
@@ -125,6 +126,7 @@ struct psp_dev {
 	struct {
 		unsigned long rotations;
 		unsigned long stales;
+		unsigned long tx_key_cnt;
 	} stats;
 
 	struct rcu_head rcu;
diff --git a/include/uapi/linux/psp.h b/include/uapi/linux/psp.h
index 1c8899cd4da5..21810e8a7229 100644
--- a/include/uapi/linux/psp.h
+++ b/include/uapi/linux/psp.h
@@ -69,6 +69,7 @@ enum {
 	PSP_A_STATS_TX_PACKETS,
 	PSP_A_STATS_TX_BYTES,
 	PSP_A_STATS_TX_ERROR,
+	PSP_A_STATS_TX_KEY_COUNT,
 
 	__PSP_A_STATS_MAX,
 	PSP_A_STATS_MAX = (__PSP_A_STATS_MAX - 1)
diff --git a/net/psp/psp_main.c b/net/psp/psp_main.c
index e118b8019e7e..9f256157d88b 100644
--- a/net/psp/psp_main.c
+++ b/net/psp/psp_main.c
@@ -169,6 +169,9 @@ void psp_dev_unregister(struct psp_dev *psd)
 		list_del(&pas->assocs_list);
 	}
 
+	WARN(psd->stats.tx_key_cnt, "psp: %s: %lu Tx keys still installed\n",
+	     netdev_name(psd->main_netdev), psd->stats.tx_key_cnt);
+
 	list_for_each_entry_safe(entry, entry_tmp, &psd->assoc_dev_list,
 				 dev_list) {
 		list_del(&entry->dev_list);
diff --git a/net/psp/psp_nl.c b/net/psp/psp_nl.c
index cdfc2d72fb39..494022ab5a71 100644
--- a/net/psp/psp_nl.c
+++ b/net/psp/psp_nl.c
@@ -934,6 +934,10 @@ psp_nl_stats_fill(struct psp_dev *psd, struct sk_buff *rsp,
 	    nla_put_uint(rsp, PSP_A_STATS_TX_ERROR, stats.tx_error))
 		goto err_cancel_msg;
 
+	if (psp_dev_has_sadb(psd) &&
+	    nla_put_uint(rsp, PSP_A_STATS_TX_KEY_COUNT, psd->stats.tx_key_cnt))
+		goto err_cancel_msg;
+
 	genlmsg_end(rsp, hdr);
 	return 0;
 
diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
index d21f7afd830a..c458bc551b4b 100644
--- a/net/psp/psp_sock.c
+++ b/net/psp/psp_sock.c
@@ -94,9 +94,11 @@ static int psp_dev_tx_key_add(struct psp_dev *psd, struct psp_assoc *pas,
 
 	memcpy(&dummy->tx, key, sizeof(*key));
 	err = psd->ops->tx_key_add(psd, dummy, extack);
-	if (!err)
+	if (!err) {
+		psd->stats.tx_key_cnt++;
 		memcpy(pas->drv_data, dummy->drv_data,
 		       psd->caps->assoc_drv_spc);
+	}
 
 	kfree(dummy);
 	return err;
@@ -105,6 +107,8 @@ static int psp_dev_tx_key_add(struct psp_dev *psd, struct psp_assoc *pas,
 void psp_dev_tx_key_del(struct psp_dev *psd, struct psp_assoc *pas)
 {
 	psd->ops->tx_key_del(psd, pas);
+	if (!WARN_ON_ONCE(!psd->stats.tx_key_cnt))
+		psd->stats.tx_key_cnt--;
 }
 
 static void psp_assoc_free(struct work_struct *work)

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
                   ` (3 preceding siblings ...)
  2026-10-09 20:46 ` [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests Daniel Zahka
  2026-10-09 20:46 ` [PATCH net-next v2 7/7] selftests: drv-net: psp: add a tx rekey drain test for SADB drivers Daniel Zahka
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

Lift out connection setup and PSP assoc uapi code in psp.py and
responder, so that the key exchange sequence can be used in other test
cases that start with a key exchange, but might choose to perform
additional tx/rx assoc calls afterwards.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
 tools/testing/selftests/drivers/net/psp.py         | 14 +++-
 .../testing/selftests/drivers/net/psp_responder.c  | 87 +++++++++++++++-------
 2 files changed, 73 insertions(+), 28 deletions(-)

diff --git a/tools/testing/selftests/drivers/net/psp.py b/tools/testing/selftests/drivers/net/psp.py
index 473500901879..71eeade62ad2 100755
--- a/tools/testing/selftests/drivers/net/psp.py
+++ b/tools/testing/selftests/drivers/net/psp.py
@@ -447,10 +447,11 @@ def assoc_twice(cfg):
             ksft_eq(len(tx), 0)
 
 
-def _data_basic_send(cfg, version, ipver):
-    """ Test basic data send """
-    _init_psp_dev(cfg)
+def _establish_psp_conn(cfg, version, ipver=None):
+    """Establish a PSP connection and return after key exchange
 
+    Requires _init_psp_dev() to have been called first.
+    """
     # Version 0 is required by spec, don't let it skip
     if version:
         name = cfg.pspnl.consts["version"].entries_by_val[version].name
@@ -475,7 +476,14 @@ def _data_basic_send(cfg, version, ipver):
                         "version": version,
                         "tx-key": tx,
                         "sock-fd": s.fileno()})
+    return s
+
+
+def _data_basic_send(cfg, version, ipver):
+    """ Test basic data send """
+    _init_psp_dev(cfg)
 
+    s = _establish_psp_conn(cfg, version, ipver)
     data_len = _send_careful(cfg, s, 100)
     _check_data_rx(cfg, data_len)
     _close_psp_conn(cfg, s)
diff --git a/tools/testing/selftests/drivers/net/psp_responder.c b/tools/testing/selftests/drivers/net/psp_responder.c
index a26e7628bbb1..57425ecb9561 100644
--- a/tools/testing/selftests/drivers/net/psp_responder.c
+++ b/tools/testing/selftests/drivers/net/psp_responder.c
@@ -37,20 +37,24 @@ static struct {
 	unsigned char rx;
 } psp_vers;
 
-static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
+static unsigned int psp_key_len(unsigned char version)
 {
+	switch (version) {
+	case PSP_VERSION_HDR0_AES_GCM_256:
+	case PSP_VERSION_HDR0_AES_GMAC_256:
+		return 32;
+	default:
+		return 16;
+	}
+}
+
+static int
+rx_assoc(struct ynl_sock *ys, __u32 *spi, char *key, int data_sock)
+{
+	unsigned int key_len = psp_key_len(psp_vers.rx);
 	struct psp_rx_assoc_rsp *rsp;
 	struct psp_rx_assoc_req *req;
-	struct psp_tx_assoc_rsp *tsp;
-	struct psp_tx_assoc_req *teq;
-	char info[300];
-	int key_len;
-	ssize_t sz;
-	__u32 spi;
 
-	dbg("create PSP connection\n");
-
-	// Rx assoc alloc
 	req = psp_rx_assoc_req_alloc();
 
 	psp_rx_assoc_req_set_sock_fd(req, data_sock);
@@ -64,29 +68,34 @@ static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
 		return -1;
 	}
 
-	// SPI exchange
-	key_len = rsp->rx_key._len.key;
-	memcpy(info, &rsp->rx_key.spi, sizeof(spi));
-	memcpy(&info[sizeof(spi)], rsp->rx_key.key, key_len);
-	sz = sizeof(spi) + key_len;
+	if (rsp->rx_key._len.key != key_len) {
+		fprintf(stderr, "ERROR: unexpected Rx key length %u\n",
+			rsp->rx_key._len.key);
+		psp_rx_assoc_rsp_free(rsp);
+		return -1;
+	}
+
+	memcpy(spi, &rsp->rx_key.spi, sizeof(*spi));
+	memcpy(key, rsp->rx_key.key, key_len);
 
-	send(data_sock, info, sz, MSG_WAITALL);
 	psp_rx_assoc_rsp_free(rsp);
 
-	sz = recv(data_sock, info, sz, MSG_WAITALL);
-	if (sz < 0) {
-		perror("ERROR: failed to read PSP key from sock");
-		return -1;
-	}
-	memcpy(&spi, info, sizeof(spi));
+	return 0;
+}
+
+static int
+tx_assoc(struct ynl_sock *ys, __u8 version, __u32 spi, char *key,
+	 int data_sock)
+{
+	struct psp_tx_assoc_rsp *tsp;
+	struct psp_tx_assoc_req *teq;
 
-	// Setup Tx assoc
 	teq = psp_tx_assoc_req_alloc();
 
 	psp_tx_assoc_req_set_sock_fd(teq, data_sock);
-	psp_tx_assoc_req_set_version(teq, psp_vers.tx);
+	psp_tx_assoc_req_set_version(teq, version);
 	psp_tx_assoc_req_set_tx_key_spi(teq, spi);
-	psp_tx_assoc_req_set_tx_key_key(teq, &info[sizeof(spi)], key_len);
+	psp_tx_assoc_req_set_tx_key_key(teq, key, psp_key_len(version));
 
 	tsp = psp_tx_assoc(ys, teq);
 	psp_tx_assoc_req_free(teq);
@@ -99,6 +108,34 @@ static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
 	return 0;
 }
 
+static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
+{
+	unsigned int key_len = psp_key_len(psp_vers.rx);
+	char info[300];
+	ssize_t sz;
+	__u32 spi;
+
+	dbg("create PSP connection\n");
+
+	if (rx_assoc(ys, &spi, &info[sizeof(spi)], data_sock))
+		return -1;
+
+	// SPI exchange
+	memcpy(info, &spi, sizeof(spi));
+	sz = sizeof(spi) + key_len;
+
+	send(data_sock, info, sz, MSG_WAITALL);
+
+	sz = recv(data_sock, info, sz, MSG_WAITALL);
+	if (sz < 0) {
+		perror("ERROR: failed to read PSP key from sock");
+		return -1;
+	}
+	memcpy(&spi, info, sizeof(spi));
+
+	return tx_assoc(ys, psp_vers.tx, spi, &info[sizeof(spi)], data_sock);
+}
+
 static void send_ack(int sock)
 {
 	send(sock, "ack", 4, MSG_WAITALL);

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
                   ` (4 preceding siblings ...)
  2026-10-09 20:46 ` [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  2026-10-10 21:43   ` netdev-bot+sashiko
  2026-10-09 20:46 ` [PATCH net-next v2 7/7] selftests: drv-net: psp: add a tx rekey drain test for SADB drivers Daniel Zahka
  6 siblings, 1 reply; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

Add testcases for rekeying psp connections.

Tests include coverage for different cases where rx-assoc and tx-assoc
should fail based on mismatched psp version and psp state.

Add rpc endpoints for tx-assoc, rx-assoc, and key-rotate netlink calls
to the psp_responder.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
v2:
- add tests for rekeys rejected due to psp state and version mismatch
- add test for rx rekey rejected without a device key rotation
- add rx-assoc, tx-assoc, and key-rotate rpcs to psp_responder
---
 tools/testing/selftests/drivers/net/psp.py         | 273 ++++++++++++++++++++-
 .../testing/selftests/drivers/net/psp_responder.c  | 111 +++++++++
 2 files changed, 382 insertions(+), 2 deletions(-)

diff --git a/tools/testing/selftests/drivers/net/psp.py b/tools/testing/selftests/drivers/net/psp.py
index 71eeade62ad2..af67c9e8c40f 100755
--- a/tools/testing/selftests/drivers/net/psp.py
+++ b/tools/testing/selftests/drivers/net/psp.py
@@ -28,6 +28,10 @@ from lib.py import ip
 TCP_ULP = 31
 
 
+_PSP_MAX_KEY_LEN = 32
+_PSP_ASSOC_MSG = f'!IB3x{_PSP_MAX_KEY_LEN}s'
+
+
 def _get_outq(s):
     one = b'\0' * 4
     outq = fcntl.ioctl(s.fileno(), termios.TIOCOUTQ, one)
@@ -80,6 +84,36 @@ def _close_psp_conn(cfg, s):
     _close_conn(cfg, s)
 
 
+def _recv_all(s, target):
+    data = b''
+    while len(data) < target:
+        chunk = s.recv(target - len(data))
+        if not chunk:
+            raise KsftFailEx(f"peer closed after {len(data)} of {target} bytes")
+        data += chunk
+    return data
+
+
+def _psp_key_len(version):
+    return 16 if version in (0, 2) else 32
+
+
+def _remote_rx_assoc(cfg):
+    _send_with_ack(cfg, b'rx assoc\0')
+    msg = _recv_all(cfg.comm_sock, struct.calcsize(_PSP_ASSOC_MSG))
+    spi, version, key = struct.unpack(_PSP_ASSOC_MSG, msg)
+    return version, {'spi': spi, 'key': key[:_psp_key_len(version)]}
+
+
+def _remote_tx_assoc(cfg, version, rx):
+    msg = struct.pack(_PSP_ASSOC_MSG, rx['spi'], version, rx['key'])
+    _send_with_ack(cfg, b'tx assoc\0' + msg)
+
+
+def _remote_key_rotate(cfg):
+    _send_with_ack(cfg, b'key rotate\0')
+
+
 def _spi_xchg(s, rx):
     s.send(struct.pack('I', rx['spi']) + rx['key'])
     tx = s.recv(4 + len(rx['key']))
@@ -130,6 +164,41 @@ def _check_data_outq(s, exp_len, force_wait=False):
     ksft_eq(outq, exp_len)
 
 
+def _recv_careful(s, target, timeout=2):
+    """Read exactly target bytes, tolerating short reads"""
+    data = b''
+    end = time.monotonic() + timeout
+    while time.monotonic() < end:
+        try:
+            data += s.recv(target - len(data), socket.MSG_DONTWAIT)
+            if len(data) == target:
+                return data
+        except BlockingIOError:
+            time.sleep(0.001)
+    raise KsftFailEx(f"short read, got {len(data)} of {target} bytes")
+
+
+def _req_echo(cfg, s):
+    """Ask the peer to echo, and check the reply arrives intact"""
+    _send_with_ack(cfg, b'data echo\0')
+    ksft_eq(_recv_careful(s, 5), b'echo\0')
+
+
+def _psp_txrx(cfg, s, rounds, sent=0):
+    """Send data both ways, and return the total bytes sent to the peer"""
+    sent += _send_careful(cfg, s, rounds)
+    _check_data_rx(cfg, sent)
+    _req_echo(cfg, s)
+    return sent
+
+
+def _require_version(cfg, version):
+    """Skip the test unless the device supports the given PSP version"""
+    name = cfg.pspnl.consts["version"].entries_by_val[version].name
+    if name not in cfg.psp_info['psp-versions-cap']:
+        raise KsftSkipEx("PSP version not supported", name)
+
+
 def _get_stat(cfg, key):
     return cfg.pspnl.get_stats({'dev-id': cfg.psp_dev_id})[key]
 
@@ -489,6 +558,204 @@ def _data_basic_send(cfg, version, ipver):
     _close_psp_conn(cfg, s)
 
 
+def _rekey_rx(cfg, s, version, sent):
+    """Rekey the Rx direction, running traffic after each step"""
+    rx_assoc = cfg.pspnl.rx_assoc({"version": version,
+                                   "dev-id": cfg.psp_dev_id,
+                                   "sock-fd": s.fileno()})
+    sent = _psp_txrx(cfg, s, 1, sent)
+
+    _remote_tx_assoc(cfg, version, rx_assoc['rx-key'])
+    return _psp_txrx(cfg, s, 1, sent)
+
+
+def _rekey_tx(cfg, s, sent):
+    """Rekey the Tx direction, running traffic after each step"""
+    version, tx = _remote_rx_assoc(cfg)
+    sent = _psp_txrx(cfg, s, 1, sent)
+
+    cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
+                        "version": version,
+                        "tx-key": tx,
+                        "sock-fd": s.fileno()})
+    return _psp_txrx(cfg, s, 1, sent)
+
+
+def rekey_rx_incomplete(cfg):
+    """Test rejecting Rx rekey until the PSP connection is established"""
+    psp_ver = 0
+    _init_psp_dev(cfg)
+
+    with _make_lo_conn() as s:
+        assoc = cfg.pspnl.rx_assoc({"version": psp_ver,
+                                    "dev-id": cfg.psp_dev_id,
+                                    "sock-fd": s.fileno()})
+
+        # Reject rekey before Tx state is configured.
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.rx_assoc({"version": psp_ver,
+                                "dev-id": cfg.psp_dev_id,
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EBUSY)
+
+        cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
+                            "version": psp_ver,
+                            "tx-key": assoc['rx-key'],
+                            "sock-fd": s.fileno()})
+
+        # Reject rekey until authenticated PSP traffic has been received.
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.rx_assoc({"version": psp_ver,
+                                "dev-id": cfg.psp_dev_id,
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EBUSY)
+
+
+def rekey_tx_incomplete(cfg):
+    """Test rejecting Tx rekey until the PSP connection is established"""
+    psp_ver = 0
+    _init_psp_dev(cfg)
+
+    with _make_lo_conn() as s:
+        assoc = cfg.pspnl.rx_assoc({"version": psp_ver,
+                                    "dev-id": cfg.psp_dev_id,
+                                    "sock-fd": s.fileno()})
+
+        cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
+                            "version": psp_ver,
+                            "tx-key": assoc['rx-key'],
+                            "sock-fd": s.fileno()})
+
+        # Reject rekey until authenticated PSP traffic has been received.
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
+                                "version": psp_ver,
+                                "tx-key": assoc['rx-key'],
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EBUSY)
+
+
+def rekey_rx_same_gen(cfg):
+    """Test rejecting an Rx rekey without an intervening key rotation"""
+    _init_psp_dev(cfg)
+
+    s = _establish_psp_conn(cfg, 0)
+    try:
+        _psp_txrx(cfg, s, 1)
+
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.rx_assoc({"version": 0,
+                                "dev-id": cfg.psp_dev_id,
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EINVAL)
+
+        cfg.pspnl.key_rotate({"id": cfg.psp_dev_id})
+        cfg.pspnl.rx_assoc({"version": 0,
+                            "dev-id": cfg.psp_dev_id,
+                            "sock-fd": s.fileno()})
+    finally:
+        _close_psp_conn(cfg, s)
+
+
+def rekey_rx_version_mismatch(cfg):
+    """Test rejecting an Rx rekey whose version does not match socket state"""
+    _init_psp_dev(cfg)
+    _require_version(cfg, 1)
+
+    s = _establish_psp_conn(cfg, 0)
+    try:
+        _psp_txrx(cfg, s, 1)
+
+        cfg.pspnl.key_rotate({"id": cfg.psp_dev_id})
+
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.rx_assoc({"dev-id": cfg.psp_dev_id,
+                                "version": 1,
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EINVAL)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
+def rekey_tx_version_mismatch(cfg):
+    """Test rejecting a Tx rekey whose version does not match socket state"""
+    _init_psp_dev(cfg)
+    _require_version(cfg, 1)
+
+    s = _establish_psp_conn(cfg, 0)
+    try:
+        _psp_txrx(cfg, s, 1)
+
+        with ksft_raises(NlError) as cm:
+            cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
+                                "version": 1,
+                                "tx-key": {'spi': 0x12345678,
+                                           'key': b'\xa5' * 32},
+                                "sock-fd": s.fileno()})
+        ksft_eq(cm.exception.nl_msg.error, -errno.EINVAL)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
+def _get_psp_ver_variants():
+    for ver in range(4):
+        yield KsftNamedVariant(f"v{ver}", ver)
+
+
+@ksft_variants(_get_psp_ver_variants())
+def rekey_rx_basic(cfg, version):
+    """Test rekeying the Rx key of an established connection"""
+    _init_psp_dev(cfg)
+
+    s = _establish_psp_conn(cfg, version)
+    try:
+        data_len = _psp_txrx(cfg, s, 10)
+        cfg.pspnl.key_rotate({"id": cfg.psp_dev_id})
+        _rekey_rx(cfg, s, version, data_len)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
+@ksft_variants(_get_psp_ver_variants())
+def rekey_tx_basic(cfg, version):
+    """Test rekeying the Tx key of an established connection"""
+    _init_psp_dev(cfg)
+
+    s = _establish_psp_conn(cfg, version)
+    try:
+        data_len = _psp_txrx(cfg, s, 10)
+        _remote_key_rotate(cfg)
+        _rekey_tx(cfg, s, data_len)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
+@ksft_variants(_get_psp_ver_variants())
+def rekey_both_sides(cfg, version):
+    """Test rekeying both directions"""
+    _init_psp_dev(cfg)
+
+    s = _establish_psp_conn(cfg, version)
+    try:
+        data_len = _psp_txrx(cfg, s, 10)
+
+        cfg.pspnl.key_rotate({"id": cfg.psp_dev_id})
+        _remote_key_rotate(cfg)
+
+        # first, rx then tx
+        data_len = _rekey_rx(cfg, s, version, data_len)
+        data_len = _rekey_tx(cfg, s, data_len)
+
+        cfg.pspnl.key_rotate({"id": cfg.psp_dev_id})
+        _remote_key_rotate(cfg)
+
+        # now, tx then rx
+        data_len = _rekey_tx(cfg, s, data_len)
+        _rekey_rx(cfg, s, version, data_len)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
 def __bad_xfer_do(cfg, s, tx, version='hdr0-aes-gcm-128'):
     # Make sure we accept the ACK for the SPI before we seal with the bad assoc
     _check_data_outq(s, 0)
@@ -1181,7 +1448,8 @@ def main() -> None:
                                                           cfg.comm_port),
                                                          timeout=1)
 
-                cases = [data_basic_send, data_mss_adjust]
+                cases = [data_basic_send, data_mss_adjust,
+                         rekey_rx_basic, rekey_tx_basic, rekey_both_sides]
 
                 if has_cont:
                     cases += [
@@ -1197,7 +1465,8 @@ def main() -> None:
                     ]
 
                 ksft_run(cases=cases, globs=globals(),
-                         case_pfx={"dev_", "data_", "assoc_", "removal_"},
+                         case_pfx={"dev_", "data_", "assoc_", "rekey_",
+                                   "removal_"},
                          args=(cfg, ))
 
                 cfg.comm_sock.send(b"exit\0")
diff --git a/tools/testing/selftests/drivers/net/psp_responder.c b/tools/testing/selftests/drivers/net/psp_responder.c
index 57425ecb9561..b3e0fe6d6975 100644
--- a/tools/testing/selftests/drivers/net/psp_responder.c
+++ b/tools/testing/selftests/drivers/net/psp_responder.c
@@ -5,6 +5,7 @@
 #include <sys/poll.h>
 #include <sys/socket.h>
 #include <sys/time.h>
+#include <arpa/inet.h>
 #include <netinet/in.h>
 #include <unistd.h>
 
@@ -23,6 +24,7 @@ static bool should_quit;
 struct opts {
 	int port;
 	int ifindex;
+	int devid;
 	bool verbose;
 };
 
@@ -155,6 +157,100 @@ static void send_str(int sock, int value)
 	send(sock, buf, ret + 1, MSG_WAITALL);
 }
 
+#define PSP_MAX_KEY_LEN		32
+
+struct assoc_msg {
+	__be32 spi;
+	__u8 version;
+	__u8 pad[3];
+	char key[PSP_MAX_KEY_LEN];
+};
+
+static void
+handle_rx_assoc(struct ynl_sock *ys, int data_sock, int comm_sock)
+{
+	struct assoc_msg msg = {};
+	__u32 spi;
+
+	if (data_sock < 0) {
+		fprintf(stderr, "WARN: rx assoc but no data sock\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	if (rx_assoc(ys, &spi, msg.key, data_sock)) {
+		fprintf(stderr, "ERROR: rx_assoc() failed\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	msg.spi = htonl(spi);
+	msg.version = psp_vers.rx;
+	send_ack(comm_sock);
+	send(comm_sock, &msg, sizeof(msg), MSG_WAITALL);
+}
+
+static void
+handle_tx_assoc(struct ynl_sock *ys, char *data, int data_sock, int comm_sock)
+{
+	struct assoc_msg msg;
+
+	if (data_sock < 0) {
+		fprintf(stderr, "WARN: tx assoc but no data sock\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	memcpy(&msg, data, sizeof(msg));
+	if (tx_assoc(ys, msg.version, ntohl(msg.spi), msg.key, data_sock)) {
+		fprintf(stderr, "ERROR: tx_assoc() failed!\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	send_ack(comm_sock);
+}
+
+static int rotate_key(struct ynl_sock *ys, int devid)
+{
+	struct psp_key_rotate_rsp *rsp;
+	struct psp_key_rotate_req *req;
+
+	req = psp_key_rotate_req_alloc();
+
+	psp_key_rotate_req_set_id(req, devid);
+
+	rsp = psp_key_rotate(ys, req);
+	psp_key_rotate_req_free(req);
+
+	if (!rsp) {
+		perror("ERROR: failed to rotate key");
+		return -1;
+	}
+
+	psp_key_rotate_rsp_free(rsp);
+
+	return 0;
+}
+
+static void
+handle_key_rotate(struct ynl_sock *ys, struct opts *opts, int comm_sock)
+{
+	if (opts->devid < 0) {
+		fprintf(stderr, "WARN: key rotate but no PSP device\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	if (rotate_key(ys, opts->devid)) {
+		fprintf(stderr, "ERROR: rotate_key() failed\n");
+		send_err(comm_sock);
+		return;
+	}
+
+	send_ack(comm_sock);
+}
+
 static void
 run_session(struct ynl_sock *ys, struct opts *opts,
 	    int server_sock, int comm_sock)
@@ -247,6 +343,9 @@ run_session(struct ynl_sock *ys, struct opts *opts,
 			match;						\
 		})
 
+#define cmd_w_msg(_name, _type)						\
+		(off >= sizeof(_name) + sizeof(_type) && cmd(_name))
+
 			do {
 				consumed = false;
 
@@ -261,6 +360,16 @@ run_session(struct ynl_sock *ys, struct opts *opts,
 						fprintf(stderr, "WARN: echo but no data sock\n");
 					send_ack(comm_sock);
 				}
+				if (cmd("rx assoc"))
+					handle_rx_assoc(ys, data_sock,
+							comm_sock);
+				if (cmd_w_msg("tx assoc", struct assoc_msg)) {
+					handle_tx_assoc(ys, buf, data_sock,
+							comm_sock);
+					__consume(sizeof(struct assoc_msg));
+				}
+				if (cmd("key rotate"))
+					handle_key_rotate(ys, opts, comm_sock);
 				if (cmd("data close")) {
 					if (data_sock >= 0) {
 						close(data_sock);
@@ -291,6 +400,7 @@ run_session(struct ynl_sock *ys, struct opts *opts,
 				}
 				if (cmd("exit"))
 					should_quit = true;
+#undef cmd_w_msg
 #undef cmd
 
 				if (!consumed) {
@@ -486,6 +596,7 @@ int main(int argc, char **argv)
 		}
 	}
 	psp_dev_get_list_free(dev_list);
+	opts.devid = devid;
 
 	if (opts.ifindex && devid < 0)
 		fprintf(stderr,

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v2 7/7] selftests: drv-net: psp: add a tx rekey drain test for SADB drivers
  2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
                   ` (5 preceding siblings ...)
  2026-10-09 20:46 ` [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests Daniel Zahka
@ 2026-10-09 20:46 ` Daniel Zahka
  6 siblings, 0 replies; 14+ messages in thread
From: Daniel Zahka @ 2026-10-09 20:46 UTC (permalink / raw)
  To: Jakub Kicinski, Willem de Bruijn, David S. Miller, Eric Dumazet,
	Paolo Abeni, Simon Horman, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Donald Hunter, Andrew Lunn, Shuah Khan
  Cc: netdev, linux-kernel, linux-doc, linux-kselftest

Drivers that export the tx-key-count stat, should drain queued tx key
deletions in a timely manner.

Add a test that performs multiple tx rekeys on a connection and waits
to make sure tx-key-count returns to pre-rekey level.

There is a bit of noise in tx-key-count due to stale timewait sockets
from prior tests. If these sockets die between the initial and final
readings, it could only serve to cover up for a queued key deletion
that is not actually drained, so the test should not flake, but may
fail to detect an actual driver bug.

This test will likely not work if other processes on the system are
using psp on the DUT, because then the tx-key-count will be completely
unrelated to what we are doing in our tests.

Signed-off-by: Daniel Zahka <daniel.zahka@gmail.com>
---
 tools/testing/selftests/drivers/net/psp.py | 38 +++++++++++++++++++++++++++++-
 1 file changed, 37 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/drivers/net/psp.py b/tools/testing/selftests/drivers/net/psp.py
index af67c9e8c40f..b22891b7f5ce 100755
--- a/tools/testing/selftests/drivers/net/psp.py
+++ b/tools/testing/selftests/drivers/net/psp.py
@@ -15,7 +15,7 @@ from contextlib import contextmanager
 
 from lib.py import defer
 from lib.py import ksft_run, ksft_exit, ksft_pr
-from lib.py import ksft_true, ksft_eq, ksft_ne, ksft_gt, ksft_raises
+from lib.py import ksft_true, ksft_eq, ksft_ne, ksft_ge, ksft_gt, ksft_raises
 from lib.py import ksft_not_none
 from lib.py import ksft_variants, KsftNamedVariant
 from lib.py import KsftSkipEx, KsftFailEx
@@ -30,6 +30,8 @@ TCP_ULP = 31
 
 _PSP_MAX_KEY_LEN = 32
 _PSP_ASSOC_MSG = f'!IB3x{_PSP_MAX_KEY_LEN}s'
+_TX_REKEY_ROUNDS = 20
+_TX_KEY_DRAIN_TIMEOUT = 5
 
 
 def _get_outq(s):
@@ -202,6 +204,19 @@ def _require_version(cfg, version):
 def _get_stat(cfg, key):
     return cfg.pspnl.get_stats({'dev-id': cfg.psp_dev_id})[key]
 
+
+def _get_tx_key_count(cfg):
+    return cfg.pspnl.get_stats({'dev-id': cfg.psp_dev_id}).get('tx-key-count')
+
+
+def _wait_tx_key_drain(cfg, base):
+    cnt = _get_tx_key_count(cfg)
+    end = time.monotonic() + _TX_KEY_DRAIN_TIMEOUT
+    while cnt > base and time.monotonic() < end:
+        time.sleep(0.1)
+        cnt = _get_tx_key_count(cfg)
+    ksft_ge(base, cnt, comment="tx keys not removed from device after rekey")
+
 #
 # Test case boiler plate
 #
@@ -730,6 +745,27 @@ def rekey_tx_basic(cfg, version):
         _close_psp_conn(cfg, s)
 
 
+def rekey_tx_drain(cfg):
+    """Test that Tx rekeys do not leak keys on the device"""
+    _init_psp_dev(cfg)
+
+    if _get_tx_key_count(cfg) is None:
+        raise KsftSkipEx("Device does not track Tx keys")
+
+    s = _establish_psp_conn(cfg, 0)
+    try:
+        data_len = _psp_txrx(cfg, s, 1)
+        base = _get_tx_key_count(cfg)
+
+        for _ in range(_TX_REKEY_ROUNDS):
+            _remote_key_rotate(cfg)
+            data_len = _rekey_tx(cfg, s, data_len)
+
+        _wait_tx_key_drain(cfg, base)
+    finally:
+        _close_psp_conn(cfg, s)
+
+
 @ksft_variants(_get_psp_ver_variants())
 def rekey_both_sides(cfg, version):
     """Test rekeying both directions"""

-- 
2.52.0


^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 1/7] psp: support rx rekey operation
  2026-10-09 20:46 ` [PATCH net-next v2 1/7] psp: support rx rekey operation Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Data race on the new field psp_assoc::prev_generation.

--

psp: support rx rekey operation

This commit lets a socket that already has PSP state install a new rx
key. The new psp_assoc copies the existing tx state and keeps the
previous rx spi and generation as a fallback. Packets matching either
the current or the previous rx spi/generation are accepted.

> diff --git a/include/net/psp/functions.h b/include/net/psp/functions.h
> index b23c30898389c..fb80e15e43681 100644
> --- a/include/net/psp/functions.h
> +++ b/include/net/psp/functions.h
> @@ -77,10 +77,12 @@ psp_is_allowed_nondata(struct sk_buff *skb, struct psp_assoc *pas)
>  static inline bool
>  psp_pse_matches_pas(struct psp_skb_ext *pse, struct psp_assoc *pas)
>  {
> -	return pse && pas->rx.spi == pse->spi &&
> -	       pas->generation == pse->generation &&
> -	       pas->version == pse->version &&
> -	       pas->dev_id == pse->dev_id;
> +	return pse && pas->version == pse->version &&
> +	       pas->dev_id == pse->dev_id &&
> +	       ((pas->rx.spi == pse->spi &&
> +		 pas->generation == pse->generation) ||
> +		(pas->prev_spi && pas->prev_spi == pse->spi &&
> +		 pas->prev_generation == pse->generation));
>  }

[ ... ]

> diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
> index a6b1c42dd626f..b5887171c84e7 100644
> --- a/net/psp/psp_sock.c
> +++ b/net/psp/psp_sock.c

[ ... ]

> @@ -304,6 +338,10 @@ void psp_assocs_key_rotated(struct psp_dev *psd)
>  		pas->generation |= ~PSP_GEN_VALID_MASK;
>  		psd->stats.stales++;
>  	}
> +
> +	list_for_each_entry(pas, &psd->active_assocs, assocs_list)
> +		pas->prev_generation |= ~PSP_GEN_VALID_MASK;
> +

[Severity: Low]
Is this new write to prev_generation racing with the reader in
psp_pse_matches_pas()?

psp_assocs_key_rotated() runs with only psd->lock held, which is taken
in the key-rotate pre_doit before psp_nl_key_rotate_doit() calls it.
It does a plain read-modify-write on associations that are already
published through sk->psp_assoc.

The new prev_generation comparison in psp_pse_matches_pas() runs on the
receive path without psd->lock:

  tcp_v4_do_rcv() / tcp_v6_do_rcv()
    psp_sk_rx_policy_check()
      __psp_sk_rx_policy_check()
        psp_pse_matches_pas()

psp_twsk_rx_policy_check() also reads it with only RCU held.

Neither side uses READ_ONCE()/WRITE_ONCE() or data_race(). KCSAN would
probably report this whenever a device key rotation overlaps PSP rx on
a socket that has been rekeyed. Since this is a u8 there is no tearing,
so the worst effect is that a reader briefly still accepts packets on
the previous spi during a rotation.

pas->generation already has the same unannotated pattern, but this
patch adds a second field and a new write loop over active, published
associations. Should these new accesses be annotated with
WRITE_ONCE()/READ_ONCE() or data_race()?

>  	list_splice_init(&psd->prev_assocs, &psd->stale_assocs);
>  	list_splice_init(&psd->active_assocs, &psd->prev_assocs);

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 2/7] psp: support tx rekey operation
  2026-10-09 20:46 ` [PATCH net-next v2 2/7] psp: support tx " Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] In this commit, psp_sock_tx_rekey() (net/psp/psp_sock.c) refuses
  a Tx rekey with -EOPNOTSUPP on every SADB device, i.e. any device…
- [Low] The new 'Rekeying' section in Documentation/networking/psp.rst
  says a connection in 'PSP Full' state can be rekeyed in place with rx-…
- [Low] The new 'Rekeying' section in Documentation/networking/psp.rst
  says the SPI in use before an Rx rekey "remains acceptable on the…

--

psp: support tx rekey operation

This patch adds a Tx rekey path to psp_sock_assoc_set_tx(). The new
psp_sock_tx_rekey() allocates a new psp_assoc that copies the Rx state of
the current assoc and installs the new Tx key. It then swaps the new assoc
into sk->psp_assoc with RCU. The patch also adds a Rekeying section to
Documentation/networking/psp.rst.

> diff --git a/Documentation/networking/psp.rst b/Documentation/networking/psp.rst
> index 0f9b6b73f2447..5c4e4215d906f 100644
> --- a/Documentation/networking/psp.rst
> +++ b/Documentation/networking/psp.rst
> @@ -140,14 +140,36 @@ The PSP assoc state of a socket is not reset when the connection is
>  torn down. ``connect()`` on a socket that has PSP assoc state will
>  return ``-EINVAL``.
>  
> +Rekeying
> +--------
> +
> +A connection which has completed the exchange described above and is
> +in the "PSP Full" state can be rekeyed in place without being torn
> +down. The Tx and Rx directions are rekeyed separately using the same
> +netlink calls as in connection setup.
> +
> +``rx-assoc`` allocates a new Rx key and SPI, which should be passed
> +to the peer exactly as during the initial exchange. The SPI which was
> +in use before the rekey remains acceptable on the socket until the
> +next Rx rekey, so packets the peer sent before it learned the new SPI
> +are still received.

[Severity: Low]
Is "until the next Rx rekey" accurate? A device key rotation seems to end
acceptance of the previous SPI as well.

After an Rx rekey, psp_assoc_create() puts the new assoc on
psd->active_assocs, with prev_generation copied from the old assoc. On the
next rotation, psp_assocs_key_rotated() runs:

	list_for_each_entry(pas, &psd->active_assocs, assocs_list)
		pas->prev_generation |= ~PSP_GEN_VALID_MASK;

psp_pse_matches_pas() only accepts the previous SPI when the generation
also matches:

	(pas->prev_spi && pas->prev_spi == pse->spi &&
	 pas->prev_generation == pse->generation)

So after the next device key rotation, packets that use the previous SPI
would be dropped, even if there has been no further Rx rekey.

Should the doc say that the old SPI also stops being accepted at the next
device key rotation?

> +
> +``tx-assoc`` installs a new Tx key, which takes effect immediately.
> +The peer will accept data encrypted from the old and new SPI.
> +
> +A rekey is rejected with ``-EINVAL`` if it would change the PSP device
> +or the PSP version of the association. It is rejected with ``-EBUSY``
> +if the socket has not completed its initial exchange in both
> +directions.

[Severity: Low]
Does this list cover every -EINVAL case for an Rx rekey?
psp_sock_rx_rekey() also rejects a new SPI whose key phase bit matches the
current one:

	if (!((pas->rx.spi ^ prev->rx.spi) & cpu_to_be32(PSP_SPI_KEY_PHASE))) {
		NL_SET_ERR_MSG(extack, "New and prev SPI have same phase bit");
		return -EINVAL;
	}

Drivers set the phase bit from the current device key. In netdevsim, for
example, nsim_rx_spi_alloc() stays inside the current phase, and only
nsim_key_rotate() flips PSP_SPI_KEY_PHASE.

If the device key has not rotated since the current SPI was allocated, an
rx-assoc rekey will always fail with -EINVAL. An application that rekeys on
its own schedule, instead of after a rotation notification, would hit this.

Could this requirement be documented here? The text is still the same at
the end of the series.

[ ... ]

> diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
> index b5887171c84e7..060149f3d72a7 100644
> --- a/net/psp/psp_sock.c
> +++ b/net/psp/psp_sock.c
> @@ -283,6 +283,52 @@ psp_sock_set_tx_key(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
>  	return err;
>  }
>  
> +static int
> +psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
> +		  struct psp_key_parsed *key, struct netlink_ext_ack *extack)
> +{
> +	struct psp_assoc *new;
> +	int err;
> +
> +	if (psp_dev_has_sadb(psd)) {
> +		NL_SET_ERR_MSG(extack, "Tx rekey not supported on this device");
> +		return -EOPNOTSUPP;
> +	}

[Severity: Low]
At this commit, should this limitation be mentioned in the commit message
or in the new Rekeying section?

psp_dev_has_sadb() returns true whenever psd->ops->tx_key_del is set. That
includes mlx5, the only in-tree hardware driver: mlx5_psp_ops sets
.tx_key_del = mlx5e_psp_assoc_del.

So on mlx5 this call chain returns -EOPNOTSUPP:

psp_nl_tx_assoc_doit()->psp_sock_assoc_set_tx()->psp_sock_tx_rekey()

However, psp.rst lists only -EINVAL and -EBUSY as reasons a rekey can be
rejected.

The next patch in the series, "psp: defer tx key deletions for SADB
drivers", removes this check and sets PSP_ASSOC_DEFER_TX_KEY_DEL on the
old assoc instead. The mismatch therefore only exists partway through the
series.

> +	if (!pas->peer_tx) {
> +		NL_SET_ERR_MSG(extack, "Socket PSP state is not fully established");
> +		return -EBUSY;
> +	}

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers
  2026-10-09 20:46 ` [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] Removing the SADB gate in psp_sock_tx_rekey() lets an
  unprivileged user who owns an established PSP TCP socket rekey in a…
- [Low] psp_dev_unregister() (net/psp/psp_main.c) disables the grace-
  period work with psp_deferred_del_stop() and then calls…

--

psp: defer tx key deletions for SADB drivers

When a Tx rekey detaches an old psp_assoc from its socket, this patch
waits for a BQL-based descriptor grace period before deleting the old key
from an SADB device. Old keys are queued on psd->tx_del, and a delayed
work polls the dql counters before calling psp_dev_tx_key_del(). The patch
also allows Tx rekey on SADB devices and makes INET_PSP depend on BQL.

> diff --git a/net/psp/psp_deferred_del.c b/net/psp/psp_deferred_del.c
> new file mode 100644
> index 0000000000000..cfbac72577e96
> --- /dev/null
> +++ b/net/psp/psp_deferred_del.c
> @@ -0,0 +1,231 @@

[ ... ]

> +void psp_deferred_del_stop(struct psp_dev *psd)
> +{
> +	disable_delayed_work_sync(&psd->tx_del.work);
> +}
> +
> +void psp_deferred_del_uninit(struct psp_dev *psd, bool unpublished)
> +{
> +	struct psp_assoc *pas, *next;
> +
> +	lockdep_assert(unpublished || lockdep_is_held(&psd->lock));
> +
> +	psp_txq_state_free(psd->tx_del.txq_state);
> +	psd->tx_del.txq_state = NULL;
> +
> +	list_splice_init(&psd->tx_del.active, &psd->tx_del.next);
> +	list_for_each_entry_safe(pas, next, &psd->tx_del.next, assocs_list) {
> +		list_del(&pas->assocs_list);
> +		psp_dev_tx_key_del(psd, pas);
> +		psp_dev_put(psd);
> +		kfree(pas);
> +	}
> +}

[Severity: Low]

psp_dev_unregister() calls psp_deferred_del_stop() to disable the poll
work. It then calls psp_deferred_del_uninit(psd, false), which runs
psp_dev_tx_key_del() on every key on tx_del.active and tx_del.next. It
does not check the BQL counters first.

These keys were deferred because Tx descriptors might still reference
them. If an SADB driver calls psp_dev_unregister() while PSP Tx
descriptors are still in flight, would the SADB entries be removed while
the hardware can still use them?

No in-tree driver hits this today. mlx5e_nic_disable() calls mlx5e_close()
before mlx5e_psp_unregister(), and netdevsim has no tx_key_add.

Before this patch, psp_dev_unregister() already deleted keys for live
sockets on the active, prev and stale lists. So the unwritten rule "quiesce
Tx before unregister" already existed.

However, the new "Deferred Tx key deletion" section in
Documentation/networking/psp.rst says PSP core waits until descriptors
are completed. The commit message sends driver authors to its list of
requirements, and that list does not include this precondition.

Could a requirement be added? It would say that all PSP Tx descriptors
must be completed or discarded before psp_dev_unregister() is called.

[ ... ]

> diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
> index 060149f3d72a7..d21f7afd830a3 100644
> --- a/net/psp/psp_sock.c
> +++ b/net/psp/psp_sock.c

[ ... ]

> @@ -290,10 +296,6 @@ psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
>  	struct psp_assoc *new;
>  	int err;
>  
> -	if (psp_dev_has_sadb(psd)) {
> -		NL_SET_ERR_MSG(extack, "Tx rekey not supported on this device");
> -		return -EOPNOTSUPP;
> -	}
>  	if (!pas->peer_tx) {
>  		NL_SET_ERR_MSG(extack, "Socket PSP state is not fully established");
>  		return -EBUSY;

[ ... ]

> @@ -324,6 +326,7 @@ psp_sock_tx_rekey(struct sock *sk, struct psp_dev *psd, struct psp_assoc *pas,
>  	list_add(&new->assocs_list, &pas->assocs_list);
>  
>  	rcu_assign_pointer(sk->psp_assoc, new);
> +	pas->flags |= PSP_ASSOC_DEFER_TX_KEY_DEL;
>  	psp_assoc_put(pas);
>  
>  	return 0;

[Severity: Medium]

With the SADB check removed, can an unprivileged user fill the device
SADB by rekeying one socket in a loop?

PSP_CMD_TX_ASSOC has no GENL_ADMIN_PERM. psp_assoc_device_get_locked()
only checks socket ownership and netns visibility, and userspace supplies
the Tx key and SPI. Each call installs a new hardware key:

psp_sock_assoc_set_tx()
  psp_sock_tx_rekey()
    psp_assoc_set_tx()
      psp_dev_tx_key_add()

The old assoc is marked PSP_ASSOC_DEFER_TX_KEY_DEL. After the RCU
callback, psp_assoc_free() only moves it onto psd->tx_del.next via
psp_deferred_del_queue():

	list_move_tail(&pas->assocs_list, &psd->tx_del.next);
	schedule_delayed_work(&psd->tx_del.work, 0);

While the 1000 ms poll is pending, schedule_delayed_work() does nothing.
So a key queued during an active grace period waits for that grace period
and then one more full grace period. If any Tx queue's BQL counters stop
advancing, the keys stay installed forever.

Nothing appears to limit how many deferred keys can be outstanding per
socket, per netns or per psp_dev. One connection could then hold about
(rekey rate x 1-2 s) SADB entries. psp_dev_tx_key_add() calls from other
sockets or netns sharing the psp_dev could then fail.

Before this patch, SADB exhaustion was already partly possible with many
connections. TIME_WAIT sockets also hold psp_assoc references for up to
60 s, and psp_assoc memory is charged to the memcg. The rekey path still
adds a new way for a single connection to multiply its usage.

Would it make sense to cap or rate-limit pending deferred deletions?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys
  2026-10-09 20:46 ` [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] psp_dev_tx_key_del() (net/psp/psp_sock.c) calls
  psd->ops->tx_key_del(psd, pas) and then always decrements…

--

psp: add core tracked stat for outstanding tx keys

This commit adds a core-maintained tx_key_cnt to psd->stats. The counter
goes up when tx_key_add() succeeds and down in psp_dev_tx_key_del(). It is
reported to userspace as tx-key-count, but only for drivers that use an
SADB, and psp_dev_unregister() warns if any keys are still outstanding.

> diff --git a/Documentation/netlink/specs/psp.yaml b/Documentation/netlink/specs/psp.yaml
> index f3266763c3253..d14c11d098914 100644
> --- a/Documentation/netlink/specs/psp.yaml
> +++ b/Documentation/netlink/specs/psp.yaml
> @@ -190,6 +190,13 @@ attribute-sets:
>          doc: |
>            Number of PSP packets for transmission with errors.
>            Device statistic (from the PSP spec).
> +      -
> +        name: tx-key-count
> +        type: uint
> +        doc: |
> +          Current number of Tx keys installed on the device.
> +          Only visible if the driver stores keys on device.
> +          Kernel statistic.

[ ... ]

> diff --git a/net/psp/psp_sock.c b/net/psp/psp_sock.c
> index d21f7afd830a3..c458bc551b4b9 100644
> --- a/net/psp/psp_sock.c
> +++ b/net/psp/psp_sock.c

[ ... ]

> @@ -105,6 +107,8 @@ static int psp_dev_tx_key_add(struct psp_dev *psd, struct psp_assoc *pas,
>  void psp_dev_tx_key_del(struct psp_dev *psd, struct psp_assoc *pas)
>  {
>  	psd->ops->tx_key_del(psd, pas);
> +	if (!WARN_ON_ONCE(!psd->stats.tx_key_cnt))
> +		psd->stats.tx_key_cnt--;
>  }

[Severity: Low]
Is tx-key-count really the number of keys installed on the device? The
YAML doc says "Current number of Tx keys installed on the device". The
commit message says it counts keys "from the hw's perspective" and can show
whether "a device is approaching its SADB capacity".

struct psp_dev_ops::tx_key_del returns void, so psp_dev_tx_key_del()
decrements the counter whether or not the device actually freed the entry.

In mlx5, which is the only SADB driver, the delete path drops any firmware
error:

drivers/net/ethernet/mellanox/mlx5/core/en_accel/psp.c:mlx5e_psp_assoc_del() {
    ...
	nkey = (struct psp_key *)pas->drv_data;
	mlx5_destroy_encryption_key(priv->mdev, nkey->id);
	atomic_dec(&psp->tx_key_cnt);
}

drivers/net/ethernet/mellanox/mlx5/core/lib/crypto.c:mlx5_crypto_destroy_dek_key() {
    ...
	mlx5_cmd_exec(mdev, in, sizeof(in), out, sizeof(out));
}

The call path is psp_assoc_free(), psp_deferred_del_work() or
psp_deferred_del_uninit(), then psp_dev_tx_key_del(), then
mlx5e_psp_assoc_del(), then mlx5_destroy_encryption_key().

If MLX5_CMD_OP_DESTROY_GENERAL_OBJECT fails (bad status or timeout), the
DEK can stay allocated on the device while tx_key_cnt still goes down. In
that case the stat would read lower than real SADB occupancy.

The counter still catches missing tx_key_del calls in the core. It does
not catch failed deletes on the device side.

Should the doc and commit message describe this as a count of keys the
stack believes are installed? The alternative would be for tx_key_del to
return an error, so the core can skip the decrement when the device-side
delete fails.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup
  2026-10-09 20:46 ` [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] In psp_responder.c, conn_setup_psp() fills and exchanges key data
  sized for the Rx version, then hands the same buffer to the new…

--

selftests: drv-net: psp: factor out psp connection setup

This commit moves the PSP Rx and Tx association code in psp_responder.c
into new rx_assoc() and tx_assoc() helpers. It also moves the psp.py key
exchange into _establish_psp_conn(), so other tests can reuse the key
exchange and then do more assoc calls.

> diff --git a/tools/testing/selftests/drivers/net/psp_responder.c b/tools/testing/selftests/drivers/net/psp_responder.c
> index a26e7628bbb13..57425ecb9561f 100644
> --- a/tools/testing/selftests/drivers/net/psp_responder.c
> +++ b/tools/testing/selftests/drivers/net/psp_responder.c

[ ... ]

> @@ -64,29 +68,34 @@ static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
>  		return -1;
>  	}
>  
> -	// SPI exchange
> -	key_len = rsp->rx_key._len.key;
> -	memcpy(info, &rsp->rx_key.spi, sizeof(spi));
> -	memcpy(&info[sizeof(spi)], rsp->rx_key.key, key_len);
> -	sz = sizeof(spi) + key_len;
> +	if (rsp->rx_key._len.key != key_len) {
> +		fprintf(stderr, "ERROR: unexpected Rx key length %u\n",
> +			rsp->rx_key._len.key);
> +		psp_rx_assoc_rsp_free(rsp);
> +		return -1;
> +	}

[ ... ]

> -	// Setup Tx assoc
>  	teq = psp_tx_assoc_req_alloc();
>  
>  	psp_tx_assoc_req_set_sock_fd(teq, data_sock);
> -	psp_tx_assoc_req_set_version(teq, psp_vers.tx);
> +	psp_tx_assoc_req_set_version(teq, version);
>  	psp_tx_assoc_req_set_tx_key_spi(teq, spi);
> -	psp_tx_assoc_req_set_tx_key_key(teq, &info[sizeof(spi)], key_len);
> +	psp_tx_assoc_req_set_tx_key_key(teq, key, psp_key_len(version));

[ ... ]

> @@ -99,6 +108,34 @@ static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
>  	return 0;
>  }
>  
> +static int conn_setup_psp(struct ynl_sock *ys, struct opts *opts, int data_sock)
> +{
> +	unsigned int key_len = psp_key_len(psp_vers.rx);
> +	char info[300];
> +	ssize_t sz;
> +	__u32 spi;
> +
> +	dbg("create PSP connection\n");
> +
> +	if (rx_assoc(ys, &spi, &info[sizeof(spi)], data_sock))
> +		return -1;
> +
> +	// SPI exchange
> +	memcpy(info, &spi, sizeof(spi));
> +	sz = sizeof(spi) + key_len;
> +
> +	send(data_sock, info, sz, MSG_WAITALL);
> +
> +	sz = recv(data_sock, info, sz, MSG_WAITALL);
> +	if (sz < 0) {
> +		perror("ERROR: failed to read PSP key from sock");
> +		return -1;
> +	}
> +	memcpy(&spi, info, sizeof(spi));
> +
> +	return tx_assoc(ys, psp_vers.tx, spi, &info[sizeof(spi)], data_sock);

[Severity: Low]
Can tx_assoc() read uninitialized bytes from info[] here if psp_vers.tx
and psp_vers.rx have different key sizes?

conn_setup_psp() sizes the exchange from the Rx version:

	unsigned int key_len = psp_key_len(psp_vers.rx);
	...
	sz = sizeof(spi) + key_len;

So rx_assoc() and recv() fill only key_len bytes after the SPI. Then
tx_assoc() copies a length based on the Tx version:

	psp_tx_assoc_req_set_tx_key_key(teq, key, psp_key_len(version));

The two versions are separate bytes read in run_session(), and nothing
checks that they match:

	memcpy(&psp_vers, buf, 2);

With a 256-bit Tx version and a 128-bit Rx version, 16 bytes of the
uninitialized info[] stack array would go into the tx-key attribute.
Before this patch, the Tx key length was rsp->rx_key._len.key, which is
the same length used for the exchange.

The kernel should reject that request in psp_sock_assoc_set_tx() because
the versions don't match. Also, the only in-tree sender, _make_psp_conn()
in psp.py, always sends equal versions. Later patches in the series pass a
fully initialized 32-byte key to tx_assoc() from handle_tx_assoc(), but
they leave conn_setup_psp() as is. So the mismatch is still there at the
end of the series.

Would it make sense to reject Tx/Rx versions with different key sizes in
conn_setup_psp(), or to pass the valid key length to tx_assoc()?

Separately, the commit message calls this a factor-out. However,
rx_assoc() now fails when rsp->rx_key._len.key doesn't match
psp_key_len(psp_vers.rx), and the Tx key length is now computed
differently. Could the commit message mention these behaviour changes?

> +}
> +
>  static void send_ack(int sock)
>  {
>  	send(sock, "ack", 4, MSG_WAITALL);

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests
  2026-10-09 20:46 ` [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests Daniel Zahka
@ 2026-10-10 21:43   ` netdev-bot+sashiko
  0 siblings, 0 replies; 14+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 21:43 UTC (permalink / raw)
  To: daniel.zahka
  Cc: kuba, willemdebruijn.kernel, davem, edumazet, pabeni, horms,
	corbet, skhan, rdunlap, donald.hunter, andrew+netdev, shuah,
	netdev, linux-kernel, linux-doc, linux-kselftest

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] In run_session()
  (tools/testing/selftests/drivers/net/psp_responder.c), the cmd() and…
- [Low] _recv_careful() in tools/testing/selftests/drivers/net/psp.py
  never detects that the peer has closed.
- [Low] rekey_tx_basic() (and, for the responder's tx-assoc,
  rekey_rx_basic()) would still pass if the Tx rekey left the old…
- [Low] selftests: drv-net: NULL pointer dereference if
  psp_key_rotate_req_alloc() fails

--

selftests: drv-net: psp: add rekey tests

This adds selftests for rekeying PSP connections. They include cases
where rx-assoc and tx-assoc should be rejected because of a PSP version
or state mismatch. It also adds rx assoc, tx assoc and key rotate
commands to psp_responder so the remote side can be rekeyed.

> diff --git a/tools/testing/selftests/drivers/net/psp.py b/tools/testing/selftests/drivers/net/psp.py
> index 71eeade62ad24..af67c9e8c40fb 100755
> --- a/tools/testing/selftests/drivers/net/psp.py
> +++ b/tools/testing/selftests/drivers/net/psp.py

[ ... ]

> @@ -130,6 +164,41 @@ def _check_data_outq(s, exp_len, force_wait=False):
>      ksft_eq(outq, exp_len)
>  
>  
> +def _recv_careful(s, target, timeout=2):
> +    """Read exactly target bytes, tolerating short reads"""
> +    data = b''
> +    end = time.monotonic() + timeout
> +    while time.monotonic() < end:
> +        try:
> +            data += s.recv(target - len(data), socket.MSG_DONTWAIT)
> +            if len(data) == target:
> +                return data
> +        except BlockingIOError:
> +            time.sleep(0.001)
> +    raise KsftFailEx(f"short read, got {len(data)} of {target} bytes")

[Severity: Low]
What happens here if the peer closes the data connection cleanly?
s.recv() with MSG_DONTWAIT then returns b'' right away instead of raising
BlockingIOError. So data never grows and the time.sleep() branch never
runs.

Won't the loop then busy-spin for the whole 2 second timeout and report
"short read, got 0 of 5 bytes" instead of reporting the close?

_recv_all(), added in this same patch, handles this case:

    if not chunk:
        raise KsftFailEx(f"peer closed after {len(data)} of {target} bytes")

Could _recv_careful() do the same? Every new rekey test reaches it via
_psp_txrx()->_req_echo()->_recv_careful().

[ ... ]

> +def _rekey_tx(cfg, s, sent):
> +    """Rekey the Tx direction, running traffic after each step"""
> +    version, tx = _remote_rx_assoc(cfg)
> +    sent = _psp_txrx(cfg, s, 1, sent)
> +
> +    cfg.pspnl.tx_assoc({"dev-id": cfg.psp_dev_id,
> +                        "version": version,
> +                        "tx-key": tx,
> +                        "sock-fd": s.fileno()})
> +    return _psp_txrx(cfg, s, 1, sent)

[Severity: Low]
Can this check tell whether the new Tx key is actually in use?

After the receiver's rx-assoc rekey, psp_sock_rx_rekey() keeps the
previous association valid:

net/psp/psp_sock.c:psp_sock_rx_rekey() {
    ...
    pas->prev_spi = prev->rx.spi;
    pas->prev_generation = prev->generation;
    ...
}

psp_pse_matches_pas() then accepts packets that match it:

    (pas->prev_spi && pas->prev_spi == pse->spi &&
     pas->prev_generation == pse->generation)

After a single key rotation the old generation is still valid. If
tx-assoc left the old SPI and key on the wire, the byte count and echo
in _psp_txrx() would still pass. This applies to rekey_tx_basic(), and
also to the responder's tx assoc step in rekey_rx_basic().

rekey_both_sides() does catch a stale Tx key indirectly, because its
second rotation invalidates the original generation. Should the basic
variants also check that the new SPI is in use?

[ ... ]

> diff --git a/tools/testing/selftests/drivers/net/psp_responder.c b/tools/testing/selftests/drivers/net/psp_responder.c
> index 57425ecb9561f..b3e0fe6d69750 100644
> --- a/tools/testing/selftests/drivers/net/psp_responder.c
> +++ b/tools/testing/selftests/drivers/net/psp_responder.c

[ ... ]

> @@ -247,6 +343,9 @@ run_session(struct ynl_sock *ys, struct opts *opts,
>  			match;						\
>  		})
>  
> +#define cmd_w_msg(_name, _type)						\
> +		(off >= sizeof(_name) + sizeof(_type) && cmd(_name))
> +
>  			do {
>  				consumed = false;
>  
> @@ -261,6 +360,16 @@ run_session(struct ynl_sock *ys, struct opts *opts,
>  						fprintf(stderr, "WARN: echo but no data sock\n");
>  					send_ack(comm_sock);
>  				}
> +				if (cmd("rx assoc"))
> +					handle_rx_assoc(ys, data_sock,
> +							comm_sock);
> +				if (cmd_w_msg("tx assoc", struct assoc_msg)) {
> +					handle_tx_assoc(ys, buf, data_sock,
> +							comm_sock);
> +					__consume(sizeof(struct assoc_msg));
> +				}

[Severity: Medium]
Could a later cmd() check match the binary assoc_msg payload that is
left in buf?

In run_session(), the cmd() and __consume() macros compare against n.
n is set only once, right after recv():

    off += n;
    n = off;

__consume() never updates it:

    if (n == (sz)) {
        off = 0;
    } else {
        off -= (sz);
        memmove(buf, &buf[(sz)], off);
    }

Suppose "tx assoc" is the last request in the batch, so n == 49.
cmd("tx assoc") leaves off at 40, and __consume(40) leaves off at 0.
The 40 byte assoc_msg is still in buf[0..39], though, and n is still 49.

The cmd() checks that follow in the same pass ("key rotate", "data
close", "conn psp", "conn clr", "exit") still see n >= sz, so their
memcmp() runs against that payload. With an SPI of 0x65786974 and
version 0, buf[0..4] is "exit\0". cmd("exit") then matches, sets
should_quit, and its __consume(5) does:

    off -= 5;                       /* off == -5 */
    memmove(buf, &buf[5], off);     /* size is (size_t)-5 */

Won't that memmove() run far out of bounds and crash the responder? The
same thing can happen when handle_tx_assoc() fails and returns after
send_err(), because parsing continues either way.

cmd_w_msg() already checks off rather than n. Should cmd() and
__consume() use off too?

[ ... ]

> @@ -155,6 +157,100 @@ static void send_str(int sock, int value)
>  	send(sock, buf, ret + 1, MSG_WAITALL);
>  }
>  
> +#define PSP_MAX_KEY_LEN		32
> +
> +struct assoc_msg {
> +	__be32 spi;
> +	__u8 version;
> +	__u8 pad[3];
> +	char key[PSP_MAX_KEY_LEN];
> +};
> +
> +static void
> +handle_rx_assoc(struct ynl_sock *ys, int data_sock, int comm_sock)
> +{
> +	struct assoc_msg msg = {};
> +	__u32 spi;
> +
> +	if (data_sock < 0) {
> +		fprintf(stderr, "WARN: rx assoc but no data sock\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	if (rx_assoc(ys, &spi, msg.key, data_sock)) {
> +		fprintf(stderr, "ERROR: rx_assoc() failed\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	msg.spi = htonl(spi);
> +	msg.version = psp_vers.rx;
> +	send_ack(comm_sock);
> +	send(comm_sock, &msg, sizeof(msg), MSG_WAITALL);
> +}
> +
> +static void
> +handle_tx_assoc(struct ynl_sock *ys, char *data, int data_sock, int comm_sock)
> +{
> +	struct assoc_msg msg;
> +
> +	if (data_sock < 0) {
> +		fprintf(stderr, "WARN: tx assoc but no data sock\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	memcpy(&msg, data, sizeof(msg));
> +	if (tx_assoc(ys, msg.version, ntohl(msg.spi), msg.key, data_sock)) {
> +		fprintf(stderr, "ERROR: tx_assoc() failed!\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	send_ack(comm_sock);
> +}
> +
> +static int rotate_key(struct ynl_sock *ys, int devid)
> +{
> +	struct psp_key_rotate_rsp *rsp;
> +	struct psp_key_rotate_req *req;
> +
> +	req = psp_key_rotate_req_alloc();
> +
> +	psp_key_rotate_req_set_id(req, devid);
> +
> +	rsp = psp_key_rotate(ys, req);
> +	psp_key_rotate_req_free(req);
> +
> +	if (!rsp) {
> +		perror("ERROR: failed to rotate key");
> +		return -1;
> +	}
> +
> +	psp_key_rotate_rsp_free(rsp);
> +
> +	return 0;
> +}
> +
> +static void
> +handle_key_rotate(struct ynl_sock *ys, struct opts *opts, int comm_sock)
> +{
> +	if (opts->devid < 0) {
> +		fprintf(stderr, "WARN: key rotate but no PSP device\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	if (rotate_key(ys, opts->devid)) {
> +		fprintf(stderr, "ERROR: rotate_key() failed\n");
> +		send_err(comm_sock);
> +		return;
> +	}
> +
> +	send_ack(comm_sock);
> +}
> +
>  static void
>  run_session(struct ynl_sock *ys, struct opts *opts,
>  	    int server_sock, int comm_sock)

[Severity: Low]
psp_key_rotate_req_alloc() is a YNL generated helper that allocates
the request with calloc(), so it can return NULL. rotate_key() passes
the result straight to psp_key_rotate_req_set_id(), which writes the
id field through req without checking it.

If the allocation fails, won't the responder segfault here instead of
taking the existing error path? That path prints a message and calls
send_err() from handle_key_rotate(). On the psp.py side the crash would
show up as a lost comm connection in the middle of a rekey test, not
as a clean failure.

This is unlikely in practice. Would it be enough to check req for NULL
and return -1, the same way the !rsp case already does?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-psp-v2-0-5596ab50f677%40gmail.com

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-10 21:43 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 20:46 [PATCH net-next v2 0/7] psp: support rekeying psp protected tcp connections Daniel Zahka
2026-10-09 20:46 ` [PATCH net-next v2 1/7] psp: support rx rekey operation Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 2/7] psp: support tx " Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 3/7] psp: defer tx key deletions for SADB drivers Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 4/7] psp: add core tracked stat for outstanding tx keys Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 5/7] selftests: drv-net: psp: factor out psp connection setup Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 6/7] selftests: drv-net: psp: add rekey tests Daniel Zahka
2026-10-10 21:43   ` netdev-bot+sashiko
2026-10-09 20:46 ` [PATCH net-next v2 7/7] selftests: drv-net: psp: add a tx rekey drain test for SADB drivers Daniel Zahka

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®