From: liushike <liushike123@gmail.com>
To: netdev@vger.kernel.org
Cc: jmaloy@redhat.com, tung.quang.nguyen@est.tech,
davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, shuah@kernel.org,
tipc-discussion@lists.sourceforge.net,
linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [RFC net-next 2/2] selftests: net: cover pending TIPC connects on peer node loss
Date: Tue, 29 Sep 2026 06:47:42 +0000 [thread overview]
Message-ID: <20260929064742.576651-3-liushike123@gmail.com> (raw)
In-Reply-To: <20260929064742.576651-1-liushike123@gmail.com>
From: liushike <liushike@ruijie.com.cn>
Exercise a client waiting for a listening server to accept a connection
using two network namespaces and one or two Ethernet veth bearers.
Cover normal acceptance, last-link loss, survival with one remaining
link, loss of both links, rejection, connection-wait timeout followed by
node loss, nonblocking connects, and service-name addressing. Run each
case with SOCK_STREAM and SOCK_SEQPACKET.
Wait for actual unicast links at both peers; the always-UP broadcast
link does not establish peer reachability. Require node-loss completion
within two seconds and check EHOSTUNREACH. For asynchronous failure,
check POLLHUP with poll(), matching TIPC_DISCONNECTING semantics rather
than assuming the socket becomes writable. Save topology diagnostics
when a test raises an exception.
Signed-off-by: liushike <liushike@ruijie.com.cn>
---
MAINTAINERS | 1 +
tools/testing/selftests/net/Makefile | 1 +
tools/testing/selftests/net/config | 1 +
tools/testing/selftests/net/tipc_connect.py | 278 ++++++++++++++++++++
4 files changed, 281 insertions(+)
create mode 100755 tools/testing/selftests/net/tipc_connect.py
diff --git a/MAINTAINERS b/MAINTAINERS
index df8ab9b82402..4d595e60a22d 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -27561,6 +27561,7 @@ S: Maintained
W: http://tipc.sourceforge.net/
F: include/uapi/linux/tipc*.h
F: net/tipc/
+F: tools/testing/selftests/net/tipc_connect.py
TLAN NETWORK DRIVER
M: Samuel Chessman <chessman@tux.org>
diff --git a/tools/testing/selftests/net/Makefile b/tools/testing/selftests/net/Makefile
index 3ee3378f8b26..c453ff976386 100644
--- a/tools/testing/selftests/net/Makefile
+++ b/tools/testing/selftests/net/Makefile
@@ -118,6 +118,7 @@ TEST_PROGS := \
test_vxlan_vnifilter_notify.sh \
test_vxlan_vnifiltering.sh \
tfo_passive.sh \
+ tipc_connect.py \
traceroute.sh \
txtimestamp.sh \
udpgro.sh \
diff --git a/tools/testing/selftests/net/config b/tools/testing/selftests/net/config
index 30d5fcb09a83..725c8fbc5f03 100644
--- a/tools/testing/selftests/net/config
+++ b/tools/testing/selftests/net/config
@@ -128,6 +128,7 @@ CONFIG_TCP_CONG_DCTCP=y
CONFIG_TCP_MD5SIG=y
CONFIG_TEST_BLACKHOLE_DEV=m
CONFIG_TEST_BPF=m
+CONFIG_TIPC=m
CONFIG_TLS=m
CONFIG_TRACEPOINTS=y
CONFIG_TUN=y
diff --git a/tools/testing/selftests/net/tipc_connect.py b/tools/testing/selftests/net/tipc_connect.py
new file mode 100755
index 000000000000..02685b3ce8dd
--- /dev/null
+++ b/tools/testing/selftests/net/tipc_connect.py
@@ -0,0 +1,278 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+
+"""Exercise pending TIPC connects across bearer loss, before accept()."""
+
+import errno
+import os
+import select
+import shutil
+import socket
+import threading
+import time
+from contextlib import ExitStack, contextmanager
+
+from lib.py import (KsftNamedVariant, KsftSkipEx, NetNS, NetNSEnter,
+ cmd, ip, ksft_eq, ksft_exit, ksft_run, ksft_true,
+ ksft_variants)
+
+
+SOCKET_TYPES = [KsftNamedVariant("stream", socket.SOCK_STREAM),
+ KsftNamedVariant("seqpacket", socket.SOCK_SEQPACKET)]
+CONNECT_TIMEOUT = 10000
+RETURN_TIMEOUT = 2
+
+
+def tipc(ns, args):
+ return cmd("tipc " + args, ns=ns).stdout
+
+
+def unicast_links_up(output):
+ """Return distinct UP links, excluding the always-UP broadcast link."""
+ links = set()
+ for line in output.splitlines():
+ name, separator, state = line.strip().rpartition(": ")
+ if (separator and state.lower() == "up" and
+ not name.startswith("broadcast-link")):
+ links.add(name)
+ return links
+
+
+def wait_for_links(client_ns, server_ns, links):
+ deadline = time.monotonic() + 10
+ while True:
+ # These isolated namespaces contain only the two test peers. Count
+ # their unicast links, not the namespace's always-UP broadcast link.
+ outputs = [tipc(ns, "link list") for ns in (client_ns, server_ns)]
+ if all(len(unicast_links_up(output)) == links for output in outputs):
+ return
+ if time.monotonic() >= deadline:
+ raise TimeoutError(
+ f"Expected {links} UP unicast links at each peer; "
+ f"client link list: {outputs[0]!r}; "
+ f"server link list: {outputs[1]!r}")
+ time.sleep(0.05)
+
+
+@contextmanager
+def topology(sock_type, links=1):
+ if not hasattr(socket, "AF_TIPC"):
+ raise KsftSkipEx("Python lacks AF_TIPC support")
+ if os.geteuid() != 0:
+ raise KsftSkipEx("root privileges required for network namespaces")
+ for tool in ("ip", "tipc"):
+ if not shutil.which(tool):
+ raise KsftSkipEx(f"{tool} is required")
+ try:
+ with socket.socket(socket.AF_TIPC, sock_type):
+ pass
+ except OSError as error:
+ if error.errno in (errno.EAFNOSUPPORT, errno.EPROTONOSUPPORT):
+ raise KsftSkipEx("TIPC support is unavailable; load tipc") from error
+ raise
+
+ with ExitStack() as stack:
+ client_ns = stack.enter_context(NetNS())
+ server_ns = stack.enter_context(NetNS())
+ for ns, address in ((client_ns, "1.1.1"), (server_ns, "1.1.2")):
+ tipc(ns, "node set netid 4711")
+ tipc(ns, f"node set address {address}")
+ for index in range(links):
+ ip(f"link add c{index} type veth peer name s{index} "
+ f"netns {server_ns}", ns=client_ns)
+ for ns, dev in ((client_ns, f"c{index}"),
+ (server_ns, f"s{index}")):
+ ip(f"link set {dev} up", ns=ns)
+ tipc(ns, f"bearer enable media eth device {dev}")
+
+ wait_for_links(client_ns, server_ns, links)
+
+ with NetNSEnter(server_ns):
+ listener = stack.enter_context(socket.socket(socket.AF_TIPC,
+ sock_type))
+ listener.listen(8)
+ with NetNSEnter(client_ns):
+ client = stack.enter_context(socket.socket(socket.AF_TIPC,
+ sock_type))
+ client.setsockopt(socket.SOL_TIPC, socket.TIPC_CONN_TIMEOUT,
+ CONNECT_TIMEOUT)
+ try:
+ yield client_ns, client, listener
+ except Exception:
+ # Collect diagnostics before ExitStack destroys the topology.
+ for label, ns in (("client", client_ns), ("server", server_ns)):
+ for args in ("link list", "node list"):
+ try:
+ output = tipc(ns, args)
+ print(f"# {label} {args}: {output!r}")
+ except Exception as error:
+ print(f"# {label} {args} failed: {error}")
+ raise
+
+
+@contextmanager
+def pending_connect(client, listener, address=None):
+ result = []
+ if address is None:
+ address = listener.getsockname()
+
+ def connect():
+ result.append(client.connect_ex(address))
+
+ worker = threading.Thread(target=connect)
+ worker.start()
+ try:
+ # Readability proves SYN is queued on the listener, without accepting it.
+ if not select.select([listener], [], [], 5)[0]:
+ raise TimeoutError(f"SYN did not reach {address!r}: {result}")
+ ksft_eq(result, [], "connect must wait for accept")
+ yield worker, result
+ finally:
+ # Bound cleanup even on the unfixed kernel or a failed assertion.
+ if worker.is_alive():
+ try:
+ client.shutdown(socket.SHUT_RDWR)
+ except OSError:
+ pass
+ worker.join(CONNECT_TIMEOUT / 1000 + 1)
+
+
+def drop_link(ns, index):
+ tipc(ns, f"bearer disable media eth device c{index}")
+
+
+def check_node_loss(client):
+ # TIPC_DISCONNECTING reports POLLHUP, not POLLOUT. Unlike select()'s
+ # write set, poll() reports hangup even when only POLLOUT is requested.
+ poller = select.poll()
+ poller.register(client, select.POLLOUT)
+ events = poller.poll(int(RETURN_TIMEOUT * 1000))
+ ksft_true(events, "node loss must notify pending connect within 2s")
+ if events:
+ ksft_eq(events[0][0], client.fileno())
+ ksft_true(events[0][1] & select.POLLHUP,
+ f"node loss must report hangup, got {events!r}")
+ ksft_eq(client.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR),
+ errno.EHOSTUNREACH)
+
+
+@ksft_variants(SOCKET_TYPES)
+def normal_accept(sock_type):
+ with topology(sock_type) as (_, client, listener):
+ with pending_connect(client, listener) as (worker, result):
+ with listener.accept()[0] as accepted:
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [0])
+ if result == [0]:
+ accepted.settimeout(RETURN_TIMEOUT)
+ client.sendall(b"connected")
+ ksft_eq(accepted.recv(32), b"connected")
+
+
+@ksft_variants(SOCKET_TYPES)
+def last_link_down(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ with pending_connect(client, listener) as (worker, result):
+ drop_link(ns, 0)
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [errno.EHOSTUNREACH],
+ "node loss must abort connect before its 10s timeout")
+
+
+@ksft_variants(SOCKET_TYPES)
+def surviving_link(sock_type):
+ with topology(sock_type, links=2) as (ns, client, listener):
+ with pending_connect(client, listener) as (worker, result):
+ drop_link(ns, 0)
+ worker.join(0.2)
+ ksft_eq(result, [], "one remaining link must preserve connect")
+ with listener.accept()[0]:
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [0])
+
+
+@ksft_variants(SOCKET_TYPES)
+def both_links_down(sock_type):
+ with topology(sock_type, links=2) as (ns, client, listener):
+ with pending_connect(client, listener) as (worker, result):
+ drop_link(ns, 0)
+ worker.join(0.2)
+ ksft_eq(result, [])
+ drop_link(ns, 1)
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [errno.EHOSTUNREACH])
+
+
+@ksft_variants(SOCKET_TYPES)
+def rejected_connect(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ with pending_connect(client, listener) as (worker, result):
+ listener.close()
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [errno.ECONNREFUSED])
+ drop_link(ns, 0)
+ ksft_eq(client.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR), 0,
+ "node loss must not change a completed rejection")
+
+
+@ksft_variants(SOCKET_TYPES)
+def connect_timeout_then_node_loss(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ client.setsockopt(socket.SOL_TIPC, socket.TIPC_CONN_TIMEOUT, 500)
+ with pending_connect(client, listener) as (worker, result):
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [errno.ETIMEDOUT])
+ # A wait timeout leaves the handshake pending, just as EINPROGRESS.
+ drop_link(ns, 0)
+ check_node_loss(client)
+
+
+@ksft_variants(SOCKET_TYPES)
+def nonblocking_connect(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ client.setblocking(False)
+ ksft_eq(client.connect_ex(listener.getsockname()), errno.EINPROGRESS)
+ if not select.select([listener], [], [], 5)[0]:
+ raise TimeoutError("SYN did not reach listener")
+ drop_link(ns, 0)
+ check_node_loss(client)
+
+
+def publish_service(ns, listener):
+ service = 18888
+ listener.bind((socket.TIPC_ADDR_NAMESEQ, service, 1, 1,
+ socket.TIPC_CLUSTER_SCOPE))
+ deadline = time.monotonic() + 10
+ while str(service) not in tipc(ns, "nametable show").split():
+ if time.monotonic() >= deadline:
+ raise TimeoutError("TIPC service publication did not arrive")
+ time.sleep(0.05)
+ return socket.TIPC_ADDR_NAME, service, 1, 0
+
+
+@ksft_variants(SOCKET_TYPES)
+def named_accept(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ address = publish_service(ns, listener)
+ with pending_connect(client, listener, address) as (worker, result):
+ with listener.accept()[0]:
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [0])
+
+
+@ksft_variants(SOCKET_TYPES)
+def named_node_loss(sock_type):
+ with topology(sock_type) as (ns, client, listener):
+ address = publish_service(ns, listener)
+ with pending_connect(client, listener, address) as (worker, result):
+ drop_link(ns, 0)
+ worker.join(RETURN_TIMEOUT)
+ ksft_eq(result, [errno.EHOSTUNREACH])
+
+
+if __name__ == "__main__":
+ ksft_run(cases=[normal_accept, last_link_down, surviving_link,
+ both_links_down, rejected_connect,
+ connect_timeout_then_node_loss, nonblocking_connect,
+ named_accept, named_node_loss])
+ ksft_exit()
--
2.34.1
next prev parent reply other threads:[~2026-09-29 6:48 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 6:47 [RFC net-next 0/2] tipc: notify pending connects of " liushike
2026-09-29 6:47 ` [RFC net-next 1/2] tipc: abort pending connects when the peer node is lost liushike
2026-09-29 6:47 ` liushike [this message]
2026-09-29 6:54 ` [RFC net-next 0/2] tipc: notify pending connects of peer node loss netdev-bot+sinfo
[not found] ` <CAL9mPHsf_10zX4Z_2WFHNJob7Tkfnaw6sp03q2+ujnTygn1Ucg@mail.gmail.com>
2026-09-30 2:08 ` shike liu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929064742.576651-3-liushike123@gmail.com \
--to=liushike123@gmail.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=jmaloy@redhat.com \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=shuah@kernel.org \
--cc=tipc-discussion@lists.sourceforge.net \
--cc=tung.quang.nguyen@est.tech \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®