mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: liushike <liushike123@gmail.com>
To: netdev@vger.kernel.org
Cc: jmaloy@redhat.com, tung.quang.nguyen@est.tech,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, horms@kernel.org, shuah@kernel.org,
	tipc-discussion@lists.sourceforge.net,
	linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [RFC net-next 2/2] selftests: net: cover pending TIPC connects on peer node loss
Date: Tue, 29 Sep 2026 06:47:42 +0000	[thread overview]
Message-ID: <20260929064742.576651-3-liushike123@gmail.com> (raw)
In-Reply-To: <20260929064742.576651-1-liushike123@gmail.com>

From: liushike <liushike@ruijie.com.cn>

Exercise a client waiting for a listening server to accept a connection
using two network namespaces and one or two Ethernet veth bearers.

Cover normal acceptance, last-link loss, survival with one remaining
link, loss of both links, rejection, connection-wait timeout followed by
node loss, nonblocking connects, and service-name addressing. Run each
case with SOCK_STREAM and SOCK_SEQPACKET.

Wait for actual unicast links at both peers; the always-UP broadcast
link does not establish peer reachability. Require node-loss completion
within two seconds and check EHOSTUNREACH. For asynchronous failure,
check POLLHUP with poll(), matching TIPC_DISCONNECTING semantics rather
than assuming the socket becomes writable. Save topology diagnostics
when a test raises an exception.

Signed-off-by: liushike <liushike@ruijie.com.cn>
---
 MAINTAINERS                                 |   1 +
 tools/testing/selftests/net/Makefile        |   1 +
 tools/testing/selftests/net/config          |   1 +
 tools/testing/selftests/net/tipc_connect.py | 278 ++++++++++++++++++++
 4 files changed, 281 insertions(+)
 create mode 100755 tools/testing/selftests/net/tipc_connect.py

diff --git a/MAINTAINERS b/MAINTAINERS
index df8ab9b82402..4d595e60a22d 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -27561,6 +27561,7 @@ S:	Maintained
 W:	http://tipc.sourceforge.net/
 F:	include/uapi/linux/tipc*.h
 F:	net/tipc/
+F:	tools/testing/selftests/net/tipc_connect.py
 
 TLAN NETWORK DRIVER
 M:	Samuel Chessman <chessman@tux.org>
diff --git a/tools/testing/selftests/net/Makefile b/tools/testing/selftests/net/Makefile
index 3ee3378f8b26..c453ff976386 100644
--- a/tools/testing/selftests/net/Makefile
+++ b/tools/testing/selftests/net/Makefile
@@ -118,6 +118,7 @@ TEST_PROGS := \
 	test_vxlan_vnifilter_notify.sh \
 	test_vxlan_vnifiltering.sh \
 	tfo_passive.sh \
+	tipc_connect.py \
 	traceroute.sh \
 	txtimestamp.sh \
 	udpgro.sh \
diff --git a/tools/testing/selftests/net/config b/tools/testing/selftests/net/config
index 30d5fcb09a83..725c8fbc5f03 100644
--- a/tools/testing/selftests/net/config
+++ b/tools/testing/selftests/net/config
@@ -128,6 +128,7 @@ CONFIG_TCP_CONG_DCTCP=y
 CONFIG_TCP_MD5SIG=y
 CONFIG_TEST_BLACKHOLE_DEV=m
 CONFIG_TEST_BPF=m
+CONFIG_TIPC=m
 CONFIG_TLS=m
 CONFIG_TRACEPOINTS=y
 CONFIG_TUN=y
diff --git a/tools/testing/selftests/net/tipc_connect.py b/tools/testing/selftests/net/tipc_connect.py
new file mode 100755
index 000000000000..02685b3ce8dd
--- /dev/null
+++ b/tools/testing/selftests/net/tipc_connect.py
@@ -0,0 +1,278 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+
+"""Exercise pending TIPC connects across bearer loss, before accept()."""
+
+import errno
+import os
+import select
+import shutil
+import socket
+import threading
+import time
+from contextlib import ExitStack, contextmanager
+
+from lib.py import (KsftNamedVariant, KsftSkipEx, NetNS, NetNSEnter,
+                    cmd, ip, ksft_eq, ksft_exit, ksft_run, ksft_true,
+                    ksft_variants)
+
+
+SOCKET_TYPES = [KsftNamedVariant("stream", socket.SOCK_STREAM),
+                KsftNamedVariant("seqpacket", socket.SOCK_SEQPACKET)]
+CONNECT_TIMEOUT = 10000
+RETURN_TIMEOUT = 2
+
+
+def tipc(ns, args):
+    return cmd("tipc " + args, ns=ns).stdout
+
+
+def unicast_links_up(output):
+    """Return distinct UP links, excluding the always-UP broadcast link."""
+    links = set()
+    for line in output.splitlines():
+        name, separator, state = line.strip().rpartition(": ")
+        if (separator and state.lower() == "up" and
+                not name.startswith("broadcast-link")):
+            links.add(name)
+    return links
+
+
+def wait_for_links(client_ns, server_ns, links):
+    deadline = time.monotonic() + 10
+    while True:
+        # These isolated namespaces contain only the two test peers. Count
+        # their unicast links, not the namespace's always-UP broadcast link.
+        outputs = [tipc(ns, "link list") for ns in (client_ns, server_ns)]
+        if all(len(unicast_links_up(output)) == links for output in outputs):
+            return
+        if time.monotonic() >= deadline:
+            raise TimeoutError(
+                f"Expected {links} UP unicast links at each peer; "
+                f"client link list: {outputs[0]!r}; "
+                f"server link list: {outputs[1]!r}")
+        time.sleep(0.05)
+
+
+@contextmanager
+def topology(sock_type, links=1):
+    if not hasattr(socket, "AF_TIPC"):
+        raise KsftSkipEx("Python lacks AF_TIPC support")
+    if os.geteuid() != 0:
+        raise KsftSkipEx("root privileges required for network namespaces")
+    for tool in ("ip", "tipc"):
+        if not shutil.which(tool):
+            raise KsftSkipEx(f"{tool} is required")
+    try:
+        with socket.socket(socket.AF_TIPC, sock_type):
+            pass
+    except OSError as error:
+        if error.errno in (errno.EAFNOSUPPORT, errno.EPROTONOSUPPORT):
+            raise KsftSkipEx("TIPC support is unavailable; load tipc") from error
+        raise
+
+    with ExitStack() as stack:
+        client_ns = stack.enter_context(NetNS())
+        server_ns = stack.enter_context(NetNS())
+        for ns, address in ((client_ns, "1.1.1"), (server_ns, "1.1.2")):
+            tipc(ns, "node set netid 4711")
+            tipc(ns, f"node set address {address}")
+        for index in range(links):
+            ip(f"link add c{index} type veth peer name s{index} "
+               f"netns {server_ns}", ns=client_ns)
+            for ns, dev in ((client_ns, f"c{index}"),
+                            (server_ns, f"s{index}")):
+                ip(f"link set {dev} up", ns=ns)
+                tipc(ns, f"bearer enable media eth device {dev}")
+
+        wait_for_links(client_ns, server_ns, links)
+
+        with NetNSEnter(server_ns):
+            listener = stack.enter_context(socket.socket(socket.AF_TIPC,
+                                                        sock_type))
+            listener.listen(8)
+        with NetNSEnter(client_ns):
+            client = stack.enter_context(socket.socket(socket.AF_TIPC,
+                                                      sock_type))
+            client.setsockopt(socket.SOL_TIPC, socket.TIPC_CONN_TIMEOUT,
+                              CONNECT_TIMEOUT)
+        try:
+            yield client_ns, client, listener
+        except Exception:
+            # Collect diagnostics before ExitStack destroys the topology.
+            for label, ns in (("client", client_ns), ("server", server_ns)):
+                for args in ("link list", "node list"):
+                    try:
+                        output = tipc(ns, args)
+                        print(f"# {label} {args}: {output!r}")
+                    except Exception as error:
+                        print(f"# {label} {args} failed: {error}")
+            raise
+
+
+@contextmanager
+def pending_connect(client, listener, address=None):
+    result = []
+    if address is None:
+        address = listener.getsockname()
+
+    def connect():
+        result.append(client.connect_ex(address))
+
+    worker = threading.Thread(target=connect)
+    worker.start()
+    try:
+        # Readability proves SYN is queued on the listener, without accepting it.
+        if not select.select([listener], [], [], 5)[0]:
+            raise TimeoutError(f"SYN did not reach {address!r}: {result}")
+        ksft_eq(result, [], "connect must wait for accept")
+        yield worker, result
+    finally:
+        # Bound cleanup even on the unfixed kernel or a failed assertion.
+        if worker.is_alive():
+            try:
+                client.shutdown(socket.SHUT_RDWR)
+            except OSError:
+                pass
+        worker.join(CONNECT_TIMEOUT / 1000 + 1)
+
+
+def drop_link(ns, index):
+    tipc(ns, f"bearer disable media eth device c{index}")
+
+
+def check_node_loss(client):
+    # TIPC_DISCONNECTING reports POLLHUP, not POLLOUT. Unlike select()'s
+    # write set, poll() reports hangup even when only POLLOUT is requested.
+    poller = select.poll()
+    poller.register(client, select.POLLOUT)
+    events = poller.poll(int(RETURN_TIMEOUT * 1000))
+    ksft_true(events, "node loss must notify pending connect within 2s")
+    if events:
+        ksft_eq(events[0][0], client.fileno())
+        ksft_true(events[0][1] & select.POLLHUP,
+                  f"node loss must report hangup, got {events!r}")
+    ksft_eq(client.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR),
+            errno.EHOSTUNREACH)
+
+
+@ksft_variants(SOCKET_TYPES)
+def normal_accept(sock_type):
+    with topology(sock_type) as (_, client, listener):
+        with pending_connect(client, listener) as (worker, result):
+            with listener.accept()[0] as accepted:
+                worker.join(RETURN_TIMEOUT)
+                ksft_eq(result, [0])
+                if result == [0]:
+                    accepted.settimeout(RETURN_TIMEOUT)
+                    client.sendall(b"connected")
+                    ksft_eq(accepted.recv(32), b"connected")
+
+
+@ksft_variants(SOCKET_TYPES)
+def last_link_down(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        with pending_connect(client, listener) as (worker, result):
+            drop_link(ns, 0)
+            worker.join(RETURN_TIMEOUT)
+            ksft_eq(result, [errno.EHOSTUNREACH],
+                    "node loss must abort connect before its 10s timeout")
+
+
+@ksft_variants(SOCKET_TYPES)
+def surviving_link(sock_type):
+    with topology(sock_type, links=2) as (ns, client, listener):
+        with pending_connect(client, listener) as (worker, result):
+            drop_link(ns, 0)
+            worker.join(0.2)
+            ksft_eq(result, [], "one remaining link must preserve connect")
+            with listener.accept()[0]:
+                worker.join(RETURN_TIMEOUT)
+                ksft_eq(result, [0])
+
+
+@ksft_variants(SOCKET_TYPES)
+def both_links_down(sock_type):
+    with topology(sock_type, links=2) as (ns, client, listener):
+        with pending_connect(client, listener) as (worker, result):
+            drop_link(ns, 0)
+            worker.join(0.2)
+            ksft_eq(result, [])
+            drop_link(ns, 1)
+            worker.join(RETURN_TIMEOUT)
+            ksft_eq(result, [errno.EHOSTUNREACH])
+
+
+@ksft_variants(SOCKET_TYPES)
+def rejected_connect(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        with pending_connect(client, listener) as (worker, result):
+            listener.close()
+            worker.join(RETURN_TIMEOUT)
+            ksft_eq(result, [errno.ECONNREFUSED])
+            drop_link(ns, 0)
+            ksft_eq(client.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR), 0,
+                    "node loss must not change a completed rejection")
+
+
+@ksft_variants(SOCKET_TYPES)
+def connect_timeout_then_node_loss(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        client.setsockopt(socket.SOL_TIPC, socket.TIPC_CONN_TIMEOUT, 500)
+        with pending_connect(client, listener) as (worker, result):
+            worker.join(RETURN_TIMEOUT)
+            ksft_eq(result, [errno.ETIMEDOUT])
+            # A wait timeout leaves the handshake pending, just as EINPROGRESS.
+            drop_link(ns, 0)
+            check_node_loss(client)
+
+
+@ksft_variants(SOCKET_TYPES)
+def nonblocking_connect(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        client.setblocking(False)
+        ksft_eq(client.connect_ex(listener.getsockname()), errno.EINPROGRESS)
+        if not select.select([listener], [], [], 5)[0]:
+            raise TimeoutError("SYN did not reach listener")
+        drop_link(ns, 0)
+        check_node_loss(client)
+
+
+def publish_service(ns, listener):
+    service = 18888
+    listener.bind((socket.TIPC_ADDR_NAMESEQ, service, 1, 1,
+                   socket.TIPC_CLUSTER_SCOPE))
+    deadline = time.monotonic() + 10
+    while str(service) not in tipc(ns, "nametable show").split():
+        if time.monotonic() >= deadline:
+            raise TimeoutError("TIPC service publication did not arrive")
+        time.sleep(0.05)
+    return socket.TIPC_ADDR_NAME, service, 1, 0
+
+
+@ksft_variants(SOCKET_TYPES)
+def named_accept(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        address = publish_service(ns, listener)
+        with pending_connect(client, listener, address) as (worker, result):
+            with listener.accept()[0]:
+                worker.join(RETURN_TIMEOUT)
+                ksft_eq(result, [0])
+
+
+@ksft_variants(SOCKET_TYPES)
+def named_node_loss(sock_type):
+    with topology(sock_type) as (ns, client, listener):
+        address = publish_service(ns, listener)
+        with pending_connect(client, listener, address) as (worker, result):
+            drop_link(ns, 0)
+            worker.join(RETURN_TIMEOUT)
+            ksft_eq(result, [errno.EHOSTUNREACH])
+
+
+if __name__ == "__main__":
+    ksft_run(cases=[normal_accept, last_link_down, surviving_link,
+                    both_links_down, rejected_connect,
+                    connect_timeout_then_node_loss, nonblocking_connect,
+                    named_accept, named_node_loss])
+    ksft_exit()
-- 
2.34.1


  parent reply	other threads:[~2026-09-29  6:48 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  6:47 [RFC net-next 0/2] tipc: notify pending connects of " liushike
2026-09-29  6:47 ` [RFC net-next 1/2] tipc: abort pending connects when the peer node is lost liushike
2026-09-29  6:47 ` liushike [this message]
2026-09-29  6:54 ` [RFC net-next 0/2] tipc: notify pending connects of peer node loss netdev-bot+sinfo
     [not found]   ` <CAL9mPHsf_10zX4Z_2WFHNJob7Tkfnaw6sp03q2+ujnTygn1Ucg@mail.gmail.com>
2026-09-30  2:08     ` shike liu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929064742.576651-3-liushike123@gmail.com \
    --to=liushike123@gmail.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=jmaloy@redhat.com \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=shuah@kernel.org \
    --cc=tipc-discussion@lists.sourceforge.net \
    --cc=tung.quang.nguyen@est.tech \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®