mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress
@ 2026-09-29  8:34 chenyuan_fl
  2026-09-29  8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl
  2026-09-29  8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl
  0 siblings, 2 replies; 5+ messages in thread
From: chenyuan_fl @ 2026-09-29  8:34 UTC (permalink / raw)
  To: netdev, bpf
  Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang,
	linux-kernel

From: Yuan Chen <chenyuan@kylinos.cn>

udp_bpf_recvmsg() re-arms its receive loop on backlog-only skbs that
sk_msg_recvmsg() can never consume: whenever the backlog stays
populated, the loop re-arms on the same skb forever and recvmsg()
spins while holding the socket lock instead of sleeping.

Patch 1 aligns the re-arm condition with TCP and unix_bpf_recvmsg();
patch 2 adds a selftest that reproduces the hang.

Yuan Chen (2):
  bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress
  selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog

 net/ipv4/udp_bpf.c                            |   2 +-
 .../bpf/prog_tests/sockmap_udp_backlog.c      | 141 ++++++++++++++++++
 .../bpf/progs/test_sockmap_udp_backlog.c      |  21 +++
 3 files changed, 163 insertions(+), 1 deletion(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
 create mode 100644 tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c

-- 
2.54.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH bpf-next 1/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress
  2026-09-29  8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl
@ 2026-09-29  8:34 ` chenyuan_fl
  2026-09-30  9:01   ` Alexei Starovoitov
  2026-09-29  8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl
  1 sibling, 1 reply; 5+ messages in thread
From: chenyuan_fl @ 2026-09-29  8:34 UTC (permalink / raw)
  To: netdev, bpf
  Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang,
	linux-kernel

From: Yuan Chen <chenyuan@kylinos.cn>

udp_bpf_recvmsg() re-arms its msg_bytes_ready loop whenever
psock_has_data() is true, but that predicate also covers an skb parked
in psock->ingress_skb, which sk_msg_recvmsg() can never consume: it
only walks psock->ingress_msg.

When the backlog stays populated, e.g. while sk_psock_handle_skb()
keeps returning -EAGAIN, every round gets copied == 0 and
udp_msg_wait_data() returns 1 again from that same skb, so the loop
never sleeps: recvmsg() spins at 100% CPU while holding the socket
lock, ignores SO_RCVTIMEO and never returns to user space.

The wide probe is intended for the entry check and the wait condition,
but re-arming can only make progress from ingress_msg. TCP and
unix_bpf_recvmsg() therefore re-arm on !sk_psock_queue_empty(psock)
and otherwise fall back to the plain receive path. Do the same for UDP
so the reader sleeps and returns -EAGAIN on timeout as expected.

Fixes: 9f2470fbc4cb ("skmsg: Improve udp_bpf_recvmsg() accuracy")
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
---
 net/ipv4/udp_bpf.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/ipv4/udp_bpf.c b/net/ipv4/udp_bpf.c
index ad57c4c9eaab..8aca9fb89334 100644
--- a/net/ipv4/udp_bpf.c
+++ b/net/ipv4/udp_bpf.c
@@ -91,7 +91,7 @@ static int udp_bpf_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
 		timeo = sock_rcvtimeo(sk, flags & MSG_DONTWAIT);
 		data = udp_msg_wait_data(sk, psock, timeo);
 		if (data) {
-			if (psock_has_data(psock))
+			if (!sk_psock_queue_empty(psock))
 				goto msg_bytes_ready;
 
 			release_sock(sk);
-- 
2.54.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog
  2026-09-29  8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl
  2026-09-29  8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl
@ 2026-09-29  8:34 ` chenyuan_fl
  2026-09-29  9:26   ` bot+bpf-ci
  1 sibling, 1 reply; 5+ messages in thread
From: chenyuan_fl @ 2026-09-29  8:34 UTC (permalink / raw)
  To: netdev, bpf
  Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang,
	linux-kernel

From: Yuan Chen <chenyuan@kylinos.cn>

A sk_skb verdict program that redirects every skb back to the socket
itself keeps psock->ingress_skb populated: the backlog work re-sends
the skb out of that same socket, so it keeps coming back through the
receive path, and ingress_msg never receives anything.

Redirect targets must be in TCP_ESTABLISHED for non-TCP sockets
(sock_map_redirect_allowed()), so the test connects the UDP socket to
its own address; without that bpf_sk_redirect_map() turns every
redirect into SK_DROP and the backlog never fills up.

Reading on such a socket must fall back to the plain UDP receive path,
honor SO_RCVTIMEO and return -EAGAIN instead of spinning.  A re-sent
copy may still win the race against the verdict and be delivered; the
test accepts either outcome.

Run the scenario in a fork()ed child supervised by the parent, since
the spinning reader holds the socket lock and even SIGKILL cannot
reclaim it on an unfixed kernel.

Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
---
 .../bpf/prog_tests/sockmap_udp_backlog.c      | 141 ++++++++++++++++++
 .../bpf/progs/test_sockmap_udp_backlog.c      |  21 +++
 2 files changed, 162 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
 create mode 100644 tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c

diff --git a/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
new file mode 100644
index 000000000000..bbfe2623bd9e
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
@@ -0,0 +1,141 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 KylinSoft */
+
+#include <sys/types.h>
+#include <sys/socket.h>
+#include <sys/wait.h>
+#include <arpa/inet.h>
+#include <errno.h>
+#include <string.h>
+#include <time.h>
+#include <unistd.h>
+
+#include "test_progs.h"
+#include "test_sockmap_udp_backlog.skel.h"
+
+#define RCV_TIMEOUT_MS	1000
+#define HANG_LIMIT_MS	5000
+
+static int run_child(void)
+{
+	struct test_sockmap_udp_backlog *skel;
+	struct timeval tv = { .tv_sec = RCV_TIMEOUT_MS / 1000 };
+	struct sockaddr_in addr = {};
+	struct timespec t0, t1;
+	socklen_t addrlen = sizeof(addr);
+	int zero = 0, sfd, ret, err, exit_code = 1;
+	double elapsed_ms;
+	char byte = 0;
+
+	skel = test_sockmap_udp_backlog__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+		return 1;
+
+	sfd = socket(AF_INET, SOCK_DGRAM, 0);
+	if (!ASSERT_GE(sfd, 0, "socket"))
+		goto out;
+
+	addr.sin_family = AF_INET;
+	addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
+	addr.sin_port = 0;
+	if (!ASSERT_OK(bind(sfd, (struct sockaddr *)&addr, sizeof(addr)), "bind"))
+		goto close;
+	addrlen = sizeof(addr);
+	if (!ASSERT_OK(getsockname(sfd, (struct sockaddr *)&addr, &addrlen),
+		       "getsockname"))
+		goto close;
+
+	/* Non-TCP redirect targets need TCP_ESTABLISHED: connect to self. */
+	if (!ASSERT_OK(connect(sfd, (struct sockaddr *)&addr, sizeof(addr)),
+		       "connect"))
+		goto close;
+
+	err = bpf_prog_attach(bpf_program__fd(skel->progs.redir_to_self),
+			      bpf_map__fd(skel->maps.sock_map),
+			      BPF_SK_SKB_VERDICT, 0);
+	if (!ASSERT_OK(err, "prog_attach"))
+		goto close;
+
+	err = bpf_map_update_elem(bpf_map__fd(skel->maps.sock_map),
+				  &zero, &sfd, BPF_ANY);
+	if (!ASSERT_OK(err, "map_update"))
+		goto close;
+
+	if (!ASSERT_EQ(send(sfd, &byte, 1, 0), 1, "send"))
+		goto close;
+
+	/* Let the backlog pick the skb up. */
+	usleep(100 * 1000);
+
+	err = setsockopt(sfd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
+	if (!ASSERT_OK(err, "set_rcvtimeo"))
+		goto close;
+
+	/* A re-sent copy may be read back; the reader must not spin. */
+	clock_gettime(CLOCK_MONOTONIC, &t0);
+	errno = 0;
+	ret = recv(sfd, &byte, 1, 0);
+	clock_gettime(CLOCK_MONOTONIC, &t1);
+	elapsed_ms = (t1.tv_sec - t0.tv_sec) * 1000.0 +
+		     (t1.tv_nsec - t0.tv_nsec) / 1000000.0;
+
+	if (ret != 1) {
+		if (!ASSERT_EQ(ret, -1, "recv"))
+			goto close;
+		if (!ASSERT_EQ(errno, EAGAIN, "recv_errno"))
+			goto close;
+		if (!ASSERT_GE(elapsed_ms, RCV_TIMEOUT_MS * 0.9, "recv_blocked"))
+			goto close;
+		if (!ASSERT_LT(elapsed_ms, HANG_LIMIT_MS, "recv_timely"))
+			goto close;
+	}
+
+	exit_code = 0;
+close:
+	close(sfd);
+out:
+	test_sockmap_udp_backlog__destroy(skel);
+	return exit_code;
+}
+
+void serial_test_sockmap_udp_backlog(void)
+{
+	pid_t pid;
+	int status = 0;
+	int i;
+
+	pid = fork();
+	if (!ASSERT_GE(pid, 0, "fork"))
+		return;
+
+	if (pid == 0)
+		_exit(run_child());
+
+	/* The child may survive SIGKILL: only a bounded wait is safe. */
+	for (i = 0; i < HANG_LIMIT_MS / 100; i++) {
+		if (waitpid(pid, &status, WNOHANG) == pid)
+			break;
+		usleep(100 * 1000);
+	}
+
+	if (i == HANG_LIMIT_MS / 100) {
+		kill(pid, SIGKILL);
+		for (i = 0; i < 10; i++) {
+			if (waitpid(pid, &status, WNOHANG) == pid)
+				break;
+			usleep(100 * 1000);
+		}
+		fprintf(stderr,
+			"udp_bpf_recvmsg() spins on backlog-only ingress (timeout %dms)\n",
+			HANG_LIMIT_MS);
+		test__fail();
+		return;
+	}
+
+	if (WIFEXITED(status)) {
+		ASSERT_EQ(WEXITSTATUS(status), 0, "child_exit_code");
+	} else {
+		fprintf(stderr, "child terminated abnormally (status=%d)\n", status);
+		test__fail();
+	}
+}
diff --git a/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c b/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c
new file mode 100644
index 000000000000..3459a66da3ba
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c
@@ -0,0 +1,21 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 KylinSoft */
+
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+
+struct {
+	__uint(type, BPF_MAP_TYPE_SOCKMAP);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, __u64);
+} sock_map SEC(".maps");
+
+/* Self-redirect: the backlog work keeps re-sending the skb. */
+SEC("sk_skb/verdict")
+int redir_to_self(struct __sk_buff *skb)
+{
+	return bpf_sk_redirect_map(skb, &sock_map, 0, 0);
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.54.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog
  2026-09-29  8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl
@ 2026-09-29  9:26   ` bot+bpf-ci
  0 siblings, 0 replies; 5+ messages in thread
From: bot+bpf-ci @ 2026-09-29  9:26 UTC (permalink / raw)
  To: chenyuan_fl, netdev, bpf
  Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang,
	linux-kernel, ast, andrii, daniel, martin.lau, eddyz87,
	yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 5425 bytes --]

> diff --git a/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
> new file mode 100644
> index 000000000000..bbfe2623bd9e
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c
> @@ -0,0 +1,141 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 KylinSoft */
> +
> +#include <sys/types.h>
> +#include <sys/socket.h>
> +#include <sys/wait.h>
> +#include <arpa/inet.h>
> +#include <errno.h>
> +#include <string.h>
> +#include <time.h>
> +#include <unistd.h>
> +
> +#include "test_progs.h"
> +#include "test_sockmap_udp_backlog.skel.h"
> +
> +#define RCV_TIMEOUT_MS	1000
> +#define HANG_LIMIT_MS	5000
> +
> +static int run_child(void)
> +{
> +	struct test_sockmap_udp_backlog *skel;
> +	struct timeval tv = { .tv_sec = RCV_TIMEOUT_MS / 1000 };
> +	struct sockaddr_in addr = {};
> +	struct timespec t0, t1;
> +	socklen_t addrlen = sizeof(addr);
> +	int zero = 0, sfd, ret, err, exit_code = 1;
> +	double elapsed_ms;
> +	char byte = 0;
> +
> +	skel = test_sockmap_udp_backlog__open_and_load();
> +	if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
> +		return 1;
> +
> +	sfd = socket(AF_INET, SOCK_DGRAM, 0);
> +	if (!ASSERT_GE(sfd, 0, "socket"))
> +		goto out;
> +
> +	addr.sin_family = AF_INET;
> +	addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
> +	addr.sin_port = 0;
> +	if (!ASSERT_OK(bind(sfd, (struct sockaddr *)&addr, sizeof(addr)), "bind"))
> +		goto close;
> +	addrlen = sizeof(addr);
> +	if (!ASSERT_OK(getsockname(sfd, (struct sockaddr *)&addr, &addrlen),
> +		       "getsockname"))
> +		goto close;
> +
> +	/* Non-TCP redirect targets need TCP_ESTABLISHED: connect to self. */
> +	if (!ASSERT_OK(connect(sfd, (struct sockaddr *)&addr, sizeof(addr)),
> +		       "connect"))
> +		goto close;
> +
> +	err = bpf_prog_attach(bpf_program__fd(skel->progs.redir_to_self),
> +			      bpf_map__fd(skel->maps.sock_map),
> +			      BPF_SK_SKB_VERDICT, 0);
> +	if (!ASSERT_OK(err, "prog_attach"))
> +		goto close;
> +
> +	err = bpf_map_update_elem(bpf_map__fd(skel->maps.sock_map),
> +				  &zero, &sfd, BPF_ANY);
> +	if (!ASSERT_OK(err, "map_update"))
> +		goto close;
> +
> +	if (!ASSERT_EQ(send(sfd, &byte, 1, 0), 1, "send"))
> +		goto close;
> +
> +	/* Let the backlog pick the skb up. */
> +	usleep(100 * 1000);
> +
> +	err = setsockopt(sfd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
> +	if (!ASSERT_OK(err, "set_rcvtimeo"))
> +		goto close;
> +
> +	/* A re-sent copy may be read back; the reader must not spin. */
> +	clock_gettime(CLOCK_MONOTONIC, &t0);
> +	errno = 0;
> +	ret = recv(sfd, &byte, 1, 0);
> +	clock_gettime(CLOCK_MONOTONIC, &t1);
> +	elapsed_ms = (t1.tv_sec - t0.tv_sec) * 1000.0 +
> +		     (t1.tv_nsec - t0.tv_nsec) / 1000000.0;
> +
> +	if (ret != 1) {
> +		if (!ASSERT_EQ(ret, -1, "recv"))
> +			goto close;
> +		if (!ASSERT_EQ(errno, EAGAIN, "recv_errno"))
> +			goto close;
> +		if (!ASSERT_GE(elapsed_ms, RCV_TIMEOUT_MS * 0.9, "recv_blocked"))
> +			goto close;
> +		if (!ASSERT_LT(elapsed_ms, HANG_LIMIT_MS, "recv_timely"))
> +			goto close;
> +	}
> +
> +	exit_code = 0;
> +close:
> +	close(sfd);
> +out:
> +	test_sockmap_udp_backlog__destroy(skel);
> +	return exit_code;
> +}
> +
> +void serial_test_sockmap_udp_backlog(void)
> +{

A subsystem pattern flags this as potentially concerning: this adds a new
standalone prog_tests file plus a new BPF object for one narrow internal
code path (udp_bpf_recvmsg() re-arming its loop when only
psock->ingress_skb is populated).

tools/testing/selftests/bpf/prog_tests/sockmap_basic.c already covers
sockmap UDP receive behavior, for example the "sockmap udp multi
channels" subtest and the SOCK_DGRAM FIONREAD subtest, both built on
test_sockmap_pass_prog.

Would it be better to add this as a new case in the existing test instead
of a new file?

There is a plausible reason for a separate file: on an unfixed kernel the
reader spins, holding lock_sock with no signal check, so the scenario has
to run in a fork()ed child with a bounded wait from the parent. It also
should not run as a parallel test, because the unkillable child pins a
CPU. test_sockmap_basic() is a parallel (non-serial_) test, so a separate
serial_ entry point is defensible.

> +	pid_t pid;
> +	int status = 0;
> +	int i;
> +
> +	pid = fork();
> +	if (!ASSERT_GE(pid, 0, "fork"))
> +		return;
> +
> +	if (pid == 0)
> +		_exit(run_child());
> +
> +	/* The child may survive SIGKILL: only a bounded wait is safe. */
> +	for (i = 0; i < HANG_LIMIT_MS / 100; i++) {
> +		if (waitpid(pid, &status, WNOHANG) == pid)
> +			break;
> +		usleep(100 * 1000);
> +	}
> +
> +	if (i == HANG_LIMIT_MS / 100) {
> +		kill(pid, SIGKILL);
> +		for (i = 0; i < 10; i++) {
> +			if (waitpid(pid, &status, WNOHANG) == pid)
> +				break;
> +			usleep(100 * 1000);
> +		}
> +		fprintf(stderr,
> +			"udp_bpf_recvmsg() spins on backlog-only ingress (timeout %dms)\n",
> +			HANG_LIMIT_MS);
> +		test__fail();
> +		return;
> +	}
> +
> +	if (WIFEXITED(status)) {
> +		ASSERT_EQ(WEXITSTATUS(status), 0, "child_exit_code");
> +	} else {
> +		fprintf(stderr, "child terminated abnormally (status=%d)\n", status);
> +		test__fail();
> +	}
> +}

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36545669306

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf-next 1/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress
  2026-09-29  8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl
@ 2026-09-30  9:01   ` Alexei Starovoitov
  0 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30  9:01 UTC (permalink / raw)
  To: chenyuan_fl, netdev, bpf
  Cc: john.fastabend, jakub, jiayuan.chen, daniel, cong.wang, linux-kernel

On Tue, Sep 29, 2026 at 04:34 PM chenyuan_fl@163.com <chenyuan_fl@163.com> wrote:
> @@ -91,7 +91,7 @@ static int udp_bpf_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
>  		timeo = sock_rcvtimeo(sk, flags & MSG_DONTWAIT);
>  		data = udp_msg_wait_data(sk, psock, timeo);
>  		if (data) {
> -			if (psock_has_data(psock))
> +			if (!sk_psock_queue_empty(psock))
>  				goto msg_bytes_ready;
>
>  			release_sock(sk);

sashiko is right. This replaces a spin with a hang.

pw-bot: cr

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-30  9:02 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl
2026-09-29  8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl
2026-09-30  9:01   ` Alexei Starovoitov
2026-09-29  8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl
2026-09-29  9:26   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®