* [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress @ 2026-09-29 8:34 chenyuan_fl 2026-09-29 8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl 2026-09-29 8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl 0 siblings, 2 replies; 5+ messages in thread From: chenyuan_fl @ 2026-09-29 8:34 UTC (permalink / raw) To: netdev, bpf Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang, linux-kernel From: Yuan Chen <chenyuan@kylinos.cn> udp_bpf_recvmsg() re-arms its receive loop on backlog-only skbs that sk_msg_recvmsg() can never consume: whenever the backlog stays populated, the loop re-arms on the same skb forever and recvmsg() spins while holding the socket lock instead of sleeping. Patch 1 aligns the re-arm condition with TCP and unix_bpf_recvmsg(); patch 2 adds a selftest that reproduces the hang. Yuan Chen (2): bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog net/ipv4/udp_bpf.c | 2 +- .../bpf/prog_tests/sockmap_udp_backlog.c | 141 ++++++++++++++++++ .../bpf/progs/test_sockmap_udp_backlog.c | 21 +++ 3 files changed, 163 insertions(+), 1 deletion(-) create mode 100644 tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c create mode 100644 tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c -- 2.54.0 ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH bpf-next 1/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress 2026-09-29 8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl @ 2026-09-29 8:34 ` chenyuan_fl 2026-09-30 9:01 ` Alexei Starovoitov 2026-09-29 8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl 1 sibling, 1 reply; 5+ messages in thread From: chenyuan_fl @ 2026-09-29 8:34 UTC (permalink / raw) To: netdev, bpf Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang, linux-kernel From: Yuan Chen <chenyuan@kylinos.cn> udp_bpf_recvmsg() re-arms its msg_bytes_ready loop whenever psock_has_data() is true, but that predicate also covers an skb parked in psock->ingress_skb, which sk_msg_recvmsg() can never consume: it only walks psock->ingress_msg. When the backlog stays populated, e.g. while sk_psock_handle_skb() keeps returning -EAGAIN, every round gets copied == 0 and udp_msg_wait_data() returns 1 again from that same skb, so the loop never sleeps: recvmsg() spins at 100% CPU while holding the socket lock, ignores SO_RCVTIMEO and never returns to user space. The wide probe is intended for the entry check and the wait condition, but re-arming can only make progress from ingress_msg. TCP and unix_bpf_recvmsg() therefore re-arm on !sk_psock_queue_empty(psock) and otherwise fall back to the plain receive path. Do the same for UDP so the reader sleeps and returns -EAGAIN on timeout as expected. Fixes: 9f2470fbc4cb ("skmsg: Improve udp_bpf_recvmsg() accuracy") Signed-off-by: Yuan Chen <chenyuan@kylinos.cn> --- net/ipv4/udp_bpf.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/net/ipv4/udp_bpf.c b/net/ipv4/udp_bpf.c index ad57c4c9eaab..8aca9fb89334 100644 --- a/net/ipv4/udp_bpf.c +++ b/net/ipv4/udp_bpf.c @@ -91,7 +91,7 @@ static int udp_bpf_recvmsg(struct sock *sk, struct msghdr *msg, size_t len, timeo = sock_rcvtimeo(sk, flags & MSG_DONTWAIT); data = udp_msg_wait_data(sk, psock, timeo); if (data) { - if (psock_has_data(psock)) + if (!sk_psock_queue_empty(psock)) goto msg_bytes_ready; release_sock(sk); -- 2.54.0 ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH bpf-next 1/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress 2026-09-29 8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl @ 2026-09-30 9:01 ` Alexei Starovoitov 0 siblings, 0 replies; 5+ messages in thread From: Alexei Starovoitov @ 2026-09-30 9:01 UTC (permalink / raw) To: chenyuan_fl, netdev, bpf Cc: john.fastabend, jakub, jiayuan.chen, daniel, cong.wang, linux-kernel On Tue, Sep 29, 2026 at 04:34 PM chenyuan_fl@163.com <chenyuan_fl@163.com> wrote: > @@ -91,7 +91,7 @@ static int udp_bpf_recvmsg(struct sock *sk, struct msghdr *msg, size_t len, > timeo = sock_rcvtimeo(sk, flags & MSG_DONTWAIT); > data = udp_msg_wait_data(sk, psock, timeo); > if (data) { > - if (psock_has_data(psock)) > + if (!sk_psock_queue_empty(psock)) > goto msg_bytes_ready; > > release_sock(sk); sashiko is right. This replaces a spin with a hang. pw-bot: cr ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog 2026-09-29 8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl 2026-09-29 8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl @ 2026-09-29 8:34 ` chenyuan_fl 2026-09-29 9:26 ` bot+bpf-ci 1 sibling, 1 reply; 5+ messages in thread From: chenyuan_fl @ 2026-09-29 8:34 UTC (permalink / raw) To: netdev, bpf Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang, linux-kernel From: Yuan Chen <chenyuan@kylinos.cn> A sk_skb verdict program that redirects every skb back to the socket itself keeps psock->ingress_skb populated: the backlog work re-sends the skb out of that same socket, so it keeps coming back through the receive path, and ingress_msg never receives anything. Redirect targets must be in TCP_ESTABLISHED for non-TCP sockets (sock_map_redirect_allowed()), so the test connects the UDP socket to its own address; without that bpf_sk_redirect_map() turns every redirect into SK_DROP and the backlog never fills up. Reading on such a socket must fall back to the plain UDP receive path, honor SO_RCVTIMEO and return -EAGAIN instead of spinning. A re-sent copy may still win the race against the verdict and be delivered; the test accepts either outcome. Run the scenario in a fork()ed child supervised by the parent, since the spinning reader holds the socket lock and even SIGKILL cannot reclaim it on an unfixed kernel. Signed-off-by: Yuan Chen <chenyuan@kylinos.cn> --- .../bpf/prog_tests/sockmap_udp_backlog.c | 141 ++++++++++++++++++ .../bpf/progs/test_sockmap_udp_backlog.c | 21 +++ 2 files changed, 162 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c create mode 100644 tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c diff --git a/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c new file mode 100644 index 000000000000..bbfe2623bd9e --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c @@ -0,0 +1,141 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft */ + +#include <sys/types.h> +#include <sys/socket.h> +#include <sys/wait.h> +#include <arpa/inet.h> +#include <errno.h> +#include <string.h> +#include <time.h> +#include <unistd.h> + +#include "test_progs.h" +#include "test_sockmap_udp_backlog.skel.h" + +#define RCV_TIMEOUT_MS 1000 +#define HANG_LIMIT_MS 5000 + +static int run_child(void) +{ + struct test_sockmap_udp_backlog *skel; + struct timeval tv = { .tv_sec = RCV_TIMEOUT_MS / 1000 }; + struct sockaddr_in addr = {}; + struct timespec t0, t1; + socklen_t addrlen = sizeof(addr); + int zero = 0, sfd, ret, err, exit_code = 1; + double elapsed_ms; + char byte = 0; + + skel = test_sockmap_udp_backlog__open_and_load(); + if (!ASSERT_OK_PTR(skel, "skel_open_and_load")) + return 1; + + sfd = socket(AF_INET, SOCK_DGRAM, 0); + if (!ASSERT_GE(sfd, 0, "socket")) + goto out; + + addr.sin_family = AF_INET; + addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); + addr.sin_port = 0; + if (!ASSERT_OK(bind(sfd, (struct sockaddr *)&addr, sizeof(addr)), "bind")) + goto close; + addrlen = sizeof(addr); + if (!ASSERT_OK(getsockname(sfd, (struct sockaddr *)&addr, &addrlen), + "getsockname")) + goto close; + + /* Non-TCP redirect targets need TCP_ESTABLISHED: connect to self. */ + if (!ASSERT_OK(connect(sfd, (struct sockaddr *)&addr, sizeof(addr)), + "connect")) + goto close; + + err = bpf_prog_attach(bpf_program__fd(skel->progs.redir_to_self), + bpf_map__fd(skel->maps.sock_map), + BPF_SK_SKB_VERDICT, 0); + if (!ASSERT_OK(err, "prog_attach")) + goto close; + + err = bpf_map_update_elem(bpf_map__fd(skel->maps.sock_map), + &zero, &sfd, BPF_ANY); + if (!ASSERT_OK(err, "map_update")) + goto close; + + if (!ASSERT_EQ(send(sfd, &byte, 1, 0), 1, "send")) + goto close; + + /* Let the backlog pick the skb up. */ + usleep(100 * 1000); + + err = setsockopt(sfd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv)); + if (!ASSERT_OK(err, "set_rcvtimeo")) + goto close; + + /* A re-sent copy may be read back; the reader must not spin. */ + clock_gettime(CLOCK_MONOTONIC, &t0); + errno = 0; + ret = recv(sfd, &byte, 1, 0); + clock_gettime(CLOCK_MONOTONIC, &t1); + elapsed_ms = (t1.tv_sec - t0.tv_sec) * 1000.0 + + (t1.tv_nsec - t0.tv_nsec) / 1000000.0; + + if (ret != 1) { + if (!ASSERT_EQ(ret, -1, "recv")) + goto close; + if (!ASSERT_EQ(errno, EAGAIN, "recv_errno")) + goto close; + if (!ASSERT_GE(elapsed_ms, RCV_TIMEOUT_MS * 0.9, "recv_blocked")) + goto close; + if (!ASSERT_LT(elapsed_ms, HANG_LIMIT_MS, "recv_timely")) + goto close; + } + + exit_code = 0; +close: + close(sfd); +out: + test_sockmap_udp_backlog__destroy(skel); + return exit_code; +} + +void serial_test_sockmap_udp_backlog(void) +{ + pid_t pid; + int status = 0; + int i; + + pid = fork(); + if (!ASSERT_GE(pid, 0, "fork")) + return; + + if (pid == 0) + _exit(run_child()); + + /* The child may survive SIGKILL: only a bounded wait is safe. */ + for (i = 0; i < HANG_LIMIT_MS / 100; i++) { + if (waitpid(pid, &status, WNOHANG) == pid) + break; + usleep(100 * 1000); + } + + if (i == HANG_LIMIT_MS / 100) { + kill(pid, SIGKILL); + for (i = 0; i < 10; i++) { + if (waitpid(pid, &status, WNOHANG) == pid) + break; + usleep(100 * 1000); + } + fprintf(stderr, + "udp_bpf_recvmsg() spins on backlog-only ingress (timeout %dms)\n", + HANG_LIMIT_MS); + test__fail(); + return; + } + + if (WIFEXITED(status)) { + ASSERT_EQ(WEXITSTATUS(status), 0, "child_exit_code"); + } else { + fprintf(stderr, "child terminated abnormally (status=%d)\n", status); + test__fail(); + } +} diff --git a/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c b/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c new file mode 100644 index 000000000000..3459a66da3ba --- /dev/null +++ b/tools/testing/selftests/bpf/progs/test_sockmap_udp_backlog.c @@ -0,0 +1,21 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 KylinSoft */ + +#include "vmlinux.h" +#include <bpf/bpf_helpers.h> + +struct { + __uint(type, BPF_MAP_TYPE_SOCKMAP); + __uint(max_entries, 1); + __type(key, __u32); + __type(value, __u64); +} sock_map SEC(".maps"); + +/* Self-redirect: the backlog work keeps re-sending the skb. */ +SEC("sk_skb/verdict") +int redir_to_self(struct __sk_buff *skb) +{ + return bpf_sk_redirect_map(skb, &sock_map, 0, 0); +} + +char _license[] SEC("license") = "GPL"; -- 2.54.0 ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog 2026-09-29 8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl @ 2026-09-29 9:26 ` bot+bpf-ci 0 siblings, 0 replies; 5+ messages in thread From: bot+bpf-ci @ 2026-09-29 9:26 UTC (permalink / raw) To: chenyuan_fl, netdev, bpf Cc: john.fastabend, jakub, jiayuan.chen, daniel, ast, cong.wang, linux-kernel, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai [-- Attachment #1: Type: text/plain, Size: 5425 bytes --] > diff --git a/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c > new file mode 100644 > index 000000000000..bbfe2623bd9e > --- /dev/null > +++ b/tools/testing/selftests/bpf/prog_tests/sockmap_udp_backlog.c > @@ -0,0 +1,141 @@ > +// SPDX-License-Identifier: GPL-2.0 > +/* Copyright (c) 2026 KylinSoft */ > + > +#include <sys/types.h> > +#include <sys/socket.h> > +#include <sys/wait.h> > +#include <arpa/inet.h> > +#include <errno.h> > +#include <string.h> > +#include <time.h> > +#include <unistd.h> > + > +#include "test_progs.h" > +#include "test_sockmap_udp_backlog.skel.h" > + > +#define RCV_TIMEOUT_MS 1000 > +#define HANG_LIMIT_MS 5000 > + > +static int run_child(void) > +{ > + struct test_sockmap_udp_backlog *skel; > + struct timeval tv = { .tv_sec = RCV_TIMEOUT_MS / 1000 }; > + struct sockaddr_in addr = {}; > + struct timespec t0, t1; > + socklen_t addrlen = sizeof(addr); > + int zero = 0, sfd, ret, err, exit_code = 1; > + double elapsed_ms; > + char byte = 0; > + > + skel = test_sockmap_udp_backlog__open_and_load(); > + if (!ASSERT_OK_PTR(skel, "skel_open_and_load")) > + return 1; > + > + sfd = socket(AF_INET, SOCK_DGRAM, 0); > + if (!ASSERT_GE(sfd, 0, "socket")) > + goto out; > + > + addr.sin_family = AF_INET; > + addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK); > + addr.sin_port = 0; > + if (!ASSERT_OK(bind(sfd, (struct sockaddr *)&addr, sizeof(addr)), "bind")) > + goto close; > + addrlen = sizeof(addr); > + if (!ASSERT_OK(getsockname(sfd, (struct sockaddr *)&addr, &addrlen), > + "getsockname")) > + goto close; > + > + /* Non-TCP redirect targets need TCP_ESTABLISHED: connect to self. */ > + if (!ASSERT_OK(connect(sfd, (struct sockaddr *)&addr, sizeof(addr)), > + "connect")) > + goto close; > + > + err = bpf_prog_attach(bpf_program__fd(skel->progs.redir_to_self), > + bpf_map__fd(skel->maps.sock_map), > + BPF_SK_SKB_VERDICT, 0); > + if (!ASSERT_OK(err, "prog_attach")) > + goto close; > + > + err = bpf_map_update_elem(bpf_map__fd(skel->maps.sock_map), > + &zero, &sfd, BPF_ANY); > + if (!ASSERT_OK(err, "map_update")) > + goto close; > + > + if (!ASSERT_EQ(send(sfd, &byte, 1, 0), 1, "send")) > + goto close; > + > + /* Let the backlog pick the skb up. */ > + usleep(100 * 1000); > + > + err = setsockopt(sfd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv)); > + if (!ASSERT_OK(err, "set_rcvtimeo")) > + goto close; > + > + /* A re-sent copy may be read back; the reader must not spin. */ > + clock_gettime(CLOCK_MONOTONIC, &t0); > + errno = 0; > + ret = recv(sfd, &byte, 1, 0); > + clock_gettime(CLOCK_MONOTONIC, &t1); > + elapsed_ms = (t1.tv_sec - t0.tv_sec) * 1000.0 + > + (t1.tv_nsec - t0.tv_nsec) / 1000000.0; > + > + if (ret != 1) { > + if (!ASSERT_EQ(ret, -1, "recv")) > + goto close; > + if (!ASSERT_EQ(errno, EAGAIN, "recv_errno")) > + goto close; > + if (!ASSERT_GE(elapsed_ms, RCV_TIMEOUT_MS * 0.9, "recv_blocked")) > + goto close; > + if (!ASSERT_LT(elapsed_ms, HANG_LIMIT_MS, "recv_timely")) > + goto close; > + } > + > + exit_code = 0; > +close: > + close(sfd); > +out: > + test_sockmap_udp_backlog__destroy(skel); > + return exit_code; > +} > + > +void serial_test_sockmap_udp_backlog(void) > +{ A subsystem pattern flags this as potentially concerning: this adds a new standalone prog_tests file plus a new BPF object for one narrow internal code path (udp_bpf_recvmsg() re-arming its loop when only psock->ingress_skb is populated). tools/testing/selftests/bpf/prog_tests/sockmap_basic.c already covers sockmap UDP receive behavior, for example the "sockmap udp multi channels" subtest and the SOCK_DGRAM FIONREAD subtest, both built on test_sockmap_pass_prog. Would it be better to add this as a new case in the existing test instead of a new file? There is a plausible reason for a separate file: on an unfixed kernel the reader spins, holding lock_sock with no signal check, so the scenario has to run in a fork()ed child with a bounded wait from the parent. It also should not run as a parallel test, because the unkillable child pins a CPU. test_sockmap_basic() is a parallel (non-serial_) test, so a separate serial_ entry point is defensible. > + pid_t pid; > + int status = 0; > + int i; > + > + pid = fork(); > + if (!ASSERT_GE(pid, 0, "fork")) > + return; > + > + if (pid == 0) > + _exit(run_child()); > + > + /* The child may survive SIGKILL: only a bounded wait is safe. */ > + for (i = 0; i < HANG_LIMIT_MS / 100; i++) { > + if (waitpid(pid, &status, WNOHANG) == pid) > + break; > + usleep(100 * 1000); > + } > + > + if (i == HANG_LIMIT_MS / 100) { > + kill(pid, SIGKILL); > + for (i = 0; i < 10; i++) { > + if (waitpid(pid, &status, WNOHANG) == pid) > + break; > + usleep(100 * 1000); > + } > + fprintf(stderr, > + "udp_bpf_recvmsg() spins on backlog-only ingress (timeout %dms)\n", > + HANG_LIMIT_MS); > + test__fail(); > + return; > + } > + > + if (WIFEXITED(status)) { > + ASSERT_EQ(WEXITSTATUS(status), 0, "child_exit_code"); > + } else { > + fprintf(stderr, "child terminated abnormally (status=%d)\n", status); > + test__fail(); > + } > +} [ ... ] --- AI reviewed your patch. Please fix the bug or email reply why it's not a bug. See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36545669306 ^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-30 9:02 UTC | newest] Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-29 8:34 [PATCH bpf-next 0/2] bpf, sockmap: Fix udp_bpf_recvmsg() spinning on backlog-only ingress chenyuan_fl 2026-09-29 8:34 ` [PATCH bpf-next 1/2] " chenyuan_fl 2026-09-30 9:01 ` Alexei Starovoitov 2026-09-29 8:34 ` [PATCH bpf-next 2/2] selftests/bpf: Add a test for udp_bpf_recvmsg() with a stuck backlog chenyuan_fl 2026-09-29 9:26 ` bot+bpf-ci
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®