* [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP
@ 2026-02-16 13:48 Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 1/2] clone: " Christian Brauner
` (2 more replies)
0 siblings, 3 replies; 5+ messages in thread
From: Christian Brauner @ 2026-02-16 13:48 UTC (permalink / raw)
To: Oleg Nesterov
Cc: Jann Horn, Linus Torvalds, Ingo Molnar, Peter Zijlstra,
linux-kernel, linux-fsdevel, Christian Brauner
Add a new clone3() flag CLONE_AUTOREAP that makes a child process
auto-reap on exit without ever becoming a zombie. This is a per-process
property in contrast to the existing auto-reap mechanism via
SA_NOCLDWAIT or SIG_IGN for SIGCHLD which applies to all children of a
given parent.
With pidfds this is very useful as the parent can monitor the pidfd via
poll and retrieve the exit status from the pidfd.
Currently the only way to automatically reap children is to set
SA_NOCLDWAIT or SIG_IGN on SIGCHLD. This is a parent-scoped property
affecting all children which makes it unsuitable for libraries or
applications that need selective auto-reaping of specific children while
still being able to wait() on others.
CLONE_AUTOREAP stores an autoreap flag in the child's signal_struct.
When the child exits do_notify_parent() checks this flag and returns
autoreap=true causing exit_notify() to transition the task directly to
EXIT_DEAD. Since the flag lives on the child it survives reparenting: if
the original parent exits and the child is reparented to a subreaper or
init the child still auto-reaps when it eventually exits. This is
cleaner then forcing the subreaper to get SIGCHLD and then reaping it.
If the parent doesn't care the subreaper won't care. If there's a
subreaper that would care it would be easy enough to add a prctl() that
either just turns back on SIGCHLD and turns of auto-reaping or a prctl()
that just notifies the subreaper whenever a child is reparented to it.
CLONE_AUTOREAP requires CLONE_PIDFD because the process will never be
visible to wait(). The parent must use the pidfd to monitor exit via
poll() and retrieve exit status via PIDFD_GET_INFO. No exit signal is
delivered so exit_signal must be zero.
The flag is not inherited by the autoreap process's own children. Each
child that should be autoreaped must be explicitly created with
CLONE_AUTOREAP.
(Later on we can augment this with another addition CLONE_PIDFD_AUTOKILL
which would SIGKILL the child process when the pidfd that was returned
from clone3() is closed. Specifically, when the file referenced by the
fd from clone3() is closed. The wrinkly here is that it would either
have to be reset on privilege gaining exec - like pdeath signal - or we
enforce that autokill only works when no-new-privileges is set.)
Signed-off-by: Christian Brauner <brauner@kernel.org>
---
Christian Brauner (2):
clone: add CLONE_AUTOREAP
selftests/pidfd: add CLONE_AUTOREAP tests
include/linux/sched/signal.h | 1 +
include/uapi/linux/sched.h | 1 +
kernel/fork.c | 16 +-
kernel/ptrace.c | 3 +-
kernel/signal.c | 4 +
tools/testing/selftests/pidfd/.gitignore | 1 +
tools/testing/selftests/pidfd/Makefile | 2 +-
.../testing/selftests/pidfd/pidfd_autoreap_test.c | 475 +++++++++++++++++++++
8 files changed, 500 insertions(+), 3 deletions(-)
---
base-commit: 72c395024dac5e215136cbff793455f065603b06
change-id: 20260214-work-pidfs-autoreap-3ee677e240a8
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH RFC 1/2] clone: add CLONE_AUTOREAP
2026-02-16 13:48 [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Christian Brauner
@ 2026-02-16 13:48 ` Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 2/2] selftests/pidfd: add CLONE_AUTOREAP tests Christian Brauner
2026-02-16 15:25 ` [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Linus Torvalds
2 siblings, 0 replies; 5+ messages in thread
From: Christian Brauner @ 2026-02-16 13:48 UTC (permalink / raw)
To: Oleg Nesterov
Cc: Jann Horn, Linus Torvalds, Ingo Molnar, Peter Zijlstra,
linux-kernel, linux-fsdevel, Christian Brauner
Add a new clone3() flag CLONE_AUTOREAP that makes a child process
auto-reap on exit without ever becoming a zombie. This is a per-process
property in contrast to the existing auto-reap mechanism via
SA_NOCLDWAIT or SIG_IGN for SIGCHLD which applies to all children of a
given parent.
Currently the only way to automatically reap children is to set
SA_NOCLDWAIT or SIG_IGN on SIGCHLD. This is a parent-scoped property
affecting all children which makes it unsuitable for libraries or
applications that need selective auto-reaping of specific children while
still being able to wait() on others.
CLONE_AUTOREAP stores an autoreap flag in the child's signal_struct.
When the child exits do_notify_parent() checks this flag and returns
autoreap=true causing exit_notify() to transition the task directly to
EXIT_DEAD. Since the flag lives on the child it survives reparenting: if
the original parent exits and the child is reparented to a subreaper or
init the child still auto-reaps when it eventually exits.
CLONE_AUTOREAP requires CLONE_PIDFD because the process will never be
visible to wait(). The parent must use the pidfd to monitor exit via
poll() and retrieve exit status via PIDFD_GET_INFO. No exit signal is
delivered so exit_signal must be zero.
The flag is not inherited by the autoreap process's own children. Each
child that should be autoreaped must be explicitly created with
CLONE_AUTOREAP.
Link: https://github.com/uapi-group/kernel-features/issues/45
Signed-off-by: Christian Brauner <brauner@kernel.org>
---
include/linux/sched/signal.h | 1 +
include/uapi/linux/sched.h | 1 +
kernel/fork.c | 16 +++++++++++++++-
kernel/ptrace.c | 3 ++-
kernel/signal.c | 4 ++++
5 files changed, 23 insertions(+), 2 deletions(-)
diff --git a/include/linux/sched/signal.h b/include/linux/sched/signal.h
index 7d6449982822..346ecbad4c2b 100644
--- a/include/linux/sched/signal.h
+++ b/include/linux/sched/signal.h
@@ -132,6 +132,7 @@ struct signal_struct {
*/
unsigned int is_child_subreaper:1;
unsigned int has_child_subreaper:1;
+ unsigned int autoreap:1;
#ifdef CONFIG_POSIX_TIMERS
diff --git a/include/uapi/linux/sched.h b/include/uapi/linux/sched.h
index 359a14cc76a4..e6fc5ae621e2 100644
--- a/include/uapi/linux/sched.h
+++ b/include/uapi/linux/sched.h
@@ -36,6 +36,7 @@
/* Flags for the clone3() syscall. */
#define CLONE_CLEAR_SIGHAND 0x100000000ULL /* Clear any signal handler and reset to SIG_DFL. */
#define CLONE_INTO_CGROUP 0x200000000ULL /* Clone into a specific cgroup given the right permissions. */
+#define CLONE_AUTOREAP 0x400000000ULL /* Auto-reap child on exit, requires CLONE_PIDFD. */
/*
* cloning flags intersect with CSIGNAL so can be used with unshare and clone3
diff --git a/kernel/fork.c b/kernel/fork.c
index 9c5effbdbdc1..a803bdad2805 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2028,6 +2028,15 @@ __latent_entropy struct task_struct *copy_process(
return ERR_PTR(-EINVAL);
}
+ if (clone_flags & CLONE_AUTOREAP) {
+ if (!(clone_flags & CLONE_PIDFD))
+ return ERR_PTR(-EINVAL);
+ if (clone_flags & CLONE_THREAD)
+ return ERR_PTR(-EINVAL);
+ if (args->exit_signal)
+ return ERR_PTR(-EINVAL);
+ }
+
/*
* Force any signals received before this point to be delivered
* before the fork happens. Collect up signals sent to multiple
@@ -2374,6 +2383,8 @@ __latent_entropy struct task_struct *copy_process(
p->parent_exec_id = current->parent_exec_id;
if (clone_flags & CLONE_THREAD)
p->exit_signal = -1;
+ else if (clone_flags & CLONE_AUTOREAP)
+ p->exit_signal = 0;
else
p->exit_signal = current->group_leader->exit_signal;
} else {
@@ -2435,6 +2446,8 @@ __latent_entropy struct task_struct *copy_process(
*/
p->signal->has_child_subreaper = p->real_parent->signal->has_child_subreaper ||
p->real_parent->signal->is_child_subreaper;
+ if (clone_flags & CLONE_AUTOREAP)
+ p->signal->autoreap = 1;
list_add_tail(&p->sibling, &p->real_parent->children);
list_add_tail_rcu(&p->tasks, &init_task.tasks);
attach_pid(p, PIDTYPE_TGID);
@@ -2897,7 +2910,8 @@ static bool clone3_args_valid(struct kernel_clone_args *kargs)
{
/* Verify that no unknown flags are passed along. */
if (kargs->flags &
- ~(CLONE_LEGACY_FLAGS | CLONE_CLEAR_SIGHAND | CLONE_INTO_CGROUP))
+ ~(CLONE_LEGACY_FLAGS | CLONE_CLEAR_SIGHAND | CLONE_INTO_CGROUP |
+ CLONE_AUTOREAP))
return false;
/*
diff --git a/kernel/ptrace.c b/kernel/ptrace.c
index 392ec2f75f01..68c17daef8d4 100644
--- a/kernel/ptrace.c
+++ b/kernel/ptrace.c
@@ -549,7 +549,8 @@ static bool __ptrace_detach(struct task_struct *tracer, struct task_struct *p)
if (!dead && thread_group_empty(p)) {
if (!same_thread_group(p->real_parent, tracer))
dead = do_notify_parent(p, p->exit_signal);
- else if (ignoring_children(tracer->sighand)) {
+ else if (ignoring_children(tracer->sighand) ||
+ p->signal->autoreap) {
__wake_up_parent(p, tracer);
dead = true;
}
diff --git a/kernel/signal.c b/kernel/signal.c
index e42b8bd6922f..2fb206c84c07 100644
--- a/kernel/signal.c
+++ b/kernel/signal.c
@@ -2251,6 +2251,10 @@ bool do_notify_parent(struct task_struct *tsk, int sig)
if (psig->action[SIGCHLD-1].sa.sa_handler == SIG_IGN)
sig = 0;
}
+ if (!tsk->ptrace && tsk->signal->autoreap) {
+ autoreap = true;
+ sig = 0;
+ }
/*
* Send with __send_signal as si_pid and si_uid are in the
* parent's namespaces.
--
2.47.3
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH RFC 2/2] selftests/pidfd: add CLONE_AUTOREAP tests
2026-02-16 13:48 [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 1/2] clone: " Christian Brauner
@ 2026-02-16 13:48 ` Christian Brauner
2026-02-16 15:25 ` [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Linus Torvalds
2 siblings, 0 replies; 5+ messages in thread
From: Christian Brauner @ 2026-02-16 13:48 UTC (permalink / raw)
To: Oleg Nesterov
Cc: Jann Horn, Linus Torvalds, Ingo Molnar, Peter Zijlstra,
linux-kernel, linux-fsdevel, Christian Brauner
Add tests for the new CLONE_AUTOREAP clone3() flag:
- autoreap_requires_pidfd: CLONE_AUTOREAP without CLONE_PIDFD fails
- autoreap_rejects_exit_signal: CLONE_AUTOREAP with non-zero
exit_signal fails
- autoreap_rejects_thread: CLONE_AUTOREAP with CLONE_THREAD fails
- autoreap_basic: child exits, pidfd poll works, PIDFD_GET_INFO returns
correct exit code, waitpid() returns -ECHILD
- autoreap_signaled: child killed by signal, exit info correct via pidfd
- autoreap_reparent: autoreap grandchild reparented to subreaper still
auto-reaps
- autoreap_multithreaded: autoreap process with sub-threads auto-reaps
after last thread exits
- autoreap_no_inherit: grandchild forked without CLONE_AUTOREAP becomes
a regular zombie
Signed-off-by: Christian Brauner <brauner@kernel.org>
---
tools/testing/selftests/pidfd/.gitignore | 1 +
tools/testing/selftests/pidfd/Makefile | 2 +-
.../testing/selftests/pidfd/pidfd_autoreap_test.c | 475 +++++++++++++++++++++
3 files changed, 477 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/pidfd/.gitignore b/tools/testing/selftests/pidfd/.gitignore
index 144e7ff65d6a..4cd8ec7fd349 100644
--- a/tools/testing/selftests/pidfd/.gitignore
+++ b/tools/testing/selftests/pidfd/.gitignore
@@ -12,3 +12,4 @@ pidfd_info_test
pidfd_exec_helper
pidfd_xattr_test
pidfd_setattr_test
+pidfd_autoreap_test
diff --git a/tools/testing/selftests/pidfd/Makefile b/tools/testing/selftests/pidfd/Makefile
index 764a8f9ecefa..4211f91e9af8 100644
--- a/tools/testing/selftests/pidfd/Makefile
+++ b/tools/testing/selftests/pidfd/Makefile
@@ -4,7 +4,7 @@ CFLAGS += -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES) -pthread -Wall
TEST_GEN_PROGS := pidfd_test pidfd_fdinfo_test pidfd_open_test \
pidfd_poll_test pidfd_wait pidfd_getfd_test pidfd_setns_test \
pidfd_file_handle_test pidfd_bind_mount pidfd_info_test \
- pidfd_xattr_test pidfd_setattr_test
+ pidfd_xattr_test pidfd_setattr_test pidfd_autoreap_test
TEST_GEN_PROGS_EXTENDED := pidfd_exec_helper
diff --git a/tools/testing/selftests/pidfd/pidfd_autoreap_test.c b/tools/testing/selftests/pidfd/pidfd_autoreap_test.c
new file mode 100644
index 000000000000..b904abf79f33
--- /dev/null
+++ b/tools/testing/selftests/pidfd/pidfd_autoreap_test.c
@@ -0,0 +1,475 @@
+// SPDX-License-Identifier: GPL-2.0
+// Copyright (c) 2026 Christian Brauner <brauner@kernel.org>
+
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/types.h>
+#include <poll.h>
+#include <pthread.h>
+#include <sched.h>
+#include <signal.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <syscall.h>
+#include <sys/ioctl.h>
+#include <sys/prctl.h>
+#include <sys/socket.h>
+#include <sys/types.h>
+#include <sys/wait.h>
+#include <unistd.h>
+
+#include "pidfd.h"
+#include "kselftest_harness.h"
+
+#ifndef CLONE_AUTOREAP
+#define CLONE_AUTOREAP 0x400000000ULL
+#endif
+
+static pid_t create_autoreap_child(int *pidfd)
+{
+ struct __clone_args args = {
+ .flags = CLONE_PIDFD | CLONE_AUTOREAP,
+ .exit_signal = 0,
+ .pidfd = ptr_to_u64(pidfd),
+ };
+
+ return sys_clone3(&args, sizeof(args));
+}
+
+/*
+ * Test that CLONE_AUTOREAP without CLONE_PIDFD fails.
+ */
+TEST(autoreap_requires_pidfd)
+{
+ struct __clone_args args = {
+ .flags = CLONE_AUTOREAP,
+ .exit_signal = SIGCHLD,
+ };
+ pid_t pid;
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_EQ(pid, -1);
+ ASSERT_EQ(errno, EINVAL);
+}
+
+/*
+ * Test that CLONE_AUTOREAP with a non-zero exit_signal fails.
+ */
+TEST(autoreap_rejects_exit_signal)
+{
+ struct __clone_args args = {
+ .flags = CLONE_PIDFD | CLONE_AUTOREAP,
+ .exit_signal = SIGCHLD,
+ };
+ int pidfd = -1;
+ pid_t pid;
+
+ args.pidfd = ptr_to_u64(&pidfd);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_EQ(pid, -1);
+ ASSERT_EQ(errno, EINVAL);
+}
+
+/*
+ * Test that CLONE_AUTOREAP with CLONE_THREAD fails.
+ */
+TEST(autoreap_rejects_thread)
+{
+ struct __clone_args args = {
+ .flags = CLONE_PIDFD | CLONE_AUTOREAP |
+ CLONE_THREAD | CLONE_SIGHAND |
+ CLONE_VM,
+ .exit_signal = 0,
+ };
+ int pidfd = -1;
+ pid_t pid;
+
+ args.pidfd = ptr_to_u64(&pidfd);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_EQ(pid, -1);
+ ASSERT_EQ(errno, EINVAL);
+}
+
+/*
+ * Basic test: create an autoreap child, let it exit, verify:
+ * - pidfd becomes readable (poll returns POLLIN)
+ * - PIDFD_GET_INFO returns the correct exit code
+ * - waitpid() returns -1/ECHILD (no zombie)
+ */
+TEST(autoreap_basic)
+{
+ struct pidfd_info info = { .mask = PIDFD_INFO_EXIT };
+ int pidfd = -1, ret;
+ struct pollfd pfd;
+ pid_t pid;
+
+ pid = create_autoreap_child(&pidfd);
+ if (pid < 0 && errno == EINVAL)
+ SKIP(return, "CLONE_AUTOREAP not supported");
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0)
+ _exit(42);
+
+ ASSERT_GE(pidfd, 0);
+
+ /* Wait for the child to exit via pidfd poll. */
+ pfd.fd = pidfd;
+ pfd.events = POLLIN;
+ ret = poll(&pfd, 1, 5000);
+ ASSERT_EQ(ret, 1);
+ ASSERT_TRUE(pfd.revents & POLLIN);
+
+ /* Verify exit info via PIDFD_GET_INFO. */
+ ret = ioctl(pidfd, PIDFD_GET_INFO, &info);
+ ASSERT_EQ(ret, 0);
+ ASSERT_TRUE(info.mask & PIDFD_INFO_EXIT);
+ /*
+ * exit_code is in waitpid format: for _exit(42),
+ * WIFEXITED is true and WEXITSTATUS is 42.
+ */
+ ASSERT_TRUE(WIFEXITED(info.exit_code));
+ ASSERT_EQ(WEXITSTATUS(info.exit_code), 42);
+
+ /* Verify no zombie: waitpid should fail with ECHILD. */
+ ret = waitpid(pid, NULL, WNOHANG);
+ ASSERT_EQ(ret, -1);
+ ASSERT_EQ(errno, ECHILD);
+
+ close(pidfd);
+}
+
+/*
+ * Test that an autoreap child killed by a signal reports
+ * the correct exit info.
+ */
+TEST(autoreap_signaled)
+{
+ struct pidfd_info info = { .mask = PIDFD_INFO_EXIT };
+ int pidfd = -1, ret;
+ struct pollfd pfd;
+ pid_t pid;
+
+ pid = create_autoreap_child(&pidfd);
+ if (pid < 0 && errno == EINVAL)
+ SKIP(return, "CLONE_AUTOREAP not supported");
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ pause();
+ _exit(1);
+ }
+
+ ASSERT_GE(pidfd, 0);
+
+ /* Kill the child. */
+ ret = sys_pidfd_send_signal(pidfd, SIGKILL, NULL, 0);
+ ASSERT_EQ(ret, 0);
+
+ /* Wait for exit via pidfd. */
+ pfd.fd = pidfd;
+ pfd.events = POLLIN;
+ ret = poll(&pfd, 1, 5000);
+ ASSERT_EQ(ret, 1);
+ ASSERT_TRUE(pfd.revents & POLLIN);
+
+ /* Verify signal info. */
+ ret = ioctl(pidfd, PIDFD_GET_INFO, &info);
+ ASSERT_EQ(ret, 0);
+ ASSERT_TRUE(info.mask & PIDFD_INFO_EXIT);
+ ASSERT_TRUE(WIFSIGNALED(info.exit_code));
+ ASSERT_EQ(WTERMSIG(info.exit_code), SIGKILL);
+
+ /* No zombie. */
+ ret = waitpid(pid, NULL, WNOHANG);
+ ASSERT_EQ(ret, -1);
+ ASSERT_EQ(errno, ECHILD);
+
+ close(pidfd);
+}
+
+/*
+ * Test autoreap survives reparenting: middle process creates an
+ * autoreap grandchild, then exits. The grandchild gets reparented
+ * to us (the grandparent, which is a subreaper). When the grandchild
+ * exits, it should still be autoreaped - no zombie under us.
+ */
+TEST(autoreap_reparent)
+{
+ int ipc_sockets[2], ret;
+ int pidfd = -1;
+ struct pollfd pfd;
+ pid_t mid_pid, grandchild_pid;
+ char buf[32] = {};
+
+ /* Make ourselves a subreaper so reparented children come to us. */
+ ret = prctl(PR_SET_CHILD_SUBREAPER, 1);
+ ASSERT_EQ(ret, 0);
+
+ ret = socketpair(AF_LOCAL, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets);
+ ASSERT_EQ(ret, 0);
+
+ mid_pid = fork();
+ ASSERT_GE(mid_pid, 0);
+
+ if (mid_pid == 0) {
+ /* Middle child: create an autoreap grandchild. */
+ int gc_pidfd = -1;
+
+ close(ipc_sockets[0]);
+
+ grandchild_pid = create_autoreap_child(&gc_pidfd);
+ if (grandchild_pid < 0) {
+ write_nointr(ipc_sockets[1], "E", 1);
+ close(ipc_sockets[1]);
+ _exit(1);
+ }
+
+ if (grandchild_pid == 0) {
+ /* Grandchild: wait for signal to exit. */
+ close(ipc_sockets[1]);
+ if (gc_pidfd >= 0)
+ close(gc_pidfd);
+ pause();
+ _exit(0);
+ }
+
+ /* Send grandchild PID to grandparent. */
+ snprintf(buf, sizeof(buf), "%d", grandchild_pid);
+ write_nointr(ipc_sockets[1], buf, strlen(buf));
+ close(ipc_sockets[1]);
+ if (gc_pidfd >= 0)
+ close(gc_pidfd);
+
+ /* Middle child exits, grandchild gets reparented. */
+ _exit(0);
+ }
+
+ close(ipc_sockets[1]);
+
+ /* Read grandchild's PID. */
+ ret = read_nointr(ipc_sockets[0], buf, sizeof(buf) - 1);
+ close(ipc_sockets[0]);
+ ASSERT_GT(ret, 0);
+
+ if (buf[0] == 'E') {
+ waitpid(mid_pid, NULL, 0);
+ prctl(PR_SET_CHILD_SUBREAPER, 0);
+ SKIP(return, "CLONE_AUTOREAP not supported");
+ }
+
+ grandchild_pid = atoi(buf);
+ ASSERT_GT(grandchild_pid, 0);
+
+ /* Wait for the middle child to exit. */
+ ret = waitpid(mid_pid, NULL, 0);
+ ASSERT_EQ(ret, mid_pid);
+
+ /*
+ * Now the grandchild is reparented to us (subreaper).
+ * Open a pidfd for the grandchild and kill it.
+ */
+ pidfd = sys_pidfd_open(grandchild_pid, 0);
+ ASSERT_GE(pidfd, 0);
+
+ ret = sys_pidfd_send_signal(pidfd, SIGKILL, NULL, 0);
+ ASSERT_EQ(ret, 0);
+
+ /* Wait for it to exit via pidfd poll. */
+ pfd.fd = pidfd;
+ pfd.events = POLLIN;
+ ret = poll(&pfd, 1, 5000);
+ ASSERT_EQ(ret, 1);
+ ASSERT_TRUE(pfd.revents & POLLIN);
+
+ /*
+ * The grandchild should have been autoreaped even though
+ * we (the new parent) haven't set SA_NOCLDWAIT.
+ * waitpid should return -1/ECHILD.
+ */
+ ret = waitpid(grandchild_pid, NULL, WNOHANG);
+ EXPECT_EQ(ret, -1);
+ EXPECT_EQ(errno, ECHILD);
+
+ close(pidfd);
+
+ /* Clean up subreaper status. */
+ prctl(PR_SET_CHILD_SUBREAPER, 0);
+}
+
+static int thread_sock_fd;
+
+static void *thread_func(void *arg)
+{
+ /* Signal parent we're running. */
+ write_nointr(thread_sock_fd, "1", 1);
+
+ /* Give main thread time to call _exit() first. */
+ usleep(200000);
+
+ return NULL;
+}
+
+/*
+ * Test that an autoreap child with multiple threads is properly
+ * autoreaped only after all threads have exited.
+ */
+TEST(autoreap_multithreaded)
+{
+ struct pidfd_info info = { .mask = PIDFD_INFO_EXIT };
+ int ipc_sockets[2], ret;
+ int pidfd = -1;
+ struct pollfd pfd;
+ pid_t pid;
+ char c;
+
+ ret = socketpair(AF_LOCAL, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets);
+ ASSERT_EQ(ret, 0);
+
+ pid = create_autoreap_child(&pidfd);
+ if (pid < 0 && errno == EINVAL) {
+ close(ipc_sockets[0]);
+ close(ipc_sockets[1]);
+ SKIP(return, "CLONE_AUTOREAP not supported");
+ }
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ pthread_t thread;
+
+ close(ipc_sockets[0]);
+
+ /*
+ * Create a sub-thread that outlives the main thread.
+ * The thread signals readiness, then sleeps.
+ * The main thread waits briefly, then calls _exit().
+ */
+ thread_sock_fd = ipc_sockets[1];
+ pthread_create(&thread, NULL, thread_func, NULL);
+ pthread_detach(thread);
+
+ /* Wait for thread to be running. */
+ usleep(100000);
+
+ /* Main thread exits; sub-thread is still alive. */
+ _exit(99);
+ }
+
+ close(ipc_sockets[1]);
+
+ /* Wait for the sub-thread to signal readiness. */
+ ret = read_nointr(ipc_sockets[0], &c, 1);
+ close(ipc_sockets[0]);
+ ASSERT_EQ(ret, 1);
+
+ /* Wait for the process to fully exit via pidfd poll. */
+ pfd.fd = pidfd;
+ pfd.events = POLLIN;
+ ret = poll(&pfd, 1, 5000);
+ ASSERT_EQ(ret, 1);
+ ASSERT_TRUE(pfd.revents & POLLIN);
+
+ /* Verify exit info. */
+ ret = ioctl(pidfd, PIDFD_GET_INFO, &info);
+ ASSERT_EQ(ret, 0);
+ ASSERT_TRUE(info.mask & PIDFD_INFO_EXIT);
+ ASSERT_TRUE(WIFEXITED(info.exit_code));
+ ASSERT_EQ(WEXITSTATUS(info.exit_code), 99);
+
+ /* No zombie. */
+ ret = waitpid(pid, NULL, WNOHANG);
+ ASSERT_EQ(ret, -1);
+ ASSERT_EQ(errno, ECHILD);
+
+ close(pidfd);
+}
+
+/*
+ * Test that autoreap is NOT inherited by grandchildren.
+ */
+TEST(autoreap_no_inherit)
+{
+ int ipc_sockets[2], ret;
+ int pidfd = -1;
+ pid_t pid;
+ char buf[2] = {};
+ struct pollfd pfd;
+
+ ret = socketpair(AF_LOCAL, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets);
+ ASSERT_EQ(ret, 0);
+
+ pid = create_autoreap_child(&pidfd);
+ if (pid < 0 && errno == EINVAL) {
+ close(ipc_sockets[0]);
+ close(ipc_sockets[1]);
+ SKIP(return, "CLONE_AUTOREAP not supported");
+ }
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ pid_t gc;
+ int status;
+
+ close(ipc_sockets[0]);
+
+ /* Autoreap child forks a grandchild (without autoreap). */
+ gc = fork();
+ if (gc < 0) {
+ write_nointr(ipc_sockets[1], "E", 1);
+ _exit(1);
+ }
+ if (gc == 0) {
+ /* Grandchild: exit immediately. */
+ close(ipc_sockets[1]);
+ _exit(77);
+ }
+
+ /*
+ * The grandchild should become a regular zombie
+ * since it was NOT created with CLONE_AUTOREAP.
+ * Wait for it to verify.
+ */
+ ret = waitpid(gc, &status, 0);
+ if (ret == gc && WIFEXITED(status) &&
+ WEXITSTATUS(status) == 77) {
+ write_nointr(ipc_sockets[1], "P", 1);
+ } else {
+ write_nointr(ipc_sockets[1], "F", 1);
+ }
+ close(ipc_sockets[1]);
+ _exit(0);
+ }
+
+ close(ipc_sockets[1]);
+
+ ret = read_nointr(ipc_sockets[0], buf, 1);
+ close(ipc_sockets[0]);
+ ASSERT_EQ(ret, 1);
+
+ /*
+ * 'P' means the autoreap child was able to waitpid() its
+ * grandchild (correct - grandchild should be a normal zombie,
+ * not autoreaped).
+ */
+ ASSERT_EQ(buf[0], 'P');
+
+ /* Wait for the autoreap child to exit. */
+ pfd.fd = pidfd;
+ pfd.events = POLLIN;
+ ret = poll(&pfd, 1, 5000);
+ ASSERT_EQ(ret, 1);
+
+ /* Autoreap child itself should be autoreaped. */
+ ret = waitpid(pid, NULL, WNOHANG);
+ ASSERT_EQ(ret, -1);
+ ASSERT_EQ(errno, ECHILD);
+
+ close(pidfd);
+}
+
+TEST_HARNESS_MAIN
--
2.47.3
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP
2026-02-16 13:48 [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 1/2] clone: " Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 2/2] selftests/pidfd: add CLONE_AUTOREAP tests Christian Brauner
@ 2026-02-16 15:25 ` Linus Torvalds
2026-02-17 8:16 ` Christian Brauner
2 siblings, 1 reply; 5+ messages in thread
From: Linus Torvalds @ 2026-02-16 15:25 UTC (permalink / raw)
To: Christian Brauner
Cc: Oleg Nesterov, Jann Horn, Ingo Molnar, Peter Zijlstra,
linux-kernel, linux-fsdevel
On Mon, 16 Feb 2026 at 05:49, Christian Brauner <brauner@kernel.org> wrote:
>
> CLONE_AUTOREAP requires CLONE_PIDFD because the process will never be
> visible to wait().
This seems an unnecessary and counter-productive limitation.
The very *traditional* unix way to do auto-reaping is to fork twice,
and have the "middle" parent just exit.
That makes the final child be re-parented to init, and it is invisible
to wait() - all very much on purpose.
This was (perhaps still is?) very commonly used for starting up
background daemons (together with disassociating from the tty etc).
So I don't mind th enew flag, but I think the restriction is
unnecessary and not logical. Sometimes you simply don't *want*
processes visible to wait - or care about a pidfd.
Linus
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP
2026-02-16 15:25 ` [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Linus Torvalds
@ 2026-02-17 8:16 ` Christian Brauner
0 siblings, 0 replies; 5+ messages in thread
From: Christian Brauner @ 2026-02-17 8:16 UTC (permalink / raw)
To: Linus Torvalds
Cc: Oleg Nesterov, Jann Horn, Ingo Molnar, Peter Zijlstra,
linux-kernel, linux-fsdevel
On Mon, Feb 16, 2026 at 07:25:48AM -0800, Linus Torvalds wrote:
> On Mon, 16 Feb 2026 at 05:49, Christian Brauner <brauner@kernel.org> wrote:
> >
> > CLONE_AUTOREAP requires CLONE_PIDFD because the process will never be
> > visible to wait().
>
> This seems an unnecessary and counter-productive limitation.
>
> The very *traditional* unix way to do auto-reaping is to fork twice,
> and have the "middle" parent just exit.
>
> That makes the final child be re-parented to init, and it is invisible
> to wait() - all very much on purpose.
>
> This was (perhaps still is?) very commonly used for starting up
> background daemons (together with disassociating from the tty etc).
>
> So I don't mind th enew flag, but I think the restriction is
> unnecessary and not logical. Sometimes you simply don't *want*
> processes visible to wait - or care about a pidfd.
I'm completely fine removing that restriction and supporting autoreap
without pidfd.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-02-17 8:16 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-02-16 13:48 [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 1/2] clone: " Christian Brauner
2026-02-16 13:48 ` [PATCH RFC 2/2] selftests/pidfd: add CLONE_AUTOREAP tests Christian Brauner
2026-02-16 15:25 ` [PATCH RFC 0/2] pidfd: add CLONE_AUTOREAP Linus Torvalds
2026-02-17 8:16 ` Christian Brauner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®