From: Oleg Nesterov <oleg@redhat.com>
To: "Rafael J. Wysocki (Intel)" <rafael@kernel.org>
Cc: Aviv Vaknin <vaknins33@gmail.com>,
Christian Brauner <brauner@kernel.org>,
Pavel Tikhomirov <ptikhomirov@virtuozzo.com>,
"Eric W. Biederman" <ebiederm@xmission.com>,
linux-pm@vger.kernel.org, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH] pid_namespace: make zap_pid_ns_processes() freezable
Date: Fri, 2 Oct 2026 20:06:24 +0200 [thread overview]
Message-ID: <ar_yoLIkXpEgwPFr@redhat.com> (raw)
In-Reply-To: <CAJZ5v0g83_OCjh+RjG3H45mADE48vJdrEYahWGJShv62VAVvkA@mail.gmail.com>
On 10/02, Rafael J. Wysocki (Intel) wrote:
>
> On Fri, Oct 2, 2026 at 4:26 PM Oleg Nesterov <oleg@redhat.com> wrote:
> >
> > > @@ -243,6 +244,11 @@ void zap_pid_ns_processes(struct pid_namespace *pid_ns)
> > > do {
> > > clear_thread_flag(TIF_SIGPENDING);
> > > rc = kernel_wait4(-1, NULL, __WALL, NULL);
> > > + /*
> > > + * The freezer's fake signal ends the wait with -ERESTARTSYS;
> > > + * freeze here, or one stuck pid namespace blocks suspend.
> > > + */
> > > + try_to_freeze();
> > > } while (rc != -ECHILD);
> >
> > Somehow I still think it would be better to change do_wait() to use
> > TASK_INTERRUPTIBLE | TASK_FREEZABLE. Slightly less robust in theory,
> > but I think should work in practice...
> >
> > Rafael, what do you think?
>
> Well, why exactly do you think that TASK_INTERRUPTIBLE would be better
> than TASK_IDLE here?
Hmm... it seems we don't understand each other. At least I certainly don't
understand your "than TASK_IDLE here".
I don't see TASK_IDLE "here", I see that this patch adds try_to_freeze()
"here". And this is fine. But, I was thinking about this change instead:
--- x/kernel/exit.c
+++ x/kernel/exit.c
@@ -1729,7 +1729,7 @@ static long do_wait(struct wait_opts *wo
add_wait_queue(¤t->signal->wait_chldexit, &wo->child_wait);
do {
- set_current_state(TASK_INTERRUPTIBLE);
+ set_current_state(TASK_INTERRUPTIBLE | TASK_FREEZABLE);
retval = __do_wait(wo);
if (retval != -ERESTARTSYS)
break;
because IMO it makes sense anyway. This way try_to_freeze_tasks() -> __freeze_task()
path can freeze the tasks sleeping in do_wait() without waiting until they react to
fake_signal_wake_up() and call get_signal() -> try_to_freeze().
Plus this looks simpler, zap_pid_ns_processes() doesn't need another try_to_freeze().
Sorry if I missed your point...
> > > for (;;) {
> > > - set_current_state(TASK_INTERRUPTIBLE);
> > > + /*
> > > + * TASK_IDLE: no hung task warning or load for a wait that can
> > > + * last as long as a tracer keeps a zombie, and a pending signal
> > > + * (e.g. the freezer's fake one) can't turn it into a busy loop.
> > > + * TASK_FREEZABLE: let the freezer freeze us while we wait.
> > > + */
> > > + set_current_state(TASK_IDLE | TASK_FREEZABLE);
> >
> > This looks like overdocumentation to me. I guess it was added by AI. Other users
> > of IDLE/FREEZABLE do not try to document the meaning of these task states.
>
> It looks a bit like a note for self TBH.
True. But...
Ok, firstly the comment about TASK_IDLE looks a bit misleading to me. I mean the
"as long as a tracer keeps a zombie" part. At this point the descedants traced from
the parent namespace have already gone. The huge comment above this loop tries to
explain the reasons for this "wait for pid_allocated == init_pids" in more details.
Secondly. The comment about TASK_FREEZABLE looks fine. But why should the user,
zap_pid_ns_processes(), explain the semantics of TASK_FREEZABLE? Probably this
documentation makes sense, but then it should be moved to sched.h. Just my IMHO.
> > But this is subjective, I won't insist.
>
> Same here.
OK ;)
Oleg.
next prev parent reply other threads:[~2026-10-02 18:06 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 12:12 Aviv Vaknin
2026-10-01 13:03 ` Rafael J. Wysocki (Intel)
2026-10-01 15:35 ` Oleg Nesterov
2026-10-01 16:53 ` Bradley Morgan
2026-10-02 14:26 ` Oleg Nesterov
2026-10-02 16:24 ` Rafael J. Wysocki (Intel)
2026-10-02 18:06 ` Oleg Nesterov [this message]
2026-10-02 18:13 ` Rafael J. Wysocki (Intel)
2026-10-02 20:28 ` Aviv Vaknin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar_yoLIkXpEgwPFr@redhat.com \
--to=oleg@redhat.com \
--cc=brauner@kernel.org \
--cc=ebiederm@xmission.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pm@vger.kernel.org \
--cc=ptikhomirov@virtuozzo.com \
--cc=rafael@kernel.org \
--cc=stable@vger.kernel.org \
--cc=vaknins33@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®