Re: [PATCH] pid_namespace: make zap_pid_ns_processes() freezable

From: Rafael J. Wysocki (Intel)

Date: Fri Oct 02 2026 - 14:15:19 EST


On Fri, Oct 2, 2026 at 8:06 PM Oleg Nesterov <oleg@xxxxxxxxxx> wrote:
>
> On 10/02, Rafael J. Wysocki (Intel) wrote:
> >
> > On Fri, Oct 2, 2026 at 4:26 PM Oleg Nesterov <oleg@xxxxxxxxxx> wrote:
> > >
> > > > @@ -243,6 +244,11 @@ void zap_pid_ns_processes(struct pid_namespace *pid_ns)
> > > > do {
> > > > clear_thread_flag(TIF_SIGPENDING);
> > > > rc = kernel_wait4(-1, NULL, __WALL, NULL);
> > > > + /*
> > > > + * The freezer's fake signal ends the wait with -ERESTARTSYS;
> > > > + * freeze here, or one stuck pid namespace blocks suspend.
> > > > + */
> > > > + try_to_freeze();
> > > > } while (rc != -ECHILD);
> > >
> > > Somehow I still think it would be better to change do_wait() to use
> > > TASK_INTERRUPTIBLE | TASK_FREEZABLE. Slightly less robust in theory,
> > > but I think should work in practice...
> > >
> > > Rafael, what do you think?
> >
> > Well, why exactly do you think that TASK_INTERRUPTIBLE would be better
> > than TASK_IDLE here?
>
> Hmm... it seems we don't understand each other. At least I certainly don't
> understand your "than TASK_IDLE here".
>
> I don't see TASK_IDLE "here", I see that this patch adds try_to_freeze()
> "here". And this is fine. But, I was thinking about this change instead:
>
> --- x/kernel/exit.c
> +++ x/kernel/exit.c
> @@ -1729,7 +1729,7 @@ static long do_wait(struct wait_opts *wo
> add_wait_queue(&current->signal->wait_chldexit, &wo->child_wait);
>
> do {
> - set_current_state(TASK_INTERRUPTIBLE);
> + set_current_state(TASK_INTERRUPTIBLE | TASK_FREEZABLE);
> retval = __do_wait(wo);
> if (retval != -ERESTARTSYS)
> break;
>
> because IMO it makes sense anyway. This way try_to_freeze_tasks() -> __freeze_task()
> path can freeze the tasks sleeping in do_wait() without waiting until they react to
> fake_signal_wake_up() and call get_signal() -> try_to_freeze().
>
> Plus this looks simpler, zap_pid_ns_processes() doesn't need another try_to_freeze().
>
> Sorry if I missed your point...

No worries and thanks for the explanation!

Yes, it looks simpler and it should actually work.

Aviv, what do you think?

> > > > for (;;) {
> > > > - set_current_state(TASK_INTERRUPTIBLE);
> > > > + /*
> > > > + * TASK_IDLE: no hung task warning or load for a wait that can
> > > > + * last as long as a tracer keeps a zombie, and a pending signal
> > > > + * (e.g. the freezer's fake one) can't turn it into a busy loop.
> > > > + * TASK_FREEZABLE: let the freezer freeze us while we wait.
> > > > + */
> > > > + set_current_state(TASK_IDLE | TASK_FREEZABLE);
> > >
> > > This looks like overdocumentation to me. I guess it was added by AI. Other users
> > > of IDLE/FREEZABLE do not try to document the meaning of these task states.
> >
> > It looks a bit like a note for self TBH.
>
> True. But...
>
> Ok, firstly the comment about TASK_IDLE looks a bit misleading to me. I mean the
> "as long as a tracer keeps a zombie" part. At this point the descedants traced from
> the parent namespace have already gone. The huge comment above this loop tries to
> explain the reasons for this "wait for pid_allocated == init_pids" in more details.
>
> Secondly. The comment about TASK_FREEZABLE looks fine. But why should the user,
> zap_pid_ns_processes(), explain the semantics of TASK_FREEZABLE? Probably this
> documentation makes sense, but then it should be moved to sched.h. Just my IMHO.
>
> > > But this is subjective, I won't insist.
> >
> > Same here.
>
> OK ;)
>
> Oleg.
>