Re: [PATCH v2] taskstats: route exit listener records through their netns
From: Greg KH
Date: Fri Oct 02 2026 - 04:34:51 EST
On Fri, Oct 02, 2026 at 04:58:42PM +0900, tjdqudcks0424@xxxxxxxxx wrote:
> From: 성병찬 <tjdqudcks0424@xxxxxxxxx>
>
> Commit edc73c7261ca ("kernel: make taskstats available from all net
> namespaces") made the taskstats Generic Netlink family available from all
> network namespaces. CPU-mask listener registrations, however,
> still store only a bare netlink port ID in a global per-CPU list and send
> exit records through init_net.
>
> Netlink port IDs are namespace-local. An administrator can register an
> exit listener after unsharing only the network namespace, while an
> unprivileged init_net socket binds the same numeric port ID. The latter
> then receives taskstats for exiting tasks of other UIDs despite being
> unable to register a listener or issue a direct taskstats query.
>
> Do not address this by rejecting listeners outside init_net. That was
> the v1 approach. No confirmed deployment relying on this combination
> was found, but taskstats has accepted the documented CPU-mask listener
> command there since v5.19 and in-tree tools use this interface. Preserve
> that behavior to avoid an unnecessary compatibility risk.
As this is the first public version of the patch, there's no need to
talk about a v1 here, it just confuses everyone involved.
> Associate each listener with taskstats family-private storage for the
> exact Generic Netlink socket. Record that socket's network namespace and
> port ID and use both for unicast. On socket release, the family-private
> destructor removes every listener owned by that socket. The socket pins
> its namespace until the destructor returns, so no additional net
> reference is needed.
>
> The per-CPU rwsem protects the listener-to-owner pointer from registration
> through unicast and removal. It also makes explicit deregistration,
> failed-send cleanup, and socket destruction mutually safe. Allocate a
> complete multi-CPU registration batch before publishing it so an
> allocation failure neither leaves a partial registration nor removes an
> older one.
>
> A purpose-built reproducer found and validated the issue in disposable
> QEMU guests. On the unmodified kernel, the child listener missed the
> record and the colliding init_net socket received it. In three fixed
> runs, the child listener received the record, the colliding socket timed
> out, and its direct query and registration returned EPERM. Init-net and
> child-net listeners, PID/TGID queries, deregistration, same-port listeners
> in two child netns, close and netns teardown races, KASAN, UBSAN, lockdep,
> listener counts, and kmemleak also passed.
>
> The per-socket Generic Netlink API exists in v6.8 and later. Older stable
> trees affected by the Fixes commit need a tailored backport.
>
> Fixes: edc73c7261ca ("kernel: make taskstats available from all net namespaces")
> Cc: stable@xxxxxxxxxxxxxxx # 6.8+
> Link: https://lore.kernel.org/all/20110630120831.GB7707@albatros/
> Link: https://lore.kernel.org/all/87v8x678ph.fsf@xxxxxxxxxxxxxxxxxxxxxxxxxxxxxx/
> Assisted-by: OpenAI Codex
> Signed-off-by: 성병찬 <tjdqudcks0424@xxxxxxxxx>
> ---
> Changes in v2:
> - Preserve CPU-mask listener registration in non-initial network
> namespaces.
> - Associate listeners with their registration network namespace.
> - Deliver exit records through the listener's namespace.
> - Handle listener cleanup across deregistration, socket close, and
> network namespace teardown.
> - Add the requested Assisted-by trailer.
> - Add cross-netns collision and teardown A/B test results.
>
> v1: https://lore.kernel.org/r/20261001223721.458667-2-tjdqudcks0424@xxxxxxxxx
>
> kernel/taskstats.c | 135 ++++++++++++++++++++++++++++++---------------
> 1 file changed, 91 insertions(+), 44 deletions(-)
This is a lot of change, is there a selftest to verify this all still
works properly somewhere? How did you test it?
thanks,
greg k-h