Re: [PATCH net] netlink: do not copy to user space under netlink_lock_table()

From: Eric Dumazet

Date: Thu Oct 08 2026 - 16:42:46 EST


Le jeu. 8 oct. 2026 à 21:48, Cen Zhang (Microsoft)
<cenzhang@xxxxxxxxxxxxxxxxxxx> a écrit :
>
> netlink_getsockopt() handles NETLINK_LIST_MEMBERSHIPS by copying the
> socket's group bitmap to user space with copy_to_iter() under
> netlink_lock_table(). However, copy_to_iter() writes to a user address
> and can trigger a page fault, and how long that fault takes is up to the
> owner of the page. An unprivileged user can make it take as long as they
> want, e.g. by mapping a memfd page and keeping it hole punched from
> another thread. While the fault is pending, the caller still counts as a
> holder of nl_table_users.
>
> nl_table_users is one global counter shared by every netlink protocol
> and every network namespace, and netlink_table_grab() waits for it to
> drop to zero in TASK_UNINTERRUPTIBLE with no timeout and no signal
> check. Therefore, by holding one page fault open, an unprivileged user
> blocks every task on the machine that needs netlink_table_grab(): e.g.,
> close of a netlink socket, multicast group join, and network namespace
> creation. The blocked tasks stay in an unkillable D state until the
> attacker lets the fault finish, and hung task detector reports is as
> following:
>
> INFO: task V1-nl-close:103 blocked for more than 122 seconds.
> task:V1-nl-close state:D stack:14448 pid:103 ...
> Call Trace:
> schedule+0x36/0xf0
> netlink_table_grab.part.0+0x66/0xe0
> netlink_release+0x6d4/0x780
> __x64_sys_close+0x38/0x80
>
> Fix this by taking a kmemdup() snapshot of the bitmap under
> netlink_lock_table() and copying the snapshot to user space after the
> lock is released.
>
> Fixes: 47191d65b647 ("netlink: fix locking around NETLINK_LIST_MEMBERSHIPS")
> Reported-by: AutonomousCodeSecurity@xxxxxxxxxxxxx
> Signed-off-by: Cen Zhang (Microsoft) <cenzhang@xxxxxxxxxxxxxxxxxxx>
> ---

This is certainly not a net candidate. I call this AI hallucination.

This is not specific to netlink_lock_table().

We have many places where copy_{from,to}_user() runs while a mutex or
a socket lock is held, and some of these locks are global.

One example among many: do_ip_getsockopt() handles IP_MSFILTER with
copy_from_sockptr(), then copy_to_sockptr() in ip_mc_msfget(), all of
it under rtnl_lock(). No privilege needed.

Even in af_netlink.c, your patch does not close the issue you describe:
netlink_bind() calls netlink_insert() under netlink_lock_table(), and
netlink_insert() starts with lock_sock(sk). Another thread can hold
that socket lock across a user copy, e.g. setsockopt(SO_ATTACH_FILTER),
where __get_filter() calls copy_from_user() under sockopt_lock_sock().

So if "user controlled page fault while holding a lock others wait on"
is the bug, a real fix is probably hundreds of patches all over the
tree, not this one.

I do not think we want to do this one call site at a time, and
certainly not as fixes for net and stable.

What is the plan here?