Re: [PATCH net v1] pfcp: fix socket lifetime on netdevice registration failure

From: Xuanqiang Luo

Date: Wed Sep 23 2026 - 06:03:50 EST


On 2026/9/23 at 17:29 Simon Horman wrote:
On Fri, Sep 18, 2026 at 08:31:58PM +0800, Xuanqiang Luo wrote:
From: Xuanqiang Luo <luoxuanqiang@xxxxxxxxxx>

pfcp_newlink() creates the UDP socket before register_netdevice().
If registration fails after pfcp_dev_init() succeeds, the core calls
pfcp_dev_uninit(), which releases the socket and clears pfcp->sk.
The newlink error path then releases it again, causing a NULL pointer
dereference. This was observed with failslab fault injection:

FAULT_INJECTION: forcing a failure.
kobject: kobject_add_internal failed for pfcp0 (error: -12 parent: net)
BUG: KASAN: null-ptr-deref in udp_tunnel_sock_release+0x1c/0x50
Read of size 8 at addr 0000000000000120 by task ip/1037
Call trace:
show_stack+0x18/0x24 (C)
dump_stack_lvl+0x78/0x90
print_report+0x468/0x5cc
kasan_report+0xa4/0xf0
__asan_load8+0x7c/0xd0
udp_tunnel_sock_release+0x1c/0x50
pfcp_newlink+0x128/0x184
rtnl_newlink+0x848/0xe44
rtnetlink_rcv_msg+0x468/0x514
netlink_rcv_skb+0xc0/0x1f0
rtnetlink_rcv+0x18/0x24
netlink_unicast+0x4b8/0x558
netlink_sendmsg+0x2b8/0x584
...

Create the socket in pfcp_dev_init() after initializing the GRO cells,
and let pfcp_dev_uninit() release it on registration failure or removal.

Move the RCU wait into pfcp_dev_uninit(), using synchronize_net() after
socket release to drain receive callbacks before destroying the GRO cells.

Fixes: 76c8764ef36a5 ("pfcp: add PFCP module")
Signed-off-by: Xuanqiang Luo <luoxuanqiang@xxxxxxxxxx>
---
A NULL check would fix the crash. Moving socket creation into ndo_init
also removes the duplicate cleanup and ensures the receive state is
initialized before enabling the socket's receive callback.

Assuming the NULL check is significantly simpler than this patch,
and that it resolves the bug, think that we should consider:

1. A patch that minimal NULL check for net
2. Follow-up with the approach taken by this patch in net-next

I say this because this seems to be lower risk than skipping to applying
2 to net.

Also, please consider CCing stable on bug fixes.

Link: https://docs.kernel.org/process/maintainer-netdev.html#stable-tree

That makes sense. I left stable off Cc because the lifecycle changes
seemed too broad for a stable fix. Splitting the two makes sense now.

I'll send a v2 for net with just the NULL check and Cc stable, then
follow up with the lifecycle changes once the fix reaches net-next.

Thanks for the suggestion!!

Xuanqiang