Re: [PATCH net v2] amt: do not store tunnel pointer in skb control block

From: Omar Ramadan

Date: Mon Sep 28 2026 - 14:28:58 EST


On Sat, Sep 26, 2026, Sashiko wrote:
> [Severity: Medium]
> Can this lookup return a tunnel that amt_request_handler() has published but
> not finished initializing?

Yes, I think this one is real for v2. amt_request_handler() makes a
new tunnel visible with list_add_tail_rcu() before it writes nonce and
mac at the send: label. A query left in the qdisc from an expired
tunnel can therefore match a re-created tunnel for the same gateway
and port during that window.

The direct-send approach proposed earlier in this thread avoids it by
construction. amt_send_igmp_gq() and amt_send_mld_gq() run at the end
of amt_request_handler(), after the same context has written nonce and
mac. They call amt_send_membership_query() for that tunnel, and nothing
is looked up again at dequeue. I've posted that as v3:
https://lore.kernel.org/netdev/20260928181601.85857-1-omar@xxxxxxxxxxxxx/

I ran Cen's reproducer on net at 17741334d00 (KASAN, slub_debug=FZU):
the unpatched tree reports the slab-use-after-free in amt_dev_xmit(),
and both v2 and the direct-send diff run clean. The reproducer does
not exercise the re-creation window above, so this confirms the UAF
fix, not the v2 race.

> [Severity: Medium]
> [...] a query that was successfully handed to udp_tunnel_xmit_skb() is
> counted as tx_dropped

v3 removes the relay query branch from amt_dev_xmit(), so a General
Query that was sent is no longer counted as dropped. The gateway report
path (a successful amt_send_membership_update() followed by goto
unlock) has the same miscount. That is independent of the UAF; I'll
send a separate patch for it once v3 is in, since it applies on top.

> [Severity: High]
> amt_dev_stop() deletes the same entries with no lock at all

This predates the fix and is independent of it. Reading net, it looks
right to me: nothing disables the per-tunnel gc_wq before the unlocked
loop. cancel_delayed_work_sync() comes after list_del_rcu(), so it can
wait for a running amt_tunnel_expire() that then deletes the entry and
calls kfree_rcu() on it a second time. Cen already posted a fix for
this, "amt: fix tunnel list corruption on device stop":
https://patchwork.kernel.org/project/netdevbpf/patch/20260822045407.28983-1-blbllhy@xxxxxxxxx/
so I'll leave that one to that thread.

pw-bot: cr