Re: [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop()
From: Andrea Mayer
Date: Thu Oct 08 2026 - 20:44:37 EST
On Mon, 05 Oct 2026 14:31:53 +0900
Yuya Kusakabe <yuya.kusakabe@xxxxxxxxx> wrote:
> ip6_route_input() attaches the input route to the skb without taking a
> reference, relying on the RCU read-side section of the receive path.
> seg6_lookup_any_nexthop() is called from that same section, either
> directly from the seg6local input handler or from a netfilter okfn,
> but still takes and drops a route reference for every packet.
Hi Yuya,
Thanks for the patch and the throughput numbers. The Sashiko review asked
whether this also holds for the BPF callers in filter.c. Under
BPF_PROG_TEST_RUN the call does not come from the receive path: LWT_IN
programs reach seg6_lookup_nexthop() through bpf_push_seg6_encap(),
bpf_test_run() runs them many times on the same skb, and the RCU section
can end between runs. With the patch, on 32-bit ARM
(CONFIG_PREEMPT_VOLUNTARY, KASAN):
BUG: KASAN: slab-use-after-free in __seg6_do_srh_encap+0x3c/0x52c
Read of size 4 at addr c46b8100 by task bpftool/119
CPU: 0 UID: 0 PID: 119 Comm: bpftool [snip] 7.3.0-rc5 #1 VOLUNTARY
Tainted: [O]=OOT_MODULE
Hardware name: Generic DT based system
Call trace:
[snip]
__seg6_do_srh_encap from seg6_do_srh_encap+0x1c/0x24
seg6_do_srh_encap from bpf_push_seg6_encap+0x184/0x194
bpf_push_seg6_encap from bpf_lwt_in_push_encap+0x48/0x5c
bpf_lwt_in_push_encap from ___bpf_prog_run+0x1aa0/0x4334
___bpf_prog_run from __bpf_prog_run32+0xd4/0x114
__bpf_prog_run32 from bpf_test_run+0x2bc/0x57c
bpf_test_run from bpf_prog_test_run_skb+0x948/0x1100
bpf_prog_test_run_skb from __sys_bpf+0x10a8/0x2eac
__sys_bpf from sys_bpf+0x50/0x64
sys_bpf from ret_fast_syscall+0x0/0x54
[snip]
Allocated by task 119:
[snip]
dst_alloc+0x50/0xac
ip6_dst_alloc+0x24/0x7c
ip6_pol_route+0x468/0x780
ip6_pol_route_input+0x48/0x50
fib6_rule_action+0x16c/0x2e8
fib_rules_lookup+0x2b8/0x3ec
fib6_rule_lookup+0x228/0x308
seg6_lookup_any_nexthop+0x3e8/0x4d0
seg6_lookup_nexthop+0x1c/0x24
bpf_lwt_in_push_encap+0x48/0x5c
___bpf_prog_run+0x1aa0/0x4334
__bpf_prog_run32+0xd4/0x114
bpf_test_run+0x2bc/0x57c
bpf_prog_test_run_skb+0x948/0x1100
__sys_bpf+0x10a8/0x2eac
sys_bpf+0x50/0x64
ret_fast_syscall+0x0/0x54
Freed by task 123:
[snip]
kmem_cache_free+0xd8/0x3cc
rcu_core+0x434/0x1154
handle_softirqs+0x1d4/0x660
__irq_exit_rcu+0xe8/0x198
irq_exit+0x10/0x18
call_with_stack+0x18/0x20
__irq_svc+0x98/0xb0
finish_task_switch+0x158/0x4f4
__schedule+0x6c8/0x1710
schedule+0x38/0x124
schedule_hrtimeout_range_clock+0x170/0x1e0
usleep_range_state+0xf4/0x178
trigger_fn+0x1dc/0x238 [seg6_noref_trigger]
kthread+0x1c4/0x200
ret_from_fork+0x14/0x28
Last potentially related work creation:
[snip]
__call_rcu_common.constprop.0+0x4c/0x608
ip6_pol_route+0x348/0x780
ip6_pol_route_output+0x48/0x50
fib6_rule_action+0x16c/0x2e8
fib_rules_lookup+0x2b8/0x3ec
fib6_rule_lookup+0x228/0x308
ip6_route_output_flags+0x100/0x220
trigger_fn+0x20c/0x238 [seg6_noref_trigger]
kthread+0x1c4/0x200
ret_from_fork+0x14/0x28
The buggy address belongs to the object at c46b8100
which belongs to the cache ip6_dst_cache of size 156
The buggy address is located 0 bytes inside of
freed 156-byte region [c46b8100, c46b819c)
[snip]
The module in the trace is a small test helper. In a loop it calls
rt_genid_bump_ipv6(), as "ip -6 rule add" or a nexthop replace does, and
then does an output lookup on a route that shares the nexthop. I used the
module to make the window easier to hit.
Before the patch seg6_lookup_nexthop() used skb_dst_set(), so the skb held
a reference on its dst between runs.
> Look the route up with RT6_LOOKUP_F_DST_NOREF, as ip6_route_input()
> does, for the input and table lookups.
> [snip]
This also changes the contract of seg6_lookup_nexthop(). It does not
always take a reference now, and when it does not, the dst on the skb is
valid only inside the RCU section of the caller. seg6_lookup_nexthop() is
shared: it is declared in include/net/seg6.h and seg6_local.h, filter.c
uses it too, and neither the call sites nor the declarations say that the
lifetime changed.
One possibility would be to split it into two functions, conceptually
similar to ip6_route_output_flags() and ip6_route_output_flags_noref(): a
noref primitive for the callers in the seg6local input path, and a
refcounted wrapper around the primitive, which bpf_push_seg6_encap() would
call. The split would put the contract in the name and fix the
use-after-free above.
Ciao,
Andrea