Re: [PATCH net-next v3 0/3] net: hash uncached route lists by device

From: Kuniyuki Iwashima

Date: Fri Sep 18 2026 - 14:48:09 EST


On Fri, Sep 18, 2026 at 11:24 AM Chris Arges <carges@xxxxxxxxxxxxxx> wrote:
>
> On 2026-09-18 04:19:34, Kuniyuki Iwashima wrote:
> > From: Chris Arges <carges@xxxxxxxxxxxxxx>
> > Date: Thu, 17 Sep 2026 19:38:40 -0500
> > > On 2026-09-17 22:10:22, Kuniyuki Iwashima wrote:
> > > > From: Chris J Arges <carges@xxxxxxxxxxxxxx>
> > > > Date: Thu, 17 Sep 2026 14:38:21 -0500
> > > > > We have observed hung tasks blocked on rtnl_mutex while network namespaces
> > > > > were being removed. The namespaces contained many network devices, and the
> > > > > host had accumulated a large population of entries on the global per-CPU
> > > > > uncached route lists. A perf profile collected during one incident
> > > > > attributed most of the cleanup worker's samples to rt_flush_dev():
> > > > >
> > > > > ```
> > > > > 99.92% kworker/u384:3- worker_thread
> > > > > `-88.71% process_one_work
> > > > > `-81.02% cleanup_net
> > > > > `-81.00% unregister_netdevice_many_notify
> > > > > `-79.42% notifier_call_chain
> > > > > `-78.05% fib_netdev_event
> > > > > `-77.92% rt_flush_dev
> > > > > ```
> > > > >
> > > > > For each device, rt_flush_dev() visits every possible CPU and scans the
> > > > > global uncached route population while its caller holds rtnl_mutex. If N is
> > > > > the number of devices, C the number of possible CPUs, and R the number of
> > > > > uncached routes, the cost is O(N * (C + R)).
> > > > >
> > > > > During namespace cleanup, other processes that issue RTNETLINK operations
> > > > > requiring the RTNL lock can stall until cleanup releases the lock.
> > > > >
> > > > > A minimal reproducer is available here:
> > > > > https://github.com/arges/linux-reproducers/tree/main/rtnl-flush-storm
> > > > >
> > > > > This series replaces each per-CPU uncached route list with a hash table
> > > > > using the network device as its key. Each table uses 64 buckets.
> > > >
> > > > This sounds a bit overkill. Also, this series still leaves
> > > > O(N * C) loops.
> > > >
> > > > Given unregistering a single device is less common than
> > > > destroying netns, I think the right approach should be to
> > > > make the route flush once in cleanup_net() + outside RTNL.
> > > >
> > > > Could you try this change ? (only compile-tested)
> > > >
> > > Excellent, I'll test this and report back.
> >
> > I found a pre-existing issue, which affects the previous
> > diff, so on top of it, please apply this patch
> >
> > https://lore.kernel.org/netdev/20260918041439.2575935-1-kuniyu@xxxxxxxxxx/T/#u
> >
> > and this diff :
> >
> > ---8<---
> > diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> > index d35b66b33bbc..c12e20e07749 100644
> > --- a/net/ipv4/route.c
> > +++ b/net/ipv4/route.c
> > @@ -1567,12 +1567,13 @@ void rt_add_uncached_list(struct rtable *rt)
> > {
> > struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
> >
> > + rt->dst.rt_uncached_list = ul;
> > +
> > spin_lock_bh(&ul->lock);
> >
> > if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> > rt_replace_uncached_list(rt);
> > } else {
> > - rt->dst.rt_uncached_list = ul;
> > list_add_tail(&rt->dst.rt_uncached, &ul->head);
> > }
> >
> > diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> > index 6cffe8440b44..f22793abbbb8 100644
> > --- a/net/ipv6/route.c
> > +++ b/net/ipv6/route.c
> > @@ -155,12 +155,13 @@ void rt6_uncached_list_add(struct rt6_info *rt)
> > {
> > struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
> >
> > + rt->dst.rt_uncached_list = ul;
> > +
> > spin_lock_bh(&ul->lock);
> >
> > if (!check_net(dst_dev_net_rcu(&rt->dst))) {
> > rt6_uncached_list_replace(rt);
> > } else {
> > - rt->dst.rt_uncached_list = ul;
> > list_add_tail(&rt->dst.rt_uncached, &ul->head);
> > }
> >
> > ---8<---
>
> Kuniyuki,
>
> I was able to test this diff, the previous diff you sent plus the fixup
> mentioned above. I was able to confirm even greater reduction in contention
> as measured by how much latency an unrelated process takes when waiting for
> cleanup_net to complete. This makes sense since we don't even need to hold the
> lock when processing those routing entries with your patch.
>
> Some rough average latency numbers with 36 devices, 160k routes, 8 vCPUs:
> - main: 137ms
> - my hashing proposal: 13ms
> - your patchset: 1.8ms
>
> I'd be happy to retest any proposed patches.

Great, thank you for testing !

I will post patches officially once my fix lands in net-next
(so should be after next Thursday)

Thanks !