Re: [PATCH net] ipv4: fib_trie: prevent speculative out-of-bounds child access
From: Eric Dumazet
Date: Thu Oct 08 2026 - 18:26:06 EST
Le ven. 9 oct. 2026 à 00:10, Daniël Trujillo via B4 Relay
<devnull+datrujillo.nvidia.com@xxxxxxxxxx> a écrit :
>
> From: Daniël Trujillo <datrujillo@xxxxxxxxxx>
>
> fib_table_lookup() checks the child index against the bounds before
> passing it to get_child_rcu().
>
> The index is derived from the lookup key, which in turn is derived from
> flp->daddr. A userspace process can therefore influence the index by
> triggering lookups for chosen destination IPv4 addresses.
>
> If this bounds check is mispredicted as in-bounds, an out-of-bounds index
> can be transiently passed to get_child_rcu(). The value loaded beyond the
> child array is then treated as a node pointer. The next loop iteration
> dereferences that transient pointer through get_cindex(), causing a
> page-table walk whose cache footprint can leak (parts of) the pointer.
Has this been demonstrated, or is it the output of a static tool ?
key, pos and bits all live in the first 8 bytes of struct key_vector,
so the bound can not be made slow relative to the index by evicting
a cache line. The conditional branch should resolve a couple of cycles
after the out-of-bounds load is issued, before a dependent dereference
can start.
For a patch targeting net and stable, in the IPv4 lookup fast path,
we would like the changelog to say what was actually observed.
>
> Sanitize the index with array_index_nospec() before accessing the child
> array.
>
> Fixes: 9f9e636d4f89 ("fib_trie: Optimize fib_table_lookup to avoid wasting time on loops/variables")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Daniël Trujillo <datrujillo@xxxxxxxxxx>
> ---
> net/ipv4/fib_trie.c | 2 ++
> 1 file changed, 2 insertions(+)
>
> diff --git a/net/ipv4/fib_trie.c b/net/ipv4/fib_trie.c
> index 248514dce0cd..bd8a368c5502 100644
> --- a/net/ipv4/fib_trie.c
> +++ b/net/ipv4/fib_trie.c
> @@ -61,6 +61,7 @@
> #include <linux/export.h>
> #include <linux/vmalloc.h>
> #include <linux/notifier.h>
> +#include <linux/nospec.h>
> #include <net/net_namespace.h>
> #include <net/inet_dscp.h>
> #include <net/ip.h>
> @@ -1474,6 +1475,7 @@ int fib_table_lookup(struct fib_table *tb, const struct flowi4 *flp,
> cindex = index;
> }
>
> + index = array_index_nospec(index, 1ul << n->bits);
> n = get_child_rcu(n, index);
> if (unlikely(!n))
> goto backtrace;
This is not complete.
cindex has been set from the unclamped index just above. If the
(!n) branch is predicted taken under the same mispredicted bounds
check, the backtrace path does :
cindex &= cindex - 1;
cptr = &pn->tnode[cindex];
...
n = rcu_dereference(*cptr);
and the result is dereferenced at the top of step 2
(prefix_mismatch(key, n), n->slen, n->pos).
This is the same pattern you describe : one out-of-bounds load, then
a dereference of the loaded value.
The clamp needs to be done before cindex is recorded, and can stay
after the IS_LEAF() test so that leaf hits do not pay for it :
@@ -1465,7 +1466,9 @@ int fib_table_lookup(struct fib_table *tb, const
struct flowi4 *flp,
/* we have found a leaf. Prefixes have already been compared */
if (IS_LEAF(n))
goto found;
+ index = array_index_nospec(index, 1ul << n->bits);
+
/* only record pn and cindex if we are going to be chopping
* bits later. Otherwise we are just wasting cycles.
*/
Also, fib_find_node() has the exact same construct (same bounds check,
then get_child_rcu(n, index) on the next iteration), and is called
with a user provided key from fib_table_insert() and fib_table_delete().
RTM_NEWROUTE / RTM_DELROUTE only need CAP_NET_ADMIN in the netns owner
user namespace. This runs under RTNL, there is no performance concern
there.
Please also take a look at leaf_walk_rcu() (cindex >> pn->bits), and
say in the changelog why it is or is not affected.
>
> ---
> base-commit: 37f12441f557468a56c1e27790413aa78c82afa2
> change-id: 20261008-fib-trie-nospec-5c76028c8e7c
>
> Best regards,
> --
> Daniël Trujillo <datrujillo@xxxxxxxxxx>
>
>