[PATCH v2] sched/topology: Free NUMA masks on topology allocation failure

From: Fengyu Wang

Date: Wed Aug 12 2026 - 02:23:57 EST


sched_init_numa() publishes sched_domains_numa_masks before it
allocates the topology array. When that allocation fails, the early
return leaves the masks published while sched_domains_numa_levels is
still zero: nothing dereferences them, but nothing can free them
either, and the topology they were built for is never installed.

Free the masks on that path, and publish them only once the topology
array they were built for has been allocated.

Fixes: cb83b629bae0 ("sched/numa: Rewrite the CONFIG_NUMA sched domain support")
Signed-off-by: Fengyu Wang <wangfengyu@xxxxxxxx>
---
v2:
- Publish sched_domains_numa_masks only after the topology array has
been allocated, instead of publishing it early and unpublishing it
on the failure path. This drops the rcu_assign_pointer(NULL) and
the synchronize_rcu() from the error path (Tim Chen).

v1: https://lore.kernel.org/lkml/20260731081413.5505-1-wangfengyu@xxxxxxxx/

Tested by hardcoding tl to NULL right after the kzalloc() to force the
failure path; the masks are released and the machine boots normally.

kernel/sched/topology.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)

diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c
index 21e816ad23ee..50457f720808 100644
--- a/kernel/sched/topology.c
+++ b/kernel/sched/topology.c
@@ -2392,15 +2392,23 @@ void sched_init_numa(int offline_node)
}
}
}
- rcu_assign_pointer(sched_domains_numa_masks, masks);

/* Compute default topology size */
for (i = 0; sched_domain_topology[i].mask; i++);

tl = kzalloc((i + nr_levels + 1) *
sizeof(struct sched_domain_topology_level), GFP_KERNEL);
- if (!tl)
+ if (!tl) {
+ for (i = 0; i < nr_levels; i++) {
+ for_each_node(j)
+ kfree(masks[i][j]);
+ kfree(masks[i]);
+ }
+ kfree(masks);
return;
+ }
+
+ rcu_assign_pointer(sched_domains_numa_masks, masks);

/*
* Copy the default topology bits..
--
2.34.1