Re: [RFC] module: init-failure path can free a module with live try_module_get() users

From: Petr Pavlu

Date: Fri Aug 28 2026 - 08:04:18 EST


On 8/24/26 8:13 AM, Mahanta Jambigi wrote:
> Hi Luis, Petr, Daniel, Sami, Aaron,
>
> I'm writing to ask about what looks like a generic module-init failure
> lifetime problem in the module loader. I ran into it while working on
> the SMC networking module (net/smc/), but after several patch
> iterations, it seems the root issue may belong in kernel/module/main.c
> rather than in SMC itself. I'd appreciate your guidance on whether this
> reading is correct, and if so, what fix direction would be preferred.
>
> THE ISSUE IN do_init_module()
> =============================
>
> include/linux/module.h has a long-standing FIXME in module_is_live():
>
> /* FIXME: It'd be nice to isolate modules during init, too, so they
> aren't used before they (may) fail. But presently too much code
> (IDE & SCSI) require entry into the module during init. */
> static inline bool module_is_live(struct module *mod)
> {
> return mod->state != MODULE_STATE_GOING;
> }
>
> Because MODULE_STATE_COMING is not MODULE_STATE_GOING, try_module_get()
> can succeed once a module's __init is executing. If __init makes the
> module externally reachable partway through and then later fails, the
> failure path in do_init_module() appears to do:
>
> fail:
> mod->state = MODULE_STATE_GOING;
> synchronize_rcu();
> module_put(mod);
> ...
> free_module(mod);
>
> synchronize_rcu() waits for RCU readers, but not for threads that
> already obtained a module reference via try_module_get() and are still
> executing module text.
>
> By contrast, the normal unload path in try_stop_module() refuses to
> proceed while the refcount is non-zero.
>
> So the asymmetry seems to be that the normal unload path waits for
> references to drain, while the init-failure path does not.
>
> A concrete race would look like:
>
> 1. Module __init registers an externally reachable interface.
> 2. User space enters through that interface and try_module_get()
> succeeds while the module is still COMING.
> 3. A later __init step fails.
> 4. do_init_module() frees the module.
> 5. The in-flight caller is still executing module text.
>
> SMC AS A CONCRETE EXAMPLE
> =========================
>
> In SMC, simply moving registration later does not appear to eliminate
> the window, because there are two separate registration points that can
> make the module reachable via socket():
>
> 1. sock_register(&smc_sock_family_ops)
> After this, socket(AF_SMC, ...) can succeed and reach
> try_module_get() via __sock_create().
>
> 2. smc_inet_init() -> inet_register_protosw()
> After this, socket(AF_INET, SOCK_STREAM, IPPROTO_SMC) can succeed
> and again reach try_module_get().
>
> Either registration point can succeed before a later init step fails.
>
> This may not be specific to SMC; other protocol modules that become
> reachable during init, such as Bluetooth, may have similar exposure and
> appear worth auditing as well.
>
> ON THE FIXME'S IDE/SCSI CONCERN
> ===============================
>
> The FIXME mentions IDE and SCSI as reasons not to isolate modules
> during init.
>
> 1. IDE was removed in Linux 5.14, so that half of the concern no
> longer applies.
>
> 2. SCSI still appears to self-reference during init
> (scsi_device_get() -> try_module_get(hostt->module) during
> scsi_scan_host()), so a blanket wait-for-refcount-to-drain
> approach in the failure path may deadlock there.
>
> Also, strong_try_module_get() already rejects MODULE_STATE_COMING with
> -EBUSY, so the infrastructure for refusing callers during init already
> exists in some form.
>
> QUESTIONS
> =========
>
> First, is my reading of this init-failure refcount/lifetime asymmetry
> correct?

Your analysis looks correct to me.

>
> If so, would one of the following directions be acceptable?
>
> 1. An opt-in mechanism (for example, a module flag) for modules that
> are safe to isolate during init and whose init-failure path should
> wait for external references to drain.

In general, it is preferred if the module loader handles all modules in
the same way.

I would say that the module loader should wait for external references
to drain after an init failure for all modules and that it should be the
responsibility of individual modules to ensure that this wait eventually
completes. Excluding some modules would mean that the module loader
could still free them while they are in use by the kernel.

Before such a wait, the module loader should cancel all idempotent
module loads. This is especially important during boot when several
udevd workers may be trying to insert the same module. In that case,
a failed module init function should block only a single udevd task, so
that the system can still boot properly.

>
> 2. Treating MODULE_STATE_COMING as non-live for normal
> try_module_get() users, with some explicit escape hatch for the
> remaining subsystems that genuinely need self-entry during init.

This might make sense, but the kernel currently has over 500 users of
try_module_get() so changing its semantics in this way would need to be
done very carefully.

I'm also not sure there is anything inherently wrong with calling
try_module_get() on a module while it is still executing its init
function. It must mean the module has reached a state in which it
properly initialized functionality to be registered with a specific
subsystem.

For instance, the __sock_create() function mentioned in your email
contains (simplified):

[...]
if (rcu_access_pointer(net_families[family]) == NULL)
request_module("net-pf-%d", family);

rcu_read_lock();
pf = rcu_dereference(net_families[family]);
err = -EAFNOSUPPORT;
if (!pf)
goto out_release;

if (!try_module_get(pf->owner))
goto out_release;
[...]

Consider two tasks that both call socket(XYZ). The first one sees
net_families[family]==NULL and calls request_module(XYZ). The module
then starts loading and its init function calls sock_register(XYZ). At
this point, the second task can see that net_families[family]!=NULL, so
it skips request_module(XYZ) and calls try_module_get(pf->owner). If
that call newly fails because the module is still in
MODULE_STATE_COMING, it breaks the autoload functionality of
__sock_create().

The scheme would need to be changed to call request_module(XYZ) if
either net_families[family]==NULL or !try_module_get(pf->owner).

>
> Any guidance on the preferred direction would be much appreciated.

Waiting for external references to drain after an init failure makes
sense to me. I'm less sure about changing the semantics of
try_module_get(), or whether strong_try_module_get() should instead be
exposed.

--
Thanks,
Petr