Re: [PATCH v3] RAS/AMD/FMPM: Fix spurious BUG when ERST record enumeration fails
From: Borislav Petkov
Date: Mon Sep 28 2026 - 01:28:40 EST
On Fri, Sep 25, 2026 at 11:09:18AM +0800, Rui Qi wrote:
> When erst_get_record_id_begin() returns an error, get_saved_records()
> jumps to the out_end label and unconditionally calls
> erst_get_record_id_end(). But begin() only bumps the
> erst_record_id_cache refcount on success, so calling end() after a
> failed begin() underflows it.
>
> This is reachable when fmpm is built as a module: if the cache lock
> is contended (by pstore, erst-dbg, or apei_read_mce) and a signal
> arrives, mutex_lock_interruptible() in begin() returns -EINTR without
> incrementing the refcount. end() then takes the lock, decrements the
> refcount below zero, and BUG_ON() fires while still holding it. The
> loader dies with the mutex held, so every later ERST user blocks
> indefinitely.
This sounds like this is something you can trigger. If so, how exactly?
> Built-in fmpm is not affected, since its initcall runs before
> userspace and mutex_lock_interruptible() can never see a signal.
>
> Fix it by jumping to the out label on begin() failure, skipping
> erst_get_record_id_end(). kfree(old) moves to that shared label so
> it still runs on every path.
Do not explain the patch in your commit message.
> Fixes: 6f15e617cc99 ("RAS: Introduce a FRU memory poison manager")
> Signed-off-by: Rui Qi <qirui.001@xxxxxxxxxxxxx>
Because if you can trigger it with the upstream kernel, then this needs to go
to stable.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette