Re: [PATCH v3] RAS/AMD/FMPM: Fix spurious BUG when ERST record enumeration fails

From: Rui Qi

Date: Thu Oct 08 2026 - 04:30:56 EST


On Sun, Sep 27, 2026 at 10:28:07PM -0700, Borislav Petkov wrote:
> This sounds like this is something you can trigger. If so, how exactly?

I found this through code review, not a runtime reproduction, and on
closer look I can't trigger it on upstream:

- The erst_disable/-ENODEV path is unreachable through fmpm.
fru_mem_poison_init() checks erst_disable and returns -ENODEV
before get_saved_records() is ever called, so begin() never
runs with erst_disable set from this caller.

- The -EINTR path needs contention on erst_record_id_cache.lock
while fmpm's begin() runs. But begin() holds that lock only
across the refcount increment, and end() across the decrement
plus an in-memory cache compaction -- a microsecond-scale
window. Another ERST user (pstore, erst-dbg, apei_read_mce)
would have to be inside its own begin()/end() at the exact
instant fmpm loads, and a signal would have to hit the modprobe
task during that same window. I don't see a realistic way in.

> Do not explain the patch in your commit message.

Agreed, that paragraph goes.

> Because if you can trigger it with the upstream kernel, then this needs to go
> to stable.

Since I can't trigger it, I won't add Cc: stable. It's still a real
bug -- begin()/end() are a refcount pair, and the old path decrements
without a matching increment -- so I'd still like to fix it, but I'll
leave that call to you.

Reworded message below; it describes the bug and lets the diff show
the fix:

RAS/AMD/FMPM: Fix spurious BUG when ERST record enumeration fails

get_saved_records() calls erst_get_record_id_end() even when
erst_get_record_id_begin() has failed. But begin() only
increments erst_record_id_cache.refcount on success -- on failure
(e.g. -EINTR from mutex_lock_interruptible()) it leaves the
refcount untouched -- so the unconditional end() decrements it
below zero and trips BUG_ON(refcount < 0).

Fixes: 6f15e617cc99 ("RAS: Introduce a FRU memory poison manager")
Signed-off-by: Rui Qi <qirui.001@xxxxxxxxxxxxx>

If the wording's fine I'll send v4 as a new thread; say the word if
you'd rather drop it.