[PATCH 3/3] x86/mce/amd: Reset deferred-register state between polling banks

From: OptoCloud

Date: Sun Aug 30 2026 - 00:45:55 EST


machine_check_poll() reuses a single struct mce_hw_err while
iterating over the MCA banks. smca_should_log_poll_error() sets
MCE_CHECK_DFR_REGS in ->kflags when an error was taken from
MCA_DESTAT rather than MCA_STATUS. Nothing clears the bit again for
the rest of the poll, including on iterations that bail out of
should_log_poll_error() early, so every later bank in the same pass
inherits it.

A stale MCE_CHECK_DFR_REGS has two effects on a later bank:

- mce_read_aux() reads MCx_DEADDR instead of MCA_ADDR, so the
address reported for that bank is wrong.

- amd_clear_bank() returns before writing 0 to MCA_STATUS, so the
bank is logged again on the next poll.

Clear MCE_CHECK_DFR_REGS at the start of each iteration, alongside
the other bank-local resets.

Found by code inspection; not reproduced on hardware.

Fixes: 7cb735d7c0cb ("x86/mce: Unify AMD DFR handler with MCA Polling")
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Eirik Bøe <git@xxxxxxxxxxxx>
---
arch/x86/kernel/cpu/mce/core.c | 1 +
1 file changed, 1 insertion(+)

diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c
index e4e588d02a18..6ed14ec171f2 100644
--- a/arch/x86/kernel/cpu/mce/core.c
+++ b/arch/x86/kernel/cpu/mce/core.c
@@ -811,6 +811,7 @@ void machine_check_poll(enum mcp_flags flags, mce_banks_t *b)
m->synd = 0;
err.vendor.amd.synd1 = 0;
err.vendor.amd.synd2 = 0;
+ m->kflags &= ~MCE_CHECK_DFR_REGS;
m->bank = i;

barrier();
--
2.47.3