Re: [PATCH 3/3] x86/mce/amd: Reset deferred-register state between polling banks

From: Yazen Ghannam

Date: Thu Sep 03 2026 - 16:46:33 EST


On Sun, Aug 30, 2026 at 04:45:25AM +0000, OptoCloud wrote:
> machine_check_poll() reuses a single struct mce_hw_err while
> iterating over the MCA banks. smca_should_log_poll_error() sets
> MCE_CHECK_DFR_REGS in ->kflags when an error was taken from
> MCA_DESTAT rather than MCA_STATUS. Nothing clears the bit again for
> the rest of the poll, including on iterations that bail out of
> should_log_poll_error() early, so every later bank in the same pass
> inherits it.
>
> A stale MCE_CHECK_DFR_REGS has two effects on a later bank:
>
> - mce_read_aux() reads MCx_DEADDR instead of MCA_ADDR, so the
> address reported for that bank is wrong.
>
> - amd_clear_bank() returns before writing 0 to MCA_STATUS, so the
> bank is logged again on the next poll.
>
> Clear MCE_CHECK_DFR_REGS at the start of each iteration, alongside
> the other bank-local resets.
>
> Found by code inspection; not reproduced on hardware.
>
> Fixes: 7cb735d7c0cb ("x86/mce: Unify AMD DFR handler with MCA Polling")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Eirik Bøe <git@xxxxxxxxxxxx>
> ---
> arch/x86/kernel/cpu/mce/core.c | 1 +
> 1 file changed, 1 insertion(+)
>
> diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c
> index e4e588d02a18..6ed14ec171f2 100644
> --- a/arch/x86/kernel/cpu/mce/core.c
> +++ b/arch/x86/kernel/cpu/mce/core.c
> @@ -811,6 +811,7 @@ void machine_check_poll(enum mcp_flags flags, mce_banks_t *b)
> m->synd = 0;
> err.vendor.amd.synd1 = 0;
> err.vendor.amd.synd2 = 0;
> + m->kflags &= ~MCE_CHECK_DFR_REGS;
> m->bank = i;
>
> barrier();
> --

Same feedback as patches 1 and 2.

Additionally, why single out the MCE_CHECK_DFR_REGS flag? Why not just
reset the entire kflags field?

Thanks,
Yazen