Re: [PATCH v2] x86/mce: Do not treat cache hierarchy errors as memory errors

From: Luck, Tony

Date: Mon Oct 05 2026 - 11:31:30 EST


On Sun, Oct 04, 2026 at 12:20:12AM +0200, Adrian Schlegel wrote:

Boris pointed out the very long lines. I think you have more details
than necessary to describe the reason for this change. Be concise when
rewriting it.

> Fixes: fa92c5869426 ("x86, mce: Support memory error recovery for both UCNA and Deferred error in machine_check_poll")
> Signed-off-by: Adrian Schlegel <me@xxxxxxxxxxxxxxxxxx>
> ---
> v2:
> - Fix mce_is_memory_error() itself instead of adding a CEC-local helper (Tony)
> - v1: https://lore.kernel.org/r/20260927151904.30524-1-me@xxxxxxxxxxxxxxxxxx
>
> Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected cache error (status 0xcc5ffc0000100179) is no longer counted and its page stays online. A corrected memory controller error (status 0x8c0000000000009f) is still counted and its page soft-offlined.
>
> arch/x86/kernel/cpu/mce/core.c | 16 ++++------------
> 1 file changed, 4 insertions(+), 12 deletions(-)
>
> diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c
> index ab469605f..4d2ace76a 100644
> --- a/arch/x86/kernel/cpu/mce/core.c
> +++ b/arch/x86/kernel/cpu/mce/core.c
> @@ -542,19 +542,11 @@ bool mce_is_memory_error(struct mce *m)
> /*
> * Intel SDM Volume 3B - 15.9.2 Compound Error Codes

Intel keeps adding new chapters and sections to the SDM. This is now
18.10.2. The split into volumes has remained stable.

Future proof this by just referring to the section name without the explicit
section number. Should be a good enough pointer for someone to find the
reference.

* Intel SDM Volume 3B "Compound Error Codes".

> *
> - * Bit 7 of the MCACOD field of IA32_MCi_STATUS is used for
> - * indicating a memory error. Bit 8 is used for indicating a
> - * cache hierarchy error. The combination of bit 2 and bit 3
> - * is used for indicating a `generic' cache hierarchy error
> - * But we can't just blindly check the above bits, because if
> - * bit 11 is set, then it is a bus/interconnect error - and
> - * either way the above bits just gives more detail on what
> - * bus/interconnect error happened. Note that bit 12 can be
> - * ignored, as it's the "filter" bit.
> + * Memory controller errors have an MCACOD of 000F 0000 1MMM CCCC:
> + * bit 7 set, bits 8-11 and 13-15 clear. Bit 12 is the "filter"
> + * bit and is ignored.
> */
> - return (m->status & 0xef80) == BIT(7) ||
> - (m->status & 0xef00) == BIT(8) ||
> - (m->status & 0xeffc) == 0xc;
> + return (m->status & 0xef80) == BIT(7);

Otherwise update to comment and code looks good.
>
> default:
> return false;
> --
> 2.55.0

-Tony
>