Re: [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel

From: Borislav Petkov

Date: Sun Sep 27 2026 - 14:15:26 EST


On Sun, Sep 27, 2026 at 05:19:04PM +0200, Adrian Schlegel wrote:
> cec_notifier() is meant to consume only correctable DRAM errors, as its
> comment says, but it relies on mce_is_memory_error() for that. On Intel
> and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
> (MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
> (000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
> Support memory error recovery for both UCNA and Deferred error in
> machine_check_poll") for uncorrected errors, where a cache error with a
> valid address can point at poisoned memory.

Looking at what calls mce_is_memory_error(), it looks to me like it should
exclude cache hierarchy errors but that's Tony's call.

> For corrected errors this does not hold. A corrected cache hierarchy
> error means a bit flipped in the cache and was fixed there. The reported
> address only names the line that happened to be cached. Counting these
> in the CEC makes it soft-offline healthy DRAM pages, and with the Intel
> action threshold of 2 from d25c6948a6aa ("RAS/CEC: Reduce offline page
> threshold for Intel systems"), two such errors on the same page are
> enough.
>
> This was observed on an i9-9900K without ECC memory and with a faulty
> core. Bank 3 on CPU 1 reports thousands of corrected cache errors
> (MCACOD 0x0135, 0x0151, 0x0179, ...) with addresses spread over the
> whole physical address space. Within one hour the CEC made 502
> soft-offline attempts on 159 distinct pages. 453 failed because the
> page was in kernel use, the others took healthy pages offline:
>
> RAS: Soft-offlining pfn: 0x100250
> mce: [Hardware Error]: TSC c43a09e330 ADDR 1002504c0 MISC 2514285
> Memory failure: 0x100250: unhandlable page.
>
> RAS: Soft-offlining pfn: 0x24f6c8
> mce: [Hardware Error]: CPU 1: Machine Check: 0 Bank 3: cc5ffdc000100151
> mce: [Hardware Error]: TSC 21903b589a2 ADDR 24f6c85c0 MISC 2516485
>
> This affects any Intel system with CONFIG_RAS_CEC=y and a core that
> repeatedly reports corrected cache errors, with or without ECC memory:
> healthy RAM is lost until reboot and HardwareCorrupted keeps growing,
> pages that cannot be offlined are retried over and over because the CEC
> drops the element after each attempt, and the messages point at failing
> DRAM while the defect is in the CPU.
>
> AMD already restricts this to DRAM ECC errors since c6708d50f166
> ("x86/MCE: Report only DRAM ECC as memory errors on AMD systems"). Do
> the equivalent for Intel and Zhaoxin, limited to the CEC.

A note for the text from here onwards:

Please, do not talk about *what* the patch is doing in the commit message
- that should be obvious from the diff itself. Rather, concentrate on the
*why* it needs to be done and why your patch exists.

It is perfectly fine to explain non-trivial aspects of the code the patch is
touching but do not regurgitate what it does.

See also https://docs.kernel.org/process/submitting-patches.html for
additional inspiration.

> Add a helper cec_is_dram_error() to drivers/ras/cec.c and use it in
> cec_notifier() instead of mce_is_memory_error(). On Intel and Zhaoxin
> it only accepts memory controller errors, i.e. MCACOD matching
> 000F 0000 1MMM CCCC, which is the first of the three checks in
> mce_is_memory_error(): (status & 0xef80) == BIT(7). For all other
> vendors it falls back to mce_is_memory_error(), so AMD and Hygon behave
> as before.
>
> Corrected cache hierarchy errors are no longer counted by the CEC.
> cec_notifier() returns NOTIFY_DONE for them, so they still reach the
> default notifier and are logged as before. mce_is_memory_error() itself
> is left unchanged because other users, like the nfit driver, rely on it
> for uncorrected errors. The action threshold and the handling of
> uncorrected errors are not touched.
>
> Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected
> cache error (status 0xcc5ffc0000100179) is no longer counted and its
> page stays online. A corrected memory controller error (status
> 0x8c0000000000009f) is still counted and its page soft-offlined.

Testing notes come...

> Fixes: 011d82611172 ("RAS: Add a Corrected Errors Collector")
> Signed-off-by: Adrian Schlegel <me@xxxxxxxxxxxxxxxxxx>
> ---

... under those three "---" of the patch so that they don't land in the commit
message.

Thx.

--
Regards/Gruss,
Boris.

https://people.kernel.org/tglx/notes-about-netiquette