Re: [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel

From: Luck, Tony

Date: Tue Sep 29 2026 - 15:01:23 EST


On Sun, Sep 27, 2026 at 11:14:56AM -0700, Borislav Petkov wrote:
> On Sun, Sep 27, 2026 at 05:19:04PM +0200, Adrian Schlegel wrote:
> > cec_notifier() is meant to consume only correctable DRAM errors, as its
> > comment says, but it relies on mce_is_memory_error() for that. On Intel
> > and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
> > (MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
> > (000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
> > Support memory error recovery for both UCNA and Deferred error in
> > machine_check_poll") for uncorrected errors, where a cache error with a
> > valid address can point at poisoned memory.
>
> Looking at what calls mce_is_memory_error(), it looks to me like it should
> exclude cache hierarchy errors but that's Tony's call.

I did some GIT and mailing list archaeology on the origin of that
code that counts cache errors as memory errors. There was a bunch
of discussion, and four versions of the patch tweaking various bits,
but the inclusion of cache errors was in the patches from v1 and
neither I nor Boris ever questioned that part.

Twelve years later ... I still can't see why we didn't push back.

I agree that looking at all the callers of mce_is_memory_error()
they all expect it to do what its name implies: just return true
for memory errors.

So the right fix would be to nuke the BIT(8) and 0xc clauses and
most of the multi-line comment.

-Tony