Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood
From: Borislav Petkov
Date: Tue Aug 25 2026 - 15:16:40 EST
On Tue, Aug 25, 2026 at 04:26:26PM +0000, Luck, Tony wrote:
> > Why isn't the strategy here page offlining and shutting up the source of the
> > error instead of doing silly counting and not doing anything to contain the
> > errors in the first place?
>
> While Breno wasn't specific about the source of the errors in these messages, I'm
> guessing that they might be coming from cache errors rather than DDR memory.
>
> There isn't a good way for software[1] to suppress these errors. Taking memory
> pages offline when the problem is the cache will just deplete available memory
> without solving the problem.
I have been thinking about this *years* ago. If it is cache errors, we should
simply offline the core or cores using that cache. We have a lot of cores
nowadays :)
In general, us being a lot more resilient and applying automatic containment
and recovery actions should be the goal IMO.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette