Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood
From: Borislav Petkov
Date: Mon Aug 31 2026 - 00:58:15 EST
On Wed, Aug 26, 2026 at 10:32:11AM -0700, Luck, Tony wrote:
> The only problem they cause is performance due to the time taken to log them.
Yeah, and I think it is noticeable when you're working on a machine which
generates an excessive CE interrupts or drops into storms.
> Absolute numbers are a problem. Maybe you choose 8000 as your threshold.
> But then some system ticks along happily logging one error per hour with
> virtually no effect on system performance. Then panics when it has been
> up for almost a year because it reaches the 8000 threshold.
>
> Perhaps it would be better to adapt the storm code to allow tuning
> the criteria for a storm and have a mce=panic_on_storm option?
But we already have the CEC. And yes, it is only for memory errors now but we
can expand it to all error types and we can use the bucket decaying behavior
paired with the page offlining or even have a per-error-type recovery callback
there to handle it all without any user interaction, no?
I mean, *if* that is really necessary.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette