Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood

From: Borislav Petkov

Date: Wed Aug 26 2026 - 12:10:22 EST


On Wed, Aug 26, 2026 at 06:22:27AM -0700, Breno Leitao wrote:
> In other words, I'm more interested in flagging that a host is a lemon
> than in keeping it limping along half-alive, thus my approach in this
> patchset.

This is your use case but I'm sure others would prefer the limping and
gracefully saving work the system has done until now...

Putting my user hat on:

+ Panic once a machine check bank has logged this many corrected
+ errors. 0, the default, disables it.
+
+ The setting is global rather than per-CPU, and it is compared
+ against a running total kept per CPU and per bank. Lowering it
+ below a total a bank has already reached takes the machine
+ down on that bank's next corrected error.

Ah, so how many CEs is "many"?

I have read that correctable errors do not have any effect on the system and
they can keep happening and do not influence execution.

So what do I put there? I don't know... hmm.

... Dropping user hat. Putting devil's advocate hat...

How many CEs do you put in your kernels?

Does the same number fit any machine and platform?

Or you simply say, after X CEs, we simply panic the machine - instead of
gracefully powering it off after leaving a helpful message in dmesg - and have
lab technicians inspect it?

Hmmm.

--
Regards/Gruss,
Boris.

https://people.kernel.org/tglx/notes-about-netiquette