Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood
From: Luck, Tony
Date: Mon Aug 24 2026 - 16:20:58 EST
On Mon, Aug 24, 2026 at 09:31:02AM -0700, Luck, Tony wrote:
> On Mon, Aug 24, 2026 at 01:29:18AM -0700, Breno Leitao wrote:
> > On Fri, Aug 21, 2026 at 09:50:19AM -0700, Luck, Tony wrote:
> >
> > > > Count corrected errors per bank and add mce=panic_on_ce_count=<count>
> > > > and panic the host if we have more events than set.
> > >
> > > FYI. I don't think this needs to be fixed, but you should be aware and
> > > perhaps document the shared bank details.
> > >
> > > This won't count accurately for banks that are shared by multiple logical
> > > CPUs (you've inherited this from the storm detection code that introduces
> > > this problem).
> > >
> > > E.g. a machine check bank reporting L2 errors is shared by both logical CPUs
> > > on a core on P-core systems, and by all cores on a module on E-core systems.
> >
> > Good point, thanks. I had not thought about the shared bank case.
> >
> > I will document it in v2, roughly like this under panic_on_ce_count=
> > in kernel-parameters.txt:
> >
> > The count is kept per CPU and per bank. A bank shared by
> > several logical CPUs -- an L2 bank is shared by the SMT
> > siblings, for instance -- has its errors split across those
> > CPUs.
> >
> > plus a line in the changelog noting that the count inherits the
> > <per-CPU,per-bank> granularity of the storm tracking. Thus, there may be
> > more errors logged than you expect before your panic fires. Would it be better?
>
> Yes. Documentation like this in the commit, and under Documentation/*
> gives fair notice to users.
>
> > One question while I am here: is the approach itself something you are
> > willing to take, or, just a just send patch [1/2]?
>
> Boris is the maintainer here. So he has the final decision. My
> opinion is that this is a bit niche for the kernel. Maybe this
> could be handled by some user agent (mcelog? rasdaemon?) to
> raise awareness to system operators about high rates of corrected
> errors.
I chatted with someone internally. They'd be very interested in being
able to see these running counts via some /sys files.
Perhaps your use case would be better for that too? Rather than:
"Everything is fine"
...
"Everything is fine"
...
"Everything is fine"
...
"Everything is fine"
...
"Count exceeded, PANIC!"
> > Thanks for the review,
> > --breno
>
> -Tony
-Tony
>