[PATCH] RAS/CEC: Only count corrected DRAM errors on Intel
From: Adrian Schlegel
Date: Sun Sep 27 2026 - 11:19:23 EST
cec_notifier() is meant to consume only correctable DRAM errors, as its
comment says, but it relies on mce_is_memory_error() for that. On Intel
and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
(MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
(000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
Support memory error recovery for both UCNA and Deferred error in
machine_check_poll") for uncorrected errors, where a cache error with a
valid address can point at poisoned memory.
For corrected errors this does not hold. A corrected cache hierarchy
error means a bit flipped in the cache and was fixed there. The reported
address only names the line that happened to be cached. Counting these
in the CEC makes it soft-offline healthy DRAM pages, and with the Intel
action threshold of 2 from d25c6948a6aa ("RAS/CEC: Reduce offline page
threshold for Intel systems"), two such errors on the same page are
enough.
This was observed on an i9-9900K without ECC memory and with a faulty
core. Bank 3 on CPU 1 reports thousands of corrected cache errors
(MCACOD 0x0135, 0x0151, 0x0179, ...) with addresses spread over the
whole physical address space. Within one hour the CEC made 502
soft-offline attempts on 159 distinct pages. 453 failed because the
page was in kernel use, the others took healthy pages offline:
RAS: Soft-offlining pfn: 0x100250
mce: [Hardware Error]: TSC c43a09e330 ADDR 1002504c0 MISC 2514285
Memory failure: 0x100250: unhandlable page.
RAS: Soft-offlining pfn: 0x24f6c8
mce: [Hardware Error]: CPU 1: Machine Check: 0 Bank 3: cc5ffdc000100151
mce: [Hardware Error]: TSC 21903b589a2 ADDR 24f6c85c0 MISC 2516485
This affects any Intel system with CONFIG_RAS_CEC=y and a core that
repeatedly reports corrected cache errors, with or without ECC memory:
healthy RAM is lost until reboot and HardwareCorrupted keeps growing,
pages that cannot be offlined are retried over and over because the CEC
drops the element after each attempt, and the messages point at failing
DRAM while the defect is in the CPU.
AMD already restricts this to DRAM ECC errors since c6708d50f166
("x86/MCE: Report only DRAM ECC as memory errors on AMD systems"). Do
the equivalent for Intel and Zhaoxin, limited to the CEC.
Add a helper cec_is_dram_error() to drivers/ras/cec.c and use it in
cec_notifier() instead of mce_is_memory_error(). On Intel and Zhaoxin
it only accepts memory controller errors, i.e. MCACOD matching
000F 0000 1MMM CCCC, which is the first of the three checks in
mce_is_memory_error(): (status & 0xef80) == BIT(7). For all other
vendors it falls back to mce_is_memory_error(), so AMD and Hygon behave
as before.
Corrected cache hierarchy errors are no longer counted by the CEC.
cec_notifier() returns NOTIFY_DONE for them, so they still reach the
default notifier and are logged as before. mce_is_memory_error() itself
is left unchanged because other users, like the nfit driver, rely on it
for uncorrected errors. The action threshold and the handling of
uncorrected errors are not touched.
Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected
cache error (status 0xcc5ffc0000100179) is no longer counted and its
page stays online. A corrected memory controller error (status
0x8c0000000000009f) is still counted and its page soft-offlined.
Fixes: 011d82611172 ("RAS: Add a Corrected Errors Collector")
Signed-off-by: Adrian Schlegel <me@xxxxxxxxxxxxxxxxxx>
---
drivers/ras/cec.c | 20 +++++++++++++++++++-
1 file changed, 19 insertions(+), 1 deletion(-)
diff --git a/drivers/ras/cec.c b/drivers/ras/cec.c
index 15f7f043c..aa3fa633a 100644
--- a/drivers/ras/cec.c
+++ b/drivers/ras/cec.c
@@ -531,6 +531,24 @@ static int __init create_debugfs_nodes(void)
return 1;
}
+/*
+ * Only corrected errors reported by the memory controller say something about
+ * the health of a DRAM page. Corrected cache hierarchy errors originate in the
+ * cache itself, their address only names the line that happened to be cached.
+ */
+static bool cec_is_dram_error(struct mce *m)
+{
+ switch (m->cpuvendor) {
+ case X86_VENDOR_INTEL:
+ case X86_VENDOR_ZHAOXIN:
+ /* Memory controller errors: MCACOD 000F 0000 1MMM CCCC */
+ return (m->status & 0xef80) == BIT(7);
+
+ default:
+ return mce_is_memory_error(m);
+ }
+}
+
static int cec_notifier(struct notifier_block *nb, unsigned long val,
void *data)
{
@@ -540,7 +558,7 @@ static int cec_notifier(struct notifier_block *nb, unsigned long val,
return NOTIFY_DONE;
/* We eat only correctable DRAM errors with usable addresses. */
- if (mce_is_memory_error(m) &&
+ if (cec_is_dram_error(m) &&
mce_is_correctable(m) &&
mce_usable_address(m)) {
if (!cec_add_elem(m->addr >> PAGE_SHIFT)) {
--
2.55.0