mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel
@ 2026-09-27 15:19 Adrian Schlegel
  2026-09-27 18:14 ` Borislav Petkov
  0 siblings, 1 reply; 3+ messages in thread
From: Adrian Schlegel @ 2026-09-27 15:19 UTC (permalink / raw)
  To: Tony Luck, Borislav Petkov; +Cc: linux-edac, linux-kernel, Adrian Schlegel

cec_notifier() is meant to consume only correctable DRAM errors, as its
comment says, but it relies on mce_is_memory_error() for that. On Intel
and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
(MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
(000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
Support memory error recovery for both UCNA and Deferred error in
machine_check_poll") for uncorrected errors, where a cache error with a
valid address can point at poisoned memory.

For corrected errors this does not hold. A corrected cache hierarchy
error means a bit flipped in the cache and was fixed there. The reported
address only names the line that happened to be cached. Counting these
in the CEC makes it soft-offline healthy DRAM pages, and with the Intel
action threshold of 2 from d25c6948a6aa ("RAS/CEC: Reduce offline page
threshold for Intel systems"), two such errors on the same page are
enough.

This was observed on an i9-9900K without ECC memory and with a faulty
core. Bank 3 on CPU 1 reports thousands of corrected cache errors
(MCACOD 0x0135, 0x0151, 0x0179, ...) with addresses spread over the
whole physical address space. Within one hour the CEC made 502
soft-offline attempts on 159 distinct pages. 453 failed because the
page was in kernel use, the others took healthy pages offline:

  RAS: Soft-offlining pfn: 0x100250
  mce: [Hardware Error]: TSC c43a09e330 ADDR 1002504c0 MISC 2514285
  Memory failure: 0x100250: unhandlable page.

  RAS: Soft-offlining pfn: 0x24f6c8
  mce: [Hardware Error]: CPU 1: Machine Check: 0 Bank 3: cc5ffdc000100151
  mce: [Hardware Error]: TSC 21903b589a2 ADDR 24f6c85c0 MISC 2516485

This affects any Intel system with CONFIG_RAS_CEC=y and a core that
repeatedly reports corrected cache errors, with or without ECC memory:
healthy RAM is lost until reboot and HardwareCorrupted keeps growing,
pages that cannot be offlined are retried over and over because the CEC
drops the element after each attempt, and the messages point at failing
DRAM while the defect is in the CPU.

AMD already restricts this to DRAM ECC errors since c6708d50f166
("x86/MCE: Report only DRAM ECC as memory errors on AMD systems"). Do
the equivalent for Intel and Zhaoxin, limited to the CEC.

Add a helper cec_is_dram_error() to drivers/ras/cec.c and use it in
cec_notifier() instead of mce_is_memory_error(). On Intel and Zhaoxin
it only accepts memory controller errors, i.e. MCACOD matching
000F 0000 1MMM CCCC, which is the first of the three checks in
mce_is_memory_error(): (status & 0xef80) == BIT(7). For all other
vendors it falls back to mce_is_memory_error(), so AMD and Hygon behave
as before.

Corrected cache hierarchy errors are no longer counted by the CEC.
cec_notifier() returns NOTIFY_DONE for them, so they still reach the
default notifier and are logged as before. mce_is_memory_error() itself
is left unchanged because other users, like the nfit driver, rely on it
for uncorrected errors. The action threshold and the handling of
uncorrected errors are not touched.

Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected
cache error (status 0xcc5ffc0000100179) is no longer counted and its
page stays online. A corrected memory controller error (status
0x8c0000000000009f) is still counted and its page soft-offlined.

Fixes: 011d82611172 ("RAS: Add a Corrected Errors Collector")
Signed-off-by: Adrian Schlegel <me@adrianschlegel.com>
---
 drivers/ras/cec.c | 20 +++++++++++++++++++-
 1 file changed, 19 insertions(+), 1 deletion(-)

diff --git a/drivers/ras/cec.c b/drivers/ras/cec.c
index 15f7f043c..aa3fa633a 100644
--- a/drivers/ras/cec.c
+++ b/drivers/ras/cec.c
@@ -531,6 +531,24 @@ static int __init create_debugfs_nodes(void)
 	return 1;
 }
 
+/*
+ * Only corrected errors reported by the memory controller say something about
+ * the health of a DRAM page. Corrected cache hierarchy errors originate in the
+ * cache itself, their address only names the line that happened to be cached.
+ */
+static bool cec_is_dram_error(struct mce *m)
+{
+	switch (m->cpuvendor) {
+	case X86_VENDOR_INTEL:
+	case X86_VENDOR_ZHAOXIN:
+		/* Memory controller errors: MCACOD 000F 0000 1MMM CCCC */
+		return (m->status & 0xef80) == BIT(7);
+
+	default:
+		return mce_is_memory_error(m);
+	}
+}
+
 static int cec_notifier(struct notifier_block *nb, unsigned long val,
 			void *data)
 {
@@ -540,7 +558,7 @@ static int cec_notifier(struct notifier_block *nb, unsigned long val,
 		return NOTIFY_DONE;
 
 	/* We eat only correctable DRAM errors with usable addresses. */
-	if (mce_is_memory_error(m) &&
+	if (cec_is_dram_error(m) &&
 	    mce_is_correctable(m)  &&
 	    mce_usable_address(m)) {
 		if (!cec_add_elem(m->addr >> PAGE_SHIFT)) {
-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel
  2026-09-27 15:19 [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel Adrian Schlegel
@ 2026-09-27 18:14 ` Borislav Petkov
  2026-09-29 19:01   ` Luck, Tony
  0 siblings, 1 reply; 3+ messages in thread
From: Borislav Petkov @ 2026-09-27 18:14 UTC (permalink / raw)
  To: Adrian Schlegel; +Cc: Tony Luck, linux-edac, linux-kernel

On Sun, Sep 27, 2026 at 05:19:04PM +0200, Adrian Schlegel wrote:
> cec_notifier() is meant to consume only correctable DRAM errors, as its
> comment says, but it relies on mce_is_memory_error() for that. On Intel
> and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
> (MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
> (000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
> Support memory error recovery for both UCNA and Deferred error in
> machine_check_poll") for uncorrected errors, where a cache error with a
> valid address can point at poisoned memory.

Looking at what calls mce_is_memory_error(), it looks to me like it should
exclude cache hierarchy errors but that's Tony's call.

> For corrected errors this does not hold. A corrected cache hierarchy
> error means a bit flipped in the cache and was fixed there. The reported
> address only names the line that happened to be cached. Counting these
> in the CEC makes it soft-offline healthy DRAM pages, and with the Intel
> action threshold of 2 from d25c6948a6aa ("RAS/CEC: Reduce offline page
> threshold for Intel systems"), two such errors on the same page are
> enough.
> 
> This was observed on an i9-9900K without ECC memory and with a faulty
> core. Bank 3 on CPU 1 reports thousands of corrected cache errors
> (MCACOD 0x0135, 0x0151, 0x0179, ...) with addresses spread over the
> whole physical address space. Within one hour the CEC made 502
> soft-offline attempts on 159 distinct pages. 453 failed because the
> page was in kernel use, the others took healthy pages offline:
> 
>   RAS: Soft-offlining pfn: 0x100250
>   mce: [Hardware Error]: TSC c43a09e330 ADDR 1002504c0 MISC 2514285
>   Memory failure: 0x100250: unhandlable page.
> 
>   RAS: Soft-offlining pfn: 0x24f6c8
>   mce: [Hardware Error]: CPU 1: Machine Check: 0 Bank 3: cc5ffdc000100151
>   mce: [Hardware Error]: TSC 21903b589a2 ADDR 24f6c85c0 MISC 2516485
> 
> This affects any Intel system with CONFIG_RAS_CEC=y and a core that
> repeatedly reports corrected cache errors, with or without ECC memory:
> healthy RAM is lost until reboot and HardwareCorrupted keeps growing,
> pages that cannot be offlined are retried over and over because the CEC
> drops the element after each attempt, and the messages point at failing
> DRAM while the defect is in the CPU.
> 
> AMD already restricts this to DRAM ECC errors since c6708d50f166
> ("x86/MCE: Report only DRAM ECC as memory errors on AMD systems"). Do
> the equivalent for Intel and Zhaoxin, limited to the CEC.

A note for the text from here onwards:

Please, do not talk about *what* the patch is doing in the commit message
- that should be obvious from the diff itself. Rather, concentrate on the
*why* it needs to be done and why your patch exists.

It is perfectly fine to explain non-trivial aspects of the code the patch is
touching but do not regurgitate what it does.

See also https://docs.kernel.org/process/submitting-patches.html for
additional inspiration.

> Add a helper cec_is_dram_error() to drivers/ras/cec.c and use it in
> cec_notifier() instead of mce_is_memory_error(). On Intel and Zhaoxin
> it only accepts memory controller errors, i.e. MCACOD matching
> 000F 0000 1MMM CCCC, which is the first of the three checks in
> mce_is_memory_error(): (status & 0xef80) == BIT(7). For all other
> vendors it falls back to mce_is_memory_error(), so AMD and Hygon behave
> as before.
> 
> Corrected cache hierarchy errors are no longer counted by the CEC.
> cec_notifier() returns NOTIFY_DONE for them, so they still reach the
> default notifier and are logged as before. mce_is_memory_error() itself
> is left unchanged because other users, like the nfit driver, rely on it
> for uncorrected errors. The action threshold and the handling of
> uncorrected errors are not touched.
> 
> Tested with mce-inject (sw) in a VM with an Intel CPU model. A corrected
> cache error (status 0xcc5ffc0000100179) is no longer counted and its
> page stays online. A corrected memory controller error (status
> 0x8c0000000000009f) is still counted and its page soft-offlined.

Testing notes come...

> Fixes: 011d82611172 ("RAS: Add a Corrected Errors Collector")
> Signed-off-by: Adrian Schlegel <me@adrianschlegel.com>
> ---

... under those three "---" of the patch so that they don't land in the commit
message.

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel
  2026-09-27 18:14 ` Borislav Petkov
@ 2026-09-29 19:01   ` Luck, Tony
  0 siblings, 0 replies; 3+ messages in thread
From: Luck, Tony @ 2026-09-29 19:01 UTC (permalink / raw)
  To: Borislav Petkov; +Cc: Adrian Schlegel, linux-edac, linux-kernel

On Sun, Sep 27, 2026 at 11:14:56AM -0700, Borislav Petkov wrote:
> On Sun, Sep 27, 2026 at 05:19:04PM +0200, Adrian Schlegel wrote:
> > cec_notifier() is meant to consume only correctable DRAM errors, as its
> > comment says, but it relies on mce_is_memory_error() for that. On Intel
> > and Zhaoxin, mce_is_memory_error() also accepts cache hierarchy errors
> > (MCACOD 000F 0001 RRRR TTLL) and generic cache hierarchy errors
> > (000F 0000 0000 11LL). That check was added in fa92c5869426 ("x86, mce:
> > Support memory error recovery for both UCNA and Deferred error in
> > machine_check_poll") for uncorrected errors, where a cache error with a
> > valid address can point at poisoned memory.
> 
> Looking at what calls mce_is_memory_error(), it looks to me like it should
> exclude cache hierarchy errors but that's Tony's call.

I did some GIT and mailing list archaeology on the origin of that
code that counts cache errors as memory errors. There was a bunch
of discussion, and four versions of the patch tweaking various bits,
but the inclusion of cache errors was in the patches from v1 and
neither I nor Boris ever questioned that part.

Twelve years later ... I still can't see why we didn't push back.

I agree that looking at all the callers of mce_is_memory_error()
they all expect it to do what its name implies: just return true
for memory errors.

So the right fix would be to nuke the BIT(8) and 0xc clauses and
most of the multi-line comment.

-Tony

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-29 19:01 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 15:19 [PATCH] RAS/CEC: Only count corrected DRAM errors on Intel Adrian Schlegel
2026-09-27 18:14 ` Borislav Petkov
2026-09-29 19:01   ` Luck, Tony

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®