From: Tony Luck <tony.luck@intel.com>
To: Tony Luck <tony.luck@intel.com>
Cc: "Borislav Petkov" <bp@alien8.de>,
"Qiuxu Zhuo" <qiuxu.zhuo@intel.com>,
"Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>,
"Breno Leitao" <leitao@debian.org>,
linux-edac@vger.kernel.org, linux-kernel@vger.kernel.org,
patches@lists.linux.dev
Subject: [PATCH v2 6/7] EDAC/intel-bff: Compute unique ID for overflowed filter
Date: Fri, 28 Aug 2026 08:30:01 -0700 [thread overview]
Message-ID: <20260828153002.10290-7-tony.luck@intel.com> (raw)
In-Reply-To: <20260828153002.10290-1-tony.luck@intel.com>
Each L2 cache instance has its own bitfix filter, but always reports
errors in machine check bank 3 (on Diamond Rapids).
Compute a unique bitfix filter instance number based on the CPU that
logged the error and the machine check bank number.
Special-case the banks associated with the Integrated Memory Hub (IMH).
Here the "even" numbered CPU modules are associated with IMH0 and the
"odd" modules with IMH1.
The unique id will be used to store a time stamp of when the bitfix
filter overflowed so that frequent overflows can be logged.
Co-developed-by: Qiuxu Zhuo <qiuxu.zhuo@intel.com>
Signed-off-by: Qiuxu Zhuo <qiuxu.zhuo@intel.com>
Signed-off-by: Tony Luck <tony.luck@intel.com>
---
drivers/edac/intel-bff.c | 80 ++++++++++++++++++++++++++++++++++++++++
1 file changed, 80 insertions(+)
diff --git a/drivers/edac/intel-bff.c b/drivers/edac/intel-bff.c
index c46d5255f34b..06a79745c2fb 100644
--- a/drivers/edac/intel-bff.c
+++ b/drivers/edac/intel-bff.c
@@ -19,13 +19,18 @@
#include <linux/bitfield.h>
#include <linux/bits.h>
+#include <linux/cacheinfo.h>
+#include <linux/cleanup.h>
#include <linux/cpufeature.h>
+#include <linux/cpuhplock.h>
#include <linux/device-id/x86_cpu.h>
#include <linux/errno.h>
#include <linux/init.h>
+#include <linux/limits.h>
#include <linux/module.h>
#include <linux/notifier.h>
#include <linux/printk.h>
+#include <linux/topology.h>
#include <linux/types.h>
#include <asm/cpu_device_id.h>
@@ -72,12 +77,87 @@ MODULE_DEVICE_TABLE(x86cpu, bff_cpu_ids);
static const enum bff_bank_type *bff_bank_types;
+/* Diamond Rapids maps APICID[2] to the IMH instance within a socket. */
+#define APICID_IMH_NUM GENMASK(2, 2)
+#define IMH_NUM(apicid) FIELD_GET(APICID_IMH_NUM, apicid)
+#define NUM_IMH_PER_SOCKET 2
+
+static void bff_set_imh_id(struct mce *mce, unsigned long *id)
+{
+ int imh_num;
+
+ imh_num = NUM_IMH_PER_SOCKET * topology_physical_package_id(mce->extcpu) +
+ IMH_NUM(mce->apicid);
+
+ *id |= imh_num;
+}
+
+static bool bff_set_cache_id(int cpu, int level, unsigned long *id)
+{
+ int cacheid;
+
+ guard(cpus_read_lock)();
+
+ cacheid = get_cpu_cacheinfo_id(cpu, level);
+ if (cacheid == -1) {
+ pr_warn("Could not get L%d cache id for CPU %d\n", level, cpu);
+ return false;
+ }
+
+ *id |= cacheid;
+
+ return true;
+}
+
+/*
+ * Cache IDs are only unique within a cache level.
+ * Include the MCA bank number so each BFF-capable hardware
+ * resource has a unique tracking ID.
+ */
+#define BFF_ID_BANK_FIELD GENMASK(63, 32)
+
+static unsigned long bff_get_id(struct mce *mce)
+{
+ unsigned long id = FIELD_PREP(BFF_ID_BANK_FIELD, mce->bank);
+
+ switch (bff_bank_types[mce->bank]) {
+ case BFF_BANK_DCU:
+ case BFF_BANK_DTLB:
+ if (!bff_set_cache_id(mce->extcpu, 1, &id))
+ return ULONG_MAX;
+ break;
+
+ case BFF_BANK_MLC:
+ if (!bff_set_cache_id(mce->extcpu, 2, &id))
+ return ULONG_MAX;
+ break;
+
+ case BFF_BANK_CCF:
+ if (!bff_set_cache_id(mce->extcpu, 3, &id))
+ return ULONG_MAX;
+ break;
+
+ case BFF_BANK_HSF:
+ case BFF_BANK_IOCACHE:
+ bff_set_imh_id(mce, &id);
+ break;
+
+ default:
+ return ULONG_MAX;
+ }
+
+ return id;
+}
+
static void bff_reset_and_report(struct mce *mce)
{
/* Reset bitfix filter using the CPU that logged the yellow status */
if (wrmsrq_on_cpu(mce->extcpu, MSR_MCx_BFF_CTL(mce->bank), MCI_BFF_RESET))
pr_warn("Failed to reset bitfix filter for CPU %d Bank %d\n",
mce->extcpu, mce->bank);
+
+ /* Placeholder use of bff_get_id() */
+ pr_debug("unique_id = 0x%lx\n", bff_get_id(mce));
}
static int bff_mce_notify(struct notifier_block *nb, unsigned long val, void *data)
--
2.55.0
next prev parent reply other threads:[~2026-08-28 15:30 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 15:29 [PATCH v2 0/7] EDAC/intel-bff: Driver to reset bitfix filters Tony Luck
2026-08-28 15:29 ` [PATCH v2 1/7] cacheinfo: Export get_cpu_cacheinfo_id() for loadable modules Tony Luck
2026-08-28 15:29 ` [PATCH v2 2/7] x86/mce: Add enumeration for Intel bitfix filter reset Tony Luck
2026-08-28 15:29 ` [PATCH v2 3/7] EDAC/intel-bff: Add stub Intel bitfix filter driver Tony Luck
2026-08-28 15:29 ` [PATCH v2 4/7] EDAC/intel-bff: Add Diamond Rapids support Tony Luck
2026-08-28 15:30 ` [PATCH v2 5/7] EDAC/intel-bff: Reset bitfix filter when it overflows Tony Luck
2026-08-28 15:30 ` Tony Luck [this message]
2026-08-28 15:30 ` [PATCH v2 7/7] EDAC/intel-bff: Report frequent filter overflows Tony Luck
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260828153002.10290-7-tony.luck@intel.com \
--to=tony.luck@intel.com \
--cc=bp@alien8.de \
--cc=ilpo.jarvinen@linux.intel.com \
--cc=leitao@debian.org \
--cc=linux-edac@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=patches@lists.linux.dev \
--cc=qiuxu.zhuo@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®