* [PATCH v3 1/3] x86/resctrl: Fix ABMC counter programming
2026-10-02 21:26 [PATCH v3 0/3] x86,fs/resctrl: Keep default MBM mode at boot and fix ABMC Babu Moger
@ 2026-10-02 21:26 ` Babu Moger
2026-10-02 21:26 ` [PATCH v3 2/3] fs/resctrl: Assign counters to existing groups when enabling mbm_event Babu Moger
2026-10-02 21:26 ` [PATCH v3 3/3] x86,fs/resctrl: Keep mbm_assign_mode in default mode at boot Babu Moger
2 siblings, 0 replies; 4+ messages in thread
From: Babu Moger @ 2026-10-02 21:26 UTC (permalink / raw)
To: tony.luck, reinette.chatre, bp
Cc: x86, Dave.Martin, james.morse, babu.moger, corbet, skhan,
rdunlap, tglx, mingo, dave.hansen, hpa, kas, rick.p.edgecombe,
linux-kernel, linux-doc, linux-coco, kvm
AMD's Assignable Bandwidth Monitoring Counters (ABMC) are configured via
MSR_IA32_L3_QOS_ABMC_CFG. The architecture [1] received an update that
expands the counter ID field (l3_qos_abmc_cfg.split.cntr_id) from 5 to 12
bits.
Use the updated field width. The number of supported counters is enumerated
separately. Mark this as a fix to the original ABMC support to avoid
misconfigurations caused by truncating counter IDs on hardware that
supports a large number of counters.
Also limit the number of supported counters to the maximum value that can
be represented by the 12-bit cntr_id field if hardware reports more than
12 bits.
The AMD64 Architecture Programmer's Manual [1], available at [2], will be
updated in a future revision to document the expanded cntr_id field.
[1] AMD64 Architecture Programmer's Manual Volume 2: System Programming,
Publication #24593, Revision 3.41, Section 19.3.3.3 "Assignable
Bandwidth Monitoring (ABMC)"
Fixes: 84ecefb76674 ("x86/resctrl: Add data structures and definitions for ABMC assignment")
Signed-off-by: Babu Moger <babu.moger@amd.com>
Cc: stable@vger.kernel.org
Link: https://bugzilla.kernel.org/show_bug.cgi?id=206537 # [2]
---
v3: Dropped the fix for truncation on 32-bit x86.
Removed the change bw_src field(RMID) width to 15 bits.
Added new check to limit the number of counters to 12 bits.
v2: Moved the link tag to the last.
v1: https://lore.kernel.org/lkml/980f39d3a0e0d9f73925e362f835aeef070a1bc5.1784322818.git.babu.moger@amd.com/
---
arch/x86/kernel/cpu/resctrl/internal.h | 4 ++--
arch/x86/kernel/cpu/resctrl/monitor.c | 3 ++-
2 files changed, 4 insertions(+), 3 deletions(-)
diff --git a/arch/x86/kernel/cpu/resctrl/internal.h b/arch/x86/kernel/cpu/resctrl/internal.h
index e3cfa0c10e92..ffd74a68671b 100644
--- a/arch/x86/kernel/cpu/resctrl/internal.h
+++ b/arch/x86/kernel/cpu/resctrl/internal.h
@@ -214,8 +214,8 @@ union l3_qos_abmc_cfg {
bw_src :12,
reserved1: 3,
is_clos : 1,
- cntr_id : 5,
- reserved : 9,
+ cntr_id :12,
+ reserved : 2,
cntr_en : 1,
cfg_en : 1;
} split;
diff --git a/arch/x86/kernel/cpu/resctrl/monitor.c b/arch/x86/kernel/cpu/resctrl/monitor.c
index 3838e0a13d36..5d0d3b18f9b8 100644
--- a/arch/x86/kernel/cpu/resctrl/monitor.c
+++ b/arch/x86/kernel/cpu/resctrl/monitor.c
@@ -470,7 +470,8 @@ int __init rdt_get_l3_mon_config(struct rdt_resource *r)
r->mon.mbm_cntr_assignable = true;
r->mon.mbm_cntr_configurable = true;
cpuid_count(0x80000020, 5, &eax, &ebx, &ecx, &edx);
- r->mon.num_mbm_cntrs = (ebx & GENMASK(15, 0)) + 1;
+ /* cntr_id is 12 bits and can only encode 4096 counters. */
+ r->mon.num_mbm_cntrs = min((ebx & GENMASK(15, 0)) + 1, BIT(12));
hw_res->mbm_cntr_assign_enabled = true;
}
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v3 2/3] fs/resctrl: Assign counters to existing groups when enabling mbm_event
2026-10-02 21:26 [PATCH v3 0/3] x86,fs/resctrl: Keep default MBM mode at boot and fix ABMC Babu Moger
2026-10-02 21:26 ` [PATCH v3 1/3] x86/resctrl: Fix ABMC counter programming Babu Moger
@ 2026-10-02 21:26 ` Babu Moger
2026-10-02 21:26 ` [PATCH v3 3/3] x86,fs/resctrl: Keep mbm_assign_mode in default mode at boot Babu Moger
2 siblings, 0 replies; 4+ messages in thread
From: Babu Moger @ 2026-10-02 21:26 UTC (permalink / raw)
To: tony.luck, reinette.chatre, bp
Cc: x86, Dave.Martin, james.morse, babu.moger, corbet, skhan,
rdunlap, tglx, mingo, dave.hansen, hpa, kas, rick.p.edgecombe,
linux-kernel, linux-doc, linux-coco, kvm
When the user enables counter assignment mode by writing "mbm_event"
to /sys/fs/resctrl/info/L3_MON/mbm_assign_mode, resctrl resets all
monitoring state and sets mbm_assign_on_mkdir for subsequent mkdir, but
does not assign counters to groups that already exist, including the
default group created at mount. The counters of those groups return
"Unassigned" until the user assigns counters by hand.
Enable mbm_assign_on_mkdir and assign counters, while there are some
available, to existing CTRL_MON and MON groups so the switch matches
mkdir auto-assignment. An event left without a counter reads
"Unassigned".
Fixes: 8004ea01cf63 ("fs/resctrl: Introduce the interface to switch between monitor modes")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/8cb66e18e32e4087a9712c1e68ee6da614efe244.1784322818.git.babu.moger%40amd.com
Cc: stable@vger.kernel.org
Signed-off-by: Babu Moger <babu.moger@amd.com>
---
v3: Changelog update.
Code comment cleanup.
Added Cc to stable.
Added Reported-by and Closes tags.
v2: New patch.
This patch addresses the Sashiko comment about documentation issue where
counters are not assigned automatically when mode is switched to mbm_event.
https://sashiko.dev/#/patchset/8cb66e18e32e4087a9712c1e68ee6da614efe244.1784322818.git.babu.moger%40amd.com
In fact, it exposed a real issue. When switching to mbm_event mode, existing
monitoring groups should be assigned counters whenever counters are available.
This provides a smooth transition between modes and aligns the behavior with
the existing auto-assignment mechanism.
---
Documentation/filesystems/resctrl.rst | 7 ++++--
fs/resctrl/monitor.c | 35 +++++++++++++++++++++++----
2 files changed, 35 insertions(+), 7 deletions(-)
diff --git a/Documentation/filesystems/resctrl.rst b/Documentation/filesystems/resctrl.rst
index b52795e03303..c8507580474a 100644
--- a/Documentation/filesystems/resctrl.rst
+++ b/Documentation/filesystems/resctrl.rst
@@ -370,8 +370,11 @@ with the following files:
of counters available is described in the "num_mbm_cntrs" file. Changing the
mode may cause all counters on the resource to reset.
- Moving to mbm_event counter assignment mode requires users to assign the counters
- to the events. Otherwise, the MBM event counters will return 'Unassigned' when read.
+ Moving to mbm_event counter assignment mode enables "mbm_assign_on_mkdir" and
+ assigns counters to the events of all existing monitoring groups, including
+ the default group, while counters remain available. Consult
+ "mbm_L3_assignments" after switching to "mbm_event" mode for counter
+ assignment states of all monitoring groups.
The mode is beneficial for AMD platforms that support more CTRL_MON
and MON groups than available hardware counters. By default, this
diff --git a/fs/resctrl/monitor.c b/fs/resctrl/monitor.c
index 73413cb128ea..fca0734bf346 100644
--- a/fs/resctrl/monitor.c
+++ b/fs/resctrl/monitor.c
@@ -1300,8 +1300,7 @@ static int rdtgroup_assign_cntr_event(struct rdt_l3_mon_domain *d, struct rdtgro
}
/*
- * rdtgroup_assign_cntrs() - Assign counters to MBM events. Called when
- * a new group is created.
+ * rdtgroup_assign_cntrs() - Assign counters to MBM events.
*
* Each group can accommodate two counters per domain: one for the total
* event and one for the local event. Assignments may fail due to the limited
@@ -1326,6 +1325,22 @@ void rdtgroup_assign_cntrs(struct rdtgroup *rdtgrp)
&mon_event_all[QOS_L3_MBM_LOCAL_EVENT_ID]);
}
+/*
+ * resctrl_assign_cntrs_allrdtgrp() - Assign counters to the MBM events of
+ * every existing group.
+ */
+static void resctrl_assign_cntrs_allrdtgrp(void)
+{
+ struct rdtgroup *prgrp, *crgrp;
+
+ list_for_each_entry(prgrp, &rdt_all_groups, rdtgroup_list) {
+ rdtgroup_assign_cntrs(prgrp);
+
+ list_for_each_entry(crgrp, &prgrp->mon.crdtgrp_list, mon.crdtgrp_list)
+ rdtgroup_assign_cntrs(crgrp);
+ }
+}
+
/*
* rdtgroup_free_unassign_cntr() - Unassign and reset the counter ID configuration
* for the event pointed to by @mevt within the domain @d and resctrl group @rdtgrp.
@@ -1599,9 +1614,6 @@ ssize_t resctrl_mbm_assign_mode_write(struct kernfs_open_file *of, char *buf,
(READS_TO_LOCAL_MEM |
READS_TO_LOCAL_S_MEM |
NON_TEMP_WRITE_TO_LOCAL_MEM);
- /* Enable auto assignment when switching to "mbm_event" mode */
- if (enable)
- r->mon.mbm_assign_on_mkdir = true;
/*
* Reset all the non-achitectural RMID state and assignable counters.
*/
@@ -1609,6 +1621,19 @@ ssize_t resctrl_mbm_assign_mode_write(struct kernfs_open_file *of, char *buf,
mbm_cntr_free_all(r, d);
resctrl_reset_rmid_all(r, d);
}
+
+ /*
+ * Counters were freed above, so both new groups (via mkdir) and
+ * the groups that already exist need assignments. Groups created
+ * while in "default" mode have no counter assigned, including the
+ * default group created when resctrl is mounted. Assign counters
+ * to them so that enabling the mode leaves the same assignments
+ * that mkdir would have made.
+ */
+ if (enable) {
+ r->mon.mbm_assign_on_mkdir = true;
+ resctrl_assign_cntrs_allrdtgrp();
+ }
}
out_unlock:
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v3 3/3] x86,fs/resctrl: Keep mbm_assign_mode in default mode at boot
2026-10-02 21:26 [PATCH v3 0/3] x86,fs/resctrl: Keep default MBM mode at boot and fix ABMC Babu Moger
2026-10-02 21:26 ` [PATCH v3 1/3] x86/resctrl: Fix ABMC counter programming Babu Moger
2026-10-02 21:26 ` [PATCH v3 2/3] fs/resctrl: Assign counters to existing groups when enabling mbm_event Babu Moger
@ 2026-10-02 21:26 ` Babu Moger
2 siblings, 0 replies; 4+ messages in thread
From: Babu Moger @ 2026-10-02 21:26 UTC (permalink / raw)
To: tony.luck, reinette.chatre, bp
Cc: x86, Dave.Martin, james.morse, babu.moger, corbet, skhan,
rdunlap, tglx, mingo, dave.hansen, hpa, kas, rick.p.edgecombe,
linux-kernel, linux-doc, linux-coco, kvm
The "mbm_event" mode assigns hardware MBM counters to RMID/event pairs.
It is intended for deployments that need to manage counter assignment on
platforms where the number of monitoring groups exceeds the available
hardware counters. Hardware counters are scarce, so mbm_event mode suits
workflows that monitor a subset of groups at a time and rotate
assignments as needed.
resctrl enables the "mbm_event" mode by default on hardware that
supports it. That breaks the pqos tool [1], which assumes the historical
default mode. pqos treats a non-numeric event read as zero bandwidth. It
creates 16 or more monitoring groups and uses two counters per group.
On a platform with 32 counters per domain, that consumes the pool, so
further groups read "Unassigned" and pqos reports 0 MB/s for them:
$sudo pqos -m all:0-4
CORE IPC MISSES LLC[KB] MBL[MB/s] MBR[MB/s]
0 0.69 16k 192.0 0.0 0.0
1 0.40 3k 0.0 0.0 0.0
2 0.44 1k 64.0 0.0 0.0
3 0.39 1k 64.0 0.0 0.0
4 0.43 1k 0.0 0.0 0.0
Leave mbm_assign_mode in "default" mode during initialization. Default
mode uses one counter per monitoring group. On existing AMD platforms
that pool is 64 counters and may be larger on newer hardware. The
"mbm_event" mode uses two counters per group, so 32 counters cover only
16 groups.
Deployments within that pool receive accurate bandwidth measurements
with default mode. The hardware supports 4096 monitoring groups. Beyond
that pool, readings may be misleading or "Unavailable", with no
user-visible indication.
Default mode is still limited once the number of monitoring groups
exceeds that counter pool. After hardware re-allocates a counter, a
read may return "Unavailable". pqos still treats "Unavailable" as zero.
Successive reads can be a count, "Unavailable", then another count.
pqos treats those jumps as wraparound and can report an inconsistent
rate:
$sudo pqos -m all:0-4
CORE IPC MISSES LLC[KB] MBL[MB/s] MBR[MB/s]
0 0.76 60k 32.0 0.0 0.0
1 0.45 1k 32.0 0.0 0.0
2 1.57 107k 64.0 0.0 17592186044184.9
3 1.62 276k 4928.0 3.2 0.5
Users that need stable measurements beyond that pool should use
"mbm_event" mode and rotate assignments as needed.
Enable it with:
$echo mbm_event > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
Fixes: 0f1576e43adc ("x86/resctrl: Configure mbm_event mode if supported")
Closes: https://github.com/intel/intel-cmt-cat/issues/311
Link: https://github.com/intel/intel-cmt-cat # [1]
Cc: stable@vger.kernel.org
Signed-off-by: Babu Moger <babu.moger@amd.com>
---
Documentation/filesystems/resctrl.rst | 48 ++++++++++++++++++---------
arch/x86/kernel/cpu/resctrl/monitor.c | 1 -
2 files changed, 33 insertions(+), 16 deletions(-)
diff --git a/Documentation/filesystems/resctrl.rst b/Documentation/filesystems/resctrl.rst
index c8507580474a..2efc99223bdc 100644
--- a/Documentation/filesystems/resctrl.rst
+++ b/Documentation/filesystems/resctrl.rst
@@ -354,8 +354,8 @@ with the following files:
::
# cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
- [mbm_event]
- default
+ [default]
+ mbm_event
"mbm_event":
@@ -363,6 +363,9 @@ with the following files:
pair and monitor the bandwidth usage as long as it is assigned. The hardware
continues to track the assigned counter until it is explicitly unassigned by
the user. Each event within a resctrl group can be assigned independently.
+ mbm_event mode uses one counter per event of one RMID. A monitoring group
+ can use two counters, one for mbm_total_bytes and one for mbm_local_bytes.
+ For example, 32 counters can monitor 16 groups.
In this mode, a monitoring event can only accumulate data while it is backed
by a hardware counter. Use "mbm_L3_assignments" found in each CTRL_MON and MON
@@ -376,20 +379,37 @@ with the following files:
"mbm_L3_assignments" after switching to "mbm_event" mode for counter
assignment states of all monitoring groups.
- The mode is beneficial for AMD platforms that support more CTRL_MON
- and MON groups than available hardware counters. By default, this
- feature is enabled on AMD platforms with the ABMC (Assignable Bandwidth
- Monitoring Counters) capability, ensuring counters remain assigned even
- when the corresponding RMID is not actively used by any processor.
+ It is intended for deployments that need to manage counter assignment on
+ platforms where the number of monitoring groups exceeds the available
+ hardware counters. Hardware counters are scarce, so mbm_event mode suits
+ workflows that monitor a subset of groups at a time and rotate assignments
+ as needed. The mode is beneficial for AMD platforms that support more
+ CTRL_MON and MON groups than available hardware counters. The mbm_event
+ mode ensures counters remain assigned even when the corresponding RMID is
+ not actively monitored.
"default":
In default mode, resctrl assumes there is a hardware counter for each
- event within every CTRL_MON and MON group. On AMD platforms, it is
- recommended to use the mbm_event mode, if supported, to prevent reset of MBM
- events between reads resulting from hardware re-allocating counters. This can
- result in misleading values or display "Unavailable" if no counter is assigned
- to the event.
+ event within every CTRL_MON and MON group. This mode is enabled by default.
+
+ On AMD platforms that support more CTRL_MON and MON groups than hardware
+ counters, hardware dynamically shares a smaller pool of counters among
+ RMIDs. One counter from this pool is used per monitoring group and counts
+ every MBM event of that group's RMID. The size of the pool is not
+ enumerated to software (unlike "num_mbm_cntrs"), and "num_rmids" may be
+ much larger. For example, 64 counters in this pool monitor 64 groups.
+ Newer hardware may provide a larger pool.
+
+ While the number of monitoring groups does not exceed that pool, a counter
+ stays attached to each RMID and readings remain accurate. Creating more
+ groups than the pool can cause hardware to re-allocate those counters
+ between successive reads of an event. Bandwidth values may then be
+ misleading, or a read may return "Unavailable" if no counter is allocated
+ to the RMID. There is no user-visible indication when this begins.
+ Users who need stable readings beyond that pool should switch to mbm_event
+ mode, if supported, and assign counters to the groups of interest
+ (rotating assignments as needed).
* To enable "mbm_event" counter assignment mode:
::
@@ -1787,11 +1807,9 @@ View the llc occupancy snapshot::
Examples on working with mbm_assign_mode
========================================
-a. Check if MBM counter assignment mode is supported.
+a. Check if MBM counter assignment mode is supported and enabled.
::
- # mount -t resctrl resctrl /sys/fs/resctrl/
-
# cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
[mbm_event]
default
diff --git a/arch/x86/kernel/cpu/resctrl/monitor.c b/arch/x86/kernel/cpu/resctrl/monitor.c
index 5d0d3b18f9b8..a6a9090c4808 100644
--- a/arch/x86/kernel/cpu/resctrl/monitor.c
+++ b/arch/x86/kernel/cpu/resctrl/monitor.c
@@ -472,7 +472,6 @@ int __init rdt_get_l3_mon_config(struct rdt_resource *r)
cpuid_count(0x80000020, 5, &eax, &ebx, &ecx, &edx);
/* cntr_id is 12 bits and can only encode 4096 counters. */
r->mon.num_mbm_cntrs = min((ebx & GENMASK(15, 0)) + 1, BIT(12));
- hw_res->mbm_cntr_assign_enabled = true;
}
r->mon_capable = true;
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread