* [PATCH v2 0/2] perf/arm-cmn: Allow userspace to select the PMU's CPU
@ 2026-09-30 23:11 Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 1/2] perf/arm-cmn: Don't schedule events on a CPU which no longer owns the PMU Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 2/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
0 siblings, 2 replies; 3+ messages in thread
From: Haris Okanovic @ 2026-09-30 23:11 UTC (permalink / raw)
To: robin.murphy, mark.rutland, will
Cc: linux-arm-kernel, linux-perf-users, linux-kernel, harisokn
arm_cmn_probe() picks the CPU that owns the PMU and nothing but the CPU
hotplug callbacks ever revisits it, so in practice all of the PMU's
recurring work stays on the first CPU local to the interconnect's NUMA
node. Patch 2 makes the existing 'cpumask' attribute writable so it can
be moved, which helps on systems that confine background kernel work to
a chosen set of housekeeping CPUs.
I made the existing 'cpumask' attribute writable rather than
adding a new one. Every other implementation of that file is read-only,
so I am happy to switch to a separate attribute if you would prefer to
keep 'cpumask' uniformly read-only across PMUs.
Patch 1 is new in v2. Robin pointed out that perf_event_open() can latch
cmn->cpu and then install the event after a migration has already moved
the PMU, leaving it on a CPU which no longer owns the PMU's shared state.
The window already exists via arm_cmn_pmu_online_cpu(), so patch 1 stands
on its own; a writable cpumask just makes it reachable on demand.
Changes since v1:
- Added patch 1 to close the perf_event_open() race.
- No change to patch 2 (the v1 patch) other than being rebased.
v1: https://lore.kernel.org/linux-arm-kernel/20260929223244.2411400-1-harisokn@amazon.com/
Tested on two platforms with CONFIG_PROVE_LOCKING=y:
AWS m9g.metal-48xl CMN S3, one mesh, 192 CPUs
AWS m8g.metal-48xl CMN-700, two meshes, 192 CPUs
- Multiplexing occurs on the configured CPU.
- Offlining the owning CPU migrates the PMU and updates the attribute.
- Writes racing CPU offline/online produced no lockdep reports.
- Hammer perf_event_open() while flipping the cpumask between two CPUs.
Haris Okanovic (2):
perf/arm-cmn: Don't schedule events on a CPU which no longer owns the
PMU
perf/arm-cmn: Allow userspace to select the PMU's CPU
Documentation/admin-guide/perf/arm-cmn.rst | 18 ++++++++
drivers/perf/arm-cmn.c | 49 ++++++++++++++++++++--
2 files changed, 63 insertions(+), 4 deletions(-)
--
Haris Okanovic
AWS Graviton
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 1/2] perf/arm-cmn: Don't schedule events on a CPU which no longer owns the PMU
2026-09-30 23:11 [PATCH v2 0/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
@ 2026-09-30 23:11 ` Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 2/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
1 sibling, 0 replies; 3+ messages in thread
From: Haris Okanovic @ 2026-09-30 23:11 UTC (permalink / raw)
To: robin.murphy, mark.rutland, will
Cc: linux-arm-kernel, linux-perf-users, linux-kernel, harisokn
perf_event_open() can latch cmn->cpu and then install the event after a
migration has already moved the PMU. The driver has no locking -- it
relies on all events living in one CPU's context -- so such an event
races the owning CPU and can corrupt the shared DTC and DTM state,
giving wrong counts. arm_cmn_event_add() now rejects an event which is
not on the owning CPU.
arm_cmn_migrate() now claims cmn->cpu before migrating, so that events
it reinstalls via perf_pmu_migrate_context() still pass that check.
Signed-off-by: Haris Okanovic <harisokn@amazon.com>
---
drivers/perf/arm-cmn.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/drivers/perf/arm-cmn.c b/drivers/perf/arm-cmn.c
index 5378fba916cf5..a5593e5821925 100644
--- a/drivers/perf/arm-cmn.c
+++ b/drivers/perf/arm-cmn.c
@@ -2112,6 +2112,9 @@ static int arm_cmn_event_add(struct perf_event *event, int flags)
enum cmn_node_type type = CMN_EVENT_TYPE(event);
unsigned int input_sel, i = 0;
+ if (cmn->cpu != smp_processor_id())
+ return -ENOENT;
+
if (type == CMN_TYPE_DTC) {
while (cmn->dtc[i].cycles)
if (++i == cmn->num_dtcs)
@@ -2245,12 +2248,12 @@ static int arm_cmn_commit_txn(struct pmu *pmu)
static void arm_cmn_migrate(struct arm_cmn *cmn, unsigned int cpu)
{
- unsigned int i;
+ unsigned int i, old = cmn->cpu;
- perf_pmu_migrate_context(&cmn->pmu, cmn->cpu, cpu);
+ cmn->cpu = cpu;
+ perf_pmu_migrate_context(&cmn->pmu, old, cpu);
for (i = 0; i < cmn->num_dtcs; i++)
irq_set_affinity(cmn->dtc[i].irq, cpumask_of(cpu));
- cmn->cpu = cpu;
}
static int arm_cmn_pmu_online_cpu(unsigned int cpu, struct hlist_node *cpuhp_node)
--
Haris Okanovic
AWS Graviton
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 2/2] perf/arm-cmn: Allow userspace to select the PMU's CPU
2026-09-30 23:11 [PATCH v2 0/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 1/2] perf/arm-cmn: Don't schedule events on a CPU which no longer owns the PMU Haris Okanovic
@ 2026-09-30 23:11 ` Haris Okanovic
1 sibling, 0 replies; 3+ messages in thread
From: Haris Okanovic @ 2026-09-30 23:11 UTC (permalink / raw)
To: robin.murphy, mark.rutland, will
Cc: linux-arm-kernel, linux-perf-users, linux-kernel, harisokn
arm_cmn_probe() picks the CPU that owns the PMU with
cmn->cpu = cpumask_local_spread(0, dev_to_node(cmn->dev));
and arm_cmn_event_init() then binds every event to it unconditionally.
That choice is only revisited by the CPU hotplug callbacks, so in
practice the PMU stays on the first CPU local to the interconnect's NUMA
node, which is usually CPU 0.
All of the PMU's recurring work therefore lands on one CPU. This is
problematic on systems which reserve particular CPUs for latency
sensitive work or confine background activity to a chosen set of
housekeeping CPUs.
There is no way to set it explicitly. perf_event_open() and 'perf stat
-C' have no effect because arm_cmn_event_init() overwrites event->cpu;
/proc/irq/*/smp_affinity is refused for the DTC interrupts, which are
requested with IRQF_NOBALANCING because their affinity has to follow the
owning CPU.
Make the existing 'cpumask' attribute writable. It already reports the
CPU which owns the PMU; writing a CPU number now migrates the PMU there
via the existing arm_cmn_migrate(), which moves the perf contexts and
the DTC interrupt affinity together. Since the PMU has a single owning
CPU, anything other than one CPU number is rejected.
The attribute expresses a preference rather than a guarantee. The
hotplug callbacks may still move the PMU, for example when the chosen
CPU is offlined.
Signed-off-by: Haris Okanovic <harisokn@amazon.com>
---
Documentation/admin-guide/perf/arm-cmn.rst | 18 ++++++++++
drivers/perf/arm-cmn.c | 40 +++++++++++++++++++++-
2 files changed, 57 insertions(+), 1 deletion(-)
diff --git a/Documentation/admin-guide/perf/arm-cmn.rst b/Documentation/admin-guide/perf/arm-cmn.rst
index 796e25b7027b2..bf75b686cef76 100644
--- a/Documentation/admin-guide/perf/arm-cmn.rst
+++ b/Documentation/admin-guide/perf/arm-cmn.rst
@@ -44,6 +44,24 @@ given type. To target a specific node, "bynodeid" must be set to 1 and
"nodeid" to the appropriate value derived from the CMN configuration
(as defined in the "Node ID Mapping" section of the TRM).
+CPU affinity
+------------
+
+The driver also provides a "cpumask" sysfs attribute, which contains a
+single CPU ID, of the processor which will be used to handle all the CMN
+PMU events.
+
+The attribute is writable, and accepts a single CPU ID to move the PMU
+to that processor, for instance to keep counter reads, interrupt handling
+and event rotation away from CPUs reserved for latency sensitive work::
+
+ $# echo 5 > /sys/bus/event_source/devices/arm_cmn_0/cpumask
+
+This expresses a preference rather than a guarantee. In case of the
+chosen processor being offlined, or a processor local to the
+interconnect's NUMA node coming online while the chosen one is not, the
+events and interrupts are migrated and the attribute is updated.
+
Watchpoints
-----------
diff --git a/drivers/perf/arm-cmn.c b/drivers/perf/arm-cmn.c
index a5593e5821925..1f114ac3f0700 100644
--- a/drivers/perf/arm-cmn.c
+++ b/drivers/perf/arm-cmn.c
@@ -5,6 +5,7 @@
#include <linux/acpi.h>
#include <linux/bitfield.h>
#include <linux/bitops.h>
+#include <linux/cpu.h>
#include <linux/debugfs.h>
#include <linux/interrupt.h>
#include <linux/io.h>
@@ -12,6 +13,7 @@
#include <linux/kernel.h>
#include <linux/list.h>
#include <linux/module.h>
+#include <linux/mutex.h>
#include <linux/of.h>
#include <linux/perf_event.h>
#include <linux/platform_device.h>
@@ -404,6 +406,8 @@ struct arm_cmn_nodeid {
u8 dev;
};
+static void arm_cmn_migrate(struct arm_cmn *cmn, unsigned int cpu);
+
static int arm_cmn_xyidbits(const struct arm_cmn *cmn)
{
return fls((cmn->mesh_x - 1) | (cmn->mesh_y - 1));
@@ -1519,8 +1523,42 @@ static ssize_t arm_cmn_cpumask_show(struct device *dev,
return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpumask_of(cmn->cpu)));
}
+static ssize_t arm_cmn_cpumask_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
+{
+ static DEFINE_MUTEX(cpumask_mutex);
+
+ struct arm_cmn *cmn = to_cmn(dev_get_drvdata(dev));
+ unsigned int cpu;
+ int err;
+
+ err = kstrtouint(buf, 0, &cpu);
+ if (err)
+ return err;
+
+ if (cpu >= nr_cpu_ids)
+ return -EINVAL;
+
+ /* Serialises multiple writers against each other */
+ mutex_lock(&cpumask_mutex);
+ /* Blocks hotplug during write */
+ cpus_read_lock();
+
+ if (!cpu_online(cpu))
+ err = -EINVAL;
+ else if (cpu != cmn->cpu)
+ arm_cmn_migrate(cmn, cpu);
+
+ cpus_read_unlock();
+ mutex_unlock(&cpumask_mutex);
+
+ return err ?: count;
+}
+
static struct device_attribute arm_cmn_cpumask_attr =
- __ATTR(cpumask, 0444, arm_cmn_cpumask_show, NULL);
+ __ATTR(cpumask, 0644, arm_cmn_cpumask_show,
+ arm_cmn_cpumask_store);
static ssize_t arm_cmn_identifier_show(struct device *dev,
struct device_attribute *attr, char *buf)
--
Haris Okanovic
AWS Graviton
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-30 23:11 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 23:11 [PATCH v2 0/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 1/2] perf/arm-cmn: Don't schedule events on a CPU which no longer owns the PMU Haris Okanovic
2026-09-30 23:11 ` [PATCH v2 2/2] perf/arm-cmn: Allow userspace to select the PMU's CPU Haris Okanovic
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®