From: Arnaldo Carvalho de Melo <acme@kernel.org>
To: Namhyung Kim <namhyung@kernel.org>
Cc: Ingo Molnar <mingo@kernel.org>,
Thomas Gleixner <tglx@linutronix.de>,
James Clark <james.clark@linaro.org>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
Clark Williams <williams@redhat.com>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
Arnaldo Carvalho de Melo <acme@redhat.com>
Subject: [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available
Date: Wed, 16 Sep 2026 08:47:39 -0300 [thread overview]
Message-ID: <20260916114740.48230-13-acme@kernel.org> (raw)
In-Reply-To: <20260916114740.48230-1-acme@kernel.org>
From: Arnaldo Carvalho de Melo <acme@redhat.com>
IBS events with exclude_{user,kernel} bits, as used in per-thread mode
when kernel samples are not allowed, are rejected by the kernel on
hardware without the privilege filter, so a per-thread 'perf mem
record' fails:
$ perf mem record -o /dev/null -- true
Error:
Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
Kernel v6.14 added swfilt, a software privilege filter that makes those
events usable per-thread, exposed as the 'swfilt' format term
(d29e744c71673a71 "perf/x86: Relax privilege filter restriction on AMD
IBS"):
$ perf record -e ibs_op/ldlat=0,swfilt=1/ -- true
Give the ibs_op memory events a name variant with the swfilt term, used
when the PMU exposes it as a format term, keeping the names that need
system wide mode otherwise.
The variant is used even for records that end up without exclude bits,
e.g. a plain system wide one: perf record adds those bits itself when
the first open fails with EACCES on an unprivileged setup, after the
event name has been built, so the term has to be in it for that retry:
$ perf record -v -e ibs_op/ldlat=0/ -- true
kernel.perf_event_paranoid=2, trying to fall back to excluding kernel and hypervisor samples
Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
With no exclude bits the kernel discards nothing.
This makes the per-thread 'perf mem record' that the data type profiling
shell test does work: the skip added by patch 1 is driven by that record
failing, so it now runs on AMD kernels with swfilt, and keeps skipping on
kernels without it. The 'Test data symbol' shell test matches the exact
ibs_op event string, so its regex now accepts terms added after
ldlat=150, and it only checks ldlat when the captured term has it: on
uarch without ldlat the swfilt variant makes the capture non-empty
("swfilt=1"), that the old 'ldlat=0' check would reject.
Suggested-by: Namhyung Kim <namhyung@kernel.org>
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/arch/x86/util/mem-events.c | 17 ++++++++++++++---
tools/perf/tests/shell/test_data_symbol.sh | 6 ++++--
tools/perf/util/mem-events.c | 21 +++++++++++++++++----
tools/perf/util/mem-events.h | 2 ++
4 files changed, 37 insertions(+), 9 deletions(-)
diff --git a/tools/perf/arch/x86/util/mem-events.c b/tools/perf/arch/x86/util/mem-events.c
index b38f519020ff8c6f..034053dc762101dc 100644
--- a/tools/perf/arch/x86/util/mem-events.c
+++ b/tools/perf/arch/x86/util/mem-events.c
@@ -7,7 +7,10 @@
#define MEM_LOADS_AUX 0x8203
-#define E(t, n, s, l, a) { .tag = t, .name = n, .event_name = s, .ldlat = l, .aux_event = a }
+#define E_INIT(t, n, s, l, a, sf) { \
+ .tag = t, .name = n, .event_name = s, .swfilt_name = sf, \
+ .ldlat = l, .aux_event = a }
+#define E(t, n, s, l, a) E_INIT(t, n, s, l, a, NULL)
struct perf_mem_event perf_mem_events_intel[PERF_MEM_EVENTS__MAX] = {
E("ldlat-loads", "%s/mem-loads,ldlat=%u/P", "mem-loads", true, 0),
@@ -21,14 +24,22 @@ struct perf_mem_event perf_mem_events_intel_aux[PERF_MEM_EVENTS__MAX] = {
E(NULL, NULL, NULL, false, 0),
};
+/*
+ * IBS events with exclude_{user,kernel} bits set, as used by perf to
+ * record per-thread when kernel samples are not allowed, are rejected
+ * by the kernel on hardware without the privilege filter unless the
+ * swfilt software filter is used, so these events carry a variant of
+ * their names with the swfilt term, used by perf_pmu__mem_events_name()
+ * when the kernel exposes the term.
+ */
struct perf_mem_event perf_mem_events_amd[PERF_MEM_EVENTS__MAX] = {
E(NULL, NULL, NULL, false, 0),
E(NULL, NULL, NULL, false, 0),
- E("mem-ldst", "%s//", NULL, false, 0),
+ E_INIT("mem-ldst", "%s//", NULL, false, 0, "%s/swfilt=1/"),
};
struct perf_mem_event perf_mem_events_amd_ldlat[PERF_MEM_EVENTS__MAX] = {
E(NULL, NULL, NULL, false, 0),
E(NULL, NULL, NULL, false, 0),
- E("mem-ldst", "%s/ldlat=%u/", NULL, true, 0),
+ E_INIT("mem-ldst", "%s/ldlat=%u/", NULL, true, 0, "%s/ldlat=%u,swfilt=1/"),
};
diff --git a/tools/perf/tests/shell/test_data_symbol.sh b/tools/perf/tests/shell/test_data_symbol.sh
index d61b5659a46d9a77..52c837fddb639595 100755
--- a/tools/perf/tests/shell/test_data_symbol.sh
+++ b/tools/perf/tests/shell/test_data_symbol.sh
@@ -65,15 +65,17 @@ if (($is_amd >= 1)); then
# --ldlat on AMD:
# o Zen4 and earlier uarch does not support ldlat
# o Even on supported platforms, it's disabled (--ldlat=0) by default.
+ # o Kernels with the swfilt term add it even when ldlat is not
+ # supported, so only check ldlat when the term is present.
ldlat=${BASH_REMATCH[1]}
- if [[ -n $ldlat ]]; then
+ if [[ $ldlat == *ldlat=* ]]; then
if ! [[ "$ldlat" =~ ldlat=0 ]]; then
echo "ERROR: ldlat not initialized to 0?"
exit 1
fi
mem_events="$(perf mem record -v --ldlat=150 -e list 2>&1)"
- if ! [[ "$mem_events" =~ ^mem-ldst.*ibs_op/ldlat=150/.*available ]]; then
+ if ! [[ "$mem_events" =~ ^mem-ldst.*ibs_op/ldlat=150[,/].*available ]]; then
echo "ERROR: --ldlat not honored?"
exit 1
fi
diff --git a/tools/perf/util/mem-events.c b/tools/perf/util/mem-events.c
index 0b49fce251fcc184..8f74cd085e500231 100644
--- a/tools/perf/util/mem-events.c
+++ b/tools/perf/util/mem-events.c
@@ -82,6 +82,7 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
char *buf, size_t buf_size)
{
struct perf_mem_event *e;
+ const char *name;
if (i >= PERF_MEM_EVENTS__MAX || !pmu)
return NULL;
@@ -90,24 +91,36 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
if (!e || !e->name)
return NULL;
+ /*
+ * Use the swfilt variant of the name when the PMU exposes the term.
+ * It is not conditional on the event already having exclude bits:
+ * perf record adds those bits itself when the first open fails with
+ * EACCES on an unprivileged setup, after this name has been built,
+ * and that retry only succeeds with the term in the name. With no
+ * exclude bits the kernel doesn't discard anything.
+ */
+ name = e->name;
+ if (e->swfilt_name && perf_pmu__has_format(pmu, "swfilt"))
+ name = e->swfilt_name;
+
if (i == PERF_MEM_EVENTS__LOAD || i == PERF_MEM_EVENTS__LOAD_STORE) {
if (e->ldlat) {
if (!e->aux_event) {
/* ARM and Most of Intel */
scnprintf(buf, buf_size,
- e->name, pmu->name,
+ name, pmu->name,
perf_mem_events__loads_ldlat);
} else {
/* Intel with mem-loads-aux event */
scnprintf(buf, buf_size,
- e->name, pmu->name, pmu->name,
+ name, pmu->name, pmu->name,
perf_mem_events__loads_ldlat);
}
} else {
if (!e->aux_event) {
/* AMD and POWER */
scnprintf(buf, buf_size,
- e->name, pmu->name);
+ name, pmu->name);
} else {
return NULL;
}
@@ -117,7 +130,7 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
if (i == PERF_MEM_EVENTS__STORE) {
scnprintf(buf, buf_size,
- e->name, pmu->name);
+ name, pmu->name);
return buf;
}
diff --git a/tools/perf/util/mem-events.h b/tools/perf/util/mem-events.h
index 5b98076904b0b689..41f628fad10709b9 100644
--- a/tools/perf/util/mem-events.h
+++ b/tools/perf/util/mem-events.h
@@ -11,6 +11,8 @@ struct perf_mem_event {
u32 aux_event;
const char *tag;
const char *name;
+ /* Name with the swfilt software privilege filter, when supported. */
+ const char *swfilt_name;
const char *event_name;
};
--
2.55.0
next prev parent reply other threads:[~2026-09-16 11:48 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 11:47 [PATCH v6 0/12] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 01/12] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 02/12] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
2026-09-16 17:59 ` Ian Rogers
2026-09-16 19:02 ` Arnaldo Carvalho de Melo
2026-09-16 21:28 ` Ian Rogers
2026-09-16 18:42 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 03/12] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 04/12] perf debuginfo: Let the user skip and disable debuginfod fetches Arnaldo Carvalho de Melo
2026-09-16 18:53 ` Namhyung Kim
2026-09-16 21:27 ` Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 05/12] perf debuginfo: Show the debuginfod fetch progress and keys in the TUI Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 06/12] perf symbol: Fall back to fetching the vmlinux by build ID Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 07/12] perf annotate-data: Show the sample count in the data-type browser Arnaldo Carvalho de Melo
2026-09-16 21:28 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 08/12] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 09/12] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 10/12] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
2026-09-16 21:44 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 11/12] perf mem record: Request PERF_SAMPLE_CPU by default Arnaldo Carvalho de Melo
2026-09-16 21:50 ` Namhyung Kim
2026-09-16 11:47 ` Arnaldo Carvalho de Melo [this message]
2026-09-16 21:59 ` [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available Namhyung Kim
2026-09-16 22:27 ` [PATCH v6 0/12] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Namhyung Kim
2026-09-16 18:32 Arnaldo Carvalho de Melo
2026-09-16 18:32 ` [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available Arnaldo Carvalho de Melo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260916114740.48230-13-acme@kernel.org \
--to=acme@kernel.org \
--cc=acme@redhat.com \
--cc=adrian.hunter@intel.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=namhyung@kernel.org \
--cc=tglx@linutronix.de \
--cc=williams@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®