From: Namhyung Kim <namhyung@kernel.org>
To: Arnaldo Carvalho de Melo <acme@kernel.org>
Cc: Ingo Molnar <mingo@kernel.org>,
Thomas Gleixner <tglx@linutronix.de>,
James Clark <james.clark@linaro.org>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
Clark Williams <williams@redhat.com>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
Arnaldo Carvalho de Melo <acme@redhat.com>,
Ravi Bangoria <ravi.bangoria@amd.com>
Subject: Re: [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available
Date: Wed, 16 Sep 2026 14:59:48 -0700 [thread overview]
Message-ID: <aqsRVABj23jaaiSq@google.com> (raw)
In-Reply-To: <20260916114740.48230-13-acme@kernel.org>
+Ravi
On Wed, Sep 16, 2026 at 08:47:39AM -0300, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
>
> IBS events with exclude_{user,kernel} bits, as used in per-thread mode
> when kernel samples are not allowed, are rejected by the kernel on
> hardware without the privilege filter, so a per-thread 'perf mem
> record' fails:
>
> $ perf mem record -o /dev/null -- true
> Error:
> Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
> Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
>
> Kernel v6.14 added swfilt, a software privilege filter that makes those
> events usable per-thread, exposed as the 'swfilt' format term
> (d29e744c71673a71 "perf/x86: Relax privilege filter restriction on AMD
> IBS"):
>
> $ perf record -e ibs_op/ldlat=0,swfilt=1/ -- true
>
> Give the ibs_op memory events a name variant with the swfilt term, used
> when the PMU exposes it as a format term, keeping the names that need
> system wide mode otherwise.
>
> The variant is used even for records that end up without exclude bits,
> e.g. a plain system wide one: perf record adds those bits itself when
> the first open fails with EACCES on an unprivileged setup, after the
> event name has been built, so the term has to be in it for that retry:
>
> $ perf record -v -e ibs_op/ldlat=0/ -- true
> kernel.perf_event_paranoid=2, trying to fall back to excluding kernel and hypervisor samples
> Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
> Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
>
> With no exclude bits the kernel discards nothing.
>
> This makes the per-thread 'perf mem record' that the data type profiling
> shell test does work: the skip added by patch 1 is driven by that record
> failing, so it now runs on AMD kernels with swfilt, and keeps skipping on
> kernels without it. The 'Test data symbol' shell test matches the exact
> ibs_op event string, so its regex now accepts terms added after
> ldlat=150, and it only checks ldlat when the captured term has it: on
> uarch without ldlat the swfilt variant makes the capture non-empty
> ("swfilt=1"), that the old 'ldlat=0' check would reject.
>
> Suggested-by: Namhyung Kim <namhyung@kernel.org>
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
> tools/perf/arch/x86/util/mem-events.c | 17 ++++++++++++++---
> tools/perf/tests/shell/test_data_symbol.sh | 6 ++++--
> tools/perf/util/mem-events.c | 21 +++++++++++++++++----
> tools/perf/util/mem-events.h | 2 ++
> 4 files changed, 37 insertions(+), 9 deletions(-)
>
> diff --git a/tools/perf/arch/x86/util/mem-events.c b/tools/perf/arch/x86/util/mem-events.c
> index b38f519020ff8c6f..034053dc762101dc 100644
> --- a/tools/perf/arch/x86/util/mem-events.c
> +++ b/tools/perf/arch/x86/util/mem-events.c
> @@ -7,7 +7,10 @@
>
> #define MEM_LOADS_AUX 0x8203
>
> -#define E(t, n, s, l, a) { .tag = t, .name = n, .event_name = s, .ldlat = l, .aux_event = a }
> +#define E_INIT(t, n, s, l, a, sf) { \
> + .tag = t, .name = n, .event_name = s, .swfilt_name = sf, \
> + .ldlat = l, .aux_event = a }
> +#define E(t, n, s, l, a) E_INIT(t, n, s, l, a, NULL)
>
> struct perf_mem_event perf_mem_events_intel[PERF_MEM_EVENTS__MAX] = {
> E("ldlat-loads", "%s/mem-loads,ldlat=%u/P", "mem-loads", true, 0),
> @@ -21,14 +24,22 @@ struct perf_mem_event perf_mem_events_intel_aux[PERF_MEM_EVENTS__MAX] = {
> E(NULL, NULL, NULL, false, 0),
> };
>
> +/*
> + * IBS events with exclude_{user,kernel} bits set, as used by perf to
> + * record per-thread when kernel samples are not allowed, are rejected
> + * by the kernel on hardware without the privilege filter unless the
> + * swfilt software filter is used, so these events carry a variant of
> + * their names with the swfilt term, used by perf_pmu__mem_events_name()
> + * when the kernel exposes the term.
> + */
> struct perf_mem_event perf_mem_events_amd[PERF_MEM_EVENTS__MAX] = {
> E(NULL, NULL, NULL, false, 0),
> E(NULL, NULL, NULL, false, 0),
> - E("mem-ldst", "%s//", NULL, false, 0),
> + E_INIT("mem-ldst", "%s//", NULL, false, 0, "%s/swfilt=1/"),
> };
>
> struct perf_mem_event perf_mem_events_amd_ldlat[PERF_MEM_EVENTS__MAX] = {
> E(NULL, NULL, NULL, false, 0),
> E(NULL, NULL, NULL, false, 0),
> - E("mem-ldst", "%s/ldlat=%u/", NULL, true, 0),
> + E_INIT("mem-ldst", "%s/ldlat=%u/", NULL, true, 0, "%s/ldlat=%u,swfilt=1/"),
> };
> diff --git a/tools/perf/tests/shell/test_data_symbol.sh b/tools/perf/tests/shell/test_data_symbol.sh
> index d61b5659a46d9a77..52c837fddb639595 100755
> --- a/tools/perf/tests/shell/test_data_symbol.sh
> +++ b/tools/perf/tests/shell/test_data_symbol.sh
> @@ -65,15 +65,17 @@ if (($is_amd >= 1)); then
> # --ldlat on AMD:
> # o Zen4 and earlier uarch does not support ldlat
> # o Even on supported platforms, it's disabled (--ldlat=0) by default.
> + # o Kernels with the swfilt term add it even when ldlat is not
> + # supported, so only check ldlat when the term is present.
> ldlat=${BASH_REMATCH[1]}
> - if [[ -n $ldlat ]]; then
> + if [[ $ldlat == *ldlat=* ]]; then
> if ! [[ "$ldlat" =~ ldlat=0 ]]; then
> echo "ERROR: ldlat not initialized to 0?"
> exit 1
> fi
>
> mem_events="$(perf mem record -v --ldlat=150 -e list 2>&1)"
> - if ! [[ "$mem_events" =~ ^mem-ldst.*ibs_op/ldlat=150/.*available ]]; then
> + if ! [[ "$mem_events" =~ ^mem-ldst.*ibs_op/ldlat=150[,/].*available ]]; then
> echo "ERROR: --ldlat not honored?"
> exit 1
> fi
> diff --git a/tools/perf/util/mem-events.c b/tools/perf/util/mem-events.c
> index 0b49fce251fcc184..8f74cd085e500231 100644
> --- a/tools/perf/util/mem-events.c
> +++ b/tools/perf/util/mem-events.c
> @@ -82,6 +82,7 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
> char *buf, size_t buf_size)
> {
> struct perf_mem_event *e;
> + const char *name;
>
> if (i >= PERF_MEM_EVENTS__MAX || !pmu)
> return NULL;
> @@ -90,24 +91,36 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
> if (!e || !e->name)
> return NULL;
>
> + /*
> + * Use the swfilt variant of the name when the PMU exposes the term.
> + * It is not conditional on the event already having exclude bits:
> + * perf record adds those bits itself when the first open fails with
> + * EACCES on an unprivileged setup, after this name has been built,
> + * and that retry only succeeds with the term in the name. With no
> + * exclude bits the kernel doesn't discard anything.
> + */
> + name = e->name;
> + if (e->swfilt_name && perf_pmu__has_format(pmu, "swfilt"))
> + name = e->swfilt_name;
It's a bit unfortunate we hardcode "swfilt" here. Probably better to
have it in the mem-events array and add it to the name format
dynamically. But as it's only needed for IBS, I think it's ok for now.
Thanks,
Namhyung
> +
> if (i == PERF_MEM_EVENTS__LOAD || i == PERF_MEM_EVENTS__LOAD_STORE) {
> if (e->ldlat) {
> if (!e->aux_event) {
> /* ARM and Most of Intel */
> scnprintf(buf, buf_size,
> - e->name, pmu->name,
> + name, pmu->name,
> perf_mem_events__loads_ldlat);
> } else {
> /* Intel with mem-loads-aux event */
> scnprintf(buf, buf_size,
> - e->name, pmu->name, pmu->name,
> + name, pmu->name, pmu->name,
> perf_mem_events__loads_ldlat);
> }
> } else {
> if (!e->aux_event) {
> /* AMD and POWER */
> scnprintf(buf, buf_size,
> - e->name, pmu->name);
> + name, pmu->name);
> } else {
> return NULL;
> }
> @@ -117,7 +130,7 @@ static const char *perf_pmu__mem_events_name(struct perf_pmu *pmu, int i,
>
> if (i == PERF_MEM_EVENTS__STORE) {
> scnprintf(buf, buf_size,
> - e->name, pmu->name);
> + name, pmu->name);
> return buf;
> }
>
> diff --git a/tools/perf/util/mem-events.h b/tools/perf/util/mem-events.h
> index 5b98076904b0b689..41f628fad10709b9 100644
> --- a/tools/perf/util/mem-events.h
> +++ b/tools/perf/util/mem-events.h
> @@ -11,6 +11,8 @@ struct perf_mem_event {
> u32 aux_event;
> const char *tag;
> const char *name;
> + /* Name with the swfilt software privilege filter, when supported. */
> + const char *swfilt_name;
> const char *event_name;
> };
>
> --
> 2.55.0
>
next prev parent reply other threads:[~2026-09-16 21:59 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 11:47 [PATCH v6 0/12] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 01/12] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 02/12] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
2026-09-16 17:59 ` Ian Rogers
2026-09-16 19:02 ` Arnaldo Carvalho de Melo
2026-09-16 21:28 ` Ian Rogers
2026-09-16 18:42 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 03/12] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 04/12] perf debuginfo: Let the user skip and disable debuginfod fetches Arnaldo Carvalho de Melo
2026-09-16 18:53 ` Namhyung Kim
2026-09-16 21:27 ` Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 05/12] perf debuginfo: Show the debuginfod fetch progress and keys in the TUI Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 06/12] perf symbol: Fall back to fetching the vmlinux by build ID Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 07/12] perf annotate-data: Show the sample count in the data-type browser Arnaldo Carvalho de Melo
2026-09-16 21:28 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 08/12] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 09/12] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-09-16 11:47 ` [PATCH 10/12] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
2026-09-16 21:44 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 11/12] perf mem record: Request PERF_SAMPLE_CPU by default Arnaldo Carvalho de Melo
2026-09-16 21:50 ` Namhyung Kim
2026-09-16 11:47 ` [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available Arnaldo Carvalho de Melo
2026-09-16 21:59 ` Namhyung Kim [this message]
2026-09-16 22:27 ` [PATCH v6 0/12] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Namhyung Kim
2026-09-16 18:32 Arnaldo Carvalho de Melo
2026-09-16 18:32 ` [PATCH 12/12] perf mem record: Use the IBS swfilt filter when available Arnaldo Carvalho de Melo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqsRVABj23jaaiSq@google.com \
--to=namhyung@kernel.org \
--cc=acme@kernel.org \
--cc=acme@redhat.com \
--cc=adrian.hunter@intel.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=ravi.bangoria@amd.com \
--cc=tglx@linutronix.de \
--cc=williams@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®