* [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping
@ 2026-09-28 7:42 Dapeng Mi
2026-09-28 7:42 ` [PATCH 01/15] perf/x86/intel: Guard leader sibling walk on nr_siblings Dapeng Mi
` (14 more replies)
0 siblings, 15 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
This series fixes several recently discovered sampling issues:
- a warning about a missing ctx->mutex hold during LBR counter event
initialization.
- PEBS sampling failures after CPU offline/online for both legacy PEBS
and arch-PEBS.
- invalid counts reported by non-PEBS ACR events in SAMPLE_READ mode.
- stale or incorrect counts reported by PEBS events without counter
snapshot support.
- incorrect event counts are seen when multiple PEBS events are active
in a group.
It also relaxes the slots-event restriction for topdown event groups:
the slots event no longer has to be the group leader, which makes
sampling for groups containing topdown metric events more flexible.
Patch layout:
- patch 01: fix the warning triggered by for_each_sibling_event() while
iterating LBR counter events.
- patches 02-04: fix PEBS sampling failures after CPU offline/online for
both legacy PEBS and arch-PEBS.
- patches 05-06: fix minor PEBS handler issues.
- patch 07: fix invalid counts reported by non-PEBS ACR events in
SAMPLE_READ mode.
- patch 08: fix stale counts reported by PEBS events without counter
snapshot support.
- patches 09-11: refactor PEBS handlers to remove duplicated code and
fix incorrect counts when multiple PEBS events are active in group.
- patch 12: make ACR static_call updates architectural to avoid
model-specific updates on every platform
- patch 13: add is_x86_event() checks before testing for
slots/topdown/mem-loads events
- patches 14-15: allow a slots event to be a non-group leader, making
sampling for groups containing topdown metric events more flexible.
Dapeng Mi (15):
perf/x86/intel: Guard leader sibling walk on nr_siblings
perf/x86/intel: Reset active_fixed_ctrl_val on CPU teardown
perf/x86/intel: Reset active_pebs_data_cfg on CPU teardown
perf/x86/intel: Reset cached acr_cfg_b[] and cfg_c_val[] on CPU
teardown
perf/x86/intel: Pass correct PEBS counter mask to no-drain update path
perf/x86/intel: Limit PEBS counter iteration to valid array bounds
perf/x86/intel: Reject SAMPLE_READ for no-counter-snapshot ACR events
perf/x86/intel: Fix stale PEBS count without counter-group support
perf/x86/intel: Refactor intel_pmu_drain_arch_pebs()
perf/x86/intel: Refactor intel_pmu_drain_pebs_icl()
perf/x86/intel: Fix invalid PEBS counts with counter-group support
perf/x86/intel: Make ACR static_call update architectural
perf/x86: Validate event type before topdown/mem-loads classification
perf/core: Add event_caps dependency flags
perf/x86/intel: Allow Topdown metrics with a non-leader slots event
arch/x86/events/intel/core.c | 216 +++++++++++++++++--
arch/x86/events/intel/ds.c | 390 ++++++++++++++++++++++++-----------
arch/x86/events/perf_event.h | 14 +-
include/linux/perf_event.h | 8 +
kernel/events/core.c | 52 ++++-
5 files changed, 535 insertions(+), 145 deletions(-)
base-commit: 6350de8671b94afb7691d110f63bcda42f658a69
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 01/15] perf/x86/intel: Guard leader sibling walk on nr_siblings
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
@ 2026-09-28 7:42 ` Dapeng Mi
2026-09-28 7:42 ` [PATCH 02/15] perf/x86/intel: Reset active_fixed_ctrl_val on CPU teardown Dapeng Mi
` (13 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
The leader event's ctx mutex is only held before calling
pmu->event_init() for sibling events, not for a newly-created group
leader.
Some paths in intel_pmu_hw_config() call for_each_sibling_event() on the
leader itself. At that point, leader->ctx is not yet initialized, so the
lockdep assertion in for_each_sibling_event() can trigger WARN_ON_ONCE(),
and the access to leader->ctx is unsafe.
This loop is unnecessary when the event is the group leader, because the
leader has no siblings yet. Guard the iteration with leader->nr_siblings
before calling for_each_sibling_event().
Also move setting PERF_X86_EVENT_BRANCH_COUNTERS to the end of the
branch-counter validation path so the flag is set only after all checks
succeed.
Fixes: 33744916196b ("perf/x86/intel: Support branch counters logging")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 13 ++++++++-----
1 file changed, 8 insertions(+), 5 deletions(-)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 0a34d674df59..65815d13ae3f 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -5046,11 +5046,12 @@ static int intel_pmu_hw_config(struct perf_event *event)
leader = event->group_leader;
if (intel_set_branch_counter_constr(leader, &num))
return -EINVAL;
- leader->hw.flags |= PERF_X86_EVENT_BRANCH_COUNTERS;
- for_each_sibling_event(sibling, leader) {
- if (intel_set_branch_counter_constr(sibling, &num))
- return -EINVAL;
+ if (leader->nr_siblings) {
+ for_each_sibling_event(sibling, leader) {
+ if (intel_set_branch_counter_constr(sibling, &num))
+ return -EINVAL;
+ }
}
/* event isn't installed as a sibling yet. */
@@ -5069,7 +5070,7 @@ static int intel_pmu_hw_config(struct perf_event *event)
if (0 == (event->attr.branch_sample_type &
~(PERF_SAMPLE_BRANCH_PLM_ALL |
PERF_SAMPLE_BRANCH_COUNTERS)))
- event->hw.flags &= ~PERF_X86_EVENT_NEEDS_BRANCH_STACK;
+ event->hw.flags &= ~PERF_X86_EVENT_NEEDS_BRANCH_STACK;
/*
* Force the leader to be a LBR event. So LBRs can be reset
@@ -5077,6 +5078,8 @@ static int intel_pmu_hw_config(struct perf_event *event)
*/
if (!intel_pmu_needs_branch_stack(leader))
return -EINVAL;
+
+ leader->hw.flags |= PERF_X86_EVENT_BRANCH_COUNTERS;
}
if (intel_pmu_needs_branch_stack(event)) {
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 02/15] perf/x86/intel: Reset active_fixed_ctrl_val on CPU teardown
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
2026-09-28 7:42 ` [PATCH 01/15] perf/x86/intel: Guard leader sibling walk on nr_siblings Dapeng Mi
@ 2026-09-28 7:42 ` Dapeng Mi
2026-09-28 7:42 ` [PATCH 03/15] perf/x86/intel: Reset active_pebs_data_cfg " Dapeng Mi
` (12 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
cpuc->active_fixed_ctrl_val tracks the value currently programmed into
MSR_CORE_PERF_FIXED_CTR_CTRL, and reprogramming is skipped when the new
value matches the cached one.
After a CPU offline/online cycle, MSR_CORE_PERF_FIXED_CTR_CTRL may return
to its reset state, but active_fixed_ctrl_val may still hold the
pre-offline value. If the next computed fixed_ctrl_val matches that stale
cache entry, the MSR update is incorrectly skipped, leaving
MSR_CORE_PERF_FIXED_CTR_CTRL uninitialized for fixed counters and
resulting in errors in fixed counter counting or samplig.
Clear active_fixed_ctrl_val and MSR_CORE_PERF_FIXED_CTR_CTRL during CPU
teardown so MSR_CORE_PERF_FIXED_CTR_CTRL is always reprogrammed when the
CPU comes back online.
Fixes: fae9ebde9696 ("perf/x86/intel: Optimize FIXED_CTR_CTRL access")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 65815d13ae3f..b8dd7b73c31a 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -6591,8 +6591,25 @@ static void free_excl_cntrs(struct cpu_hw_events *cpuc)
cpuc->constraint_list = NULL;
}
+static void fini_fixed_cntrs_on_cpu(int cpu)
+{
+ struct cpu_hw_events *cpuc;
+
+ if (x86_pmu.version < 2)
+ return;
+
+ /*
+ * Clear active_fixed_ctrl_val so MSR_CORE_PERF_FIXED_CTR_CTRL
+ * can be reprogrammed after CPU online.
+ */
+ cpuc = &per_cpu(cpu_hw_events, cpu);
+ cpuc->active_fixed_ctrl_val = 0;
+ wrmsrq_on_cpu(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
+}
+
static void intel_pmu_cpu_dying(int cpu)
{
+ fini_fixed_cntrs_on_cpu(cpu);
fini_debug_store_on_cpu(cpu);
fini_arch_pebs_on_cpu(cpu);
}
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 03/15] perf/x86/intel: Reset active_pebs_data_cfg on CPU teardown
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
2026-09-28 7:42 ` [PATCH 01/15] perf/x86/intel: Guard leader sibling walk on nr_siblings Dapeng Mi
2026-09-28 7:42 ` [PATCH 02/15] perf/x86/intel: Reset active_fixed_ctrl_val on CPU teardown Dapeng Mi
@ 2026-09-28 7:42 ` Dapeng Mi
2026-09-28 7:42 ` [PATCH 04/15] perf/x86/intel: Reset cached acr_cfg_b[] and cfg_c_val[] " Dapeng Mi
` (11 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
cpuc->active_pebs_data_cfg tracks the PEBS data configuration currently
programmed into MSR_PEBS_DATA_CFG, and reprogramming is skipped when the
new value matches the cached one.
After a CPU offline/online cycle, MSR_PEBS_DATA_CFG may return to its
reset state, but active_pebs_data_cfg may still hold the pre-offline
value. If the next computed PEBS data configuration matches that stale
cache entry, the MSR update is incorrectly skipped, leaving
MSR_PEBS_DATA_CFG uninitialized for PEBS sampling and resulting in
incorrect PEBS records.
Clear active_pebs_data_cfg and MSR_PEBS_DATA_CFG during CPU teardown so
MSR_PEBS_DATA_CFG is always reprogrammed when the CPU comes back online.
Fixes: c22497f5838c ("perf/x86/intel: Support adaptive PEBS v4")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index b8dd7b73c31a..2f3eaf96daf5 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -6607,9 +6607,26 @@ static void fini_fixed_cntrs_on_cpu(int cpu)
wrmsrq_on_cpu(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
}
+static void fini_adaptive_pebs_on_cpu(int cpu)
+{
+ struct cpu_hw_events *cpuc;
+
+ if (!x86_pmu.ds_pebs || !x86_pmu.intel_cap.pebs_baseline)
+ return;
+
+ /*
+ * Clear active_pebs_data_cfg so MSR_PEBS_DATA_CFG can
+ * be reprogrammed after CPU online.
+ */
+ cpuc = &per_cpu(cpu_hw_events, cpu);
+ cpuc->active_pebs_data_cfg = 0;
+ wrmsrq_on_cpu(cpu, MSR_PEBS_DATA_CFG, 0);
+}
+
static void intel_pmu_cpu_dying(int cpu)
{
fini_fixed_cntrs_on_cpu(cpu);
+ fini_adaptive_pebs_on_cpu(cpu);
fini_debug_store_on_cpu(cpu);
fini_arch_pebs_on_cpu(cpu);
}
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 04/15] perf/x86/intel: Reset cached acr_cfg_b[] and cfg_c_val[] on CPU teardown
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (2 preceding siblings ...)
2026-09-28 7:42 ` [PATCH 03/15] perf/x86/intel: Reset active_pebs_data_cfg " Dapeng Mi
@ 2026-09-28 7:42 ` Dapeng Mi
2026-09-28 7:42 ` [PATCH 05/15] perf/x86/intel: Pass correct PEBS counter mask to no-drain update path Dapeng Mi
` (10 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Reset cpuc->acr_cfg_b[] and cpuc->cfg_c_val[] and corresponding MSRs in
the CPU offline path.
These arrays cache the last programmed values for the *_CFG_B and
*_CFG_C MSRs, and matching values can cause reprogramming to be skipped.
After CPU hotplug, those MSRs may return to reset defaults, but the cached
values may still reflect the pre-offline state. This can incorrectly
skip MSR writes when events are enabled again, leaving the hardware with
reset MSR contents.
Clear both caches and corresponding *_CFG_B/*_CFG_C MSRs on CPU teardown,
so *_CFG_B and *_CFG_C are always reprogrammed after CPU online.
Fixes: 52448a0a7390 ("perf/x86/intel: Setup PEBS data configuration and enable legacy groups")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 38 ++++++++++++++++++++++++++++++++++++
1 file changed, 38 insertions(+)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 2f3eaf96daf5..377ff3912420 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -6623,9 +6623,47 @@ static void fini_adaptive_pebs_on_cpu(int cpu)
wrmsrq_on_cpu(cpu, MSR_PEBS_DATA_CFG, 0);
}
+#define clear_pmu_ext_msrs_on_cpu(cpu, mask, gp_base, fixed_base) \
+do { \
+ int idx, msr; \
+ \
+ for_each_set_bit(idx, (unsigned long *)&(mask), X86_PMC_IDX_MAX) { \
+ if (idx < INTEL_PMC_IDX_FIXED) { \
+ msr = (gp_base) + x86_pmu.addr_offset(idx, false); \
+ } else { \
+ msr = (fixed_base) + \
+ x86_pmu.addr_offset(idx - INTEL_PMC_IDX_FIXED, false); \
+ } \
+ wrmsrq_on_cpu((cpu), msr, 0); \
+ } \
+} while (0)
+
+static void fini_pmu_ext_msrs_on_cpu(int cpu)
+{
+ struct cpu_hw_events *cpuc = &per_cpu(cpu_hw_events, cpu);
+ u64 cfg_b_mask = hybrid(cpuc->pmu, acr_cntr_mask64);
+ u64 cfg_c_mask = cfg_b_mask |
+ hybrid(cpuc->pmu, arch_pebs_cap).counters;
+
+ if (x86_pmu.version < 6)
+ return;
+
+ /*
+ * Clear cached acr_cfg_b[] and cfg_c_val[] so *_CFG_B and
+ *_CFG_C MSRs can always be reprogrammed after CPU online.
+ */
+ memset(cpuc->acr_cfg_b, 0, sizeof(cpuc->acr_cfg_b));
+ memset(cpuc->cfg_c_val, 0, sizeof(cpuc->cfg_c_val));
+ clear_pmu_ext_msrs_on_cpu(cpu, cfg_b_mask, MSR_IA32_PMC_V6_GP0_CFG_B,
+ MSR_IA32_PMC_V6_FX0_CFG_B);
+ clear_pmu_ext_msrs_on_cpu(cpu, cfg_c_mask, MSR_IA32_PMC_V6_GP0_CFG_C,
+ MSR_IA32_PMC_V6_FX0_CFG_C);
+}
+
static void intel_pmu_cpu_dying(int cpu)
{
fini_fixed_cntrs_on_cpu(cpu);
+ fini_pmu_ext_msrs_on_cpu(cpu);
fini_adaptive_pebs_on_cpu(cpu);
fini_debug_store_on_cpu(cpu);
fini_arch_pebs_on_cpu(cpu);
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 05/15] perf/x86/intel: Pass correct PEBS counter mask to no-drain update path
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (3 preceding siblings ...)
2026-09-28 7:42 ` [PATCH 04/15] perf/x86/intel: Reset cached acr_cfg_b[] and cfg_c_val[] " Dapeng Mi
@ 2026-09-28 7:42 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 06/15] perf/x86/intel: Limit PEBS counter iteration to valid array bounds Dapeng Mi
` (9 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:42 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
intel_pmu_drain_arch_pebs() calls intel_pmu_pebs_event_update_no_drain()
with X86_PMC_IDX_MAX, but that parameter is a counter bitmask, not a
max index value. As a result, PEBS event state is not updated/restored
correctly on the no-drain path.
Pass the PEBS counter mask to intel_pmu_pebs_event_update_no_drain().
Fixes: d21954c8a0ff ("perf/x86/intel: Process arch-PEBS records or record fragments")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 5 ++---
1 file changed, 2 insertions(+), 3 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index 8444670cee4a..0f24098587bf 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -3358,8 +3358,9 @@ static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
rdmsrq(MSR_IA32_PEBS_INDEX, index.whole);
+ mask = hybrid(cpuc->pmu, arch_pebs_cap).counters & cpuc->pebs_enabled;
if (unlikely(!index.wr)) {
- intel_pmu_pebs_event_update_no_drain(cpuc, X86_PMC_IDX_MAX);
+ intel_pmu_pebs_event_update_no_drain(cpuc, mask);
return 0;
}
@@ -3375,8 +3376,6 @@ static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
index.thresh = ARCH_PEBS_THRESH_SINGLE;
wrmsrq(MSR_IA32_PEBS_INDEX, index.whole);
- mask = hybrid(cpuc->pmu, arch_pebs_cap).counters & cpuc->pebs_enabled;
-
if (!iregs)
iregs = &dummy_iregs;
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 06/15] perf/x86/intel: Limit PEBS counter iteration to valid array bounds
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (4 preceding siblings ...)
2026-09-28 7:42 ` [PATCH 05/15] perf/x86/intel: Pass correct PEBS counter mask to no-drain update path Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 07/15] perf/x86/intel: Reject SAMPLE_READ for no-counter-snapshot ACR events Dapeng Mi
` (8 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
__intel_pmu_handle_{,last_}pebs_record() iterate counter indexes up to
X86_PMC_IDX_MAX (64), while counts[] and *last[] are sized only for
INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS (48) entries.
If pebs_status were ever to contain bits above the highest valid PEBS
counter index, the loop could access those arrays out of bounds.
Although that should not happen in practice, limit the iteration bound
to INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS to match the array
sizes and keep the code consistent.
Fixes: 8807d922705f ("perf/x86/intel/ds: Factor out PEBS record processing code to functions")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index 0f24098587bf..f68391d60b2d 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -3239,7 +3239,8 @@ __intel_pmu_handle_pebs_record(struct pt_regs *iregs,
struct perf_event *event;
int bit;
- for_each_set_bit(bit, (unsigned long *)&pebs_status, X86_PMC_IDX_MAX) {
+ for_each_set_bit(bit, (unsigned long *)&pebs_status,
+ INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS) {
event = cpuc->events[bit];
if (WARN_ON_ONCE(!event) ||
@@ -3268,7 +3269,8 @@ __intel_pmu_handle_last_pebs_record(struct pt_regs *iregs,
bool handled = false;
int bit;
- for_each_set_bit(bit, (unsigned long *)&mask, X86_PMC_IDX_MAX) {
+ for_each_set_bit(bit, (unsigned long *)&mask,
+ INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS) {
if (!counts[bit])
continue;
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 07/15] perf/x86/intel: Reject SAMPLE_READ for no-counter-snapshot ACR events
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (5 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 06/15] perf/x86/intel: Limit PEBS counter iteration to valid array bounds Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 08/15] perf/x86/intel: Fix stale PEBS count without counter-group support Dapeng Mi
` (7 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
ACR events cannot always provide a reliable value through SAMPLE_READ.
For non-PEBS ACR events, another ACR overflow can auto-reload the counter
before software reads it, so software cannot sample the exact count before
the hardware reload.
PEBS-backed ACR events are safe only when counter snapshot support is
available, because the value is captured in the PEBS record before the
counter is reloaded.
Reject SAMPLE_READ for ACR event groups unless the event has PEBS
counter snapshot support, so perf does not report invalid counts.
Reported-by: Andi Kleen <ak@linux.intel.com>
Fixes: ec980e4facef ("perf/x86/intel: Support auto counter reload")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 59 ++++++++++++++++++++++++++++++++++++
1 file changed, 59 insertions(+)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 377ff3912420..3fc3795534bf 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -4974,6 +4974,53 @@ static inline int intel_set_branch_counter_constr(struct perf_event *event,
return 0;
}
+static inline bool is_acr_sample_read_allowed(struct perf_event *event,
+ bool group_has_sample_read)
+{
+ /*
+ * ACR events cannot report an accurate count for non-PEBS events
+ * or for PEBS events without counter snapshots: another ACR event
+ * may overflow andauto-reload the counter before software can read
+ * the precise value.
+ *
+ * We keep the check simple and do not validate the acr_mask precisely
+ * to determine whether the SAMPLE_READ event is actually auto-reloaded
+ * by another ACR event. If a SAMPLE_READ event is in the group, the
+ * ACR event must be a PEBS event with counter snapshots; otherwise it
+ * is rejected.
+ */
+ if (group_has_sample_read && is_sampling_event(event) &&
+ (!event->attr.precise_ip || !is_pebs_counter_event_group(event)))
+ return false;
+
+ return true;
+}
+
+static bool intel_pmu_allow_acr_sample_read(struct perf_event *event,
+ bool group_has_sample_read)
+{
+ struct perf_event *leader = event->group_leader;
+ struct perf_event *sibling;
+
+ if (!is_acr_sample_read_allowed(leader, group_has_sample_read))
+ return false;
+
+ if (leader->nr_siblings) {
+ for_each_sibling_event(sibling, leader) {
+ if (!is_acr_sample_read_allowed(sibling,
+ group_has_sample_read))
+ return false;
+ }
+ }
+
+ /* event isn't installed as a sibling yet. */
+ if ((event != leader) &&
+ !is_acr_sample_read_allowed(event, group_has_sample_read))
+ return false;
+
+ return true;
+}
+
static int intel_pmu_hw_config(struct perf_event *event)
{
int ret = x86_pmu_hw_config(event);
@@ -5117,6 +5164,7 @@ static int intel_pmu_hw_config(struct perf_event *event)
struct perf_event *sibling, *leader = event->group_leader;
struct pmu *pmu = event->pmu;
bool has_sw_event = false;
+ bool has_sample_read = false;
int num = 0, idx = 0;
u64 cause_mask = 0;
@@ -5162,8 +5210,14 @@ static int intel_pmu_hw_config(struct perf_event *event)
if (leader->attr.config2)
intel_pmu_set_acr_cntr_constr(leader, &cause_mask, &num);
+ if ((leader->attr.sample_type & PERF_SAMPLE_READ) ||
+ (event->attr.sample_type & PERF_SAMPLE_READ))
+ has_sample_read = true;
+
if (leader->nr_siblings) {
for_each_sibling_event(sibling, leader) {
+ if (sibling->attr.sample_type & PERF_SAMPLE_READ)
+ has_sample_read = true;
if (!is_x86_event(sibling)) {
has_sw_event = true;
continue;
@@ -5175,6 +5229,7 @@ static int intel_pmu_hw_config(struct perf_event *event)
intel_pmu_set_acr_cntr_constr(sibling, &cause_mask, &num);
}
}
+
if (leader != event && event->attr.config2) {
if (has_sw_event)
return -EINVAL;
@@ -5184,6 +5239,10 @@ static int intel_pmu_hw_config(struct perf_event *event)
if (hweight64(cause_mask) > hweight64(hybrid(pmu, acr_cause_mask64)) ||
num > hweight64(hybrid(event->pmu, acr_cntr_mask64)))
return -EINVAL;
+
+ if (!intel_pmu_allow_acr_sample_read(event, has_sample_read))
+ return -EINVAL;
+
/*
* In the second round, apply the counter-constraints for
* the events which can cause other events reload.
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 08/15] perf/x86/intel: Fix stale PEBS count without counter-group support
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (6 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 07/15] perf/x86/intel: Reject SAMPLE_READ for no-counter-snapshot ACR events Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 09/15] perf/x86/intel: Refactor intel_pmu_drain_arch_pebs() Dapeng Mi
` (6 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
On platforms without PEBS counter-group support (for example, Sapphire
Rapids), PEBS group read samples can report mismatched counts between the
precise event and its sibling, e.g.,
$ perf record -e \
'{cpu/instructions,period=100000/pu,cpu/instructions/u}:S' -- sleep 1
The invalid event counts are seen,
1st PEBS record:
.... group nr 2
..... id 0000000000006189, value 0000000000000000, lost 0
..... id 0000000000006269, value 00000000000186c1, lost 0
2nd PEBS record:
.... group nr 2
..... id 0000000000006189, value 00000000000186a0, lost 0
..... id 0000000000006269, value 0000000000030d6b, lost 0
In PEBS records, the first event count can lag behind the second event by
about one sample period, even though both counts should be close or equal.
The root cause is ordering in __intel_pmu_pebs_last_event(): PEBS event
counts are updated after perf_event_output()/perf_event_overflow(). As a
result, perf_output_read_group() snapshots stale PEBS counts.
Move the PEBS count update before
perf_event_output()/perf_event_overflow(), matching the non-PEBS ordering
in handle_pmi_common(), so group reads report correct PEBS event counts.
Fixes: 3c00ed344cef ("perf/x86/intel/ds: Factor out functions for PEBS records processing")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 33 +++++++++++++++++----------------
1 file changed, 17 insertions(+), 16 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index f68391d60b2d..1a32f4168a84 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -2956,22 +2956,6 @@ __intel_pmu_pebs_last_event(struct perf_event *event,
setup_sample(event, iregs, at, data, regs);
}
- if (iregs == &dummy_iregs) {
- /*
- * The PEBS records may be drained in the non-overflow context,
- * e.g., large PEBS + context switch. Perf should treat the
- * last record the same as other PEBS records, and doesn't
- * invoke the generic overflow handler.
- */
- perf_event_output(event, data, regs);
- } else {
- /*
- * All but the last records are processed.
- * The last one is left to be able to call the overflow handler.
- */
- perf_event_overflow(event, data, regs);
- }
-
if (hwc->flags & PERF_X86_EVENT_AUTO_RELOAD) {
if (is_pebs_counter_event_group(event)) {
if (corrupted) {
@@ -3011,6 +2995,23 @@ __intel_pmu_pebs_last_event(struct perf_event *event,
else
intel_pmu_save_and_restart(event);
}
+
+ if (iregs == &dummy_iregs) {
+ /*
+ * The PEBS records may be drained in the non-overflow context,
+ * e.g., large PEBS + context switch. Perf should treat the
+ * last record the same as other PEBS records, and doesn't
+ * invoke the generic overflow handler.
+ */
+ perf_event_output(event, data, regs);
+ } else {
+ /*
+ * All but the last records are processed.
+ * The last one is left to be able to call the overflow handler.
+ */
+ perf_event_overflow(event, data, regs);
+ }
+
}
static DEFINE_PER_CPU(struct x86_perf_regs, x86_pebs_regs);
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 09/15] perf/x86/intel: Refactor intel_pmu_drain_arch_pebs()
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (7 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 08/15] perf/x86/intel: Fix stale PEBS count without counter-group support Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 10/15] perf/x86/intel: Refactor intel_pmu_drain_pebs_icl() Dapeng Mi
` (5 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
intel_pmu_drain_arch_pebs() and intel_pmu_drain_pebs_icl() use the same
high-level PEBS drain flow. To reduce duplication, extract the shared
logic into a common helper, intel_pmu_handle_pebs_records(), and have
the arch-PEBS path use it.
Add small helper wrappers, including get_pebs_cntr_mask(),
get_pebs_cntr_status(), and find_next_pebs_record(), to hide
record-layout differences between ICL adaptive PEBS and arch PEBS.
This patch refactors only intel_pmu_drain_arch_pebs(); converting
intel_pmu_drain_pebs_icl() to the same helper is done in the next patch.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 163 ++++++++++++++++++++++---------------
1 file changed, 96 insertions(+), 67 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index 1a32f4168a84..961f7138387a 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -3267,7 +3267,6 @@ __intel_pmu_handle_last_pebs_record(struct pt_regs *iregs,
{
struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
struct perf_event *event;
- bool handled = false;
int bit;
for_each_set_bit(bit, (unsigned long *)&mask,
@@ -3275,15 +3274,101 @@ __intel_pmu_handle_last_pebs_record(struct pt_regs *iregs,
if (!counts[bit])
continue;
- handled = true;
event = cpuc->events[bit];
__intel_pmu_pebs_last_event(event, iregs, regs, data, last[bit],
counts[bit], corrupted, setup_sample);
}
+}
+
+static inline u64 get_pebs_cntr_mask(struct cpu_hw_events *cpuc)
+{
+ return hybrid(cpuc->pmu, arch_pebs_cap).counters &
+ cpuc->pebs_enabled;
+}
+
+static inline u64 get_pebs_cntr_status(void *at)
+{
+ struct arch_pebs_basic *basic;
+
+ basic = at + sizeof(struct arch_pebs_header);
+ return basic->applicable_counters;
+}
+
+static void *find_next_pebs_record(void *at, void *top)
+{
+ struct arch_pebs_header *header;
+
+ header = at;
+ if (WARN_ON_ONCE(!header->size))
+ return NULL;
+
+ /* 1st fragment or single record must have basic group */
+ if (WARN_ON_ONCE(!header->basic))
+ return NULL;
+
+ /* Skip non-last fragments */
+ while (arch_pebs_record_continued(header)) {
+ if (!header->size)
+ return NULL;
+ at += header->size;
+ if (WARN_ON_ONCE(at >= top))
+ return NULL;
+ header = at;
+ }
+
+ /* Skip last fragment or the single record */
+ at += header->size;
+ if (WARN_ON_ONCE(at > top))
+ return NULL;
+
+ return at;
+}
+
+static __always_inline int
+intel_pmu_handle_pebs_records(struct pt_regs *iregs,
+ struct perf_sample_data *data,
+ void *base, void *top,
+ setup_fn setup_sample)
+{
+ struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
+ short counts[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS] = {};
+ void *last[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS];
+ struct x86_perf_regs *perf_regs = this_cpu_ptr(&x86_pebs_regs);
+ struct pt_regs *regs = &perf_regs->regs;
+ u64 mask = get_pebs_cntr_mask(cpuc);
+ u64 events_bitmap = 0;
+ bool corrupted = false;
+ void *at, *next;
+
+ if (!iregs)
+ iregs = &dummy_iregs;
+
+ /* Process all but the last event for each counter. */
+ for (at = base; at < top;) {
+ u64 pebs_status;
+
+ next = find_next_pebs_record(at, top);
+ if (!next) {
+ corrupted = true;
+ break;
+ }
+
+ pebs_status = mask & get_pebs_cntr_status(at);
+ events_bitmap |= pebs_status;
+ __intel_pmu_handle_pebs_record(iregs, regs, data, at,
+ pebs_status, counts, last,
+ setup_sample);
+ at = next;
+ }
+
+ __intel_pmu_handle_last_pebs_record(iregs, regs, data, mask,
+ counts, last, corrupted,
+ setup_sample);
- /* All records are corrupted, reset sampling period. */
- if (!handled)
+ if (!events_bitmap)
intel_pmu_pebs_event_update_no_drain(cpuc, mask);
+
+ return hweight64(events_bitmap);
}
static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs, struct perf_sample_data *data)
@@ -3342,22 +3427,19 @@ static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs, struct perf_sample_da
__intel_pmu_handle_last_pebs_record(iregs, regs, data, mask, counts, last,
corrupted, setup_pebs_adaptive_sample_data);
+ if (!events_bitmap)
+ intel_pmu_pebs_event_update_no_drain(cpuc, mask);
+
return hweight64(events_bitmap);
}
static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
struct perf_sample_data *data)
{
- short counts[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS] = {};
- void *last[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS];
struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
+ u64 mask = get_pebs_cntr_mask(cpuc);
union arch_pebs_index index;
- struct x86_perf_regs *perf_regs = this_cpu_ptr(&x86_pebs_regs);
- struct pt_regs *regs = &perf_regs->regs;
- void *base, *at, *top;
- u64 events_bitmap = 0;
- bool corrupted = false;
- u64 mask;
+ void *base, *top;
rdmsrq(MSR_IA32_PEBS_INDEX, index.whole);
@@ -3379,61 +3461,8 @@ static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
index.thresh = ARCH_PEBS_THRESH_SINGLE;
wrmsrq(MSR_IA32_PEBS_INDEX, index.whole);
- if (!iregs)
- iregs = &dummy_iregs;
-
- /* Process all but the last event for each counter. */
- for (at = base; at < top;) {
- struct arch_pebs_header *header;
- struct arch_pebs_basic *basic;
- u64 pebs_status;
-
- header = at;
-
- if (WARN_ON_ONCE(!header->size)) {
- corrupted = true;
- goto done;
- }
-
- /* 1st fragment or single record must have basic group */
- if (!header->basic) {
- at += header->size;
- continue;
- }
-
- basic = at + sizeof(struct arch_pebs_header);
- pebs_status = mask & basic->applicable_counters;
- events_bitmap |= pebs_status;
- __intel_pmu_handle_pebs_record(iregs, regs, data, at,
- pebs_status, counts, last,
- setup_arch_pebs_sample_data);
-
- /* Skip non-last fragments */
- while (arch_pebs_record_continued(header)) {
- if (!header->size)
- break;
- at += header->size;
- if (WARN_ON_ONCE(at >= top)) {
- corrupted = true;
- goto done;
- }
- header = at;
- }
-
- /* Skip last fragment or the single record */
- at += header->size;
- if (WARN_ON_ONCE(at > top)) {
- corrupted = true;
- goto done;
- }
- }
-
-done:
- __intel_pmu_handle_last_pebs_record(iregs, regs, data, mask,
- counts, last, corrupted,
- setup_arch_pebs_sample_data);
-
- return hweight64(events_bitmap);
+ return intel_pmu_handle_pebs_records(iregs, data, base, top,
+ setup_arch_pebs_sample_data);
}
static void __init intel_arch_pebs_init(void)
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 10/15] perf/x86/intel: Refactor intel_pmu_drain_pebs_icl()
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (8 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 09/15] perf/x86/intel: Refactor intel_pmu_drain_arch_pebs() Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 11/15] perf/x86/intel: Fix invalid PEBS counts with counter-group support Dapeng Mi
` (4 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Switch intel_pmu_drain_pebs_icl() to the common PEBS drain path provided
by intel_pmu_handle_pebs_records(), reducing duplicated logic between
PEBS drain implementations.
Extend the helper layer, including get_pebs_cntr_mask(),
get_pebs_cntr_status(), and find_next_pebs_record(), so it also handles
the ICL adaptive PEBS record layout.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 108 +++++++++++++++++++------------------
1 file changed, 56 insertions(+), 52 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index 961f7138387a..1562d4cb1903 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -3282,19 +3282,53 @@ __intel_pmu_handle_last_pebs_record(struct pt_regs *iregs,
static inline u64 get_pebs_cntr_mask(struct cpu_hw_events *cpuc)
{
- return hybrid(cpuc->pmu, arch_pebs_cap).counters &
- cpuc->pebs_enabled;
+ u64 mask;
+
+ if (x86_pmu.arch_pebs)
+ mask = hybrid(cpuc->pmu, arch_pebs_cap).counters;
+ else {
+ mask = hybrid(cpuc->pmu, pebs_events_mask) |
+ hybrid(cpuc->pmu, fixed_cntr_mask64) << INTEL_PMC_IDX_FIXED;
+ }
+
+ return mask & cpuc->pebs_enabled;
}
static inline u64 get_pebs_cntr_status(void *at)
{
- struct arch_pebs_basic *basic;
+ u64 status;
+
+ if (x86_pmu.arch_pebs) {
+ struct arch_pebs_basic *basic;
+
+ basic = at + sizeof(struct arch_pebs_header);
+ status = basic->applicable_counters;
+ } else {
+ struct pebs_basic *basic = at;
- basic = at + sizeof(struct arch_pebs_header);
- return basic->applicable_counters;
+ status = basic->applicable_counters;
+ }
+
+ return status;
}
-static void *find_next_pebs_record(void *at, void *top)
+static void *icl_pebs_find_next_pebs_record(void *at, void *top)
+{
+ struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
+ struct pebs_basic *basic = at;
+
+ if (WARN_ON_ONCE(!basic->format_size))
+ return NULL;
+
+ if (WARN_ON_ONCE(basic->format_size != cpuc->pebs_record_size))
+ return NULL;
+
+ at += basic->format_size;
+
+ return at;
+}
+
+static void *arch_pebs_find_next_pebs_record(void *at, void *top)
{
struct arch_pebs_header *header;
@@ -3324,6 +3358,14 @@ static void *find_next_pebs_record(void *at, void *top)
return at;
}
+static void *find_next_pebs_record(void *at, void *top)
+{
+ if (x86_pmu.arch_pebs)
+ return arch_pebs_find_next_pebs_record(at, top);
+ else
+ return icl_pebs_find_next_pebs_record(at, top);
+}
+
static __always_inline int
intel_pmu_handle_pebs_records(struct pt_regs *iregs,
struct perf_sample_data *data,
@@ -3371,66 +3413,28 @@ intel_pmu_handle_pebs_records(struct pt_regs *iregs,
return hweight64(events_bitmap);
}
-static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs, struct perf_sample_data *data)
+static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs,
+ struct perf_sample_data *data)
{
- short counts[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS] = {};
- void *last[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS];
struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
+ u64 mask = get_pebs_cntr_mask(cpuc);
struct debug_store *ds = cpuc->ds;
- struct x86_perf_regs *perf_regs = this_cpu_ptr(&x86_pebs_regs);
- struct pt_regs *regs = &perf_regs->regs;
- struct pebs_basic *basic;
- void *base, *at, *top;
- u64 events_bitmap = 0;
- bool corrupted = false;
- u64 mask;
+ void *base, *top;
if (!x86_pmu.pebs_active)
return 0;
- base = (struct pebs_basic *)(unsigned long)ds->pebs_buffer_base;
- top = (struct pebs_basic *)(unsigned long)ds->pebs_index;
-
+ base = (void *)(unsigned long)ds->pebs_buffer_base;
+ top = (void *)(unsigned long)ds->pebs_index;
ds->pebs_index = ds->pebs_buffer_base;
- mask = hybrid(cpuc->pmu, pebs_events_mask) |
- (hybrid(cpuc->pmu, fixed_cntr_mask64) << INTEL_PMC_IDX_FIXED);
- mask &= cpuc->pebs_enabled;
-
if (unlikely(base >= top)) {
intel_pmu_pebs_event_update_no_drain(cpuc, mask);
return 0;
}
- if (!iregs)
- iregs = &dummy_iregs;
-
- /* Process all but the last event for each counter. */
- for (at = base; at < top; at += basic->format_size) {
- u64 pebs_status;
-
- basic = at;
- if (WARN_ON_ONCE(!basic->format_size)) {
- corrupted = true;
- break;
- }
- if (basic->format_size != cpuc->pebs_record_size)
- continue;
-
- pebs_status = mask & basic->applicable_counters;
- events_bitmap |= pebs_status;
- __intel_pmu_handle_pebs_record(iregs, regs, data, at,
- pebs_status, counts, last,
- setup_pebs_adaptive_sample_data);
- }
-
- __intel_pmu_handle_last_pebs_record(iregs, regs, data, mask, counts, last,
- corrupted, setup_pebs_adaptive_sample_data);
-
- if (!events_bitmap)
- intel_pmu_pebs_event_update_no_drain(cpuc, mask);
-
- return hweight64(events_bitmap);
+ return intel_pmu_handle_pebs_records(iregs, data, base, top,
+ setup_pebs_adaptive_sample_data);
}
static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 11/15] perf/x86/intel: Fix invalid PEBS counts with counter-group support
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (9 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 10/15] perf/x86/intel: Refactor intel_pmu_drain_pebs_icl() Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 12/15] perf/x86/intel: Make ACR static_call update architectural Dapeng Mi
` (3 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
On platforms with PEBS counter-group support (for example ARL and NVL),
running mixed PEBS events can occasionally produce an invalidly large
count for the instructions event, even exceeding the hardware counter
width (48 bits).
$ perf record -e '{cpu_core/instructions,period=100000/u,\
cpu_core/l1-icache-load-misses,period=100000/u}:pS' -- ./foo
The issue is caused by PEBS drain ordering. To avoid scanning the PEBS
buffer twice, current drain helpers defer each event's last PEBS record
and do not process all records strictly in sampling order. That is safe
when no counter-group PEBS event is involved, but it can corrupt PEBS
count updates when two or more active PEBS events use counter-group
snapshotting.
Add intel_pmu_handle_pebs_records_in_order() and use it when two or more
counter-group PEBS events are active, so records are processed strictly in
sampling order for both adaptive PEBS and arch PEBS. This preserves
correct count updates and prevents invalid PEBS counts.
Fixes: 8807d922705f ("perf/x86/intel/ds: Factor out PEBS record processing code to functions")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/ds.c | 131 +++++++++++++++++++++++++++++++++++--
1 file changed, 127 insertions(+), 4 deletions(-)
diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c
index 1562d4cb1903..06ce3129e4a9 100644
--- a/arch/x86/events/intel/ds.c
+++ b/arch/x86/events/intel/ds.c
@@ -3413,6 +3413,113 @@ intel_pmu_handle_pebs_records(struct pt_regs *iregs,
return hweight64(events_bitmap);
}
+static __always_inline int
+intel_pmu_handle_pebs_records_in_order(struct pt_regs *iregs,
+ struct perf_sample_data *data,
+ void *base, void *top,
+ setup_fn setup_sample)
+{
+ struct cpu_hw_events *cpuc = this_cpu_ptr(&cpu_hw_events);
+ short counts[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS] = {};
+ void *last[INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS] = {NULL};
+ struct x86_perf_regs *perf_regs = this_cpu_ptr(&x86_pebs_regs);
+ struct pt_regs *regs = &perf_regs->regs;
+ u64 mask = get_pebs_cntr_mask(cpuc);
+ u64 events_bitmap = 0;
+ bool corrupted = false;
+ void *at, *next;
+ int bit;
+
+ if (!iregs)
+ iregs = &dummy_iregs;
+
+ /* Find the last PEBS record for each PEBS event. */
+ for (at = base; at < top;) {
+ u64 pebs_status;
+
+ next = find_next_pebs_record(at, top);
+ if (!next) {
+ corrupted = true;
+ break;
+ }
+
+ pebs_status = mask & get_pebs_cntr_status(at);
+ for_each_set_bit(bit, (unsigned long *)&pebs_status,
+ INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS) {
+ struct perf_event *event = cpuc->events[bit];
+
+ if (WARN_ON_ONCE(!event) ||
+ WARN_ON_ONCE(!event->attr.precise_ip))
+ continue;
+
+ last[bit] = at;
+ }
+ at = next;
+ }
+
+ /* Process all PEBS records in sampling order. */
+ for (at = base; at < top;) {
+ u64 pebs_status;
+
+ next = find_next_pebs_record(at, top);
+ if (!next)
+ break;
+
+ pebs_status = mask & get_pebs_cntr_status(at);
+ events_bitmap |= pebs_status;
+
+ for_each_set_bit(bit, (unsigned long *)&pebs_status,
+ INTEL_PMC_IDX_FIXED + MAX_FIXED_PEBS_EVENTS) {
+ struct perf_event *event = cpuc->events[bit];
+
+ if (WARN_ON_ONCE(!event) ||
+ WARN_ON_ONCE(!event->attr.precise_ip))
+ continue;
+
+ if (last[bit] && (at != last[bit])) {
+ counts[bit]++;
+ __intel_pmu_pebs_event(event, iregs, regs, data,
+ at, setup_sample);
+ } else {
+ __intel_pmu_pebs_last_event(event, iregs, regs, data,
+ at, counts[bit] + 1,
+ corrupted,
+ setup_sample);
+ }
+ }
+ at = next;
+ }
+
+ if (!events_bitmap)
+ intel_pmu_pebs_event_update_no_drain(cpuc, mask);
+
+ return hweight64(events_bitmap);
+}
+
+static bool must_process_pebs_in_order(struct cpu_hw_events *cpuc)
+{
+ int count = 0;
+ int bit;
+
+ /*
+ * When multiple PEBS events with counter-group are active, PEBS
+ * records must be processed in sampling order; otherwise counter-group
+ * updates can corrupt PEBS event counts.
+ */
+ for_each_set_bit(bit, (unsigned long *)&cpuc->pebs_enabled, X86_PMC_IDX_MAX) {
+ struct perf_event *event = cpuc->events[bit];
+
+ if (!event || !event->attr.precise_ip)
+ continue;
+ if (is_pebs_counter_event_group(event))
+ count++;
+ if (count > 1)
+ return true;
+ }
+
+ return false;
+}
+
static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs,
struct perf_sample_data *data)
{
@@ -3420,6 +3527,7 @@ static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs,
u64 mask = get_pebs_cntr_mask(cpuc);
struct debug_store *ds = cpuc->ds;
void *base, *top;
+ int handled;
if (!x86_pmu.pebs_active)
return 0;
@@ -3433,8 +3541,15 @@ static int intel_pmu_drain_pebs_icl(struct pt_regs *iregs,
return 0;
}
- return intel_pmu_handle_pebs_records(iregs, data, base, top,
- setup_pebs_adaptive_sample_data);
+ if (must_process_pebs_in_order(cpuc)) {
+ handled = intel_pmu_handle_pebs_records_in_order(iregs, data,
+ base, top, setup_pebs_adaptive_sample_data);
+ } else {
+ handled = intel_pmu_handle_pebs_records(iregs, data,
+ base, top, setup_pebs_adaptive_sample_data);
+ }
+
+ return handled;
}
static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
@@ -3444,6 +3559,7 @@ static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
u64 mask = get_pebs_cntr_mask(cpuc);
union arch_pebs_index index;
void *base, *top;
+ int handled;
rdmsrq(MSR_IA32_PEBS_INDEX, index.whole);
@@ -3465,8 +3581,15 @@ static int intel_pmu_drain_arch_pebs(struct pt_regs *iregs,
index.thresh = ARCH_PEBS_THRESH_SINGLE;
wrmsrq(MSR_IA32_PEBS_INDEX, index.whole);
- return intel_pmu_handle_pebs_records(iregs, data, base, top,
- setup_arch_pebs_sample_data);
+ if (must_process_pebs_in_order(cpuc)) {
+ handled = intel_pmu_handle_pebs_records_in_order(iregs, data,
+ base, top, setup_arch_pebs_sample_data);
+ } else {
+ handled = intel_pmu_handle_pebs_records(iregs, data,
+ base, top, setup_arch_pebs_sample_data);
+ }
+
+ return handled;
}
static void __init intel_arch_pebs_init(void)
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 12/15] perf/x86/intel: Make ACR static_call update architectural
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (10 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 11/15] perf/x86/intel: Fix invalid PEBS counts with counter-group support Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 13/15] perf/x86: Validate event type before topdown/mem-loads classification Dapeng Mi
` (2 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Auto counter reload (ACR) is an architectural feature, but the
intel_pmu_enable_acr_event() static_call update is currently done in a
model-specific way.
This is only necessary because older platforms such as PTL/WCL only
support ACR on E-core, which made the boot-time capability check on
P-core impossible in the generic init path. Later platforms such as
DMR/NVL support ACR on booting P-core and can detect it architecturally.
Make the intel_pmu_enable_acr_event() static_call update architectural
so ACR can be enabled correctly on supported CPUs without relying on
model-specific setup.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 27 +++++++++++++++++++++++++--
1 file changed, 25 insertions(+), 2 deletions(-)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 3fc3795534bf..0d360fba4b79 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -6365,6 +6365,25 @@ static inline void __intel_update_large_pebs_flags(struct pmu *pmu)
#define counter_mask(_gp, _fixed) ((_gp) | ((u64)(_fixed) << INTEL_PMC_IDX_FIXED))
+static bool intel_pmu_support_acr(void)
+{
+ unsigned int eax, ebx, ecx, edx;
+ union cpuid35_eax eax_0;
+ union cpuid35_ebx ebx_0;
+
+ if (!this_cpu_has(X86_FEATURE_ARCH_PERFMON_EXT))
+ return false;
+
+ cpuid(ARCH_PERFMON_EXT_LEAF, &eax_0.full, &ebx_0.full, &ecx, &edx);
+ if (eax_0.split.acr_subleaf) {
+ cpuid_count(ARCH_PERFMON_EXT_LEAF, ARCH_PERFMON_ACR_LEAF,
+ &eax, &ebx, &ecx, &edx);
+ return !!counter_mask(eax, ebx);
+ }
+
+ return false;
+}
+
static void update_pmu_cap_from_perfmonext(struct pmu *pmu)
{
unsigned int eax, ebx, ecx, edx;
@@ -8112,7 +8131,6 @@ static __always_inline void intel_pmu_init_pnc(struct pmu *pmu)
hybrid(pmu, event_constraints) = intel_pnc_event_constraints;
hybrid(pmu, pebs_constraints) = intel_pnc_pebs_event_constraints;
hybrid(pmu, extra_regs) = intel_pnc_extra_regs;
- static_call_update(intel_pmu_enable_acr_event, intel_pmu_enable_acr);
}
static __always_inline void intel_pmu_init_cyc(struct pmu *pmu)
@@ -8160,7 +8178,6 @@ static __always_inline void intel_pmu_init_arw(struct pmu *pmu)
hybrid(pmu, event_constraints) = intel_arw_event_constraints;
hybrid(pmu, pebs_constraints) = intel_dkt_pebs_event_constraints;
hybrid(pmu, extra_regs) = intel_arw_extra_regs;
- static_call_update(intel_pmu_enable_acr_event, intel_pmu_enable_acr);
}
__init int intel_pmu_init(void)
@@ -9127,6 +9144,12 @@ __init int intel_pmu_init(void)
if (!is_hybrid())
intel_update_pmu_caps(NULL);
+ if (x86_pmu.acr_cntr_mask64 || intel_pmu_support_acr()) {
+ static_call_update(intel_pmu_enable_acr_event,
+ intel_pmu_enable_acr);
+ pr_cont("Auto counter reload, ");
+ }
+
if (x86_pmu.arch_pebs) {
static_call_update(intel_pmu_disable_event_ext,
intel_pmu_disable_event_ext);
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 13/15] perf/x86: Validate event type before topdown/mem-loads classification
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (11 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 12/15] perf/x86/intel: Make ACR static_call update architectural Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 14/15] perf/core: Add event_caps dependency flags Dapeng Mi
2026-09-28 7:43 ` [PATCH 15/15] perf/x86/intel: Allow Topdown metrics with a non-leader slots event Dapeng Mi
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Software events can be grouped with x86 PMU events. While iterating an
x86 event group, a non-x86 event may be misclassified as a SLOTS,
topdown metric or mem-loads{,-aux} events if its config bits happen to
match the x86 masks.
Require is_x86_event() before classifying SLOTS, topdown metric or
mem-loads{,-aux} events to avoid false matches on software events.
Fixes: 7b2c05a15d29 ("perf/x86/intel: Generic support for hardware TopDown metrics")
Fixes: 61b985e3e775 ("perf/x86/intel: Add perf core PMU support for Sapphire Rapids")
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 8 ++++++--
arch/x86/events/perf_event.h | 14 ++++++++------
2 files changed, 14 insertions(+), 8 deletions(-)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 0d360fba4b79..6cc1462c9c82 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -4809,12 +4809,16 @@ static bool is_available_metric_event(struct perf_event *event)
static inline bool is_mem_loads_event(struct perf_event *event)
{
- return (event->attr.config & INTEL_ARCH_EVENT_MASK) == X86_CONFIG(.event=0xcd, .umask=0x01);
+ return is_x86_event(event) &&
+ ((event->attr.config & INTEL_ARCH_EVENT_MASK) ==
+ X86_CONFIG(.event = 0xcd, .umask = 0x01));
}
static inline bool is_mem_loads_aux_event(struct perf_event *event)
{
- return (event->attr.config & INTEL_ARCH_EVENT_MASK) == X86_CONFIG(.event=0x03, .umask=0x82);
+ return is_x86_event(event) &&
+ ((event->attr.config & INTEL_ARCH_EVENT_MASK) ==
+ X86_CONFIG(.event = 0x03, .umask = 0x82));
}
static inline bool require_mem_loads_aux_event(struct perf_event *event)
diff --git a/arch/x86/events/perf_event.h b/arch/x86/events/perf_event.h
index e274802ef062..55f9684152c9 100644
--- a/arch/x86/events/perf_event.h
+++ b/arch/x86/events/perf_event.h
@@ -96,18 +96,22 @@ static inline bool is_topdown_count(struct perf_event *event)
return event->hw.flags & PERF_X86_EVENT_TOPDOWN;
}
+int is_x86_event(struct perf_event *event);
+
static inline bool is_metric_event(struct perf_event *event)
{
u64 config = event->attr.config;
- return ((config & ARCH_PERFMON_EVENTSEL_EVENT) == 0) &&
- ((config & INTEL_ARCH_EVENT_MASK) >= INTEL_TD_METRIC_RETIRING) &&
- ((config & INTEL_ARCH_EVENT_MASK) <= INTEL_TD_METRIC_MAX);
+ return is_x86_event(event) &&
+ ((config & ARCH_PERFMON_EVENTSEL_EVENT) == 0) &&
+ ((config & INTEL_ARCH_EVENT_MASK) >= INTEL_TD_METRIC_RETIRING) &&
+ ((config & INTEL_ARCH_EVENT_MASK) <= INTEL_TD_METRIC_MAX);
}
static inline bool is_slots_event(struct perf_event *event)
{
- return (event->attr.config & INTEL_ARCH_EVENT_MASK) == INTEL_TD_SLOTS;
+ return is_x86_event(event) &&
+ (event->attr.config & INTEL_ARCH_EVENT_MASK) == INTEL_TD_SLOTS;
}
static inline bool is_topdown_event(struct perf_event *event)
@@ -115,8 +119,6 @@ static inline bool is_topdown_event(struct perf_event *event)
return is_metric_event(event) || is_slots_event(event);
}
-int is_x86_event(struct perf_event *event);
-
static inline bool check_leader_group(struct perf_event *leader, int flags)
{
return is_x86_event(leader) ? !!(leader->hw.flags & flags) : false;
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 14/15] perf/core: Add event_caps dependency flags
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (12 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 13/15] perf/x86: Validate event type before topdown/mem-loads classification Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
2026-09-28 7:43 ` [PATCH 15/15] perf/x86/intel: Allow Topdown metrics with a non-leader slots event Dapeng Mi
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Add two new event capability flags:
- PERF_EV_CAP_RELIED_ON: An event that other group events rely on. Must
not be set on a group leader.
- PERF_EV_CAP_RELIANT: An event that relies on a group event carrying
PERF_EV_CAP_RELIED_ON. Must not be combined with PERF_EV_CAP_RELIED_ON.
If the dependency is on the group leader, use PERF_EV_CAP_SIBLING
instead of PERF_EV_CAP_RELIANT.
When an event with PERF_EV_CAP_RELIED_ON is detached from the group, all
events in the group carrying PERF_EV_CAP_RELIANT can no longer function
and must be transitioned to the ERROR state.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
include/linux/perf_event.h | 8 ++++++
kernel/events/core.c | 52 +++++++++++++++++++++++++++++++++-----
2 files changed, 53 insertions(+), 7 deletions(-)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 7797ce207555..3031fe8dab37 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -706,11 +706,19 @@ typedef void (*perf_overflow_handler_t)(struct perf_event *,
* group it is scheduled out and moved into an unrecoverable ERROR state.
* PERF_EV_CAP_READ_SCOPE: A CPU event that can be read from any CPU of the
* PMU scope where it is active.
+ * PERF_EV_CAP_RELIED_ON: An event on which other group members depend.
+ * This flag must not be set on a group leader.
+ * PERF_EV_CAP_RELIANT: A member that depends on another group event carrying
+ * PERF_EV_CAP_RELIED_ON. It must not be combined with PERF_EV_CAP_RELIED_ON
+ * and must not be set on a group leader. If the dependency is on the group
+ * leader, use PERF_EV_CAP_SIBLING instead.
*/
#define PERF_EV_CAP_SOFTWARE BIT(0)
#define PERF_EV_CAP_READ_ACTIVE_PKG BIT(1)
#define PERF_EV_CAP_SIBLING BIT(2)
#define PERF_EV_CAP_READ_SCOPE BIT(3)
+#define PERF_EV_CAP_RELIED_ON BIT(4)
+#define PERF_EV_CAP_RELIANT BIT(5)
#define SWEVENT_HLIST_BITS 8
#define SWEVENT_HLIST_SIZE (1 << SWEVENT_HLIST_BITS)
diff --git a/kernel/events/core.c b/kernel/events/core.c
index de05df65ab3d..55db0ad8b93a 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -2349,13 +2349,12 @@ static void perf_promote_sibling_to_leader(struct perf_event *sibling,
int group_caps)
{
/*
- * Events that have PERF_EV_CAP_SIBLING require being part of
- * a group and cannot exist on their own, schedule them out
- * and move them into the ERROR state. Also see
- * _perf_event_enable(), it will not be able to recover this
- * ERROR state.
+ * Events with PERF_EV_CAP_SIBLING or PERF_EV_CAP_RELIANT
+ * must stay in a group or depend on another event in the
+ * group; otherwise schedule them out and move them to ERROR.
+ * _perf_event_enable() cannot recover from this state.
*/
- if (sibling->event_caps & PERF_EV_CAP_SIBLING)
+ if (sibling->event_caps & (PERF_EV_CAP_SIBLING | PERF_EV_CAP_RELIANT))
__event_disable(sibling, ctx, PERF_EVENT_STATE_ERROR);
sibling->group_leader = sibling;
@@ -2371,6 +2370,43 @@ static void perf_promote_sibling_to_leader(struct perf_event *sibling,
perf_event__header_size(sibling);
}
+static void perf_group_disable_reliants(struct perf_event *event)
+{
+ struct perf_event *leader = event->group_leader;
+ struct perf_event *sibling, *tmp;
+ struct perf_event_context *ctx = event->ctx;
+ u32 mask = PERF_EV_CAP_RELIED_ON | PERF_EV_CAP_RELIANT;
+
+ if (!(event->event_caps & PERF_EV_CAP_RELIED_ON))
+ return;
+
+ /* Group leader should not carry PERF_EV_CAP_RELIED_ON. */
+ WARN_ON_ONCE(leader->event_caps & PERF_EV_CAP_RELIED_ON);
+ /* Group leader should not carry PERF_EV_CAP_RELIANT. */
+ WARN_ON_ONCE(leader->event_caps & PERF_EV_CAP_RELIANT);
+
+ /*
+ * Disable all siblings that rely on this event. Without it,
+ * they cannot function and would otherwise fall into ERROR state.
+ */
+ list_for_each_entry_safe(sibling, tmp, &leader->sibling_list, sibling_list) {
+ if (!(sibling->event_caps & PERF_EV_CAP_RELIANT))
+ continue;
+
+ /*
+ * An event must not set both PERF_EV_CAP_RELIED_ON and
+ * PERF_EV_CAP_RELIANT simultaneously.
+ */
+ if (WARN_ON_ONCE((sibling->event_caps & mask) == mask))
+ continue;
+
+ list_del_init(&sibling->sibling_list);
+ leader->nr_siblings--;
+ leader->group_generation++;
+ perf_promote_sibling_to_leader(sibling, ctx, leader->group_caps);
+ }
+}
+
static void perf_group_detach(struct perf_event *event)
{
struct perf_event *leader = event->group_leader;
@@ -2393,6 +2429,7 @@ static void perf_group_detach(struct perf_event *event)
* If this is a sibling, remove it from its group.
*/
if (leader != event) {
+ perf_group_disable_reliants(event);
list_del_init(&event->sibling_list);
leader->nr_siblings--;
leader->group_generation++;
@@ -3305,7 +3342,8 @@ static void _perf_event_enable(struct perf_event *event)
/*
* Detached SIBLING events cannot leave ERROR state.
*/
- if (event->event_caps & PERF_EV_CAP_SIBLING &&
+ if (event->event_caps &
+ (PERF_EV_CAP_SIBLING | PERF_EV_CAP_RELIANT) &&
event->group_leader == event)
goto out;
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 15/15] perf/x86/intel: Allow Topdown metrics with a non-leader slots event
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
` (13 preceding siblings ...)
2026-09-28 7:43 ` [PATCH 14/15] perf/core: Add event_caps dependency flags Dapeng Mi
@ 2026-09-28 7:43 ` Dapeng Mi
14 siblings, 0 replies; 16+ messages in thread
From: Dapeng Mi @ 2026-09-28 7:43 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Ian Rogers, Adrian Hunter, Alexander Shishkin,
Andi Kleen, Eranian Stephane
Cc: linux-kernel, linux-perf-users, Dapeng Mi, Zide Chen,
Falcon Thomas, Xudong Hao, Dapeng Mi
Topdown metric events currently require the slots event to be the group
leader. That is unnecessarily strict. For metric counting, it is
sufficient that a valid slots event exists in the same group, regardless
of whether it is the leader.
Relax the validation in intel_pmu_hw_config() to accept a slots sibling
and validate metric constraints against that slots event. This enables
leader sampling for groups such as:
-e '{cycles:p,slots,topdown-retiring}:S'
Keep the existing ordering rule: the slots event must appear before all
metric events. Group validation is incremental, and intel_pmu_hw_config()
cannot determine whether the current event is the last event in the
group.
Also set the slots event with PERF_EV_CAP_RELIED_ON and metric events
with PERF_EV_CAP_RELIANT so that detaching the slots event will place
the dependent metric events into the ERROR state if slots event is not
the group leader.
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/events/intel/core.c | 37 +++++++++++++++++++++++++++++-------
1 file changed, 30 insertions(+), 7 deletions(-)
diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c
index 6cc1462c9c82..9025f29f625d 100644
--- a/arch/x86/events/intel/core.c
+++ b/arch/x86/events/intel/core.c
@@ -5300,30 +5300,53 @@ static int intel_pmu_hw_config(struct perf_event *event)
if (is_available_metric_event(event)) {
struct perf_event *leader = event->group_leader;
+ struct perf_event *slots = NULL;
+ struct perf_event *sibling;
/* The metric events don't support sampling. */
if (is_sampling_event(event))
return -EINVAL;
- /* The metric events require a slots group leader. */
- if (!is_slots_event(leader))
+ /*
+ * intel_pmu_hw_config() cannot tell whether the current
+ * event is the last one in the group. Require the slots
+ * event to appear before all metric events.
+ */
+ if (is_slots_event(leader)) {
+ slots = leader;
+ } else if (leader->nr_siblings) {
+ for_each_sibling_event(sibling, leader) {
+ if (is_slots_event(sibling)) {
+ slots = sibling;
+ break;
+ }
+ }
+ }
+
+ /* The metric events require a slots event. */
+ if (!slots)
return -EINVAL;
/*
- * The leader/SLOTS must not be a sampling event for
+ * The slots event must not be a sampling event for
* metric use; hardware requires it starts at 0 when used
* in conjunction with MSR_PERF_METRICS.
*/
- if (is_sampling_event(leader))
+ if (is_sampling_event(slots))
return -EINVAL;
- event->event_caps |= PERF_EV_CAP_SIBLING;
+ if (slots == leader) {
+ event->event_caps |= PERF_EV_CAP_SIBLING;
+ } else {
+ slots->event_caps |= PERF_EV_CAP_RELIED_ON;
+ event->event_caps |= PERF_EV_CAP_RELIANT;
+ }
/*
* Only once we have a METRICs sibling do we
* need TopDown magic.
*/
- leader->hw.flags |= PERF_X86_EVENT_TOPDOWN;
- event->hw.flags |= PERF_X86_EVENT_TOPDOWN;
+ slots->hw.flags |= PERF_X86_EVENT_TOPDOWN;
+ event->hw.flags |= PERF_X86_EVENT_TOPDOWN;
}
}
--
2.34.1
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-09-28 7:51 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 7:42 [PATCH 00/15] perf/x86: Fix sampling bugs and relax slots-event grouping Dapeng Mi
2026-09-28 7:42 ` [PATCH 01/15] perf/x86/intel: Guard leader sibling walk on nr_siblings Dapeng Mi
2026-09-28 7:42 ` [PATCH 02/15] perf/x86/intel: Reset active_fixed_ctrl_val on CPU teardown Dapeng Mi
2026-09-28 7:42 ` [PATCH 03/15] perf/x86/intel: Reset active_pebs_data_cfg " Dapeng Mi
2026-09-28 7:42 ` [PATCH 04/15] perf/x86/intel: Reset cached acr_cfg_b[] and cfg_c_val[] " Dapeng Mi
2026-09-28 7:42 ` [PATCH 05/15] perf/x86/intel: Pass correct PEBS counter mask to no-drain update path Dapeng Mi
2026-09-28 7:43 ` [PATCH 06/15] perf/x86/intel: Limit PEBS counter iteration to valid array bounds Dapeng Mi
2026-09-28 7:43 ` [PATCH 07/15] perf/x86/intel: Reject SAMPLE_READ for no-counter-snapshot ACR events Dapeng Mi
2026-09-28 7:43 ` [PATCH 08/15] perf/x86/intel: Fix stale PEBS count without counter-group support Dapeng Mi
2026-09-28 7:43 ` [PATCH 09/15] perf/x86/intel: Refactor intel_pmu_drain_arch_pebs() Dapeng Mi
2026-09-28 7:43 ` [PATCH 10/15] perf/x86/intel: Refactor intel_pmu_drain_pebs_icl() Dapeng Mi
2026-09-28 7:43 ` [PATCH 11/15] perf/x86/intel: Fix invalid PEBS counts with counter-group support Dapeng Mi
2026-09-28 7:43 ` [PATCH 12/15] perf/x86/intel: Make ACR static_call update architectural Dapeng Mi
2026-09-28 7:43 ` [PATCH 13/15] perf/x86: Validate event type before topdown/mem-loads classification Dapeng Mi
2026-09-28 7:43 ` [PATCH 14/15] perf/core: Add event_caps dependency flags Dapeng Mi
2026-09-28 7:43 ` [PATCH 15/15] perf/x86/intel: Allow Topdown metrics with a non-leader slots event Dapeng Mi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®