From: Kunwu Chan <kunwu.chan@gmail.com>
To: peterz@infradead.org, mingo@redhat.com, acme@kernel.org,
namhyung@kernel.org
Cc: sj@kernel.org, corbet@lwn.net, skhan@linuxfoundation.org,
rdunlap@infradead.org, mark.rutland@arm.com,
alexander.shishkin@linux.intel.com, jolsa@kernel.org,
irogers@google.com, adrian.hunter@intel.com,
james.clark@linaro.org, akpm@linux-foundation.org,
lianux.mm@gmail.com, kunwu.chan@gmail.com,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-perf-users@vger.kernel.org,
linux-kselftest@vger.kernel.org
Subject: [PATCH 1/5] perf/core: add AUX buffer ownership for kernel events
Date: Mon, 5 Oct 2026 01:34:53 +0800 [thread overview]
Message-ID: <20261004173458.837842-2-kunwu.chan@gmail.com> (raw)
In-Reply-To: <20261004173458.837842-1-kunwu.chan@gmail.com>
Add an in-kernel AUX owner reference and setup/release helpers for
kernel-created perf events that need an AUX buffer without a userspace
mmap. The setup path validates that the event is a kernel event with
no parent, rejects non-power-of-two page counts and negative watermark
values, allocates the perf buffer and AUX pages, records the kernel
owner, and attaches the buffer under the same lock.
perf_event_release_aux() stops AUX writers, frees AUX storage, and
detaches the buffer, in that order, matching the AUX teardown ordering
of perf_mmap_close(): perf_pmu_output_stop() walks event->rb->event_list
and rb_free_aux() must run while the buffer is still referenced by the
event. The single event->rb reference is dropped by
ring_buffer_attach(event, NULL) itself, so the release path does not
put it again.
Both functions run under event->mmap_mutex, the lock that already
serialises ring-buffer attach/detach transitions for an event:
perf_mmap(), perf_mmap_close(), _perf_event_set_output() and
_free_event() all hold it around ring_buffer_attach(). Concurrent
perf_event_release_aux() callers therefore go through the same
serialisation point, and a second release call, or a release of a
buffer this API does not own, is a no-op.
Keep aux_mmap_count dedicated to userspace mappings.
perf_aux_output_begin() accepts a writer while either a userspace or kernel
owner remains, preserving the existing teardown ordering.
Co-developed-by: Lian Wang <lianux.mm@gmail.com>
Signed-off-by: Lian Wang <lianux.mm@gmail.com>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
include/linux/perf_event.h | 11 +++
kernel/events/core.c | 164 ++++++++++++++++++++++++++++++++++++
kernel/events/internal.h | 6 ++
kernel/events/ring_buffer.c | 11 ++-
4 files changed, 188 insertions(+), 4 deletions(-)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 7797ce207555..78fd2ed11fcf 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -1255,6 +1255,17 @@ perf_event_create_kernel_counter(struct perf_event_attr *attr,
perf_overflow_handler_t callback,
void *context);
+/*
+ * AUX ring-buffer support for kernel-created events (no user mmap).
+ * perf_event_setup_aux() allocates the buffer the PMU writes into via
+ * perf_aux_output_begin()/perf_aux_output_end(); the paired
+ * perf_event_release_aux() must be called before
+ * perf_event_release_kernel().
+ */
+extern int perf_event_setup_aux(struct perf_event *event, int nr_pages,
+ long watermark);
+extern void perf_event_release_aux(struct perf_event *event);
+
extern void perf_pmu_migrate_context(struct pmu *pmu,
int src_cpu, int dst_cpu);
extern int perf_event_read_local(struct perf_event *event, u64 *value,
diff --git a/kernel/events/core.c b/kernel/events/core.c
index de05df65ab3d..966e74645cc9 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -7139,6 +7139,170 @@ static void perf_mmap_close(struct vm_area_struct *vma)
ring_buffer_put(rb); /* could be last */
}
+/*
+ * perf_event_setup_aux()/perf_event_release_aux() - AUX buffer support for
+ * kernel-created perf events (perf_event_create_kernel_counter()).
+ *
+ * perf_mmap()/perf_mmap_close() build and tear down AUX buffers for
+ * user-space events; kernel consumers (e.g. DAMON's ARM SPE backend) have no
+ * mmap, so they need this symmetric pair to get a buffer the PMU can write
+ * into via perf_aux_output_begin()/perf_aux_output_end().
+ *
+ * event->mmap_mutex is not an mmap-specific lock here: it is the lock that
+ * already serialises ring-buffer attach/detach transitions for an event
+ * (perf_mmap(), perf_mmap_close(), _perf_event_set_output() and
+ * _free_event() all hold it around ring_buffer_attach()). Both functions
+ * below use it so that the attach and detach they perform are atomic with
+ * respect to those paths and to each other.
+ *
+ * The buffer holds an AUX owner reference (aux_kernel_count = 1) so
+ * perf_aux_output_begin() admits writers and so it is torn down here
+ * rather than by perf_mmap_close(). perf_event_release_aux() also detaches
+ * the event's normal ring-buffer reference before the event is released.
+ */
+
+/**
+ * perf_event_setup_aux() - Allocate an AUX ring buffer for an event.
+ * @event: Kernel-created perf event (must not have an rb yet).
+ * @nr_pages: AUX buffer size in pages; power of two, >= 1.
+ * @watermark: AUX watermark; 0 selects the perf default (half the buffer).
+ *
+ * The buffer is non-overwrite streaming (RING_BUFFER_WRITABLE): the PMU
+ * pauses when the buffer is full until the consumer advances the tail.
+ * Must be called before the event is enabled. The PMU's ->setup_aux()
+ * callback (e.g. arm_spe_pmu_setup_aux) validates the size and maps the
+ * pages for hardware writes. The ring buffer is allocated with zero
+ * data pages, so PERF_RECORD_AUX events are not recorded; this is
+ * intentional for kernel consumers that drain trace data directly.
+ * Returns 0 on success, -errno otherwise.
+ */
+int perf_event_setup_aux(struct perf_event *event, int nr_pages,
+ long watermark)
+{
+ struct perf_buffer *rb;
+ int ret;
+
+ if (!is_kernel_event(event) || event->parent ||
+ !is_power_of_2(nr_pages) || watermark < 0)
+ return -EINVAL;
+
+ mutex_lock(&event->mmap_mutex);
+ if (event->rb) {
+ ret = -EBUSY;
+ goto out_unlock;
+ }
+
+ rb = rb_alloc(0, 0, event->cpu, 0);
+ if (!rb) {
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+
+ /* AUX area sits right after the user (control) page. */
+ ret = rb_alloc_aux(rb, event, 1, nr_pages, watermark,
+ RING_BUFFER_WRITABLE);
+ if (ret)
+ goto err_put;
+
+ /*
+ * No user-space mmap exists for this buffer; record an in-kernel
+ * AUX owner (aux_kernel_count) instead so perf_aux_output_begin()
+ * admits writers and perf_event_release_aux() performs the teardown
+ * that perf_mmap_close() would otherwise do.
+ */
+ refcount_set(&rb->aux_kernel_count, 1);
+
+ /* Transfer the allocation reference to the event. */
+ ring_buffer_attach(event, rb);
+ mutex_unlock(&event->mmap_mutex);
+
+ return 0;
+
+err_put:
+ ring_buffer_put(rb);
+out_unlock:
+ mutex_unlock(&event->mmap_mutex);
+ return ret;
+}
+EXPORT_SYMBOL_GPL(perf_event_setup_aux);
+
+/**
+ * perf_event_release_aux() - Tear down an AUX buffer for an event.
+ * @event: Event with an AUX buffer allocated by perf_event_setup_aux().
+ *
+ * Stops any active AUX writers, frees the AUX pages, and detaches the ring
+ * buffer from the event. Must be called before perf_event_release_kernel().
+ * Safe to call with a missing rb (no-op); a second call is also a no-op.
+ * A ring buffer not allocated by perf_event_setup_aux() is left attached.
+ *
+ * The whole teardown runs under event->mmap_mutex, the lock that already
+ * serialises ring-buffer attach/detach transitions for this event
+ * (perf_mmap(), perf_mmap_close() and _perf_event_set_output() hold it
+ * around ring_buffer_attach()): concurrent callers on the same event
+ * serialise here, so exactly one of them drops the single
+ * aux_kernel_count owner reference and performs the teardown, and the
+ * others then observe event->rb == NULL and return. Holding it also
+ * keeps event->rb valid across perf_pmu_output_stop(), which walks
+ * event->rb->event_list, until the detach is done.
+ *
+ * The ordering matches the AUX teardown of perf_mmap_close(): the PMU
+ * output is stopped and the AUX pages freed before the buffer is
+ * detached from the event, because perf_pmu_output_stop() walks
+ * event->rb->event_list and rb_free_aux() must run while the buffer is
+ * still referenced by the event.
+ *
+ * The caller must quiesce AUX consumers before releasing the buffer;
+ * active AUX references at release time indicate a violated lifetime
+ * contract (see the aux_refcount WARN below).
+ */
+void perf_event_release_aux(struct perf_event *event)
+{
+ struct perf_buffer *rb;
+
+ if (!is_kernel_event(event) || event->parent)
+ return;
+
+ mutex_lock(&event->mmap_mutex);
+
+ rb = event->rb;
+ if (!rb)
+ goto out_unlock;
+
+ /* Only tear down a ring buffer that this API owns. */
+ if (!rb_has_aux(rb) || !refcount_read(&rb->aux_kernel_count))
+ goto out_unlock;
+
+ /*
+ * Drop the kernel AUX-owner reference; writers are refused once it
+ * reaches zero (perf_aux_output_begin()). Since concurrent
+ * releases serialise on mmap_mutex and setup rejects a second
+ * buffer with -EBUSY, the count can only be 1 here.
+ */
+ WARN_ON_ONCE(!refcount_dec_and_test(&rb->aux_kernel_count));
+
+ /*
+ * Stop all AUX events writing to this buffer so the pages can be
+ * freed; the aux_mutex serialisation matches perf_mmap_close().
+ */
+ mutex_lock(&rb->aux_mutex);
+ perf_pmu_output_stop(event);
+ rb_free_aux(rb);
+ WARN_ON_ONCE(refcount_read(&rb->aux_refcount));
+ mutex_unlock(&rb->aux_mutex);
+
+ /*
+ * Detach the buffer from the event. ring_buffer_attach(event, NULL)
+ * drops the event->rb reference itself (the allocation reference
+ * transferred by perf_event_setup_aux()), so no extra
+ * ring_buffer_put() is needed here.
+ */
+ ring_buffer_attach(event, NULL);
+
+out_unlock:
+ mutex_unlock(&event->mmap_mutex);
+}
+EXPORT_SYMBOL_GPL(perf_event_release_aux);
+
static vm_fault_t perf_mmap_pfn_mkwrite(struct vm_fault *vmf)
{
/* The first page is the user control page, others are read-only. */
diff --git a/kernel/events/internal.h b/kernel/events/internal.h
index c03c4f2eea57..760b7659ad87 100644
--- a/kernel/events/internal.h
+++ b/kernel/events/internal.h
@@ -48,6 +48,12 @@ struct perf_buffer {
int aux_nr_pages;
int aux_overwrite;
refcount_t aux_mmap_count;
+ /*
+ * In-kernel AUX owner reference, set by perf_event_setup_aux():
+ * admits writers and ties the AUX teardown to
+ * perf_event_release_aux() when no userspace mmap holds the buffer.
+ */
+ refcount_t aux_kernel_count;
unsigned long aux_mmap_locked;
void (*free_aux)(void *);
refcount_t aux_refcount;
diff --git a/kernel/events/ring_buffer.c b/kernel/events/ring_buffer.c
index 1b1ffe0533e5..eae28e42346c 100644
--- a/kernel/events/ring_buffer.c
+++ b/kernel/events/ring_buffer.c
@@ -395,14 +395,17 @@ void *perf_aux_output_begin(struct perf_output_handle *handle,
goto err;
/*
- * If aux_mmap_count is zero, the aux buffer is in perf_mmap_close(),
- * about to get freed, so we leave immediately.
+ * If no AUX owner remains, the buffer is in perf_mmap_close() or
+ * perf_event_release_aux(), about to get freed, so we leave
+ * immediately. aux_mmap_count tracks user-space mmap owners;
+ * aux_kernel_count tracks in-kernel owners (perf_event_setup_aux()).
*
- * Checking rb::aux_mmap_count and rb::refcount has to be done in
+ * Checking the AUX owner counts and rb::refcount has to be done in
* the same order, see perf_mmap_close. Otherwise we end up freeing
* aux pages in this path, which is a bug, because in_atomic().
*/
- if (!refcount_read(&rb->aux_mmap_count))
+ if (!refcount_read(&rb->aux_mmap_count) &&
+ !refcount_read(&rb->aux_kernel_count))
goto err;
if (!refcount_inc_not_zero(&rb->aux_refcount))
--
2.43.0
next prev parent reply other threads:[~2026-10-04 17:35 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-04 17:34 [PATCH 0/5] perf/core: add AUX buffer kernel-consumer API Kunwu Chan
2026-10-04 17:34 ` Kunwu Chan [this message]
2026-10-04 17:34 ` [PATCH 2/5] perf/core: add AUX ring accessors for kernel consumers Kunwu Chan
2026-10-04 17:34 ` [PATCH 3/5] perf/core: add KUnit tests for AUX kernel-consumer API Kunwu Chan
2026-10-04 17:34 ` [PATCH 4/5] selftests/perf_events: add userspace AUX regression test Kunwu Chan
2026-10-04 17:34 ` [PATCH 5/5] selftests/perf_events: add AUX kernel API selftest script Kunwu Chan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261004173458.837842-2-kunwu.chan@gmail.com \
--to=kunwu.chan@gmail.com \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=akpm@linux-foundation.org \
--cc=alexander.shishkin@linux.intel.com \
--cc=corbet@lwn.net \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=lianux.mm@gmail.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
--cc=rdunlap@infradead.org \
--cc=sj@kernel.org \
--cc=skhan@linuxfoundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®