mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Kunwu Chan <kunwu.chan@gmail.com>
To: peterz@infradead.org, mingo@redhat.com, acme@kernel.org,
	namhyung@kernel.org
Cc: sj@kernel.org, corbet@lwn.net, skhan@linuxfoundation.org,
	rdunlap@infradead.org, mark.rutland@arm.com,
	alexander.shishkin@linux.intel.com, jolsa@kernel.org,
	irogers@google.com, adrian.hunter@intel.com,
	james.clark@linaro.org, akpm@linux-foundation.org,
	lianux.mm@gmail.com, kunwu.chan@gmail.com,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-perf-users@vger.kernel.org,
	linux-kselftest@vger.kernel.org
Subject: [PATCH 1/5] perf/core: add AUX buffer ownership for kernel events
Date: Mon,  5 Oct 2026 01:34:53 +0800	[thread overview]
Message-ID: <20261004173458.837842-2-kunwu.chan@gmail.com> (raw)
In-Reply-To: <20261004173458.837842-1-kunwu.chan@gmail.com>

Add an in-kernel AUX owner reference and setup/release helpers for
kernel-created perf events that need an AUX buffer without a userspace
mmap.  The setup path validates that the event is a kernel event with
no parent, rejects non-power-of-two page counts and negative watermark
values, allocates the perf buffer and AUX pages, records the kernel
owner, and attaches the buffer under the same lock.

perf_event_release_aux() stops AUX writers, frees AUX storage, and
detaches the buffer, in that order, matching the AUX teardown ordering
of perf_mmap_close(): perf_pmu_output_stop() walks event->rb->event_list
and rb_free_aux() must run while the buffer is still referenced by the
event.  The single event->rb reference is dropped by
ring_buffer_attach(event, NULL) itself, so the release path does not
put it again.

Both functions run under event->mmap_mutex, the lock that already
serialises ring-buffer attach/detach transitions for an event:
perf_mmap(), perf_mmap_close(), _perf_event_set_output() and
_free_event() all hold it around ring_buffer_attach().  Concurrent
perf_event_release_aux() callers therefore go through the same
serialisation point, and a second release call, or a release of a
buffer this API does not own, is a no-op.

Keep aux_mmap_count dedicated to userspace mappings.
perf_aux_output_begin() accepts a writer while either a userspace or kernel
owner remains, preserving the existing teardown ordering.

Co-developed-by: Lian Wang <lianux.mm@gmail.com>
Signed-off-by: Lian Wang <lianux.mm@gmail.com>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
 include/linux/perf_event.h  |  11 +++
 kernel/events/core.c        | 164 ++++++++++++++++++++++++++++++++++++
 kernel/events/internal.h    |   6 ++
 kernel/events/ring_buffer.c |  11 ++-
 4 files changed, 188 insertions(+), 4 deletions(-)

diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 7797ce207555..78fd2ed11fcf 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -1255,6 +1255,17 @@ perf_event_create_kernel_counter(struct perf_event_attr *attr,
 				 perf_overflow_handler_t callback,
 				 void *context);
 
+/*
+ * AUX ring-buffer support for kernel-created events (no user mmap).
+ * perf_event_setup_aux() allocates the buffer the PMU writes into via
+ * perf_aux_output_begin()/perf_aux_output_end(); the paired
+ * perf_event_release_aux() must be called before
+ * perf_event_release_kernel().
+ */
+extern int perf_event_setup_aux(struct perf_event *event, int nr_pages,
+				long watermark);
+extern void perf_event_release_aux(struct perf_event *event);
+
 extern void perf_pmu_migrate_context(struct pmu *pmu,
 				     int src_cpu, int dst_cpu);
 extern int perf_event_read_local(struct perf_event *event, u64 *value,
diff --git a/kernel/events/core.c b/kernel/events/core.c
index de05df65ab3d..966e74645cc9 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -7139,6 +7139,170 @@ static void perf_mmap_close(struct vm_area_struct *vma)
 	ring_buffer_put(rb); /* could be last */
 }
 
+/*
+ * perf_event_setup_aux()/perf_event_release_aux() - AUX buffer support for
+ * kernel-created perf events (perf_event_create_kernel_counter()).
+ *
+ * perf_mmap()/perf_mmap_close() build and tear down AUX buffers for
+ * user-space events; kernel consumers (e.g. DAMON's ARM SPE backend) have no
+ * mmap, so they need this symmetric pair to get a buffer the PMU can write
+ * into via perf_aux_output_begin()/perf_aux_output_end().
+ *
+ * event->mmap_mutex is not an mmap-specific lock here: it is the lock that
+ * already serialises ring-buffer attach/detach transitions for an event
+ * (perf_mmap(), perf_mmap_close(), _perf_event_set_output() and
+ * _free_event() all hold it around ring_buffer_attach()).  Both functions
+ * below use it so that the attach and detach they perform are atomic with
+ * respect to those paths and to each other.
+ *
+ * The buffer holds an AUX owner reference (aux_kernel_count = 1) so
+ * perf_aux_output_begin() admits writers and so it is torn down here
+ * rather than by perf_mmap_close().  perf_event_release_aux() also detaches
+ * the event's normal ring-buffer reference before the event is released.
+ */
+
+/**
+ * perf_event_setup_aux() - Allocate an AUX ring buffer for an event.
+ * @event:	Kernel-created perf event (must not have an rb yet).
+ * @nr_pages:	AUX buffer size in pages; power of two, >= 1.
+ * @watermark:	AUX watermark; 0 selects the perf default (half the buffer).
+ *
+ * The buffer is non-overwrite streaming (RING_BUFFER_WRITABLE): the PMU
+ * pauses when the buffer is full until the consumer advances the tail.
+ * Must be called before the event is enabled.  The PMU's ->setup_aux()
+ * callback (e.g. arm_spe_pmu_setup_aux) validates the size and maps the
+ * pages for hardware writes.  The ring buffer is allocated with zero
+ * data pages, so PERF_RECORD_AUX events are not recorded; this is
+ * intentional for kernel consumers that drain trace data directly.
+ * Returns 0 on success, -errno otherwise.
+ */
+int perf_event_setup_aux(struct perf_event *event, int nr_pages,
+			 long watermark)
+{
+	struct perf_buffer *rb;
+	int ret;
+
+	if (!is_kernel_event(event) || event->parent ||
+	    !is_power_of_2(nr_pages) || watermark < 0)
+		return -EINVAL;
+
+	mutex_lock(&event->mmap_mutex);
+	if (event->rb) {
+		ret = -EBUSY;
+		goto out_unlock;
+	}
+
+	rb = rb_alloc(0, 0, event->cpu, 0);
+	if (!rb) {
+		ret = -ENOMEM;
+		goto out_unlock;
+	}
+
+	/* AUX area sits right after the user (control) page. */
+	ret = rb_alloc_aux(rb, event, 1, nr_pages, watermark,
+			   RING_BUFFER_WRITABLE);
+	if (ret)
+		goto err_put;
+
+	/*
+	 * No user-space mmap exists for this buffer; record an in-kernel
+	 * AUX owner (aux_kernel_count) instead so perf_aux_output_begin()
+	 * admits writers and perf_event_release_aux() performs the teardown
+	 * that perf_mmap_close() would otherwise do.
+	 */
+	refcount_set(&rb->aux_kernel_count, 1);
+
+	/* Transfer the allocation reference to the event. */
+	ring_buffer_attach(event, rb);
+	mutex_unlock(&event->mmap_mutex);
+
+	return 0;
+
+err_put:
+	ring_buffer_put(rb);
+out_unlock:
+	mutex_unlock(&event->mmap_mutex);
+	return ret;
+}
+EXPORT_SYMBOL_GPL(perf_event_setup_aux);
+
+/**
+ * perf_event_release_aux() - Tear down an AUX buffer for an event.
+ * @event:	Event with an AUX buffer allocated by perf_event_setup_aux().
+ *
+ * Stops any active AUX writers, frees the AUX pages, and detaches the ring
+ * buffer from the event.  Must be called before perf_event_release_kernel().
+ * Safe to call with a missing rb (no-op); a second call is also a no-op.
+ * A ring buffer not allocated by perf_event_setup_aux() is left attached.
+ *
+ * The whole teardown runs under event->mmap_mutex, the lock that already
+ * serialises ring-buffer attach/detach transitions for this event
+ * (perf_mmap(), perf_mmap_close() and _perf_event_set_output() hold it
+ * around ring_buffer_attach()): concurrent callers on the same event
+ * serialise here, so exactly one of them drops the single
+ * aux_kernel_count owner reference and performs the teardown, and the
+ * others then observe event->rb == NULL and return.  Holding it also
+ * keeps event->rb valid across perf_pmu_output_stop(), which walks
+ * event->rb->event_list, until the detach is done.
+ *
+ * The ordering matches the AUX teardown of perf_mmap_close(): the PMU
+ * output is stopped and the AUX pages freed before the buffer is
+ * detached from the event, because perf_pmu_output_stop() walks
+ * event->rb->event_list and rb_free_aux() must run while the buffer is
+ * still referenced by the event.
+ *
+ * The caller must quiesce AUX consumers before releasing the buffer;
+ * active AUX references at release time indicate a violated lifetime
+ * contract (see the aux_refcount WARN below).
+ */
+void perf_event_release_aux(struct perf_event *event)
+{
+	struct perf_buffer *rb;
+
+	if (!is_kernel_event(event) || event->parent)
+		return;
+
+	mutex_lock(&event->mmap_mutex);
+
+	rb = event->rb;
+	if (!rb)
+		goto out_unlock;
+
+	/* Only tear down a ring buffer that this API owns. */
+	if (!rb_has_aux(rb) || !refcount_read(&rb->aux_kernel_count))
+		goto out_unlock;
+
+	/*
+	 * Drop the kernel AUX-owner reference; writers are refused once it
+	 * reaches zero (perf_aux_output_begin()).  Since concurrent
+	 * releases serialise on mmap_mutex and setup rejects a second
+	 * buffer with -EBUSY, the count can only be 1 here.
+	 */
+	WARN_ON_ONCE(!refcount_dec_and_test(&rb->aux_kernel_count));
+
+	/*
+	 * Stop all AUX events writing to this buffer so the pages can be
+	 * freed; the aux_mutex serialisation matches perf_mmap_close().
+	 */
+	mutex_lock(&rb->aux_mutex);
+	perf_pmu_output_stop(event);
+	rb_free_aux(rb);
+	WARN_ON_ONCE(refcount_read(&rb->aux_refcount));
+	mutex_unlock(&rb->aux_mutex);
+
+	/*
+	 * Detach the buffer from the event.  ring_buffer_attach(event, NULL)
+	 * drops the event->rb reference itself (the allocation reference
+	 * transferred by perf_event_setup_aux()), so no extra
+	 * ring_buffer_put() is needed here.
+	 */
+	ring_buffer_attach(event, NULL);
+
+out_unlock:
+	mutex_unlock(&event->mmap_mutex);
+}
+EXPORT_SYMBOL_GPL(perf_event_release_aux);
+
 static vm_fault_t perf_mmap_pfn_mkwrite(struct vm_fault *vmf)
 {
 	/* The first page is the user control page, others are read-only. */
diff --git a/kernel/events/internal.h b/kernel/events/internal.h
index c03c4f2eea57..760b7659ad87 100644
--- a/kernel/events/internal.h
+++ b/kernel/events/internal.h
@@ -48,6 +48,12 @@ struct perf_buffer {
 	int				aux_nr_pages;
 	int				aux_overwrite;
 	refcount_t			aux_mmap_count;
+	/*
+	 * In-kernel AUX owner reference, set by perf_event_setup_aux():
+	 * admits writers and ties the AUX teardown to
+	 * perf_event_release_aux() when no userspace mmap holds the buffer.
+	 */
+	refcount_t			aux_kernel_count;
 	unsigned long			aux_mmap_locked;
 	void				(*free_aux)(void *);
 	refcount_t			aux_refcount;
diff --git a/kernel/events/ring_buffer.c b/kernel/events/ring_buffer.c
index 1b1ffe0533e5..eae28e42346c 100644
--- a/kernel/events/ring_buffer.c
+++ b/kernel/events/ring_buffer.c
@@ -395,14 +395,17 @@ void *perf_aux_output_begin(struct perf_output_handle *handle,
 		goto err;
 
 	/*
-	 * If aux_mmap_count is zero, the aux buffer is in perf_mmap_close(),
-	 * about to get freed, so we leave immediately.
+	 * If no AUX owner remains, the buffer is in perf_mmap_close() or
+	 * perf_event_release_aux(), about to get freed, so we leave
+	 * immediately.  aux_mmap_count tracks user-space mmap owners;
+	 * aux_kernel_count tracks in-kernel owners (perf_event_setup_aux()).
 	 *
-	 * Checking rb::aux_mmap_count and rb::refcount has to be done in
+	 * Checking the AUX owner counts and rb::refcount has to be done in
 	 * the same order, see perf_mmap_close. Otherwise we end up freeing
 	 * aux pages in this path, which is a bug, because in_atomic().
 	 */
-	if (!refcount_read(&rb->aux_mmap_count))
+	if (!refcount_read(&rb->aux_mmap_count) &&
+	    !refcount_read(&rb->aux_kernel_count))
 		goto err;
 
 	if (!refcount_inc_not_zero(&rb->aux_refcount))
-- 
2.43.0


  reply	other threads:[~2026-10-04 17:35 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04 17:34 [PATCH 0/5] perf/core: add AUX buffer kernel-consumer API Kunwu Chan
2026-10-04 17:34 ` Kunwu Chan [this message]
2026-10-04 17:34 ` [PATCH 2/5] perf/core: add AUX ring accessors for kernel consumers Kunwu Chan
2026-10-04 17:34 ` [PATCH 3/5] perf/core: add KUnit tests for AUX kernel-consumer API Kunwu Chan
2026-10-04 17:34 ` [PATCH 4/5] selftests/perf_events: add userspace AUX regression test Kunwu Chan
2026-10-04 17:34 ` [PATCH 5/5] selftests/perf_events: add AUX kernel API selftest script Kunwu Chan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261004173458.837842-2-kunwu.chan@gmail.com \
    --to=kunwu.chan@gmail.com \
    --cc=acme@kernel.org \
    --cc=adrian.hunter@intel.com \
    --cc=akpm@linux-foundation.org \
    --cc=alexander.shishkin@linux.intel.com \
    --cc=corbet@lwn.net \
    --cc=irogers@google.com \
    --cc=james.clark@linaro.org \
    --cc=jolsa@kernel.org \
    --cc=lianux.mm@gmail.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=mingo@redhat.com \
    --cc=namhyung@kernel.org \
    --cc=peterz@infradead.org \
    --cc=rdunlap@infradead.org \
    --cc=sj@kernel.org \
    --cc=skhan@linuxfoundation.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®