From: Namhyung Kim <namhyung@kernel.org>
To: Ian Rogers <irogers@google.com>
Cc: acme@kernel.org, howardchu95@gmail.com, adrian.hunter@intel.com,
james.clark@linaro.org, jolsa@kernel.org,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
mingo@redhat.com, peterz@infradead.org
Subject: Re: [PATCH v5 01/23] perf trace: Set the augmented arg header in the augmenters that omit it
Date: Wed, 23 Sep 2026 22:39:11 -0700 [thread overview]
Message-ID: <arS3f8bio-m1WNML@z2> (raw)
In-Reply-To: <aa751d26ad55f8a7d72db30dc23a67fae66e4b6b.1790145937.git.irogers@google.com>
On Wed, Sep 23, 2026 at 12:13:41AM -0700, Ian Rogers wrote:
> An augmented argument is a struct augmented_arg header, holding the
> length of the payload and an error code, followed by the payload. The
> augmenters build it in augmented_args_tmp, a single entry per-CPU array
> reused by every syscall on that CPU, so a field left unassigned holds
> whatever the previous syscall on that CPU put there rather than anything
> about this one.
>
> The string augmenters get both fields from augmented_arg__read_str(),
> and sys_enter_perf_event_open() bails out when its read fails, but
> sys_enter_connect(), sys_enter_sendto(), sys_enter_clock_nanosleep() and
> sys_enter_nanosleep() copy the payload and output the record without
> ever describing it, and augment_arg() sets the length but not the error.
>
> This has gone unnoticed because the beautifiers for those payloads take
> a fixed sized type and so read the value without consulting the header:
> syscall_arg__scnprintf_augmented_sockaddr(),
> syscall_arg__scnprintf_augmented_timespec() and
> syscall_arg__scnprintf_augmented_perf_event_attr() all cast
> augmented.args->value directly. Nothing has yet read a length that was
> never written.
>
> Describe the payload everywhere one is produced, so that a reader can
> bound it by its length.
>
> These reads can fail, and none of the four checked whether they had.
> Rather than claim a payload that was not read, report a length of zero
> and the error, so the record describes what it holds and leaves the
> scratch out of it. A reader that bounds the payload by the header then
> shows the pointer, as it does for a syscall with no augmentation at
> all, rather than another task's data.
>
> The error is set for the same reason the length is, that the header is
> in a buffer the next syscall on this CPU will reuse and so carries the
> previous one's value if it is not written, rather than because anything
> reads it yet: beauty.h still calls the field int_arg and no beautifier
> looks at it. augment_arg() reported a string it could not read as an
> empty one before this and still does.
>
> Assisted-by: Antigravity:gemini-3.1-pro
> Signed-off-by: Ian Rogers <irogers@google.com>
> ---
> .../bpf_skel/augmented_raw_syscalls.bpf.c | 53 ++++++++++++++++---
> 1 file changed, 47 insertions(+), 6 deletions(-)
>
> diff --git a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c b/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c
> index 3bc9e28a9b8a..dd3aa5bd910b 100644
> --- a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c
> +++ b/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c
> @@ -210,6 +210,7 @@ int sys_enter_connect(struct syscall_enter_args *args)
> const void *sockaddr_arg = (const void *)args->args[1];
> unsigned int socklen = args->args[2];
> unsigned int len = sizeof(u64) + sizeof(augmented_args->args); // the size + err in all 'augmented_arg' structs
> + int err;
>
> if (augmented_args == NULL)
> return 1; /* Failure: don't filter */
> @@ -217,9 +218,16 @@ int sys_enter_connect(struct syscall_enter_args *args)
> _Static_assert(is_power_of_2(sizeof(augmented_args->arg.saddr)), "sizeof(augmented_args->arg.saddr) needs to be a power of two");
> socklen &= sizeof(augmented_args->arg.saddr) - 1;
>
> - bpf_probe_read_user(&augmented_args->arg.saddr, socklen, sockaddr_arg);
> + err = bpf_probe_read_user(&augmented_args->arg.saddr, socklen, sockaddr_arg);
> + /*
> + * A failed read leaves the scratch holding whatever the previous
> + * syscall on this CPU put there, so say there is no payload and why,
> + * rather than describing another task's data as this task's sockaddr.
> + */
> + if (err < 0)
> + socklen = 0;
> augmented_args->arg.size = socklen;
> - augmented_args->arg.err = 0;
> + augmented_args->arg.err = err < 0 ? err : 0;
It seems we can simply use the return value of bpf_probe_read_user()
as it only returns 0 or a negative error according to the doc.
https://docs.ebpf.io/linux/helper-function/bpf_probe_read_user/
>
> return augmented__output(args, augmented_args, len + socklen);
> }
> @@ -231,13 +239,19 @@ int sys_enter_sendto(struct syscall_enter_args *args)
> const void *sockaddr_arg = (const void *)args->args[4];
> unsigned int socklen = args->args[5];
> unsigned int len = sizeof(u64) + sizeof(augmented_args->args); // the size + err in all 'augmented_arg' structs
> + int err;
>
> if (augmented_args == NULL)
> return 1; /* Failure: don't filter */
>
> socklen &= sizeof(augmented_args->arg.saddr) - 1;
>
> - bpf_probe_read_user(&augmented_args->arg.saddr, socklen, sockaddr_arg);
> + err = bpf_probe_read_user(&augmented_args->arg.saddr, socklen, sockaddr_arg);
> + /* As in sys_enter_connect(), do not describe scratch as a sockaddr. */
> + if (err < 0)
> + socklen = 0;
> + augmented_args->arg.size = socklen;
> + augmented_args->arg.err = err < 0 ? err : 0;
>
> return augmented__output(args, augmented_args, len + socklen);
> }
> @@ -372,6 +386,9 @@ int sys_enter_perf_event_open(struct syscall_enter_args *args)
> if (bpf_probe_read_user(&augmented_args->arg.value, size, attr) < 0)
> goto failure;
>
> + augmented_args->arg.size = size;
> + augmented_args->arg.err = 0;
> +
> return augmented__output(args, augmented_args, len + size);
> failure:
> return 1; /* Failure: don't filter */
> @@ -384,6 +401,7 @@ int sys_enter_clock_nanosleep(struct syscall_enter_args *args)
> const void *rqtp_arg = (const void *)args->args[2];
> unsigned int len = sizeof(u64) + sizeof(augmented_args->args); // the size + err in all 'augmented_arg' structs
> __u32 size = sizeof(struct timespec64);
> + int err;
>
> if (augmented_args == NULL)
> goto failure;
> @@ -391,7 +409,12 @@ int sys_enter_clock_nanosleep(struct syscall_enter_args *args)
> if (size > sizeof(augmented_args->arg.value))
> goto failure;
>
> - bpf_probe_read_user(&augmented_args->arg.value, size, rqtp_arg);
> + err = bpf_probe_read_user(&augmented_args->arg.value, size, rqtp_arg);
> + /* As in sys_enter_connect(), do not describe scratch as a timespec. */
> + if (err < 0)
> + size = 0;
> + augmented_args->arg.size = size;
> + augmented_args->arg.err = err < 0 ? err : 0;
>
> return augmented__output(args, augmented_args, len + size);
> failure:
> @@ -405,6 +428,7 @@ int sys_enter_nanosleep(struct syscall_enter_args *args)
> const void *req_arg = (const void *)args->args[0];
> unsigned int len = sizeof(augmented_args->args);
> __u32 size = sizeof(struct timespec64);
> + int err;
>
> if (augmented_args == NULL)
> goto failure;
> @@ -412,7 +436,12 @@ int sys_enter_nanosleep(struct syscall_enter_args *args)
> if (size > sizeof(augmented_args->arg.value))
> goto failure;
>
> - bpf_probe_read_user(&augmented_args->arg.value, size, req_arg);
> + err = bpf_probe_read_user(&augmented_args->arg.value, size, req_arg);
> + /* As in sys_enter_connect(), do not describe scratch as a timespec. */
> + if (err < 0)
> + size = 0;
> + augmented_args->arg.size = size;
> + augmented_args->arg.err = err < 0 ? err : 0;
>
> return augmented__output(args, augmented_args, len + size);
> failure:
> @@ -445,6 +474,7 @@ static inline int augment_arg(struct syscall_enter_args *args, int i,
> struct beauty_payload_enter *payload, u64 offset)
> {
> int index, value_size = sizeof(struct augmented_arg) - offsetof(struct augmented_arg, value);
> + int read_err = 0;
> struct augmented_arg *payload_offset;
> s64 aug_size, size;
> bool augmented;
> @@ -467,8 +497,18 @@ static inline int augment_arg(struct syscall_enter_args *args, int i,
> if (size == 1) { /* string */
> aug_size = bpf_probe_read_user_str(payload_offset->value, value_size, arg);
> /* minimum of 0 to pass the verifier */
> - if (aug_size < 0)
> + if (aug_size < 0) {
> + /*
> + * Record why nothing was read. The header sits in
> + * scratch that the next syscall on this CPU reuses,
> + * so an error left unwritten is the previous one's.
> + * No beautifier reads it yet, beauty.h still calls
> + * the field int_arg, so the string is still shown
> + * as an empty one.
> + */
Do we really need this comment? Looks too verbose.
Thanks,
Namhyung
> + read_err = aug_size;
> aug_size = 0;
> + }
>
> augmented = true;
> } else if (size > 0 && size <= value_size) { /* struct */
> @@ -498,6 +538,7 @@ static inline int augment_arg(struct syscall_enter_args *args, int i,
> return -1;
>
> payload_offset->size = aug_size;
> + payload_offset->err = read_err;
> return written;
> }
>
> --
> 2.56.0.rc1.315.gc6ed9934b7-goog
>
next prev parent reply other threads:[~2026-09-24 5:39 UTC|newest]
Thread overview: 96+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 6:42 [PATCH v1 00/13] perf trace: Fix BPF filtering and make tracing tests non-exclusive Ian Rogers
2026-09-17 6:42 ` [PATCH v1 01/13] perf trace: Start BPF summary before starting workload Ian Rogers
2026-09-17 6:42 ` [PATCH v1 02/13] perf trace: Skip internal tracepoint fields in formatting and beauty map Ian Rogers
2026-09-17 6:42 ` [PATCH v1 03/13] perf trace: Do not set unaugmented BPF program on sys_exit map Ian Rogers
2026-09-17 6:42 ` [PATCH v1 04/13] perf trace: Filter events in BPF and avoid tracepoint vetoes Ian Rogers
2026-09-17 6:42 ` [PATCH v1 05/13] perf trace: Handle fork and exit directly in BPF filter maps Ian Rogers
2026-09-17 6:42 ` [PATCH v1 06/13] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive Ian Rogers
2026-09-17 6:42 ` [PATCH v1 07/13] perf test common: Do not globally disable tracing events in clear_all_probes Ian Rogers
2026-09-17 6:42 ` [PATCH v1 08/13] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive Ian Rogers
2026-09-17 6:42 ` [PATCH v1 09/13] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, " Ian Rogers
2026-09-17 6:42 ` [PATCH v1 10/13] perf test trace_summary: Improve error diagnostics Ian Rogers
2026-09-17 6:42 ` [PATCH v1 11/13] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive Ian Rogers
2026-09-17 6:42 ` [PATCH v1 12/13] perf test trace_summary: Make non-exclusive Ian Rogers
2026-09-17 6:42 ` [PATCH v1 13/13] perf test uprobe_from_different_cu: Scope probe name to PID Ian Rogers
2026-09-17 16:38 ` [PATCH v2 00/14] perf trace: Fix BPF filtering and make tracing tests non-exclusive Ian Rogers
2026-09-17 16:38 ` [PATCH v2 01/14] perf trace: Include the headers declaring pid_t and strcmp Ian Rogers
2026-09-17 16:38 ` [PATCH v2 02/14] perf trace: Start BPF summary before starting workload Ian Rogers
2026-09-17 16:38 ` [PATCH v2 03/14] perf trace: Skip internal tracepoint fields in formatting and beauty map Ian Rogers
2026-09-17 16:38 ` [PATCH v2 04/14] perf trace: Do not set unaugmented BPF program on sys_exit map Ian Rogers
2026-09-17 16:38 ` [PATCH v2 05/14] perf trace: Filter events in BPF and avoid tracepoint vetoes Ian Rogers
2026-09-17 16:38 ` [PATCH v2 06/14] perf trace: Handle fork and exit directly in BPF filter maps Ian Rogers
2026-09-17 16:38 ` [PATCH v2 07/14] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive Ian Rogers
2026-09-17 16:38 ` [PATCH v2 08/14] perf test common: Do not globally disable tracing events in clear_all_probes Ian Rogers
2026-09-17 16:38 ` [PATCH v2 09/14] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive Ian Rogers
2026-09-17 16:38 ` [PATCH v2 10/14] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, " Ian Rogers
2026-09-17 16:38 ` [PATCH v2 11/14] perf test trace_summary: Improve error diagnostics Ian Rogers
2026-09-17 16:39 ` [PATCH v2 12/14] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive Ian Rogers
2026-09-17 16:39 ` [PATCH v2 13/14] perf test trace_summary: Make non-exclusive Ian Rogers
2026-09-17 16:39 ` [PATCH v2 14/14] perf test uprobe_from_different_cu: Scope probe name to PID Ian Rogers
2026-09-18 14:06 ` [PATCH v3 00/16] perf trace: Fix BPF filtering and make tracing tests non-exclusive Ian Rogers
2026-09-18 14:06 ` [PATCH v3 01/16] perf trace: Include the headers declaring pid_t, strcmp and assert Ian Rogers
2026-09-18 14:06 ` [PATCH v3 02/16] perf trace: Free the whole evsel_trace in evsel__put_and_free_priv Ian Rogers
2026-09-18 14:06 ` [PATCH v3 03/16] perf trace: Start BPF summary before starting workload Ian Rogers
2026-09-18 14:06 ` [PATCH v3 04/16] perf trace: Skip internal tracepoint fields in formatting and beauty map Ian Rogers
2026-09-18 14:06 ` [PATCH v3 05/16] perf trace: Do not set unaugmented BPF program on sys_exit map Ian Rogers
2026-09-18 14:06 ` [PATCH v3 06/16] perf trace: Filter events in BPF and avoid tracepoint vetoes Ian Rogers
2026-09-18 14:06 ` [PATCH v3 07/16] perf trace: Handle fork and exit directly in BPF filter maps Ian Rogers
2026-09-18 14:06 ` [PATCH v3 08/16] perf trace: Enumerate the target again once BPF is attached Ian Rogers
2026-09-18 14:06 ` [PATCH v3 09/16] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive Ian Rogers
2026-09-18 14:06 ` [PATCH v3 10/16] perf test common: Only disable probes in clear_all_probes Ian Rogers
2026-09-18 14:06 ` [PATCH v3 11/16] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive Ian Rogers
2026-09-18 14:06 ` [PATCH v3 12/16] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, " Ian Rogers
2026-09-18 14:06 ` [PATCH v3 13/16] perf test trace_summary: Improve error diagnostics Ian Rogers
2026-09-18 14:06 ` [PATCH v3 14/16] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive Ian Rogers
2026-09-18 14:06 ` [PATCH v3 15/16] perf test trace_summary: Make non-exclusive Ian Rogers
2026-09-18 14:06 ` [PATCH v3 16/16] perf test uprobe_from_different_cu: Scope probe name to PID Ian Rogers
2026-09-18 21:19 ` [PATCH v4 00/18] perf trace: Fix BPF filtering and make tracing tests non-exclusive Ian Rogers
2026-09-18 21:19 ` [PATCH v4 01/18] perf trace: Include the headers declaring pid_t, strcmp and assert Ian Rogers
2026-09-18 21:19 ` [PATCH v4 02/18] perf trace: Free the whole evsel_trace in evsel__put_and_free_priv Ian Rogers
2026-09-18 21:19 ` [PATCH v4 03/18] perf evsel: Report an allocation failure as ENOMEM when setting filters Ian Rogers
2026-09-18 21:19 ` [PATCH v4 04/18] perf trace: Start BPF summary before starting workload Ian Rogers
2026-09-18 21:19 ` [PATCH v4 05/18] perf trace: Skip internal tracepoint fields in formatting and beauty map Ian Rogers
2026-09-18 21:19 ` [PATCH v4 06/18] perf trace: Do not set unaugmented BPF program on sys_exit map Ian Rogers
2026-09-18 21:19 ` [PATCH v4 07/18] perf trace: Filter events in BPF and avoid tracepoint vetoes Ian Rogers
2026-09-18 21:19 ` [PATCH v4 08/18] perf trace: Handle fork and exit directly in BPF filter maps Ian Rogers
2026-09-18 21:19 ` [PATCH v4 09/18] perf trace: Enumerate the target again once BPF is attached Ian Rogers
2026-09-18 21:19 ` [PATCH v4 10/18] perf trace: Drop targets that died before they were filtered Ian Rogers
2026-09-18 21:19 ` [PATCH v4 11/18] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive Ian Rogers
2026-09-18 21:19 ` [PATCH v4 12/18] perf test common: Only disable probes in clear_all_probes Ian Rogers
2026-09-18 21:19 ` [PATCH v4 13/18] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive Ian Rogers
2026-09-18 21:19 ` [PATCH v4 14/18] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, " Ian Rogers
2026-09-18 21:19 ` [PATCH v4 15/18] perf test trace_summary: Improve error diagnostics Ian Rogers
2026-09-18 21:19 ` [PATCH v4 16/18] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive Ian Rogers
2026-09-18 21:19 ` [PATCH v4 17/18] perf test trace_summary: Make non-exclusive Ian Rogers
2026-09-18 21:19 ` [PATCH v4 18/18] perf test uprobe_from_different_cu: Scope probe name to PID Ian Rogers
2026-09-22 13:11 ` [PATCH v4 00/18] perf trace: Fix BPF filtering and make tracing tests non-exclusive Arnaldo Carvalho de Melo
2026-09-22 22:43 ` Ian Rogers
2026-09-23 7:13 ` [PATCH v5 00/23] " Ian Rogers
2026-09-23 7:13 ` [PATCH v5 01/23] perf trace: Set the augmented arg header in the augmenters that omit it Ian Rogers
2026-09-24 5:39 ` Namhyung Kim [this message]
2026-09-23 7:13 ` [PATCH v5 02/23] perf trace: Include the augmented arg header in nanosleep's payload length Ian Rogers
2026-09-23 7:13 ` [PATCH v5 03/23] perf trace: Include the headers declaring pid_t, strcmp and assert Ian Rogers
2026-09-24 5:44 ` Namhyung Kim
2026-09-23 7:13 ` [PATCH v5 04/23] perf trace: Free the whole evsel_trace in evsel__put_and_free_priv Ian Rogers
2026-09-23 7:13 ` [PATCH v5 05/23] perf evsel: Report an allocation failure as ENOMEM when setting filters Ian Rogers
2026-09-23 7:13 ` [PATCH v5 06/23] perf trace: Start BPF summary before starting workload Ian Rogers
2026-09-23 7:13 ` [PATCH v5 07/23] perf trace: Skip internal tracepoint fields in formatting and beauty map Ian Rogers
2026-09-24 6:44 ` Namhyung Kim
2026-09-23 7:13 ` [PATCH v5 08/23] perf trace: Bounds check augmented arguments before reading them Ian Rogers
2026-09-24 6:58 ` Namhyung Kim
2026-09-23 7:13 ` [PATCH v5 09/23] perf trace: Bound the fixed size augmented argument beautifiers Ian Rogers
2026-09-23 7:13 ` [PATCH v5 10/23] perf trace: Do not read sample padding as an augmented argument Ian Rogers
2026-09-23 7:13 ` [PATCH v5 11/23] perf trace: Do not set unaugmented BPF program on sys_exit map Ian Rogers
2026-09-23 7:13 ` [PATCH v5 12/23] perf trace: Filter events in BPF and avoid tracepoint vetoes Ian Rogers
2026-09-23 7:13 ` [PATCH v5 13/23] perf trace: Handle fork and exit directly in BPF filter maps Ian Rogers
2026-09-23 7:13 ` [PATCH v5 14/23] perf trace: Enumerate the target again once BPF is attached Ian Rogers
2026-09-23 7:13 ` [PATCH v5 15/23] perf trace: Drop targets that died before they were filtered Ian Rogers
2026-09-23 7:13 ` [PATCH v5 16/23] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive Ian Rogers
2026-09-23 7:13 ` [PATCH v5 17/23] perf test common: Only disable probes in clear_all_probes Ian Rogers
2026-09-23 7:13 ` [PATCH v5 18/23] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive Ian Rogers
2026-09-23 7:13 ` [PATCH v5 19/23] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, " Ian Rogers
2026-09-23 7:14 ` [PATCH v5 20/23] perf test trace_summary: Improve error diagnostics Ian Rogers
2026-09-23 7:14 ` [PATCH v5 21/23] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive Ian Rogers
2026-09-23 7:14 ` [PATCH v5 22/23] perf test trace_summary: Make non-exclusive Ian Rogers
2026-09-23 7:14 ` [PATCH v5 23/23] perf test uprobe_from_different_cu: Scope probe name to PID Ian Rogers
2026-09-24 7:05 ` [PATCH v5 00/23] perf trace: Fix BPF filtering and make tracing tests non-exclusive Namhyung Kim
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arS3f8bio-m1WNML@z2 \
--to=namhyung@kernel.org \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=howardchu95@gmail.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®