Re: [PATCH v6 06/26] perf trace: Copy sockaddr arguments by their length
From: Ian Rogers
Date: Thu Oct 01 2026 - 02:01:48 EST
On Wed, Sep 30, 2026 at 4:08 PM Namhyung Kim <namhyung@xxxxxxxxxx> wrote:
>
> On Wed, Sep 30, 2026 at 09:14:42PM +0200, Arnaldo Carvalho de Melo wrote:
> > On Mon, Sep 28, 2026 at 11:25:45AM -0700, Ian Rogers wrote:
> > > The BTF augmenter copies sizeof(struct sockaddr), 16 bytes, for the
> > > sockaddr arguments of connect, bind and sendto, and runs before the
> > > sys_enter_connect and sys_enter_sendto programs that copy their length.
> > > An IPv6 address needs 24 or 28 bytes, so the rest was read from past the
> > > payload, and now that the printers are bounded only its family is shown.
> > >
> > > Copy them as buffers sized by the length argument after them, up to the
> > > 32 bytes augmented buffers are limited to. That holds an IPv6 address
> > > and doubles the AF_LOCAL path shown.
> > >
> > > Fixes: a68fd6a6cdd3 ("perf trace: Collect augmented data using BPF")
> > > Assisted-by: Antigravity:gemini-3.1-pro
> > > Signed-off-by: Ian Rogers <irogers@xxxxxxxxxx>
> > > ---
> > > tools/perf/builtin-trace.c | 7 ++++++-
> > > 1 file changed, 6 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c
> > > index 85db74965280..f94745a60f4a 100644
> > > --- a/tools/perf/builtin-trace.c
> > > +++ b/tools/perf/builtin-trace.c
> > > @@ -4159,7 +4159,12 @@ static int trace__bpf_sys_enter_beauty_map(struct trace *trace, int e_machine, i
> > > continue;
> > >
> > > bt = sc->arg_fmt[i].type;
> > > - beauty_array[i] = bt->size;
> > > + /* Copy a sockaddr as a buffer sized by the next argument, e.g. addrlen. */
> > > + if (strcmp(name, "sockaddr") == 0 && field->next &&
> > > + strstr(field->next->name, "len"))
> > > + beauty_array[i] = -((i + 1) + 1);
> > > + else
> > > + beauty_array[i] = bt->size;
> >
> >
> > Humm, I thought that this would be called in BPF handlers like:
> >
> > SEC("tp/syscalls/sys_enter_connect")
> > int sys_enter_connect(struct syscall_enter_args *args)
> > {
> > struct augmented_args_payload *augmented_args = augmented_args_payload();
> > const void *sockaddr_arg = (const void *)args->args[1];
> > unsigned int socklen = args->args[2];
> > unsigned int len = sizeof(u64) + sizeof(augmented_args->args); // the size + err in all 'augmented_arg' structs
> >
> > if (augmented_args == NULL)
> > return 1; /* Failure: don't filter */
> >
> > _Static_assert(is_power_of_2(sizeof(augmented_args->arg.saddr)), "sizeof(augmented_args->arg.saddr) needs to be a power of two");
> > socklen &= sizeof(augmented_args->arg.saddr) - 1;
> >
> > bpf_probe_read_user(&augmented_args->arg.saddr, socklen, sockaddr_arg);
> > augmented_args->arg.size = socklen;
> > augmented_args->arg.err = 0;
> >
> > return augmented__output(args, augmented_args, len + socklen);
> > }
> >
> > And it knows how many bytes to read by looking at socklen
> > (args->args[2]), i.e. not use the generic BPF handler that uses this
> > beauty_array, because knowing how many bytes to read in this case is
> > dynamic, varies with each syscall, according to one of its arguments :-\
> >
> > What am I missing?
>
> I think Ian's patch update the beauty map which is used by
> augment_sys_enter() before tail-calling syscall-specific functions.
>
> It'd be great if we cover all syscalls in the BPF skeleton and switch
> to the beauty-map and discard the functions.
Right, as Namhyung says, sys_enter() runs augment_sys_enter() before
tail-calling syscalls_sys_enter, and only falls back to the tail call
if augment_sys_enter() returns non-zero:
if (augment_sys_enter(args, &augmented_args->args))
bpf_tail_call(args, &syscalls_sys_enter, augmented_args->args.syscall_nr);
trace__init_syscalls_bpf_prog_array_maps() populates beauty_map_enter
for every enabled syscall, even ones with a dedicated sys_enter_*
program. In trace__bpf_sys_enter_beauty_map(), connect, bind and
sendto's "struct sockaddr *" argument matched the struct branch, which
looked up "struct sockaddr" in BTF and put bt->size (16 bytes) in
beauty_array. So whenever BTF was available, augment_sys_enter()
handled connect, bind and sendto by copying 16 bytes and returned 0,
and sys_enter_connect / sys_enter_sendto were never tail-called.
beauty_array already supports dynamic lengths from another syscall
argument: a negative entry -(j + 1) tells augment_arg() to read
args->args[j] bytes (clamped to TRACE_AUG_MAX_BUF = 32), which was
added for buffer arguments like write's buf/count. This patch uses
that encoding for struct sockaddr * when the next argument is its
length.
If you'd rather have connect, bind and sendto fall back to
sys_enter_connect and sys_enter_sendto (which copies up to
sizeof(struct sockaddr_storage) = 128 bytes, so AF_LOCAL paths over 30
bytes aren't truncated), we could instead skip "struct sockaddr" in
trace__bpf_sys_enter_beauty_map().
On Namhyung's point: all 9 syscalls with dedicated sys_enter_*
programs (connect, sendto, open, openat, rename, renameat2,
clock_nanosleep, nanosleep, perf_event_open) already get
beauty_map_enter entries when BTF is loaded, so the tail-call programs
and trace__find_usable_bpf_prog_entry() are mostly shadowed today. The
main gaps before removing them are:
1. trace__bpf_sys_enter_beauty_map() bails out early if trace->btf is
NULL, even for strings/buffers that don't need BTF.
2. TRACE_AUG_MAX_BUF is 32 bytes vs sizeof(struct sockaddr_storage)
(128 bytes) for AF_LOCAL sockaddrs.
3. perf_event_open reads attr->size from userspace in
sys_enter_perf_event_open if the BTF-sized read of perf_event_attr
faults on an older, smaller struct.
That seems like a good follow-up cleanup.
Thanks,
Ian