Re: [PATCH v4] tracing: Fix race between update_event_fields and, event_define_fields
From: Steven Rostedt
Date: Sat Aug 08 2026 - 22:13:09 EST
On Mon, 3 Aug 2026 17:40:55 +0800
Michael Wu <michael@xxxxxxxxxxxxxxxxx> wrote:
> event_define_fields() (pri=1 MODULE_STATE_COMING notifier, locked by
> event_mutex) populates class->fields via list_add(), while
> update_event_fields() (called from the pri=0 notifier path via
> trace_event_update_all) traverses class->fields protected only by
> trace_event_sem. These are two different locks guarding the same
> data structure, so during cross-module loading a reader on one CPU can
> observe partially initialized list nodes being concurrently added by a
> writer on another CPU.
>
> On arm64 with weak memory ordering, __list_add() writes to two
> different cache lines:
>
> next->prev = new; // (1) ordinary store
> new->next = next; // (2) ordinary store
> new->prev = prev; // (3) ordinary store
> WRITE_ONCE(prev->next, new); // (4) release store
>
> The store buffer can drain (2) and (4) independently since they target
> different cache lines. A remote CPU may observe (4) before (2): it
> sees prev->next pointing to the new node, but the new node's link.next
> is still zero (kmem_cache_alloc zero-initialized via KMEM_CACHE with
> SLAB_PANIC). Since offsetof(struct ftrace_event_field, link) == 0,
> list_for_each_entry() derives field == NULL from link.next == 0 and
> crashes at field->type (offset 0x18):
The above description is way too verbose. What exactly is the race?
Was the above written by AI? It looks like it .
I already tested it but when I went to write the log for Linus, I
realized this description isn't acceptable for the commit itself.
-- Steve