Re: [RFC] tracing: Try user copies with page faults disabled first
From: Usama Arif
Date: Fri Jul 17 2026 - 13:38:22 EST
On Wed, 15 Jul 2026 08:54:54 -0700 Usama Arif <usama.arif@xxxxxxxxx> wrote:
> trace_user_fault_read() is called with preemption disabled to copy user
> memory into a per-cpu scratch buffer. The existing implementation enables
> preemption around the copy because faulting user memory can sleep. That
> opens a window where another task can run on the same CPU and clobber the
> per-cpu buffer, so the copy is wrapped in a retry loop: sample
> nr_context_switches_cpu(), do the preempt-enabled copy, and retry if the
> counter changed. If this fails to complete 100 times, the function gives up
> with a warning.
>
> nr_context_switches_cpu() reads rq->nr_switches. That counter increments
> for every context switch on the CPU, not only for switches to tasks that
> use this tracing scratch buffer. On a heavily loaded system, unrelated
> scheduler activity can move the counter during every preempt-enabled copy
> attempt, exhaust the retry guard, and trigger the warning.
>
> This is showing up across the Meta fleet around 100 times a day since the
> kernel began upgrading to 7.1, mostly on arm servers:
>
> Error: Too many tries to read user space
> WARNING: kernel/trace/trace.c:6244 at trace_user_fault_read+0x284/0x2c8, CPU#28: Collection-18/677527
> CPU: 28 UID: 0 PID: 677527 Comm: Collection-18 Kdump: loaded Not tainted 7.1.0-.... #1 PREEMPTLAZY
> Hardware name: Quanta Java Island MP 29F0EMA08CH/Java Island, BIOS F0EJ3A16 03/12/2026
> Call trace:
> trace_user_fault_read+0x284/0x2c8 (P)
> syscall_get_data+0x144/0x2c0
> perf_syscall_enter+0xc0/0x2d8
> syscall_trace_enter+0x1a0/0x270
> do_el0_svc+0x54/0xb8
> el0_svc+0x44/0x268
> el0t_64_sync_handler+0x7c/0x120
> el0t_64_sync+0x17c/0x180
> ---[ end trace 0000000000000000 ]---
>
I think the actual problem might be:
https://lore.kernel.org/all/20260717173252.3431565-1-usama.arif@xxxxxxxxx/