Re: [Patch v4 01/22] sched/cache: Introduce infrastructure for cache-aware load balancing

From: Tim Chen

Date: Mon Sep 14 2026 - 19:14:39 EST


On Tue, 2026-09-15 at 01:50 +0800, Zenghui Yu wrote:
>

[snip]

> I sporadically hit the SLUB "Poison overwritten" reports on the mm_struct
> cache while running mm-new:
>
> [Poison overwritten] 0xffff8001076ec8e8-0xffff8001076ec8eb @offset=51432. First byte 0xff instead of 0x6b
> =============================================================================
> BUG mm_struct (Tainted: G N ): Object corrupt
> -----------------------------------------------------------------------------
>
> Allocated in copy_process+0x1e48/0x2078 age=2 cpu=7 pid=11866
> copy_process+0x1e48/0x2078
> kernel_clone+0xa4/0x498
> __do_sys_clone+0x5c/0x88
> __arm64_sys_clone+0x1c/0x28
> invoke_syscall+0x54/0x110
> el0_svc_common.constprop.0+0x40/0xe0
> do_el0_svc+0x1c/0x28
> el0_svc+0x54/0x424
> el0t_64_sync_handler+0xa0/0xe4
> el0t_64_sync+0x1b0/0x1b4
> Freed in __mmdrop+0x108/0x180 age=2 cpu=3 pid=11955
> kmem_cache_free+0x290/0x53c
> __mmdrop+0x108/0x180
> __mmput+0x150/0x154
> mmput+0x50/0x5c
> exec_mm_put_old+0x74/0x84
> setup_new_exec+0x7c/0x90
> load_elf_binary+0x4b0/0x1914
> bprm_execve+0x300/0x83c
> do_execveat_common+0x168/0x1cc
> __arm64_sys_execve+0x44/0x68
> invoke_syscall+0x54/0x110
> el0_svc_common.constprop.0+0x40/0xe0
> do_el0_svc+0x1c/0x28
> el0_svc+0x54/0x424
> el0t_64_sync_handler+0xa0/0xe4
> el0t_64_sync+0x1b0/0x1b4
> Slab 0xffffffbfc1076e00 objects=23 used=18 fp=0xffff8001076e2140 flags=0x13fffe0000000240(workingset|head|node=1|zone=0|lastcpupid=0x1ffff)
> Object 0xffff8001076ec640 @offset=50752 fp=0xffff8001076e2140
>
> [...]
>
> The corruption is always exactly 4 bytes (0xffffffff) with everything
> around still being intact poison. The in-object offset (51432 - 50752 =
> 680) resolves to &mm->sc_stat.cpu, and 0xffffffff is just -1. My AI model
> points me to this write in account_mm_sched():
>
> if (READ_ONCE(mm->sc_stat.cpu) != -1)
> WRITE_ONCE(mm->sc_stat.cpu, -1);
>
> and helps with analyzing and fixing the issue like below :-) . Please have
> a look.
>
> Thanks,
> Zenghui
>
> ---8<---
>
> From 992b515f18710e77308cf5f88943cc3ce918a525 Mon Sep 17 00:00:00 2001
> From: "Zenghui Yu (Huawei)" <zenghui.yu@xxxxxxxxx>
> Date: Mon, 14 Sep 2026 22:00:18 +0800
> Subject: [PATCH] sched/cache: Fix use-after-free of mm in account_mm_sched()

I think you have hit a similar use after free issue that was discussed in this
thread.
https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/

Can you try the last two patches in this 4 patch series
that address this issue in a comprehensive way
https://lore.kernel.org/lkml/cover.1789061845.git.tim.c.chen@xxxxxxxxxxxxxxx/

Thanks.

Tim

>
> account_mm_sched() accounts runtime against rq->curr and dereferences its
> ->mm: it updates the percpu chunk mm->sc_stat.pcpu_sched and may write
> mm->sc_stat.cpu = -1.
>
> update_se(), which samples rq->curr and calls account_mm_sched(), is not
> only called from local contexts (tick, context switch) but also through
> update_curr() from enqueue/dequeue paths, which frequently run on a remote
> CPU while holding this rq's lock (cross-CPU try_to_wake_up(), load
> balancing).
>
> In those remote contexts rq->curr is a task concurrently running on its
> home CPU. The rq lock guarantees that rq->curr's identity does not change,
> but it says nothing about the lifetime of rq->curr->mm: that task does not
> need the rq lock to execute execve or exit, and switches and drops its ->mm
> under task_lock() and mmput(), neither of which orders against the remote
> CPU. A remote CPU can therefore sample a valid mm pointer right before it
> is freed and write to it afterwards, corrupting the freed mm_struct (and
> the pcpu_sched percpu chunk, which mm_destroy_sched() frees even earlier).
>
> Observed with CONFIG_SLUB_DEBUG=y as a sporadic "Poison overwritten" report
> on the mm_struct cache, with the overwritten bytes resolving to
> &mm->sc_stat.cpu.
>
> Only account the physically running task (p == current), whose ->mm cannot
> go away while it is the one executing this code. Local tick, context
> switch and sched_ttwu_pending() paths are unaffected; updates skipped in
> remote contexts only cause minor under-accounting of the sc_stat runtime
> heuristics.
>
> Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
> Assisted-by: GLM-5.3 OpenCode
> Signed-off-by: Zenghui Yu (Huawei) <zenghui.yu@xxxxxxxxx>
> ---
> kernel/sched/fair.c | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index ade1eceb39b8..2bbf59370d23 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -1731,6 +1731,9 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec)
> int mm_sched_llc = -1;
> unsigned long epoch;
>
> + if (p != current)
> + return;
> +
> if (!sched_cache_enabled())
> return;
>