Re: [Patch v4 01/22] sched/cache: Introduce infrastructure for cache-aware load balancing

From: Zenghui Yu

Date: Fri Sep 18 2026 - 03:18:52 EST


On 9/15/26 7:12 AM, Tim Chen wrote:
> On Tue, 2026-09-15 at 01:50 +0800, Zenghui Yu wrote:
> >
>
> [snip]
>
> > I sporadically hit the SLUB "Poison overwritten" reports on the mm_struct
> > cache while running mm-new:
> >
> > [Poison overwritten] 0xffff8001076ec8e8-0xffff8001076ec8eb @offset=51432. First byte 0xff instead of 0x6b
> > =============================================================================
> > BUG mm_struct (Tainted: G N ): Object corrupt
> > -----------------------------------------------------------------------------
> >
> > Allocated in copy_process+0x1e48/0x2078 age=2 cpu=7 pid=11866
> > copy_process+0x1e48/0x2078
> > kernel_clone+0xa4/0x498
> > __do_sys_clone+0x5c/0x88
> > __arm64_sys_clone+0x1c/0x28
> > invoke_syscall+0x54/0x110
> > el0_svc_common.constprop.0+0x40/0xe0
> > do_el0_svc+0x1c/0x28
> > el0_svc+0x54/0x424
> > el0t_64_sync_handler+0xa0/0xe4
> > el0t_64_sync+0x1b0/0x1b4
> > Freed in __mmdrop+0x108/0x180 age=2 cpu=3 pid=11955
> > kmem_cache_free+0x290/0x53c
> > __mmdrop+0x108/0x180
> > __mmput+0x150/0x154
> > mmput+0x50/0x5c
> > exec_mm_put_old+0x74/0x84
> > setup_new_exec+0x7c/0x90
> > load_elf_binary+0x4b0/0x1914
> > bprm_execve+0x300/0x83c
> > do_execveat_common+0x168/0x1cc
> > __arm64_sys_execve+0x44/0x68
> > invoke_syscall+0x54/0x110
> > el0_svc_common.constprop.0+0x40/0xe0
> > do_el0_svc+0x1c/0x28
> > el0_svc+0x54/0x424
> > el0t_64_sync_handler+0xa0/0xe4
> > el0t_64_sync+0x1b0/0x1b4
> > Slab 0xffffffbfc1076e00 objects=23 used=18 fp=0xffff8001076e2140 flags=0x13fffe0000000240(workingset|head|node=1|zone=0|lastcpupid=0x1ffff)
> > Object 0xffff8001076ec640 @offset=50752 fp=0xffff8001076e2140
> >
> > [...]
> >
> > The corruption is always exactly 4 bytes (0xffffffff) with everything
> > around still being intact poison. The in-object offset (51432 - 50752 =
> > 680) resolves to &mm->sc_stat.cpu, and 0xffffffff is just -1. My AI model
> > points me to this write in account_mm_sched():
> >
> > if (READ_ONCE(mm->sc_stat.cpu) != -1)
> > WRITE_ONCE(mm->sc_stat.cpu, -1);
> >
> > and helps with analyzing and fixing the issue like below :-) . Please have
> > a look.
> >
> > ---8<---
> >
> > From 992b515f18710e77308cf5f88943cc3ce918a525 Mon Sep 17 00:00:00 2001
> > From: "Zenghui Yu (Huawei)" <zenghui.yu@xxxxxxxxx>
> > Date: Mon, 14 Sep 2026 22:00:18 +0800
> > Subject: [PATCH] sched/cache: Fix use-after-free of mm in account_mm_sched()
>
> I think you have hit a similar use after free issue that was discussed in this
> thread.
> https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/

I agree.

> Can you try the last two patches in this 4 patch series
> that address this issue in a comprehensive way
> https://lore.kernel.org/lkml/cover.1789061845.git.tim.c.chen@xxxxxxxxxxxxxxx/

I'll have a try. Thanks for the heads up!

Thanks,
Zenghui