Re: [PATCH v2] kasan: fix cache shrink race with CPU hotplug

From: Andrew Morton

Date: Sat Aug 08 2026 - 03:39:33 EST


On Sat, 8 Aug 2026 11:14:59 +0800 Hui Su <sh_def@xxxxxxx> wrote:

> kasan_quarantine_remove_cache() first invokes per_cpu_remove_cache() on
> all online CPUs. Each callback moves objects belonging to the cache from
> cpu_quarantine to the CPU's shrink_qlist, where they can later be freed
> from task context.
>
> kmem_cache_destroy() invokes the quarantine removal path while holding
> cpus_read_lock(), but kmem_cache_shrink() does not. The latter can
> therefore race with CPU offlining as follows:
>
> kmem_cache_shrink() CPU hotplug
> ------------------- -----------
> on_each_cpu()
> CPU1 moves objects to
> CPU1's shrink_qlist
> on_each_cpu() returns
> CPU1 goes offline
> kasan_cpu_offline()
> drains cpu_quarantine
> leaves shrink_qlist untouched
> for_each_online_cpu()
> skips CPU1
>
> The objects left on CPU1's shrink_qlist are not returned to the slab
> allocator. This may prevent kmem_cache_shrink() from releasing slabs
> that would otherwise become empty. If CPU1 remains offline, a later
> kmem_cache_destroy() also skips the list and can report that the cache
> still contains objects.
>
> An intermittent occurrence was observed with a virtio-9p filesystem.
> The mount and umount commands both returned 0, but the kernel logged
> the following during the userspace-triggered teardown:
>
> ...
>

Thanks, I'll queue this for testing while we await maintainer review.

AI review suggests that there's a pre-existing quarantine_size
accounting flaw later in this function:

https://sashiko.dev/#/patchset/20260808031459.3032812-1-sh_def@xxxxxxx

If true, I'm surprised this hasn't yet been reported.


Also, I'd like to see a need_resched() wrapping that expensive

/* Scanning whole quarantine can take a while. */
raw_spin_unlock_irqrestore(&quarantine_lock, flags);
cond_resched();
raw_spin_lock_irqsave(&quarantine_lock, flags);