Re: [BUG] mm/memory_hotplug: panic due to race between compaction and memory hot-unplug

From: David Hildenbrand (Arm)

Date: Mon Sep 07 2026 - 10:48:17 EST


On 9/3/26 11:55, Yuan Liu wrote:
> Hi all,

Hi!

>
> While stress testing memory hotplug on a VM guest running an
> unmodified vanilla mainline kernel (7.3.0-rc1, as reported by
> uname -r), we hit a kernel panic in the guest when memory
> hot-unplug runs concurrently with memory compaction.
>
> The kernel was built from mainline at commit:
>
> cee9395acd80 ("Linux 7.3-rc1")
>
> To be more specific, after a large virtio-mem hot-unplug, the guest
> kernel takes a fatal page fault in suitable_migration_target(), called
> from isolate_freepages() during compaction.

Sounds like a real problem we should tackle.

>
> We are not sure whether this race is reachable under realistic
> workloads or only under this synthetic stress test. Sharing it here
> in case it is useful, and in case this is already a known issue.
> Thanks.
>
>
> Call trace (top to bottom)
> ==========================
> - RIP: suitable_migration_target+0x5/0x70
> isolate_freepages() <- compaction_alloc() <-
> migrate_pages() <- compact_zone() <- compact_node() <-
> sysctl_compaction_handler().
>
>
> Why the race happens
> ====================
> CPU0 (compaction free-scanner) CPU1 (virtio-mem hot-unplug)
> ---- ----
> page = pageblock_pfn_to_page()
> /* checks pass, section ONLINE */
> /* returns valid struct page* */
>
> offline_pages()
> /* section -> offline */
> __remove_pages()
> vmemmap_free()
> /* struct page UNMAPPED */
>
> suitable_migration_target(page)
> PageBuddy(page)
> read page->page_type
> *** not-present fault -> panic ***
>

If it can be hit with virtio-mem, it can be hit with any other memory hotunplug
code path (e.g., dimm, dax).

It is known that pfn_to_online_page() is racy. We usually expect the race window
to be extremely small. But for compaction the race window is much larger.

We do have get_online_mems()/put_online_mems(), big its the big hammer.

We once discussed using RCU to protect pfn_to_online_page(), but I suspect for
comapction that's not actually helpful (again, large race window).

We'd have to use the memory notifier or a dedicated callback to let memory
offlining sync with compaction.

That's where it gets tricky :)


It would be sufficient to let MEM_OFFLINE wait until any previous compaction
users would be done with the range. In that case, the sections would be offline,
but the memmap and zone range would not have been adjusted yet.

--
Cheers,

David