Re: [PATCH] mm/hugetlb: fix missing migratable flag on same-node hugetlb migration

From: mawupeng

Date: Tue Jul 07 2026 - 21:14:48 EST




On 周三 2026-7-8 04:21, Andrew Morton wrote:
> On Tue, 7 Jul 2026 19:02:54 +0800 Wupeng Ma <mawupeng1@xxxxxxxxxx> wrote:
>
>> Commit ba23f58de896 ("mm/migrate: don't call
>> folio_putback_active_hugetlb() on dst hugetlb folio") moved setting of
>> the migratable flag and active-list placement from
>> folio_putback_active_hugetlb(dst) into move_hugetlb_state(), so that
>> the freshly allocated destination folio is handled where allocation is
>> known to have succeeded.
>>
>> Unfortunately, the new code was appended after the existing
>> temporary-folio block in move_hugetlb_state(), which contains an early
>> return added earlier by commit 5af1ab1d24e08 ("mm/hugetlb: optimize
>> the surplus state transfer code in move_hugetlb_state()"):
>>
>> if (folio_test_hugetlb_temporary(new_folio)) {
>> ...
>> if (new_nid == old_nid)
>> return; <-- skips the new code
>> ...
>> }
>>
>> /* added by ba23f58 */
>> folio_set_hugetlb_migratable(new_folio);
>> list_move_tail(&new_folio->lru, ...&h->hugepage_activelist);
>>
>> When the destination folio is temporary (i.e. the hugetlb pool was
>> exhausted and the migration callback fell back to
>> alloc_migrate_hugetlb_folio()) and the migration does not cross a
>> node -- the common case, and always true on a single-NUMA system --
>> move_hugetlb_state() returns before setting the migratable flag or
>> adding the new folio to the active list. The destination folio is
>> then installed in the page table but cannot be isolated afterwards,
>> since folio_isolate_hugetlb() rejects folios without the migratable
>> flag; a subsequent soft-offline, hard-offline or memory-hotplug
>> offline of that folio fails with -EBUSY.
>>
>> This was reproduced on a single-NUMA arm64 VM: a second
>> MADV_SOFT_OFFLINE on an already-migrated hugetlb page returned EBUSY
>> and logged "hugepage isolation failed".
>>
>> Keep the surplus adjustment, which is the only part that depends on
>> the node crossing, guarded by `if (new_nid != old_nid)', while making
>> the migratable flag and active-list placement unconditional. This
>> preserves the cleanup intent of ba23f58 and closes the early-return
>> hole.
>
> Thanks, I'll queue this for test and review.
>
>> Fixes: ba23f58de896 ("mm/migrate: don't call folio_putback_active_hugetlb() on dst hugetlb folio")
>
> It sounds like we should backport this into -stable kernels?

Yes, it is precisely because this patch was backported to 6.6 stable that
we discovered this issue.

>
>