Re: [PATCH v3 3/4] mm: hugetlb: Fix subpool usage leak on allocation failure
From: Karl Mehltretter
Date: Mon Sep 28 2026 - 02:30:16 EST
On Sun, Sep 27, 2026 at 10:19:08PM +0100, Ackerley Tng wrote:
> == Questions for Karl
>
> (Independent of the "Path forward for 7.3/7.4")
>
> 1. Does squashing patch [4/4] of this series into [2/4] [3/4] resolve
> any of what you found?
> 2. Does patch [1/4] introduce new issues, or does it just not fully fix
> all the issues? At this point I think we have so many bugs, it's more
> about fixing bugs progressively (not regressing) than finding a
> complete fix.
> 3. Would it be ok if you integrate your findings across your two replies
> on [2/4] and [3/4]? They're similar yet slightly different so it's
> kind of confusing. If you have reproducers, it would help to share
> them! :)
>
1. Squashing the unchanged patches does not fix it. Full v3 already includes
4/4, and both controlled tests still fail:
Path Cleanup After files / after unmount
----------------------- ------------ ---------------------------
hugetlb_reserve_pages() +2, -ENOMEM 2 / ULONG_MAX-1
alloc_hugetlb_folio() +1, -ENOMEM 3 / ULONG_MAX
Both should end at 4 / 0. I got the same result with one and four vCPUs.
Patch 4 fixes the later region_add() failure path, where the global
reservations have already been acquired. These two failures happen earlier,
after global accounting or folio allocation has failed, so 4/4 does not
cover them.
2. The exact v3 patch 1 does introduce regressions as an intermediate
commit. It starts tracking used_hpages on min_size-only mounts, but the old
failure paths do not undo the new charge. In the ordinary tests, with or
without the two queued fixes below it:
- after the SIGBUS test with no files, HugePages_Rsvd is 0 instead of 2;
- after the failed mmap and unmount, it is 3 instead of 0.
This does not mean always tracking used_hpages is wrong, but 1/4 is not safe
on its own. A reworked series on top of Zhao's and Jinmeng's fixes will need
new testing.
3. In both races, hugepage_subpool_get_pages() first changes the local
subpool state. The later global charge or folio allocation fails. Another
operation then releases capacity and a separate mount consumes it before
rollback. Rollback restores the local minimum, but its positive global
correction fails with -ENOMEM and the error is ignored.
I put the exact test hooks, both userspace controllers, QEMU helpers and
validators in this test-only commit:
https://github.com/kmehltretter82/linux/commit/3ed79d756dc3845b9af9dc1dec3daa8cb141a500
I support merging Zhao's max-only fix and Jinmeng's combined min/max fix
now. They are useful, narrower improvements and do not try to cover these
min-only races.
Karl