Re: [PATCH v3 0/4] Fix HugeTLB subpool used_hpages tracking
From: Ackerley Tng
Date: Wed Sep 23 2026 - 20:27:36 EST
Karl Mehltretter <kmehltretter@xxxxxxxxx> writes:
> On Wed, 16 Sep 2026 20:13:07 -0700 Andrew Morton <akpm@xxxxxxxxxxxxxxxxxxxx> wrote:
>> So if downstream people (-stable maintainers, others) follow our
>> recommendations, some kernels will get two of these patches, other
>> kernel versions will get three and some lucky kernels might get all
>> four. Are you confident that the patches can be split apart in this
>> fashion and still produce a good result?
>
> Two things I noticed while testing this series on v7.3-rc3 (x86_64
> QEMU, one CPU):
>
Thank you so much for testing!
> 1. It overlaps with two fixes already queued in mm.git, so I've added
> their authors to Cc:
>
> - 3/4 rewrites the same out_subpool_put: block as Zhao Li's
> "mm/hugetlb: fix max-only subpool accounting on
> alloc_hugetlb_folio failure" (mm-hotfixes-unstable), and has the
> same Fixes: tag.
Andrew also brought this up on v2 of this series [3].
> - 2/4 rewrites the same out_put_pages: block as Jinmeng Zhou's
> "mm/hugetlb: fix subpool minimum reservation rollback"
> (mm-unstable).
>
Just took a quick look, this fix doesn't address the case where there
might be another CPU that updates subpool state. In that race, the
global state will be messed up.
> 3/4 doesn't apply to mm-hotfixes-unstable, and 2/4-4/4 don't apply
> to mm-unstable. If I read them right, the queued fixes only handle
> mounts with size=, while 2/4 and 3/4 also cover min_size mounts, so
> they would probably replace them.
>
Yes, this set of fixes should be a superset of what's fixed in the other
patches. My vote is to switch over to this set. Would it be okay to drop
the other fixes from mm-*unstable?
> 2. On splitting: 1/4 applies cleanly on top of both queued fixes, but
> in my tests it made min_size-only mounts worse on its own. As far
> as I can tell, that's because it starts tracking used_hpages on
> those mounts, while the two error paths only release it after 2/4
> and 3/4.
>
> HugePages_Rsvd, expected value in parentheses. Q = the two queued
> fixes, wrap = 18446744073709551615:
>
> rc3 +Q +1/4 +Q+1/4 +1/4..4/4
> min_size=4M only:
> SIGBUS faults [1], no files (2) 2 2 0 0 2
> min_size=8M only:
> failed mmap [2], umounted (0) 0 0 3 3 0
> size=8M,min_size=4M:
> SIGBUS faults [1], no files (2) 0 0 0 0 2
> size=10M,min_size=8M:
> failed mmap [2], umounted (0) wrap 0 wrap 0 0
>
> With 1/4 alone, the min_size-only mount loses its reservation, and
> in the failed-mmap case three huge pages stay reserved after umount,
> presumably because the subpool is never freed.
>
> So it looks to me like 1/4 shouldn't go anywhere without 2/4 and 3/4.
Thanks for testing this! IIUC from your table, 1/4 introduces
regressions.
How do we handle this, should I squash patches 1/2/3 together, or is
there some way to use the Fixes: tags to require those 3 to go together?
> I haven't tested older stable trees. With all four patches applied,
> all the cases above give the expected values, and so does the
> partial-truncate case from the 1/4 changelog (Rsvd drops to 0 after
> the truncate instead of staying at 1).
>
> I'm happy to rerun these tests on a rebased v4.
>
> [1] https://lore.kernel.org/r/20260923065714.20781-1-kmehltretter@xxxxxxxxx/
> [2] the scenario from Jinmeng's changelog:
> https://lore.kernel.org/20260907132055.26696-1-zhoujinmeng@xxxxxxxxxxxxx
>
> Thanks,
> Karl
[3] https://lore.kernel.org/all/CAEvNRgHbzY30n1hz2vv7BiBMMhci+NEv_ubP1LjQNmBuu6Kjgw@xxxxxxxxxxxxxx/