Re: [PATCH v2 00/17] mm/huge_memory: clean up folio split and lift swapcache split limits
From: David Hildenbrand (Arm)
Date: Thu Aug 13 2026 - 15:40:37 EST
On 8/12/26 20:48, Kairui Song via B4 Relay wrote:
> This series clean up the split code, add better swap cache split support
> for mappingless, large order, uniform and non-uniform split. Generic
> performance is on par or slightly better, and stack usage is reduced.
>
> The swap cache infrastructure can handle non-uniform or high order folio
> replace, so there is no reason for either restriction from the THP side.
> What stands in the way is the mixed anon/file folio split routine,
> which makes lifting the restrictions hard to follow, and it already
> carries some buggy or redundant checks.
>
> So this series cleans up the split path and separates anon and file
> splitting into two helpers. The file split path never sees a swap
> cache folio, and that is now enforced up front: a folio that is both
> in the page cache and the swap cache can only be a shmem folio, which
> remains unsupported and is rejected early. That helps to rule out swap
> cache handling in that part completely. Only the anon split path
> handles swap cache folios, with an anon mapping or mappingless:
> either way the splitting is similar, and non-uniform split is
> supported as well.
>
> Order-1 is still forbidden for swap cache splitting. In theory it is
> doable for shmem swap cache folios, but a mappingless swap cache
> folio cannot currently be told apart from a shmem one, so forbid it
> for all swap cache folios for now.
>
> Testing:
>
> The in-tree split_huge_page_test selftest (uniform, non-uniform and
> in-folio-offset splits of anon and pagecache folios) passes 62/62 on
> the patched kernel.
>
> ftrace function_graph tracing filtered on __folio_split() was used to
> compare per-call durations between the base and the patched kernel on
> the same x86-64 box (interleaved runs across alternating reboots;
> mean +- stddev of the per-run averages, 135 split calls per run):
>
> base: 24 runs, 69.6 +- 0.7 us per __folio_split()
> patched: 26 runs, 68.8 +- 1.3 us per __folio_split()
>
> The patched kernel is consistently ~1% faster; with this sample
> count the difference is outside run-to-run noise.
>
> On x86-64 with gcc 12 (-fstack-usage), the stack frame of
> __folio_split() shrinks from 240 to 96 bytes, and the worst-case
> split call chain from ~544 to ~384 (anon) or ~464 (file) bytes.
>
> Bloat-o-meter shows a tiny growth of huge_memory.o:
> before=58419 after=58446, chg +0.05% (+27 bytes).
>
> Signed-off-by: Kairui Song <kasong@xxxxxxxxxxx>
> ---
I might need a bit to get to this; but the merge window is about to open either
way so, so this is material for the one afterwards.
--
Cheers,
David