Re: [PATCH] x86/mm: avoid a reclaiming allocation in pud_free_pmd_page()
From: Mikhail Gavrilov
Date: Wed Sep 23 2026 - 18:28:46 EST
On Wed, Sep 23, 2026 at 07:38:55PM +0100, Pedro Falcato wrote:
> So, the question is: why the heck do we need a copy? PMD is still allocated
> by the time we flush the TLB. Why doesn't a simple pud_clear() + flush_tlb +
> free over the pmd Just Work? Am I missing something? The git log isn't
> clueing me in.
I don't see anything you are missing. The copy came with 5e0fb5df2ee8
("x86/mm: Add TLB purge to free pmd/pte page interfaces"). Its changelog
explains the flush but not the copy, and the only discussion of the copy
in that thread was Joerg objecting to the allocation and suggesting a
list_head on the stack instead:
https://lore.kernel.org/all/20180529144438.GM18595@xxxxxxxxxx/
The existing code already frees the PMD table itself after pud_clear()
and that flush, so it already relies on the table being out of reach of
the page walker at that point. If it is safe to free it then, it is
safe to read it then; clearing the PMD entries up front buys nothing,
because nothing is freed before the flush. Nobody else writes to the
table either: vmap_try_huge_pud() only gets here for a range covering
the whole PUD, and ptdump is kept out by the init_mm lock the caller
holds.
> All-in-all, I would much prefer not having a copy of the PMD at all. Perhaps,
> if this isn't workable, then a linked list of PTEs would work. But I would rather
> not have tricky logic at all.
Agreed. It also removes the allocation instead of weakening it, so no
fallback path is left behind. I'll send a v2 that does pud_clear(), the
flush, and then frees the PTE tables straight from the detached PMD
table - the same order pmd_free_pte_page() already uses one level down.
Thanks for looking at it.
--
Thanks,
Mikhail