Re: [PATCH v2] x86/mm: Drop the page allocation from pud_free_pmd_page()

From: Pedro Falcato

Date: Thu Sep 24 2026 - 13:33:40 EST


On Thu, Sep 24, 2026 at 03:31:16AM +0500, Mikhail Gavrilov wrote:
> On a box with a discrete GPU, lockdep reports a possible deadlock as soon
> as kswapd shrinks the TTM page pool:
>
> WARNING: possible circular locking dependency detected
> 7.3.0-rc3-f6e7b42bf05b+ #183 Tainted: G U
> ------------------------------------------------------
> kswapd0/269 is trying to acquire lock:
> ((init_mm).mmap_lock){++++}-{4:4}, at: change_page_attr_set_clr+0x29a/0x4a0
> but task is already holding lock:
> (pool_shrink_rwsem){.+.+}-{4:4}, at: ttm_pool_shrink+0xb2/0x330 [ttm]
> Chain exists of:
> (init_mm).mmap_lock --> fs_reclaim --> pool_shrink_rwsem
>
> The cycle is built from three edges:
>
> 1) pool_shrink_rwsem -> (init_mm).mmap_lock
>
> The TTM shrinker restores the caching attribute of every page it
> frees, while holding pool_shrink_rwsem:
>
> ttm_pool_shrink()
> -> ttm_pool_dispose_list()
> -> ttm_pool_free_page()
> -> set_pages_wb()
> -> change_page_attr_set_clr() [ init_mm mmap read lock ]
>
> 2) fs_reclaim -> pool_shrink_rwsem
>
> The same shrinker, called from reclaim.
>
> 3) (init_mm).mmap_lock -> fs_reclaim
>
> ioremap() installing a huge PUD mapping over an existing PMD table:
>
> ioremap_page_range()
> -> vmap_range_noflush()
> -> vmap_try_huge_pud() [ init_mm mmap read lock ]
> -> pud_free_pmd_page()
> -> __get_free_page(GFP_KERNEL) [ enters reclaim ]
>
> Edge 3 is the one that should not exist. Now that the attribute-change
> path takes the init_mm mmap lock, reclaim can acquire it, so the lock
> must not be held over an allocation which can enter reclaim. CPA itself
> follows this rule: split_large_page() drops the lock around
> pagetable_alloc(). The huge vmap path, which has held the same lock
> since commit 26444eb71465
> ("mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF"),
> does not: pud_free_pmd_page() allocates a scratch page underneath it.
>
> That page does not need to exist. It only holds a copy of the PMD
> entries, so that they can be cleared before the PUD is. But the PMD
> table itself is freed after pud_clear() and the flush, so the code
> already relies on the table being out of reach of the page walker at
> that point - and if it is safe to free it then, it is safe to read it
> then. Nobody else writes to it either: vmap_try_huge_pud() only gets
> here for a range covering the whole PUD, and ptdump is kept out by the
> init_mm lock the caller holds.
>
> So clear the PUD, flush, and free the PTE tables straight from the
> detached PMD table - the same order pmd_free_pte_page() uses one level
> down. With no allocation left the cycle is gone, and so is the only way
> this function could fail.
>
> The copy came with commit 5e0fb5df2ee8
> ("x86/mm: Add TLB purge to free pmd/pte page interfaces"), whose
> changelog explains the flush but not the copy; the allocation itself was
> already questioned in review back then [1]. The same lock cycle was
> also reported from the i915 shrinker, with &vm->mutex in place of
> pool_shrink_rwsem [2].
>
> Fixes: d5d8b8662e6e ("x86/mm/pat: Acquire init_mm read lock on attribute changes to avoid UAF")
> Suggested-by: Pedro Falcato <pfalcato@xxxxxxx>
> Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@xxxxxxxxx>

Reviewed-by: Pedro Falcato <pfalcato@xxxxxxx>

Thanks for the fix!

--
Pedro