Re: [PATCH 1/1] x86/mm: fix incomplete page-table invalidation with TCE
From: Nadav Amit
Date: Mon Oct 05 2026 - 02:39:33 EST
>
>
> On 5 Oct 2026, at 8:23, Lance Yang <lance.yang@xxxxxxxxx> wrote:
>
> pud_free_pmd_page() uses a single-address invalidation to flush the
> paging-structure caches before freeing the page tables. With AMD TCE
> enabled, this only invalidates upper-level entries associated with the
> target address. Cached PMD entries for other addresses in the PUD range can
> still reference the PTE pages being freed.
>
> The AMD manual quoted in the commit enabling TCE says these instructions
> remove
>
> "only those upper-level entries that lead to the target PTE in the page
> table hierarchy, leaving unrelated upper-level entries intact."
>
> Even with all PTEs cleared, speculative page walks can cache present PMD
> entries after the earlier TLB purge.
>
> Use a full TLB flush before freeing the page tables on CPUs with TCE. Keep
> the single-address invalidation otherwise.
>
> Fixes: 440a65b7d25f ("x86/mm: Enable AMD translation cache extensions")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Lance Yang <lance.yang@xxxxxxxxx>
> ---
> arch/x86/mm/pgtable.c | 11 ++++++++++-
> 1 file changed, 10 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/mm/pgtable.c b/arch/x86/mm/pgtable.c
> index 4a105f283cfb..6b7fa44f1bf6 100644
> --- a/arch/x86/mm/pgtable.c
> +++ b/arch/x86/mm/pgtable.c
> @@ -727,7 +727,16 @@ int pud_free_pmd_page(pud_t *pud, unsigned long addr)
> * via normal page walks. Make them unreachable
> * in cached mid-level walks too:
> */
> - flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
> + if (boot_cpu_has(X86_FEATURE_TCE)) {
> + /*
> + * With TCE enabled, a single-address flush does not invalidate
> + * cached PMD entries for the rest of the PUD range.
> + */
> + flush_tlb_all();
> + } else {
> + /* INVLPG to clear all paging-structure caches */
> + flush_tlb_kernel_range(addr, addr + PAGE_SIZE-1);
> + }
>
It might be cleaner to replace flush_tlb_all() with:
flush_tlb_kernel_range(addr, addr + PUD_SIZE - 1);
While the flush-ceiling would usually end up doing a full flush, the
code would be easier to follow (the very least). Maybe adding stride
to kernel TLB range flushing would make sense in the future.