Re: [PATCH] mm: don't defer freeing kernel page tables while booting

From: David Hildenbrand (Arm)

Date: Thu Sep 24 2026 - 03:13:17 EST


On 9/24/26 08:43, Mikhail Gavrilov wrote:
> Booting with a boot-time function tracer and a filter, for example
>
> ftrace=function ftrace_filter=pud_free_pmd_page
>
> panics on 7.3-rc4 as soon as the tracer starts:
>
> [ 23.531178] Starting tracer 'function'
> [ 23.675800] Oops: general protection fault, probably for non-canonical address 0xdffffc0000000038: 0000 [#1] SMP KASAN NOPTI
> [ 23.819917] KASAN: null-ptr-deref in range [0x00000000000001c0-0x00000000000001c7]
> [ 23.964025] CPU: 0 UID: 0 PID: 0 Comm: swapper Not tainted 7.3.0-rc4-fe2ec83746e5-with-fixes-v2+ #195 PREEMPT(undef)
> [ 24.252248] RIP: 0010:__queue_work+0xab/0xf00
> [ 25.981629] Call Trace:
> [ 26.125727] <TASK>
> [ 26.413912] ? pagetable_free_kernel+0x20/0x120
> [ 26.990283] queue_work_on+0x97/0xf0
> [ 27.134382] __cpa_collapse_large_pages+0x501/0x6f0
> [ 27.566662] cpa_flush+0x394/0x620
> [ 27.998953] change_page_attr_set_clr+0x321/0x4a0
> [ 29.151729] set_memory_rox+0xa2/0xf0
> [ 29.584018] create_trampoline+0x431/0x6f0
> ...
> [ 44.343347] Kernel panic - not syncing: Attempted to kill the idle task!
>
> The boot-time tracer is started from early_trace_init(), which runs
> before workqueue_init_early(). Making its trampoline read-only splits a
> large page, and CPA collapses it again right away. The split table has
> been a kernel page table since commit 9e4a3ec3411b
> ("x86/mm/pat: Allocate split page tables as kernel page tables"), so the
> collapse frees it through pagetable_free_kernel(), which queues work on
> system_percpu_wq - still NULL at that point.
>
> The deferral exists so that IOMMUs using SVA can have their paging
> structure caches flushed before a kernel page table is freed. While
> system_state is still SYSTEM_BOOTING only the boot CPU runs and no IOMMU
> has been initialised yet (on x86 that happens from pci_iommu_init(), a
> rootfs_initcall), so there is nothing to flush. Free the table directly
> in that case.
>
> Fixes: 9e4a3ec3411b ("x86/mm/pat: Allocate split page tables as kernel page tables")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Mikhail Gavrilov <mikhail.v.gavrilov@xxxxxxxxx>
> ---
> #regzbot introduced: 9e4a3ec3411b
>
> Tested on a Ryzen 9 7950X with a Radeon RX 7900 XTX, lockdep and KASAN
> enabled, booting with the command line above plus earlycon=efifb so
> that the oops is visible:
>
> 7.3-rc4 (fe2ec83746e5) panics as quoted
> + revert of 9e4a3ec3411b boots
> + this patch, 9e4a3ec3411b kept boots
>
> With this patch the boot-time tracer works as intended and records
> pud_free_pmd_page() being called from vmap_p4d_range() during boot.
> All three kernels also carry my pud_free_pmd_page() v2 patch and
> unrelated local changes elsewhere (HID, sound/usb, Bluetooth, debugfs);
> none of them is on this path.
>
> 9e4a3ec3411b is also in 7.2.7 and 6.18.53, together with the deferred
> pagetable_free_kernel(); I have not tried those.
>
> mm/pgtable-generic.c | 11 +++++++++++
> 1 file changed, 11 insertions(+)
>
> diff --git a/mm/pgtable-generic.c b/mm/pgtable-generic.c
> index b91b1a98029c..2d9b4ba33b46 100644
> --- a/mm/pgtable-generic.c
> +++ b/mm/pgtable-generic.c
> @@ -440,6 +440,17 @@ static void kernel_pgtable_work_func(struct work_struct *work)
>
> void pagetable_free_kernel(struct ptdesc *pt)
> {
> + /*
> + * While the system is still booting only the boot CPU runs and no
> + * IOMMU has been set up, so nothing can be caching this table and
> + * there is nothing to flush. The workqueue this defers to may not
> + * exist yet either.
> + */
> + if (system_state == SYSTEM_BOOTING) {
> + __pagetable_free(pt);
> + return;
> + }
> +
> spin_lock(&kernel_pgtable_work.lock);
> list_add(&pt->pt_list, &kernel_pgtable_work.list);
> spin_unlock(&kernel_pgtable_work.lock);

Should we instead simply skip the

schedule_work(&kernel_pgtable_work.work);

and rely on anybody freeing stuff later to just free that one alongside?

That avoids throwing in more freeing handling.

--
Cheers,

David