Re: [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()

From: Edgecombe, Rick P

Date: Thu Oct 01 2026 - 12:40:45 EST


On Thu, 2026-10-01 at 19:12 +0300, Hitesh Murali via B4 Relay wrote:
> From: Hitesh Murali <hitesh.murali@xxxxxxxxxxx>
>
> spurious_kernel_fault() only considers faults whose error code is exactly
> X86_PF_WRITE | X86_PF_PROT or X86_PF_INSTR | X86_PF_PROT. A write fault
> that reaches spurious_kernel_fault_check() is therefore a normal store,
> never a shadow stack access, and with CR0.WP set, which the kernel pins,
> the hardware permits a normal store only if _PAGE_RW is set. The check is
> applied to 4K, 2M and 1G leaves and to the PMD table entry, and _PAGE_RW
> means the same at each of these levels.
>
> Commit bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> made pte_write() return true for Write=0,Dirty=1 entries, the encoding of
> shadow stack memory, so that core mm treats shadow stacks as writable.
> pte_shstk() decides this from X86_FEATURE_SHSTK alone, so it applies on
> every shadow stack capable CPU, also with CONFIG_X86_USER_SHADOW_STACK=n.
> For a kernel mapping the hardware does not agree: the kernel never sets
> IA32_S_CET.SH_STK_EN, so a supervisor Write=0,Dirty=1 entry is an
> ordinary read-only entry.

I've always wondered if we really need spurious_kernel_fault(). Does anyone know
what functionality depends on it? Also, always seemed strange that it works from
user accesses to the kernel half of the address space.

>
> Kernel read-only mappings are created without Dirty since commit
> f788b71768ff ("x86/mm: Remove _PAGE_DIRTY from kernel RO pages"), but a
> store made while _PAGE_RW is temporarily set leaves Dirty behind, and the
> kernel does not clear it again.

The intention was to prevent RO,Dirty mappings all together unless they were
really intended to be shadow stack. So maybe that is the right fix. Do you know
what flow left the PTE in this state? It was this test module? Or something in
the upstream kernel?

> A later normal store to such an entry
> raises a protection fault, which is then classified as spurious. The
> store is restarted, faults again, and the CPU makes no progress.
>
> When the soft lockup watchdog reports it, it reports a stuck CPU at the
> store, which reads as a long running loop rather than as a write to a
> read-only page. This was found with an out-of-tree module that writes to
> .rodata; on a distribution kernel carrying bb3aadf7d446, on Sapphire
> Rapids, one CPU looped on the store for days.
>
> Test _PAGE_RW directly. The fault then takes the regular
> bad_area_nosemaphore() path: an exception table fixup where the access
> has one, as for any other read-only page, and otherwise an oops that
> reports the faulting address and the page table entry.
>
> Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> Assisted-by: LLM
> Signed-off-by: Hitesh Murali <hitesh.murali@xxxxxxxxxxx>
> Cc: stable@xxxxxxxxxxxxxxx