Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP
From: Borislav Petkov
Date: Fri Oct 09 2026 - 00:14:52 EST
On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote:
> With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a
> user address with EFLAGS.AC clear as a kernel bug: it does not consult the
> exception table and oopses right away. That is correct for ordinary kernel
> code, where get_kernel_nofault() and the other nofault accessors never let
> a user address reach a faulting instruction.
>
> JITed BPF programs are different. A privileged program may dereference a
> pointer the verifier cannot prove valid, and the verifier marks such loads
> PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM
> load, so that a fault on an unmapped kernel address zeroes the destination
> register and the program continues. Since a user address would oops
> instead, the x86 JIT also emits an address range check in front of every
> PROBE_MEM load, which keeps user addresses, the guard page above
> TASK_SIZE_MAX and the vsyscall page away from the load. The check
> duplicates the fault handler's knowledge of the address space layout, got
> the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM
> runtime load check"), and costs nine instructions and 32 to 39 bytes of
> code per load.
I can't parse that. Why do bpf memory accesses need to be handled differently
than any other memory accesses when SMAP is enabled?
Perhaps you should give a concrete example.
And why can't all that gunk be resolved at program load instead of going all
the way in the #PF handler?
> When the faulting instruction belongs to a BPF program, resolve its
> exception table entry instead of oopsing, exactly as is done for faults on
> kernel addresses. The is_bpf_text_address() lookup sits inside the
> unlikely() SMAP branch that currently ends in page_fault_oops(), so no path
> that does not oops today executes any additional code, and the oops itself
> is unchanged when no entry matches. Non-BPF code keeps the existing
> behaviour: a kernel-mode user access without STAC still oopses, extable
> entry or not.
>
> The next patch uses this to drop the range check from the JIT when SMAP is
There's no next patch and previous patch in git history.
> enabled. Nothing changes about which addresses a BPF program may read: a
> user address never becomes readable, since SMAP forbids the access, and a
> PROBE_MEM load of a kernel address is handled as before. Without SMAP the
> JIT keeps its range check and this path is never reached.
>
> Acked-by: Puranjay Mohan <puranjay@xxxxxxxxxx>
> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@xxxxxxxxx>
> ---
> arch/x86/mm/fault.c | 11 +++++++++++
> 1 file changed, 11 insertions(+)
>
> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
> index aa88370ce739..2060e5f35d77 100644
> --- a/arch/x86/mm/fault.c
> +++ b/arch/x86/mm/fault.c
> @@ -20,6 +20,7 @@
> #include <linux/mm_types.h>
> #include <linux/mm.h> /* find_and_lock_vma() */
> #include <linux/vmalloc.h>
> +#include <linux/filter.h> /* is_bpf_text_address() */
>
> #include <asm/cpufeature.h> /* boot_cpu_has, ... */
> #include <asm/traps.h> /* dotraplinkage, ... */
> @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs,
> if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) &&
> !(error_code & X86_PF_USER) &&
> !(regs->flags & X86_EFLAGS_AC))) {
> + /*
> + * JITed BPF programs dereference untrusted pointers with loads
> + * that carry an exception table entry (PROBE_MEM). SMAP makes
> + * sure such a load cannot read user memory, so resolve the
> + * fault through the exception table, as for an unmapped kernel
> + * address, instead of oopsing.
> + */
> + if (is_bpf_text_address(regs->ip) &&
> + fixup_exception(regs, X86_TRAP_PF, error_code, address))
> + return;
I'm not at all amused from this adding bpf-specific handling to the #PF
handler, TBH...
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette