[RFC] x86: usermode IBT and signal handling
From: Richard Patel
Date: Wed Oct 07 2026 - 10:14:23 EST
I'd like to request feedback for usermode IBT (indirect branch tracking)
support from the x86 and libc people before sending another series.
Also, thank you for the chat at LPC!
Questions:
1. Should the initial version preserve WAIT_FOR_ENDBR across signal
delivery? (OpenBSD does not do that)
2. Where to preserve WAIT_FOR_ENDBR across signal? Options:
- shstk signal frame
- signal frame fpstate (carve out a bit in _fpx_sw_bytes)
- uc_flags
3. OK to break legacy 32-bit sigreturn if IBT enabled?
For context, here the previous series:
- Intel (2021): https://lore.kernel.org/all/20210830182221.3535-1-yu-cheng.yu@xxxxxxxxx/
- my v1: https://lore.kernel.org/lkml/20260517183024.16292-1-ripatel@xxxxxxx/T/
- my v2: https://lore.kernel.org/lkml/20260605184715.3383415-2-ripatel@xxxxxxx/T/
The uncontroversial parts of the kernel-side changes:
- usermode enables and locks IBT via RISC-V's
prctl(PR_SET_CFI, PR_CFI_BRANCH_LANDING_PADS, ...) API
- IBT enablement is all-or-nothing process-wide (could extend later)
- IBT enablement sets both ENDBR_EN and NO_TRACK_EN
(enforce endbr64 on indirect jumps, allow 'notrack' prefix to opt-out)
- vDSO polishing needed (add missing endbr64 markers, GNU property note)
- signal handler entrypoint does not need 'endbr64' (the signal handler
entrypoint cannot easily be changed)
The libc side is analogous to shadow stack: a new tunable, tracking of
DSOs opting into IBT, prctl() on startup, etc.
Florian seems fine with the prctl() approach as opposed to enabling IBT
automatically in the kernel.
Next, context switching:
1. kernel<->usermode (e.g. process switching, page faults, syscalls)
require no changes (usermode shadow stack does all the work already)
2. usermode<->usermode (longjmp) is a glibc affair, no kernel changes
3. usermode<->kernel<->usermode (signal handling) is annoying
Signal handling (entering the handler and rt_sigreturn) is annoying
because x86 does not have hw support for unprivileged IBT state backup/
restore.
500: jmp rax ; rax=1000
1000: nop ; WAIT_FOR_ENDBR=1
** CET violation **
In OpenBSD and my v2 series, the signal frame does not back up IBT state
(the WAIT_FOR_ENDBR bit), and resets it to zero instead.
500: jmp rax ; rax=1000
** Interrupt, signal handler called **
100: syscall ; rt_sigreturn
** Return from signal handler **
1000: nop ; WAIT_FOR_ENDBR=0
** CET bypassed! **
This race can be made deterministic for some apps (e.g. SIGBUS or
whatever).
But IBT with signals can be done, with two more pieces (see v1 series):
- space to back up IBT state (a single bit, WAIT_FOR_ENDBR).
options are:
- uc_flags, which is a uapi change (Intel's original series)
- sigframe fpstate
- decoupled from SHSTK
- new bit in _fpx_sw_bytes (uapi/asm/sigcontext.h)
- U_CET is supervisor state, so not saved by signal frame XSAVE
- supports legacy ia32 sigframe
- shadow stack signal frame
- the LSB of the saved SSP is always zero, so we can use it
- this makes IBT require SHSTK
- awkward situation where SHSTK arch_prctl needs to be enabled
before BRANCH_LANDING_PADS prctl is allowed
- automatically bans 32-bit mode since SSP > 4G raises #GP
- may confuse libgcc, CRIU, etc, depending on how we do it
- a primitive to change saved user state from kernel mode.
when hardware switches from kernel to user, it loads IBT state from
one of 3 places. Since the kernel modifies this user state, it has
to know what to modify.
1. FRED exception frame (wfe bit)
2. U_CET MSR (TIF_NEED_FPU_LOAD=0)
3. task's fpstate save area
Let me know what you all think. My personal preference is:
1. preserve WAIT_FOR_ENDBR across signals.
2. back up the WAIT_FOR_ENDBR bit in signal frame fpstate
3. no special handling for 32-bit mode
Separately, I'll take a look at kernel shadow stacks unless someone else
is already working on it ...
Cheers,
-- Richard