mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC] x86: usermode IBT and signal handling
@ 2026-10-07 13:32 Richard Patel
  2026-10-07 15:28 ` Edgecombe, Rick P
  0 siblings, 1 reply; 2+ messages in thread
From: Richard Patel @ 2026-10-07 13:32 UTC (permalink / raw)
  To: Dave Hansen, Florian Weimer, x86
  Cc: Thomas Gleixner, Ingo Molnar, David Laight, Borislav Petkov,
	H. Peter Anvin, Xin Li, Rick Edgecombe, H.J. Lu, Peter Zijlstra,
	David Rubin, linux-kernel, libc-alpha

I'd like to request feedback for usermode IBT (indirect branch tracking)
support from the x86 and libc people before sending another series.
Also, thank you for the chat at LPC!

Questions:
1. Should the initial version preserve WAIT_FOR_ENDBR across signal
   delivery? (OpenBSD does not do that)
2. Where to preserve WAIT_FOR_ENDBR across signal? Options:
   - shstk signal frame
   - signal frame fpstate (carve out a bit in _fpx_sw_bytes)
   - uc_flags
3. OK to break legacy 32-bit sigreturn if IBT enabled?

For context, here the previous series:
- Intel (2021): https://lore.kernel.org/all/20210830182221.3535-1-yu-cheng.yu@intel.com/ 
- my v1: https://lore.kernel.org/lkml/20260517183024.16292-1-ripatel@wii.dev/T/
- my v2: https://lore.kernel.org/lkml/20260605184715.3383415-2-ripatel@wii.dev/T/

The uncontroversial parts of the kernel-side changes:
- usermode enables and locks IBT via RISC-V's
  prctl(PR_SET_CFI, PR_CFI_BRANCH_LANDING_PADS, ...) API
- IBT enablement is all-or-nothing process-wide (could extend later)
- IBT enablement sets both ENDBR_EN and NO_TRACK_EN
  (enforce endbr64 on indirect jumps, allow 'notrack' prefix to opt-out)
- vDSO polishing needed (add missing endbr64 markers, GNU property note)
- signal handler entrypoint does not need 'endbr64' (the signal handler
  entrypoint cannot easily be changed)

The libc side is analogous to shadow stack: a new tunable, tracking of
DSOs opting into IBT, prctl() on startup, etc. 
Florian seems fine with the prctl() approach as opposed to enabling IBT
automatically in the kernel.

Next, context switching:
1. kernel<->usermode (e.g. process switching, page faults, syscalls)
   require no changes (usermode shadow stack does all the work already)
2. usermode<->usermode (longjmp) is a glibc affair, no kernel changes
3. usermode<->kernel<->usermode (signal handling) is annoying

Signal handling (entering the handler and rt_sigreturn) is annoying
because x86 does not have hw support for unprivileged IBT state backup/
restore.

   500:  jmp rax   ; rax=1000
  1000:  nop       ; WAIT_FOR_ENDBR=1
  ** CET violation **

In OpenBSD and my v2 series, the signal frame does not back up IBT state
(the WAIT_FOR_ENDBR bit), and resets it to zero instead.

   500:  jmp rax   ; rax=1000
  ** Interrupt, signal handler called **
   100:  syscall   ; rt_sigreturn
  ** Return from signal handler **
  1000:  nop       ; WAIT_FOR_ENDBR=0
  ** CET bypassed! **

This race can be made deterministic for some apps (e.g. SIGBUS or
whatever).

But IBT with signals can be done, with two more pieces (see v1 series):

- space to back up IBT state (a single bit, WAIT_FOR_ENDBR).
  options are:
  - uc_flags, which is a uapi change (Intel's original series)
  - sigframe fpstate
    - decoupled from SHSTK
    - new bit in _fpx_sw_bytes (uapi/asm/sigcontext.h)
    - U_CET is supervisor state, so not saved by signal frame XSAVE
    - supports legacy ia32 sigframe
  - shadow stack signal frame
    - the LSB of the saved SSP is always zero, so we can use it
    - this makes IBT require SHSTK
    - awkward situation where SHSTK arch_prctl needs to be enabled
      before BRANCH_LANDING_PADS prctl is allowed
    - automatically bans 32-bit mode since SSP > 4G raises #GP
    - may confuse libgcc, CRIU, etc, depending on how we do it

- a primitive to change saved user state from kernel mode.
  when hardware switches from kernel to user, it loads IBT state from
  one of 3 places.  Since the kernel modifies this user state, it has
  to know what to modify.
  1. FRED exception frame (wfe bit)
  2. U_CET MSR (TIF_NEED_FPU_LOAD=0)
  3. task's fpstate save area

Let me know what you all think. My personal preference is:
1. preserve WAIT_FOR_ENDBR across signals.
2. back up the WAIT_FOR_ENDBR bit in signal frame fpstate
3. no special handling for 32-bit mode

Separately, I'll take a look at kernel shadow stacks unless someone else
is already working on it ...

Cheers,
-- Richard

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-07 15:28 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-07 13:32 [RFC] x86: usermode IBT and signal handling Richard Patel
2026-10-07 15:28 ` Edgecombe, Rick P

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®