* [RFC] x86: usermode IBT and signal handling
@ 2026-10-07 13:32 Richard Patel
2026-10-07 15:28 ` Edgecombe, Rick P
0 siblings, 1 reply; 2+ messages in thread
From: Richard Patel @ 2026-10-07 13:32 UTC (permalink / raw)
To: Dave Hansen, Florian Weimer, x86
Cc: Thomas Gleixner, Ingo Molnar, David Laight, Borislav Petkov,
H. Peter Anvin, Xin Li, Rick Edgecombe, H.J. Lu, Peter Zijlstra,
David Rubin, linux-kernel, libc-alpha
I'd like to request feedback for usermode IBT (indirect branch tracking)
support from the x86 and libc people before sending another series.
Also, thank you for the chat at LPC!
Questions:
1. Should the initial version preserve WAIT_FOR_ENDBR across signal
delivery? (OpenBSD does not do that)
2. Where to preserve WAIT_FOR_ENDBR across signal? Options:
- shstk signal frame
- signal frame fpstate (carve out a bit in _fpx_sw_bytes)
- uc_flags
3. OK to break legacy 32-bit sigreturn if IBT enabled?
For context, here the previous series:
- Intel (2021): https://lore.kernel.org/all/20210830182221.3535-1-yu-cheng.yu@intel.com/
- my v1: https://lore.kernel.org/lkml/20260517183024.16292-1-ripatel@wii.dev/T/
- my v2: https://lore.kernel.org/lkml/20260605184715.3383415-2-ripatel@wii.dev/T/
The uncontroversial parts of the kernel-side changes:
- usermode enables and locks IBT via RISC-V's
prctl(PR_SET_CFI, PR_CFI_BRANCH_LANDING_PADS, ...) API
- IBT enablement is all-or-nothing process-wide (could extend later)
- IBT enablement sets both ENDBR_EN and NO_TRACK_EN
(enforce endbr64 on indirect jumps, allow 'notrack' prefix to opt-out)
- vDSO polishing needed (add missing endbr64 markers, GNU property note)
- signal handler entrypoint does not need 'endbr64' (the signal handler
entrypoint cannot easily be changed)
The libc side is analogous to shadow stack: a new tunable, tracking of
DSOs opting into IBT, prctl() on startup, etc.
Florian seems fine with the prctl() approach as opposed to enabling IBT
automatically in the kernel.
Next, context switching:
1. kernel<->usermode (e.g. process switching, page faults, syscalls)
require no changes (usermode shadow stack does all the work already)
2. usermode<->usermode (longjmp) is a glibc affair, no kernel changes
3. usermode<->kernel<->usermode (signal handling) is annoying
Signal handling (entering the handler and rt_sigreturn) is annoying
because x86 does not have hw support for unprivileged IBT state backup/
restore.
500: jmp rax ; rax=1000
1000: nop ; WAIT_FOR_ENDBR=1
** CET violation **
In OpenBSD and my v2 series, the signal frame does not back up IBT state
(the WAIT_FOR_ENDBR bit), and resets it to zero instead.
500: jmp rax ; rax=1000
** Interrupt, signal handler called **
100: syscall ; rt_sigreturn
** Return from signal handler **
1000: nop ; WAIT_FOR_ENDBR=0
** CET bypassed! **
This race can be made deterministic for some apps (e.g. SIGBUS or
whatever).
But IBT with signals can be done, with two more pieces (see v1 series):
- space to back up IBT state (a single bit, WAIT_FOR_ENDBR).
options are:
- uc_flags, which is a uapi change (Intel's original series)
- sigframe fpstate
- decoupled from SHSTK
- new bit in _fpx_sw_bytes (uapi/asm/sigcontext.h)
- U_CET is supervisor state, so not saved by signal frame XSAVE
- supports legacy ia32 sigframe
- shadow stack signal frame
- the LSB of the saved SSP is always zero, so we can use it
- this makes IBT require SHSTK
- awkward situation where SHSTK arch_prctl needs to be enabled
before BRANCH_LANDING_PADS prctl is allowed
- automatically bans 32-bit mode since SSP > 4G raises #GP
- may confuse libgcc, CRIU, etc, depending on how we do it
- a primitive to change saved user state from kernel mode.
when hardware switches from kernel to user, it loads IBT state from
one of 3 places. Since the kernel modifies this user state, it has
to know what to modify.
1. FRED exception frame (wfe bit)
2. U_CET MSR (TIF_NEED_FPU_LOAD=0)
3. task's fpstate save area
Let me know what you all think. My personal preference is:
1. preserve WAIT_FOR_ENDBR across signals.
2. back up the WAIT_FOR_ENDBR bit in signal frame fpstate
3. no special handling for 32-bit mode
Separately, I'll take a look at kernel shadow stacks unless someone else
is already working on it ...
Cheers,
-- Richard
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [RFC] x86: usermode IBT and signal handling
2026-10-07 13:32 [RFC] x86: usermode IBT and signal handling Richard Patel
@ 2026-10-07 15:28 ` Edgecombe, Rick P
0 siblings, 0 replies; 2+ messages in thread
From: Edgecombe, Rick P @ 2026-10-07 15:28 UTC (permalink / raw)
To: fweimer, ripatel, dave.hansen, x86
Cc: libc-alpha, bp, peterz, hpa, mingo, david.laight.linux,
hjl.tools, tglx, david, xin, linux-kernel
On Wed, 2026-10-07 at 13:32 +0000, Richard Patel wrote:
> I'd like to request feedback for usermode IBT (indirect branch tracking)
> support from the x86 and libc people before sending another series.
> Also, thank you for the chat at LPC!
>
> Questions:
> 1. Should the initial version preserve WAIT_FOR_ENDBR across signal
> delivery? (OpenBSD does not do that)
I think it should, but if it becomes a major technical hurdle, then having
something is better than nothing.
> 2. Where to preserve WAIT_FOR_ENDBR across signal? Options:
> - shstk signal frame
> - signal frame fpstate (carve out a bit in _fpx_sw_bytes)
> - uc_flags
The shadow stack signal frame is extensible to fit things like this and gives a
nice security protection for the bit. But I'd consider pursuing the simplest
option first. This option requires making IBT require shadow stack though. It
probably is the common case.
The original patches found a bit in the normal signal frame, but I don't think
it ever got fully settled? In any case it would need to be revisited at this
point.
But the best option is probably the one with the least resistance. It might be
shadow stack, since that is off to the side. We can always harden things with
further changes (add to shadow stack), or extend to IBT-only later.
> 3. OK to break legacy 32-bit sigreturn if IBT enabled?
IBT needs userspace enabling to work, and legacy 32 bit userspace is basically,
well, legacy. So unless there is any serious user that pops up, we should just
not supporting user IBT for 32 bit. This simplifies things. BTW we don't support
32 bit shadow stack.
>
> For context, here the previous series:
> - Intel (2021):
> https://lore.kernel.org/all/20210830182221.3535-1-yu-cheng.yu@intel.com/
> - my v1:
> https://lore.kernel.org/lkml/20260517183024.16292-1-ripatel@wii.dev/T/
> - my v2:
> https://lore.kernel.org/lkml/20260605184715.3383415-2-ripatel@wii.dev/T/
>
> The uncontroversial parts of the kernel-side changes:
> - usermode enables and locks IBT via RISC-V's
> prctl(PR_SET_CFI, PR_CFI_BRANCH_LANDING_PADS, ...) API
> - IBT enablement is all-or-nothing process-wide (could extend later)
What happened to the discussion of a PROT_IBT (like PROT_BTI)? I had POCed two
ways to do it and one was not that bad. The legacy bitmap thing is not easy to
use here, and I don't recommend it. But handling IBT #CPs by looking up the VMA
of RIP was pretty compact. If the VMA has !PROT_IBT, then clear the TRACKER bit
and proceed. This punishes mixed mode apps indirect calls (not *that* horrible
IIRC, I've lost my test results), but doesn't hurt fully enabled apps.
Now, the security question of whether a mixed mode makes sense, is valid I
think. But the issue we struggled with a lot for shadow stack was compatibility
of apps made of multiple packages (which is a lot of them), or which have JITs.
So I think a mixed mode would be valuable if it can be a stepping stone to a
fully locked down mode eventually. For example, start with a more permissive
prctl that turns on mixed (PROT_IBT) mode. Distros and other wider enablers can
use this without fear of crashing apps with JITs, etc. Emit a pr_info() or
something that incentives people to fix their libs. Then later distros can
switch to a fully enforced on thing.
But would be good to hear updated opinions from the distros on this. (maybe I
missed it?)
> - IBT enablement sets both ENDBR_EN and NO_TRACK_EN
> (enforce endbr64 on indirect jumps, allow 'notrack' prefix to opt-out)
> - vDSO polishing needed (add missing endbr64 markers, GNU property note)
> - signal handler entrypoint does not need 'endbr64' (the signal handler
> entrypoint cannot easily be changed)
I think the signal delivery should manually check for endbr in SW. Using
something like the speculate loop in shstk_pop_sigframe(). What is the problem?
At the same time, something small that actually upstream is better than nothing.
>
> The libc side is analogous to shadow stack: a new tunable, tracking of
> DSOs opting into IBT, prctl() on startup, etc.
> Florian seems fine with the prctl() approach as opposed to enabling IBT
> automatically in the kernel.
>
> Next, context switching:
> 1. kernel<->usermode (e.g. process switching, page faults, syscalls)
> require no changes (usermode shadow stack does all the work already)
> 2. usermode<->usermode (longjmp) is a glibc affair, no kernel changes
> 3. usermode<->kernel<->usermode (signal handling) is annoying
>
> Signal handling (entering the handler and rt_sigreturn) is annoying
> because x86 does not have hw support for unprivileged IBT state backup/
> restore.
>
> 500: jmp rax ; rax=1000
> 1000: nop ; WAIT_FOR_ENDBR=1
> ** CET violation **
>
> In OpenBSD and my v2 series, the signal frame does not back up IBT state
> (the WAIT_FOR_ENDBR bit), and resets it to zero instead.
>
> 500: jmp rax ; rax=1000
> ** Interrupt, signal handler called **
> 100: syscall ; rt_sigreturn
> ** Return from signal handler **
> 1000: nop ; WAIT_FOR_ENDBR=0
> ** CET bypassed! **
>
> This race can be made deterministic for some apps (e.g. SIGBUS or
> whatever).
>
> But IBT with signals can be done, with two more pieces (see v1 series):
>
> - space to back up IBT state (a single bit, WAIT_FOR_ENDBR).
> options are:
> - uc_flags, which is a uapi change (Intel's original series)
> - sigframe fpstate
> - decoupled from SHSTK
> - new bit in _fpx_sw_bytes (uapi/asm/sigcontext.h)
> - U_CET is supervisor state, so not saved by signal frame XSAVE
> - supports legacy ia32 sigframe
> - shadow stack signal frame
> - the LSB of the saved SSP is always zero, so we can use it
> - this makes IBT require SHSTK
> - awkward situation where SHSTK arch_prctl needs to be enabled
> before BRANCH_LANDING_PADS prctl is allowed
> - automatically bans 32-bit mode since SSP > 4G raises #GP
> - may confuse libgcc, CRIU, etc, depending on how we do it
>
> - a primitive to change saved user state from kernel mode.
> when hardware switches from kernel to user, it loads IBT state from
> one of 3 places. Since the kernel modifies this user state, it has
> to know what to modify.
> 1. FRED exception frame (wfe bit)
> 2. U_CET MSR (TIF_NEED_FPU_LOAD=0)
> 3. task's fpstate save area
>
> Let me know what you all think. My personal preference is:
> 1. preserve WAIT_FOR_ENDBR across signals.
> 2. back up the WAIT_FOR_ENDBR bit in signal frame fpstate
> 3. no special handling for 32-bit mode
>
> Separately, I'll take a look at kernel shadow stacks unless someone else
> is already working on it ...
>
> Cheers,
> -- Richard
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-07 15:28 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-07 13:32 [RFC] x86: usermode IBT and signal handling Richard Patel
2026-10-07 15:28 ` Edgecombe, Rick P
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®