From: "Kumar Kartikeya Dwivedi" <memxor@gmail.com>
To: "Borislav Petkov" <bp@alien8.de>
Cc: <bpf@vger.kernel.org>, "Puranjay Mohan" <puranjay@kernel.org>,
"Alexei Starovoitov" <ast@kernel.org>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Emil Tsalapatis" <emil@etsalapatis.com>,
"Ihor Solodrai" <ihor.solodrai@linux.dev>,
"Dave Hansen" <dave.hansen@linux.intel.com>,
"Andy Lutomirski" <luto@kernel.org>,
"Peter Zijlstra" <peterz@infradead.org>,
"Thomas Gleixner" <tglx@kernel.org>,
"Ingo Molnar" <mingo@redhat.com>, <kkd@meta.com>,
<kernel-team@meta.com>, <x86@kernel.org>,
<linux-kernel@vger.kernel.org>
Subject: Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP
Date: Fri, 09 Oct 2026 17:23:13 +0200 [thread overview]
Message-ID: <DM0ESVBT9553.2S1RWK1W7M8WW@gmail.com> (raw)
In-Reply-To: <20261009041346.GHashp-hYpLl-o-l3M@fat_crate.local>
On Fri Oct 9, 2026 at 6:13 AM CEST, Borislav Petkov wrote:
> On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote:
>> With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a
>> user address with EFLAGS.AC clear as a kernel bug: it does not consult the
>> exception table and oopses right away. That is correct for ordinary kernel
>> code, where get_kernel_nofault() and the other nofault accessors never let
>> a user address reach a faulting instruction.
>>
>> JITed BPF programs are different. A privileged program may dereference a
>> pointer the verifier cannot prove valid, and the verifier marks such loads
>> PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM
>> load, so that a fault on an unmapped kernel address zeroes the destination
>> register and the program continues. Since a user address would oops
>> instead, the x86 JIT also emits an address range check in front of every
>> PROBE_MEM load, which keeps user addresses, the guard page above
>> TASK_SIZE_MAX and the vsyscall page away from the load. The check
>> duplicates the fault handler's knowledge of the address space layout, got
>> the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM
>> runtime load check"), and costs nine instructions and 32 to 39 bytes of
>> code per load.
>
> I can't parse that. Why do bpf memory accesses need to be handled differently
> than any other memory accesses when SMAP is enabled?
>
It requires the same handling as get_kernel_nofault() as if that were inlined.
The point of the patch is to be able to avoid bounds checking in the JITed code.
We can probably extend this to other nofault accessors as well, but I wanted to
keep the scope limited to BPF for now, because the cost of the bounds checking
is significant in BPF programs.
For background: Privileged programs are allowed to trace kernel code and
dereference pointers where establishing whether the pointer is pointing to a
valid object is not possible during verification. Program authors know about
this behavior, most of the use cases for this are reading data, collecting
statistics, and so on.
It is a more convenient and efficient way to walk a chain of pointers than to
use the equivalent of copy_from_kernel_nofault() in a loop.
get_kernel_nofault() does user address checks in software, in
copy_from_kernel_nofault_allowed(), because the SMAP branch in
do_user_addr_fault() treats any kernel-mode fault on a user address as a missing
STAC and oopses without looking at the extable. The x86 JIT does the same today,
inlined, in front of every such load in a BPF program.
> Perhaps you should give a concrete example.
Consider this example:
SEC("fentry/tcp_retransmit_skb") <- attaches to the entry of tcp_retransmit_skb()
int BPF_PROG(retrans, struct sock *sk)
{
struct file *f = sk->sk_socket->file;
This program was written for doing a TCP related investigation. The first load
of sk is fine, but sk_socket can be NULL (orphaned socket) or stale, so the
second load becomes a mov with an extable entry. Below is the JITed sequence.
movq $-10485760, %r10
movq %rax, %r11
addq $2200, %r11
subq %r10, %r11
movabsq $72057594048413696, %r10
cmpq %r10, %r11
ja load
xorl %edi, %edi
jmp done
load:
movq 2200(%rax), %rdi
done:
Nine instructions and 32 to 39 bytes to guard a every load instruction against
an access that SMAP, when enabled, will block for us, and which we could fixup,
if we had exception handling in its fault path.
So the overall idea proposal is to eliminate this sequence and invoke the
fixup_exception() handler for the faulting instruction instead, which will zero
the destination register and continue execution, just like it does for faults on
kernel addresses already.
>
> And why can't all that gunk be resolved at program load instead of going all
> the way in the #PF handler?
>
This is what we do now, but it is a significant amount of code preceding each
load instruction. If you look at some examples in the cover letter, we can shave
the text size of programs by more than half in extreme cases.
Do note that we already do fixup_exception() for kernel address faults, so this
is mostly trying to mirror that behavior for the rest and eliminate the bounds
checking.
>> When the faulting instruction belongs to a BPF program, resolve its
>> exception table entry instead of oopsing, exactly as is done for faults on
>> kernel addresses. The is_bpf_text_address() lookup sits inside the
>> unlikely() SMAP branch that currently ends in page_fault_oops(), so no path
>> that does not oops today executes any additional code, and the oops itself
>> is unchanged when no entry matches. Non-BPF code keeps the existing
>> behaviour: a kernel-mode user access without STAC still oopses, extable
>> entry or not.
>>
>> The next patch uses this to drop the range check from the JIT when SMAP is
>
> There's no next patch and previous patch in git history.
Yeah, I will reword this bit in the commit log for v2.
>
>> enabled. Nothing changes about which addresses a BPF program may read: a
>> user address never becomes readable, since SMAP forbids the access, and a
>> PROBE_MEM load of a kernel address is handled as before. Without SMAP the
>> JIT keeps its range check and this path is never reached.
>>
>> Acked-by: Puranjay Mohan <puranjay@kernel.org>
>> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
>> ---
>> arch/x86/mm/fault.c | 11 +++++++++++
>> 1 file changed, 11 insertions(+)
>>
>> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
>> index aa88370ce739..2060e5f35d77 100644
>> --- a/arch/x86/mm/fault.c
>> +++ b/arch/x86/mm/fault.c
>> @@ -20,6 +20,7 @@
>> #include <linux/mm_types.h>
>> #include <linux/mm.h> /* find_and_lock_vma() */
>> #include <linux/vmalloc.h>
>> +#include <linux/filter.h> /* is_bpf_text_address() */
>>
>> #include <asm/cpufeature.h> /* boot_cpu_has, ... */
>> #include <asm/traps.h> /* dotraplinkage, ... */
>> @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs,
>> if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) &&
>> !(error_code & X86_PF_USER) &&
>> !(regs->flags & X86_EFLAGS_AC))) {
>> + /*
>> + * JITed BPF programs dereference untrusted pointers with loads
>> + * that carry an exception table entry (PROBE_MEM). SMAP makes
>> + * sure such a load cannot read user memory, so resolve the
>> + * fault through the exception table, as for an unmapped kernel
>> + * address, instead of oopsing.
>> + */
>> + if (is_bpf_text_address(regs->ip) &&
>> + fixup_exception(regs, X86_TRAP_PF, error_code, address))
>> + return;
>
> I'm not at all amused from this adding bpf-specific handling to the #PF
> handler, TBH...
I understand, and perhaps this can be adjusted to be a little different, but on
the flip side, at this point, the kernel is about to print an OOPS. I made it
BPF specific because we really don't expect to fixup exceptions for anything
else here. We already invoke fixup_exception() for kernel address faults.
The advantage is losing significant amount of code size in JITed programs and
the associated runtime overhead (you can see the numbers in the cover letter).
Do note that we can make it more generic, we can follow up and optimize
copy_from_kernel_nofault() with cpu_feature_enabled(X86_FEATURE_SMAP) as well.
But I didn't want to go there just yet. I am happy to work on a follow up if so
desired.
next prev parent reply other threads:[~2026-10-09 15:23 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
2026-10-09 4:13 ` Borislav Petkov
2026-10-09 15:23 ` Kumar Kartikeya Dwivedi [this message]
2026-10-09 2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
2026-10-09 3:42 ` bot+bpf-ci
2026-10-09 2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DM0ESVBT9553.2S1RWK1W7M8WW@gmail.com \
--to=memxor@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bp@alien8.de \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=dave.hansen@linux.intel.com \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=ihor.solodrai@linux.dev \
--cc=kernel-team@meta.com \
--cc=kkd@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=luto@kernel.org \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=puranjay@kernel.org \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®