* [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP
@ 2026-10-09 2:49 Kumar Kartikeya Dwivedi
2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
` (2 more replies)
0 siblings, 3 replies; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen,
Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar,
Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86,
linux-kernel
Privileged BPF programs may dereference pointers the verifier cannot prove
valid: fields walked through untrusted pointers, bpf_core_cast() results,
legacy kptr map values. The verifier marks those loads PROBE_MEM and the
JIT attaches an exception table entry to each, so a fault on an invalid
kernel address zeroes the destination register and the program continues.
On x86 that is not enough. With SMAP enabled, do_user_addr_fault() treats a
kernel-mode fault on a user address as a kernel bug and oopses without
consulting the exception table. The JIT therefore precedes every PROBE_MEM
load with a range check against TASK_SIZE_MAX and the vsyscall page: nine
instructions and 32 to 39 bytes in front of a load of four to seven bytes.
Patch 1 lets do_user_addr_fault() resolve the exception table entry when
the faulting instruction belongs to a BPF program. The lookup sits on the
unlikely() branch that ends in page_fault_oops() today, so no fault that is
handled today runs any new code, and a kernel-mode user access without STAC
from anything but JITed BPF code still oopses. Patch 2 drops the range
check from the JIT when SMAP is enabled, leaving the bare load and its
exception table entry, which is what the arm64, riscv, s390 and loongarch
JITs already emit. Kernels without SMAP keep the check.
The series does not change which addresses a program may read. SMAP still
forbids reading user memory, PROBE_MEM loads of kernel addresses are
handled as before, and confidential guests, where reading an unaccepted
page is fatal at any kernel address, are no better and no worse off than
with copy_from_kernel_nofault() today.
Patch 3 adds a selftest that reads through NULL, a low user address, the
last user page, a non-canonical address and the vsyscall page. It passes
before and after the series and oopses on a kernel with only patch 2
applied.
The fault.c change needs an x86 Ack. It can go through tip with the JIT
change following later, or through bpf-next with the Ack, whichever the x86
maintainers prefer.
Results
-------
JITed code size measured with veristat over every BPF selftests object and
over 466 production BPF objects from Meta's fleet, on an x86-64 guest with
SMAP, bpf-next at e1d84a37cba9 with and without the series:
programs with PROBE_MEM reduction per program
mean median max
BPF selftests 3324 91 35.0% 37.5% 74.4%
Meta production programs 1710 228 7.0% 1.8% 69.2%
Programs that walk untrusted pointers lose up to three quarters of their
code and no program grows. Programs without PROBE_MEM loads are unchanged,
which is most of them: loads through trusted pointers and the probe_read
helpers do not use PROBE_MEM.
BPF selftests: 91 of 3324 programs contain PROBE_MEM loads. Their JITed code
shrinks by 35.0% on average, 37.5% median and 74.4% at most; no program
grows.
Largest relative savings (JITed bytes before->after, saved, percent):
type_cast/md_xdp 1039-> 266 -773 -74.4%
test_lwt_ip_encap/fexit_lwt_push_ip_encap 565-> 223 -342 -60.5%
sock_iter_batch/iter_udp_soreuse 851-> 347 -504 -59.2%
connect_unix_prog/connect_unix_prog 1927-> 791 -1136 -59.0%
sendmsg_unix_prog/sendmsg_unix_prog 1927-> 791 -1136 -59.0%
test_skc_to_unix_sock/unix_listen 492-> 212 -280 -56.9%
cgrp_ls_attach_cgroup/update_cookie_tracing 272-> 119 -153 -56.2%
bpf_smc/bpf_smc_release 139-> 61 -78 -56.1%
verifier_typedef/resolve_typedef 71-> 32 -39 -54.9%
type_cast/md_skb 291-> 135 -156 -53.6%
kfree_skb/fentry_eth_type_trans 148-> 70 -78 -52.7%
bpf_iter_vma_offset/get_vma_offset 527-> 254 -273 -51.8%
bpf_iter_tcp6/dump_tcp6 4386-> 2124 -2262 -51.6%
socket_cookie_prog/update_cookie_tracing 223-> 109 -114 -51.1%
mptcp_sock/_sockops 571-> 287 -284 -49.7%
sock_iter_batch/iter_tcp_soreuse 1437-> 733 -704 -49.0%
mptcp_subflow/_getsockopt_subflow 2079-> 1067 -1012 -48.7%
bpf_iter_bpf_sk_storage_helpers/fill_socket_owner 180-> 94 -86 -47.8%
bpf_iter_task_file/dump_task_file 656-> 344 -312 -47.6%
verifier_kfunc_perfmon/rdonly_cast_noperfmon 82-> 43 -39 -47.6%
bpf_iter_ipv6_route/dump_ipv6_route 997-> 527 -470 -47.1%
tracing_struct/test_struct_arg_11 83-> 44 -39 -47.0%
bpf_iter_ksym/dump_ksym 1219-> 649 -570 -46.8%
kfree_skb/fexit_eth_type_trans 168-> 90 -78 -46.4%
bpf_iter_unix/dump_unix 1268-> 686 -582 -45.9%
test_signed_loader_lsm/inspect_prog_load 273-> 148 -125 -45.8%
bpf_iter_tcp4/dump_tcp4 3485-> 1904 -1581 -45.4%
bpf_iter_netlink/dump_netlink 887-> 497 -390 -44.0%
sk_storage_omem_uncharge/bpf_sk_storage_free 190-> 108 -82 -43.2%
bpf_iter_setsockopt/change_tcp_cc 562-> 324 -238 -42.4%
bpf_iter_task_vmas/proc_maps 1012-> 588 -424 -41.9%
test_module_attach/handle_fexit_ret 170-> 99 -71 -41.8%
mptcp_sock/trace_mptcp_pm_new_connection 94-> 55 -39 -41.5%
rcu_read_lock/rcu_untrusted_union_ld 97-> 58 -39 -40.2%
test_subprogs_extable/handle_fexit_ret_subprogs 177-> 106 -71 -40.1%
test_subprogs_extable/handle_fexit_ret_subprogs2 177-> 106 -71 -40.1%
test_subprogs_extable/handle_fexit_ret_subprogs3 177-> 106 -71 -40.1%
core_kern/fentry_eth_type_trans 292-> 175 -117 -40.1%
core_kern/fexit_eth_type_trans 292-> 175 -117 -40.1%
bpf_smc/smc_run 226-> 136 -90 -39.8%
test_wakeup_source/iterate_wakeupsources 1150-> 696 -454 -39.5%
lsm/test_sys_setdomainname 199-> 121 -78 -39.2%
lsm_bdev/bdev_free_security 100-> 61 -39 -39.0%
bpf_smc/bpf_smc_switch_to_fallback 103-> 64 -39 -37.9%
test_ldsx_insn/test_ptr_struct_arg 85-> 53 -32 -37.6%
bpf_iter_setsockopt_unix/change_sndbuf 1060-> 662 -398 -37.5%
fentry_test/test8 88-> 56 -32 -36.4%
fexit_test/test8 88-> 56 -32 -36.4%
find_vma/handle_getpid 443-> 287 -156 -35.2%
find_vma/handle_pe 443-> 287 -156 -35.2%
Largest absolute savings (JITed bytes before->after, saved, percent):
bpf_iter_tcp6/dump_tcp6 4386-> 2124 -2262 -51.6%
bpf_iter_tcp4/dump_tcp4 3485-> 1904 -1581 -45.4%
connect_unix_prog/connect_unix_prog 1927-> 791 -1136 -59.0%
sendmsg_unix_prog/sendmsg_unix_prog 1927-> 791 -1136 -59.0%
mptcp_subflow/_getsockopt_subflow 2079-> 1067 -1012 -48.7%
type_cast/md_xdp 1039-> 266 -773 -74.4%
sock_iter_batch/iter_tcp_soreuse 1437-> 733 -704 -49.0%
setget_sockopt/skops_sockopt 5279-> 4602 -677 -12.8%
bpf_iter_unix/dump_unix 1268-> 686 -582 -45.9%
bpf_iter_ksym/dump_ksym 1219-> 649 -570 -46.8%
sock_iter_batch/iter_udp_soreuse 851-> 347 -504 -59.2%
bpf_iter_ipv6_route/dump_ipv6_route 997-> 527 -470 -47.1%
test_wakeup_source/iterate_wakeupsources 1150-> 696 -454 -39.5%
bpf_iter_task_vmas/proc_maps 1012-> 588 -424 -41.9%
bpf_iter_setsockopt_unix/change_sndbuf 1060-> 662 -398 -37.5%
bpf_iter_netlink/dump_netlink 887-> 497 -390 -44.0%
test_lwt_ip_encap/fexit_lwt_push_ip_encap 565-> 223 -342 -60.5%
bpf_iter_task_file/dump_task_file 656-> 344 -312 -47.6%
mptcp_sock/_sockops 571-> 287 -284 -49.7%
test_skc_to_unix_sock/unix_listen 492-> 212 -280 -56.9%
bpf_iter_vma_offset/get_vma_offset 527-> 254 -273 -51.8%
test_tc_tunnel/decap_f 1262-> 989 -273 -21.6%
bpf_iter_setsockopt/change_tcp_cc 562-> 324 -238 -42.4%
net_timestamping/skops_sockopt 1153-> 993 -160 -13.9%
cgroup_hierarchical_stats/flusher 534-> 377 -157 -29.4%
bpf_ma_ttrace/check_ttrace 665-> 509 -156 -23.5%
find_vma/handle_getpid 443-> 287 -156 -35.2%
find_vma/handle_pe 443-> 287 -156 -35.2%
type_cast/md_skb 291-> 135 -156 -53.6%
cgrp_ls_attach_cgroup/update_cookie_tracing 272-> 119 -153 -56.2%
kfree_skb/trace_kfree_skb 569-> 418 -151 -26.5%
test_signed_loader_lsm/inspect_prog_load 273-> 148 -125 -45.8%
core_kern/fentry_eth_type_trans 292-> 175 -117 -40.1%
core_kern/fexit_eth_type_trans 292-> 175 -117 -40.1%
test_ksyms_btf/handler 376-> 259 -117 -31.1%
socket_cookie_prog/update_cookie_tracing 223-> 109 -114 -51.1%
bpf_qdisc_fq/bpf_fq_init 845-> 734 -111 -13.1%
sk_bypass_prot_mem/fentry_tcp_init_sock 429-> 319 -110 -25.6%
sk_bypass_prot_mem/fentry_udp_init_sock 429-> 319 -110 -25.6%
verifier_global_ptr_args/anything_to_untrusted_m~ 362-> 266 -96 -26.5%
bpf_smc/smc_run 226-> 136 -90 -39.8%
bpf_iter_bpf_sk_storage_helpers/fill_socket_owner 180-> 94 -86 -47.8%
cgroup_ancestor/log_cgroup_id 248-> 162 -86 -34.7%
test_btf_skc_cls_ingress/cls_ingress 1211-> 1126 -85 -7.0%
test_tcpbpf_kern/bpf_testcb 1113-> 1029 -84 -7.5%
sk_storage_omem_uncharge/bpf_sk_storage_free 190-> 108 -82 -43.2%
test_cgroup1_hierarchy/lsm_s_run 363-> 284 -79 -21.8%
bpf_smc/bpf_smc_release 139-> 61 -78 -56.1%
kfree_skb/fentry_eth_type_trans 148-> 70 -78 -52.7%
kfree_skb/fexit_eth_type_trans 168-> 90 -78 -46.4%
Meta production programs: 228 of 1710 programs contain PROBE_MEM loads.
Their JITed code shrinks by 7.0% on average, 1.8% median and 69.2% at most;
no program grows. Programs are listed by category only.
Largest relative savings (JITed bytes before->after, saved, percent):
security enforcement 7 452-> 139 -313 -69.2%
TCP congestion control 5 841-> 267 -574 -68.2%
container runtime 1 789-> 315 -474 -60.1%
network tuning 2 300-> 132 -168 -56.0%
security monitoring 71 147-> 69 -78 -53.1%
profiling 3 604-> 319 -285 -47.2%
load balancing 3 426-> 270 -156 -36.6%
host monitoring 3 1170-> 816 -354 -30.3%
security monitoring 15 422-> 296 -126 -29.9%
profiling 2 531-> 375 -156 -29.4%
security monitoring 75 822-> 588 -234 -28.5%
security monitoring 76 861-> 627 -234 -27.2%
security enforcement 2 1781-> 1341 -440 -24.7%
load balancing 1 1441-> 1093 -348 -24.1%
load balancing 2 1441-> 1093 -348 -24.1%
security monitoring 31 137-> 105 -32 -23.4%
network telemetry 3 488-> 377 -111 -22.8%
TCP congestion control 1 2760-> 2136 -624 -22.6%
network telemetry 2 493-> 382 -111 -22.5%
security monitoring 21 360-> 288 -72 -20.0%
security monitoring 23 360-> 288 -72 -20.0%
network telemetry 6 1130-> 907 -223 -19.7%
network telemetry 1 407-> 328 -79 -19.4%
security enforcement 52 201-> 162 -39 -19.4%
security enforcement 54 206-> 167 -39 -18.9%
security enforcement 55 206-> 167 -39 -18.9%
security monitoring 80 1095-> 896 -199 -18.2%
TCP congestion control 6 4153-> 3529 -624 -15.0%
security monitoring 19 272-> 233 -39 -14.3%
storage telemetry 7 1433-> 1230 -203 -14.2%
storage telemetry 1 1436-> 1233 -203 -14.1%
storage telemetry 5 1436-> 1233 -203 -14.1%
storage telemetry 9 1436-> 1233 -203 -14.1%
storage telemetry 11 1436-> 1233 -203 -14.1%
storage telemetry 17 1436-> 1233 -203 -14.1%
storage telemetry 2 1444-> 1241 -203 -14.1%
storage telemetry 3 1444-> 1241 -203 -14.1%
storage telemetry 10 1444-> 1241 -203 -14.1%
storage telemetry 12 1444-> 1241 -203 -14.1%
storage telemetry 13 1444-> 1241 -203 -14.1%
storage telemetry 14 1444-> 1241 -203 -14.1%
storage telemetry 15 1444-> 1241 -203 -14.1%
storage telemetry 16 1444-> 1241 -203 -14.1%
security enforcement 22 680-> 586 -94 -13.8%
security monitoring 81 1420-> 1224 -196 -13.8%
security monitoring 67 896-> 786 -110 -12.3%
host monitoring 4 966-> 849 -117 -12.1%
host monitoring 5 288-> 256 -32 -11.1%
host monitoring 1 1785-> 1588 -197 -11.0%
network tuning 4 2929-> 2613 -316 -10.8%
Largest absolute savings (JITed bytes before->after, saved, percent):
security enforcement 40 12217->11588 -629 -5.2%
TCP congestion control 1 2760-> 2136 -624 -22.6%
TCP congestion control 3 8456-> 7832 -624 -7.4%
TCP congestion control 6 4153-> 3529 -624 -15.0%
security enforcement 46 14109->13486 -623 -4.4%
security enforcement 48 13941->13319 -622 -4.5%
security enforcement 42 13783->13193 -590 -4.3%
TCP congestion control 5 841-> 267 -574 -68.2%
security enforcement 26 11883->11332 -551 -4.6%
security enforcement 44 9468-> 8917 -551 -5.8%
security monitoring 68 10867->10324 -543 -5.0%
security enforcement 8 34695->34181 -514 -1.5%
security enforcement 9 25952->25440 -512 -2.0%
security enforcement 12 25993->25481 -512 -2.0%
security enforcement 13 25993->25481 -512 -2.0%
security enforcement 14 25993->25481 -512 -2.0%
security enforcement 10 26290->25779 -511 -1.9%
container runtime 1 789-> 315 -474 -60.1%
security enforcement 3 25670->25202 -468 -1.8%
security enforcement 2 1781-> 1341 -440 -24.7%
security enforcement 58 9190-> 8831 -359 -3.9%
security enforcement 11 29176->28821 -355 -1.2%
security enforcement 57 23466->23111 -355 -1.5%
host monitoring 3 1170-> 816 -354 -30.3%
load balancing 1 1441-> 1093 -348 -24.1%
load balancing 2 1441-> 1093 -348 -24.1%
security enforcement 16 25842->25526 -316 -1.2%
network tuning 3 2957-> 2641 -316 -10.7%
network tuning 4 2929-> 2613 -316 -10.8%
security enforcement 7 452-> 139 -313 -69.2%
security monitoring 3 20315->20004 -311 -1.5%
security enforcement 45 13025->12715 -310 -2.4%
security enforcement 47 12857->12548 -309 -2.4%
security enforcement 60 7751-> 7445 -306 -4.0%
profiling 3 604-> 319 -285 -47.2%
sched_ext scheduler 9 3389-> 3105 -284 -8.4%
security enforcement 1 10406->10129 -277 -2.7%
security enforcement 4 9267-> 8990 -277 -3.0%
security enforcement 5 20127->19850 -277 -1.4%
security enforcement 17 31234->30957 -277 -0.9%
security enforcement 18 8439-> 8162 -277 -3.3%
security enforcement 19 8439-> 8162 -277 -3.3%
security enforcement 20 8439-> 8162 -277 -3.3%
security enforcement 27 10654->10377 -277 -2.6%
security enforcement 28 11089->10812 -277 -2.5%
security enforcement 29 8482-> 8205 -277 -3.3%
security enforcement 30 8479-> 8202 -277 -3.3%
security enforcement 31 8492-> 8215 -277 -3.3%
security enforcement 32 16529->16252 -277 -1.7%
security enforcement 33 16387->16110 -277 -1.7%
The branchless mask Peter suggested for v3 (lea, mov, sar, inc, and, dec in
front of the load, no x86/mm change), measured the same way, gives less
than half of that: 16.2% mean and 35.7% max for the selftests programs,
3.3% mean and 31.9% max for the production programs, with five dependent
ALU instructions left in front of every load. It also keeps the vsyscall
problem: an address with the top bit set passes the mask unchanged, so a
PROBE_MEM load from the vsyscall page still oopses under SMAP, and the
selftest from patch 3 crashes a kernel with that variant at exactly that
address.
Changelog:
----------
v3 -> v4
v3: https://lore.kernel.org/bpf/20241103193512.4076710-1-memxor@gmail.com
* Rebase on bpf-next.
* Drop "zero overhead" from the title and state the fault handler cost up
front. (Dave)
* Spell out what the series does and does not change for confidential
guests. (Dave)
* Measure code size over the selftests and Meta production objects, for
this series and for the branchless mask alternative. (Peter)
* Add a selftest for PROBE_MEM loads from user, vsyscall and non-canonical
addresses.
* Call fixup_exception() directly in the SMAP branch and fall through to
the existing page_fault_oops() instead of kernelmode_fixup_or_oops().
* Restructure the JIT change around probe_mem and bounds_check flags
instead of a goto; drop Puranjay's Ack on it since the code changed.
v2 -> v3
v2: https://lore.kernel.org/bpf/20240619092216.1780946-1-memxor@gmail.com
* Rebase on bpf-next
* Add Puranjay's Acks
v1 -> v2
v1: https://lore.kernel.org/bpf/20240515233932.3733815-1-memxor@gmail.com
* Rebase on bpf-next
Kumar Kartikeya Dwivedi (3):
x86/mm: Resolve BPF exception fixups for user address faults under
SMAP
bpf, x86: Skip the PROBE_MEM address range check under SMAP
selftests/bpf: Test PROBE_MEM loads from invalid addresses
arch/x86/mm/fault.c | 11 +++
arch/x86/net/bpf_jit_comp.c | 21 ++++--
.../bpf/prog_tests/probe_mem_fault.c | 69 +++++++++++++++++++
.../selftests/bpf/progs/probe_mem_fault.c | 41 +++++++++++
4 files changed, 137 insertions(+), 5 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c
create mode 100644 tools/testing/selftests/bpf/progs/probe_mem_fault.c
base-commit: e1d84a37cba984388988d2f1ddc84561413f0db2
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP 2026-10-09 2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 ` Kumar Kartikeya Dwivedi 2026-10-09 4:13 ` Borislav Petkov 2026-10-09 2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi 2 siblings, 1 reply; 7+ messages in thread From: Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 UTC (permalink / raw) To: bpf Cc: Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Borislav Petkov, kkd, kernel-team, x86, linux-kernel With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a user address with EFLAGS.AC clear as a kernel bug: it does not consult the exception table and oopses right away. That is correct for ordinary kernel code, where get_kernel_nofault() and the other nofault accessors never let a user address reach a faulting instruction. JITed BPF programs are different. A privileged program may dereference a pointer the verifier cannot prove valid, and the verifier marks such loads PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM load, so that a fault on an unmapped kernel address zeroes the destination register and the program continues. Since a user address would oops instead, the x86 JIT also emits an address range check in front of every PROBE_MEM load, which keeps user addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away from the load. The check duplicates the fault handler's knowledge of the address space layout, got the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM runtime load check"), and costs nine instructions and 32 to 39 bytes of code per load. When the faulting instruction belongs to a BPF program, resolve its exception table entry instead of oopsing, exactly as is done for faults on kernel addresses. The is_bpf_text_address() lookup sits inside the unlikely() SMAP branch that currently ends in page_fault_oops(), so no path that does not oops today executes any additional code, and the oops itself is unchanged when no entry matches. Non-BPF code keeps the existing behaviour: a kernel-mode user access without STAC still oopses, extable entry or not. The next patch uses this to drop the range check from the JIT when SMAP is enabled. Nothing changes about which addresses a BPF program may read: a user address never becomes readable, since SMAP forbids the access, and a PROBE_MEM load of a kernel address is handled as before. Without SMAP the JIT keeps its range check and this path is never reached. Acked-by: Puranjay Mohan <puranjay@kernel.org> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> --- arch/x86/mm/fault.c | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c index aa88370ce739..2060e5f35d77 100644 --- a/arch/x86/mm/fault.c +++ b/arch/x86/mm/fault.c @@ -20,6 +20,7 @@ #include <linux/mm_types.h> #include <linux/mm.h> /* find_and_lock_vma() */ #include <linux/vmalloc.h> +#include <linux/filter.h> /* is_bpf_text_address() */ #include <asm/cpufeature.h> /* boot_cpu_has, ... */ #include <asm/traps.h> /* dotraplinkage, ... */ @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs, if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) && !(error_code & X86_PF_USER) && !(regs->flags & X86_EFLAGS_AC))) { + /* + * JITed BPF programs dereference untrusted pointers with loads + * that carry an exception table entry (PROBE_MEM). SMAP makes + * sure such a load cannot read user memory, so resolve the + * fault through the exception table, as for an unmapped kernel + * address, instead of oopsing. + */ + if (is_bpf_text_address(regs->ip) && + fixup_exception(regs, X86_TRAP_PF, error_code, address)) + return; /* * No extable entry here. This was a kernel access to an * invalid pointer. get_kernel_nofault() will not get here. -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP 2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi @ 2026-10-09 4:13 ` Borislav Petkov 2026-10-09 15:23 ` Kumar Kartikeya Dwivedi 0 siblings, 1 reply; 7+ messages in thread From: Borislav Petkov @ 2026-10-09 4:13 UTC (permalink / raw) To: Kumar Kartikeya Dwivedi Cc: bpf, Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar, kkd, kernel-team, x86, linux-kernel On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote: > With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a > user address with EFLAGS.AC clear as a kernel bug: it does not consult the > exception table and oopses right away. That is correct for ordinary kernel > code, where get_kernel_nofault() and the other nofault accessors never let > a user address reach a faulting instruction. > > JITed BPF programs are different. A privileged program may dereference a > pointer the verifier cannot prove valid, and the verifier marks such loads > PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM > load, so that a fault on an unmapped kernel address zeroes the destination > register and the program continues. Since a user address would oops > instead, the x86 JIT also emits an address range check in front of every > PROBE_MEM load, which keeps user addresses, the guard page above > TASK_SIZE_MAX and the vsyscall page away from the load. The check > duplicates the fault handler's knowledge of the address space layout, got > the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM > runtime load check"), and costs nine instructions and 32 to 39 bytes of > code per load. I can't parse that. Why do bpf memory accesses need to be handled differently than any other memory accesses when SMAP is enabled? Perhaps you should give a concrete example. And why can't all that gunk be resolved at program load instead of going all the way in the #PF handler? > When the faulting instruction belongs to a BPF program, resolve its > exception table entry instead of oopsing, exactly as is done for faults on > kernel addresses. The is_bpf_text_address() lookup sits inside the > unlikely() SMAP branch that currently ends in page_fault_oops(), so no path > that does not oops today executes any additional code, and the oops itself > is unchanged when no entry matches. Non-BPF code keeps the existing > behaviour: a kernel-mode user access without STAC still oopses, extable > entry or not. > > The next patch uses this to drop the range check from the JIT when SMAP is There's no next patch and previous patch in git history. > enabled. Nothing changes about which addresses a BPF program may read: a > user address never becomes readable, since SMAP forbids the access, and a > PROBE_MEM load of a kernel address is handled as before. Without SMAP the > JIT keeps its range check and this path is never reached. > > Acked-by: Puranjay Mohan <puranjay@kernel.org> > Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> > --- > arch/x86/mm/fault.c | 11 +++++++++++ > 1 file changed, 11 insertions(+) > > diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c > index aa88370ce739..2060e5f35d77 100644 > --- a/arch/x86/mm/fault.c > +++ b/arch/x86/mm/fault.c > @@ -20,6 +20,7 @@ > #include <linux/mm_types.h> > #include <linux/mm.h> /* find_and_lock_vma() */ > #include <linux/vmalloc.h> > +#include <linux/filter.h> /* is_bpf_text_address() */ > > #include <asm/cpufeature.h> /* boot_cpu_has, ... */ > #include <asm/traps.h> /* dotraplinkage, ... */ > @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs, > if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) && > !(error_code & X86_PF_USER) && > !(regs->flags & X86_EFLAGS_AC))) { > + /* > + * JITed BPF programs dereference untrusted pointers with loads > + * that carry an exception table entry (PROBE_MEM). SMAP makes > + * sure such a load cannot read user memory, so resolve the > + * fault through the exception table, as for an unmapped kernel > + * address, instead of oopsing. > + */ > + if (is_bpf_text_address(regs->ip) && > + fixup_exception(regs, X86_TRAP_PF, error_code, address)) > + return; I'm not at all amused from this adding bpf-specific handling to the #PF handler, TBH... -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP 2026-10-09 4:13 ` Borislav Petkov @ 2026-10-09 15:23 ` Kumar Kartikeya Dwivedi 0 siblings, 0 replies; 7+ messages in thread From: Kumar Kartikeya Dwivedi @ 2026-10-09 15:23 UTC (permalink / raw) To: Borislav Petkov Cc: bpf, Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar, kkd, kernel-team, x86, linux-kernel On Fri Oct 9, 2026 at 6:13 AM CEST, Borislav Petkov wrote: > On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote: >> With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a >> user address with EFLAGS.AC clear as a kernel bug: it does not consult the >> exception table and oopses right away. That is correct for ordinary kernel >> code, where get_kernel_nofault() and the other nofault accessors never let >> a user address reach a faulting instruction. >> >> JITed BPF programs are different. A privileged program may dereference a >> pointer the verifier cannot prove valid, and the verifier marks such loads >> PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM >> load, so that a fault on an unmapped kernel address zeroes the destination >> register and the program continues. Since a user address would oops >> instead, the x86 JIT also emits an address range check in front of every >> PROBE_MEM load, which keeps user addresses, the guard page above >> TASK_SIZE_MAX and the vsyscall page away from the load. The check >> duplicates the fault handler's knowledge of the address space layout, got >> the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM >> runtime load check"), and costs nine instructions and 32 to 39 bytes of >> code per load. > > I can't parse that. Why do bpf memory accesses need to be handled differently > than any other memory accesses when SMAP is enabled? > It requires the same handling as get_kernel_nofault() as if that were inlined. The point of the patch is to be able to avoid bounds checking in the JITed code. We can probably extend this to other nofault accessors as well, but I wanted to keep the scope limited to BPF for now, because the cost of the bounds checking is significant in BPF programs. For background: Privileged programs are allowed to trace kernel code and dereference pointers where establishing whether the pointer is pointing to a valid object is not possible during verification. Program authors know about this behavior, most of the use cases for this are reading data, collecting statistics, and so on. It is a more convenient and efficient way to walk a chain of pointers than to use the equivalent of copy_from_kernel_nofault() in a loop. get_kernel_nofault() does user address checks in software, in copy_from_kernel_nofault_allowed(), because the SMAP branch in do_user_addr_fault() treats any kernel-mode fault on a user address as a missing STAC and oopses without looking at the extable. The x86 JIT does the same today, inlined, in front of every such load in a BPF program. > Perhaps you should give a concrete example. Consider this example: SEC("fentry/tcp_retransmit_skb") <- attaches to the entry of tcp_retransmit_skb() int BPF_PROG(retrans, struct sock *sk) { struct file *f = sk->sk_socket->file; This program was written for doing a TCP related investigation. The first load of sk is fine, but sk_socket can be NULL (orphaned socket) or stale, so the second load becomes a mov with an extable entry. Below is the JITed sequence. movq $-10485760, %r10 movq %rax, %r11 addq $2200, %r11 subq %r10, %r11 movabsq $72057594048413696, %r10 cmpq %r10, %r11 ja load xorl %edi, %edi jmp done load: movq 2200(%rax), %rdi done: Nine instructions and 32 to 39 bytes to guard a every load instruction against an access that SMAP, when enabled, will block for us, and which we could fixup, if we had exception handling in its fault path. So the overall idea proposal is to eliminate this sequence and invoke the fixup_exception() handler for the faulting instruction instead, which will zero the destination register and continue execution, just like it does for faults on kernel addresses already. > > And why can't all that gunk be resolved at program load instead of going all > the way in the #PF handler? > This is what we do now, but it is a significant amount of code preceding each load instruction. If you look at some examples in the cover letter, we can shave the text size of programs by more than half in extreme cases. Do note that we already do fixup_exception() for kernel address faults, so this is mostly trying to mirror that behavior for the rest and eliminate the bounds checking. >> When the faulting instruction belongs to a BPF program, resolve its >> exception table entry instead of oopsing, exactly as is done for faults on >> kernel addresses. The is_bpf_text_address() lookup sits inside the >> unlikely() SMAP branch that currently ends in page_fault_oops(), so no path >> that does not oops today executes any additional code, and the oops itself >> is unchanged when no entry matches. Non-BPF code keeps the existing >> behaviour: a kernel-mode user access without STAC still oopses, extable >> entry or not. >> >> The next patch uses this to drop the range check from the JIT when SMAP is > > There's no next patch and previous patch in git history. Yeah, I will reword this bit in the commit log for v2. > >> enabled. Nothing changes about which addresses a BPF program may read: a >> user address never becomes readable, since SMAP forbids the access, and a >> PROBE_MEM load of a kernel address is handled as before. Without SMAP the >> JIT keeps its range check and this path is never reached. >> >> Acked-by: Puranjay Mohan <puranjay@kernel.org> >> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> >> --- >> arch/x86/mm/fault.c | 11 +++++++++++ >> 1 file changed, 11 insertions(+) >> >> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c >> index aa88370ce739..2060e5f35d77 100644 >> --- a/arch/x86/mm/fault.c >> +++ b/arch/x86/mm/fault.c >> @@ -20,6 +20,7 @@ >> #include <linux/mm_types.h> >> #include <linux/mm.h> /* find_and_lock_vma() */ >> #include <linux/vmalloc.h> >> +#include <linux/filter.h> /* is_bpf_text_address() */ >> >> #include <asm/cpufeature.h> /* boot_cpu_has, ... */ >> #include <asm/traps.h> /* dotraplinkage, ... */ >> @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs, >> if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) && >> !(error_code & X86_PF_USER) && >> !(regs->flags & X86_EFLAGS_AC))) { >> + /* >> + * JITed BPF programs dereference untrusted pointers with loads >> + * that carry an exception table entry (PROBE_MEM). SMAP makes >> + * sure such a load cannot read user memory, so resolve the >> + * fault through the exception table, as for an unmapped kernel >> + * address, instead of oopsing. >> + */ >> + if (is_bpf_text_address(regs->ip) && >> + fixup_exception(regs, X86_TRAP_PF, error_code, address)) >> + return; > > I'm not at all amused from this adding bpf-specific handling to the #PF > handler, TBH... I understand, and perhaps this can be adjusted to be a little different, but on the flip side, at this point, the kernel is about to print an OOPS. I made it BPF specific because we really don't expect to fixup exceptions for anything else here. We already invoke fixup_exception() for kernel address faults. The advantage is losing significant amount of code size in JITed programs and the associated runtime overhead (you can see the numbers in the cover letter). Do note that we can make it more generic, we can follow up and optimize copy_from_kernel_nofault() with cpu_feature_enabled(X86_FEATURE_SMAP) as well. But I didn't want to go there just yet. I am happy to work on a follow up if so desired. ^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check under SMAP 2026-10-09 2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 ` Kumar Kartikeya Dwivedi 2026-10-09 3:42 ` bot+bpf-ci 2026-10-09 2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi 2 siblings, 1 reply; 7+ messages in thread From: Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 UTC (permalink / raw) To: bpf Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86, linux-kernel The x86 JIT guards every PROBE_MEM load with a range check that keeps user addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away from the load, because a kernel-mode fault on those addresses would oops under SMAP instead of reaching the load's exception table entry. The check is nine instructions and 39 bytes (32 when the offset is zero) in front of a load of a few bytes. For "r7 = *(u64 *)(r0 + 2200)" on a 5-level paging kernel, with VSYSCALL_ADDR and TASK_SIZE_MAX + PAGE_SIZE - VSYSCALL_ADDR as the two constants: movq $-10485760, %r10 movq %rax, %r11 addq $2200, %r11 subq %r10, %r11 movabsq $72057594048413696, %r10 cmpq %r10, %r11 ja load xorl %edi, %edi jmp done load: movq 2200(%rax), %rdi done: The previous patch made do_user_addr_fault() resolve the exception table entries of BPF programs for faults on user addresses when SMAP is enabled. On such kernels, emit the bare load with its exception table entry, as the arm64, riscv, s390 and loongarch JITs already do. All BPF programs run with SMAP active once the CPU feature is enabled, and the feature cannot change after boot, so checking it at JIT time is sufficient. Kernels without SMAP, including those booted with nosmap, keep the range check. Measured with veristat over every object of the BPF selftests and over 466 production objects from Meta's fleet, on an x86-64 guest with SMAP, with and without this series on top of bpf-next: programs with PROBE_MEM reduction per program mean median max BPF selftests 3324 91 35.0% 37.5% 74.4% Meta production programs 1710 228 7.0% 1.8% 69.2% No program grows. The socket and task iterators of the selftests lose about half of their code, dump_tcp6 goes from 4386 to 2124 bytes, and the smallest production programs lose 60% to 69%. Loads through trusted pointers and the probe_read helpers do not use PROBE_MEM, which is why most programs are unaffected. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> --- arch/x86/net/bpf_jit_comp.c | 21 ++++++++++++++++----- 1 file changed, 16 insertions(+), 5 deletions(-) diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c index 083fcd6cf15b..793e7cd5a5c4 100644 --- a/arch/x86/net/bpf_jit_comp.c +++ b/arch/x86/net/bpf_jit_comp.c @@ -2133,6 +2133,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * const s32 imm32 = insn->imm; u32 dst_reg = insn->dst_reg; u32 src_reg = insn->src_reg; + bool probe_mem, bounds_check; bool accesses_stack_only; u8 b2 = 0, b3 = 0; u8 *start_of_ldx; @@ -2709,6 +2710,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * case BPF_LDX | BPF_PROBE_MEMSX | BPF_B: case BPF_LDX | BPF_PROBE_MEMSX | BPF_H: case BPF_LDX | BPF_PROBE_MEMSX | BPF_W: + probe_mem = BPF_MODE(insn->code) == BPF_PROBE_MEM || + BPF_MODE(insn->code) == BPF_PROBE_MEMSX; + /* + * With SMAP enabled, a load from a user address faults and + * do_user_addr_fault() resolves the exception table entry of the + * program, as for an unmapped kernel address, so the address range + * check is only needed without SMAP. + */ + bounds_check = probe_mem && !cpu_feature_enabled(X86_FEATURE_SMAP); insn_off = insn->off; if (src_reg == BPF_REG_PARAMS) { if (insn_off == 8) { @@ -2724,8 +2734,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * */ } - if (BPF_MODE(insn->code) == BPF_PROBE_MEM || - BPF_MODE(insn->code) == BPF_PROBE_MEMSX) { + if (bounds_check) { /* Conservatively check that src_reg + insn->off is a kernel address: * src_reg + insn->off > TASK_SIZE_MAX + PAGE_SIZE * and @@ -2772,6 +2781,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * /* populate jmp_offset for JAE above to jump to start_of_ldx */ start_of_ldx = prog; end_of_jmp[-1] = start_of_ldx - end_of_jmp; + } else if (probe_mem) { + start_of_ldx = prog; } else if (!accesses_stack_only) { err = emit_kasan_check(env, &prog, src_reg, insn_off, @@ -2785,14 +2796,14 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * emit_ldsx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off); else emit_ldx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off); - if (BPF_MODE(insn->code) == BPF_PROBE_MEM || - BPF_MODE(insn->code) == BPF_PROBE_MEMSX) { + if (probe_mem) { struct exception_table_entry *ex; u8 *_insn = image + proglen + (start_of_ldx - temp); s64 delta; /* populate jmp_offset for JMP above */ - start_of_ldx[-1] = prog - start_of_ldx; + if (bounds_check) + start_of_ldx[-1] = prog - start_of_ldx; if (!bpf_prog->aux->extable) break; -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check under SMAP 2026-10-09 2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi @ 2026-10-09 3:42 ` bot+bpf-ci 0 siblings, 0 replies; 7+ messages in thread From: bot+bpf-ci @ 2026-10-09 3:42 UTC (permalink / raw) To: memxor, bpf Cc: ast, andrii, daniel, eddyz87, emil, ihor.solodrai, dave.hansen, luto, peterz, tglx, mingo, bp, puranjay, kkd, kernel-team, x86, linux-kernel, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai [-- Attachment #1: Type: text/plain, Size: 8727 bytes --] > bpf, x86: Skip the PROBE_MEM address range check under SMAP TL;DR: - Possible kernel oops: unprivileged (CAP_PERFMON, no CAP_BPF) programs aren't in kallsyms, so a NULL PROBE_MEM load under SMAP may not be fixed up - Commit message wrongly credits arm64/riscv/loongarch; runtime cost unmentioned - Stale comments (LASS, uaddress_limit) > bpf, x86: Skip the PROBE_MEM address range check under SMAP > > The x86 JIT guards every PROBE_MEM load with a range check that keeps user > addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away > from the load, because a kernel-mode fault on those addresses would oops > under SMAP instead of reaching the load's exception table entry. The check > is nine instructions and 39 bytes (32 when the offset is zero) in front of > a load of a few bytes. For "r7 = *(u64 *)(r0 + 2200)" on a 5-level paging > kernel, with VSYSCALL_ADDR and TASK_SIZE_MAX + PAGE_SIZE - VSYSCALL_ADDR as > the two constants: > > movq $-10485760, %r10 > movq %rax, %r11 > addq $2200, %r11 > subq %r10, %r11 > movabsq $72057594048413696, %r10 > cmpq %r10, %r11 > ja load > xorl %edi, %edi > jmp done > load: > movq 2200(%rax), %rdi > done: > > The previous patch made do_user_addr_fault() resolve the exception table > entries of BPF programs for faults on user addresses when SMAP is enabled. > On such kernels, emit the bare load with its exception table entry, as the > arm64, riscv, s390 and loongarch JITs already do. All BPF programs run with > SMAP active once the CPU feature is enabled, and the feature cannot change > after boot, so checking it at JIT time is sufficient. Kernels without SMAP, > including those booted with nosmap, keep the range check. > > Measured with veristat over every object of the BPF selftests and over 466 > production objects from Meta's fleet, on an x86-64 guest with SMAP, with > and without this series on top of bpf-next: > > programs with PROBE_MEM reduction per program > mean median max > BPF selftests 3324 91 35.0% 37.5% 74.4% > Meta production programs 1710 228 7.0% 1.8% 69.2% > > No program grows. The socket and task iterators of the selftests lose about > half of their code, dump_tcp6 goes from 4386 to 2124 bytes, and the > smallest production programs lose 60% to 69%. Loads through trusted > pointers and the probe_read helpers do not use PROBE_MEM, which is why most > programs are unaffected. > > Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> The commit message says the JIT emits the bare load "as the arm64, riscv, s390 and loongarch JITs already do". Is that accurate for arm64, riscv and loongarch? Those three select ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE and use the weak bpf_arch_uaddress_limit(), which returns TASK_SIZE. For them bpf_do_misc_fixups() rewrites every PROBE_MEM and PROBE_MEMSX load into a guarded sequence before the JIT sees it: *patch++ = BPF_MOV64_REG(BPF_REG_AX, insn->src_reg); if (insn->off) *patch++ = BPF_ALU64_IMM(BPF_ADD, BPF_REG_AX, insn->off); *patch++ = BPF_ALU64_IMM(BPF_RSH, BPF_REG_AX, 32); *patch++ = BPF_JMP_IMM(BPF_JLE, BPF_REG_AX, uaddress_limit >> 32, 2); *patch++ = *insn; *patch++ = BPF_JMP_IMM(BPF_JA, 0, 0, 1); *patch++ = BPF_MOV64_IMM(insn->dst_reg, 0); So those JITs never emit a bare PROBE_MEM load for a user address. Only s390, which has separate address spaces, relies on the fault path in the way the changelog describes. With this patch, x86 with X86_FEATURE_SMAP set takes a full page fault, is_bpf_text_address() and the exception table search for a load of NULL plus an offset, where it used to execute a compare and a branch. The veristat numbers only measure code size. Could the changelog correct the arm64/riscv/loongarch statement and say something about the runtime cost of loads that now fault where they used to branch? > diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c > index 083fcd6cf15b..793e7cd5a5c4 100644 > --- a/arch/x86/net/bpf_jit_comp.c > +++ b/arch/x86/net/bpf_jit_comp.c [ ... ] > @@ -2709,6 +2710,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * > case BPF_LDX | BPF_PROBE_MEMSX | BPF_B: > case BPF_LDX | BPF_PROBE_MEMSX | BPF_H: > case BPF_LDX | BPF_PROBE_MEMSX | BPF_W: > + probe_mem = BPF_MODE(insn->code) == BPF_PROBE_MEM || > + BPF_MODE(insn->code) == BPF_PROBE_MEMSX; > + /* > + * With SMAP enabled, a load from a user address faults and > + * do_user_addr_fault() resolves the exception table entry of the > + * program, as for an unmapped kernel address, so the address range > + * check is only needed without SMAP. > + */ this isn't a bug, but is this comment accurate when X86_FEATURE_LASS is enabled? With LASS, a supervisor load from the user half of the address space raises a general protection fault rather than a page fault, and that is handled on a different path: exc_general_protection() -> gp_try_fixup_and_notify() -> fixup_exception() The outcome is the same, but the comment names do_user_addr_fault() as the place where the entry is resolved. > + bounds_check = probe_mem && !cpu_feature_enabled(X86_FEATURE_SMAP); > insn_off = insn->off; Can this oops for a program that is not in kallsyms? With bounds_check false, a PROBE_MEM load of a user address such as NULL plus a field offset depends only on the new branch in do_user_addr_fault(): if (is_bpf_text_address(regs->ip) && fixup_exception(regs, X86_TRAP_PF, error_code, address)) return; ... page_fault_oops(regs, error_code, address); Both is_bpf_text_address() and search_bpf_extables() only find programs that are registered in bpf_tree, and bpf_prog_kallsyms_add() skips programs whose loader lacks CAP_BPF: void bpf_prog_kallsyms_add(struct bpf_prog *fp) { if (!bpf_prog_kallsyms_candidate(fp) || !bpf_token_capable(fp->aux->token, CAP_BPF)) return; Such programs can still contain PROBE_MEM loads. With kernel.unprivileged_bpf_disabled=0, bpf_prog_load() accepts BPF_PROG_TYPE_SOCKET_FILTER and BPF_PROG_TYPE_CGROUP_SKB without CAP_BPF. With CAP_PERFMON, env->allow_ptr_leaks is true and bpf_sk_base_func_proto() exposes bpf_skc_to_tcp_sock() and the other skc_to_* helpers. skb->sk is readable as PTR_TO_SOCK_COMMON_OR_NULL, and sk_filter_trim_cap() sets skb->sk while the filter runs. The helper returns a plain PTR_TO_BTF_ID, and walking a pointer field from it, for example tp->inet_conn.icsk_ulp_ops->..., or sk_socket on an orphaned socket, gets rewritten to BPF_PROBE_MEM by bpf_convert_ctx_accesses(). The x86 JIT then emits an extable entry for it. Before this patch, a NULL plus offset load in such a program was caught by the JIT range check and dst was set to 0. After it, on a CPU with X86_FEATURE_SMAP the sequence would be: handle_page_fault() -> do_user_addr_fault() -> is_bpf_text_address() returns false -> page_fault_oops() That would be an oops in softirq context (sk_filter on the receive path) triggered by a user with CAP_PERFMON but without CAP_BPF. arm64, riscv and loongarch are not exposed to this, because bpf_do_misc_fixups() inserts the generic user address guard for them. Should the range check be kept when the program will not be in kallsyms (!bpf_token_capable(prog->aux->token, CAP_BPF)), or should every JITed program with exception entries be registered in bpf_tree? [ ... ] Separately, the comment above bpf_arch_uaddress_limit() in this file: arch/x86/net/bpf_jit_comp.c: /* x86-64 JIT emits its own code to filter user addresses so return 0 here */ u64 bpf_arch_uaddress_limit(void) { return 0; } is the stated reason x86 opts out of the generic verifier user address guard in bpf_do_misc_fixups(). Is it still correct? After this patch do_jit() emits no filter at all when cpu_feature_enabled(X86_FEATURE_SMAP), and user addresses are rejected only by the hardware check plus the fault handler fixup. Should the comment be updated, or should bpf_arch_uaddress_limit() be made consistent with the new JIT behavior? --- AI reviewed your patch. Please fix the bug or email reply why it's not a bug. See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37877475188 ^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses 2026-10-09 2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 ` Kumar Kartikeya Dwivedi 2 siblings, 0 replies; 7+ messages in thread From: Kumar Kartikeya Dwivedi @ 2026-10-09 2:49 UTC (permalink / raw) To: bpf Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86, linux-kernel Add a test that performs PROBE_MEM loads of three sizes through a bpf_core_cast() pointer whose value is chosen by userspace: NULL, a low user address, the last user page, a non-canonical address and, on x86-64, the vsyscall page and an offset into it. Each load must read zero and the kernel must survive. A load through the current task pointer checks that the same loads read real values when the address is valid. The test passes on kernels that reject the addresses with the JIT's range check and on kernels that rely on the fault handler; with only the JIT change of the previous patch applied, the first subtest oopses, which is what the fault handler change prevents. Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com> --- .../bpf/prog_tests/probe_mem_fault.c | 69 +++++++++++++++++++ .../selftests/bpf/progs/probe_mem_fault.c | 41 +++++++++++ 2 files changed, 110 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c create mode 100644 tools/testing/selftests/bpf/progs/probe_mem_fault.c diff --git a/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c new file mode 100644 index 000000000000..57a313e35eb6 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c @@ -0,0 +1,69 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#include <test_progs.h> +#include "probe_mem_fault.skel.h" + +#if defined(__x86_64__) +#include <asm/vsyscall.h> +#endif + +/* + * Addresses a PROBE_MEM load has to survive. Either the JIT's address check + * or the fault handler must turn each load into a zero result. + */ +static const struct { + const char *name; + unsigned long addr; +} bad_addrs[] = { + { "null", 0 }, + { "low_user", 4096 }, + { "last_user_page", (1UL << 47) - 4096 }, + { "non_canonical", 1UL << 63 }, +#if defined(__x86_64__) + { "vsyscall", VSYSCALL_ADDR }, + { "vsyscall_tail", VSYSCALL_ADDR + 0x800 }, +#endif +}; + +static void trigger(struct probe_mem_fault *skel, int *runs) +{ + skel->bss->val_dw = ~0ULL; + skel->bss->val_w = ~0U; + skel->bss->val_b = ~0; + usleep(1); + ASSERT_EQ(skel->bss->runs, ++*runs, "runs"); +} + +void test_probe_mem_fault(void) +{ + struct probe_mem_fault *skel; + int runs = 0, i; + + skel = probe_mem_fault__open_and_load(); + if (!ASSERT_OK_PTR(skel, "open_and_load")) + return; + + skel->bss->target_pid = getpid(); + if (!ASSERT_OK(probe_mem_fault__attach(skel), "attach")) + goto out; + + /* A valid kernel address is read for real. */ + skel->bss->use_current_task = true; + trigger(skel, &runs); + ASSERT_EQ(skel->bss->val_w, getpid(), "pid"); + ASSERT_NEQ(skel->bss->val_dw, 0, "start_time"); + ASSERT_NEQ(skel->bss->val_b, 0, "comm"); + + skel->bss->use_current_task = false; + for (i = 0; i < ARRAY_SIZE(bad_addrs); i++) { + if (!test__start_subtest(bad_addrs[i].name)) + continue; + skel->bss->addr = bad_addrs[i].addr; + trigger(skel, &runs); + ASSERT_EQ(skel->bss->val_dw, 0, "start_time"); + ASSERT_EQ(skel->bss->val_w, 0, "pid"); + ASSERT_EQ(skel->bss->val_b, 0, "comm"); + } +out: + probe_mem_fault__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/probe_mem_fault.c b/tools/testing/selftests/bpf/progs/probe_mem_fault.c new file mode 100644 index 000000000000..748be45418f5 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/probe_mem_fault.c @@ -0,0 +1,41 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#include <vmlinux.h> +#include <bpf/bpf_helpers.h> +#include <bpf/bpf_tracing.h> +#include <bpf/bpf_core_read.h> +#include "bpf_misc.h" + +char _license[] SEC("license") = "GPL"; + +int target_pid; +bool use_current_task; +unsigned long addr; +int runs; +__u64 val_dw; +__u32 val_w; +__u8 val_b; + +SEC("fentry/" SYS_PREFIX "sys_nanosleep") +int probe_mem_fault(void *ctx) +{ + struct task_struct *task; + unsigned long p = addr; + + if ((bpf_get_current_pid_tgid() >> 32) != target_pid) + return 0; + + if (use_current_task) + p = (unsigned long)bpf_get_current_task_btf(); + /* + * bpf_core_cast() yields an untrusted pointer, so every load through it + * is a PROBE_MEM load. Whatever the address is, the load must either + * read the field or produce zero; the kernel must not oops. + */ + task = bpf_core_cast((void *)p, struct task_struct); + val_dw = task->start_time; + val_w = task->pid; + val_b = task->comm[0]; + runs++; + return 0; +} -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-09 15:23 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-10-09 2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi 2026-10-09 4:13 ` Borislav Petkov 2026-10-09 15:23 ` Kumar Kartikeya Dwivedi 2026-10-09 2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi 2026-10-09 3:42 ` bot+bpf-ci 2026-10-09 2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®