mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP
@ 2026-10-09  2:49 Kumar Kartikeya Dwivedi
  2026-10-09  2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09  2:49 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen,
	Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86,
	linux-kernel

Privileged BPF programs may dereference pointers the verifier cannot prove
valid: fields walked through untrusted pointers, bpf_core_cast() results,
legacy kptr map values. The verifier marks those loads PROBE_MEM and the
JIT attaches an exception table entry to each, so a fault on an invalid
kernel address zeroes the destination register and the program continues.

On x86 that is not enough. With SMAP enabled, do_user_addr_fault() treats a
kernel-mode fault on a user address as a kernel bug and oopses without
consulting the exception table. The JIT therefore precedes every PROBE_MEM
load with a range check against TASK_SIZE_MAX and the vsyscall page: nine
instructions and 32 to 39 bytes in front of a load of four to seven bytes.

Patch 1 lets do_user_addr_fault() resolve the exception table entry when
the faulting instruction belongs to a BPF program. The lookup sits on the
unlikely() branch that ends in page_fault_oops() today, so no fault that is
handled today runs any new code, and a kernel-mode user access without STAC
from anything but JITed BPF code still oopses. Patch 2 drops the range
check from the JIT when SMAP is enabled, leaving the bare load and its
exception table entry, which is what the arm64, riscv, s390 and loongarch
JITs already emit. Kernels without SMAP keep the check.

The series does not change which addresses a program may read. SMAP still
forbids reading user memory, PROBE_MEM loads of kernel addresses are
handled as before, and confidential guests, where reading an unaccepted
page is fatal at any kernel address, are no better and no worse off than
with copy_from_kernel_nofault() today.

Patch 3 adds a selftest that reads through NULL, a low user address, the
last user page, a non-canonical address and the vsyscall page. It passes
before and after the series and oopses on a kernel with only patch 2
applied.

The fault.c change needs an x86 Ack. It can go through tip with the JIT
change following later, or through bpf-next with the Ack, whichever the x86
maintainers prefer.

Results
-------

JITed code size measured with veristat over every BPF selftests object and
over 466 production BPF objects from Meta's fleet, on an x86-64 guest with
SMAP, bpf-next at e1d84a37cba9 with and without the series:

                        programs  with PROBE_MEM  reduction per program
                                                  mean  median    max
  BPF selftests             3324              91  35.0%  37.5%  74.4%
  Meta production programs  1710             228   7.0%   1.8%  69.2%

Programs that walk untrusted pointers lose up to three quarters of their
code and no program grows. Programs without PROBE_MEM loads are unchanged,
which is most of them: loads through trusted pointers and the probe_read
helpers do not use PROBE_MEM.

BPF selftests: 91 of 3324 programs contain PROBE_MEM loads. Their JITed code
shrinks by 35.0% on average, 37.5% median and 74.4% at most; no program
grows.

  Largest relative savings (JITed bytes before->after, saved, percent):
  type_cast/md_xdp                                    1039->  266   -773  -74.4%
  test_lwt_ip_encap/fexit_lwt_push_ip_encap            565->  223   -342  -60.5%
  sock_iter_batch/iter_udp_soreuse                     851->  347   -504  -59.2%
  connect_unix_prog/connect_unix_prog                 1927->  791  -1136  -59.0%
  sendmsg_unix_prog/sendmsg_unix_prog                 1927->  791  -1136  -59.0%
  test_skc_to_unix_sock/unix_listen                    492->  212   -280  -56.9%
  cgrp_ls_attach_cgroup/update_cookie_tracing          272->  119   -153  -56.2%
  bpf_smc/bpf_smc_release                              139->   61    -78  -56.1%
  verifier_typedef/resolve_typedef                      71->   32    -39  -54.9%
  type_cast/md_skb                                     291->  135   -156  -53.6%
  kfree_skb/fentry_eth_type_trans                      148->   70    -78  -52.7%
  bpf_iter_vma_offset/get_vma_offset                   527->  254   -273  -51.8%
  bpf_iter_tcp6/dump_tcp6                             4386-> 2124  -2262  -51.6%
  socket_cookie_prog/update_cookie_tracing             223->  109   -114  -51.1%
  mptcp_sock/_sockops                                  571->  287   -284  -49.7%
  sock_iter_batch/iter_tcp_soreuse                    1437->  733   -704  -49.0%
  mptcp_subflow/_getsockopt_subflow                   2079-> 1067  -1012  -48.7%
  bpf_iter_bpf_sk_storage_helpers/fill_socket_owner    180->   94    -86  -47.8%
  bpf_iter_task_file/dump_task_file                    656->  344   -312  -47.6%
  verifier_kfunc_perfmon/rdonly_cast_noperfmon          82->   43    -39  -47.6%
  bpf_iter_ipv6_route/dump_ipv6_route                  997->  527   -470  -47.1%
  tracing_struct/test_struct_arg_11                     83->   44    -39  -47.0%
  bpf_iter_ksym/dump_ksym                             1219->  649   -570  -46.8%
  kfree_skb/fexit_eth_type_trans                       168->   90    -78  -46.4%
  bpf_iter_unix/dump_unix                             1268->  686   -582  -45.9%
  test_signed_loader_lsm/inspect_prog_load             273->  148   -125  -45.8%
  bpf_iter_tcp4/dump_tcp4                             3485-> 1904  -1581  -45.4%
  bpf_iter_netlink/dump_netlink                        887->  497   -390  -44.0%
  sk_storage_omem_uncharge/bpf_sk_storage_free         190->  108    -82  -43.2%
  bpf_iter_setsockopt/change_tcp_cc                    562->  324   -238  -42.4%
  bpf_iter_task_vmas/proc_maps                        1012->  588   -424  -41.9%
  test_module_attach/handle_fexit_ret                  170->   99    -71  -41.8%
  mptcp_sock/trace_mptcp_pm_new_connection              94->   55    -39  -41.5%
  rcu_read_lock/rcu_untrusted_union_ld                  97->   58    -39  -40.2%
  test_subprogs_extable/handle_fexit_ret_subprogs      177->  106    -71  -40.1%
  test_subprogs_extable/handle_fexit_ret_subprogs2     177->  106    -71  -40.1%
  test_subprogs_extable/handle_fexit_ret_subprogs3     177->  106    -71  -40.1%
  core_kern/fentry_eth_type_trans                      292->  175   -117  -40.1%
  core_kern/fexit_eth_type_trans                       292->  175   -117  -40.1%
  bpf_smc/smc_run                                      226->  136    -90  -39.8%
  test_wakeup_source/iterate_wakeupsources            1150->  696   -454  -39.5%
  lsm/test_sys_setdomainname                           199->  121    -78  -39.2%
  lsm_bdev/bdev_free_security                          100->   61    -39  -39.0%
  bpf_smc/bpf_smc_switch_to_fallback                   103->   64    -39  -37.9%
  test_ldsx_insn/test_ptr_struct_arg                    85->   53    -32  -37.6%
  bpf_iter_setsockopt_unix/change_sndbuf              1060->  662   -398  -37.5%
  fentry_test/test8                                     88->   56    -32  -36.4%
  fexit_test/test8                                      88->   56    -32  -36.4%
  find_vma/handle_getpid                               443->  287   -156  -35.2%
  find_vma/handle_pe                                   443->  287   -156  -35.2%

  Largest absolute savings (JITed bytes before->after, saved, percent):
  bpf_iter_tcp6/dump_tcp6                             4386-> 2124  -2262  -51.6%
  bpf_iter_tcp4/dump_tcp4                             3485-> 1904  -1581  -45.4%
  connect_unix_prog/connect_unix_prog                 1927->  791  -1136  -59.0%
  sendmsg_unix_prog/sendmsg_unix_prog                 1927->  791  -1136  -59.0%
  mptcp_subflow/_getsockopt_subflow                   2079-> 1067  -1012  -48.7%
  type_cast/md_xdp                                    1039->  266   -773  -74.4%
  sock_iter_batch/iter_tcp_soreuse                    1437->  733   -704  -49.0%
  setget_sockopt/skops_sockopt                        5279-> 4602   -677  -12.8%
  bpf_iter_unix/dump_unix                             1268->  686   -582  -45.9%
  bpf_iter_ksym/dump_ksym                             1219->  649   -570  -46.8%
  sock_iter_batch/iter_udp_soreuse                     851->  347   -504  -59.2%
  bpf_iter_ipv6_route/dump_ipv6_route                  997->  527   -470  -47.1%
  test_wakeup_source/iterate_wakeupsources            1150->  696   -454  -39.5%
  bpf_iter_task_vmas/proc_maps                        1012->  588   -424  -41.9%
  bpf_iter_setsockopt_unix/change_sndbuf              1060->  662   -398  -37.5%
  bpf_iter_netlink/dump_netlink                        887->  497   -390  -44.0%
  test_lwt_ip_encap/fexit_lwt_push_ip_encap            565->  223   -342  -60.5%
  bpf_iter_task_file/dump_task_file                    656->  344   -312  -47.6%
  mptcp_sock/_sockops                                  571->  287   -284  -49.7%
  test_skc_to_unix_sock/unix_listen                    492->  212   -280  -56.9%
  bpf_iter_vma_offset/get_vma_offset                   527->  254   -273  -51.8%
  test_tc_tunnel/decap_f                              1262->  989   -273  -21.6%
  bpf_iter_setsockopt/change_tcp_cc                    562->  324   -238  -42.4%
  net_timestamping/skops_sockopt                      1153->  993   -160  -13.9%
  cgroup_hierarchical_stats/flusher                    534->  377   -157  -29.4%
  bpf_ma_ttrace/check_ttrace                           665->  509   -156  -23.5%
  find_vma/handle_getpid                               443->  287   -156  -35.2%
  find_vma/handle_pe                                   443->  287   -156  -35.2%
  type_cast/md_skb                                     291->  135   -156  -53.6%
  cgrp_ls_attach_cgroup/update_cookie_tracing          272->  119   -153  -56.2%
  kfree_skb/trace_kfree_skb                            569->  418   -151  -26.5%
  test_signed_loader_lsm/inspect_prog_load             273->  148   -125  -45.8%
  core_kern/fentry_eth_type_trans                      292->  175   -117  -40.1%
  core_kern/fexit_eth_type_trans                       292->  175   -117  -40.1%
  test_ksyms_btf/handler                               376->  259   -117  -31.1%
  socket_cookie_prog/update_cookie_tracing             223->  109   -114  -51.1%
  bpf_qdisc_fq/bpf_fq_init                             845->  734   -111  -13.1%
  sk_bypass_prot_mem/fentry_tcp_init_sock              429->  319   -110  -25.6%
  sk_bypass_prot_mem/fentry_udp_init_sock              429->  319   -110  -25.6%
  verifier_global_ptr_args/anything_to_untrusted_m~    362->  266    -96  -26.5%
  bpf_smc/smc_run                                      226->  136    -90  -39.8%
  bpf_iter_bpf_sk_storage_helpers/fill_socket_owner    180->   94    -86  -47.8%
  cgroup_ancestor/log_cgroup_id                        248->  162    -86  -34.7%
  test_btf_skc_cls_ingress/cls_ingress                1211-> 1126    -85   -7.0%
  test_tcpbpf_kern/bpf_testcb                         1113-> 1029    -84   -7.5%
  sk_storage_omem_uncharge/bpf_sk_storage_free         190->  108    -82  -43.2%
  test_cgroup1_hierarchy/lsm_s_run                     363->  284    -79  -21.8%
  bpf_smc/bpf_smc_release                              139->   61    -78  -56.1%
  kfree_skb/fentry_eth_type_trans                      148->   70    -78  -52.7%
  kfree_skb/fexit_eth_type_trans                       168->   90    -78  -46.4%

Meta production programs: 228 of 1710 programs contain PROBE_MEM loads.
Their JITed code shrinks by 7.0% on average, 1.8% median and 69.2% at most;
no program grows. Programs are listed by category only.

  Largest relative savings (JITed bytes before->after, saved, percent):
  security enforcement 7      452->  139   -313  -69.2%
  TCP congestion control 5    841->  267   -574  -68.2%
  container runtime 1         789->  315   -474  -60.1%
  network tuning 2            300->  132   -168  -56.0%
  security monitoring 71      147->   69    -78  -53.1%
  profiling 3                 604->  319   -285  -47.2%
  load balancing 3            426->  270   -156  -36.6%
  host monitoring 3          1170->  816   -354  -30.3%
  security monitoring 15      422->  296   -126  -29.9%
  profiling 2                 531->  375   -156  -29.4%
  security monitoring 75      822->  588   -234  -28.5%
  security monitoring 76      861->  627   -234  -27.2%
  security enforcement 2     1781-> 1341   -440  -24.7%
  load balancing 1           1441-> 1093   -348  -24.1%
  load balancing 2           1441-> 1093   -348  -24.1%
  security monitoring 31      137->  105    -32  -23.4%
  network telemetry 3         488->  377   -111  -22.8%
  TCP congestion control 1   2760-> 2136   -624  -22.6%
  network telemetry 2         493->  382   -111  -22.5%
  security monitoring 21      360->  288    -72  -20.0%
  security monitoring 23      360->  288    -72  -20.0%
  network telemetry 6        1130->  907   -223  -19.7%
  network telemetry 1         407->  328    -79  -19.4%
  security enforcement 52     201->  162    -39  -19.4%
  security enforcement 54     206->  167    -39  -18.9%
  security enforcement 55     206->  167    -39  -18.9%
  security monitoring 80     1095->  896   -199  -18.2%
  TCP congestion control 6   4153-> 3529   -624  -15.0%
  security monitoring 19      272->  233    -39  -14.3%
  storage telemetry 7        1433-> 1230   -203  -14.2%
  storage telemetry 1        1436-> 1233   -203  -14.1%
  storage telemetry 5        1436-> 1233   -203  -14.1%
  storage telemetry 9        1436-> 1233   -203  -14.1%
  storage telemetry 11       1436-> 1233   -203  -14.1%
  storage telemetry 17       1436-> 1233   -203  -14.1%
  storage telemetry 2        1444-> 1241   -203  -14.1%
  storage telemetry 3        1444-> 1241   -203  -14.1%
  storage telemetry 10       1444-> 1241   -203  -14.1%
  storage telemetry 12       1444-> 1241   -203  -14.1%
  storage telemetry 13       1444-> 1241   -203  -14.1%
  storage telemetry 14       1444-> 1241   -203  -14.1%
  storage telemetry 15       1444-> 1241   -203  -14.1%
  storage telemetry 16       1444-> 1241   -203  -14.1%
  security enforcement 22     680->  586    -94  -13.8%
  security monitoring 81     1420-> 1224   -196  -13.8%
  security monitoring 67      896->  786   -110  -12.3%
  host monitoring 4           966->  849   -117  -12.1%
  host monitoring 5           288->  256    -32  -11.1%
  host monitoring 1          1785-> 1588   -197  -11.0%
  network tuning 4           2929-> 2613   -316  -10.8%

  Largest absolute savings (JITed bytes before->after, saved, percent):
  security enforcement 40   12217->11588   -629   -5.2%
  TCP congestion control 1   2760-> 2136   -624  -22.6%
  TCP congestion control 3   8456-> 7832   -624   -7.4%
  TCP congestion control 6   4153-> 3529   -624  -15.0%
  security enforcement 46   14109->13486   -623   -4.4%
  security enforcement 48   13941->13319   -622   -4.5%
  security enforcement 42   13783->13193   -590   -4.3%
  TCP congestion control 5    841->  267   -574  -68.2%
  security enforcement 26   11883->11332   -551   -4.6%
  security enforcement 44    9468-> 8917   -551   -5.8%
  security monitoring 68    10867->10324   -543   -5.0%
  security enforcement 8    34695->34181   -514   -1.5%
  security enforcement 9    25952->25440   -512   -2.0%
  security enforcement 12   25993->25481   -512   -2.0%
  security enforcement 13   25993->25481   -512   -2.0%
  security enforcement 14   25993->25481   -512   -2.0%
  security enforcement 10   26290->25779   -511   -1.9%
  container runtime 1         789->  315   -474  -60.1%
  security enforcement 3    25670->25202   -468   -1.8%
  security enforcement 2     1781-> 1341   -440  -24.7%
  security enforcement 58    9190-> 8831   -359   -3.9%
  security enforcement 11   29176->28821   -355   -1.2%
  security enforcement 57   23466->23111   -355   -1.5%
  host monitoring 3          1170->  816   -354  -30.3%
  load balancing 1           1441-> 1093   -348  -24.1%
  load balancing 2           1441-> 1093   -348  -24.1%
  security enforcement 16   25842->25526   -316   -1.2%
  network tuning 3           2957-> 2641   -316  -10.7%
  network tuning 4           2929-> 2613   -316  -10.8%
  security enforcement 7      452->  139   -313  -69.2%
  security monitoring 3     20315->20004   -311   -1.5%
  security enforcement 45   13025->12715   -310   -2.4%
  security enforcement 47   12857->12548   -309   -2.4%
  security enforcement 60    7751-> 7445   -306   -4.0%
  profiling 3                 604->  319   -285  -47.2%
  sched_ext scheduler 9      3389-> 3105   -284   -8.4%
  security enforcement 1    10406->10129   -277   -2.7%
  security enforcement 4     9267-> 8990   -277   -3.0%
  security enforcement 5    20127->19850   -277   -1.4%
  security enforcement 17   31234->30957   -277   -0.9%
  security enforcement 18    8439-> 8162   -277   -3.3%
  security enforcement 19    8439-> 8162   -277   -3.3%
  security enforcement 20    8439-> 8162   -277   -3.3%
  security enforcement 27   10654->10377   -277   -2.6%
  security enforcement 28   11089->10812   -277   -2.5%
  security enforcement 29    8482-> 8205   -277   -3.3%
  security enforcement 30    8479-> 8202   -277   -3.3%
  security enforcement 31    8492-> 8215   -277   -3.3%
  security enforcement 32   16529->16252   -277   -1.7%
  security enforcement 33   16387->16110   -277   -1.7%

The branchless mask Peter suggested for v3 (lea, mov, sar, inc, and, dec in
front of the load, no x86/mm change), measured the same way, gives less
than half of that: 16.2% mean and 35.7% max for the selftests programs,
3.3% mean and 31.9% max for the production programs, with five dependent
ALU instructions left in front of every load. It also keeps the vsyscall
problem: an address with the top bit set passes the mask unchanged, so a
PROBE_MEM load from the vsyscall page still oopses under SMAP, and the
selftest from patch 3 crashes a kernel with that variant at exactly that
address.

Changelog:
----------
v3 -> v4
v3: https://lore.kernel.org/bpf/20241103193512.4076710-1-memxor@gmail.com

 * Rebase on bpf-next.
 * Drop "zero overhead" from the title and state the fault handler cost up
   front. (Dave)
 * Spell out what the series does and does not change for confidential
   guests. (Dave)
 * Measure code size over the selftests and Meta production objects, for
   this series and for the branchless mask alternative. (Peter)
 * Add a selftest for PROBE_MEM loads from user, vsyscall and non-canonical
   addresses.
 * Call fixup_exception() directly in the SMAP branch and fall through to
   the existing page_fault_oops() instead of kernelmode_fixup_or_oops().
 * Restructure the JIT change around probe_mem and bounds_check flags
   instead of a goto; drop Puranjay's Ack on it since the code changed.

v2 -> v3
v2: https://lore.kernel.org/bpf/20240619092216.1780946-1-memxor@gmail.com

 * Rebase on bpf-next
 * Add Puranjay's Acks

v1 -> v2
v1: https://lore.kernel.org/bpf/20240515233932.3733815-1-memxor@gmail.com

 * Rebase on bpf-next

Kumar Kartikeya Dwivedi (3):
  x86/mm: Resolve BPF exception fixups for user address faults under
    SMAP
  bpf, x86: Skip the PROBE_MEM address range check under SMAP
  selftests/bpf: Test PROBE_MEM loads from invalid addresses

 arch/x86/mm/fault.c                           | 11 +++
 arch/x86/net/bpf_jit_comp.c                   | 21 ++++--
 .../bpf/prog_tests/probe_mem_fault.c          | 69 +++++++++++++++++++
 .../selftests/bpf/progs/probe_mem_fault.c     | 41 +++++++++++
 4 files changed, 137 insertions(+), 5 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c
 create mode 100644 tools/testing/selftests/bpf/progs/probe_mem_fault.c


base-commit: e1d84a37cba984388988d2f1ddc84561413f0db2
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP
  2026-10-09  2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi
@ 2026-10-09  2:49 ` Kumar Kartikeya Dwivedi
  2026-10-09  4:13   ` Borislav Petkov
  2026-10-09  2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
  2026-10-09  2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi
  2 siblings, 1 reply; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09  2:49 UTC (permalink / raw)
  To: bpf
  Cc: Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko,
	Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis,
	Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra,
	Thomas Gleixner, Ingo Molnar, Borislav Petkov, kkd, kernel-team,
	x86, linux-kernel

With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a
user address with EFLAGS.AC clear as a kernel bug: it does not consult the
exception table and oopses right away. That is correct for ordinary kernel
code, where get_kernel_nofault() and the other nofault accessors never let
a user address reach a faulting instruction.

JITed BPF programs are different. A privileged program may dereference a
pointer the verifier cannot prove valid, and the verifier marks such loads
PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM
load, so that a fault on an unmapped kernel address zeroes the destination
register and the program continues. Since a user address would oops
instead, the x86 JIT also emits an address range check in front of every
PROBE_MEM load, which keeps user addresses, the guard page above
TASK_SIZE_MAX and the vsyscall page away from the load. The check
duplicates the fault handler's knowledge of the address space layout, got
the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM
runtime load check"), and costs nine instructions and 32 to 39 bytes of
code per load.

When the faulting instruction belongs to a BPF program, resolve its
exception table entry instead of oopsing, exactly as is done for faults on
kernel addresses. The is_bpf_text_address() lookup sits inside the
unlikely() SMAP branch that currently ends in page_fault_oops(), so no path
that does not oops today executes any additional code, and the oops itself
is unchanged when no entry matches. Non-BPF code keeps the existing
behaviour: a kernel-mode user access without STAC still oopses, extable
entry or not.

The next patch uses this to drop the range check from the JIT when SMAP is
enabled. Nothing changes about which addresses a BPF program may read: a
user address never becomes readable, since SMAP forbids the access, and a
PROBE_MEM load of a kernel address is handled as before. Without SMAP the
JIT keeps its range check and this path is never reached.

Acked-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 arch/x86/mm/fault.c | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
index aa88370ce739..2060e5f35d77 100644
--- a/arch/x86/mm/fault.c
+++ b/arch/x86/mm/fault.c
@@ -20,6 +20,7 @@
 #include <linux/mm_types.h>
 #include <linux/mm.h>			/* find_and_lock_vma() */
 #include <linux/vmalloc.h>
+#include <linux/filter.h>		/* is_bpf_text_address()	*/
 
 #include <asm/cpufeature.h>		/* boot_cpu_has, ...		*/
 #include <asm/traps.h>			/* dotraplinkage, ...		*/
@@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs,
 	if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) &&
 		     !(error_code & X86_PF_USER) &&
 		     !(regs->flags & X86_EFLAGS_AC))) {
+		/*
+		 * JITed BPF programs dereference untrusted pointers with loads
+		 * that carry an exception table entry (PROBE_MEM). SMAP makes
+		 * sure such a load cannot read user memory, so resolve the
+		 * fault through the exception table, as for an unmapped kernel
+		 * address, instead of oopsing.
+		 */
+		if (is_bpf_text_address(regs->ip) &&
+		    fixup_exception(regs, X86_TRAP_PF, error_code, address))
+			return;
 		/*
 		 * No extable entry here.  This was a kernel access to an
 		 * invalid pointer.  get_kernel_nofault() will not get here.
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check under SMAP
  2026-10-09  2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi
  2026-10-09  2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
@ 2026-10-09  2:49 ` Kumar Kartikeya Dwivedi
  2026-10-09  3:42   ` bot+bpf-ci
  2026-10-09  2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi
  2 siblings, 1 reply; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09  2:49 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen,
	Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86,
	linux-kernel

The x86 JIT guards every PROBE_MEM load with a range check that keeps user
addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away
from the load, because a kernel-mode fault on those addresses would oops
under SMAP instead of reaching the load's exception table entry. The check
is nine instructions and 39 bytes (32 when the offset is zero) in front of
a load of a few bytes. For "r7 = *(u64 *)(r0 + 2200)" on a 5-level paging
kernel, with VSYSCALL_ADDR and TASK_SIZE_MAX + PAGE_SIZE - VSYSCALL_ADDR as
the two constants:

  movq    $-10485760, %r10
  movq    %rax, %r11
  addq    $2200, %r11
  subq    %r10, %r11
  movabsq $72057594048413696, %r10
  cmpq    %r10, %r11
  ja      load
  xorl    %edi, %edi
  jmp     done
load:
  movq    2200(%rax), %rdi
done:

The previous patch made do_user_addr_fault() resolve the exception table
entries of BPF programs for faults on user addresses when SMAP is enabled.
On such kernels, emit the bare load with its exception table entry, as the
arm64, riscv, s390 and loongarch JITs already do. All BPF programs run with
SMAP active once the CPU feature is enabled, and the feature cannot change
after boot, so checking it at JIT time is sufficient. Kernels without SMAP,
including those booted with nosmap, keep the range check.

Measured with veristat over every object of the BPF selftests and over 466
production objects from Meta's fleet, on an x86-64 guest with SMAP, with
and without this series on top of bpf-next:

                        programs  with PROBE_MEM  reduction per program
                                                  mean  median    max
  BPF selftests             3324              91  35.0%  37.5%  74.4%
  Meta production programs  1710             228   7.0%   1.8%  69.2%

No program grows. The socket and task iterators of the selftests lose about
half of their code, dump_tcp6 goes from 4386 to 2124 bytes, and the
smallest production programs lose 60% to 69%. Loads through trusted
pointers and the probe_read helpers do not use PROBE_MEM, which is why most
programs are unaffected.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 arch/x86/net/bpf_jit_comp.c | 21 ++++++++++++++++-----
 1 file changed, 16 insertions(+), 5 deletions(-)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 083fcd6cf15b..793e7cd5a5c4 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -2133,6 +2133,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 		const s32 imm32 = insn->imm;
 		u32 dst_reg = insn->dst_reg;
 		u32 src_reg = insn->src_reg;
+		bool probe_mem, bounds_check;
 		bool accesses_stack_only;
 		u8 b2 = 0, b3 = 0;
 		u8 *start_of_ldx;
@@ -2709,6 +2710,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 		case BPF_LDX | BPF_PROBE_MEMSX | BPF_B:
 		case BPF_LDX | BPF_PROBE_MEMSX | BPF_H:
 		case BPF_LDX | BPF_PROBE_MEMSX | BPF_W:
+			probe_mem = BPF_MODE(insn->code) == BPF_PROBE_MEM ||
+				    BPF_MODE(insn->code) == BPF_PROBE_MEMSX;
+			/*
+			 * With SMAP enabled, a load from a user address faults and
+			 * do_user_addr_fault() resolves the exception table entry of the
+			 * program, as for an unmapped kernel address, so the address range
+			 * check is only needed without SMAP.
+			 */
+			bounds_check = probe_mem && !cpu_feature_enabled(X86_FEATURE_SMAP);
 			insn_off = insn->off;
 			if (src_reg == BPF_REG_PARAMS) {
 				if (insn_off == 8) {
@@ -2724,8 +2734,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 				 */
 			}
 
-			if (BPF_MODE(insn->code) == BPF_PROBE_MEM ||
-			    BPF_MODE(insn->code) == BPF_PROBE_MEMSX) {
+			if (bounds_check) {
 				/* Conservatively check that src_reg + insn->off is a kernel address:
 				 *   src_reg + insn->off > TASK_SIZE_MAX + PAGE_SIZE
 				 *   and
@@ -2772,6 +2781,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 				/* populate jmp_offset for JAE above to jump to start_of_ldx */
 				start_of_ldx = prog;
 				end_of_jmp[-1] = start_of_ldx - end_of_jmp;
+			} else if (probe_mem) {
+				start_of_ldx = prog;
 			} else if (!accesses_stack_only) {
 				err = emit_kasan_check(env, &prog, src_reg,
 						       insn_off,
@@ -2785,14 +2796,14 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 				emit_ldsx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off);
 			else
 				emit_ldx(&prog, BPF_SIZE(insn->code), dst_reg, src_reg, insn_off);
-			if (BPF_MODE(insn->code) == BPF_PROBE_MEM ||
-			    BPF_MODE(insn->code) == BPF_PROBE_MEMSX) {
+			if (probe_mem) {
 				struct exception_table_entry *ex;
 				u8 *_insn = image + proglen + (start_of_ldx - temp);
 				s64 delta;
 
 				/* populate jmp_offset for JMP above */
-				start_of_ldx[-1] = prog - start_of_ldx;
+				if (bounds_check)
+					start_of_ldx[-1] = prog - start_of_ldx;
 
 				if (!bpf_prog->aux->extable)
 					break;
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses
  2026-10-09  2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi
  2026-10-09  2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
  2026-10-09  2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
@ 2026-10-09  2:49 ` Kumar Kartikeya Dwivedi
  2 siblings, 0 replies; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09  2:49 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, Ihor Solodrai, Dave Hansen,
	Andy Lutomirski, Peter Zijlstra, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Puranjay Mohan, kkd, kernel-team, x86,
	linux-kernel

Add a test that performs PROBE_MEM loads of three sizes through a
bpf_core_cast() pointer whose value is chosen by userspace: NULL, a low
user address, the last user page, a non-canonical address and, on x86-64,
the vsyscall page and an offset into it. Each load must read zero and the
kernel must survive. A load through the current task pointer checks that
the same loads read real values when the address is valid.

The test passes on kernels that reject the addresses with the JIT's range
check and on kernels that rely on the fault handler; with only the JIT
change of the previous patch applied, the first subtest oopses, which is
what the fault handler change prevents.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../bpf/prog_tests/probe_mem_fault.c          | 69 +++++++++++++++++++
 .../selftests/bpf/progs/probe_mem_fault.c     | 41 +++++++++++
 2 files changed, 110 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c
 create mode 100644 tools/testing/selftests/bpf/progs/probe_mem_fault.c

diff --git a/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c
new file mode 100644
index 000000000000..57a313e35eb6
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/probe_mem_fault.c
@@ -0,0 +1,69 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "probe_mem_fault.skel.h"
+
+#if defined(__x86_64__)
+#include <asm/vsyscall.h>
+#endif
+
+/*
+ * Addresses a PROBE_MEM load has to survive. Either the JIT's address check
+ * or the fault handler must turn each load into a zero result.
+ */
+static const struct {
+	const char *name;
+	unsigned long addr;
+} bad_addrs[] = {
+	{ "null", 0 },
+	{ "low_user", 4096 },
+	{ "last_user_page", (1UL << 47) - 4096 },
+	{ "non_canonical", 1UL << 63 },
+#if defined(__x86_64__)
+	{ "vsyscall", VSYSCALL_ADDR },
+	{ "vsyscall_tail", VSYSCALL_ADDR + 0x800 },
+#endif
+};
+
+static void trigger(struct probe_mem_fault *skel, int *runs)
+{
+	skel->bss->val_dw = ~0ULL;
+	skel->bss->val_w = ~0U;
+	skel->bss->val_b = ~0;
+	usleep(1);
+	ASSERT_EQ(skel->bss->runs, ++*runs, "runs");
+}
+
+void test_probe_mem_fault(void)
+{
+	struct probe_mem_fault *skel;
+	int runs = 0, i;
+
+	skel = probe_mem_fault__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_and_load"))
+		return;
+
+	skel->bss->target_pid = getpid();
+	if (!ASSERT_OK(probe_mem_fault__attach(skel), "attach"))
+		goto out;
+
+	/* A valid kernel address is read for real. */
+	skel->bss->use_current_task = true;
+	trigger(skel, &runs);
+	ASSERT_EQ(skel->bss->val_w, getpid(), "pid");
+	ASSERT_NEQ(skel->bss->val_dw, 0, "start_time");
+	ASSERT_NEQ(skel->bss->val_b, 0, "comm");
+
+	skel->bss->use_current_task = false;
+	for (i = 0; i < ARRAY_SIZE(bad_addrs); i++) {
+		if (!test__start_subtest(bad_addrs[i].name))
+			continue;
+		skel->bss->addr = bad_addrs[i].addr;
+		trigger(skel, &runs);
+		ASSERT_EQ(skel->bss->val_dw, 0, "start_time");
+		ASSERT_EQ(skel->bss->val_w, 0, "pid");
+		ASSERT_EQ(skel->bss->val_b, 0, "comm");
+	}
+out:
+	probe_mem_fault__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/probe_mem_fault.c b/tools/testing/selftests/bpf/progs/probe_mem_fault.c
new file mode 100644
index 000000000000..748be45418f5
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/probe_mem_fault.c
@@ -0,0 +1,41 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+#include "bpf_misc.h"
+
+char _license[] SEC("license") = "GPL";
+
+int target_pid;
+bool use_current_task;
+unsigned long addr;
+int runs;
+__u64 val_dw;
+__u32 val_w;
+__u8 val_b;
+
+SEC("fentry/" SYS_PREFIX "sys_nanosleep")
+int probe_mem_fault(void *ctx)
+{
+	struct task_struct *task;
+	unsigned long p = addr;
+
+	if ((bpf_get_current_pid_tgid() >> 32) != target_pid)
+		return 0;
+
+	if (use_current_task)
+		p = (unsigned long)bpf_get_current_task_btf();
+	/*
+	 * bpf_core_cast() yields an untrusted pointer, so every load through it
+	 * is a PROBE_MEM load. Whatever the address is, the load must either
+	 * read the field or produce zero; the kernel must not oops.
+	 */
+	task = bpf_core_cast((void *)p, struct task_struct);
+	val_dw = task->start_time;
+	val_w = task->pid;
+	val_b = task->comm[0];
+	runs++;
+	return 0;
+}
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check under SMAP
  2026-10-09  2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
@ 2026-10-09  3:42   ` bot+bpf-ci
  0 siblings, 0 replies; 7+ messages in thread
From: bot+bpf-ci @ 2026-10-09  3:42 UTC (permalink / raw)
  To: memxor, bpf
  Cc: ast, andrii, daniel, eddyz87, emil, ihor.solodrai, dave.hansen,
	luto, peterz, tglx, mingo, bp, puranjay, kkd, kernel-team, x86,
	linux-kernel, ast, andrii, daniel, martin.lau, eddyz87,
	yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 8727 bytes --]

> bpf, x86: Skip the PROBE_MEM address range check under SMAP

TL;DR:
- Possible kernel oops: unprivileged (CAP_PERFMON, no CAP_BPF) programs aren't
  in kallsyms, so a NULL PROBE_MEM load under SMAP may not be fixed up
- Commit message wrongly credits arm64/riscv/loongarch; runtime cost
  unmentioned
- Stale comments (LASS, uaddress_limit)

>     bpf, x86: Skip the PROBE_MEM address range check under SMAP
>
>     The x86 JIT guards every PROBE_MEM load with a range check that keeps user
>     addresses, the guard page above TASK_SIZE_MAX and the vsyscall page away
>     from the load, because a kernel-mode fault on those addresses would oops
>     under SMAP instead of reaching the load's exception table entry. The check
>     is nine instructions and 39 bytes (32 when the offset is zero) in front of
>     a load of a few bytes. For "r7 = *(u64 *)(r0 + 2200)" on a 5-level paging
>     kernel, with VSYSCALL_ADDR and TASK_SIZE_MAX + PAGE_SIZE - VSYSCALL_ADDR as
>     the two constants:
>
>       movq    $-10485760, %r10
>       movq    %rax, %r11
>       addq    $2200, %r11
>       subq    %r10, %r11
>       movabsq $72057594048413696, %r10
>       cmpq    %r10, %r11
>       ja      load
>       xorl    %edi, %edi
>       jmp     done
>     load:
>       movq    2200(%rax), %rdi
>     done:
>
>     The previous patch made do_user_addr_fault() resolve the exception table
>     entries of BPF programs for faults on user addresses when SMAP is enabled.
>     On such kernels, emit the bare load with its exception table entry, as the
>     arm64, riscv, s390 and loongarch JITs already do. All BPF programs run with
>     SMAP active once the CPU feature is enabled, and the feature cannot change
>     after boot, so checking it at JIT time is sufficient. Kernels without SMAP,
>     including those booted with nosmap, keep the range check.
>
>     Measured with veristat over every object of the BPF selftests and over 466
>     production objects from Meta's fleet, on an x86-64 guest with SMAP, with
>     and without this series on top of bpf-next:
>
>                             programs  with PROBE_MEM  reduction per program
>                                                       mean  median    max
>       BPF selftests             3324              91  35.0%  37.5%  74.4%
>       Meta production programs  1710             228   7.0%   1.8%  69.2%
>
>     No program grows. The socket and task iterators of the selftests lose about
>     half of their code, dump_tcp6 goes from 4386 to 2124 bytes, and the
>     smallest production programs lose 60% to 69%. Loads through trusted
>     pointers and the probe_read helpers do not use PROBE_MEM, which is why most
>     programs are unaffected.
>
>     Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>

The commit message says the JIT emits the bare load "as the arm64, riscv,
s390 and loongarch JITs already do".  Is that accurate for arm64, riscv and
loongarch?

Those three select ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE and use the weak
bpf_arch_uaddress_limit(), which returns TASK_SIZE.  For them
bpf_do_misc_fixups() rewrites every PROBE_MEM and PROBE_MEMSX load into a
guarded sequence before the JIT sees it:

    *patch++ = BPF_MOV64_REG(BPF_REG_AX, insn->src_reg);
    if (insn->off)
            *patch++ = BPF_ALU64_IMM(BPF_ADD, BPF_REG_AX, insn->off);
    *patch++ = BPF_ALU64_IMM(BPF_RSH, BPF_REG_AX, 32);
    *patch++ = BPF_JMP_IMM(BPF_JLE, BPF_REG_AX, uaddress_limit >> 32, 2);
    *patch++ = *insn;
    *patch++ = BPF_JMP_IMM(BPF_JA, 0, 0, 1);
    *patch++ = BPF_MOV64_IMM(insn->dst_reg, 0);

So those JITs never emit a bare PROBE_MEM load for a user address.  Only
s390, which has separate address spaces, relies on the fault path in the way
the changelog describes.

With this patch, x86 with X86_FEATURE_SMAP set takes a full page fault,
is_bpf_text_address() and the exception table search for a load of NULL plus
an offset, where it used to execute a compare and a branch.  The veristat
numbers only measure code size.

Could the changelog correct the arm64/riscv/loongarch statement and say
something about the runtime cost of loads that now fault where they used to
branch?

> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index 083fcd6cf15b..793e7cd5a5c4 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c

[ ... ]

> @@ -2709,6 +2710,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>  		case BPF_LDX | BPF_PROBE_MEMSX | BPF_B:
>  		case BPF_LDX | BPF_PROBE_MEMSX | BPF_H:
>  		case BPF_LDX | BPF_PROBE_MEMSX | BPF_W:
> +			probe_mem = BPF_MODE(insn->code) == BPF_PROBE_MEM ||
> +				    BPF_MODE(insn->code) == BPF_PROBE_MEMSX;
> +			/*
> +			 * With SMAP enabled, a load from a user address faults and
> +			 * do_user_addr_fault() resolves the exception table entry of the
> +			 * program, as for an unmapped kernel address, so the address range
> +			 * check is only needed without SMAP.
> +			 */

this isn't a bug, but is this comment accurate when X86_FEATURE_LASS is
enabled?  With LASS, a supervisor load from the user half of the address
space raises a general protection fault rather than a page fault, and that is
handled on a different path:

    exc_general_protection() -> gp_try_fixup_and_notify() -> fixup_exception()

The outcome is the same, but the comment names do_user_addr_fault() as the
place where the entry is resolved.

> +			bounds_check = probe_mem && !cpu_feature_enabled(X86_FEATURE_SMAP);
>  			insn_off = insn->off;

Can this oops for a program that is not in kallsyms?  With bounds_check
false, a PROBE_MEM load of a user address such as NULL plus a field offset
depends only on the new branch in do_user_addr_fault():

    if (is_bpf_text_address(regs->ip) &&
        fixup_exception(regs, X86_TRAP_PF, error_code, address))
            return;
    ...
    page_fault_oops(regs, error_code, address);

Both is_bpf_text_address() and search_bpf_extables() only find programs that
are registered in bpf_tree, and bpf_prog_kallsyms_add() skips programs whose
loader lacks CAP_BPF:

    void bpf_prog_kallsyms_add(struct bpf_prog *fp)
    {
            if (!bpf_prog_kallsyms_candidate(fp) ||
                !bpf_token_capable(fp->aux->token, CAP_BPF))
                    return;

Such programs can still contain PROBE_MEM loads.  With
kernel.unprivileged_bpf_disabled=0, bpf_prog_load() accepts
BPF_PROG_TYPE_SOCKET_FILTER and BPF_PROG_TYPE_CGROUP_SKB without CAP_BPF.
With CAP_PERFMON, env->allow_ptr_leaks is true and bpf_sk_base_func_proto()
exposes bpf_skc_to_tcp_sock() and the other skc_to_* helpers.

skb->sk is readable as PTR_TO_SOCK_COMMON_OR_NULL, and sk_filter_trim_cap()
sets skb->sk while the filter runs.  The helper returns a plain
PTR_TO_BTF_ID, and walking a pointer field from it, for example
tp->inet_conn.icsk_ulp_ops->..., or sk_socket on an orphaned socket, gets
rewritten to BPF_PROBE_MEM by bpf_convert_ctx_accesses().  The x86 JIT then
emits an extable entry for it.

Before this patch, a NULL plus offset load in such a program was caught by
the JIT range check and dst was set to 0.  After it, on a CPU with
X86_FEATURE_SMAP the sequence would be:

    handle_page_fault() -> do_user_addr_fault() -> is_bpf_text_address()
    returns false -> page_fault_oops()

That would be an oops in softirq context (sk_filter on the receive path)
triggered by a user with CAP_PERFMON but without CAP_BPF.

arm64, riscv and loongarch are not exposed to this, because
bpf_do_misc_fixups() inserts the generic user address guard for them.

Should the range check be kept when the program will not be in kallsyms
(!bpf_token_capable(prog->aux->token, CAP_BPF)), or should every JITed
program with exception entries be registered in bpf_tree?

[ ... ]

Separately, the comment above bpf_arch_uaddress_limit() in this file:

arch/x86/net/bpf_jit_comp.c:
    /* x86-64 JIT emits its own code to filter user addresses so return 0 here */
    u64 bpf_arch_uaddress_limit(void)
    {
            return 0;
    }

is the stated reason x86 opts out of the generic verifier user address guard
in bpf_do_misc_fixups().  Is it still correct?  After this patch do_jit()
emits no filter at all when cpu_feature_enabled(X86_FEATURE_SMAP), and user
addresses are rejected only by the hardware check plus the fault handler
fixup.  Should the comment be updated, or should bpf_arch_uaddress_limit() be
made consistent with the new JIT behavior?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37877475188

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP
  2026-10-09  2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
@ 2026-10-09  4:13   ` Borislav Petkov
  2026-10-09 15:23     ` Kumar Kartikeya Dwivedi
  0 siblings, 1 reply; 7+ messages in thread
From: Borislav Petkov @ 2026-10-09  4:13 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi
  Cc: bpf, Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko,
	Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis,
	Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra,
	Thomas Gleixner, Ingo Molnar, kkd, kernel-team, x86,
	linux-kernel

On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote:
> With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a
> user address with EFLAGS.AC clear as a kernel bug: it does not consult the
> exception table and oopses right away. That is correct for ordinary kernel
> code, where get_kernel_nofault() and the other nofault accessors never let
> a user address reach a faulting instruction.
> 
> JITed BPF programs are different. A privileged program may dereference a
> pointer the verifier cannot prove valid, and the verifier marks such loads
> PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM
> load, so that a fault on an unmapped kernel address zeroes the destination
> register and the program continues. Since a user address would oops
> instead, the x86 JIT also emits an address range check in front of every
> PROBE_MEM load, which keeps user addresses, the guard page above
> TASK_SIZE_MAX and the vsyscall page away from the load. The check
> duplicates the fault handler's knowledge of the address space layout, got
> the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM
> runtime load check"), and costs nine instructions and 32 to 39 bytes of
> code per load.

I can't parse that. Why do bpf memory accesses need to be handled differently
than any other memory accesses when SMAP is enabled?

Perhaps you should give a concrete example.

And why can't all that gunk be resolved at program load instead of going all
the way in the #PF handler?

> When the faulting instruction belongs to a BPF program, resolve its
> exception table entry instead of oopsing, exactly as is done for faults on
> kernel addresses. The is_bpf_text_address() lookup sits inside the
> unlikely() SMAP branch that currently ends in page_fault_oops(), so no path
> that does not oops today executes any additional code, and the oops itself
> is unchanged when no entry matches. Non-BPF code keeps the existing
> behaviour: a kernel-mode user access without STAC still oopses, extable
> entry or not.
> 
> The next patch uses this to drop the range check from the JIT when SMAP is

There's no next patch and previous patch in git history.

> enabled. Nothing changes about which addresses a BPF program may read: a
> user address never becomes readable, since SMAP forbids the access, and a
> PROBE_MEM load of a kernel address is handled as before. Without SMAP the
> JIT keeps its range check and this path is never reached.
> 
> Acked-by: Puranjay Mohan <puranjay@kernel.org>
> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
> ---
>  arch/x86/mm/fault.c | 11 +++++++++++
>  1 file changed, 11 insertions(+)
> 
> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
> index aa88370ce739..2060e5f35d77 100644
> --- a/arch/x86/mm/fault.c
> +++ b/arch/x86/mm/fault.c
> @@ -20,6 +20,7 @@
>  #include <linux/mm_types.h>
>  #include <linux/mm.h>			/* find_and_lock_vma() */
>  #include <linux/vmalloc.h>
> +#include <linux/filter.h>		/* is_bpf_text_address()	*/
>  
>  #include <asm/cpufeature.h>		/* boot_cpu_has, ...		*/
>  #include <asm/traps.h>			/* dotraplinkage, ...		*/
> @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs,
>  	if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) &&
>  		     !(error_code & X86_PF_USER) &&
>  		     !(regs->flags & X86_EFLAGS_AC))) {
> +		/*
> +		 * JITed BPF programs dereference untrusted pointers with loads
> +		 * that carry an exception table entry (PROBE_MEM). SMAP makes
> +		 * sure such a load cannot read user memory, so resolve the
> +		 * fault through the exception table, as for an unmapped kernel
> +		 * address, instead of oopsing.
> +		 */
> +		if (is_bpf_text_address(regs->ip) &&
> +		    fixup_exception(regs, X86_TRAP_PF, error_code, address))
> +			return;

I'm not at all amused from this adding bpf-specific handling to the #PF
handler, TBH...

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults under SMAP
  2026-10-09  4:13   ` Borislav Petkov
@ 2026-10-09 15:23     ` Kumar Kartikeya Dwivedi
  0 siblings, 0 replies; 7+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-10-09 15:23 UTC (permalink / raw)
  To: Borislav Petkov
  Cc: bpf, Puranjay Mohan, Alexei Starovoitov, Andrii Nakryiko,
	Daniel Borkmann, Eduard Zingerman, Emil Tsalapatis,
	Ihor Solodrai, Dave Hansen, Andy Lutomirski, Peter Zijlstra,
	Thomas Gleixner, Ingo Molnar, kkd, kernel-team, x86,
	linux-kernel

On Fri Oct 9, 2026 at 6:13 AM CEST, Borislav Petkov wrote:
> On Fri, Oct 09, 2026 at 04:49:21AM +0200, Kumar Kartikeya Dwivedi wrote:
>> With SMAP enabled, do_user_addr_fault() treats a kernel-mode fault on a
>> user address with EFLAGS.AC clear as a kernel bug: it does not consult the
>> exception table and oopses right away. That is correct for ordinary kernel
>> code, where get_kernel_nofault() and the other nofault accessors never let
>> a user address reach a faulting instruction.
>>
>> JITed BPF programs are different. A privileged program may dereference a
>> pointer the verifier cannot prove valid, and the verifier marks such loads
>> PROBE_MEM. The JIT attaches an exception table entry to each PROBE_MEM
>> load, so that a fault on an unmapped kernel address zeroes the destination
>> register and the program continues. Since a user address would oops
>> instead, the x86 JIT also emits an address range check in front of every
>> PROBE_MEM load, which keeps user addresses, the guard page above
>> TASK_SIZE_MAX and the vsyscall page away from the load. The check
>> duplicates the fault handler's knowledge of the address space layout, got
>> the vsyscall page wrong until commit b599d7d26d6a ("bpf, x86: Fix PROBE_MEM
>> runtime load check"), and costs nine instructions and 32 to 39 bytes of
>> code per load.
>
> I can't parse that. Why do bpf memory accesses need to be handled differently
> than any other memory accesses when SMAP is enabled?
>

It requires the same handling as get_kernel_nofault() as if that were inlined.
The point of the patch is to be able to avoid bounds checking in the JITed code.
We can probably extend this to other nofault accessors as well, but I wanted to
keep the scope limited to BPF for now, because the cost of the bounds checking
is significant in BPF programs.

For background: Privileged programs are allowed to trace kernel code and
dereference pointers where establishing whether the pointer is pointing to a
valid object is not possible during verification. Program authors know about
this behavior, most of the use cases for this are reading data, collecting
statistics, and so on.

It is a more convenient and efficient way to walk a chain of pointers than to
use the equivalent of copy_from_kernel_nofault() in a loop.

get_kernel_nofault() does user address checks in software, in
copy_from_kernel_nofault_allowed(), because the SMAP branch in
do_user_addr_fault() treats any kernel-mode fault on a user address as a missing
STAC and oopses without looking at the extable. The x86 JIT does the same today,
inlined, in front of every such load in a BPF program.

> Perhaps you should give a concrete example.

Consider this example:

  SEC("fentry/tcp_retransmit_skb") <- attaches to the entry of tcp_retransmit_skb()
  int BPF_PROG(retrans, struct sock *sk)
  {
          struct file *f = sk->sk_socket->file;

This program was written for doing a TCP related investigation. The first load
of sk is fine, but sk_socket can be NULL (orphaned socket) or stale, so the
second load becomes a mov with an extable entry. Below is the JITed sequence.

  movq    $-10485760, %r10
  movq    %rax, %r11
  addq    $2200, %r11
  subq    %r10, %r11
  movabsq $72057594048413696, %r10
  cmpq    %r10, %r11
  ja      load
  xorl    %edi, %edi
  jmp     done
load:
  movq    2200(%rax), %rdi
done:

Nine instructions and 32 to 39 bytes to guard a every load instruction against
an access that SMAP, when enabled, will block for us, and which we could fixup,
if we had exception handling in its fault path.

So the overall idea proposal is to eliminate this sequence and invoke the
fixup_exception() handler for the faulting instruction instead, which will zero
the destination register and continue execution, just like it does for faults on
kernel addresses already.

>
> And why can't all that gunk be resolved at program load instead of going all
> the way in the #PF handler?
>

This is what we do now, but it is a significant amount of code preceding each
load instruction. If you look at some examples in the cover letter, we can shave
the text size of programs by more than half in extreme cases.

Do note that we already do fixup_exception() for kernel address faults, so this
is mostly trying to mirror that behavior for the rest and eliminate the bounds
checking.

>> When the faulting instruction belongs to a BPF program, resolve its
>> exception table entry instead of oopsing, exactly as is done for faults on
>> kernel addresses. The is_bpf_text_address() lookup sits inside the
>> unlikely() SMAP branch that currently ends in page_fault_oops(), so no path
>> that does not oops today executes any additional code, and the oops itself
>> is unchanged when no entry matches. Non-BPF code keeps the existing
>> behaviour: a kernel-mode user access without STAC still oopses, extable
>> entry or not.
>>
>> The next patch uses this to drop the range check from the JIT when SMAP is
>
> There's no next patch and previous patch in git history.

Yeah, I will reword this bit in the commit log for v2.

>
>> enabled. Nothing changes about which addresses a BPF program may read: a
>> user address never becomes readable, since SMAP forbids the access, and a
>> PROBE_MEM load of a kernel address is handled as before. Without SMAP the
>> JIT keeps its range check and this path is never reached.
>>
>> Acked-by: Puranjay Mohan <puranjay@kernel.org>
>> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
>> ---
>>  arch/x86/mm/fault.c | 11 +++++++++++
>>  1 file changed, 11 insertions(+)
>>
>> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
>> index aa88370ce739..2060e5f35d77 100644
>> --- a/arch/x86/mm/fault.c
>> +++ b/arch/x86/mm/fault.c
>> @@ -20,6 +20,7 @@
>>  #include <linux/mm_types.h>
>>  #include <linux/mm.h>			/* find_and_lock_vma() */
>>  #include <linux/vmalloc.h>
>> +#include <linux/filter.h>		/* is_bpf_text_address()	*/
>>
>>  #include <asm/cpufeature.h>		/* boot_cpu_has, ...		*/
>>  #include <asm/traps.h>			/* dotraplinkage, ...		*/
>> @@ -1262,6 +1263,16 @@ void do_user_addr_fault(struct pt_regs *regs,
>>  	if (unlikely(cpu_feature_enabled(X86_FEATURE_SMAP) &&
>>  		     !(error_code & X86_PF_USER) &&
>>  		     !(regs->flags & X86_EFLAGS_AC))) {
>> +		/*
>> +		 * JITed BPF programs dereference untrusted pointers with loads
>> +		 * that carry an exception table entry (PROBE_MEM). SMAP makes
>> +		 * sure such a load cannot read user memory, so resolve the
>> +		 * fault through the exception table, as for an unmapped kernel
>> +		 * address, instead of oopsing.
>> +		 */
>> +		if (is_bpf_text_address(regs->ip) &&
>> +		    fixup_exception(regs, X86_TRAP_PF, error_code, address))
>> +			return;
>
> I'm not at all amused from this adding bpf-specific handling to the #PF
> handler, TBH...

I understand, and perhaps this can be adjusted to be a little different, but on
the flip side, at this point, the kernel is about to print an OOPS. I made it
BPF specific because we really don't expect to fixup exceptions for anything
else here. We already invoke fixup_exception() for kernel address faults.

The advantage is losing significant amount of code size in JITed programs and
the associated runtime overhead (you can see the numbers in the cover letter).

Do note that we can make it more generic, we can follow up and optimize
copy_from_kernel_nofault() with cpu_feature_enabled(X86_FEATURE_SMAP) as well.
But I didn't want to go there just yet. I am happy to work on a follow up if so
desired.

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-09 15:23 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09  2:49 [PATCH bpf-next v1 0/3] bpf, x86: Drop the PROBE_MEM address range check under SMAP Kumar Kartikeya Dwivedi
2026-10-09  2:49 ` [PATCH bpf-next v1 1/3] x86/mm: Resolve BPF exception fixups for user address faults " Kumar Kartikeya Dwivedi
2026-10-09  4:13   ` Borislav Petkov
2026-10-09 15:23     ` Kumar Kartikeya Dwivedi
2026-10-09  2:49 ` [PATCH bpf-next v1 2/3] bpf, x86: Skip the PROBE_MEM address range check " Kumar Kartikeya Dwivedi
2026-10-09  3:42   ` bot+bpf-ci
2026-10-09  2:49 ` [PATCH bpf-next v1 3/3] selftests/bpf: Test PROBE_MEM loads from invalid addresses Kumar Kartikeya Dwivedi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®