mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [REGRESSION] secretmem: SIGBUS on memfd_secret() pages while perf record runs
@ 2026-10-02 17:56 Ameer Hamza
  2026-10-03  1:30 ` Sasha Levin
                   ` (2 more replies)
  0 siblings, 3 replies; 11+ messages in thread
From: Ameer Hamza @ 2026-10-02 17:56 UTC (permalink / raw)
  To: ljs, rppt, peterz
  Cc: akpm, david, mingo, acme, namhyung, linux-mm, linux-perf-users,
	linux-kernel, regressions, stable, 4ncienth, sashal,
	alexander.motin, caleb.stjohn, ameer.hamza

Hi,

Since commit 97d34aa65c29 ("mm/secretmem: properly account locked
pages"), in v7.3-rc2 and the 6.18.52 and 7.2.5 backports, touching a
memfd_secret() page for the first time fails with SIGBUS while the
same user is running perf record, and read() into such a page fails
with EFAULT. This happens for root as well as for an unprivileged
user running its own perf record. It reproduces on v7.3-rc4 and
6.18.52, and the same test passes without the recording or with the
commit reverted.

The commit charges secretmem pages to the per-user user->locked_vm
and checks that counter against RLIMIT_MEMLOCK, with no CAP_IPC_LOCK
exemption because the fd can be passed to other processes. perf has
charged its ring buffers to the same counter since 2009, up to
perf_event_mlock_kb per online CPU, and only what exceeds that
allowance is checked against RLIMIT_MEMLOCK. sysctl/kernel.rst
describes the allowance as not counted against the mlock limit, yet
it sits in the counter that secretmem now enforces. perf record sizes
its buffers to the allowance by default, 516 KiB per CPU, so on 16 or
more CPUs one recording uses up the default 8 MiB limit on its own,
and on fewer CPUs it leaves correspondingly less for secretmem.

Reproducer, as root on a 16 vCPU x86_64 guest with the default
ulimit -l 8192, while "perf record -o /tmp/p.data -- sleep 300" runs
in another shell:

  fd = syscall(SYS_memfd_secret, 0);
  ftruncate(fd, 4096);
  p = mmap(NULL, 4096, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
  read(open("/dev/zero", O_RDONLY), p, 4096);   /* EFAULT */
  p[0] = 1;                                     /* SIGBUS */

The commit closes a real hole, unbounded pinning of unevictable memory
by unprivileged users. The question is whether the perf_event_mlock_kb
allowance is meant to count against the RLIMIT_MEMLOCK budget that
secretmem and the other subsystems accounting to user->locked_vm
enforce, or whether perf should keep it in a counter of its own, which
would leave the hole closed and secretmem usable during a recording.

#regzbot introduced: 97d34aa65c29

Thanks,
Ameer Hamza

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-10-03 22:34 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 17:56 [REGRESSION] secretmem: SIGBUS on memfd_secret() pages while perf record runs Ameer Hamza
2026-10-03  1:30 ` Sasha Levin
2026-10-03  9:18   ` Lorenzo Stoakes (ARM)
2026-10-03 14:30     ` Sasha Levin
2026-10-03 16:23       ` Lorenzo Stoakes (ARM)
2026-10-03 22:34         ` Sasha Levin
2026-10-03 14:14 ` David Hildenbrand (Arm)
2026-10-03 17:05   ` Ameer Hamza
2026-10-03 15:56 ` Lorenzo Stoakes (ARM)
2026-10-03 16:31   ` Lorenzo Stoakes (ARM)
2026-10-03 18:02   ` Ameer Hamza

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®