From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Ameer Hamza <ameer.hamza@truenas.com>
Cc: rppt@kernel.org, peterz@infradead.org, akpm@linux-foundation.org,
david@kernel.org, mingo@redhat.com, acme@kernel.org,
namhyung@kernel.org, linux-mm@kvack.org,
linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org,
regressions@lists.linux.dev, stable@vger.kernel.org,
4ncienth@gmail.com, sashal@kernel.org,
alexander.motin@truenas.com, caleb.stjohn@truenas.com
Subject: Re: [REGRESSION] secretmem: SIGBUS on memfd_secret() pages while perf record runs
Date: Sat, 3 Oct 2026 16:56:32 +0100 [thread overview]
Message-ID: <asEllR7if-J3wteg@gremlin> (raw)
In-Reply-To: <20261002175651.811343-1-ameer.hamza@truenas.com>
On Fri, Oct 02, 2026 at 10:56:51PM +0500, Ameer Hamza wrote:
> Hi,
>
> Since commit 97d34aa65c29 ("mm/secretmem: properly account locked
> pages"), in v7.3-rc2 and the 6.18.52 and 7.2.5 backports, touching a
> memfd_secret() page for the first time fails with SIGBUS while the
> same user is running perf record, and read() into such a page fails
> with EFAULT. This happens for root as well as for an unprivileged
> user running its own perf record. It reproduces on v7.3-rc4 and
> 6.18.52, and the same test passes without the recording or with the
> commit reverted.
>
> The commit charges secretmem pages to the per-user user->locked_vm
> and checks that counter against RLIMIT_MEMLOCK, with no CAP_IPC_LOCK
> exemption because the fd can be passed to other processes. perf has
> charged its ring buffers to the same counter since 2009, up to
> perf_event_mlock_kb per online CPU, and only what exceeds that
> allowance is checked against RLIMIT_MEMLOCK. sysctl/kernel.rst
> describes the allowance as not counted against the mlock limit, yet
> it sits in the counter that secretmem now enforces. perf record sizes
> its buffers to the allowance by default, 516 KiB per CPU, so on 16 or
> more CPUs one recording uses up the default 8 MiB limit on its own,
> and on fewer CPUs it leaves correspondingly less for secretmem.
>
> Reproducer, as root on a 16 vCPU x86_64 guest with the default
> ulimit -l 8192, while "perf record -o /tmp/p.data -- sleep 300" runs
> in another shell:
>
> fd = syscall(SYS_memfd_secret, 0);
> ftruncate(fd, 4096);
> p = mmap(NULL, 4096, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
> read(open("/dev/zero", O_RDONLY), p, 4096); /* EFAULT */
> p[0] = 1; /* SIGBUS */
>
> The commit closes a real hole, unbounded pinning of unevictable memory
> by unprivileged users. The question is whether the perf_event_mlock_kb
> allowance is meant to count against the RLIMIT_MEMLOCK budget that
> secretmem and the other subsystems accounting to user->locked_vm
> enforce, or whether perf should keep it in a counter of its own, which
> would leave the hole closed and secretmem usable during a recording.
>
> #regzbot introduced: 97d34aa65c29
>
Hi,
Thanks for the report, however this is not a regression in secretmem.
The field that is being used for tracking the allowance (user->locked_vm)
is used by everything-but-perf to track against RLIMIT_MEMLOCK, but perf is
using it for something else (pinned_vm).
So this is a bug in perf, which should use its own counter for this,
AFAICT.
Actually it seems perf invented the field (user->locked_vm) to calculate
memory that actually isn't mlock()'d, but is kernel-allocated and pinned so
treated 'as if' it were.
It tracks a per-process (rather than per-mm) RLIMIT_MEMLOCK allowance plus
a 'gift' of extra available pages which exceeds this.
Every other user since is checking against RLIMIT_MEMLOCK (and you'd hit
the same issues with them too) - io_uring, MSG_ZEROCOPY, AF_XDP, iommufd,
s390 KVM zCPI and now secretmem.
So this is a long-standing bug it seems.
Other things:
The SIGBUS is unfortunately necessary - since the region is populated on
fault, it is only then that it can be determined whether the limit is
violated.
Not being unlimited on CAP_IPC_LOCK is also, unfortunately, necessary to
resolve the security hole - often privileged processes hand out secretmem
to other unprivileged processes.
If those processes then didn't apply the limit check, the mitigation can't
work.
The commit message goes into detail about all this.
The way to address this right now, as I see you've applied yourselves as
far as I can see (in [0]), is to change the limits to account for this.
I will work on a patch for perf that fixes this specific issue.
Sasha - you can re-queue the fix. I will send out another fix for perf.
--
Cheers, Lorenzo
[0]:https://github.com/truenas/truenas_ros/pull/128
next prev parent reply other threads:[~2026-10-03 15:56 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 17:56 Ameer Hamza
2026-10-03 1:30 ` Sasha Levin
2026-10-03 9:18 ` Lorenzo Stoakes (ARM)
2026-10-03 14:30 ` Sasha Levin
2026-10-03 16:23 ` Lorenzo Stoakes (ARM)
2026-10-03 22:34 ` Sasha Levin
2026-10-03 14:14 ` David Hildenbrand (Arm)
2026-10-03 17:05 ` Ameer Hamza
2026-10-03 15:56 ` Lorenzo Stoakes (ARM) [this message]
2026-10-03 16:31 ` Lorenzo Stoakes (ARM)
2026-10-03 18:02 ` Ameer Hamza
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asEllR7if-J3wteg@gremlin \
--to=ljs@kernel.org \
--cc=4ncienth@gmail.com \
--cc=acme@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=alexander.motin@truenas.com \
--cc=ameer.hamza@truenas.com \
--cc=caleb.stjohn@truenas.com \
--cc=david@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
--cc=regressions@lists.linux.dev \
--cc=rppt@kernel.org \
--cc=sashal@kernel.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®