Re: [REGRESSION] secretmem: SIGBUS on memfd_secret() pages while perf record runs
From: Lorenzo Stoakes (ARM)
Date: Sat Oct 03 2026 - 12:02:22 EST
On Fri, Oct 02, 2026 at 10:56:51PM +0500, Ameer Hamza wrote:
> Hi,
>
> Since commit 97d34aa65c29 ("mm/secretmem: properly account locked
> pages"), in v7.3-rc2 and the 6.18.52 and 7.2.5 backports, touching a
> memfd_secret() page for the first time fails with SIGBUS while the
> same user is running perf record, and read() into such a page fails
> with EFAULT. This happens for root as well as for an unprivileged
> user running its own perf record. It reproduces on v7.3-rc4 and
> 6.18.52, and the same test passes without the recording or with the
> commit reverted.
>
> The commit charges secretmem pages to the per-user user->locked_vm
> and checks that counter against RLIMIT_MEMLOCK, with no CAP_IPC_LOCK
> exemption because the fd can be passed to other processes. perf has
> charged its ring buffers to the same counter since 2009, up to
> perf_event_mlock_kb per online CPU, and only what exceeds that
> allowance is checked against RLIMIT_MEMLOCK. sysctl/kernel.rst
> describes the allowance as not counted against the mlock limit, yet
> it sits in the counter that secretmem now enforces. perf record sizes
> its buffers to the allowance by default, 516 KiB per CPU, so on 16 or
> more CPUs one recording uses up the default 8 MiB limit on its own,
> and on fewer CPUs it leaves correspondingly less for secretmem.
>
> Reproducer, as root on a 16 vCPU x86_64 guest with the default
> ulimit -l 8192, while "perf record -o /tmp/p.data -- sleep 300" runs
> in another shell:
>
> fd = syscall(SYS_memfd_secret, 0);
> ftruncate(fd, 4096);
> p = mmap(NULL, 4096, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
> read(open("/dev/zero", O_RDONLY), p, 4096); /* EFAULT */
> p[0] = 1; /* SIGBUS */
>
> The commit closes a real hole, unbounded pinning of unevictable memory
> by unprivileged users. The question is whether the perf_event_mlock_kb
> allowance is meant to count against the RLIMIT_MEMLOCK budget that
> secretmem and the other subsystems accounting to user->locked_vm
> enforce, or whether perf should keep it in a counter of its own, which
> would leave the hole closed and secretmem usable during a recording.
>
> #regzbot introduced: 97d34aa65c29
>
Hi,
Thanks for the report, however this is not a regression in secretmem.
The field that is being used for tracking the allowance (user->locked_vm)
is used by everything-but-perf to track against RLIMIT_MEMLOCK, but perf is
using it for something else (pinned_vm).
So this is a bug in perf, which should use its own counter for this,
AFAICT.
Actually it seems perf invented the field (user->locked_vm) to calculate
memory that actually isn't mlock()'d, but is kernel-allocated and pinned so
treated 'as if' it were.
It tracks a per-process (rather than per-mm) RLIMIT_MEMLOCK allowance plus
a 'gift' of extra available pages which exceeds this.
Every other user since is checking against RLIMIT_MEMLOCK (and you'd hit
the same issues with them too) - io_uring, MSG_ZEROCOPY, AF_XDP, iommufd,
s390 KVM zCPI and now secretmem.
So this is a long-standing bug it seems.
Other things:
The SIGBUS is unfortunately necessary - since the region is populated on
fault, it is only then that it can be determined whether the limit is
violated.
Not being unlimited on CAP_IPC_LOCK is also, unfortunately, necessary to
resolve the security hole - often privileged processes hand out secretmem
to other unprivileged processes.
If those processes then didn't apply the limit check, the mitigation can't
work.
The commit message goes into detail about all this.
The way to address this right now, as I see you've applied yourselves as
far as I can see (in [0]), is to change the limits to account for this.
I will work on a patch for perf that fixes this specific issue.
Sasha - you can re-queue the fix. I will send out another fix for perf.
--
Cheers, Lorenzo
[0]:https://github.com/truenas/truenas_ros/pull/128