Re: [PATCH v18] arm64: mm: Handle Granule Protection Faults (GPFs)
From: Catalin Marinas
Date: Fri Sep 25 2026 - 13:11:30 EST
On Wed, Sep 23, 2026 at 05:04:06PM +0100, Suzuki K Poulose wrote:
> On 23/09/2026 16:45, Will Deacon wrote:
> > On Wed, Sep 23, 2026 at 12:06:15PM +0100, Catalin Marinas wrote:
> > > On Tue, Sep 22, 2026 at 06:15:34PM +0100, Will Deacon wrote:
> > > > On Sun, Sep 13, 2026 at 08:04:58AM +0100, Suzuki K Poulose wrote:
> > > > > From: Steven Price <steven.price@xxxxxxx>
> > > > >
> > > > > If the host attempts to access granules that have been delegated for use
> > > > > in a realm these accesses will be caught and will trigger a Granule
> > > > > Protection Fault (GPF).
> > > > >
> > > > > A fault during a page walk signals a bug in the kernel and is handled by
> > > > > oopsing the kernel. A non-page walk fault could be caused by user space
> > > > > having access to a page which has been delegated to the kernel and will
> > > > > trigger a SIGBUS to allow debugging why user space is trying to access a
> > > > > delegated page.
> > > > >
> > > > > There is work in progress to unmap the guest_memfd backed private pages from the
> > > > > linear map. Until we get that support, we could get spurious GPFs from within
> > > > > the kernel, e.g., load_unaligned_zeropad(). So, try to fix them up for now.
> > > > >
> > > > > Reviewed-by: Suzuki K Poulose <suzuki.poulose@xxxxxxx>
> > > > > Reviewed-by: Gavin Shan <gshan@xxxxxxxxxx>
> > > > > Reviewed-by: Catalin Marinas <catalin.marinas@xxxxxxx>
> > > > > Signed-off-by: Steven Price <steven.price@xxxxxxx>
> > > > > Signed-off-by: Suzuki K Poulose <suzuki.poulose@xxxxxxx>
> > > > > ---
> > > > > Changes since v17:
> > > > > * Pass untagged address to die_kernel_fault() - Sashiko
> > > > > * Explicitly check !user_mode() for fixups - Catalin
> > > > > * Switch to BUS_OBJERR for si_code from SI_KERNEL - Catalin
> > > > > * Clarify the commit description about the upcoming work on
> > > > > unmapping guest_memfd backed pages from linear map
> > > > > Changes since v16:
> > > > > * Update the commit description to indicate why we try to fixup GPFs
> > > > > Changes since v10:
> > > > > * Don't call arm64_notify_die() in do_gpf() but simply return 1.
> > > > > Changes since v2:
> > > > > * Include missing "Granule Protection Fault at level -1"
> > > > > ---
> > > > > arch/arm64/mm/fault.c | 30 ++++++++++++++++++++++++------
> > > > > 1 file changed, 24 insertions(+), 6 deletions(-)
> > > >
> > > > I still don't think we should do this, given that the plan is to unmap
> > > > the memory from the linear map. If this thing fires, it's a kernel bug
> > > > and it should be fatal.
> > >
> > > If the linear unmapping gets merged first, I agree, no need to handle
> > > these faults. I haven't followed that series, so no idea where it is at.
>
> There doesn't seem to be much progress on that series. Brendan
> volunteered to resurrect the series, taking over from Nikita [0].
> But looks like Brendan is not working on this anymore. Will see
> if someone is really planning to look at it.
>
> [0] https://lore.kernel.org/all/DJJ35VLH2PE5.DFD8OYXEOH97@xxxxxxxxx
I just realised that this only solves part of the problem. Normal kernel
allocations are delegated RMM metadata and they'll also trigger GPF
faults. Kdump is another point raised here:
https://lore.kernel.org/all/a11aba04-eaaa-43fd-988b-a2581b1b2fdb@xxxxxxx/
I think rmi_delegate_range() covers the non-kdump cases, both for
metadata and realm memory, we could unmap the linear map there (map it
back in rmi_undelegate_range()).
For kdump, we have /proc/vmcore that may be mmap'ed or read() syscalls.
I think Alper proposed an idea to use copy_from_kernel_nofault() (with
exception recovery in do_gpf()) but we'd still need to block
mmap_vmcore() or somehow handle SIGBUS in makedumpfile unless it does
this already.
So, I think we should revisit the GPF handling. If we manage to unmap
all delegated pages, we could limit the exception handling to
is_kdump_kernel(). Otherwise we keep the exception handling as proposed
earlier, maybe with some warnings like EL0 access.
--
Catalin