Re: [PATCH v18] arm64: mm: Handle Granule Protection Faults (GPFs)

From: Suzuki K Poulose

Date: Tue Sep 29 2026 - 06:59:57 EST


On 29/09/2026 11:30, Catalin Marinas wrote:
On Mon, Sep 28, 2026 at 11:38:01AM +0100, Suzuki K Poulose wrote:
On 25/09/2026 18:07, Catalin Marinas wrote:
On Wed, Sep 23, 2026 at 05:04:06PM +0100, Suzuki K Poulose wrote:
On 23/09/2026 16:45, Will Deacon wrote:
On Wed, Sep 23, 2026 at 12:06:15PM +0100, Catalin Marinas wrote:
On Tue, Sep 22, 2026 at 06:15:34PM +0100, Will Deacon wrote:
On Sun, Sep 13, 2026 at 08:04:58AM +0100, Suzuki K Poulose wrote:
From: Steven Price <steven.price@xxxxxxx>

If the host attempts to access granules that have been delegated for use
in a realm these accesses will be caught and will trigger a Granule
Protection Fault (GPF).

A fault during a page walk signals a bug in the kernel and is handled by
oopsing the kernel. A non-page walk fault could be caused by user space
having access to a page which has been delegated to the kernel and will
trigger a SIGBUS to allow debugging why user space is trying to access a
delegated page.

There is work in progress to unmap the guest_memfd backed private pages from the
linear map. Until we get that support, we could get spurious GPFs from within
the kernel, e.g., load_unaligned_zeropad(). So, try to fix them up for now.

Reviewed-by: Suzuki K Poulose <suzuki.poulose@xxxxxxx>
Reviewed-by: Gavin Shan <gshan@xxxxxxxxxx>
Reviewed-by: Catalin Marinas <catalin.marinas@xxxxxxx>
Signed-off-by: Steven Price <steven.price@xxxxxxx>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@xxxxxxx>
---
Changes since v17:
* Pass untagged address to die_kernel_fault() - Sashiko
* Explicitly check !user_mode() for fixups - Catalin
* Switch to BUS_OBJERR for si_code from SI_KERNEL - Catalin
* Clarify the commit description about the upcoming work on
unmapping guest_memfd backed pages from linear map
Changes since v16:
* Update the commit description to indicate why we try to fixup GPFs
Changes since v10:
* Don't call arm64_notify_die() in do_gpf() but simply return 1.
Changes since v2:
* Include missing "Granule Protection Fault at level -1"
---
arch/arm64/mm/fault.c | 30 ++++++++++++++++++++++++------
1 file changed, 24 insertions(+), 6 deletions(-)

I still don't think we should do this, given that the plan is to unmap
the memory from the linear map. If this thing fires, it's a kernel bug
and it should be fatal.

If the linear unmapping gets merged first, I agree, no need to handle
these faults. I haven't followed that series, so no idea where it is at.

There doesn't seem to be much progress on that series. Brendan
volunteered to resurrect the series, taking over from Nikita [0].
But looks like Brendan is not working on this anymore. Will see
if someone is really planning to look at it.

[0] https://lore.kernel.org/all/DJJ35VLH2PE5.DFD8OYXEOH97@xxxxxxxxx

I just realised that this only solves part of the problem. Normal kernel
allocations are delegated RMM metadata and they'll also trigger GPF

Other than Kdump, kernel shouldn't try to touch these metadata pages
unless there is a kernel bug ?

Hibernate but we need to reject this as well since there's no way you
can save and restore the realm memory.

Ack. I have disabled both kexec (including kdump for now, more on that
below) and hibernate when RMM is active.


And the userspace wouldn't have a mapping for them anyways ?

That's the aim but proposed KVM/CCA support still allows non-gmem slots
to end up delegated (I haven't checked the latest if addressed).

This has been addressed in v20 integration branch here :

https://git.gitlab.arm.com/linux-arm/linux-cca/-/commit/0e697f0f3b2c11855192ae6e80d6036e9374ccef?file_path=arch%2Farm64%2Fkvm%2Fmmu.c#line_b12897615_A2307


https://lore.kernel.org/all/a11aba04-eaaa-43fd-988b-a2581b1b2fdb@xxxxxxx/

I think rmi_delegate_range() covers the non-kdump cases, both for
metadata and realm memory, we could unmap the linear map there (map it
back in rmi_undelegate_range()).

We could explore that option for the kernel allocations too.

I do wonder what we gain from unmapping. If architecturally we get a
synchronous fault when accessing delegated pages, it's only marginally
more code to do_gpf() (and we might need this function anyway for
kdump) and we avoid linear map fragmentation. So, I think we just need
to agree what's fatal and what can safely recover. We do need Alper's
patch for kdump, otherwise we can't fix it up in do_gpf().

Ack. We could add that in later series. For now we block all kexecs.

Cheers
Suzuki