Re: [PATCH v18] arm64: mm: Handle Granule Protection Faults (GPFs)
From: Catalin Marinas
Date: Thu Sep 24 2026 - 11:49:17 EST
On Wed, Sep 23, 2026 at 05:04:06PM +0100, Suzuki K Poulose wrote:
> On 23/09/2026 16:45, Will Deacon wrote:
> > On Wed, Sep 23, 2026 at 12:06:15PM +0100, Catalin Marinas wrote:
> > > However, I'd still keep part of this patch - the reporting and panic but
> > > without the actual exception table recovery. There's some value in
> > > killing user-space and WARN (or pr_ratelimited) without a full panic, it
> > > helps with debugging. That's what do_bad() via arm64_notify_die() gives
> > > us currently anyway.
> > >
> > > So maybe we can keep it to just:
> > >
> > > static int do_gpf(unsigned long far, unsigned long esr, struct pt_regs *regs)
> > > {
> > > if (user_mode(regs)) {
> > > pr_alert_ratelimited("%s[%d]: granule protection fault at 0x%016lx\n",
> > > current->comm, task_pid_nr(current),
> > > untagged_addr(far));
> > > mem_abort_decode(esr);
> > > }
> > >
> > > return 1;
> > > }
> > >
> > > and we get the SIGBUS or panic via do_mem_abort(). No recovery for
> > > uaccess though, we get the same kernel panic.
> >
> > But how can this ever occur in user mode?
Only if we have a kernel bug.
> > I'm fine with making that part unconditional.
>
> Agree, if the user mode can hit this, a page is mapped in the EL0 and
> it can as well cause the Kernel to hit a GPF.
> Also if make the handling unconditional, we end up calling
> die_kernel_fault() and that does the mem_abort_decode() causing
> duplicate logs.
If we just return 1 here without anything printed (or rely on do_bad()),
we don't get any info when the user tripped over such pages. Printing
without the user_mode() check duplicates the mem_abort_decode() for the
kernel.
If we want panic always here even if only the user triggered it, we can
do like do_gpf_ptw() (and keep a single function for both). However,
die_kernel_fault() is a bit confusing as it prints "kernel access" when
it was user.
If we go with forced signal for EL0 faults (only helpful if we want to
continue debugging), I'd keep the warning, maybe as WARN_RATELIMIT() or
a printk. It would be very similar to our current do_bad() behaviour -
kernel => panic, user => kill, but with more information when it
happened in user space.
--
Catalin