[PATCH 2/6] alpha: handle VM_FAULT_HWPOISON in do_page_fault()
From: Matt Turner
Date: Tue Oct 06 2026 - 10:04:50 EST
do_page_fault() handles VM_FAULT_OOM, VM_FAULT_SIGSEGV and
VM_FAULT_SIGBUS and then falls into BUG() for anything else in
VM_FAULT_ERROR. VM_FAULT_HWPOISON and VM_FAULT_HWPOISON_LARGE are in
that set, so a fault on a poisoned page takes down the kernel instead of
delivering a signal.
Alpha does not select ARCH_SUPPORTS_MEMORY_FAILURE, so the machine check
paths cannot produce these. UFFDIO_POISON can: it installs a poison PTE
marker without any memory failure support, and a subsequent access
returns VM_FAULT_HWPOISON. Registering an anonymous range with
userfaultfd, poisoning it and reading it back reliably hits the BUG():
Kernel bug at arch/alpha/mm/fault.c:188
poison(50809): Kernel Bug 1
pc is at do_page_fault+0x4e8/0x550
ra is at do_page_fault+0xfc/0x550
The BUG() fires with mmap_read_lock() still held, so the faulting task is left
unkillable in D state holding the lock, and shutdown stalls behind it.
Deliver SIGBUS with BUS_MCEERR_AR instead, reporting the size of the
poisoned area, as the other architectures do. That is the huge page size
for VM_FAULT_HWPOISON_LARGE, which becomes reachable once alpha
implements huge pages.
Tested on an UP1500 (EV68AL): the reproducer above now takes a SIGBUS
and the kernel logs
poison[373]: hardware memory error at 0000020000030000 pc 00000200010007ec
with no oops, no wedged task and no taint.
Fixes: fc71884a5f59 ("mm: userfaultfd: add new UFFDIO_POISON ioctl")
Signed-off-by: Matt Turner <mattst88@xxxxxxxxx>
---
arch/alpha/mm/fault.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/arch/alpha/mm/fault.c b/arch/alpha/mm/fault.c
index dfe427d93072..24408b53197c 100644
--- a/arch/alpha/mm/fault.c
+++ b/arch/alpha/mm/fault.c
@@ -8,6 +8,7 @@
#include <linux/sched/signal.h>
#include <linux/kernel.h>
#include <linux/mm.h>
+#include <linux/hugetlb.h>
#include <asm/io.h>
#define __EXTERN_INLINE inline
@@ -113,6 +114,7 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
struct mm_struct *mm = current->mm;
const struct exception_table_entry *fixup;
int si_code = SEGV_MAPERR;
+ unsigned int lsb;
vm_fault_t fault;
unsigned int flags = FAULT_FLAG_DEFAULT;
@@ -185,6 +187,8 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto bad_area;
else if (fault & VM_FAULT_SIGBUS)
goto do_sigbus;
+ else if (fault & (VM_FAULT_HWPOISON | VM_FAULT_HWPOISON_LARGE))
+ goto do_sigbus_mceerr;
BUG();
}
@@ -248,6 +252,21 @@ do_page_fault(unsigned long address, unsigned long mmcsr,
goto no_context;
return;
+ do_sigbus_mceerr:
+ mmap_read_unlock(mm);
+ if (!user_mode(regs))
+ goto no_context;
+ /*
+ * Report the size of the poisoned area, which for a hugetlb fault
+ * is the size of the huge page that could not be mapped.
+ */
+ lsb = PAGE_SHIFT;
+ if (fault & VM_FAULT_HWPOISON_LARGE)
+ lsb = hstate_index_to_shift(VM_FAULT_GET_HINDEX(fault));
+ show_signal_msg(regs, address, SIGBUS, "hardware memory error");
+ force_sig_mceerr(BUS_MCEERR_AR, (void __user *) address, lsb);
+ return;
+
do_sigsegv:
show_signal_msg(regs, address, SIGSEGV,
si_code == SEGV_MAPERR ? "unmapped access"
--
2.54.0