[PATCH 5.10.y/5.15.y 1/4] ARM: fix hash_name() fault
From: Xie Yuanbin
Date: Wed Sep 16 2026 - 04:20:21 EST
From: "Russell King (Oracle)" <rmk+kernel@xxxxxxxxxxxxxxx>
From: Russell King (Oracle) <rmk+kernel@xxxxxxxxxxxxxxx>
[ Upstream commit 7733bc7d299d682f2723dc38fc7f370b9bf973e9 ]
Zizhi Wo reports:
"During the execution of hash_name()->load_unaligned_zeropad(), a
potential memory access beyond the PAGE boundary may occur. For
example, when the filename length is near the PAGE_SIZE boundary.
This triggers a page fault, which leads to a call to
do_page_fault()->mmap_read_trylock(). If we can't acquire the lock,
we have to fall back to the mmap_read_lock() path, which calls
might_sleep(). This breaks RCU semantics because path lookup occurs
under an RCU read-side critical section."
This is seen with CONFIG_DEBUG_ATOMIC_SLEEP=y and CONFIG_KFENCE=y.
Kernel addresses (with the exception of the vectors/kuser helper
page) do not have VMAs associated with them. If the vectors/kuser
helper page faults, then there are two possibilities:
1. if the fault happened while in kernel mode, then we're basically
dead, because the CPU won't be able to vector through this page
to handle the fault.
2. if the fault happened while in user mode, that means the page was
protected from user access, and we want to fault anyway.
Thus, we can handle kernel addresses from any context entirely
separately without going anywhere near the mmap lock. This gives us
an entirely non-sleeping path for all kernel mode kernel address
faults.
As we handle the kernel address faults before interrupts are enabled,
this change has the side effect of improving the branch predictor
hardening, but does not completely solve the issue.
[ Xie Yuanbin: At the upstream, the following patches are a patch set:
1. commit dea20281ac8822661576 ("ARM: group is_permission_fault() with
is_translation_fault()")
2. commit 40b466db1dffb41f0529 ("ARM: allow __do_kernel_fault() to
report execution of memory faults")
3. commit 7733bc7d299d682f2723 ("ARM: fix hash_name() fault")
4. commit fd2dee1c6e2256f726ba ("ARM: fix branch predictor hardening")
patch 1. and 2. is unneeded for 5.10.y and 5.15.y . This patch backports
patch 3. and simply adapts to the context differences. ]
Reported-by: Zizhi Wo <wozizhi@xxxxxxxxxxxxxxx>
Reported-by: Xie Yuanbin <xieyuanbin1@xxxxxxxxxx>
Link: https://lore.kernel.org/r/20251126090505.3057219-1-wozizhi@xxxxxxxxxxxxxxx
Reviewed-by: Xie Yuanbin <xieyuanbin1@xxxxxxxxxx>
Tested-by: Xie Yuanbin <xieyuanbin1@xxxxxxxxxx>
Signed-off-by: Russell King (Oracle) <rmk+kernel@xxxxxxxxxxxxxxx>
---
arch/arm/mm/fault.c | 35 +++++++++++++++++++++++++++++++++++
1 file changed, 35 insertions(+)
diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c
index c16d6a293b97..094137cd29c8 100644
--- a/arch/arm/mm/fault.c
+++ b/arch/arm/mm/fault.c
@@ -247,6 +247,35 @@ __do_page_fault(struct mm_struct *mm, unsigned long addr, unsigned int fsr,
return fault;
}
+static int __kprobes
+do_kernel_address_page_fault(struct mm_struct *mm, unsigned long addr,
+ unsigned int fsr, struct pt_regs *regs)
+{
+ if (user_mode(regs)) {
+ /*
+ * Fault from user mode for a kernel space address. User mode
+ * should not be faulting in kernel space, which includes the
+ * vector/khelper page. Send a SIGSEGV.
+ */
+ __do_user_fault(addr, fsr, SIGSEGV, SEGV_MAPERR, regs);
+ } else {
+ /*
+ * Fault from kernel mode. Enable interrupts if they were
+ * enabled in the parent context. Section (upper page table)
+ * translation faults are handled via do_translation_fault(),
+ * so we will only get here for a non-present kernel space
+ * PTE or PTE permission fault. This may happen in exceptional
+ * circumstances and need the fixup tables to be walked.
+ */
+ if (interrupts_enabled(regs))
+ local_irq_enable();
+
+ __do_kernel_fault(mm, addr, fsr, regs);
+ }
+
+ return 0;
+}
+
static int __kprobes
do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
{
@@ -261,6 +290,12 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
tsk = current;
mm = tsk->mm;
+ /*
+ * Handle kernel addresses faults separately, which avoids touching
+ * the mmap lock from contexts that are not able to sleep.
+ */
+ if (addr >= TASK_SIZE)
+ return do_kernel_address_page_fault(mm, addr, fsr, regs);
/* Enable interrupts if they were enabled in the parent context. */
if (interrupts_enabled(regs))
--
2.55.0