[PATCH v4] riscv: lib: Fix ZBB strnlen wrap-around regression on huge counts

From: shao.mingyin

Date: Tue Sep 15 2026 - 03:35:35 EST


From: Shao Mingyin <shao.mingyin@xxxxxxxxxx>

The aligned scan boundary is derived from the last valid byte,
(s + count - 1). When count is huge (e.g. SIZE_MAX, which FORTIFY
strcat/strlcat pass when the destination size is not known at compile
time), s + count wraps around and the boundary lands before s, so the
ZBB path returns a bogus length. The original implementation
(5ba15d419fab) had the same wrap-around in its (s + count) & ~7
boundary computation; after 5d588c684833 the wrapped boundary is caught
by the pre-loop guard "bgeu t0, t4, 2f", which then always exits for
aligned strings of 8 or more characters and strnlen() returns 8
instead of the real length.

This silently truncates strings built by fortified strcat: the dm
sysfs name attribute shows "live-bas" instead of "live-base", the
truncated name pollutes the udev database, and blivet/anaconda (as
well as LVM/dm-crypt/multipath userspace) break on RISC-V systems.

Detect the wrap-around and saturate the boundary to the top of the
address space, making the scan equivalent to strlen(). The saturation
clamps the increment to ~s, so it stays branchless and wrap-free:

s + min(count - 1, ~s) == saturating_add(s, count - 1)

Normal counts are unaffected.

Fixes: 5ba15d419fab ("riscv: lib: add strnlen() implementation")
Cc: stable@xxxxxxxxxxxxxxx
Suggested-by: David Laight <david.laight.linux@xxxxxxxxx>
Suggested-by: Qingfang Deng <qingfang.deng@xxxxxxxxx>
Signed-off-by: Shao Mingyin <shao.mingyin@xxxxxxxxxx>
Acked-by: Michael Neuling <mikey@xxxxxxxxxxx>
---
Changes in v4:
- Use Zbb minu to clamp the increment (addi/not/minu/add): one
instruction less than the sltu/mask/or sequence and no extra
register (Qingfang Deng). Clobbers stays t0-t4.

Changes in v3:
- Replace the taken branch in the saturation with a branchless
sltu/mask/or sequence (David Laight).
- Update the Clobbers list for the additional t5 register.
- Michael's Acked-by is kept: the patch semantics are unchanged, only
the saturation sequence is branchless now.

Changes in v2:
- Point Fixes: at the original implementation (5ba15d419fab) and reword
the commit message accordingly: the wrap-around exists since the
original implementation, 5d588c684833 only changed how it surfaces
(Michael Neuling).
- Add Acked-by from Michael Neuling.

v3: https://lore.kernel.org/all/20260914162123230u1y1M4UHrO8E-cU-opJ_7@xxxxxxxxxx/
v2: https://lore.kernel.org/all/20260914145205778-sZJbZc1D-XBfWRXO2f-o@xxxxxxxxxx/
v1: https://lore.kernel.org/all/20260828145152578tXQPUG9lxxgbJjmfpuaQz@xxxxxxxxxx/

arch/riscv/lib/strnlen.S | 20 ++++++++++++++++++--
1 file changed, 18 insertions(+), 2 deletions(-)

diff --git a/arch/riscv/lib/strnlen.S b/arch/riscv/lib/strnlen.S
index a8911605c248..5b3bbf0a3098 100644
--- a/arch/riscv/lib/strnlen.S
+++ b/arch/riscv/lib/strnlen.S
@@ -87,9 +87,25 @@ strnlen_zbb:
* Aligned boundary. Use the address of the last valid byte
* (s + count - 1) to avoid loading a word past the count
* boundary in the loop below. count == 0 is handled above.
+ *
+ * Saturate the boundary when s + count would wrap around (very
+ * large counts, e.g. SIZE_MAX passed by FORTIFY strcat/strlcat
+ * with a destination whose size is unknown at compile time).
+ * Without this, the wrapped boundary lands before s and the
+ * pre-loop guard below always exits, returning a truncated
+ * length.
+ *
+ * Clamping the increment to ~s (== SIZE_MAX - s) keeps the
+ * computation branchless and wrap-free:
+ *
+ * s + min(count - 1, ~s) == saturating_add(s, count - 1)
+ *
+ * Saturating makes the scan equivalent to strlen().
*/
- add t4, a0, a1
- addi t4, t4, -1
+ addi t4, a1, -1 /* count - 1 */
+ not t1, a0 /* SIZE_MAX - s */
+ minu t4, t4, t1 /* clamp, so s + t4 can never wrap */
+ add t4, a0, t4 /* saturated s + count - 1 */
andi t4, t4, -SZREG

/* Get the first word. */
--
2.27.0