Re: [PATCH] ARM: mm: compare ARM926 cache range sizes as unsigned

From: Arnd Bergmann

Date: Tue Oct 06 2026 - 05:07:04 EST


On Tue, Oct 6, 2026, at 07:52, Karl Mehltretter wrote:
> arm926_flush_user_cache_range() uses BGT to compare end - start
> with CACHE_DLIMIT. Sizes of 2 GiB or more are treated as negative
> and take the cache-line loop instead of the whole-cache path.
>
> Use BHI for an unsigned comparison, keeping the line-by-line path
> at exactly CACHE_DLIMIT.
>
> With the default 3G/1G split, unmapping an untouched 2 GiB
> PROT_NONE mapping flushes the range through tlb_start_vma().
> On a SAM9X75 Curiosity board, munmap() took about 1.3 seconds
> unpatched and 20 ms patched.

Good catch!

> Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter@xxxxxxxxx>

Acked-by: Arnd Bergmann <arnd@xxxxxxxx>
CC: stable@xxxxxxxxxxxxxxx

> ---
> Based on v7.3-rc6. Tested on a Microchip EV31H43A SAM9X75 Curiosity
> LAN Kit with two kernels built from Linux4Microchip revision
> 8f2c610093aa3d5a7bd50a8d0597be4f0b7a3eda. The vendor function matches
> mainline.
>
> Both builds used Clang/LLD 22.1.3, the same DTB and an
> at91_dt_defconfig-derived config: ARM926T, MMU, CPU_CACHE_VIVT,
> VMSPLIT_3G, CPU_DCACHE_WRITETHROUGH=n. proc-arm926.o differs
> by one byte, the branch condition.
>
> An unprivileged test timed munmap() of untouched mappings made with
> mmap(PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS). Boot order was
> original, patched, original. Each cell is one munmap() call:
>
> Range Original A1 Patched B Original A2
> 16 KiB 0.0840 ms 0.0852 ms 0.0844 ms
> 20 KiB 0.0656 ms 0.0636 ms 0.0632 ms
> 1 GiB 9.7688 ms 9.7702 ms 9.7716 ms
> 2 GiB - 4 KiB 19.4054 ms 19.3928 ms 19.4066 ms
> 2 GiB 1323.2802 ms 19.6478 ms 1331.8680 ms
> 2 GiB + 4 KiB 1322.2016 ms 19.3972 ms 1322.2146 ms

I would suggest moving the extra text into the actual changelog,
as this seems useful information.

> ARM925, Feroceon and Mohawk have the same BGT but were not tested.
> Their fixes are left for a separate patch.

My feeling here would be that a single patch for all four
is best, since you have conclusively shown the solution to
be correct for one of them and it easily translates to the
others. I don't object to separate patches either, it just
seems more work for everyone.

Arnd