Re: [PATCH 0/2] arm64: csum: Add fused copy and Internet checksum
From: Jakub Kicinski
Date: Tue Sep 29 2026 - 21:33:16 EST
On Sun, 27 Sep 2026 15:17:56 +0200 Demian Shulhan wrote:
> arm64 currently uses the generic csum_partial_copy_nocheck(), which
> performs memcpy() followed by a second pass for csum_partial(). This
> double pass exerts unnecessary pressure on the L1 cache.
>
> Replace it with a single-pass implementation. The new implementation
> provides a general-purpose register path for short buffers and atomic
> contexts, and a kernel-mode NEON path for lengths >= 1024 bytes.
>
> Measured in-kernel on an Ampere Altra (Neoverse-N1):
> - Scalar path: 1.2x-1.6x faster for lengths < 1024 bytes.
> - NEON path: 1.2x faster at 1024 bytes, scaling up to 1.6x-1.8x at
> 4096 bytes.
> On Apple M-series cores, gains are 1.3-1.7x below 1024 bytes and
> 1.6-2.4x above. No length or alignment regresses on either
> microarchitecture.
>
> Patch 1 implements the fused routines and the dispatcher.
> Patch 2 adds KUnit test coverage for the new API and internal paths.
>
> Tested: in-kernel benchmark module on Neoverse-N1 with both
> implementations cross-checked (0 mismatches); KUnit suite under QEMU
> (with/without KASAN, with PREEMPT_RT), exhaustive and random userspace
> testing of both routines against a naive reference with PROT_NONE guard
> pages, gcc 13 and clang 18 W=1 builds, checkpatch --strict.
>
NIPA CI flagged a regression from this patch: the new "checksum" KUnit
suite fails on the x86-64 test kernel (ARCH=x86_64, qemu). Specifically:
test_csum_copy_small_all_alignments (len=0, src_off=0, dst_off=0)
test_csum_copy_patterns (len=1, src_off=0, dst_off=0)
test_csum_copy_zero_len
All three failures involve a zero-length (or very short) copy. Digging
into it, x86-64's csum_partial_copy_generic() (arch/x86/lib/csum-copy_64.S)
seeds its accumulator with -1 (0xffffffff) rather than 0:
movl $-1, %eax
...
cmpl $8, %ecx
jb .Lshort
For len == 0 (and other very short lengths that never execute an
add/adc against the seed) it returns that -1 unmodified, which folds to
0. The naive reference used by the new tests, and csum_partial()'s own
convention for an empty input, instead treat the "empty checksum" as raw
sum 0, which folds to 0xffff. So the new tests' expectations don't match
the pre-existing x86-64 assembly implementation for these edge cases.
It looks like the tests were validated against the new arm64
implementation and against memcpy()+csum_partial() under QEMU on arm64,
but not run against x86-64's existing csum_partial_copy_nocheck()
implementation, which is what our CI kunit runner builds by default.
Could you take a look at either:
- adjusting the zero/short-length expectations in the new test to
match the existing (documented?) x86-64 behavior, or
- treating this as a real x86-64 bug and fixing
csum_partial_copy_generic()'s handling of very short lengths,
whichever is judged correct? Happy to share the full kunit log if useful.