[PATCH] powerpc/lib: prefetch ahead in __csum_partial()

From: Rosen Penev

Date: Wed Oct 07 2026 - 23:49:42 EST


__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.

Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.

On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):

prefetch distance LAN->WAN WAN->LAN (Mbit/s)
none 370 341
64 bytes 430 381
96 bytes 438 397
128 bytes 438 385
160 bytes 434 386
256 bytes 423 383

Prefetching the first lines before the loop made no measurable
difference and is not done.

Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@xxxxxxxxx>
---
arch/powerpc/lib/checksum_32.S | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
index cd00b9bdd772..5a14d45c7b9d 100644
--- a/arch/powerpc/lib/checksum_32.S
+++ b/arch/powerpc/lib/checksum_32.S
@@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
bdnz 2b
21: srwi. r6,r4,4 /* # blocks of 4 words to do */
beq 3f
+ li r9,4*L1_CACHE_BYTES /* prefetch distance */
lwz r0,4(r3)
mtctr r6
lwz r6,8(r3)
@@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
lwzu r8,16(r3)
adde r5,r5,r7
bdz 23f
-22: lwz r0,4(r3)
+22: dcbt r3,r9
+ lwz r0,4(r3)
adde r5,r5,r8
lwz r6,8(r3)
adde r5,r5,r0
--
2.56.0