Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
From: Christophe Leroy (CS GROUP)
Date: Fri Oct 09 2026 - 12:30:34 EST
Hi Rosen,
Le 08/10/2026 à 05:49, Rosen Penev a écrit :
__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.
Very nice patch. What is the new % after the change ?
How do you test that ? I'd like to do some performance test on 8xx and 83xx.
When you say "never touches it again", do you mean the data is not used again after that ? In that case would it help to do a 'dcbi' after using the data in order to make the cache line available again and avoid possible writeback to free a new cache line ?
Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.
On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
prefetch distance LAN->WAN WAN->LAN (Mbit/s)
none 370 341
64 bytes 430 381
96 bytes 438 397
128 bytes 438 385
160 bytes 434 386
256 bytes 423 383
3x CACHE_SIZE distance seems the most efficient, why did you choose to implement 4x CACHE_SIZE ?
Prefetching the first lines before the loop made no measurable
difference and is not done.
Assisted-by: LLM
In what way did LLM help ?
Signed-off-by: Rosen Penev <rosenp@xxxxxxxxx>
---
arch/powerpc/lib/checksum_32.S | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
index cd00b9bdd772..5a14d45c7b9d 100644
--- a/arch/powerpc/lib/checksum_32.S
+++ b/arch/powerpc/lib/checksum_32.S
@@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
bdnz 2b
21: srwi. r6,r4,4 /* # blocks of 4 words to do */
beq 3f
+ li r9,4*L1_CACHE_BYTES /* prefetch distance */
lwz r0,4(r3)
mtctr r6
lwz r6,8(r3)
@@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
lwzu r8,16(r3)
adde r5,r5,r7
bdz 23f
-22: lwz r0,4(r3)
+22: dcbt r3,r9
+ lwz r0,4(r3)
adde r5,r5,r8
lwz r6,8(r3)
adde r5,r5,r0