[PATCH] MIPS: lib: prefetch ahead in csum_partial()

From: Rosen Penev

Date: Sat Oct 10 2026 - 02:57:14 EST


csum_partial() reads every byte of the buffer once. Cores like the 74K
have no hardware prefetcher, so every cache line is a demand miss, and
after a non-coherent DMA the received data is never in the cache. With
GRO enabled, validating the checksum of forwarded TCP traffic made
csum_partial() the top entry in the profile on a BCM4716 (74Kc at
480 MHz, no L2, ~80 cycles per DRAM miss): 13-16% of all cycles.

Prefetch the 128 byte block two iterations ahead in the main loop.
Unlike memcpy(), which disables prefetching on DMA_NONCOHERENT
entirely, only prefetch blocks that lie inside the buffer: the 74K does
not need a post-DMA invalidate, so a line prefetched past the end could
belong to a buffer the device is still writing, and the CPU would later
read the stale copy.

On an Asus RT-N16 (BCM4716, bgmac, single TCP stream unless noted,
iperf3 median of 3, A/B in the same boot by toggling the prefetch,
two rounds each):

without with (Mbit/s)
local rx / tx 220 / 194 243 / 212
routed up / down 169 / 140 187 / 154
routed, 4 streams 65 / 79 69 / 84
flow offload bidir 253 292

Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@xxxxxxxxx>
---
arch/mips/lib/csum_partial.S | 15 +++++++++++++++
1 file changed, 15 insertions(+)

diff --git a/arch/mips/lib/csum_partial.S b/arch/mips/lib/csum_partial.S
index 3d2ff4118d79..f0ec47259ef4 100644
--- a/arch/mips/lib/csum_partial.S
+++ b/arch/mips/lib/csum_partial.S
@@ -190,6 +190,21 @@ EXPORT_SYMBOL(csum_partial)
andi t2, a1, 0x40

.Lmove_128bytes:
+#ifdef CONFIG_CPU_HAS_PREFETCH
+ /*
+ * Fetch the block two iterations ahead, but only while it lies
+ * inside the buffer: on non-coherent systems a line prefetched
+ * past the end could belong to a buffer the device still owns.
+ */
+ sltiu t5, t8, 3
+ bnez t5, 2f
+ nop
+ pref 0, 0x100(src)
+ pref 0, 0x120(src)
+ pref 0, 0x140(src)
+ pref 0, 0x160(src)
+2:
+#endif
CSUM_BIGCHUNK(src, 0x00, sum, t0, t1, t3, t4)
CSUM_BIGCHUNK(src, 0x20, sum, t0, t1, t3, t4)
CSUM_BIGCHUNK(src, 0x40, sum, t0, t1, t3, t4)
--
2.56.0