Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
From: Christophe Leroy (CS GROUP)
Date: Sat Oct 10 2026 - 13:31:46 EST
Le 09/10/2026 à 22:13, Rosen Penev a écrit :
On Fri, Oct 9, 2026 at 9:23 AM Christophe Leroy (CS GROUP)
<chleroy@xxxxxxxxxx> wrote:
22.2% -> 13.4% of cycles, but the second profile also had a larger RX
Hi Rosen,
Le 08/10/2026 à 05:49, Rosen Penev a écrit :
__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.
Very nice patch. What is the new % after the change ?
ring (RXB 256) and was forwarding ~20% more traffic, so it understates
the reduction per byte. I can redo a clean A/B profile if useful.
Yes would be usefull
Single-stream iperf3 TCP routed through the board (LAN and WAN in
How do you test that ? I'd like to do some performance test on 8xx and 83xx.
separate netns on the host), GRO on, median of 3 runs, A/B in the same
boot. On the MX60 the TAH can't verify the checksum because of the
qca8k tag in front of the EtherType, so every forwarded packet is
checksummed by GRO in software. Any setup where GRO has to checksum
in software should show it; "perf record -a" during the run shows
the __csum_partial share.
- The 8xx distance. On 8xx, 3x is only 48 bytes, shorter than the
shortest distance tested on the 464 (64 bytes). The reviewer's 8xx
results may argue for a fixed distance in bytes instead of a multiple
of L1_CACHE_BYTES.
Let see
"Never touches it again" was poorly worded. I meant the loop reads each
When you say "never touches it again", do you mean the data is not used
again after that ? In that case would it help to do a 'dcbi' after using
the data in order to make the cache line available again and avoid
possible writeback to free a new cache line ?
byte once, so there is no reuse within the function. The caller may
well use the data (local receive followed by a copy to user space), and
csum_partial() is also called on buffers the CPU just wrote, where dcbi
would throw away dirty data. For the DMA'd RX case the lines are clean,
so evicting them costs no writeback anyway. I'll reword it in v2.
Yes of course, I should have thought twice before asking.
A reword will be welcome anyway.
AI answer:
Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.
On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
prefetch distance LAN->WAN WAN->LAN (Mbit/s)
none 370 341
64 bytes 430 381
96 bytes 438 397
128 bytes 438 385
160 bytes 434 386
256 bytes 423 383
3x CACHE_SIZE distance seems the most efficient, why did you choose to
implement 4x CACHE_SIZE ?
Why 4x
There wasn't a measured reason. In the original session's reasoning,
96 to 160 bytes were treated as the same within noise, and 128 was
picked from the middle of that range. The full log is less one-sided
than the table in the commit message:
dist=96 head=0 up=438(375-443) down=397(392-398) <- run once, one bad up run
dist=96 head=3 up=425(422-443) down=391(386-392)
dist=128 head=0 up=438(433-439) down=385(385-389)
dist=128 head=4 up=444(437-444) down=387(386-392)
dist=128 head=4 up=440(420-446) down=380(380-389)
pf128 small_rx=0 up=436(426-439) down=389(384-393)
pf128 small_rx=0 up=431(422-438) down=387(380-389)
dist=160 head=0 up=434(425-435) down=386(385-391)
128 was run five times, with WAN->LAN between 380 and 393. 96 was run
once with a clean configuration, and that run had an outlier upstream
sample of 375. The 397 vs 385 gap is about 3%, which is close to the
spread between 128 runs. So 96 may be slightly better, but the data
can't separate the two. The honest answer is "no strong reason, happy
to switch to 3x."
Ok
All of it. I set up a serial console to my Meraki MX60, sudo chmod 666
Prefetching the first lines before the loop made no measurable
difference and is not done.
Assisted-by: LLM
In what way did LLM help ?
/dev/ttyUSB0, and hooked up two ethernet cables to my desktop and its
WAN and one of its LAN ports so it could test routing performance.
The task I gave it was to speed up ethernet performance. I expected a
fair amount of work on its dated ethernet driver (IBM EMAC) but
because GRO (software checksumming actually) was unavoidable, it made
this change. It was using perf to find out where the bottlenecks were.
This is one place. I probably should have constrained it further.
All responses before the above were generated. I'll fix the patch up
once I get the MX60 set up again. I currently have a bcm47xx device
hooked up which codewise is in a really bad shape.
Nice to know.