* [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
@ 2026-10-08 3:49 Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
0 siblings, 1 reply; 3+ messages in thread
From: Rosen Penev @ 2026-10-08 3:49 UTC (permalink / raw)
To: linuxppc-dev
Cc: Madhavan Srinivasan, Michael Ellerman, Nicholas Piggin,
Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
Shrikanth Hegde, open list
__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.
Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.
On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
prefetch distance LAN->WAN WAN->LAN (Mbit/s)
none 370 341
64 bytes 430 381
96 bytes 438 397
128 bytes 438 385
160 bytes 434 386
256 bytes 423 383
Prefetching the first lines before the loop made no measurable
difference and is not done.
Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@gmail.com>
---
arch/powerpc/lib/checksum_32.S | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
index cd00b9bdd772..5a14d45c7b9d 100644
--- a/arch/powerpc/lib/checksum_32.S
+++ b/arch/powerpc/lib/checksum_32.S
@@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
bdnz 2b
21: srwi. r6,r4,4 /* # blocks of 4 words to do */
beq 3f
+ li r9,4*L1_CACHE_BYTES /* prefetch distance */
lwz r0,4(r3)
mtctr r6
lwz r6,8(r3)
@@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
lwzu r8,16(r3)
adde r5,r5,r7
bdz 23f
-22: lwz r0,4(r3)
+22: dcbt r3,r9
+ lwz r0,4(r3)
adde r5,r5,r8
lwz r6,8(r3)
adde r5,r5,r0
--
2.56.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
2026-10-08 3:49 [PATCH] powerpc/lib: prefetch ahead in __csum_partial() Rosen Penev
@ 2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
2026-10-09 20:13 ` Rosen Penev
0 siblings, 1 reply; 3+ messages in thread
From: Christophe Leroy (CS GROUP) @ 2026-10-09 16:23 UTC (permalink / raw)
To: Rosen Penev, linuxppc-dev
Cc: Madhavan Srinivasan, Michael Ellerman, Nicholas Piggin,
Ritesh Harjani (IBM),
Shrikanth Hegde, open list
Hi Rosen,
Le 08/10/2026 à 05:49, Rosen Penev a écrit :
> __csum_partial() reads every byte of the buffer once and never touches
> it again, so on cores without a hardware prefetcher every cache line is
> a demand miss. After a non-coherent DMA the received data is never in
> the cache, which makes the GRO checksum validation of forwarded TCP
> traffic the top entry in the profile on a 464FP (APM82181): 22% of all
> cycles were spent in __csum_partial() when routing with GRO enabled.
Very nice patch. What is the new % after the change ?
How do you test that ? I'd like to do some performance test on 8xx and 83xx.
When you say "never touches it again", do you mean the data is not used
again after that ? In that case would it help to do a 'dcbi' after using
the data in order to make the cache line available again and avoid
possible writeback to free a new cache line ?
>
> Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
> never faults, so prefetching past the end of the buffer is harmless.
>
> On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
> a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
>
> prefetch distance LAN->WAN WAN->LAN (Mbit/s)
> none 370 341
> 64 bytes 430 381
> 96 bytes 438 397
> 128 bytes 438 385
> 160 bytes 434 386
> 256 bytes 423 383
3x CACHE_SIZE distance seems the most efficient, why did you choose to
implement 4x CACHE_SIZE ?
>
> Prefetching the first lines before the loop made no measurable
> difference and is not done.
>
> Assisted-by: LLM
In what way did LLM help ?
> Signed-off-by: Rosen Penev <rosenp@gmail.com>
> ---
> arch/powerpc/lib/checksum_32.S | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
> index cd00b9bdd772..5a14d45c7b9d 100644
> --- a/arch/powerpc/lib/checksum_32.S
> +++ b/arch/powerpc/lib/checksum_32.S
> @@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
> bdnz 2b
> 21: srwi. r6,r4,4 /* # blocks of 4 words to do */
> beq 3f
> + li r9,4*L1_CACHE_BYTES /* prefetch distance */
> lwz r0,4(r3)
> mtctr r6
> lwz r6,8(r3)
> @@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
> lwzu r8,16(r3)
> adde r5,r5,r7
> bdz 23f
> -22: lwz r0,4(r3)
> +22: dcbt r3,r9
> + lwz r0,4(r3)
> adde r5,r5,r8
> lwz r6,8(r3)
> adde r5,r5,r0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
@ 2026-10-09 20:13 ` Rosen Penev
0 siblings, 0 replies; 3+ messages in thread
From: Rosen Penev @ 2026-10-09 20:13 UTC (permalink / raw)
To: Christophe Leroy (CS GROUP)
Cc: linuxppc-dev, Madhavan Srinivasan, Michael Ellerman,
Nicholas Piggin, Ritesh Harjani (IBM),
Shrikanth Hegde, open list
On Fri, Oct 9, 2026 at 9:23 AM Christophe Leroy (CS GROUP)
<chleroy@kernel.org> wrote:
>
> Hi Rosen,
>
> Le 08/10/2026 à 05:49, Rosen Penev a écrit :
> > __csum_partial() reads every byte of the buffer once and never touches
> > it again, so on cores without a hardware prefetcher every cache line is
> > a demand miss. After a non-coherent DMA the received data is never in
> > the cache, which makes the GRO checksum validation of forwarded TCP
> > traffic the top entry in the profile on a 464FP (APM82181): 22% of all
> > cycles were spent in __csum_partial() when routing with GRO enabled.
>
> Very nice patch. What is the new % after the change ?
22.2% -> 13.4% of cycles, but the second profile also had a larger RX
ring (RXB 256) and was forwarding ~20% more traffic, so it understates
the reduction per byte. I can redo a clean A/B profile if useful.
>
> How do you test that ? I'd like to do some performance test on 8xx and 83xx.
Single-stream iperf3 TCP routed through the board (LAN and WAN in
separate netns on the host), GRO on, median of 3 runs, A/B in the same
boot. On the MX60 the TAH can't verify the checksum because of the
qca8k tag in front of the EtherType, so every forwarded packet is
checksummed by GRO in software. Any setup where GRO has to checksum
in software should show it; "perf record -a" during the run shows
the __csum_partial share.
- The 8xx distance. On 8xx, 3x is only 48 bytes, shorter than the
shortest distance tested on the 464 (64 bytes). The reviewer's 8xx
results may argue for a fixed distance in bytes instead of a multiple
of L1_CACHE_BYTES.
>
>
> When you say "never touches it again", do you mean the data is not used
> again after that ? In that case would it help to do a 'dcbi' after using
> the data in order to make the cache line available again and avoid
> possible writeback to free a new cache line ?
"Never touches it again" was poorly worded. I meant the loop reads each
byte once, so there is no reuse within the function. The caller may
well use the data (local receive followed by a copy to user space), and
csum_partial() is also called on buffers the CPU just wrote, where dcbi
would throw away dirty data. For the DMA'd RX case the lines are clean,
so evicting them costs no writeback anyway. I'll reword it in v2.
>
> >
> > Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
> > never faults, so prefetching past the end of the buffer is harmless.
> >
> > On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
> > a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
> >
> > prefetch distance LAN->WAN WAN->LAN (Mbit/s)
> > none 370 341
> > 64 bytes 430 381
> > 96 bytes 438 397
> > 128 bytes 438 385
> > 160 bytes 434 386
> > 256 bytes 423 383
>
> 3x CACHE_SIZE distance seems the most efficient, why did you choose to
> implement 4x CACHE_SIZE ?
AI answer:
Why 4x
There wasn't a measured reason. In the original session's reasoning,
96 to 160 bytes were treated as the same within noise, and 128 was
picked from the middle of that range. The full log is less one-sided
than the table in the commit message:
dist=96 head=0 up=438(375-443) down=397(392-398) <- run once, one bad up run
dist=96 head=3 up=425(422-443) down=391(386-392)
dist=128 head=0 up=438(433-439) down=385(385-389)
dist=128 head=4 up=444(437-444) down=387(386-392)
dist=128 head=4 up=440(420-446) down=380(380-389)
pf128 small_rx=0 up=436(426-439) down=389(384-393)
pf128 small_rx=0 up=431(422-438) down=387(380-389)
dist=160 head=0 up=434(425-435) down=386(385-391)
128 was run five times, with WAN->LAN between 380 and 393. 96 was run
once with a clean configuration, and that run had an outlier upstream
sample of 375. The 397 vs 385 gap is about 3%, which is close to the
spread between 128 runs. So 96 may be slightly better, but the data
can't separate the two. The honest answer is "no strong reason, happy
to switch to 3x."
>
> >
> > Prefetching the first lines before the loop made no measurable
> > difference and is not done.
> >
> > Assisted-by: LLM
>
> In what way did LLM help ?
All of it. I set up a serial console to my Meraki MX60, sudo chmod 666
/dev/ttyUSB0, and hooked up two ethernet cables to my desktop and its
WAN and one of its LAN ports so it could test routing performance.
The task I gave it was to speed up ethernet performance. I expected a
fair amount of work on its dated ethernet driver (IBM EMAC) but
because GRO (software checksumming actually) was unavoidable, it made
this change. It was using perf to find out where the bottlenecks were.
This is one place. I probably should have constrained it further.
All responses before the above were generated. I'll fix the patch up
once I get the MX60 set up again. I currently have a bcm47xx device
hooked up which codewise is in a really bad shape.
>
> > Signed-off-by: Rosen Penev <rosenp@gmail.com>
> > ---
> > arch/powerpc/lib/checksum_32.S | 4 +++-
> > 1 file changed, 3 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
> > index cd00b9bdd772..5a14d45c7b9d 100644
> > --- a/arch/powerpc/lib/checksum_32.S
> > +++ b/arch/powerpc/lib/checksum_32.S
> > @@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
> > bdnz 2b
> > 21: srwi. r6,r4,4 /* # blocks of 4 words to do */
> > beq 3f
> > + li r9,4*L1_CACHE_BYTES /* prefetch distance */
> > lwz r0,4(r3)
> > mtctr r6
> > lwz r6,8(r3)
> > @@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
> > lwzu r8,16(r3)
> > adde r5,r5,r7
> > bdz 23f
> > -22: lwz r0,4(r3)
> > +22: dcbt r3,r9
> > + lwz r0,4(r3)
> > adde r5,r5,r8
> > lwz r6,8(r3)
> > adde r5,r5,r0
>
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-09 20:13 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 3:49 [PATCH] powerpc/lib: prefetch ahead in __csum_partial() Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
2026-10-09 20:13 ` Rosen Penev
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®