mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
@ 2026-10-08  3:49 Rosen Penev
  2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
  0 siblings, 1 reply; 3+ messages in thread
From: Rosen Penev @ 2026-10-08  3:49 UTC (permalink / raw)
  To: linuxppc-dev
  Cc: Madhavan Srinivasan, Michael Ellerman, Nicholas Piggin,
	Christophe Leroy (CS GROUP), Ritesh Harjani (IBM),
	Shrikanth Hegde, open list

__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.

Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.

On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):

  prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
  none                370        341
  64 bytes            430        381
  96 bytes            438        397
  128 bytes           438        385
  160 bytes           434        386
  256 bytes           423        383

Prefetching the first lines before the loop made no measurable
difference and is not done.

Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@gmail.com>
---
 arch/powerpc/lib/checksum_32.S | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
index cd00b9bdd772..5a14d45c7b9d 100644
--- a/arch/powerpc/lib/checksum_32.S
+++ b/arch/powerpc/lib/checksum_32.S
@@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
 	bdnz	2b
 21:	srwi.	r6,r4,4		/* # blocks of 4 words to do */
 	beq	3f
+	li	r9,4*L1_CACHE_BYTES	/* prefetch distance */
 	lwz	r0,4(r3)
 	mtctr	r6
 	lwz	r6,8(r3)
@@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
 	lwzu	r8,16(r3)
 	adde	r5,r5,r7
 	bdz	23f
-22:	lwz	r0,4(r3)
+22:	dcbt	r3,r9
+	lwz	r0,4(r3)
 	adde	r5,r5,r8
 	lwz	r6,8(r3)
 	adde	r5,r5,r0
-- 
2.56.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
  2026-10-08  3:49 [PATCH] powerpc/lib: prefetch ahead in __csum_partial() Rosen Penev
@ 2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
  2026-10-09 20:13   ` Rosen Penev
  0 siblings, 1 reply; 3+ messages in thread
From: Christophe Leroy (CS GROUP) @ 2026-10-09 16:23 UTC (permalink / raw)
  To: Rosen Penev, linuxppc-dev
  Cc: Madhavan Srinivasan, Michael Ellerman, Nicholas Piggin,
	Ritesh Harjani (IBM),
	Shrikanth Hegde, open list

Hi Rosen,

Le 08/10/2026 à 05:49, Rosen Penev a écrit :
> __csum_partial() reads every byte of the buffer once and never touches
> it again, so on cores without a hardware prefetcher every cache line is
> a demand miss. After a non-coherent DMA the received data is never in
> the cache, which makes the GRO checksum validation of forwarded TCP
> traffic the top entry in the profile on a 464FP (APM82181): 22% of all
> cycles were spent in __csum_partial() when routing with GRO enabled.

Very nice patch. What is the new % after the change ?

How do you test that ? I'd like to do some performance test on 8xx and 83xx.


When you say "never touches it again", do you mean the data is not used 
again after that ? In that case would it help to do a 'dcbi' after using 
the data in order to make the cache line available again and avoid 
possible writeback to free a new cache line ?

> 
> Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
> never faults, so prefetching past the end of the buffer is harmless.
> 
> On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
> a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
> 
>    prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
>    none                370        341
>    64 bytes            430        381
>    96 bytes            438        397
>    128 bytes           438        385
>    160 bytes           434        386
>    256 bytes           423        383

3x CACHE_SIZE distance seems the most efficient, why did you choose to 
implement 4x CACHE_SIZE ?

> 
> Prefetching the first lines before the loop made no measurable
> difference and is not done.
> 
> Assisted-by: LLM

In what way did LLM help ?

> Signed-off-by: Rosen Penev <rosenp@gmail.com>
> ---
>   arch/powerpc/lib/checksum_32.S | 4 +++-
>   1 file changed, 3 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
> index cd00b9bdd772..5a14d45c7b9d 100644
> --- a/arch/powerpc/lib/checksum_32.S
> +++ b/arch/powerpc/lib/checksum_32.S
> @@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
>   	bdnz	2b
>   21:	srwi.	r6,r4,4		/* # blocks of 4 words to do */
>   	beq	3f
> +	li	r9,4*L1_CACHE_BYTES	/* prefetch distance */
>   	lwz	r0,4(r3)
>   	mtctr	r6
>   	lwz	r6,8(r3)
> @@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
>   	lwzu	r8,16(r3)
>   	adde	r5,r5,r7
>   	bdz	23f
> -22:	lwz	r0,4(r3)
> +22:	dcbt	r3,r9
> +	lwz	r0,4(r3)
>   	adde	r5,r5,r8
>   	lwz	r6,8(r3)
>   	adde	r5,r5,r0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
  2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
@ 2026-10-09 20:13   ` Rosen Penev
  0 siblings, 0 replies; 3+ messages in thread
From: Rosen Penev @ 2026-10-09 20:13 UTC (permalink / raw)
  To: Christophe Leroy (CS GROUP)
  Cc: linuxppc-dev, Madhavan Srinivasan, Michael Ellerman,
	Nicholas Piggin, Ritesh Harjani (IBM),
	Shrikanth Hegde, open list

On Fri, Oct 9, 2026 at 9:23 AM Christophe Leroy (CS GROUP)
<chleroy@kernel.org> wrote:
>
> Hi Rosen,
>
> Le 08/10/2026 à 05:49, Rosen Penev a écrit :
> > __csum_partial() reads every byte of the buffer once and never touches
> > it again, so on cores without a hardware prefetcher every cache line is
> > a demand miss. After a non-coherent DMA the received data is never in
> > the cache, which makes the GRO checksum validation of forwarded TCP
> > traffic the top entry in the profile on a 464FP (APM82181): 22% of all
> > cycles were spent in __csum_partial() when routing with GRO enabled.
>
> Very nice patch. What is the new % after the change ?
22.2% -> 13.4% of cycles, but the second profile also had a larger RX
ring (RXB 256) and was forwarding ~20% more traffic, so it understates
the reduction per byte. I can redo a clean A/B profile if useful.
>
> How do you test that ? I'd like to do some performance test on 8xx and 83xx.
Single-stream iperf3 TCP routed through the board (LAN and WAN in
separate netns on the host), GRO on, median of 3 runs, A/B in the same
boot. On the MX60 the TAH can't verify the checksum because of the
qca8k tag in front of the EtherType, so every forwarded packet is
checksummed by GRO in software. Any setup where GRO has to checksum
in software should show it; "perf record -a" during the run shows
the __csum_partial share.

- The 8xx distance. On 8xx, 3x is only 48 bytes, shorter than the
shortest distance tested on the 464 (64 bytes). The reviewer's 8xx
results may argue for a fixed distance in bytes instead of a multiple
of L1_CACHE_BYTES.
>
>
> When you say "never touches it again", do you mean the data is not used
> again after that ? In that case would it help to do a 'dcbi' after using
> the data in order to make the cache line available again and avoid
> possible writeback to free a new cache line ?
"Never touches it again" was poorly worded. I meant the loop reads each
byte once, so there is no reuse within the function. The caller may
well use the data (local receive followed by a copy to user space), and
csum_partial() is also called on buffers the CPU just wrote, where dcbi
would throw away dirty data. For the DMA'd RX case the lines are clean,
so evicting them costs no writeback anyway. I'll reword it in v2.
>
> >
> > Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
> > never faults, so prefetching past the end of the buffer is harmless.
> >
> > On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
> > a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
> >
> >    prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
> >    none                370        341
> >    64 bytes            430        381
> >    96 bytes            438        397
> >    128 bytes           438        385
> >    160 bytes           434        386
> >    256 bytes           423        383
>
> 3x CACHE_SIZE distance seems the most efficient, why did you choose to
> implement 4x CACHE_SIZE ?
AI answer:

Why 4x

There wasn't a measured reason. In the original session's reasoning,
96 to 160 bytes were treated as the same within noise, and 128 was
picked from the middle of that range. The full log is less one-sided
than the table in the commit message:

dist=96  head=0  up=438(375-443) down=397(392-398)   <- run once, one bad up run
dist=96  head=3  up=425(422-443) down=391(386-392)
dist=128 head=0  up=438(433-439) down=385(385-389)
dist=128 head=4  up=444(437-444) down=387(386-392)
dist=128 head=4  up=440(420-446) down=380(380-389)
pf128 small_rx=0 up=436(426-439) down=389(384-393)
pf128 small_rx=0 up=431(422-438) down=387(380-389)
dist=160 head=0  up=434(425-435) down=386(385-391)

128 was run five times, with WAN->LAN between 380 and 393. 96 was run
once with a clean configuration, and that run had an outlier upstream
sample of 375. The 397 vs 385 gap is about 3%, which is close to the
spread between 128 runs. So 96 may be slightly better, but the data
can't separate the two. The honest answer is "no strong reason, happy
to switch to 3x."
>
> >
> > Prefetching the first lines before the loop made no measurable
> > difference and is not done.
> >
> > Assisted-by: LLM
>
> In what way did LLM help ?
All of it. I set up a serial console to my Meraki MX60, sudo chmod 666
/dev/ttyUSB0, and hooked up two ethernet cables to my desktop and its
WAN and one of its LAN ports so it could test routing performance.

The task I gave it was to speed up ethernet performance. I expected a
fair amount of work on its dated ethernet driver (IBM EMAC) but
because GRO (software checksumming actually) was unavoidable, it made
this change. It was using perf to find out where the bottlenecks were.
This is one place. I probably should have constrained it further.

All responses before the above were generated. I'll fix the patch up
once I get the MX60 set up again.  I currently have a bcm47xx device
hooked up which codewise is in a really bad shape.
>
> > Signed-off-by: Rosen Penev <rosenp@gmail.com>
> > ---
> >   arch/powerpc/lib/checksum_32.S | 4 +++-
> >   1 file changed, 3 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
> > index cd00b9bdd772..5a14d45c7b9d 100644
> > --- a/arch/powerpc/lib/checksum_32.S
> > +++ b/arch/powerpc/lib/checksum_32.S
> > @@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
> >       bdnz    2b
> >   21: srwi.   r6,r4,4         /* # blocks of 4 words to do */
> >       beq     3f
> > +     li      r9,4*L1_CACHE_BYTES     /* prefetch distance */
> >       lwz     r0,4(r3)
> >       mtctr   r6
> >       lwz     r6,8(r3)
> > @@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
> >       lwzu    r8,16(r3)
> >       adde    r5,r5,r7
> >       bdz     23f
> > -22:  lwz     r0,4(r3)
> > +22:  dcbt    r3,r9
> > +     lwz     r0,4(r3)
> >       adde    r5,r5,r8
> >       lwz     r6,8(r3)
> >       adde    r5,r5,r0
>

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-09 20:13 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08  3:49 [PATCH] powerpc/lib: prefetch ahead in __csum_partial() Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
2026-10-09 20:13   ` Rosen Penev

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®