mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
To: Rosen Penev <rosenp@gmail.com>, linuxppc-dev@lists.ozlabs.org
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>,
	Michael Ellerman <mpe@ellerman.id.au>,
	Nicholas Piggin <npiggin@gmail.com>,
	"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
	Shrikanth Hegde <sshegde@linux.ibm.com>,
	open list <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
Date: Fri, 9 Oct 2026 18:23:10 +0200	[thread overview]
Message-ID: <509cac24-1b3b-4b85-8ad2-ef90da521dd6@kernel.org> (raw)
In-Reply-To: <20261008034932.3180490-1-rosenp@gmail.com>

Hi Rosen,

Le 08/10/2026 à 05:49, Rosen Penev a écrit :
> __csum_partial() reads every byte of the buffer once and never touches
> it again, so on cores without a hardware prefetcher every cache line is
> a demand miss. After a non-coherent DMA the received data is never in
> the cache, which makes the GRO checksum validation of forwarded TCP
> traffic the top entry in the profile on a 464FP (APM82181): 22% of all
> cycles were spent in __csum_partial() when routing with GRO enabled.

Very nice patch. What is the new % after the change ?

How do you test that ? I'd like to do some performance test on 8xx and 83xx.


When you say "never touches it again", do you mean the data is not used 
again after that ? In that case would it help to do a 'dcbi' after using 
the data in order to make the cache line available again and avoid 
possible writeback to free a new cache line ?

> 
> Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
> never faults, so prefetching past the end of the buffer is harmless.
> 
> On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
> a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
> 
>    prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
>    none                370        341
>    64 bytes            430        381
>    96 bytes            438        397
>    128 bytes           438        385
>    160 bytes           434        386
>    256 bytes           423        383

3x CACHE_SIZE distance seems the most efficient, why did you choose to 
implement 4x CACHE_SIZE ?

> 
> Prefetching the first lines before the loop made no measurable
> difference and is not done.
> 
> Assisted-by: LLM

In what way did LLM help ?

> Signed-off-by: Rosen Penev <rosenp@gmail.com>
> ---
>   arch/powerpc/lib/checksum_32.S | 4 +++-
>   1 file changed, 3 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
> index cd00b9bdd772..5a14d45c7b9d 100644
> --- a/arch/powerpc/lib/checksum_32.S
> +++ b/arch/powerpc/lib/checksum_32.S
> @@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
>   	bdnz	2b
>   21:	srwi.	r6,r4,4		/* # blocks of 4 words to do */
>   	beq	3f
> +	li	r9,4*L1_CACHE_BYTES	/* prefetch distance */
>   	lwz	r0,4(r3)
>   	mtctr	r6
>   	lwz	r6,8(r3)
> @@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
>   	lwzu	r8,16(r3)
>   	adde	r5,r5,r7
>   	bdz	23f
> -22:	lwz	r0,4(r3)
> +22:	dcbt	r3,r9
> +	lwz	r0,4(r3)
>   	adde	r5,r5,r8
>   	lwz	r6,8(r3)
>   	adde	r5,r5,r0


  reply	other threads:[~2026-10-09 16:23 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08  3:49 Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP) [this message]
2026-10-09 20:13   ` Rosen Penev

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=509cac24-1b3b-4b85-8ad2-ef90da521dd6@kernel.org \
    --to=chleroy@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=maddy@linux.ibm.com \
    --cc=mpe@ellerman.id.au \
    --cc=npiggin@gmail.com \
    --cc=ritesh.list@gmail.com \
    --cc=rosenp@gmail.com \
    --cc=sshegde@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®