mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Rosen Penev <rosenp@gmail.com>
To: linuxppc-dev@lists.ozlabs.org
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>,
	Michael Ellerman <mpe@ellerman.id.au>,
	Nicholas Piggin <npiggin@gmail.com>,
	"Christophe Leroy (CS GROUP)" <chleroy@kernel.org>,
	"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
	Shrikanth Hegde <sshegde@linux.ibm.com>,
	linux-kernel@vger.kernel.org (open list)
Subject: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
Date: Wed,  7 Oct 2026 20:49:32 -0700	[thread overview]
Message-ID: <20261008034932.3180490-1-rosenp@gmail.com> (raw)

__csum_partial() reads every byte of the buffer once and never touches
it again, so on cores without a hardware prefetcher every cache line is
a demand miss. After a non-coherent DMA the received data is never in
the cache, which makes the GRO checksum validation of forwarded TCP
traffic the top entry in the profile on a 464FP (APM82181): 22% of all
cycles were spent in __csum_partial() when routing with GRO enabled.

Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
never faults, so prefetching past the end of the buffer is harmless.

On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):

  prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
  none                370        341
  64 bytes            430        381
  96 bytes            438        397
  128 bytes           438        385
  160 bytes           434        386
  256 bytes           423        383

Prefetching the first lines before the loop made no measurable
difference and is not done.

Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@gmail.com>
---
 arch/powerpc/lib/checksum_32.S | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/arch/powerpc/lib/checksum_32.S b/arch/powerpc/lib/checksum_32.S
index cd00b9bdd772..5a14d45c7b9d 100644
--- a/arch/powerpc/lib/checksum_32.S
+++ b/arch/powerpc/lib/checksum_32.S
@@ -43,6 +43,7 @@ _GLOBAL(__csum_partial)
 	bdnz	2b
 21:	srwi.	r6,r4,4		/* # blocks of 4 words to do */
 	beq	3f
+	li	r9,4*L1_CACHE_BYTES	/* prefetch distance */
 	lwz	r0,4(r3)
 	mtctr	r6
 	lwz	r6,8(r3)
@@ -52,7 +53,8 @@ _GLOBAL(__csum_partial)
 	lwzu	r8,16(r3)
 	adde	r5,r5,r7
 	bdz	23f
-22:	lwz	r0,4(r3)
+22:	dcbt	r3,r9
+	lwz	r0,4(r3)
 	adde	r5,r5,r8
 	lwz	r6,8(r3)
 	adde	r5,r5,r0
-- 
2.56.0


                 reply	other threads:[~2026-10-08  3:49 UTC|newest]

Thread overview: [no followups] expand[flat|nested]  mbox.gz  Atom feed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008034932.3180490-1-rosenp@gmail.com \
    --to=rosenp@gmail.com \
    --cc=chleroy@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=maddy@linux.ibm.com \
    --cc=mpe@ellerman.id.au \
    --cc=npiggin@gmail.com \
    --cc=ritesh.list@gmail.com \
    --cc=sshegde@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®