mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Rosen Penev <rosenp@gmail.com>
To: linux-mips@vger.kernel.org
Cc: Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
	linux-kernel@vger.kernel.org (open list)
Subject: [PATCH] MIPS: lib: prefetch ahead in csum_partial()
Date: Fri,  9 Oct 2026 23:54:42 -0700	[thread overview]
Message-ID: <20261010065442.23780-1-rosenp@gmail.com> (raw)

csum_partial() reads every byte of the buffer once. Cores like the 74K
have no hardware prefetcher, so every cache line is a demand miss, and
after a non-coherent DMA the received data is never in the cache. With
GRO enabled, validating the checksum of forwarded TCP traffic made
csum_partial() the top entry in the profile on a BCM4716 (74Kc at
480 MHz, no L2, ~80 cycles per DRAM miss): 13-16% of all cycles.

Prefetch the 128 byte block two iterations ahead in the main loop.
Unlike memcpy(), which disables prefetching on DMA_NONCOHERENT
entirely, only prefetch blocks that lie inside the buffer: the 74K does
not need a post-DMA invalidate, so a line prefetched past the end could
belong to a buffer the device is still writing, and the CPU would later
read the stale copy.

On an Asus RT-N16 (BCM4716, bgmac, single TCP stream unless noted,
iperf3 median of 3, A/B in the same boot by toggling the prefetch,
two rounds each):

                       without      with      (Mbit/s)
  local rx / tx        220 / 194    243 / 212
  routed up / down     169 / 140    187 / 154
  routed, 4 streams    65 / 79      69 / 84
  flow offload bidir   253          292

Assisted-by: LLM
Signed-off-by: Rosen Penev <rosenp@gmail.com>
---
 arch/mips/lib/csum_partial.S | 15 +++++++++++++++
 1 file changed, 15 insertions(+)

diff --git a/arch/mips/lib/csum_partial.S b/arch/mips/lib/csum_partial.S
index 3d2ff4118d79..f0ec47259ef4 100644
--- a/arch/mips/lib/csum_partial.S
+++ b/arch/mips/lib/csum_partial.S
@@ -190,6 +190,21 @@ EXPORT_SYMBOL(csum_partial)
 	 andi	t2, a1, 0x40
 
 .Lmove_128bytes:
+#ifdef CONFIG_CPU_HAS_PREFETCH
+	/*
+	 * Fetch the block two iterations ahead, but only while it lies
+	 * inside the buffer: on non-coherent systems a line prefetched
+	 * past the end could belong to a buffer the device still owns.
+	 */
+	sltiu	t5, t8, 3
+	bnez	t5, 2f
+	 nop
+	pref	0, 0x100(src)
+	pref	0, 0x120(src)
+	pref	0, 0x140(src)
+	pref	0, 0x160(src)
+2:
+#endif
 	CSUM_BIGCHUNK(src, 0x00, sum, t0, t1, t3, t4)
 	CSUM_BIGCHUNK(src, 0x20, sum, t0, t1, t3, t4)
 	CSUM_BIGCHUNK(src, 0x40, sum, t0, t1, t3, t4)
-- 
2.56.0


                 reply	other threads:[~2026-10-10  6:54 UTC|newest]

Thread overview: [no followups] expand[flat|nested]  mbox.gz  Atom feed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261010065442.23780-1-rosenp@gmail.com \
    --to=rosenp@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mips@vger.kernel.org \
    --cc=tsbogend@alpha.franken.de \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®