From: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
To: Rosen Penev <rosenp@gmail.com>
Cc: linuxppc-dev@lists.ozlabs.org,
Madhavan Srinivasan <maddy@linux.ibm.com>,
Michael Ellerman <mpe@ellerman.id.au>,
Nicholas Piggin <npiggin@gmail.com>,
"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
open list <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
Date: Sat, 10 Oct 2026 19:31:35 +0200 [thread overview]
Message-ID: <cbc1ae4e-3e83-4f47-91e0-f4f2913b141e@kernel.org> (raw)
In-Reply-To: <CAKxU2N9imBpPHmJSefACvSR1FJofF9JY=ziRatmnQyKK9084gw@mail.gmail.com>
Le 09/10/2026 à 22:13, Rosen Penev a écrit :
> On Fri, Oct 9, 2026 at 9:23 AM Christophe Leroy (CS GROUP)
> <chleroy@kernel.org> wrote:
>>
>> Hi Rosen,
>>
>> Le 08/10/2026 à 05:49, Rosen Penev a écrit :
>>> __csum_partial() reads every byte of the buffer once and never touches
>>> it again, so on cores without a hardware prefetcher every cache line is
>>> a demand miss. After a non-coherent DMA the received data is never in
>>> the cache, which makes the GRO checksum validation of forwarded TCP
>>> traffic the top entry in the profile on a 464FP (APM82181): 22% of all
>>> cycles were spent in __csum_partial() when routing with GRO enabled.
>>
>> Very nice patch. What is the new % after the change ?
> 22.2% -> 13.4% of cycles, but the second profile also had a larger RX
> ring (RXB 256) and was forwarding ~20% more traffic, so it understates
> the reduction per byte. I can redo a clean A/B profile if useful.
Yes would be usefull
>>
>> How do you test that ? I'd like to do some performance test on 8xx and 83xx.
> Single-stream iperf3 TCP routed through the board (LAN and WAN in
> separate netns on the host), GRO on, median of 3 runs, A/B in the same
> boot. On the MX60 the TAH can't verify the checksum because of the
> qca8k tag in front of the EtherType, so every forwarded packet is
> checksummed by GRO in software. Any setup where GRO has to checksum
> in software should show it; "perf record -a" during the run shows
> the __csum_partial share.
>
> - The 8xx distance. On 8xx, 3x is only 48 bytes, shorter than the
> shortest distance tested on the 464 (64 bytes). The reviewer's 8xx
> results may argue for a fixed distance in bytes instead of a multiple
> of L1_CACHE_BYTES.
Let see
>>
>>
>> When you say "never touches it again", do you mean the data is not used
>> again after that ? In that case would it help to do a 'dcbi' after using
>> the data in order to make the cache line available again and avoid
>> possible writeback to free a new cache line ?
> "Never touches it again" was poorly worded. I meant the loop reads each
> byte once, so there is no reuse within the function. The caller may
> well use the data (local receive followed by a copy to user space), and
> csum_partial() is also called on buffers the CPU just wrote, where dcbi
> would throw away dirty data. For the DMA'd RX case the lines are clean,
> so evicting them costs no writeback anyway. I'll reword it in v2.
Yes of course, I should have thought twice before asking.
A reword will be welcome anyway.
>>
>>>
>>> Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
>>> never faults, so prefetching past the end of the buffer is harmless.
>>>
>>> On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
>>> a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
>>>
>>> prefetch distance LAN->WAN WAN->LAN (Mbit/s)
>>> none 370 341
>>> 64 bytes 430 381
>>> 96 bytes 438 397
>>> 128 bytes 438 385
>>> 160 bytes 434 386
>>> 256 bytes 423 383
>>
>> 3x CACHE_SIZE distance seems the most efficient, why did you choose to
>> implement 4x CACHE_SIZE ?
> AI answer:
>
> Why 4x
>
> There wasn't a measured reason. In the original session's reasoning,
> 96 to 160 bytes were treated as the same within noise, and 128 was
> picked from the middle of that range. The full log is less one-sided
> than the table in the commit message:
>
> dist=96 head=0 up=438(375-443) down=397(392-398) <- run once, one bad up run
> dist=96 head=3 up=425(422-443) down=391(386-392)
> dist=128 head=0 up=438(433-439) down=385(385-389)
> dist=128 head=4 up=444(437-444) down=387(386-392)
> dist=128 head=4 up=440(420-446) down=380(380-389)
> pf128 small_rx=0 up=436(426-439) down=389(384-393)
> pf128 small_rx=0 up=431(422-438) down=387(380-389)
> dist=160 head=0 up=434(425-435) down=386(385-391)
>
> 128 was run five times, with WAN->LAN between 380 and 393. 96 was run
> once with a clean configuration, and that run had an outlier upstream
> sample of 375. The 397 vs 385 gap is about 3%, which is close to the
> spread between 128 runs. So 96 may be slightly better, but the data
> can't separate the two. The honest answer is "no strong reason, happy
> to switch to 3x."
Ok
>>
>>>
>>> Prefetching the first lines before the loop made no measurable
>>> difference and is not done.
>>>
>>> Assisted-by: LLM
>>
>> In what way did LLM help ?
> All of it. I set up a serial console to my Meraki MX60, sudo chmod 666
> /dev/ttyUSB0, and hooked up two ethernet cables to my desktop and its
> WAN and one of its LAN ports so it could test routing performance.
>
> The task I gave it was to speed up ethernet performance. I expected a
> fair amount of work on its dated ethernet driver (IBM EMAC) but
> because GRO (software checksumming actually) was unavoidable, it made
> this change. It was using perf to find out where the bottlenecks were.
> This is one place. I probably should have constrained it further.
>
> All responses before the above were generated. I'll fix the patch up
> once I get the MX60 set up again. I currently have a bcm47xx device
> hooked up which codewise is in a really bad shape.
Nice to know.
prev parent reply other threads:[~2026-10-10 17:31 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 3:49 Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
2026-10-09 20:13 ` Rosen Penev
2026-10-10 17:31 ` Christophe Leroy (CS GROUP) [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cbc1ae4e-3e83-4f47-91e0-f4f2913b141e@kernel.org \
--to=chleroy@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=ritesh.list@gmail.com \
--cc=rosenp@gmail.com \
--cc=sshegde@linux.ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®