mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
To: Rosen Penev <rosenp@gmail.com>
Cc: linuxppc-dev@lists.ozlabs.org,
	Madhavan Srinivasan <maddy@linux.ibm.com>,
	Michael Ellerman <mpe@ellerman.id.au>,
	Nicholas Piggin <npiggin@gmail.com>,
	"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
	Shrikanth Hegde <sshegde@linux.ibm.com>,
	open list <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] powerpc/lib: prefetch ahead in __csum_partial()
Date: Sat, 10 Oct 2026 19:31:35 +0200	[thread overview]
Message-ID: <cbc1ae4e-3e83-4f47-91e0-f4f2913b141e@kernel.org> (raw)
In-Reply-To: <CAKxU2N9imBpPHmJSefACvSR1FJofF9JY=ziRatmnQyKK9084gw@mail.gmail.com>



Le 09/10/2026 à 22:13, Rosen Penev a écrit :
> On Fri, Oct 9, 2026 at 9:23 AM Christophe Leroy (CS GROUP)
> <chleroy@kernel.org> wrote:
>>
>> Hi Rosen,
>>
>> Le 08/10/2026 à 05:49, Rosen Penev a écrit :
>>> __csum_partial() reads every byte of the buffer once and never touches
>>> it again, so on cores without a hardware prefetcher every cache line is
>>> a demand miss. After a non-coherent DMA the received data is never in
>>> the cache, which makes the GRO checksum validation of forwarded TCP
>>> traffic the top entry in the profile on a 464FP (APM82181): 22% of all
>>> cycles were spent in __csum_partial() when routing with GRO enabled.
>>
>> Very nice patch. What is the new % after the change ?
> 22.2% -> 13.4% of cycles, but the second profile also had a larger RX
> ring (RXB 256) and was forwarding ~20% more traffic, so it understates
> the reduction per byte. I can redo a clean A/B profile if useful.

Yes would be usefull

>>
>> How do you test that ? I'd like to do some performance test on 8xx and 83xx.
> Single-stream iperf3 TCP routed through the board (LAN and WAN in
> separate netns on the host), GRO on, median of 3 runs, A/B in the same
> boot. On the MX60 the TAH can't verify the checksum because of the
> qca8k tag in front of the EtherType, so every forwarded packet is
> checksummed by GRO in software. Any setup where GRO has to checksum
> in software should show it; "perf record -a" during the run shows
> the __csum_partial share.
> 
> - The 8xx distance. On 8xx, 3x is only 48 bytes, shorter than the
> shortest distance tested on the 464 (64 bytes). The reviewer's 8xx
> results may argue for a fixed distance in bytes instead of a multiple
> of L1_CACHE_BYTES.

Let see

>>
>>
>> When you say "never touches it again", do you mean the data is not used
>> again after that ? In that case would it help to do a 'dcbi' after using
>> the data in order to make the cache line available again and avoid
>> possible writeback to free a new cache line ?
> "Never touches it again" was poorly worded. I meant the loop reads each
> byte once, so there is no reuse within the function. The caller may
> well use the data (local receive followed by a copy to user space), and
> csum_partial() is also called on buffers the CPU just wrote, where dcbi
> would throw away dirty data. For the DMA'd RX case the lines are clean,
> so evicting them costs no writeback anyway. I'll reword it in v2.

Yes of course, I should have thought twice before asking.

A reword will be welcome anyway.

>>
>>>
>>> Issue a dcbt four cache lines ahead in the main loop. dcbt is a hint and
>>> never faults, so prefetching past the end of the buffer is harmless.
>>>
>>> On a Meraki MX60 (APM82181 at 800 MHz, single TCP stream routed through
>>> a qca8k DSA switch, iperf3 median of 3, A/B in the same boot):
>>>
>>>     prefetch distance   LAN->WAN   WAN->LAN  (Mbit/s)
>>>     none                370        341
>>>     64 bytes            430        381
>>>     96 bytes            438        397
>>>     128 bytes           438        385
>>>     160 bytes           434        386
>>>     256 bytes           423        383
>>
>> 3x CACHE_SIZE distance seems the most efficient, why did you choose to
>> implement 4x CACHE_SIZE ?
> AI answer:
> 
> Why 4x
> 
> There wasn't a measured reason. In the original session's reasoning,
> 96 to 160 bytes were treated as the same within noise, and 128 was
> picked from the middle of that range. The full log is less one-sided
> than the table in the commit message:
> 
> dist=96  head=0  up=438(375-443) down=397(392-398)   <- run once, one bad up run
> dist=96  head=3  up=425(422-443) down=391(386-392)
> dist=128 head=0  up=438(433-439) down=385(385-389)
> dist=128 head=4  up=444(437-444) down=387(386-392)
> dist=128 head=4  up=440(420-446) down=380(380-389)
> pf128 small_rx=0 up=436(426-439) down=389(384-393)
> pf128 small_rx=0 up=431(422-438) down=387(380-389)
> dist=160 head=0  up=434(425-435) down=386(385-391)
> 
> 128 was run five times, with WAN->LAN between 380 and 393. 96 was run
> once with a clean configuration, and that run had an outlier upstream
> sample of 375. The 397 vs 385 gap is about 3%, which is close to the
> spread between 128 runs. So 96 may be slightly better, but the data
> can't separate the two. The honest answer is "no strong reason, happy
> to switch to 3x."

Ok

>>
>>>
>>> Prefetching the first lines before the loop made no measurable
>>> difference and is not done.
>>>
>>> Assisted-by: LLM
>>
>> In what way did LLM help ?
> All of it. I set up a serial console to my Meraki MX60, sudo chmod 666
> /dev/ttyUSB0, and hooked up two ethernet cables to my desktop and its
> WAN and one of its LAN ports so it could test routing performance.
> 
> The task I gave it was to speed up ethernet performance. I expected a
> fair amount of work on its dated ethernet driver (IBM EMAC) but
> because GRO (software checksumming actually) was unavoidable, it made
> this change. It was using perf to find out where the bottlenecks were.
> This is one place. I probably should have constrained it further.
> 
> All responses before the above were generated. I'll fix the patch up
> once I get the MX60 set up again.  I currently have a bcm47xx device
> hooked up which codewise is in a really bad shape.

Nice to know.


      reply	other threads:[~2026-10-10 17:31 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08  3:49 Rosen Penev
2026-10-09 16:23 ` Christophe Leroy (CS GROUP)
2026-10-09 20:13   ` Rosen Penev
2026-10-10 17:31     ` Christophe Leroy (CS GROUP) [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cbc1ae4e-3e83-4f47-91e0-f4f2913b141e@kernel.org \
    --to=chleroy@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=maddy@linux.ibm.com \
    --cc=mpe@ellerman.id.au \
    --cc=npiggin@gmail.com \
    --cc=ritesh.list@gmail.com \
    --cc=rosenp@gmail.com \
    --cc=sshegde@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®