From: David Laight <David.Laight@ACULAB.COM>
To: 'Robin Murphy' <robin.murphy@arm.com>,
Yang Yingliang <yangyingliang@huawei.com>,
"linux-arm-kernel@lists.infradead.org"
<linux-arm-kernel@lists.infradead.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Cc: "catalin.marinas@arm.com" <catalin.marinas@arm.com>,
"will@kernel.org" <will@kernel.org>,
"guohanjun@huawei.com" <guohanjun@huawei.com>
Subject: RE: [PATCH 2/3] arm64: lib: improve copy performance when size is ge 128 bytes
Date: Wed, 24 Mar 2021 16:38:31 +0000 [thread overview]
Message-ID: <62602598e7b742d09c581f3fc988e487@AcuMS.aculab.com> (raw)
In-Reply-To: <03ac41af-c433-cd66-8195-afbf9c49554c@arm.com>
From: Robin Murphy
> Sent: 23 March 2021 12:09
>
> On 2021-03-23 07:34, Yang Yingliang wrote:
> > When copy over 128 bytes, src/dst is added after
> > each ldp/stp instruction, it will cost more time.
> > To improve this, we only add src/dst after load
> > or store 64 bytes.
>
> This breaks the required behaviour for copy_*_user(), since the fault
> handler expects the base address to be up-to-date at all times. Say
> you're copying 128 bytes and fault on the 4th store, it should return 80
> bytes not copied; the code below would return 128 bytes not copied, even
> though 48 bytes have actually been written to the destination.
Are there any non-superscaler amd64 cpu (that anyone cares about)?
If the cpu can execute multiple instructions in one clock
then it is usually possible to get the loop control (almost) free.
You might need to unroll once to interleave read/write
but any more may be pointless.
So something like:
a = *src++
do {
b = *src++;
*dst++ = a;
a = *src++;
*dst++ = b;
} while (src != lim);
*dst++ = b;
David
-
Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK
Registration No: 1397386 (Wales)
next prev parent reply other threads:[~2021-03-24 16:39 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2021-03-23 7:34 [PATCH 0/3] arm64: lib: improve copy performance Yang Yingliang
2021-03-23 7:34 ` [PATCH 1/3] arm64: lib: introduce ldp2/stp2 macro Yang Yingliang
2021-03-23 7:34 ` [PATCH 2/3] arm64: lib: improve copy performance when size is ge 128 bytes Yang Yingliang
2021-03-23 12:08 ` Robin Murphy
2021-03-23 13:32 ` Will Deacon
2021-03-23 14:28 ` Robin Murphy
2021-03-23 15:03 ` Catalin Marinas
2021-03-24 16:38 ` David Laight [this message]
2021-03-24 19:36 ` Robin Murphy
2021-03-23 7:34 ` [PATCH 3/3] arm64: lib: improve copy performance when size is less than 128 and ge 64 bytes Yang Yingliang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=62602598e7b742d09c581f3fc988e487@AcuMS.aculab.com \
--to=david.laight@aculab.com \
--cc=catalin.marinas@arm.com \
--cc=guohanjun@huawei.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=robin.murphy@arm.com \
--cc=will@kernel.org \
--cc=yangyingliang@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®