From: David Laight <David.Laight@ACULAB.COM>
To: 'Bibo Mao' <maobibo@loongson.cn>,
Huacai Chen <chenhuacai@kernel.org>,
WANG Xuerui <kernel@xen0n.name>
Cc: Jiaxun Yang <jiaxun.yang@flygoat.com>,
"loongarch@lists.linux.dev" <loongarch@lists.linux.dev>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: RE: [PATCH v2] LoongArch: add checksum optimization for 64-bit system
Date: Thu, 9 Feb 2023 09:35:29 +0000 [thread overview]
Message-ID: <e6bb59c32134477aa4890047ae5ad51b@AcuMS.aculab.com> (raw)
In-Reply-To: <20230209035839.2610277-1-maobibo@loongson.cn>
From: Bibo Mao
> Sent: 09 February 2023 03:59
>
> loongArch platform is 64-bit system, which supports 8 bytes memory
> accessing, generic checksum function uses 4 byte memory access.
> This patch adds 8-bytes memory access optimization for checksum
> function on loongArch. And the code comes from arm64 system.
How fast do these functions actually run (in bytes/clock)?
It is quite possible that just adding 32bit values to a
64bit register is faster.
Any non-trivial cpu will run that at 4 bytes/clock
(for suitably unrolled and pipelined code).
On a more complex cpu adding to two registers will
give 8 bytes/clock (needs two memory loads/clock).
The fastest 64bit sum you'll get on anything mips-like
(no carry flag) is probably from something like:
val = *mem++; // 64bit read
sum += val;
carry = sum < val;
carry_sum += carry;
which is 2 bytes/instruction again.
To get to 8 bytes/clock you need to execute all 4 instructions
every clock - so 1 read and 3 arithmetic.
(c/f 2 read and 2 arithmetic for 32bit adds.)
Arm has a carry flag so the code is:
val = *mem++;
temp,carry = sum + val;
sum = sum + val + carry;
There are still two dependant arithmetic instructions for
each 8-byte word.
The dependencies on the flags register also make it harder
to get any benefit from interleaving adds to two registers.
x86-64 uses 64bit 'add with carry' chains.
No one ever noticed that they take two clocks each on
Intel cpu until (about) Haswell.
It is possible to get 12 bytes/clock with some strange
loops that use (IIRC) adxo and adxc.
David
-
Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK
Registration No: 1397386 (Wales)
next prev parent reply other threads:[~2023-02-09 9:35 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-02-09 3:58 Bibo Mao
2023-02-09 9:35 ` David Laight [this message]
2023-02-09 11:55 ` maobibo
2023-02-09 12:39 ` David Laight
2023-02-10 3:21 ` Huacai Chen
2023-02-10 7:12 ` maobibo
2023-02-10 10:06 ` maobibo
2023-02-10 11:08 ` David Laight
2023-02-10 13:30 ` maobibo
2023-02-14 1:31 ` maobibo
2023-02-14 9:47 ` David Laight
2023-02-14 14:19 ` maobibo
2023-02-16 9:03 ` David Laight
2023-02-16 9:26 ` maobibo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e6bb59c32134477aa4890047ae5ad51b@AcuMS.aculab.com \
--to=david.laight@aculab.com \
--cc=chenhuacai@kernel.org \
--cc=jiaxun.yang@flygoat.com \
--cc=kernel@xen0n.name \
--cc=linux-kernel@vger.kernel.org \
--cc=loongarch@lists.linux.dev \
--cc=maobibo@loongson.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®