From: "George Spelvin" <linux@horizon.com>
To: David.Laight@ACULAB.COM, linux-kernel@vger.kernel.org,
linux@horizon.com, netdev@vger.kernel.org, tom@herbertland.com
Cc: mingo@kernel.org
Subject: RE: [PATCH v3 net-next] net: Implement fast csum_partial for x86_64
Date: 10 Feb 2016 09:43:34 -0500 [thread overview]
Message-ID: <20160210144334.23242.qmail@ns.horizon.com> (raw)
In-Reply-To: <063D6719AE5E284EB5DD2968C1650D6D1CCDCC8D@AcuExch.aculab.com>
David Laight wrote:
> Separate renaming allows:
> 1) The value to tested without waiting for pending updates to complete.
> Useful for IE and DIR.
I don't quite follow. It allows the value to be tested without waiting
for pending updates *of other bits* to complete.
Obviusly, the update of the bit being tested has to complete!
> I can't see any obvious gain from separating out O or Z (even with
> adcx and adox). You'd need some other instructions that don't set O (or Z)
> but set some other useful flags.
> (A decrement that only set Z for instance.)
I tried to describe the advantages in the previous message.
The problems arise much less often than the INC/DEC pair, but there are
instructions whick write only the O and C flags, (ROL, ROR) and only
the Z flag (CMPXCHG).
The sign, aux carry, and parity flags are *always* updated as
a group, so they can be renamed as a group.
> While LOOP could be used on Bulldozer+ an equivalently fast loop
> can be done with inc/dec and jnz.
> So you only care about LOOP/JCXZ when ADOX is supported.
>
> I think the fastest loop is:
> 10: adc %rax,0(%rdi,%rcx,8)
> inc %rcx
> jnz 10b
> but check if any cpu add an extra clock for the 'scaled' offset
> (they might be faster if %rdi is incremented).
> That loop looks like it will have no overhead on recent cpu.
Well, it should execute at 1 instruction/cycle. (No, a scaled offset
doesn't take extra time.) To break that requires ADCX/ADOX:
10: adcxq 0(%rdi,%rcx),%rax
adoxq 8(%rdi,%rcx),%rdx
leaq 16(%rcx),%rcx
jrcxz 11f
j 10b
11:
next prev parent reply other threads:[~2016-02-10 14:43 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2016-02-08 20:12 George Spelvin
2016-02-09 10:48 ` David Laight
2016-02-10 0:53 ` George Spelvin
2016-02-10 11:39 ` David Laight
2016-02-10 14:43 ` George Spelvin [this message]
2016-02-10 15:18 ` David Laight
[not found] <1454527121-4007853-1-git-send-email-tom@herbertland.com>
2016-02-04 9:30 ` Ingo Molnar
2016-02-04 10:56 ` Ingo Molnar
2016-02-04 19:24 ` Tom Herbert
2016-02-05 9:24 ` Ingo Molnar
2016-02-04 21:46 ` Linus Torvalds
2016-02-04 22:09 ` Linus Torvalds
2016-02-05 1:27 ` Linus Torvalds
2016-02-05 1:39 ` Linus Torvalds
2016-02-04 22:43 ` Tom Herbert
2016-02-04 22:57 ` Linus Torvalds
2016-02-05 8:01 ` Ingo Molnar
2016-02-05 10:07 ` David Laight
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20160210144334.23242.qmail@ns.horizon.com \
--to=linux@horizon.com \
--cc=David.Laight@ACULAB.COM \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=tom@herbertland.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®