From: Borislav Petkov <bp@alien8.de>
To: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Andy Lutomirski <luto@amacapital.net>,
Andy Lutomirski <luto@kernel.org>, X86 ML <x86@kernel.org>,
"H. Peter Anvin" <hpa@zytor.com>,
Denys Vlasenko <vda.linux@googlemail.com>,
Brian Gerst <brgerst@gmail.com>,
Denys Vlasenko <dvlasenk@redhat.com>,
Ingo Molnar <mingo@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
Oleg Nesterov <oleg@redhat.com>,
Frederic Weisbecker <fweisbec@gmail.com>,
Alexei Starovoitov <ast@plumgrid.com>,
Will Drewry <wad@chromium.org>, Kees Cook <keescook@chromium.org>,
Linux Kernel Mailing List <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] x86_64, asm: Work around AMD SYSRET SS descriptor attribute issue
Date: Mon, 27 Apr 2015 21:11:45 +0200 [thread overview]
Message-ID: <20150427191145.GJ28871@pd.tnic> (raw)
In-Reply-To: <20150427183854.GG28871@pd.tnic>
On Mon, Apr 27, 2015 at 08:38:54PM +0200, Borislav Petkov wrote:
> I'm running them now and will report numbers relative to the last run
> once it is done. And those numbers should in practice get even better if
> we revert to the simpler canonical-ness check but let's see...
Results are done. New row is F: which is with the F16h NOPs.
With all things equal and with this change ontop:
---
diff --git a/arch/x86/kernel/alternative.c b/arch/x86/kernel/alternative.c
index aef653193160..d713080005ef 100644
--- a/arch/x86/kernel/alternative.c
+++ b/arch/x86/kernel/alternative.c
@@ -227,6 +227,14 @@ void __init arch_init_ideal_nops(void)
#endif
}
break;
+
+ case X86_VENDOR_AMD:
+ if (boot_cpu_data.x86 == 0x16) {
+ ideal_nops = p6_nops;
+ return;
+ }
+
+ /* fall through */
default:
#ifdef CONFIG_X86_64
ideal_nops = k8_nops;
---
... cycles, instructions, branches, branch-misses, context-switches
drop or remain roughly the same. BUT(!) timings increases.
cpu-clock/task-clock and duration of the workload are all the worst of
all possible cases.
So either those NOPs are not really optimal (i.e., trusting the manuals
and so on :-)) or it is their alignment.
But look at the chapter in the manual - "2.7.2.1 Encoding Padding for
Loop Alignment" - those NOPs are supposed to be used as padding so
they themselves will not be necessarily aligned when you use them to pad
stuff.
Or maybe using the longer NOPs is probably worse than the shorter 4-byte
ones with 3 0x66 prefixes which should "flow" easier through the pipe
due to their smaller length.
Or something completely different...
Oh well, enough measurements for today - will do the rc1 measurement
tomorrow.
Thanks.
---
Performance counter stats for 'system wide' (10 runs):
A: 2835570.145246 cpu-clock (msec) ( +- 0.02% ) [100.00%]
B: 2833364.074970 cpu-clock (msec) ( +- 0.04% ) [100.00%]
C: 2834708.335431 cpu-clock (msec) ( +- 0.02% ) [100.00%]
D: 2835055.118431 cpu-clock (msec) ( +- 0.01% ) [100.00%]
E: 2833115.118624 cpu-clock (msec) ( +- 0.06% ) [100.00%]
F: 2835863.670798 cpu-clock (msec) ( +- 0.02% ) [100.00%]
A: 2835570.099981 task-clock (msec) # 3.996 CPUs utilized ( +- 0.02% ) [100.00%]
B: 2833364.073633 task-clock (msec) # 3.996 CPUs utilized ( +- 0.04% ) [100.00%]
C: 2834708.350387 task-clock (msec) # 3.996 CPUs utilized ( +- 0.02% ) [100.00%]
D: 2835055.094383 task-clock (msec) # 3.996 CPUs utilized ( +- 0.01% ) [100.00%]
E: 2833115.145292 task-clock (msec) # 3.996 CPUs utilized ( +- 0.06% ) [100.00%]
F: 2835863.719556 task-clock (msec) # 3.996 CPUs utilized ( +- 0.02% ) [100.00%]
A: 5,591,213,166,613 cycles # 1.972 GHz ( +- 0.03% ) [75.00%]
B: 5,585,023,802,888 cycles # 1.971 GHz ( +- 0.03% ) [75.00%]
C: 5,587,983,212,758 cycles # 1.971 GHz ( +- 0.02% ) [75.00%]
D: 5,584,838,532,936 cycles # 1.970 GHz ( +- 0.03% ) [75.00%]
E: 5,583,979,727,842 cycles # 1.971 GHz ( +- 0.05% ) [75.00%]
F: 5,581,639,840,197 cycles # 1.968 GHz ( +- 0.03% ) [75.00%]
A: 3,106,707,101,530 instructions # 0.56 insns per cycle ( +- 0.01% ) [75.00%]
B: 3,106,632,251,528 instructions # 0.56 insns per cycle ( +- 0.00% ) [75.00%]
C: 3,106,265,958,142 instructions # 0.56 insns per cycle ( +- 0.00% ) [75.00%]
D: 3,106,294,801,185 instructions # 0.56 insns per cycle ( +- 0.00% ) [75.00%]
E: 3,106,381,223,355 instructions # 0.56 insns per cycle ( +- 0.01% ) [75.00%]
F: 3,105,996,162,436 instructions # 0.56 insns per cycle ( +- 0.00% ) [75.00%]
A: 683,676,044,429 branches # 241.107 M/sec ( +- 0.01% ) [75.00%]
B: 683,670,899,595 branches # 241.293 M/sec ( +- 0.01% ) [75.00%]
C: 683,675,772,858 branches # 241.180 M/sec ( +- 0.01% ) [75.00%]
D: 683,683,533,664 branches # 241.154 M/sec ( +- 0.00% ) [75.00%]
E: 683,648,518,667 branches # 241.306 M/sec ( +- 0.01% ) [75.00%]
F: 683,663,028,656 branches # 241.078 M/sec ( +- 0.00% ) [75.00%]
A: 43,829,535,008 branch-misses # 6.41% of all branches ( +- 0.02% ) [75.00%]
B: 43,844,118,416 branch-misses # 6.41% of all branches ( +- 0.03% ) [75.00%]
C: 43,819,871,086 branch-misses # 6.41% of all branches ( +- 0.02% ) [75.00%]
D: 43,795,107,998 branch-misses # 6.41% of all branches ( +- 0.02% ) [75.00%]
E: 43,801,985,070 branch-misses # 6.41% of all branches ( +- 0.02% ) [75.00%]
F: 43,804,449,271 branch-misses # 6.41% of all branches ( +- 0.02% ) [75.00%]
A: 2,030,357 context-switches # 0.716 K/sec ( +- 0.06% ) [100.00%]
B: 2,029,313 context-switches # 0.716 K/sec ( +- 0.05% ) [100.00%]
C: 2,028,566 context-switches # 0.716 K/sec ( +- 0.06% ) [100.00%]
D: 2,028,895 context-switches # 0.716 K/sec ( +- 0.06% ) [100.00%]
E: 2,031,008 context-switches # 0.717 K/sec ( +- 0.09% ) [100.00%]
F: 2,028,132 context-switches # 0.715 K/sec ( +- 0.05% ) [100.00%]
A: 52,421 migrations # 0.018 K/sec ( +- 1.13% )
B: 52,049 migrations # 0.018 K/sec ( +- 1.02% )
C: 51,365 migrations # 0.018 K/sec ( +- 0.92% )
D: 51,766 migrations # 0.018 K/sec ( +- 1.11% )
E: 53,047 migrations # 0.019 K/sec ( +- 1.08% )
F: 51,447 migrations # 0.018 K/sec ( +- 0.86% )
A: 709.528485252 seconds time elapsed ( +- 0.02% )
B: 708.976557288 seconds time elapsed ( +- 0.04% )
C: 709.312844791 seconds time elapsed ( +- 0.02% )
D: 709.400050112 seconds time elapsed ( +- 0.01% )
E: 708.914562508 seconds time elapsed ( +- 0.06% )
F: 709.602255085 seconds time elapsed ( +- 0.02% )
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
next prev parent reply other threads:[~2015-04-27 19:12 UTC|newest]
Thread overview: 67+ messages / expand[flat|nested] mbox.gz Atom feed top
2015-04-24 2:15 Andy Lutomirski
2015-04-24 2:18 ` Andy Lutomirski
2015-04-26 12:34 ` Denys Vlasenko
2015-04-24 3:58 ` Brian Gerst
2015-04-24 9:59 ` Denys Vlasenko
2015-04-24 10:59 ` Borislav Petkov
2015-04-24 19:58 ` Borislav Petkov
2015-04-24 11:27 ` Denys Vlasenko
2015-04-24 12:00 ` Brian Gerst
2015-04-24 16:25 ` Linus Torvalds
2015-04-24 17:33 ` Brian Gerst
2015-04-24 17:41 ` Linus Torvalds
2015-04-24 17:57 ` Brian Gerst
2015-04-24 20:21 ` Andy Lutomirski
2015-04-24 20:46 ` Denys Vlasenko
2015-04-24 20:50 ` Andy Lutomirski
2015-04-24 21:45 ` H. Peter Anvin
2015-04-24 21:45 ` H. Peter Anvin
2015-04-24 21:45 ` H. Peter Anvin
2015-04-24 21:45 ` H. Peter Anvin
2015-04-24 21:45 ` H. Peter Anvin
2015-04-24 21:45 ` H. Peter Anvin
2015-04-25 2:17 ` Denys Vlasenko
2015-04-26 23:36 ` Andy Lutomirski
2015-04-24 20:53 ` Linus Torvalds
2015-04-25 21:12 ` Borislav Petkov
2015-04-26 11:22 ` perf numbers (was: Re: [PATCH] x86_64, asm: Work around AMD SYSRET SS descriptor attribute issue) Borislav Petkov
2015-04-26 23:39 ` [PATCH] x86_64, asm: Work around AMD SYSRET SS descriptor attribute issue Andy Lutomirski
2015-04-27 8:53 ` Borislav Petkov
2015-04-27 10:07 ` Denys Vlasenko
2015-04-27 10:09 ` Borislav Petkov
2015-04-27 11:35 ` Borislav Petkov
2015-04-27 12:08 ` Denys Vlasenko
2015-04-27 12:48 ` Borislav Petkov
2015-04-27 14:57 ` Linus Torvalds
2015-04-27 15:06 ` Linus Torvalds
2015-04-27 15:35 ` Borislav Petkov
2015-04-27 15:46 ` Borislav Petkov
2015-04-27 15:56 ` Andy Lutomirski
2015-04-27 16:04 ` Brian Gerst
2015-04-27 16:10 ` Denys Vlasenko
2015-04-27 16:00 ` Linus Torvalds
2015-04-27 16:40 ` Borislav Petkov
2015-04-27 18:14 ` Linus Torvalds
2015-04-27 18:38 ` Borislav Petkov
2015-04-27 18:47 ` Linus Torvalds
2015-04-27 18:53 ` Borislav Petkov
2015-04-27 19:59 ` H. Peter Anvin
2015-04-27 20:03 ` Borislav Petkov
2015-04-27 20:14 ` H. Peter Anvin
2015-04-28 15:55 ` Borislav Petkov
2015-04-28 16:28 ` Linus Torvalds
2015-04-28 16:58 ` Borislav Petkov
2015-04-28 17:16 ` Linus Torvalds
2015-04-28 18:38 ` Borislav Petkov
2015-04-30 21:39 ` H. Peter Anvin
2015-04-30 23:23 ` H. Peter Anvin
2015-05-01 9:03 ` Borislav Petkov
2015-05-03 11:51 ` Borislav Petkov
2015-04-27 19:11 ` Borislav Petkov [this message]
2015-04-27 19:21 ` Denys Vlasenko
2015-04-27 19:45 ` Borislav Petkov
2015-04-28 13:40 ` Borislav Petkov
2015-04-27 16:12 ` Denys Vlasenko
2015-04-27 18:12 ` Linus Torvalds
2015-04-27 18:47 ` Borislav Petkov
2015-04-27 14:39 ` Borislav Petkov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20150427191145.GJ28871@pd.tnic \
--to=bp@alien8.de \
--cc=ast@plumgrid.com \
--cc=brgerst@gmail.com \
--cc=dvlasenk@redhat.com \
--cc=fweisbec@gmail.com \
--cc=hpa@zytor.com \
--cc=keescook@chromium.org \
--cc=linux-kernel@vger.kernel.org \
--cc=luto@amacapital.net \
--cc=luto@kernel.org \
--cc=mingo@kernel.org \
--cc=oleg@redhat.com \
--cc=rostedt@goodmis.org \
--cc=torvalds@linux-foundation.org \
--cc=vda.linux@googlemail.com \
--cc=wad@chromium.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome