* Re: Hyper-Threading Vulnerability
2005-05-13 21:26 ` Andy Isaacson
@ 2005-05-13 21:59 ` Matt Mackall
2005-05-13 22:47 ` Alan Cox
2005-05-14 0:39 ` dean gaudet
` (3 subsequent siblings)
4 siblings, 1 reply; 84+ messages in thread
From: Matt Mackall @ 2005-05-13 21:59 UTC (permalink / raw)
To: Andy Isaacson
Cc: Andi Kleen, Richard F. Rebel, Gabor MICSKO, linux-kernel, tytso
On Fri, May 13, 2005 at 02:26:20PM -0700, Andy Isaacson wrote:
> On Fri, May 13, 2005 at 09:05:49PM +0200, Andi Kleen wrote:
> > On Fri, May 13, 2005 at 02:38:03PM -0400, Richard F. Rebel wrote:
> > > Why? It's certainly reasonable to disable it for the time being and
> > > even prudent to do so.
> >
> > No, i strongly disagree on that. The reasonable thing to do is
> > to fix the crypto code which has this vulnerability, not break
> > a useful performance enhancement for everybody else.
>
> Pardon me for saying so, but that's bullshit. You're asking the crypto
> guys to give up a 5x performance gain (that's my wild guess) by giving
> up all their data-dependent algorithms and contorting their code wildly,
> to avoid a microarchitectural problem with Intel's HT implementation.
>
> There are three places to cut off the side channel, none of which is
> obviously the right one.
> 1. The HT implementation could do the cache tricks Colin suggested in
> his paper. Fairly large performance hit to address a fairly small
> problem.
> 2. The OS could do the scheduler tricks to avoid scheduling unfriendly
> threads on the same core. You're leaving a lot of the benefit of HT
> on the floor by doing so.
> 3. Every security-sensitive app can be rigorously audited and re-written
> to avoid *ever* referencing memory with the address determined by
> private data.
>
> (3) is a complete non-starter. It's just not feasible to rewrite all
> that code. Furthermore, there's no way to know what code needs to be
> rewritten! (Until someone publishes an advisory, that is...)
>
> Hmm, I can't think of any reason that this technique wouldn't work to
> extract information from kernel secrets, as well...
>
> If SHA has plaintext-dependent memory references, Colin's technique
> would enable an adversary to extract the contents of the /dev/random
> pools. I don't *think* SHA does, based on a quick reading of
> lib/sha1.c, but someone with an actual clue should probably take a look.
SHA1 should be fine, as are the pool mixing bits. Much more
problematic is the ability to do timing attacks against the entropy
gathering itself. If an attacker can guess the TSC value that gets
mixed into the pool, that's a problem.
It might not be much of a problem though. If he's a bit off per guess
(really impressive), he'll still be many bits off by the time there's
enough entropy in the primary pool to reseed the secondary pool so he
can check his guesswork.
--
Mathematics is the supreme nostalgia of our time.
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-13 21:59 ` Matt Mackall
@ 2005-05-13 22:47 ` Alan Cox
2005-05-13 23:00 ` Lee Revell
0 siblings, 1 reply; 84+ messages in thread
From: Alan Cox @ 2005-05-13 22:47 UTC (permalink / raw)
To: Matt Mackall
Cc: Andy Isaacson, Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> It might not be much of a problem though. If he's a bit off per guess
> (really impressive), he'll still be many bits off by the time there's
> enough entropy in the primary pool to reseed the secondary pool so he
> can check his guesswork.
You can also disable the tsc to user space in the intel processors.
Thats something they anticipated as being neccessary in secure
environments long ago. This makes the attack much harder.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 22:47 ` Alan Cox
@ 2005-05-13 23:00 ` Lee Revell
2005-05-13 23:27 ` Dave Jones
0 siblings, 1 reply; 84+ messages in thread
From: Lee Revell @ 2005-05-13 23:00 UTC (permalink / raw)
To: Alan Cox
Cc: Matt Mackall, Andy Isaacson, Andi Kleen, Richard F. Rebel,
Gabor MICSKO, Linux Kernel Mailing List, tytso
On Fri, 2005-05-13 at 23:47 +0100, Alan Cox wrote:
> On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> > It might not be much of a problem though. If he's a bit off per guess
> > (really impressive), he'll still be many bits off by the time there's
> > enough entropy in the primary pool to reseed the secondary pool so he
> > can check his guesswork.
>
> You can also disable the tsc to user space in the intel processors.
> Thats something they anticipated as being neccessary in secure
> environments long ago. This makes the attack much harder.
And break the hundreds of apps that depend on rdtsc? Am I missing
something?
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 23:00 ` Lee Revell
@ 2005-05-13 23:27 ` Dave Jones
2005-05-13 23:38 ` Lee Revell
0 siblings, 1 reply; 84+ messages in thread
From: Dave Jones @ 2005-05-13 23:27 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Fri, May 13, 2005 at 07:00:12PM -0400, Lee Revell wrote:
> On Fri, 2005-05-13 at 23:47 +0100, Alan Cox wrote:
> > On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> > > It might not be much of a problem though. If he's a bit off per guess
> > > (really impressive), he'll still be many bits off by the time there's
> > > enough entropy in the primary pool to reseed the secondary pool so he
> > > can check his guesswork.
> >
> > You can also disable the tsc to user space in the intel processors.
> > Thats something they anticipated as being neccessary in secure
> > environments long ago. This makes the attack much harder.
>
> And break the hundreds of apps that depend on rdtsc? Am I missing
> something?
If those apps depend on rdtsc being a) present, and b) working
without providing fallbacks, they're already broken.
There's a reason its displayed in /proc/cpuinfo's flags field,
and visible through cpuid. Apps should be testing for presence
before assuming features are present.
Dave
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 23:27 ` Dave Jones
@ 2005-05-13 23:38 ` Lee Revell
2005-05-13 23:44 ` Dave Jones
2005-05-14 15:23 ` Alan Cox
0 siblings, 2 replies; 84+ messages in thread
From: Lee Revell @ 2005-05-13 23:38 UTC (permalink / raw)
To: Dave Jones
Cc: Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Fri, 2005-05-13 at 19:27 -0400, Dave Jones wrote:
> On Fri, May 13, 2005 at 07:00:12PM -0400, Lee Revell wrote:
> > On Fri, 2005-05-13 at 23:47 +0100, Alan Cox wrote:
> > > On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> > > > It might not be much of a problem though. If he's a bit off per guess
> > > > (really impressive), he'll still be many bits off by the time there's
> > > > enough entropy in the primary pool to reseed the secondary pool so he
> > > > can check his guesswork.
> > >
> > > You can also disable the tsc to user space in the intel processors.
> > > Thats something they anticipated as being neccessary in secure
> > > environments long ago. This makes the attack much harder.
> >
> > And break the hundreds of apps that depend on rdtsc? Am I missing
> > something?
>
> If those apps depend on rdtsc being a) present, and b) working
> without providing fallbacks, they're already broken.
>
> There's a reason its displayed in /proc/cpuinfo's flags field,
> and visible through cpuid. Apps should be testing for presence
> before assuming features are present.
>
Well yes but you would still have to recompile those apps. And take the
big performance hit from using gettimeofday vs rdtsc. Disabling HT by
default looks pretty good by comparison.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 23:38 ` Lee Revell
@ 2005-05-13 23:44 ` Dave Jones
2005-05-14 7:37 ` Lee Revell
2005-05-14 15:23 ` Alan Cox
1 sibling, 1 reply; 84+ messages in thread
From: Dave Jones @ 2005-05-13 23:44 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Fri, May 13, 2005 at 07:38:08PM -0400, Lee Revell wrote:
> On Fri, 2005-05-13 at 19:27 -0400, Dave Jones wrote:
> > On Fri, May 13, 2005 at 07:00:12PM -0400, Lee Revell wrote:
> > > On Fri, 2005-05-13 at 23:47 +0100, Alan Cox wrote:
> > > > On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> > > > > It might not be much of a problem though. If he's a bit off per guess
> > > > > (really impressive), he'll still be many bits off by the time there's
> > > > > enough entropy in the primary pool to reseed the secondary pool so he
> > > > > can check his guesswork.
> > > >
> > > > You can also disable the tsc to user space in the intel processors.
> > > > Thats something they anticipated as being neccessary in secure
> > > > environments long ago. This makes the attack much harder.
> > >
> > > And break the hundreds of apps that depend on rdtsc? Am I missing
> > > something?
> >
> > If those apps depend on rdtsc being a) present, and b) working
> > without providing fallbacks, they're already broken.
> >
> > There's a reason its displayed in /proc/cpuinfo's flags field,
> > and visible through cpuid. Apps should be testing for presence
> > before assuming features are present.
> >
>
> Well yes but you would still have to recompile those apps.
Not if the app is written correctly. See above.
Dave
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 23:44 ` Dave Jones
@ 2005-05-14 7:37 ` Lee Revell
2005-05-14 15:33 ` Andrea Arcangeli
0 siblings, 1 reply; 84+ messages in thread
From: Lee Revell @ 2005-05-14 7:37 UTC (permalink / raw)
To: Dave Jones
Cc: Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Fri, 2005-05-13 at 19:44 -0400, Dave Jones wrote:
> On Fri, May 13, 2005 at 07:38:08PM -0400, Lee Revell wrote:
> > On Fri, 2005-05-13 at 19:27 -0400, Dave Jones wrote:
> > > On Fri, May 13, 2005 at 07:00:12PM -0400, Lee Revell wrote:
> > > > On Fri, 2005-05-13 at 23:47 +0100, Alan Cox wrote:
> > > > > On Gwe, 2005-05-13 at 22:59, Matt Mackall wrote:
> > > > > > It might not be much of a problem though. If he's a bit off per guess
> > > > > > (really impressive), he'll still be many bits off by the time there's
> > > > > > enough entropy in the primary pool to reseed the secondary pool so he
> > > > > > can check his guesswork.
> > > > >
> > > > > You can also disable the tsc to user space in the intel processors.
> > > > > Thats something they anticipated as being neccessary in secure
> > > > > environments long ago. This makes the attack much harder.
> > > >
> > > > And break the hundreds of apps that depend on rdtsc? Am I missing
> > > > something?
> > >
> > > If those apps depend on rdtsc being a) present, and b) working
> > > without providing fallbacks, they're already broken.
> > >
> > > There's a reason its displayed in /proc/cpuinfo's flags field,
> > > and visible through cpuid. Apps should be testing for presence
> > > before assuming features are present.
> > >
> >
> > Well yes but you would still have to recompile those apps.
>
> Not if the app is written correctly. See above.
The apps that bother to use rdtsc vs. gettimeofday need a cheap high res
timer more than a correct one anyway - it's not guaranteed that rdtsc
provides a reliable time source at all, due to SMP and frequency scaling
issues.
I'll try to benchmark the difference. Maybe it's not that big a deal.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 7:37 ` Lee Revell
@ 2005-05-14 15:33 ` Andrea Arcangeli
2005-05-15 1:07 ` Christer Weinigel
2005-05-15 9:48 ` Andi Kleen
0 siblings, 2 replies; 84+ messages in thread
From: Andrea Arcangeli @ 2005-05-14 15:33 UTC (permalink / raw)
To: Lee Revell
Cc: Dave Jones, Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, May 14, 2005 at 03:37:18AM -0400, Lee Revell wrote:
> The apps that bother to use rdtsc vs. gettimeofday need a cheap high res
> timer more than a correct one anyway - it's not guaranteed that rdtsc
> provides a reliable time source at all, due to SMP and frequency scaling
> issues.
On x86-64 the cost of gettimeofday is the same of the tsc, turning off
tsc on x86-64 is not nice (even if we usually have HPET there, so
perhaps it wouldn't be too bad). TSC is something only the kernel (or a
person with some kernel/hardware knowledge) can do safely knowing it'll
work fine. But on x86-64 parts of the kernel runs in userland...
Preventing tasks with different uid to run on the same physical cpu was
my first idea, disabled by default via sysctl, so only if one is
paranoid can enable it.
But before touching the kernel in any way it would be really nice if
somebody could bother to demonstrate this is real because I've an hard
time to believe this is not purely vapourware. On artificial
environments a computer can recognize the difference between two faces
too, no big deal, but that doesn't mean the same software is going to
recognize million of different faces at the airport too. So nothing has
been demonstrated in practical terms yet.
Nobody runs openssl -sign thousand of times in a row on a pure idle
system without noticing the 100% load on the other cpu for months (and
he's not root so he can't hide his runaway 100% process, if he was root
and he could modify the kernel or ps/top to hide the runaway process,
he'd have faster ways to sniff).
So to me this sounds a purerly theoretical problem. Cache covert
channels are possible too as the paper states, next time somebody will
find how to sniff a letter out of a pdf document on a UP no-HT system by
opening and closing it some million of times on a otherwise idle system.
We're sure not going to flush the l2 cache because of that (at least not
by default ;).
This was an interesting read, but in practice I'd rate this to have
severity 1 on a 0-100 scale, unless somebody bothers to demonstrate it
in a remotely realistic environment.
Even if this would be real if they sniff a openssl key, unless they also
crack the dns the browser will complain (not very different from not
having a certificate authority signature on a fake key). And if the
server is remotely serious they'll notice the 100% runaway process way
before he can sniff the whole key (the 100% runaway load cannot be
hidden). Most servers have some statistics so a 100% load for weeks or
months isn't very likely to be overlooked.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 15:33 ` Andrea Arcangeli
@ 2005-05-15 1:07 ` Christer Weinigel
2005-05-15 9:48 ` Andi Kleen
1 sibling, 0 replies; 84+ messages in thread
From: Christer Weinigel @ 2005-05-15 1:07 UTC (permalink / raw)
To: Andrea Arcangeli
Cc: Lee Revell, Dave Jones, Alan Cox, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
Andrea Arcangeli <andrea@suse.de> writes:
> Nobody runs openssl -sign thousand of times in a row on a pure idle
> system without noticing the 100% load on the other cpu for months
Well, actually one does. On a normal https server, each https request
results in an operation on the private key. So if the attacker shares
the same web server as the victim it's probably rather easy for the
attacker to see when the machine is idle and launch an attack giving
him thousands of chances to spy on the victim.
But I do agree that this probably isn't all that serious, for those
who really have secrets to hide, they won't run their https server on
a machine shared with anybody else.
/Christer
--
"Just how much can I get away with and still go to heaven?"
Freelance consultant specializing in device driver programming for Linux
Christer Weinigel <christer@weinigel.se> http://www.weinigel.se
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 15:33 ` Andrea Arcangeli
2005-05-15 1:07 ` Christer Weinigel
@ 2005-05-15 9:48 ` Andi Kleen
1 sibling, 0 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 9:48 UTC (permalink / raw)
To: Andrea Arcangeli
Cc: Lee Revell, Dave Jones, Alan Cox, Matt Mackall, Andy Isaacson,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, May 14, 2005 at 05:33:07PM +0200, Andrea Arcangeli wrote:
> On Sat, May 14, 2005 at 03:37:18AM -0400, Lee Revell wrote:
> > The apps that bother to use rdtsc vs. gettimeofday need a cheap high res
> > timer more than a correct one anyway - it's not guaranteed that rdtsc
> > provides a reliable time source at all, due to SMP and frequency scaling
> > issues.
>
> On x86-64 the cost of gettimeofday is the same of the tsc, turning off
It depends, on many systems it is more costly. e.g. on many SMP
systems we have to use HPET or even the PM timer, because TSC is not
reliable.
> tsc on x86-64 is not nice (even if we usually have HPET there, so
> perhaps it wouldn't be too bad). TSC is something only the kernel (or a
> person with some kernel/hardware knowledge) can do safely knowing it'll
> work fine. But on x86-64 parts of the kernel runs in userland...
Agreed. It is quite complicated to decide if TSC is reliable or not
and I would not recommend user space to do this.
[hmm actually I already have constant_tsc fake cpuid bit, but
it only refers to single CPUs. I wonder if I should add another
one for SMP "synchronized_tsc". The latest mm code already has
this information, but it does not export it yet]
>
> Preventing tasks with different uid to run on the same physical cpu was
> my first idea, disabled by default via sysctl, so only if one is
> paranoid can enable it.
The paranoid should just fix their crypto code. And if they're
clinically paranoid they can always boot with noht or disable
it in the BIOS. But really I think they should just fix OpenSSL.
>
> But before touching the kernel in any way it would be really nice if
> somebody could bother to demonstrate this is real because I've an hard
> time to believe this is not purely vapourware. On artificial
Similar feeling here.
> Nobody runs openssl -sign thousand of times in a row on a pure idle
> system without noticing the 100% load on the other cpu for months (and
> he's not root so he can't hide his runaway 100% process, if he was root
> and he could modify the kernel or ps/top to hide the runaway process,
> he'd have faster ways to sniff).
Exactly.
>
> So to me this sounds a purerly theoretical problem. Cache covert
Perhaps not purely theoretical, but it is certainly not something
that needs drastic action like disabling HT in general.
> This was an interesting read, but in practice I'd rate this to have
> severity 1 on a 0-100 scale, unless somebody bothers to demonstrate it
> in a remotely realistic environment.
Agreed.
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 23:38 ` Lee Revell
2005-05-13 23:44 ` Dave Jones
@ 2005-05-14 15:23 ` Alan Cox
2005-05-14 15:45 ` andrea
2005-05-14 16:30 ` Lee Revell
1 sibling, 2 replies; 84+ messages in thread
From: Alan Cox @ 2005-05-14 15:23 UTC (permalink / raw)
To: Lee Revell
Cc: Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sad, 2005-05-14 at 00:38, Lee Revell wrote:
> Well yes but you would still have to recompile those apps. And take the
> big performance hit from using gettimeofday vs rdtsc. Disabling HT by
> default looks pretty good by comparison.
You cannot use rdtsc for anything but rough instruction timing. The
timers for different processors run at different speeds on some SMP
systems, the timer rates vary as processors change clock rate nowdays.
Rdtsc may also jump dramatically on a suspend/resume.
If the app uses rdtsc then generally speaking its terminally broken. The
only exception is some profiling tools.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 15:23 ` Alan Cox
@ 2005-05-14 15:45 ` andrea
2005-05-15 13:38 ` Mikulas Patocka
2005-05-14 16:30 ` Lee Revell
1 sibling, 1 reply; 84+ messages in thread
From: andrea @ 2005-05-14 15:45 UTC (permalink / raw)
To: Alan Cox
Cc: Lee Revell, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso,
Andrew Morton
On Sat, May 14, 2005 at 04:23:10PM +0100, Alan Cox wrote:
> You cannot use rdtsc for anything but rough instruction timing. The
> timers for different processors run at different speeds on some SMP
> systems, the timer rates vary as processors change clock rate nowdays.
> Rdtsc may also jump dramatically on a suspend/resume.
x86-64 uses it for vgettimeofday very safely (i386 could do too but it
doesn't).
Anyway I believe at least for seccomp it's worth to turn off the tsc,
not just for HT but for the L2 cache too. So it's up to you, either you
turn it off completely (which isn't very nice IMHO) or I recommend to
apply this below patch. This has been tested successfully on x86-64
against current cogito repository (i686 compiles so I didn't bother
testing ;). People selling the cpu through cpushare may appreciate this
bit for a peace of mind. There's no way to get any timing info anymore
with this applied (gettimeofday is forbidden of course). The seccomp
environment is completely deterministic so it can't be allowed to get
timing info, it has to be deterministic so in the future I can enable a
computing mode that does a parallel computing for each task with server
side transparent checkpointing and verification that the output is the
same from all the 2/3 seller computers for each task, without the buyer
even noticing (for now the verification is left to the buyer client
side and there's no checkpointing, since that would require more kernel
changes to track the dirty bits but it'll be easy to extend once the
basic mode is finished).
Thanks.
Signed-off-by: Andrea Arcangeli <andrea@cpushare.com>
Index: arch/i386/kernel/process.c
===================================================================
--- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/arch/i386/kernel/process.c (mode:100644)
+++ uncommitted/arch/i386/kernel/process.c (mode:100644)
@@ -561,6 +561,25 @@
}
/*
+ * This function selects if the context switch from prev to next
+ * has to tweak the TSC disable bit in the cr4.
+ */
+static void disable_tsc(struct thread_info *prev,
+ struct thread_info *next)
+{
+ if (unlikely(has_secure_computing(prev) ||
+ has_secure_computing(next))) {
+ /* slow path here */
+ if (has_secure_computing(prev) &&
+ !has_secure_computing(next)) {
+ clear_in_cr4(X86_CR4_TSD);
+ } else if (!has_secure_computing(prev) &&
+ has_secure_computing(next))
+ set_in_cr4(X86_CR4_TSD);
+ }
+}
+
+/*
* switch_to(x,yn) should switch tasks from x to y.
*
* We fsave/fwait so that an exception goes off at the right time
@@ -639,6 +658,8 @@
if (unlikely(prev->io_bitmap_ptr || next->io_bitmap_ptr))
handle_io_bitmap(next, tss);
+ disable_tsc(prev_p->thread_info, next_p->thread_info);
+
return prev_p;
}
Index: arch/x86_64/kernel/process.c
===================================================================
--- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/arch/x86_64/kernel/process.c (mode:100644)
+++ uncommitted/arch/x86_64/kernel/process.c (mode:100644)
@@ -439,6 +439,25 @@
}
/*
+ * This function selects if the context switch from prev to next
+ * has to tweak the TSC disable bit in the cr4.
+ */
+static void disable_tsc(struct thread_info *prev,
+ struct thread_info *next)
+{
+ if (unlikely(has_secure_computing(prev) ||
+ has_secure_computing(next))) {
+ /* slow path here */
+ if (has_secure_computing(prev) &&
+ !has_secure_computing(next)) {
+ clear_in_cr4(X86_CR4_TSD);
+ } else if (!has_secure_computing(prev) &&
+ has_secure_computing(next))
+ set_in_cr4(X86_CR4_TSD);
+ }
+}
+
+/*
* This special macro can be used to load a debugging register
*/
#define loaddebug(thread,r) set_debug(thread->debugreg ## r, r)
@@ -556,6 +575,8 @@
}
}
+ disable_tsc(prev_p->thread_info, next_p->thread_info);
+
return prev_p;
}
Index: include/linux/seccomp.h
===================================================================
--- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/include/linux/seccomp.h (mode:100644)
+++ uncommitted/include/linux/seccomp.h (mode:100644)
@@ -19,6 +19,11 @@
__secure_computing(this_syscall);
}
+static inline int has_secure_computing(struct thread_info *ti)
+{
+ return unlikely(test_ti_thread_flag(ti, TIF_SECCOMP));
+}
+
#else /* CONFIG_SECCOMP */
#if (__GNUC__ > 2)
@@ -28,6 +33,7 @@
#endif
#define secure_computing(x) do { } while (0)
+#define has_secure_computing(x) 0
#endif /* CONFIG_SECCOMP */
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-14 15:45 ` andrea
@ 2005-05-15 13:38 ` Mikulas Patocka
2005-05-16 7:06 ` andrea
0 siblings, 1 reply; 84+ messages in thread
From: Mikulas Patocka @ 2005-05-15 13:38 UTC (permalink / raw)
To: andrea
Cc: Alan Cox, Lee Revell, Dave Jones, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso, Andrew Morton
On Sat, 14 May 2005 andrea@cpushare.com wrote:
> On Sat, May 14, 2005 at 04:23:10PM +0100, Alan Cox wrote:
> > You cannot use rdtsc for anything but rough instruction timing. The
> > timers for different processors run at different speeds on some SMP
> > systems, the timer rates vary as processors change clock rate nowdays.
> > Rdtsc may also jump dramatically on a suspend/resume.
>
> x86-64 uses it for vgettimeofday very safely (i386 could do too but it
> doesn't).
>
> Anyway I believe at least for seccomp it's worth to turn off the tsc,
> not just for HT but for the L2 cache too. So it's up to you, either you
> turn it off completely (which isn't very nice IMHO) or I recommend to
> apply this below patch. This has been tested successfully on x86-64
> against current cogito repository (i686 compiles so I didn't bother
> testing ;). People selling the cpu through cpushare may appreciate this
> bit for a peace of mind. There's no way to get any timing info anymore
> with this applied (gettimeofday is forbidden of course).
Another possibility to get timing is from direct-io --- i.e. initiate
direct io read, wait until one cache line contains new data and you can be
sure that the next will contain new data in certain time. IDE controller
bus master operation acts here as a timer.
Mikulas
> The seccomp
> environment is completely deterministic so it can't be allowed to get
> timing info, it has to be deterministic so in the future I can enable a
> computing mode that does a parallel computing for each task with server
> side transparent checkpointing and verification that the output is the
> same from all the 2/3 seller computers for each task, without the buyer
> even noticing (for now the verification is left to the buyer client
> side and there's no checkpointing, since that would require more kernel
> changes to track the dirty bits but it'll be easy to extend once the
> basic mode is finished).
>
> Thanks.
>
> Signed-off-by: Andrea Arcangeli <andrea@cpushare.com>
>
> Index: arch/i386/kernel/process.c
> ===================================================================
> --- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/arch/i386/kernel/process.c (mode:100644)
> +++ uncommitted/arch/i386/kernel/process.c (mode:100644)
> @@ -561,6 +561,25 @@
> }
>
> /*
> + * This function selects if the context switch from prev to next
> + * has to tweak the TSC disable bit in the cr4.
> + */
> +static void disable_tsc(struct thread_info *prev,
> + struct thread_info *next)
> +{
> + if (unlikely(has_secure_computing(prev) ||
> + has_secure_computing(next))) {
> + /* slow path here */
> + if (has_secure_computing(prev) &&
> + !has_secure_computing(next)) {
> + clear_in_cr4(X86_CR4_TSD);
> + } else if (!has_secure_computing(prev) &&
> + has_secure_computing(next))
> + set_in_cr4(X86_CR4_TSD);
> + }
> +}
> +
> +/*
> * switch_to(x,yn) should switch tasks from x to y.
> *
> * We fsave/fwait so that an exception goes off at the right time
> @@ -639,6 +658,8 @@
> if (unlikely(prev->io_bitmap_ptr || next->io_bitmap_ptr))
> handle_io_bitmap(next, tss);
>
> + disable_tsc(prev_p->thread_info, next_p->thread_info);
> +
> return prev_p;
> }
>
> Index: arch/x86_64/kernel/process.c
> ===================================================================
> --- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/arch/x86_64/kernel/process.c (mode:100644)
> +++ uncommitted/arch/x86_64/kernel/process.c (mode:100644)
> @@ -439,6 +439,25 @@
> }
>
> /*
> + * This function selects if the context switch from prev to next
> + * has to tweak the TSC disable bit in the cr4.
> + */
> +static void disable_tsc(struct thread_info *prev,
> + struct thread_info *next)
> +{
> + if (unlikely(has_secure_computing(prev) ||
> + has_secure_computing(next))) {
> + /* slow path here */
> + if (has_secure_computing(prev) &&
> + !has_secure_computing(next)) {
> + clear_in_cr4(X86_CR4_TSD);
> + } else if (!has_secure_computing(prev) &&
> + has_secure_computing(next))
> + set_in_cr4(X86_CR4_TSD);
> + }
> +}
> +
> +/*
> * This special macro can be used to load a debugging register
> */
> #define loaddebug(thread,r) set_debug(thread->debugreg ## r, r)
> @@ -556,6 +575,8 @@
> }
> }
>
> + disable_tsc(prev_p->thread_info, next_p->thread_info);
> +
> return prev_p;
> }
>
> Index: include/linux/seccomp.h
> ===================================================================
> --- eed337ef5e9ae7d62caa84b7974a11fddc7f06e0/include/linux/seccomp.h (mode:100644)
> +++ uncommitted/include/linux/seccomp.h (mode:100644)
> @@ -19,6 +19,11 @@
> __secure_computing(this_syscall);
> }
>
> +static inline int has_secure_computing(struct thread_info *ti)
> +{
> + return unlikely(test_ti_thread_flag(ti, TIF_SECCOMP));
> +}
> +
> #else /* CONFIG_SECCOMP */
>
> #if (__GNUC__ > 2)
> @@ -28,6 +33,7 @@
> #endif
>
> #define secure_computing(x) do { } while (0)
> +#define has_secure_computing(x) 0
>
> #endif /* CONFIG_SECCOMP */
>
> -
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at http://www.tux.org/lkml/
>
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-15 13:38 ` Mikulas Patocka
@ 2005-05-16 7:06 ` andrea
0 siblings, 0 replies; 84+ messages in thread
From: andrea @ 2005-05-16 7:06 UTC (permalink / raw)
To: Mikulas Patocka
Cc: Alan Cox, Lee Revell, Dave Jones, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso, Andrew Morton
On Sun, May 15, 2005 at 03:38:22PM +0200, Mikulas Patocka wrote:
> Another possibility to get timing is from direct-io --- i.e. initiate
> direct io read, wait until one cache line contains new data and you can be
> sure that the next will contain new data in certain time. IDE controller
> bus master operation acts here as a timer.
There's no way to do direct-io through seccomp, all the fds are pipes
with twisted userland listening the other side of the pipe. So disabling
the tsc is more than enough to give to CPUShare users a peace of mind
with HT enabled and without having to flush the l2 cache either.
CPUShare is the only case I can imagine where an untrusted and random
bytecode running at 100% system load is the normal behaviour.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 15:23 ` Alan Cox
2005-05-14 15:45 ` andrea
@ 2005-05-14 16:30 ` Lee Revell
2005-05-14 16:44 ` Arjan van de Ven
` (2 more replies)
1 sibling, 3 replies; 84+ messages in thread
From: Lee Revell @ 2005-05-14 16:30 UTC (permalink / raw)
To: Alan Cox
Cc: Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 16:23 +0100, Alan Cox wrote:
> On Sad, 2005-05-14 at 00:38, Lee Revell wrote:
> > Well yes but you would still have to recompile those apps. And take the
> > big performance hit from using gettimeofday vs rdtsc. Disabling HT by
> > default looks pretty good by comparison.
>
> You cannot use rdtsc for anything but rough instruction timing. The
> timers for different processors run at different speeds on some SMP
> systems, the timer rates vary as processors change clock rate nowdays.
> Rdtsc may also jump dramatically on a suspend/resume.
>
> If the app uses rdtsc then generally speaking its terminally broken. The
> only exception is some profiling tools.
That is basically all JACK and mplayer use it for. They have RT
constraints and the tsc is used to know if we got woken up too late and
should just drop some frames. The developers are aware of the issues
with rdtsc and have chosen to use it anyway because these apps need
every ounce of CPU and cannot tolerate the overhead of gettimeofday().
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-14 16:30 ` Lee Revell
@ 2005-05-14 16:44 ` Arjan van de Ven
2005-05-14 17:56 ` Lee Revell
2005-05-14 17:04 ` Jindrich Makovicka
2005-05-15 9:58 ` Andi Kleen
2 siblings, 1 reply; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-14 16:44 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 12:30 -0400, Lee Revell wrote:
> On Sat, 2005-05-14 at 16:23 +0100, Alan Cox wrote:
> > On Sad, 2005-05-14 at 00:38, Lee Revell wrote:
> > > Well yes but you would still have to recompile those apps. And take the
> > > big performance hit from using gettimeofday vs rdtsc. Disabling HT by
> > > default looks pretty good by comparison.
> >
> > You cannot use rdtsc for anything but rough instruction timing. The
> > timers for different processors run at different speeds on some SMP
> > systems, the timer rates vary as processors change clock rate nowdays.
> > Rdtsc may also jump dramatically on a suspend/resume.
> >
> > If the app uses rdtsc then generally speaking its terminally broken. The
> > only exception is some profiling tools.
>
> That is basically all JACK and mplayer use it for. They have RT
> constraints and the tsc is used to know if we got woken up too late and
> should just drop some frames. The developers are aware of the issues
> with rdtsc and have chosen to use it anyway because these apps need
> every ounce of CPU and cannot tolerate the overhead of gettimeofday().
then JACK is terminally broken if it doesn't have a fallback for non-
rdtsc cpus.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 16:44 ` Arjan van de Ven
@ 2005-05-14 17:56 ` Lee Revell
2005-05-14 18:01 ` Arjan van de Ven
2005-05-15 9:33 ` Adrian Bunk
0 siblings, 2 replies; 84+ messages in thread
From: Lee Revell @ 2005-05-14 17:56 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> then JACK is terminally broken if it doesn't have a fallback for non-
> rdtsc cpus.
It does have a fallback, but the selection is done at compile time. It
uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
Maybe we should check at runtime, but this has always worked.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 17:56 ` Lee Revell
@ 2005-05-14 18:01 ` Arjan van de Ven
2005-05-14 19:21 ` Lee Revell
2005-05-15 10:01 ` Andi Kleen
2005-05-15 9:33 ` Adrian Bunk
1 sibling, 2 replies; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-14 18:01 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 13:56 -0400, Lee Revell wrote:
> On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > then JACK is terminally broken if it doesn't have a fallback for non-
> > rdtsc cpus.
>
> It does have a fallback, but the selection is done at compile time. It
> uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
>
> Maybe we should check at runtime,
it's probably a sign that JACK isn't used on SMP systems much, at least
not on the bigger systems (like IBM's x440's) where the tsc *will*
differ wildly between cpus...
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 18:01 ` Arjan van de Ven
@ 2005-05-14 19:21 ` Lee Revell
2005-05-14 19:48 ` Arjan van de Ven
2005-05-15 10:01 ` Andi Kleen
1 sibling, 1 reply; 84+ messages in thread
From: Lee Revell @ 2005-05-14 19:21 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 20:01 +0200, Arjan van de Ven wrote:
> On Sat, 2005-05-14 at 13:56 -0400, Lee Revell wrote:
> > On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > > then JACK is terminally broken if it doesn't have a fallback for non-
> > > rdtsc cpus.
> >
> > It does have a fallback, but the selection is done at compile time. It
> > uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
> >
> > Maybe we should check at runtime,
>
> it's probably a sign that JACK isn't used on SMP systems much, at least
> not on the bigger systems (like IBM's x440's) where the tsc *will*
> differ wildly between cpus...
Correct. The only bug reports we have seen related to the use of the
TSC is due to CPU frequency scaling. The fix is to not use it - people
who want to use their PC as a DSP for audio probably don't want their
processor slowing down anyway. And JACK is targeted at desktop and
smaller systems, it would be kind of crazy to run it on a big iron.
Well, maybe there are people who like to record sessions or practice
guitar in the server room...
If gettimeofday is really as cheap as rdtsc on x86_64, we should use it.
But it's too expensive for slower x86 systems. Anyway, Andi's fix
disables *all* high res timing including gettimeofday. Obviously no
multimedia app can tolerate this, so discussing rdtsc is really a red
herring. But multimedia apps aren't much in seccomp environments
either.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 19:21 ` Lee Revell
@ 2005-05-14 19:48 ` Arjan van de Ven
2005-05-14 23:40 ` Lee Revell
2005-05-15 3:19 ` dean gaudet
0 siblings, 2 replies; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-14 19:48 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 15:21 -0400, Lee Revell wrote:
> On Sat, 2005-05-14 at 20:01 +0200, Arjan van de Ven wrote:
> > On Sat, 2005-05-14 at 13:56 -0400, Lee Revell wrote:
> > > On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > > > then JACK is terminally broken if it doesn't have a fallback for non-
> > > > rdtsc cpus.
> > >
> > > It does have a fallback, but the selection is done at compile time. It
> > > uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
> > >
> > > Maybe we should check at runtime,
> >
> > it's probably a sign that JACK isn't used on SMP systems much, at least
> > not on the bigger systems (like IBM's x440's) where the tsc *will*
> > differ wildly between cpus...
>
> Correct. The only bug reports we have seen related to the use of the
> TSC is due to CPU frequency scaling. The fix is to not use it - people
> who want to use their PC as a DSP for audio probably don't want their
> processor slowing down anyway.
it's a matter of time (my estimate is a year or two) before processors
get variable frequencies based on temperature targets etc...
and then rdtsc is really useless for this kind of thing..
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 19:48 ` Arjan van de Ven
@ 2005-05-14 23:40 ` Lee Revell
2005-05-15 7:30 ` Arjan van de Ven
2005-05-15 9:37 ` Andi Kleen
2005-05-15 3:19 ` dean gaudet
1 sibling, 2 replies; 84+ messages in thread
From: Lee Revell @ 2005-05-14 23:40 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 21:48 +0200, Arjan van de Ven wrote:
> On Sat, 2005-05-14 at 15:21 -0400, Lee Revell wrote:
> > On Sat, 2005-05-14 at 20:01 +0200, Arjan van de Ven wrote:
> > > On Sat, 2005-05-14 at 13:56 -0400, Lee Revell wrote:
> > > > On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > > > > then JACK is terminally broken if it doesn't have a fallback for non-
> > > > > rdtsc cpus.
> > > >
> > > > It does have a fallback, but the selection is done at compile time. It
> > > > uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
> > > >
> > > > Maybe we should check at runtime,
> > >
> > > it's probably a sign that JACK isn't used on SMP systems much, at least
> > > not on the bigger systems (like IBM's x440's) where the tsc *will*
> > > differ wildly between cpus...
> >
> > Correct. The only bug reports we have seen related to the use of the
> > TSC is due to CPU frequency scaling. The fix is to not use it - people
> > who want to use their PC as a DSP for audio probably don't want their
> > processor slowing down anyway.
>
> it's a matter of time (my estimate is a year or two) before processors
> get variable frequencies based on temperature targets etc...
> and then rdtsc is really useless for this kind of thing..
I was under the impression that P4 and later processors do not vary the
TSC rate when doing frequency scaling. This is mentioned in the
documentation for the high res timers patch.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 23:40 ` Lee Revell
@ 2005-05-15 7:30 ` Arjan van de Ven
2005-05-15 20:41 ` Alan Cox
2005-05-15 9:37 ` Andi Kleen
1 sibling, 1 reply; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-15 7:30 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, 2005-05-14 at 19:40 -0400, Lee Revell wrote:
> > it's a matter of time (my estimate is a year or two) before processors
> > get variable frequencies based on temperature targets etc...
> > and then rdtsc is really useless for this kind of thing..
>
> I was under the impression that P4 and later processors do not vary the
> TSC rate when doing frequency scaling. This is mentioned in the
> documentation for the high res timers patch.
seems not the case, and worse, during idle time the clock is allowed to
stop entirely.... (and that is also happening more and more and linux is
getting more agressive idle support (eg no timer tick and such patches)
which will trigger bios thresholds for this even more too.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 7:30 ` Arjan van de Ven
@ 2005-05-15 20:41 ` Alan Cox
2005-05-15 20:48 ` Arjan van de Ven
0 siblings, 1 reply; 84+ messages in thread
From: Alan Cox @ 2005-05-15 20:41 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Lee Revell, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> stop entirely.... (and that is also happening more and more and linux is
> getting more agressive idle support (eg no timer tick and such patches)
> which will trigger bios thresholds for this even more too.
Cyrix did TSC stop on halt a long long time ago, back when it was worth
the power difference.
Alan
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 20:41 ` Alan Cox
@ 2005-05-15 20:48 ` Arjan van de Ven
2005-05-15 21:10 ` Lee Revell
0 siblings, 1 reply; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-15 20:48 UTC (permalink / raw)
To: Alan Cox
Cc: Lee Revell, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sun, 2005-05-15 at 21:41 +0100, Alan Cox wrote:
> On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> > stop entirely.... (and that is also happening more and more and linux is
> > getting more agressive idle support (eg no timer tick and such patches)
> > which will trigger bios thresholds for this even more too.
>
> Cyrix did TSC stop on halt a long long time ago, back when it was worth
> the power difference.
With linux going to ACPI C2 mode more... tsc is defined to halt in C2...
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 20:48 ` Arjan van de Ven
@ 2005-05-15 21:10 ` Lee Revell
2005-05-15 22:55 ` Dave Jones
0 siblings, 1 reply; 84+ messages in thread
From: Lee Revell @ 2005-05-15 21:10 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sun, 2005-05-15 at 22:48 +0200, Arjan van de Ven wrote:
> On Sun, 2005-05-15 at 21:41 +0100, Alan Cox wrote:
> > On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> > > stop entirely.... (and that is also happening more and more and linux is
> > > getting more agressive idle support (eg no timer tick and such patches)
> > > which will trigger bios thresholds for this even more too.
> >
> > Cyrix did TSC stop on halt a long long time ago, back when it was worth
> > the power difference.
>
> With linux going to ACPI C2 mode more... tsc is defined to halt in C2...
JACK doesn't care about any of this now, the behavior when you
suspend/resume with a running jackd is undefined. Eventually we should
handle it, but there's no point until the ALSA drivers get proper
suspend/resume support.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 21:10 ` Lee Revell
@ 2005-05-15 22:55 ` Dave Jones
2005-05-15 23:10 ` Lee Revell
0 siblings, 1 reply; 84+ messages in thread
From: Dave Jones @ 2005-05-15 22:55 UTC (permalink / raw)
To: Lee Revell
Cc: Arjan van de Ven, Alan Cox, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
On Sun, May 15, 2005 at 05:10:59PM -0400, Lee Revell wrote:
> On Sun, 2005-05-15 at 22:48 +0200, Arjan van de Ven wrote:
> > On Sun, 2005-05-15 at 21:41 +0100, Alan Cox wrote:
> > > On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> > > > stop entirely.... (and that is also happening more and more and linux is
> > > > getting more agressive idle support (eg no timer tick and such patches)
> > > > which will trigger bios thresholds for this even more too.
> > >
> > > Cyrix did TSC stop on halt a long long time ago, back when it was worth
> > > the power difference.
> >
> > With linux going to ACPI C2 mode more... tsc is defined to halt in C2...
>
> JACK doesn't care about any of this now, the behavior when you
> suspend/resume with a running jackd is undefined. Eventually we should
> handle it, but there's no point until the ALSA drivers get proper
> suspend/resume support.
suspend/resume are S states, not C states. C states are occuring
during runtime.
Dave
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 22:55 ` Dave Jones
@ 2005-05-15 23:10 ` Lee Revell
2005-05-16 7:25 ` Arjan van de Ven
0 siblings, 1 reply; 84+ messages in thread
From: Lee Revell @ 2005-05-15 23:10 UTC (permalink / raw)
To: Dave Jones
Cc: Arjan van de Ven, Alan Cox, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
On Sun, 2005-05-15 at 18:55 -0400, Dave Jones wrote:
> On Sun, May 15, 2005 at 05:10:59PM -0400, Lee Revell wrote:
> > On Sun, 2005-05-15 at 22:48 +0200, Arjan van de Ven wrote:
> > > On Sun, 2005-05-15 at 21:41 +0100, Alan Cox wrote:
> > > > On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> > > > > stop entirely.... (and that is also happening more and more and linux is
> > > > > getting more agressive idle support (eg no timer tick and such patches)
> > > > > which will trigger bios thresholds for this even more too.
> > > >
> > > > Cyrix did TSC stop on halt a long long time ago, back when it was worth
> > > > the power difference.
> > >
> > > With linux going to ACPI C2 mode more... tsc is defined to halt in C2...
> >
> > JACK doesn't care about any of this now, the behavior when you
> > suspend/resume with a running jackd is undefined. Eventually we should
> > handle it, but there's no point until the ALSA drivers get proper
> > suspend/resume support.
>
> suspend/resume are S states, not C states. C states are occuring
> during runtime.
It should never go into C2 if jackd is running, because you're getting
interrupts from the audio interface at least every 100ms or so (usually
much more often) which will wake up jackd and any clients.
Lee
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 23:10 ` Lee Revell
@ 2005-05-16 7:25 ` Arjan van de Ven
0 siblings, 0 replies; 84+ messages in thread
From: Arjan van de Ven @ 2005-05-16 7:25 UTC (permalink / raw)
To: Lee Revell
Cc: Dave Jones, Alan Cox, Matt Mackall, Andy Isaacson, Andi Kleen,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sun, 2005-05-15 at 19:10 -0400, Lee Revell wrote:
> On Sun, 2005-05-15 at 18:55 -0400, Dave Jones wrote:
> > On Sun, May 15, 2005 at 05:10:59PM -0400, Lee Revell wrote:
> > > On Sun, 2005-05-15 at 22:48 +0200, Arjan van de Ven wrote:
> > > > On Sun, 2005-05-15 at 21:41 +0100, Alan Cox wrote:
> > > > > On Sul, 2005-05-15 at 08:30, Arjan van de Ven wrote:
> > > > > > stop entirely.... (and that is also happening more and more and linux is
> > > > > > getting more agressive idle support (eg no timer tick and such patches)
> > > > > > which will trigger bios thresholds for this even more too.
> > > > >
> > > > > Cyrix did TSC stop on halt a long long time ago, back when it was worth
> > > > > the power difference.
> > > >
> > > > With linux going to ACPI C2 mode more... tsc is defined to halt in C2...
> > >
> > > JACK doesn't care about any of this now, the behavior when you
> > > suspend/resume with a running jackd is undefined. Eventually we should
> > > handle it, but there's no point until the ALSA drivers get proper
> > > suspend/resume support.
> >
> > suspend/resume are S states, not C states. C states are occuring
> > during runtime.
>
> It should never go into C2 if jackd is running, because you're getting
> interrupts from the audio interface at least every 100ms or so (usually
> much more often) which will wake up jackd and any clients.
you're not guaranteed to not enter C2 in that case. C2 can happen after
a few ms already
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 23:40 ` Lee Revell
2005-05-15 7:30 ` Arjan van de Ven
@ 2005-05-15 9:37 ` Andi Kleen
1 sibling, 0 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 9:37 UTC (permalink / raw)
To: Lee Revell
Cc: Arjan van de Ven, Alan Cox, Dave Jones, Matt Mackall,
Andy Isaacson, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
> I was under the impression that P4 and later processors do not vary the
> TSC rate when doing frequency scaling. This is mentioned in the
> documentation for the high res timers patch.
Prescott and later do not vary TSC, but P4s before that do.
On x86-64 it is true because only Nocona is supported which has a
pstate invariant TSC.
The latest x86-64 kernel has a special X86_CONSTANT_TSC internal
CPUID bit, which is set in that case. If some other subsystem
uses it I would recommend to port that to i386 too.
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 19:48 ` Arjan van de Ven
2005-05-14 23:40 ` Lee Revell
@ 2005-05-15 3:19 ` dean gaudet
1 sibling, 0 replies; 84+ messages in thread
From: dean gaudet @ 2005-05-15 3:19 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Lee Revell, Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson,
Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
On Sat, 14 May 2005, Arjan van de Ven wrote:
> it's a matter of time (my estimate is a year or two) before processors
> get variable frequencies based on temperature targets etc...
> and then rdtsc is really useless for this kind of thing..
what do you mean "a year or two"? processors have been doing this for
many years now.
i'm biased, but i still think transmeta did this the right way... the tsc
operates at the top frequency of the processor always.
i do a hell of a lot of microbenchmarking on various processors and i
always use tsc -- but i'm just smart enough to take multiple samples and i
try to make each sample smaller than a time slice... which avoids most of
the pitfalls, and would even work on smp boxes with tsc differences.
-dean
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 18:01 ` Arjan van de Ven
2005-05-14 19:21 ` Lee Revell
@ 2005-05-15 10:01 ` Andi Kleen
1 sibling, 0 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 10:01 UTC (permalink / raw)
To: Arjan van de Ven
Cc: Lee Revell, Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, May 14, 2005 at 08:01:33PM +0200, Arjan van de Ven wrote:
> On Sat, 2005-05-14 at 13:56 -0400, Lee Revell wrote:
> > On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > > then JACK is terminally broken if it doesn't have a fallback for non-
> > > rdtsc cpus.
> >
> > It does have a fallback, but the selection is done at compile time. It
> > uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
> >
> > Maybe we should check at runtime,
>
> it's probably a sign that JACK isn't used on SMP systems much, at least
> not on the bigger systems (like IBM's x440's) where the tsc *will*
> differ wildly between cpus...
It does not even need SMP, just use a Centrino laptop.
I suppose what the Jack guys are doing is to recommend to disable
frequency scaling then the sound guys complain again
that sound on linux is so hard to use. I wonder where this comes from? :)
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 17:56 ` Lee Revell
2005-05-14 18:01 ` Arjan van de Ven
@ 2005-05-15 9:33 ` Adrian Bunk
1 sibling, 0 replies; 84+ messages in thread
From: Adrian Bunk @ 2005-05-15 9:33 UTC (permalink / raw)
To: Lee Revell
Cc: Arjan van de Ven, Alan Cox, Dave Jones, Matt Mackall,
Andy Isaacson, Andi Kleen, Richard F. Rebel, Gabor MICSKO,
Linux Kernel Mailing List, tytso
On Sat, May 14, 2005 at 01:56:36PM -0400, Lee Revell wrote:
> On Sat, 2005-05-14 at 18:44 +0200, Arjan van de Ven wrote:
> > then JACK is terminally broken if it doesn't have a fallback for non-
> > rdtsc cpus.
>
> It does have a fallback, but the selection is done at compile time. It
> uses rdtsc for all x86 CPUs except pre-i586 SMP systems.
>
> Maybe we should check at runtime, but this has always worked.
If this is critical for JACK, runtime selection was an improvement for
distributions like Debian that support both pre-i586 SMP systems and
current hardware.
> Lee
cu
Adrian
--
"Is there not promise of rain?" Ling Tan asked suddenly out
of the darkness. There had been need of rain for many days.
"Only a promise," Lao Er said.
Pearl S. Buck - Dragon Seed
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 16:30 ` Lee Revell
2005-05-14 16:44 ` Arjan van de Ven
@ 2005-05-14 17:04 ` Jindrich Makovicka
2005-05-14 18:27 ` Lee Revell
2005-05-15 9:58 ` Andi Kleen
2 siblings, 1 reply; 84+ messages in thread
From: Jindrich Makovicka @ 2005-05-14 17:04 UTC (permalink / raw)
To: linux-kernel
Lee Revell wrote:
> On Sat, 2005-05-14 at 16:23 +0100, Alan Cox wrote:
>
>>On Sad, 2005-05-14 at 00:38, Lee Revell wrote:
>>
>>>Well yes but you would still have to recompile those apps. And take the
>>>big performance hit from using gettimeofday vs rdtsc. Disabling HT by
>>>default looks pretty good by comparison.
>>
>>You cannot use rdtsc for anything but rough instruction timing. The
>>timers for different processors run at different speeds on some SMP
>>systems, the timer rates vary as processors change clock rate nowdays.
>>Rdtsc may also jump dramatically on a suspend/resume.
>>
>>If the app uses rdtsc then generally speaking its terminally broken. The
>>only exception is some profiling tools.
>
>
> That is basically all JACK and mplayer use it for. They have RT
> constraints and the tsc is used to know if we got woken up too late and
> should just drop some frames. The developers are aware of the issues
> with rdtsc and have chosen to use it anyway because these apps need
> every ounce of CPU and cannot tolerate the overhead of gettimeofday().
AFAIK, mplayer actually uses gettimeofday(). rdtsc is used in some
places for profiling and debugging purposes and not compiled in by default.
--
Jindrich Makovicka
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-14 16:30 ` Lee Revell
2005-05-14 16:44 ` Arjan van de Ven
2005-05-14 17:04 ` Jindrich Makovicka
@ 2005-05-15 9:58 ` Andi Kleen
2 siblings, 0 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 9:58 UTC (permalink / raw)
To: Lee Revell
Cc: Alan Cox, Dave Jones, Matt Mackall, Andy Isaacson,
Richard F. Rebel, Gabor MICSKO, Linux Kernel Mailing List, tytso
On Sat, May 14, 2005 at 12:30:28PM -0400, Lee Revell wrote:
> On Sat, 2005-05-14 at 16:23 +0100, Alan Cox wrote:
> > On Sad, 2005-05-14 at 00:38, Lee Revell wrote:
> > > Well yes but you would still have to recompile those apps. And take the
> > > big performance hit from using gettimeofday vs rdtsc. Disabling HT by
> > > default looks pretty good by comparison.
> >
> > You cannot use rdtsc for anything but rough instruction timing. The
> > timers for different processors run at different speeds on some SMP
> > systems, the timer rates vary as processors change clock rate nowdays.
> > Rdtsc may also jump dramatically on a suspend/resume.
> >
> > If the app uses rdtsc then generally speaking its terminally broken. The
> > only exception is some profiling tools.
>
> That is basically all JACK and mplayer use it for. They have RT
> constraints and the tsc is used to know if we got woken up too late and
> should just drop some frames. The developers are aware of the issues
> with rdtsc and have chosen to use it anyway because these apps need
> every ounce of CPU and cannot tolerate the overhead of gettimeofday().
I would consider jack broken then. For once it breaks
on Centrinos and on AMD systems with PowerNow and some others which all
have frequency scaling with non pstate invariant TSC.
As an additional problem the modern Opterons which support SMP
powernow can even have completely different TSC frequencies
on different CPUs.
All I can recommend is to use gettimeofday() for this. The kernel
goes to considerable pains to make gettimeofday() fast, and when
it is not fast then the system in general cannot do it better.
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 21:26 ` Andy Isaacson
2005-05-13 21:59 ` Matt Mackall
@ 2005-05-14 0:39 ` dean gaudet
2005-05-16 13:41 ` Andrea Arcangeli
2005-05-15 9:43 ` Andi Kleen
` (2 subsequent siblings)
4 siblings, 1 reply; 84+ messages in thread
From: dean gaudet @ 2005-05-14 0:39 UTC (permalink / raw)
To: Andy Isaacson
Cc: Andi Kleen, Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
On Fri, 13 May 2005, Andy Isaacson wrote:
> On Fri, May 13, 2005 at 09:05:49PM +0200, Andi Kleen wrote:
> > On Fri, May 13, 2005 at 02:38:03PM -0400, Richard F. Rebel wrote:
> > > Why? It's certainly reasonable to disable it for the time being and
> > > even prudent to do so.
> >
> > No, i strongly disagree on that. The reasonable thing to do is
> > to fix the crypto code which has this vulnerability, not break
> > a useful performance enhancement for everybody else.
>
> Pardon me for saying so, but that's bullshit. You're asking the crypto
> guys to give up a 5x performance gain (that's my wild guess) by giving
> up all their data-dependent algorithms and contorting their code wildly,
> to avoid a microarchitectural problem with Intel's HT implementation.
i think your wild guess is way off. i can think of several approaches to
fix these problems which won't be anywhere near 5x.
the problem is that an attacker can observe which cache indices (rows) are
in use. one workaround is to overload the possible secrets which each
index represents.
you can overload the secrets in each cache line: for example when doing
exponentiation there is an array of bignums x**(2*n). bignums themselves
are arrays (which span multiple cache lines). do a "row/column transpose"
on this array of arrays -- suddenly each cache line contains a number of
possible secrets. if you're operating with 32-bit words in a 64 byte line
then you've achieved a 16-fold reduction in exposed information by this
transpose. there'll be almost no performance penalty.
you can overload the secrets in each cache index: abuse the associativity
of the cache. the affected processors are all 8-way associative.
ideally you'd want to arrange your data so that it all collides within the
same cache index -- and get an 8-fold reduction in exposure. the trick
here is the L2 is physically indexed, and userland code can perform only
virtual allocations. but it's not too hard to discover physical conflicts
if you really want to (using rdtsc) -- it would be done early in the
initialization of the program because it involves asking for enough memory
until the kernel gives you enough colliding pages. (a system call could
help with this if we really wanted it.)
my not-so-wild guess is a 128-fold reduction for less than 10% perf hit...
i think there's possibly another approach involving a permuted array of
indirection pointers... which is going to affect perf a bit due to the
extra indirection required, but we're talking <10% here. (i'm just not
convinced yet you can select a permutation in a manner which doesn't leak
information when the attacker can view multiple invocations of the crypto
for example.)
> If SHA has plaintext-dependent memory references, Colin's technique
> would enable an adversary to extract the contents of the /dev/random
> pools. I don't *think* SHA does, based on a quick reading of
> lib/sha1.c, but someone with an actual clue should probably take a look.
the SHA family do not have any data-dependencies in their memory access
patterns.
-dean
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-14 0:39 ` dean gaudet
@ 2005-05-16 13:41 ` Andrea Arcangeli
0 siblings, 0 replies; 84+ messages in thread
From: Andrea Arcangeli @ 2005-05-16 13:41 UTC (permalink / raw)
To: dean gaudet
Cc: Andy Isaacson, Andi Kleen, Richard F. Rebel, Gabor MICSKO,
linux-kernel, mpm, tytso
On Fri, May 13, 2005 at 05:39:25PM -0700, dean gaudet wrote:
> same cache index -- and get an 8-fold reduction in exposure. the trick
> here is the L2 is physically indexed, and userland code can perform only
> virtual allocations. but it's not too hard to discover physical conflicts
> if you really want to (using rdtsc) -- it would be done early in the
> initialization of the program because it involves asking for enough memory
> until the kernel gives you enough colliding pages. (a system call could
> help with this if we really wanted it.)
A 8-way set associative 1M cache is guaranteed to go at l2 speed only
up to 128K (no matter what the kernel does), but even if the secret
payload is larger than 128K as long as the load is still distributed
evenly at each pass for each page, there's not going to be any covert
channel, simply the process will run slower than it could if it had a
better page coloring.
So I don't see the need of kernel support, all it needs to do is to know
the page size, and that's provided to userland already.
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 21:26 ` Andy Isaacson
2005-05-13 21:59 ` Matt Mackall
2005-05-14 0:39 ` dean gaudet
@ 2005-05-15 9:43 ` Andi Kleen
2005-05-15 18:42 ` David Schwartz
2005-05-16 7:10 ` Eric W. Biederman
2005-05-15 14:00 ` Mikulas Patocka
2005-05-15 14:26 ` Andi Kleen
4 siblings, 2 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 9:43 UTC (permalink / raw)
To: Andy Isaacson; +Cc: Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
On Fri, May 13, 2005 at 02:26:20PM -0700, Andy Isaacson wrote:
> On Fri, May 13, 2005 at 09:05:49PM +0200, Andi Kleen wrote:
> > On Fri, May 13, 2005 at 02:38:03PM -0400, Richard F. Rebel wrote:
> > > Why? It's certainly reasonable to disable it for the time being and
> > > even prudent to do so.
> >
> > No, i strongly disagree on that. The reasonable thing to do is
> > to fix the crypto code which has this vulnerability, not break
> > a useful performance enhancement for everybody else.
>
> Pardon me for saying so, but that's bullshit. You're asking the crypto
> guys to give up a 5x performance gain (that's my wild guess) by giving
> up all their data-dependent algorithms and contorting their code wildly,
> to avoid a microarchitectural problem with Intel's HT implementation.
And what you're doing is to ask all the non crypto guys to give
up an useful optimization just to fix a problem in the crypto guy's
code. The cache line information leak is just a information leak
bug in the crypto code, not a general problem.
There is much more non crypto code than crypto code around - you
are proposing to screw the majority of codes to solve a relatively
obscure problem of only a few functions, which seems like the totally
wrong approach to me.
BTW the crypto guys are always free to check for hyperthreading
themselves and use different functions. However there is a catch
there - the modern dual core processors which actually have
separated L1 and L2 caches set these too to stay compatible
with old code and license managers.
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* RE: Hyper-Threading Vulnerability
2005-05-15 9:43 ` Andi Kleen
@ 2005-05-15 18:42 ` David Schwartz
2005-05-15 18:56 ` Dr. David Alan Gilbert
2005-05-16 7:10 ` Eric W. Biederman
1 sibling, 1 reply; 84+ messages in thread
From: David Schwartz @ 2005-05-15 18:42 UTC (permalink / raw)
To: linux-kernel
Andi Kleen wrote:
> And what you're doing is to ask all the non crypto guys to give
> up an useful optimization just to fix a problem in the crypto guy's
> code. The cache line information leak is just a information leak
> bug in the crypto code, not a general problem.
Portable code shouldn't even have to know that there is such a thing as a
cache line. It should be able to rely on the operating system not to let
other tasks with a different security context spy on the details of its
operation.
> There is much more non crypto code than crypto code around - you
> are proposing to screw the majority of codes to solve a relatively
> obscure problem of only a few functions, which seems like the totally
> wrong approach to me.
That I do agree with.
> BTW the crypto guys are always free to check for hyperthreading
> themselves and use different functions. However there is a catch
> there - the modern dual core processors which actually have
> separated L1 and L2 caches set these too to stay compatible
> with old code and license managers.
This is just a recipe for making it impossible to write correct code. If
you don't believe the operating system or the hardware is at all at fault
for this problem, then it would follow that they could repeat this same
problem with some new mechanism and still not be at fault. So even if the
program checked for hyper-threading, it would still not be correct. It would
have to check for every possible future way this same type of problem could
arise and hide every type of trace that they could create, even if that
trace is in optimization mechanisms and potential channels over which the
programmer has no knowledge because they don't exist yet.
Let's try a rudctio ad absurdum. Surely you would agree that something
other than than the crypto software is at fault if the operating system or
hardware allowed another process with a different security context to see
every instruction the code executed. The crypto authors shouldn't be
expected to make the instruction flows look identical. How different is
monitoring the memory accesses?
Portable, POSIX-compliant C software shouldn't even have to know that there
is such a thing as a cache line.
I'm not going to be unreasonable though. Hyper-threading is here, and now
that we know the potential problems, it's not unreasonable to ask developers
of crypto code to work around it. But it's not a bug in their code that they
need to fix. In fact, they can't even fix it yet because there is no
portable way to determine if you're on a machine that has hyper-threading or
not.
DS
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 18:42 ` David Schwartz
@ 2005-05-15 18:56 ` Dr. David Alan Gilbert
0 siblings, 0 replies; 84+ messages in thread
From: Dr. David Alan Gilbert @ 2005-05-15 18:56 UTC (permalink / raw)
To: David Schwartz; +Cc: linux-kernel
* David Schwartz (davids@webmaster.com) wrote:
>
> Andi Kleen wrote:
>
> > And what you're doing is to ask all the non crypto guys to give
> > up an useful optimization just to fix a problem in the crypto guy's
> > code. The cache line information leak is just a information leak
> > bug in the crypto code, not a general problem.
>
> Portable code shouldn't even have to know that there is such a thing as a
> cache line. It should be able to rely on the operating system not to let
> other tasks with a different security context spy on the details of its
> operation.
I find it interesting to compare this thread with a thread from about
a week ago talking about how /proc/cpuinfo wasn't consistent
across architectures - where we come round to the view of whether
the application writers shouldn't care/are too dumb/shouldn't need
to know about/can't be trusted with knowing about what the real
hardware is.
Personally I think this is a good case of where the application
should take care of it - with whatever support the OS can really
give.
(That is if this is actually a real problem and not just
purely theoretical - my crypto knowledge isn't good enough
to answer that - but it feels very very abstract).
Dave
-----Open up your eyes, open up your mind, open up your code -------
/ Dr. David Alan Gilbert | Running GNU/Linux on Alpha,68K| Happy \
\ gro.gilbert @ treblig.org | MIPS,x86,ARM,SPARC,PPC & HPPA | In Hex /
\ _________________________|_____ http://www.treblig.org |_______/
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-15 9:43 ` Andi Kleen
2005-05-15 18:42 ` David Schwartz
@ 2005-05-16 7:10 ` Eric W. Biederman
2005-05-16 11:04 ` Andi Kleen
1 sibling, 1 reply; 84+ messages in thread
From: Eric W. Biederman @ 2005-05-16 7:10 UTC (permalink / raw)
To: Andi Kleen
Cc: Andy Isaacson, Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
Andi Kleen <ak@muc.de> writes:
> On Fri, May 13, 2005 at 02:26:20PM -0700, Andy Isaacson wrote:
> > On Fri, May 13, 2005 at 09:05:49PM +0200, Andi Kleen wrote:
> > > On Fri, May 13, 2005 at 02:38:03PM -0400, Richard F. Rebel wrote:
> > > > Why? It's certainly reasonable to disable it for the time being and
> > > > even prudent to do so.
> > >
> > > No, i strongly disagree on that. The reasonable thing to do is
> > > to fix the crypto code which has this vulnerability, not break
> > > a useful performance enhancement for everybody else.
> >
> > Pardon me for saying so, but that's bullshit. You're asking the crypto
> > guys to give up a 5x performance gain (that's my wild guess) by giving
> > up all their data-dependent algorithms and contorting their code wildly,
> > to avoid a microarchitectural problem with Intel's HT implementation.
>
> And what you're doing is to ask all the non crypto guys to give
> up an useful optimization just to fix a problem in the crypto guy's
> code. The cache line information leak is just a information leak
> bug in the crypto code, not a general problem.
It is not a problem in the crypto code, it is a mis-feature of
the hardware/kernel combination. As such you must know be intimate
about each and every flavor of the hardware to attempt to avoid
it in the software, and that way lies madness.
First this is a reminder that prefect security requires an audit
of the hardware as well as the software. As we are neither
auditing the hardware not locking it down we obviously will not
achieve perfection. The question then becomes what can be done
to decrease the likely hood that an application will inadvertently
and unavoidably leak information from timing attacks due to unknown
hardware optimizations? Attacks that do not result from hardware
micro-architecture are another problem and one an application can
anticipate and avoid.
Ideally a solution will be proposed that will allow this problem
to be avoided using the existing POSIX API or at least the current
linux kernel API. But that problem may not be the case.
The only solution I have seen proposed so far that seems to work
is to not schedule untrusted processes simultaneously with
the security code. With the current API that sounds like
a root process killing off, or at least stopping all non-root
processes until the critical process has finished.
Potentially the scheduler can be modified to do this at a finer
grain but I don't know if this would impact the scheduler fast
path. Given the rarity and uncertainty of this it should probably
be something that the process that is worried about security should
asks for instead of simply getting by default.
So it looks to me like the sanest way to handle this is to
allocate a pool of threads/processes one per cpu. Set the
affinity of each process to a particular cpu. And set priority
of the threads to run at the highest priority. And during the
critical time ensure none of the threads are sleeping.
Can someone see a better way to prevent an accidental information
leak to do to hardware architecture details?
I wish there was a better way to ensure all of the threads were
running simultaneously and other then giving them the highest priority
in the system but I don't currently see an alternative.
> There is much more non crypto code than crypto code around - you
> are proposing to screw the majority of codes to solve a relatively
> obscure problem of only a few functions, which seems like the totally
> wrong approach to me.
>
> BTW the crypto guys are always free to check for hyperthreading
> themselves and use different functions. However there is a catch
> there - the modern dual core processors which actually have
> separated L1 and L2 caches set these too to stay compatible
> with old code and license managers.
And those same processors will have the same problem if the share
significant cpu resources. Ideally the entire problem set
would fit in the cache and the cpu designers would allow cache
blocks to be locked but that is not currently the case. So a shared
L3 cache with dual core processors will have the same problem.
In addition a flavor of this attack may be made by repeatedly doing
multiplies or other activities that access functional units and seeing
how long they have to be waited for. So even hyperthreading without
sharing a L2 cache may see this problem.
Eric
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-16 7:10 ` Eric W. Biederman
@ 2005-05-16 11:04 ` Andi Kleen
2005-05-16 19:14 ` Eric W. Biederman
0 siblings, 1 reply; 84+ messages in thread
From: Andi Kleen @ 2005-05-16 11:04 UTC (permalink / raw)
To: Eric W. Biederman
Cc: Andy Isaacson, Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
> The only solution I have seen proposed so far that seems to work
> is to not schedule untrusted processes simultaneously with
> the security code. With the current API that sounds like
> a root process killing off, or at least stopping all non-root
> processes until the critical process has finished.
With virtualization and a hypervisor freely scheduling it is quite
impossible to guarantee this. Of course as always the signal
is quite noisy so it is unclear if it is exploitable in practical
settings. On virtualized environments you cannot use ps to see
if a crypto process is running.
> And those same processors will have the same problem if the share
> significant cpu resources. Ideally the entire problem set
> would fit in the cache and the cpu designers would allow cache
> blocks to be locked but that is not currently the case. So a shared
> L3 cache with dual core processors will have the same problem.
At some point the signal gets noisy enough and the assumptions
an attacker has to make too great for it being an useful attack.
For me it is not even clear it is a real attack on native Linux, at
least the setup in the paper looked highly artifical and quite impractical.
e.g. I suppose it would be quite difficult to really synchronize
to the beginning and end of the RSA encryptions on a server that
does other things too.
-Andi
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-16 11:04 ` Andi Kleen
@ 2005-05-16 19:14 ` Eric W. Biederman
2005-05-16 20:05 ` Valdis.Kletnieks
0 siblings, 1 reply; 84+ messages in thread
From: Eric W. Biederman @ 2005-05-16 19:14 UTC (permalink / raw)
To: Andi Kleen
Cc: Andy Isaacson, Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
Andi Kleen <ak@muc.de> writes:
> > The only solution I have seen proposed so far that seems to work
> > is to not schedule untrusted processes simultaneously with
> > the security code. With the current API that sounds like
> > a root process killing off, or at least stopping all non-root
> > processes until the critical process has finished.
>
> With virtualization and a hypervisor freely scheduling it is quite
> impossible to guarantee this. Of course as always the signal
> is quite noisy so it is unclear if it is exploitable in practical
> settings. On virtualized environments you cannot use ps to see
> if a crypto process is running.
Interesting. I think that is a problem for the hypervisor maintainer.
Although that is about enough to convince me to request a
OS flag that says "please give me privacy" and later that can be passed
down to the hypervisor. My gut feel is running under a hypervisor
is when things will at their most vulnerable.
Where this is a threat is when there will be a lot of RSA
key transactions. At which point it is likely that the attacker
can reproduce enough of the setup to figure out the fine details.
I think discovering a crypto process will simply be a matter
finding a https sever. As for getting the timing how about
initiating a https connection? Getting rid of the noise will certainly
be a challenge but you will have multiple attempts.
> > And those same processors will have the same problem if the share
> > significant cpu resources. Ideally the entire problem set
> > would fit in the cache and the cpu designers would allow cache
> > blocks to be locked but that is not currently the case. So a shared
> > L3 cache with dual core processors will have the same problem.
>
> At some point the signal gets noisy enough and the assumptions
> an attacker has to make too great for it being an useful attack.
> For me it is not even clear it is a real attack on native Linux, at
> least the setup in the paper looked highly artifical and quite impractical.
> e.g. I suppose it would be quite difficult to really synchronize
> to the beginning and end of the RSA encryptions on a server that
> does other things too.
Possibly. But then buffer overflow attacks when you don't know the exact
stack layout are similarly difficult and ways have been found. And if
you have multiple chances things get easier. And if you are aiming
at something easier then brute forcing a private key even the littlest
bit is a help.
When people mmap pages we zero them for the same reason so that
we don't have unintentional information leaks.
I agree that for now because little is known this is a highly specialized
attack. However the trend is now towards increasingly big SMP's.
That increases the number of resources that can be shared so the
possibility of a problem increases. At the rate Intel's cpus are
going we may see throttling of one cpu core when the other one
generates too much heat, because it is busy doing something else cpu
intensive. And other optimizations lead to much easier to imagine
vulnerabilities.
As for noise with the area cpu designers are getting into things
are becoming increasingly fine grained so information is leaking
at an increasingly fine level. As the L2 cache issue has shown
that information starts to leak below the level an application
designer has control of. At which point things get very difficult
to manage.
Information leaks are more difficult than simply gaining root on
the box where you can simply take the information you want. But
that means that is exactly where a locked down well administered
box will be vulnerable if a way is not found to avoid the problem.
I don't know what the consequences of having your private key
discovered are, but I have never heard a case where identity theft
was something pleasant to fix.
Eric
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-16 19:14 ` Eric W. Biederman
@ 2005-05-16 20:05 ` Valdis.Kletnieks
0 siblings, 0 replies; 84+ messages in thread
From: Valdis.Kletnieks @ 2005-05-16 20:05 UTC (permalink / raw)
To: Eric W. Biederman
Cc: Andi Kleen, Andy Isaacson, Richard F. Rebel, Gabor MICSKO,
linux-kernel, mpm, tytso
[-- Attachment #1: Type: text/plain, Size: 713 bytes --]
On Mon, 16 May 2005 13:14:23 MDT, Eric W. Biederman said:
> Interesting. I think that is a problem for the hypervisor maintainer.
> Although that is about enough to convince me to request a
> OS flag that says "please give me privacy" and later that can be passed
> down to the hypervisor. My gut feel is running under a hypervisor
> is when things will at their most vulnerable.
Not really, because....
> I think discovering a crypto process will simply be a matter
> finding a https sever. As for getting the timing how about
> initiating a https connection? Getting rid of the noise will certainly
> be a challenge but you will have multiple attempts.
And the hypervisor is, if anything, adding noise.
[-- Attachment #2: Type: application/pgp-signature, Size: 226 bytes --]
^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: Hyper-Threading Vulnerability
2005-05-13 21:26 ` Andy Isaacson
` (2 preceding siblings ...)
2005-05-15 9:43 ` Andi Kleen
@ 2005-05-15 14:00 ` Mikulas Patocka
2005-05-15 14:26 ` Andi Kleen
4 siblings, 0 replies; 84+ messages in thread
From: Mikulas Patocka @ 2005-05-15 14:00 UTC (permalink / raw)
To: Andy Isaacson
Cc: Andi Kleen, Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
On Fri, 13 May 2005, Andy Isaacson wrote:
> On Fri, May 13, 2005 at 09:05:49PM +0200, Andi Kleen wrote:
> > On Fri, May 13, 2005 at 02:38:03PM -0400, Richard F. Rebel wrote:
> > > Why? It's certainly reasonable to disable it for the time being and
> > > even prudent to do so.
> >
> > No, i strongly disagree on that. The reasonable thing to do is
> > to fix the crypto code which has this vulnerability, not break
> > a useful performance enhancement for everybody else.
>
> Pardon me for saying so, but that's bullshit. You're asking the crypto
> guys to give up a 5x performance gain (that's my wild guess) by giving
> up all their data-dependent algorithms and contorting their code wildly,
> to avoid a microarchitectural problem with Intel's HT implementation.
That information leak can be exploited not only on HT or SMP, but on any
CPU with L2 cache. Without HT it's much harder to get information about L2
cache footprint, but it's still possible. If an attacker can make
unlimited number of connections to ssh or http server and manages to get 1
bit in 100 connections, it's still a problem.
Possible solutions:
1) don't use branches and data-dependent memory accesses depending on
secret data
2) flush cache completely when switching to process with different EUID
(0.2ms on Pentium 4 with 1M cache, even worse on CPUs with more cache).
Disabling HT/SMP is not a solution. A year later someone may come with
something like this:
* prefill L2 cache with known pattern
* sleep on some precious timer
* make connection to security application (ssh, https)
* on wakeup, read what's in L2 cache --- get one bit with small
probability --- but when repeated many times, it's still a problem
Mikulas
> There are three places to cut off the side channel, none of which is
> obviously the right one.
> 1. The HT implementation could do the cache tricks Colin suggested in
> his paper. Fairly large performance hit to address a fairly small
> problem.
> 2. The OS could do the scheduler tricks to avoid scheduling unfriendly
> threads on the same core. You're leaving a lot of the benefit of HT
> on the floor by doing so.
> 3. Every security-sensitive app can be rigorously audited and re-written
> to avoid *ever* referencing memory with the address determined by
> private data.
>
> (3) is a complete non-starter. It's just not feasible to rewrite all
> that code. Furthermore, there's no way to know what code needs to be
> rewritten! (Until someone publishes an advisory, that is...)
>
> Hmm, I can't think of any reason that this technique wouldn't work to
> extract information from kernel secrets, as well...
>
> If SHA has plaintext-dependent memory references, Colin's technique
> would enable an adversary to extract the contents of the /dev/random
> pools. I don't *think* SHA does, based on a quick reading of
> lib/sha1.c, but someone with an actual clue should probably take a look.
>
> Andi, are you prepared to *require* that no code ever make a memory
> reference as a function of a secret? Because that's what you're
> suggesting the crypto people should do.
>
> -andy
> -
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at http://www.tux.org/lkml/
>
^ permalink raw reply [flat|nested] 84+ messages in thread* Re: Hyper-Threading Vulnerability
2005-05-13 21:26 ` Andy Isaacson
` (3 preceding siblings ...)
2005-05-15 14:00 ` Mikulas Patocka
@ 2005-05-15 14:26 ` Andi Kleen
4 siblings, 0 replies; 84+ messages in thread
From: Andi Kleen @ 2005-05-15 14:26 UTC (permalink / raw)
To: Andy Isaacson; +Cc: Richard F. Rebel, Gabor MICSKO, linux-kernel, mpm, tytso
> There are three places to cut off the side channel, none of which is
> obviously the right one.
> 1. The HT implementation could do the cache tricks Colin suggested in
> his paper. Fairly large performance hit to address a fairly small
> problem.
As Dean pointed out that is probably not true.
> 2. The OS could do the scheduler tricks to avoid scheduling unfriendly
> threads on the same core. You're leaving a lot of the benefit of HT
> on the floor by doing so.
And probably still lose badly in some workloads.
> 3. Every security-sensitive app can be rigorously audited and re-written
> to avoid *ever* referencing memory with the address determined by
> private data.
Sure after it was demonstrated that this attack is actually feasible
in practice. If yes then fix the crypto code, otherwise do nothing.
I have no problem with crypto people being paranoid (that is their
job after all), as long as they don't try to affect non crypto code
in the process. But the later seems to be clearly the case here :-(
>
> (3) is a complete non-starter. It's just not feasible to rewrite all
> that code. Furthermore, there's no way to know what code needs to be
> rewritten! (Until someone publishes an advisory, that is...)
>
> Hmm, I can't think of any reason that this technique wouldn't work to
> extract information from kernel secrets, as well...
>
> If SHA has plaintext-dependent memory references, Colin's technique
> would enable an adversary to extract the contents of the /dev/random
> pools. I don't *think* SHA does, based on a quick reading of
> lib/sha1.c, but someone with an actual clue should probably take a look.
>
> Andi, are you prepared to *require* that no code ever make a memory
> reference as a function of a secret? Because that's what you're
> suggesting the crypto people should do.
No, just not do it frequently enough that you leak enough data.
Or add dummy memory references to blend your data.
And then nobody said writing crypto code was easy. It just got a bit
harder today.
It is basically like writing smart card code, where you need
to care about such side channels. The other crypto code writers
just need to care about this too. They will probably avoid
other timing attacks on cache misses too with this approach. Although
it is doubtful enough signal is leaked in this way, e.g. if you time the
performance of a network server with RSA answering over the network -
but you see some data is always leaked - the question is just
if it is enough and accurate data to aid an attacker. The paper
has shown that it is feasible in some cases, but so far the proof
is still out this could be actually replicated in not very controlled
loads. With more noise in the data it becomes harder. And the question
is is the small amount of data with normal background workload
is really useful enough to lead to useful real world attacks. I have
severe doubts on that. Certainly the effidence is not clear enough
for a serious step like disabling an useful performance enhancement like
HT.
-Andi
P.S.: My personal opinion is that we have a far bigger crypto security
problem on many system due to weak /dev/random seeding on many systems.
If anything is done it would be better to attack that.
^ permalink raw reply [flat|nested] 84+ messages in thread