* [BUG] perf: bogus correlation of kernel symbols
@ 2011-05-12 14:48 Stephane Eranian
2011-05-12 18:06 ` David Miller
2011-05-12 20:31 ` Linus Torvalds
0 siblings, 2 replies; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 14:48 UTC (permalink / raw)
To: Arnaldo Carvalho de Melo; +Cc: LKML
Hi,
I think there is a serious problem with kernel symbol correlation
with the latest perf in 2.6.39-rc7-tip.
Here is a simple example with a stupid program that only
does open()/close on /dev/null:
$ perf record -e cycles:k openclose
$ perf report --stdio
# Events: 2K cycles
#
# Overhead Command Shared Object Symbol
# ........ ......... ................ ...............
#
99.76% openclose [binfmt_misc] [k] 0xffffffff81010fe6
0.13% openclose libc-2.12.1.so [.] __open_nocancel
0.09% openclose libc-2.12.1.so [.] __GI_close
The DSO (binfmt_misc) is bogus. That's not where time is spent.
But if I ran the same test as root:
$ sudo perf record -e cycles:k openclose
$ sudo perf report --stdio
# Events: 2K cycles
#
# Overhead Command Shared Object Symbol
# ........ ......... ................. .............................
#
17.13% openclose [kernel.kallsyms] [k] __lock_acquire
11.77% openclose [kernel.kallsyms] [k] native_sched_clock
7.36% openclose [kernel.kallsyms] [k] sched_clock_local
5.99% openclose [kernel.kallsyms] [k] lock_release
5.38% openclose [kernel.kallsyms] [k] local_clock
4.43% openclose [kernel.kallsyms] [k] lock_acquired
4.05% openclose [kernel.kallsyms] [k] lock_acquire
3.95% openclose [kernel.kallsyms] [k] lock_is_held
3.51% openclose [kernel.kallsyms] [k] sched_clock_cpu
3.24% openclose [kernel.kallsyms] [k] trace_hardirqs_off_caller
This is much more meaningful.
This is not related to the paranoid level (1 for me).
Looking at perf report -D, the same kernel address is associated to different
module based on my permission level.
first perf.data:
416749738927 0x4210 [0x28]: PERF_RECORD_SAMPLE(IP, 1): 4886/4886:
0xffffffff8107c1d8 period: 2262681
... thread: openclose:4886
...... dso: /lib/modules/2.6.39-rc7-tip/kernel/fs/binfmt_misc.ko
second perf.data:
436879910722 0xc950 [0x28]: PERF_RECORD_SAMPLE(IP, 1): 4894/4894:
0xffffffff8107c1d8 period: 2280253
... thread: openclose:4894
...... dso: vmlinux
Same address different mapping!
My path to vmlinux is all accessible to me.
If there were permission problems, I would expect perf record or perf report
to tell me and not fallback to some bogus mappings.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 14:48 [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
@ 2011-05-12 18:06 ` David Miller
2011-05-12 18:37 ` Dave Jones
2011-05-12 21:06 ` Ingo Molnar
2011-05-12 20:31 ` Linus Torvalds
1 sibling, 2 replies; 29+ messages in thread
From: David Miller @ 2011-05-12 18:06 UTC (permalink / raw)
To: eranian; +Cc: acme, linux-kernel
From: Stephane Eranian <eranian@google.com>
Date: Thu, 12 May 2011 16:48:46 +0200
> I think there is a serious problem with kernel symbol correlation
> with the latest perf in 2.6.39-rc7-tip.
The behavior seems to be intentional, so that we don't expose internal
kernel addresses to userspace.
I hate this too, and I think it's absolutely rediculous.
Also, like you, I lost an entire afternoon trying to figure out why
this started happening.
I wish we could revert this change.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 18:06 ` David Miller
@ 2011-05-12 18:37 ` Dave Jones
2011-05-12 19:01 ` David Miller
2011-05-12 21:06 ` Ingo Molnar
1 sibling, 1 reply; 29+ messages in thread
From: Dave Jones @ 2011-05-12 18:37 UTC (permalink / raw)
To: David Miller; +Cc: eranian, acme, linux-kernel
On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
> From: Stephane Eranian <eranian@google.com>
> Date: Thu, 12 May 2011 16:48:46 +0200
>
> > I think there is a serious problem with kernel symbol correlation
> > with the latest perf in 2.6.39-rc7-tip.
>
> The behavior seems to be intentional, so that we don't expose internal
> kernel addresses to userspace.
Sounds like commit 9f36e2c448007b54851e7e4fa48da97d1477a175
> I hate this too, and I think it's absolutely rediculous.
>
> Also, like you, I lost an entire afternoon trying to figure out why
> this started happening.
>
> I wish we could revert this change.
At least it can be permanently disabled..
echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
Dave
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 18:37 ` Dave Jones
@ 2011-05-12 19:01 ` David Miller
2011-05-12 19:58 ` Pekka Enberg
2011-05-12 20:24 ` Alexey Dobriyan
0 siblings, 2 replies; 29+ messages in thread
From: David Miller @ 2011-05-12 19:01 UTC (permalink / raw)
To: davej; +Cc: eranian, acme, linux-kernel
From: Dave Jones <davej@redhat.com>
Date: Thu, 12 May 2011 14:37:41 -0400
> On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
> > I hate this too, and I think it's absolutely rediculous.
> >
> > Also, like you, I lost an entire afternoon trying to figure out why
> > this started happening.
> >
> > I wish we could revert this change.
>
> At least it can be permanently disabled..
>
> echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
Regardless, what to do about all of the "perf is broken" reports?
First off, perf can find out whether this madness exists, and it
should by default print out a warning in this situation instead of
knowingly emitting garbage kernel event information.
"I'm going to knowingly give you bad data, and I'm not even going to
let you know about it."
It's really crazy that we give people these incredibly powerful tools
and they don't even work properly by default.
We've been exposing kernel pointers for 20 years, nobody's grandmother
died because of it.
This is very "Animal Farm" the way we're gradually losing little bits
of functionality, time and time again, over this "kernel pointer
exposure" issue.
Are we going to be like animals and just accept the totality of this,
or are we going to be outraged enough to push back on stuff like perf
actually working properly?
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 19:01 ` David Miller
@ 2011-05-12 19:58 ` Pekka Enberg
2011-05-13 6:12 ` Kees Cook
2011-05-12 20:24 ` Alexey Dobriyan
1 sibling, 1 reply; 29+ messages in thread
From: Pekka Enberg @ 2011-05-12 19:58 UTC (permalink / raw)
To: David Miller
Cc: davej, eranian, acme, linux-kernel, Kees Cook, Linus Torvalds,
Ingo Molnar
On Thu, May 12, 2011 at 10:01 PM, David Miller <davem@davemloft.net> wrote:
> From: Dave Jones <davej@redhat.com>
> Date: Thu, 12 May 2011 14:37:41 -0400
>
>> On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
>> > I hate this too, and I think it's absolutely rediculous.
>> >
>> > Also, like you, I lost an entire afternoon trying to figure out why
>> > this started happening.
>> >
>> > I wish we could revert this change.
>>
>> At least it can be permanently disabled..
>>
>> echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
>
> Regardless, what to do about all of the "perf is broken" reports?
Lets revert the commit 9f36e2c448007b54851e7e4fa48da97d1477a175
("printk: use %pK for /proc/kallsyms and /proc/modules"), please! I
too have been wondering what's going on with perf reporting insane
symbols and this should definitely not be enabled by default.
Pekka
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 19:01 ` David Miller
2011-05-12 19:58 ` Pekka Enberg
@ 2011-05-12 20:24 ` Alexey Dobriyan
1 sibling, 0 replies; 29+ messages in thread
From: Alexey Dobriyan @ 2011-05-12 20:24 UTC (permalink / raw)
To: David Miller; +Cc: davej, eranian, acme, linux-kernel
On Thu, May 12, 2011 at 03:01:32PM -0400, David Miller wrote:
> From: Dave Jones <davej@redhat.com>
> Date: Thu, 12 May 2011 14:37:41 -0400
>
> > On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
> > > I hate this too, and I think it's absolutely rediculous.
> > >
> > > Also, like you, I lost an entire afternoon trying to figure out why
> > > this started happening.
> > >
> > > I wish we could revert this change.
> >
> > At least it can be permanently disabled..
> >
> > echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
>
> Regardless, what to do about all of the "perf is broken" reports?
The problem is that they turned it on by default.
int kptr_restrict = 1;
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 14:48 [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
2011-05-12 18:06 ` David Miller
@ 2011-05-12 20:31 ` Linus Torvalds
2011-05-12 20:43 ` David Miller
2011-05-12 21:07 ` [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
1 sibling, 2 replies; 29+ messages in thread
From: Linus Torvalds @ 2011-05-12 20:31 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Arnaldo Carvalho de Melo, LKML, Ingo Molnar
On Thu, May 12, 2011 at 7:48 AM, Stephane Eranian <eranian@google.com> wrote:
>
> I think there is a serious problem with kernel symbol correlation
> with the latest perf in 2.6.39-rc7-tip.
Yeah. It's annoying. It's a "perf" bug, though - triggered by
/proc/sys/kernel/kptr_restrict being set to 1.
The bug is that perf doesn't say "I can't match kernel symbols", but
instead does some crazy matching and gives total crap module
information (I think it just picks the one that shows up last in
/proc/kallsyms).
That said, I have considered just reverting the thing that makes
kptr_restrict be 1 by default. I do like the security implications of
restricting visibility into kernel pointers, but I also think that
security rules that make the system less usable are dubious. So I
dunno.
Linus
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 20:31 ` Linus Torvalds
@ 2011-05-12 20:43 ` David Miller
2011-05-12 21:00 ` [PATCH] vsprintf: Turn kptr_restrict off by default Ingo Molnar
2011-05-12 21:07 ` [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
1 sibling, 1 reply; 29+ messages in thread
From: David Miller @ 2011-05-12 20:43 UTC (permalink / raw)
To: torvalds; +Cc: eranian, acme, linux-kernel, mingo
From: Linus Torvalds <torvalds@linux-foundation.org>
Date: Thu, 12 May 2011 13:31:37 -0700
> That said, I have considered just reverting the thing that makes
> kptr_restrict be 1 by default. I do like the security implications of
> restricting visibility into kernel pointers, but I also think that
> security rules that make the system less usable are dubious. So I
> dunno.
We don't have any firewalling or SELINUX rules installed by default,
even if those features are enabled in the kernel. Userspace asks for
it.
Many people would claim that use of such things are "essential" these
days.
I don't see a good reason to handle kptr_restrict any differently.
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH] vsprintf: Turn kptr_restrict off by default
2011-05-12 20:43 ` David Miller
@ 2011-05-12 21:00 ` Ingo Molnar
2011-05-12 21:08 ` David Miller
0 siblings, 1 reply; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:00 UTC (permalink / raw)
To: David Miller; +Cc: torvalds, eranian, acme, linux-kernel
* David Miller <davem@davemloft.net> wrote:
> From: Linus Torvalds <torvalds@linux-foundation.org>
> Date: Thu, 12 May 2011 13:31:37 -0700
>
> > That said, I have considered just reverting the thing that makes
> > kptr_restrict be 1 by default. I do like the security implications of
> > restricting visibility into kernel pointers, but I also think that
> > security rules that make the system less usable are dubious. So I
> > dunno.
>
> We don't have any firewalling or SELINUX rules installed by default, even if
> those features are enabled in the kernel. Userspace asks for it.
>
> Many people would claim that use of such things are "essential" these days.
>
> I don't see a good reason to handle kptr_restrict any differently.
That's a good argument.
We'll fix the perf bug - i was bitten by another incarnation of it: 'perf top'
stops showing kernel symbols and it took some time that kptr_restrict was
turned on by default. I reported it to Arnaldo knows about it but there's no
fix yet at the moment.
I didnt realize that perf diff got confused by this as well. (but it's logical)
So how about the patch below?
Thanks,
Ingo
------------------------>
Subject: vsprintf: Turn kptr_restrict off by default
kptr_restrict has been triggering bugs in apps such as perf, and it also makes
the system less useful by default, so turn it off by default.
This is how we generally handle security features that remove functionality,
such as firewall code or SELinux - they have to be configured and activated
from user-space.
Distributions can turn kptr_restrict on again via this line in
/etc/sysctrl.conf:
kernel.kptr_restrict = 1
( Also mark the variable __read_mostly while at it, as it's typically modified
only once per bootup, or not at all. )
Signed-off-by: Ingo Molnar <mingo@elte.hu>
---
diff --git a/lib/vsprintf.c b/lib/vsprintf.c
index bc0ac6b..dfd6019 100644
--- a/lib/vsprintf.c
+++ b/lib/vsprintf.c
@@ -797,7 +797,7 @@ char *uuid_string(char *buf, char *end, const u8 *addr,
return string(buf, end, uuid, spec);
}
-int kptr_restrict = 1;
+int kptr_restrict __read_mostly;
/*
* Show a '%p' thing. A kernel extension is that the '%p' is followed
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 18:06 ` David Miller
2011-05-12 18:37 ` Dave Jones
@ 2011-05-12 21:06 ` Ingo Molnar
1 sibling, 0 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:06 UTC (permalink / raw)
To: David Miller; +Cc: eranian, acme, linux-kernel, Linus Torvalds
* David Miller <davem@davemloft.net> wrote:
> From: Stephane Eranian <eranian@google.com>
> Date: Thu, 12 May 2011 16:48:46 +0200
>
> > I think there is a serious problem with kernel symbol correlation
> > with the latest perf in 2.6.39-rc7-tip.
>
> The behavior seems to be intentional, so that we don't expose internal
> kernel addresses to userspace.
>
> I hate this too, and I think it's absolutely rediculous.
>
> Also, like you, I lost an entire afternoon trying to figure out why
> this started happening.
I lost about an hour with Arnaldo on IRC to help me until we figured out that
/proc/kallsyms started having zero value entries ... I'm too running perf as an
unprivileged user.
Zero is a valid symbol address so nothing within perf tripped up explicitly,
but perf report and perf top results were nonsensical.
There was another problem with it: perf is caching and storing known kernel
buildid addresses in ~/.debug, under the (previously correct) assumption that
kernel symbols do not change for one given kernel build. But with kptr_restrict
it would cache the zero values - which were cached even after kptr_restrict was
set back to 0.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 20:31 ` Linus Torvalds
2011-05-12 20:43 ` David Miller
@ 2011-05-12 21:07 ` Stephane Eranian
2011-05-12 21:30 ` Stephane Eranian
2011-05-12 21:36 ` Ingo Molnar
1 sibling, 2 replies; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 21:07 UTC (permalink / raw)
To: Linus Torvalds; +Cc: Arnaldo Carvalho de Melo, LKML, Ingo Molnar
On Thu, May 12, 2011 at 10:31 PM, Linus Torvalds
<torvalds@linux-foundation.org> wrote:
>
> On Thu, May 12, 2011 at 7:48 AM, Stephane Eranian <eranian@google.com> wrote:
> >
> > I think there is a serious problem with kernel symbol correlation
> > with the latest perf in 2.6.39-rc7-tip.
>
> Yeah. It's annoying. It's a "perf" bug, though - triggered by
> /proc/sys/kernel/kptr_restrict being set to 1.
>
I did not know about this new masquerading of pointers in /proc/kallsyms.
That certainly explains the problem.
>
> The bug is that perf doesn't say "I can't match kernel symbols", but
> instead does some crazy matching and gives total crap module
> information (I think it just picks the one that shows up last in
> /proc/kallsyms).
>
But I agree perf must not silently return bogus information. It
should print a big warning message and/or fallback to printing the raw
addresses. So much for having perf in the kernel source tree to
keep things in sync...
>
> That said, I have considered just reverting the thing that makes
> kptr_restrict be 1 by default. I do like the security implications of
> restricting visibility into kernel pointers, but I also think that
> security rules that make the system less usable are dubious. So I
> dunno.
>
I am not clear as to what people could actually do with the addresses
taken out of /proc/kallsyms. Looks to me like we've lost functionality
for the vast majority of users. So maybe the default should be inverted.
I know of a somewhat similar issue with the file descriptor limit which
people are hitting frequently these days when monitoring apps with lots
of threads or lots of events in one run on large smp systems.
That can easily be corrected by again requires root privilege to regain
the functionality.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH] vsprintf: Turn kptr_restrict off by default
2011-05-12 21:00 ` [PATCH] vsprintf: Turn kptr_restrict off by default Ingo Molnar
@ 2011-05-12 21:08 ` David Miller
0 siblings, 0 replies; 29+ messages in thread
From: David Miller @ 2011-05-12 21:08 UTC (permalink / raw)
To: mingo; +Cc: torvalds, eranian, acme, linux-kernel
From: Ingo Molnar <mingo@elte.hu>
Date: Thu, 12 May 2011 23:00:28 +0200
> Subject: vsprintf: Turn kptr_restrict off by default
>
> kptr_restrict has been triggering bugs in apps such as perf, and it also makes
> the system less useful by default, so turn it off by default.
>
> This is how we generally handle security features that remove functionality,
> such as firewall code or SELinux - they have to be configured and activated
> from user-space.
>
> Distributions can turn kptr_restrict on again via this line in
> /etc/sysctrl.conf:
>
> kernel.kptr_restrict = 1
>
> ( Also mark the variable __read_mostly while at it, as it's typically modified
> only once per bootup, or not at all. )
>
> Signed-off-by: Ingo Molnar <mingo@elte.hu>
Acked-by: David S. Miller <davem@davemloft.net>
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:07 ` [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
@ 2011-05-12 21:30 ` Stephane Eranian
2011-05-12 21:35 ` Ingo Molnar
2011-05-12 21:36 ` Ingo Molnar
1 sibling, 1 reply; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 21:30 UTC (permalink / raw)
To: Linus Torvalds; +Cc: Arnaldo Carvalho de Melo, LKML, Ingo Molnar
The other contradiction, I see, is that you have perf_event paranoia level
and this new kptr masquerading feature which conflict with each
other.
You can be allowed to monitor at the kernel level (paranoid=1, default)
but you cannot correlate symbols:
$ perf record -e cycles:k foo
I suspect if you have this kptr thing turned on, then you need to disallow
monitoring at the kernel level too.
On Thu, May 12, 2011 at 11:07 PM, Stephane Eranian <eranian@google.com> wrote:
> On Thu, May 12, 2011 at 10:31 PM, Linus Torvalds
> <torvalds@linux-foundation.org> wrote:
>>
>> On Thu, May 12, 2011 at 7:48 AM, Stephane Eranian <eranian@google.com> wrote:
>> >
>> > I think there is a serious problem with kernel symbol correlation
>> > with the latest perf in 2.6.39-rc7-tip.
>>
>> Yeah. It's annoying. It's a "perf" bug, though - triggered by
>> /proc/sys/kernel/kptr_restrict being set to 1.
>>
> I did not know about this new masquerading of pointers in /proc/kallsyms.
> That certainly explains the problem.
>
>>
>> The bug is that perf doesn't say "I can't match kernel symbols", but
>> instead does some crazy matching and gives total crap module
>> information (I think it just picks the one that shows up last in
>> /proc/kallsyms).
>>
> But I agree perf must not silently return bogus information. It
> should print a big warning message and/or fallback to printing the raw
> addresses. So much for having perf in the kernel source tree to
> keep things in sync...
>
>>
>> That said, I have considered just reverting the thing that makes
>> kptr_restrict be 1 by default. I do like the security implications of
>> restricting visibility into kernel pointers, but I also think that
>> security rules that make the system less usable are dubious. So I
>> dunno.
>>
> I am not clear as to what people could actually do with the addresses
> taken out of /proc/kallsyms. Looks to me like we've lost functionality
> for the vast majority of users. So maybe the default should be inverted.
>
> I know of a somewhat similar issue with the file descriptor limit which
> people are hitting frequently these days when monitoring apps with lots
> of threads or lots of events in one run on large smp systems.
> That can easily be corrected by again requires root privilege to regain
> the functionality.
>
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:30 ` Stephane Eranian
@ 2011-05-12 21:35 ` Ingo Molnar
2011-05-12 21:38 ` Stephane Eranian
0 siblings, 1 reply; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:35 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> The other contradiction, I see, is that you have perf_event paranoia level
> and this new kptr masquerading feature which conflict with each
> other.
>
> You can be allowed to monitor at the kernel level (paranoid=1, default)
> but you cannot correlate symbols:
>
> $ perf record -e cycles:k foo
>
> I suspect if you have this kptr thing turned on, then you need to disallow
> monitoring at the kernel level too.
The better (and consistent) solution would be to turn the kptr_restrict thing
off - see the patch i sent.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:07 ` [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
2011-05-12 21:30 ` Stephane Eranian
@ 2011-05-12 21:36 ` Ingo Molnar
2011-05-12 21:41 ` Stephane Eranian
1 sibling, 1 reply; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:36 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> > The bug is that perf doesn't say "I can't match kernel symbols", but
> > instead does some crazy matching and gives total crap module information (I
> > think it just picks the one that shows up last in /proc/kallsyms).
>
> But I agree perf must not silently return bogus information. It should print
> a big warning message and/or fallback to printing the raw addresses. [...]
Yes, agreed, this is a bug in perf. I found out about this about two weeks ago
and reported it to Arnaldo, but he is away right now - he might be able to fix
it next week the earliest.
> [...] So much for having perf in the kernel source tree to keep things in
> sync...
What do you mean?
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:35 ` Ingo Molnar
@ 2011-05-12 21:38 ` Stephane Eranian
2011-05-12 21:50 ` Ingo Molnar
0 siblings, 1 reply; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 21:38 UTC (permalink / raw)
To: Ingo Molnar; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
On Thu, May 12, 2011 at 11:35 PM, Ingo Molnar <mingo@elte.hu> wrote:
>
> * Stephane Eranian <eranian@google.com> wrote:
>
>> The other contradiction, I see, is that you have perf_event paranoia level
>> and this new kptr masquerading feature which conflict with each
>> other.
>>
>> You can be allowed to monitor at the kernel level (paranoid=1, default)
>> but you cannot correlate symbols:
>>
>> $ perf record -e cycles:k foo
>>
>> I suspect if you have this kptr thing turned on, then you need to disallow
>> monitoring at the kernel level too.
>
> The better (and consistent) solution would be to turn the kptr_restrict thing
> off - see the patch i sent.
>
I saw that. But I think that when someone turns it back on, then you need
to increase the perf_events paranoia level to disallow kernel monitoring to
regular users such that you maintain consistency across the board.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:36 ` Ingo Molnar
@ 2011-05-12 21:41 ` Stephane Eranian
2011-05-12 21:54 ` Ingo Molnar
0 siblings, 1 reply; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 21:41 UTC (permalink / raw)
To: Ingo Molnar; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
On Thu, May 12, 2011 at 11:36 PM, Ingo Molnar <mingo@elte.hu> wrote:
>
> * Stephane Eranian <eranian@google.com> wrote:
>
>> > The bug is that perf doesn't say "I can't match kernel symbols", but
>> > instead does some crazy matching and gives total crap module information (I
>> > think it just picks the one that shows up last in /proc/kallsyms).
>>
>> But I agree perf must not silently return bogus information. It should print
>> a big warning message and/or fallback to printing the raw addresses. [...]
>
> Yes, agreed, this is a bug in perf. I found out about this about two weeks ago
> and reported it to Arnaldo, but he is away right now - he might be able to fix
> it next week the earliest.
>
>> [...] So much for having perf in the kernel source tree to keep things in
>> sync...
>
> What do you mean?
>
I meant that when this kptr feature was added, people should have scanned the
entire tree (include tools/perf) to look for potential impact on
programs relying
on /proc/kallsyms. Having perf in the tree should have made this easier to
catch. That's all.
> Thanks,
>
> Ingo
>
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:38 ` Stephane Eranian
@ 2011-05-12 21:50 ` Ingo Molnar
2011-05-12 21:56 ` Stephane Eranian
2011-05-12 22:07 ` Dave Jones
0 siblings, 2 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:50 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> On Thu, May 12, 2011 at 11:35 PM, Ingo Molnar <mingo@elte.hu> wrote:
> >
> > * Stephane Eranian <eranian@google.com> wrote:
> >
> >> The other contradiction, I see, is that you have perf_event paranoia level
> >> and this new kptr masquerading feature which conflict with each
> >> other.
> >>
> >> You can be allowed to monitor at the kernel level (paranoid=1, default)
> >> but you cannot correlate symbols:
> >>
> >> $ perf record -e cycles:k foo
> >>
> >> I suspect if you have this kptr thing turned on, then you need to disallow
> >> monitoring at the kernel level too.
> >
> > The better (and consistent) solution would be to turn the kptr_restrict thing
> > off - see the patch i sent.
>
> I saw that. But I think that when someone turns it back on, then you need to
> increase the perf_events paranoia level to disallow kernel monitoring to
> regular users such that you maintain consistency across the board.
Dunno, i would not couple them necessarily - certain users might still have
access to kernel symbols via some other channel - for example the System.map.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:41 ` Stephane Eranian
@ 2011-05-12 21:54 ` Ingo Molnar
0 siblings, 0 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 21:54 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> >> [...] So much for having perf in the kernel source tree to keep things in
> >> sync...
> >
> > What do you mean?
>
> I meant that when this kptr feature was added, people should have scanned the
> entire tree (include tools/perf) to look for potential impact on programs
> relying on /proc/kallsyms. Having perf in the tree should have made this
> easier to catch. That's all.
It was noticed in another case when there was kallsyms twiddling going on so it
depends. What wasnt noticed here was how the present but zero value symbols:
0000000000000000 D irq_stack_union
0000000000000000 D __per_cpu_start
0000000000000000 D gdt_page
0000000000000000 d exception_stacks
0000000000000000 d tlb_vector_offset
0000000000000000 d shared_msrs
0000000000000000 d cpu_tsc_khz
caused the symbol code of perf consider them non-existing.
Perf being in-tree wont magically avoid all bugs, so you should not expect that
magical effect from tool integration into the kernel tree.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:50 ` Ingo Molnar
@ 2011-05-12 21:56 ` Stephane Eranian
2011-05-12 22:00 ` Ingo Molnar
2011-05-12 22:07 ` Dave Jones
1 sibling, 1 reply; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 21:56 UTC (permalink / raw)
To: Ingo Molnar; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
On Thu, May 12, 2011 at 11:50 PM, Ingo Molnar <mingo@elte.hu> wrote:
>
> * Stephane Eranian <eranian@google.com> wrote:
>
>> On Thu, May 12, 2011 at 11:35 PM, Ingo Molnar <mingo@elte.hu> wrote:
>> >
>> > * Stephane Eranian <eranian@google.com> wrote:
>> >
>> >> The other contradiction, I see, is that you have perf_event paranoia level
>> >> and this new kptr masquerading feature which conflict with each
>> >> other.
>> >>
>> >> You can be allowed to monitor at the kernel level (paranoid=1, default)
>> >> but you cannot correlate symbols:
>> >>
>> >> $ perf record -e cycles:k foo
>> >>
>> >> I suspect if you have this kptr thing turned on, then you need to disallow
>> >> monitoring at the kernel level too.
>> >
>> > The better (and consistent) solution would be to turn the kptr_restrict thing
>> > off - see the patch i sent.
>>
>> I saw that. But I think that when someone turns it back on, then you need to
>> increase the perf_events paranoia level to disallow kernel monitoring to
>> regular users such that you maintain consistency across the board.
>
> Dunno, i would not couple them necessarily - certain users might still have
> access to kernel symbols via some other channel - for example the System.map.
>
Ok, that's true, but then you'd need to have perf print a message or refuse to
use /proc/kallsyms and suggest that the user provides a System.map.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:56 ` Stephane Eranian
@ 2011-05-12 22:00 ` Ingo Molnar
0 siblings, 0 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-12 22:00 UTC (permalink / raw)
To: Stephane Eranian; +Cc: Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> On Thu, May 12, 2011 at 11:50 PM, Ingo Molnar <mingo@elte.hu> wrote:
> >
> > * Stephane Eranian <eranian@google.com> wrote:
> >
> >> On Thu, May 12, 2011 at 11:35 PM, Ingo Molnar <mingo@elte.hu> wrote:
> >> >
> >> > * Stephane Eranian <eranian@google.com> wrote:
> >> >
> >> >> The other contradiction, I see, is that you have perf_event paranoia level
> >> >> and this new kptr masquerading feature which conflict with each
> >> >> other.
> >> >>
> >> >> You can be allowed to monitor at the kernel level (paranoid=1, default)
> >> >> but you cannot correlate symbols:
> >> >>
> >> >> $ perf record -e cycles:k foo
> >> >>
> >> >> I suspect if you have this kptr thing turned on, then you need to disallow
> >> >> monitoring at the kernel level too.
> >> >
> >> > The better (and consistent) solution would be to turn the kptr_restrict thing
> >> > off - see the patch i sent.
> >>
> >> I saw that. But I think that when someone turns it back on, then you need to
> >> increase the perf_events paranoia level to disallow kernel monitoring to
> >> regular users such that you maintain consistency across the board.
> >
> > Dunno, i would not couple them necessarily - certain users might still have
> > access to kernel symbols via some other channel - for example the System.map.
>
> Ok, that's true, but then you'd need to have perf print a message or refuse to
> use /proc/kallsyms and suggest that the user provides a System.map.
Correct - the right approach would be to just use what we had in earlier
versions, an 'unknown symbol' kind of catch-all entry that shows how much
stuff we did not recognize.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 21:50 ` Ingo Molnar
2011-05-12 21:56 ` Stephane Eranian
@ 2011-05-12 22:07 ` Dave Jones
2011-05-12 22:15 ` Stephane Eranian
2011-05-13 8:57 ` Ingo Molnar
1 sibling, 2 replies; 29+ messages in thread
From: Dave Jones @ 2011-05-12 22:07 UTC (permalink / raw)
To: Ingo Molnar
Cc: Stephane Eranian, Linus Torvalds, Arnaldo Carvalho de Melo, LKML
On Thu, May 12, 2011 at 11:50:23PM +0200, Ingo Molnar wrote:
> Dunno, i would not couple them necessarily - certain users might still have
> access to kernel symbols via some other channel - for example the System.map.
That always made this security by obscurity feature seem pointless for the bulk
of users to me. Given the majority are going to be running distro kernels,
anyone can find those addresses easily no matter how hard we hide them on the
running system.
Unless we were somehow introduced randomness into where we unpack the kernel
each boot, and using System.map as a table of offsets instead of absolute addresses.
Dave
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 22:07 ` Dave Jones
@ 2011-05-12 22:15 ` Stephane Eranian
2011-05-13 9:01 ` Ingo Molnar
2011-05-13 8:57 ` Ingo Molnar
1 sibling, 1 reply; 29+ messages in thread
From: Stephane Eranian @ 2011-05-12 22:15 UTC (permalink / raw)
To: Dave Jones, Ingo Molnar, Stephane Eranian, Linus Torvalds,
Arnaldo Carvalho de Melo, LKML
On Fri, May 13, 2011 at 12:07 AM, Dave Jones <davej@redhat.com> wrote:
> On Thu, May 12, 2011 at 11:50:23PM +0200, Ingo Molnar wrote:
>
> > Dunno, i would not couple them necessarily - certain users might still have
> > access to kernel symbols via some other channel - for example the System.map.
>
> That always made this security by obscurity feature seem pointless for the bulk
> of users to me. Given the majority are going to be running distro kernels,
> anyone can find those addresses easily no matter how hard we hide them on the
> running system.
> Unless we were somehow introduced randomness into where we unpack the kernel
> each boot, and using System.map as a table of offsets instead of absolute addresses.
>
Good point about System.map! Even if /proc/kallsyms contains zero
addresses, I can
still get them from /boot/System.map which is readable by everyone, I
think. It does
not contain the modules addresses, but you have the core functions, unless I am
somehow mistaken.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 19:58 ` Pekka Enberg
@ 2011-05-13 6:12 ` Kees Cook
2011-05-13 6:24 ` Pekka Enberg
0 siblings, 1 reply; 29+ messages in thread
From: Kees Cook @ 2011-05-13 6:12 UTC (permalink / raw)
To: Pekka Enberg
Cc: David Miller, davej, eranian, acme, linux-kernel, Linus Torvalds,
Ingo Molnar
Hi Pekka,
On Thu, May 12, 2011 at 10:58:53PM +0300, Pekka Enberg wrote:
> On Thu, May 12, 2011 at 10:01 PM, David Miller <davem@davemloft.net> wrote:
> > From: Dave Jones <davej@redhat.com>
> > Date: Thu, 12 May 2011 14:37:41 -0400
> >
> >> On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
> >> > I hate this too, and I think it's absolutely rediculous.
> >> >
> >> > Also, like you, I lost an entire afternoon trying to figure out why
> >> > this started happening.
> >> >
> >> > I wish we could revert this change.
> >>
> >> At least it can be permanently disabled..
> >>
> >> echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
> >
> > Regardless, what to do about all of the "perf is broken" reports?
>
> Lets revert the commit 9f36e2c448007b54851e7e4fa48da97d1477a175
> ("printk: use %pK for /proc/kallsyms and /proc/modules"), please! I
> too have been wondering what's going on with perf reporting insane
> symbols and this should definitely not be enabled by default.
No, reverting that is not the answer. If perf has a problem with the
kptr_restrict feature, it should just disable it in /proc/sys when it
runs and restore it when finished. Since our defaults should be secure
for the average user (who does not use perf), it's fine the way it
is. Anyone using perf can adjust this for their use-case (that is why
there is a /proc/sys tunable).
-Kees
--
Kees Cook
Ubuntu Security Team
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-13 6:12 ` Kees Cook
@ 2011-05-13 6:24 ` Pekka Enberg
0 siblings, 0 replies; 29+ messages in thread
From: Pekka Enberg @ 2011-05-13 6:24 UTC (permalink / raw)
To: Kees Cook
Cc: David Miller, davej, eranian, acme, linux-kernel, Linus Torvalds,
Ingo Molnar
Hi Kees,
On Fri, May 13, 2011 at 9:12 AM, Kees Cook <kees.cook@canonical.com> wrote:
> Hi Pekka,
>
> On Thu, May 12, 2011 at 10:58:53PM +0300, Pekka Enberg wrote:
>> On Thu, May 12, 2011 at 10:01 PM, David Miller <davem@davemloft.net> wrote:
>> > From: Dave Jones <davej@redhat.com>
>> > Date: Thu, 12 May 2011 14:37:41 -0400
>> >
>> >> On Thu, May 12, 2011 at 02:06:30PM -0400, David Miller wrote:
>> >> > I hate this too, and I think it's absolutely rediculous.
>> >> >
>> >> > Also, like you, I lost an entire afternoon trying to figure out why
>> >> > this started happening.
>> >> >
>> >> > I wish we could revert this change.
>> >>
>> >> At least it can be permanently disabled..
>> >>
>> >> echo kernel.kptr_restrict = 0 >> /etc/sysctl.conf
>> >
>> > Regardless, what to do about all of the "perf is broken" reports?
>>
>> Lets revert the commit 9f36e2c448007b54851e7e4fa48da97d1477a175
>> ("printk: use %pK for /proc/kallsyms and /proc/modules"), please! I
>> too have been wondering what's going on with perf reporting insane
>> symbols and this should definitely not be enabled by default.
>
> No, reverting that is not the answer. If perf has a problem with the
> kptr_restrict feature, it should just disable it in /proc/sys when it
> runs and restore it when finished. Since our defaults should be secure
> for the average user (who does not use perf), it's fine the way it
> is. Anyone using perf can adjust this for their use-case (that is why
> there is a /proc/sys tunable).
No, it's the other way around. See commit
411f05f123cbd7f8aa1edcae86970755a6e2a9d9 ("vsprintf: Turn
kptr_restrict off by default") in Linus' tree for details.
Pekka
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 22:07 ` Dave Jones
2011-05-12 22:15 ` Stephane Eranian
@ 2011-05-13 8:57 ` Ingo Molnar
2011-05-13 16:23 ` Andi Kleen
1 sibling, 1 reply; 29+ messages in thread
From: Ingo Molnar @ 2011-05-13 8:57 UTC (permalink / raw)
To: Dave Jones, Stephane Eranian, Linus Torvalds,
Arnaldo Carvalho de Melo, LKML
Cc: H. Peter Anvin, Thomas Gleixner, Arjan van de Ven
* Dave Jones <davej@redhat.com> wrote:
> On Thu, May 12, 2011 at 11:50:23PM +0200, Ingo Molnar wrote:
>
> > Dunno, i would not couple them necessarily - certain users might still have
> > access to kernel symbols via some other channel - for example the System.map.
>
> That always made this security by obscurity feature seem pointless for the bulk
> of users to me. Given the majority are going to be running distro kernels,
> anyone can find those addresses easily no matter how hard we hide them on the
> running system.
I certainly agree and made that argument as well, in the original thread(s)
about /proc/kallsyms.
> Unless we were somehow introduced randomness into where we unpack the kernel
> each boot, and using System.map as a table of offsets instead of absolute
> addresses.
Correct. This security feature is IMO only solving a tiny fraction of the
problem and is thus in fact hindering the implementation of a *real* layer
of protection of kernel absolute addresses:
The x86 kernel is relocatable, so slightly randomizing the position of the
kernel would be feasible with no overhead on the vast majority of exising
distro installs, with just an updated kernel.
When exposing randomized RIPs to user-space we could recalculate all RIPs back
to the 0xffffffff80000000 base, so oopses would have the usual non-randomized
form:
[ 32.946003] IP: [<ffffffff80222521>] get_cur_val+0xcc/0x106
[ 32.946003] PGD 0
[ 32.946003] Oops: 0002 [#1] SMP DEBUG_PAGEALLOC
[ 32.946003] last sysfs file:
[ 32.946003] CPU 1
[ 32.946003] Pid: 1, comm: swapper Tainted: G W 2.6.29-rc1-00190-g37a76bd #10
[ 32.946003] RIP: 0010:[<ffffffff80222521>] [<ffffffff80222521>] get_cur_val+0xcc/0x106
[ 32.946003] RSP: 0018:ffff88003f977b80 EFLAGS: 00010202
[ 32.946003] RAX: 0000000000000001 RBX: ffff8800029c8c80 RCX: 0000000000000008
[ 32.946003] RDX: 0000000000000000 RSI: ffffffff80ce0100 RDI: 0000000000000000
[ 32.946003] RBP: ffff88003f977bd0 R08: 0000000000000004 R09: 0000000000000040
[ 32.946003] R10: 0000000000000060 R11: 0000000081363fa8 R12: ffffffff81c4f0e0
[ 32.946003] R13: ffffffff80ce0100 R14: ffff88003c888a00 R15: 0000000000000000
[ 32.946003] FS: 0000000000000000(0000) GS:ffff88003f802c00(0000) knlGS:0000000000000000
[ 32.946003] CS: 0010 DS: 0018 ES: 0018 CR0: 000000008005003b
[ 32.946003] CR2: 0000000000000000 CR3: 0000000000201000 CR4: 00000000000006e0
[ 32.946003] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[ 32.946003] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[ 32.946003] Process swapper (pid: 1, threadinfo ffff88003f976000, task ffff88003f978000)
[ 32.946003] Stack:
Likewise, /proc/kallsyms could pass these addresses as well and the perf
call-chain code and other places that sample RIPs could easily convert them to
the constant address as well.
We'd still leak some information like the relative position of symbols from
each other (this can be useful to certain classes of attacks), but we could
pretty effectively hide the absolute location of the kernel - which is the most
valuable piece of information -.
Then the random base has to be protected: i.e. all information leaks of raw
kernel RIPs have to be plugged. The nice thing is that this will happen as
*bugfixes*: randomized RIPs will not be useful for anything, so any
tools/people who rely on them will notice it immediately.
I think *that* would be a maintainable and complete security feature to truly
hide the exact location of the kernel image. kptr_restrict is not.
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-12 22:15 ` Stephane Eranian
@ 2011-05-13 9:01 ` Ingo Molnar
0 siblings, 0 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-13 9:01 UTC (permalink / raw)
To: Stephane Eranian
Cc: Dave Jones, Linus Torvalds, Arnaldo Carvalho de Melo, LKML
* Stephane Eranian <eranian@google.com> wrote:
> Good point about System.map! Even if /proc/kallsyms contains zero addresses,
> I can still get them from /boot/System.map which is readable by everyone, I
> think. It does not contain the modules addresses, but you have the core
> functions, unless I am somehow mistaken.
Yes. I pointed out this and some other details a couple of months ago, for a
similarly motivated /proc/kallsyms obfuscation patch:
http://lkml.org/lkml/2010/11/4/113
http://lkml.org/lkml/2010/11/4/145
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-13 8:57 ` Ingo Molnar
@ 2011-05-13 16:23 ` Andi Kleen
2011-05-17 12:17 ` Ingo Molnar
0 siblings, 1 reply; 29+ messages in thread
From: Andi Kleen @ 2011-05-13 16:23 UTC (permalink / raw)
To: Ingo Molnar
Cc: Dave Jones, Stephane Eranian, Linus Torvalds,
Arnaldo Carvalho de Melo, LKML, H. Peter Anvin, Thomas Gleixner,
Arjan van de Ven
Ingo Molnar <mingo@elte.hu> writes:
I agree that the current %kP default is really a catastrophe, clearly
on the trajectory of "the system is only secure when nothing works anymore"
> The x86 kernel is relocatable, so slightly randomizing the position of the
> kernel would be feasible with no overhead on the vast majority of exising
> distro installs, with just an updated kernel.
Problem is that we don't have a source of secure randomness early on
when the relocation would need to happen.
You could either pass it as an option, but that option would be right
now too exposed, or just use kexec and boot twice.
But all of this has drawbacks.
> When exposing randomized RIPs to user-space we could recalculate all RIPs back
> to the 0xffffffff80000000 base, so oopses would have the usual non-randomized
> form:
This would be very confusing because the register and stack contents
would have the non relocated addresses.
I bet it would cause a lot of similar problems as the current %kP
madness, just more subtle ones.
-Andi
--
ak@linux.intel.com -- Speaking for myself only
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [BUG] perf: bogus correlation of kernel symbols
2011-05-13 16:23 ` Andi Kleen
@ 2011-05-17 12:17 ` Ingo Molnar
0 siblings, 0 replies; 29+ messages in thread
From: Ingo Molnar @ 2011-05-17 12:17 UTC (permalink / raw)
To: Andi Kleen
Cc: Dave Jones, Stephane Eranian, Linus Torvalds,
Arnaldo Carvalho de Melo, LKML, H. Peter Anvin, Thomas Gleixner,
Arjan van de Ven
* Andi Kleen <andi@firstfloor.org> wrote:
> > The x86 kernel is relocatable, so slightly randomizing the position of the
> > kernel would be feasible with no overhead on the vast majority of exising
> > distro installs, with just an updated kernel.
>
> Problem is that we don't have a source of secure randomness early on when the
> relocation would need to happen.
That's indeed a problem but not a fundamental one: we can read out the current
time (RTC CMOS is always available on most systems), mix it with the current
cycle counter value and PRNG mix it.
It could only be recovered if the attacker is local to that box, guesses the
precise cycle count on that specific hardware (and hardware has small thermal
variations) and knows the precise boot time to the second as present in the
RTC.
Note that the amount of randomization would be small to begin with: if we have
only 3 bits of randomization and can thus make ~90% of kernel exploit attempts
crash statistically then we have most of the advantages already.
[ For the really paranoid we could add a new flag to the boot protocol and
embedd a random seed in the bzImage. This could be re-set upon installation
of a new kernel package, so on the next reboot the system gains a unique
seed.
Or we could add a boot parameter to seed things and cut this particular boot
parameter from all output like /proc/cmdline or the syslog command line
parameters printout. /etc/grub.conf is already inaccessible to unprivileged
userspace on most distros so the parameter is hidden. ]
So it's a solvable problem.
> You could either pass it as an option, but that option would be right now too
> exposed, or just use kexec and boot twice.
>
> But all of this has drawbacks.
>
> > When exposing randomized RIPs to user-space we could recalculate all RIPs back
> > to the 0xffffffff80000000 base, so oopses would have the usual non-randomized
> > form:
>
> This would be very confusing because the register and stack contents
> would have the non relocated addresses.
Well, kernel crashes can expose security relevant details anyway so they better
be hidden from unprivileged attackers anyway, the important thing is to
properly decode the symbols. We can keep the rest of the oops in its raw form
(and thus expose the seed to a privileged user - which we'd do anyway), being
dependable is important for oopses.
> I bet it would cause a lot of similar problems as the current %kP madness,
> just more subtle ones.
Well, did you expect me to react to your claim of 'subtle issues'?
If yes (which i assume) then why didn't you outline what you meant with that in
more detail, why are you forcing me to ask you what you mean precisely?
Thanks,
Ingo
^ permalink raw reply [flat|nested] 29+ messages in thread
end of thread, other threads:[~2011-05-17 12:18 UTC | newest]
Thread overview: 29+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2011-05-12 14:48 [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
2011-05-12 18:06 ` David Miller
2011-05-12 18:37 ` Dave Jones
2011-05-12 19:01 ` David Miller
2011-05-12 19:58 ` Pekka Enberg
2011-05-13 6:12 ` Kees Cook
2011-05-13 6:24 ` Pekka Enberg
2011-05-12 20:24 ` Alexey Dobriyan
2011-05-12 21:06 ` Ingo Molnar
2011-05-12 20:31 ` Linus Torvalds
2011-05-12 20:43 ` David Miller
2011-05-12 21:00 ` [PATCH] vsprintf: Turn kptr_restrict off by default Ingo Molnar
2011-05-12 21:08 ` David Miller
2011-05-12 21:07 ` [BUG] perf: bogus correlation of kernel symbols Stephane Eranian
2011-05-12 21:30 ` Stephane Eranian
2011-05-12 21:35 ` Ingo Molnar
2011-05-12 21:38 ` Stephane Eranian
2011-05-12 21:50 ` Ingo Molnar
2011-05-12 21:56 ` Stephane Eranian
2011-05-12 22:00 ` Ingo Molnar
2011-05-12 22:07 ` Dave Jones
2011-05-12 22:15 ` Stephane Eranian
2011-05-13 9:01 ` Ingo Molnar
2011-05-13 8:57 ` Ingo Molnar
2011-05-13 16:23 ` Andi Kleen
2011-05-17 12:17 ` Ingo Molnar
2011-05-12 21:36 ` Ingo Molnar
2011-05-12 21:41 ` Stephane Eranian
2011-05-12 21:54 ` Ingo Molnar
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®