From: Randy Dunlap <randy.dunlap@oracle.com>
To: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Kevin Winchester <kjwinchester@gmail.com>,
"J. Bruce Fields" <bfields@fieldses.org>,
Al Viro <viro@ZenIV.linux.org.uk>,
Arjan van de Ven <arjan@linux.intel.com>,
Linux Kernel Mailing List <linux-kernel@vger.kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
NetDev <netdev@vger.kernel.org>
Subject: Re: Top 10 kernel oopses for the week ending January 5th, 2008
Date: Tue, 8 Jan 2008 08:14:01 -0800 [thread overview]
Message-ID: <20080108081401.d9576ac5.randy.dunlap@oracle.com> (raw)
In-Reply-To: <alpine.LFD.1.00.0801071851120.3148@woody.linux-foundation.org>
On Mon, 7 Jan 2008 19:26:12 -0800 (PST) Linus Torvalds wrote:
> On Mon, 7 Jan 2008, Kevin Winchester wrote:
>
> > J. Bruce Fields wrote:
> > >
> > > Is there any good basic documentation on this to point people at?
> >
> > I would second this question. I see people "decode" oops on lkml often
> > enough, but I've never been entirely sure how its done. Is it somewhere
> > in Documentation?
>
> It's actually not necessarily at all that trivial, unless you have a deep
> understanding of the code generated for the architecture in question (and
> even then, some oopses take more time to figure out than others, thanks
> to inlining and tailcalls etc).
>
> If the oops happened with a kernel you generated yourself, it's usually
> rather easy. Especially if you said "y" to the "generate debugging info"
> question at configuration time. Because, in that case, you really just do
> a simple
>
> gdb vmlinux
>
> and then you can do (for example) something like setting a breakpoint at
> the EIP that was reported for the oops, and it will tell you what line it
> came from.
>
> However, if you don't have the exact binary - which is the common case for
> random oopses reported on lkml - you will generally have to disassemble
> the hex sequence given in the oops (the "Code:" line), and try to match it
> up against the source code to try to figure out what is going on.
>
> Even just the disassembly is not entirely trivial, since the oops will
> give you the eip that it happened at, but you often want to also
> disassemble *backwards* in order to get more of a context (the "Code:"
> line will mark the particular EIP that starts the oopsing instruction by
> enclosing it in <xx>, but with non-constant instruction lengths, you need
> to use a bit of trial-and-error to figure it out.
>
> I usually just compile a small program like
>
> const char array[]="\xnn\xnn\xnn...";
>
> int main(int argc, char **argv)
> {
> printf("%p\n", array);
> *(int *)0=0;
> }
>
> and run it under gdb, and then when it gets the SIGSEGV (due to the
> obvious NULL pointer dereference), I can just ask gdb to disassemble
> around the array that contains the code[] stuff. Try a few offsets, to see
> when the disassembly makes sense (and gives the reported EIP as the
> beginning of one of the disassembled instructions).
>
> (You can do it other and smarter ways too, I'm not claiming that's a
> particularly good way to do it, and the old "ksymoops" program used to do
> a pretty good job of this, but I'm used to that particular idiotic way
> myself, since it's how I've basically always done it)
One other way to do it (at least for x86-32/64) is to use
$kerneltree/scripts/decodecode. It may work on other $arches also,
but I haven't tested it on others.
---
~Randy
next prev parent reply other threads:[~2008-01-08 16:15 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2008-01-05 21:06 Arjan van de Ven
2008-01-05 21:26 ` Al Viro
2008-01-05 21:39 ` Al Viro
2008-01-07 17:44 ` J. Bruce Fields
2008-01-08 1:19 ` Kevin Winchester
2008-01-08 3:26 ` Linus Torvalds
2008-01-08 5:59 ` Al Viro
2008-01-08 7:33 ` Jarek Poplawski
2008-01-10 4:13 ` Neil Brown
2008-01-10 5:53 ` Al Viro
2008-01-14 1:36 ` Neil Brown
2008-01-08 16:14 ` Randy Dunlap [this message]
2008-01-08 17:42 ` Arjan van de Ven
2008-01-08 18:08 ` Linus Torvalds
2008-01-08 18:16 ` Arjan van de Ven
2008-01-08 18:27 ` Linus Torvalds
2008-01-08 19:05 ` Arjan van de Ven
2008-01-08 19:31 ` Linus Torvalds
2008-01-08 22:56 ` Rafael J. Wysocki
2008-01-08 17:08 ` Andi Kleen
2008-01-06 3:30 ` Andi Kleen
2008-01-06 3:31 ` Arjan van de Ven
2008-01-06 3:50 ` Andi Kleen
2008-01-09 14:12 ` Johannes Berg
2008-01-09 15:28 ` Arjan van de Ven
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20080108081401.d9576ac5.randy.dunlap@oracle.com \
--to=randy.dunlap@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=arjan@linux.intel.com \
--cc=bfields@fieldses.org \
--cc=kjwinchester@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=torvalds@linux-foundation.org \
--cc=viro@ZenIV.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome