mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Re: [PATCH] ftrace based hard lockup detector
       [not found] <4973cff5.05a0660a.650c.2f54@mx.google.com>
@ 2009-01-19 12:59 ` Steven Rostedt
  2009-01-19 13:04   ` Ingo Molnar
  0 siblings, 1 reply; 5+ messages in thread
From: Steven Rostedt @ 2009-01-19 12:59 UTC (permalink / raw)
  To: Frederic Weisbecker; +Cc: Ingo Molnar, linux-kernel



On Sun, 18 Jan 2009, Frederic Weisbecker wrote:

> Like the NMI watchdog, this feature try to detect hard lockups by
> lurking at the non-progress of the timer interrupts.
> 
> You can enable it at boot time by passing the ftrace_hardlockup parameter.
> I plan to add a debugfs file to enable/disable at runtime.
> 
> When a hardlockup is detected, it will print a backtrace. Perhaps it
> would be good to print the locks held from lockdep too?
> 
> It only support x86 for the moment, because a kind of generic timer interrupt
> counter is needed on all archs to have it generic.
> 
> Signed-off-by: Frederic Weisbecker <fweisbec@gmail.com>

Hi Frederic,

This seems like a rewrite of the NMI lockup code. In my debugging, I 
simply put ftrace_dump in the NMI lockup, which gives me a ftrace dump as 
soon as NMI detects a lockup. I'm a bit confused at what this gives us 
over that?

-- Steve

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] ftrace based hard lockup detector
  2009-01-19 12:59 ` [PATCH] ftrace based hard lockup detector Steven Rostedt
@ 2009-01-19 13:04   ` Ingo Molnar
  2009-01-19 13:18     ` Steven Rostedt
  0 siblings, 1 reply; 5+ messages in thread
From: Ingo Molnar @ 2009-01-19 13:04 UTC (permalink / raw)
  To: Steven Rostedt; +Cc: Frederic Weisbecker, linux-kernel


* Steven Rostedt <rostedt@goodmis.org> wrote:

> On Sun, 18 Jan 2009, Frederic Weisbecker wrote:
> 
> > Like the NMI watchdog, this feature try to detect hard lockups by
> > lurking at the non-progress of the timer interrupts.
> > 
> > You can enable it at boot time by passing the ftrace_hardlockup parameter.
> > I plan to add a debugfs file to enable/disable at runtime.
> > 
> > When a hardlockup is detected, it will print a backtrace. Perhaps it
> > would be good to print the locks held from lockdep too?
> > 
> > It only support x86 for the moment, because a kind of generic timer interrupt
> > counter is needed on all archs to have it generic.
> > 
> > Signed-off-by: Frederic Weisbecker <fweisbec@gmail.com>
> 
> Hi Frederic,
> 
> This seems like a rewrite of the NMI lockup code. In my debugging, I 
> simply put ftrace_dump in the NMI lockup, which gives me a ftrace dump 
> as soon as NMI detects a lockup. I'm a bit confused at what this gives 
> us over that?

this is different from the NMI watchdog in a number of ways:

 - it works on all platforms and in all situations where the NMI watchdog 
   does not work.

 - in theory it can detect hard lockups in situations where the NMI 
   watchdog is disabled, such as suspend/resume or early bootup. 
   (especially early bootup lockups are nasty and the NMI watchdog is 
    enabled relatively late)

 - it could be extended to detect 'soft' lockups too - i.e. we could have 
   a one-stop facility to detect all kinds of "kernel does not seem to 
   progress" lockups.

But it's not as complete as the NMI watchdog: it relies on instrumented 
function calls rolling on and on during the lockup - that's not the case 
when we get a hard lockup due to a tight, infinite loop somewhere.

	Ingo

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] ftrace based hard lockup detector
  2009-01-19 13:04   ` Ingo Molnar
@ 2009-01-19 13:18     ` Steven Rostedt
  2009-01-19 13:24       ` Ingo Molnar
  2009-01-19 13:25       ` Frédéric Weisbecker
  0 siblings, 2 replies; 5+ messages in thread
From: Steven Rostedt @ 2009-01-19 13:18 UTC (permalink / raw)
  To: Ingo Molnar; +Cc: Frederic Weisbecker, linux-kernel


On Mon, 19 Jan 2009, Ingo Molnar wrote:

> 
> * Steven Rostedt <rostedt@goodmis.org> wrote:
> 
> > On Sun, 18 Jan 2009, Frederic Weisbecker wrote:
> > 
> > > Like the NMI watchdog, this feature try to detect hard lockups by
> > > lurking at the non-progress of the timer interrupts.
> > > 
> > > You can enable it at boot time by passing the ftrace_hardlockup parameter.
> > > I plan to add a debugfs file to enable/disable at runtime.
> > > 
> > > When a hardlockup is detected, it will print a backtrace. Perhaps it
> > > would be good to print the locks held from lockdep too?
> > > 
> > > It only support x86 for the moment, because a kind of generic timer interrupt
> > > counter is needed on all archs to have it generic.
> > > 
> > > Signed-off-by: Frederic Weisbecker <fweisbec@gmail.com>
> > 
> > Hi Frederic,
> > 
> > This seems like a rewrite of the NMI lockup code. In my debugging, I 
> > simply put ftrace_dump in the NMI lockup, which gives me a ftrace dump 
> > as soon as NMI detects a lockup. I'm a bit confused at what this gives 
> > us over that?
> 
> this is different from the NMI watchdog in a number of ways:
> 
>  - it works on all platforms and in all situations where the NMI watchdog 
>    does not work.
> 
>  - in theory it can detect hard lockups in situations where the NMI 
>    watchdog is disabled, such as suspend/resume or early bootup. 
>    (especially early bootup lockups are nasty and the NMI watchdog is 
>     enabled relatively late)
> 
>  - it could be extended to detect 'soft' lockups too - i.e. we could have 
>    a one-stop facility to detect all kinds of "kernel does not seem to 
>    progress" lockups.
> 
> But it's not as complete as the NMI watchdog: it relies on instrumented 
> function calls rolling on and on during the lockup - that's not the case 
> when we get a hard lockup due to a tight, infinite loop somewhere.

Ah, OK, the check is in the function tracer. Hmm, my logdev code had an 
option to enable tracing at early bootup. Instead of using the normal 
memory alloction for the ring buffer, it needed to use alloc_bootmem. I 
wonder if it would be worth it to allow for a tracer to do the same if it 
needs to be allocated early on (before memory is initialized)?

--Steve


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] ftrace based hard lockup detector
  2009-01-19 13:18     ` Steven Rostedt
@ 2009-01-19 13:24       ` Ingo Molnar
  2009-01-19 13:25       ` Frédéric Weisbecker
  1 sibling, 0 replies; 5+ messages in thread
From: Ingo Molnar @ 2009-01-19 13:24 UTC (permalink / raw)
  To: Steven Rostedt; +Cc: Frederic Weisbecker, linux-kernel


* Steven Rostedt <rostedt@goodmis.org> wrote:

> 
> On Mon, 19 Jan 2009, Ingo Molnar wrote:
> 
> > 
> > * Steven Rostedt <rostedt@goodmis.org> wrote:
> > 
> > > On Sun, 18 Jan 2009, Frederic Weisbecker wrote:
> > > 
> > > > Like the NMI watchdog, this feature try to detect hard lockups by
> > > > lurking at the non-progress of the timer interrupts.
> > > > 
> > > > You can enable it at boot time by passing the ftrace_hardlockup parameter.
> > > > I plan to add a debugfs file to enable/disable at runtime.
> > > > 
> > > > When a hardlockup is detected, it will print a backtrace. Perhaps it
> > > > would be good to print the locks held from lockdep too?
> > > > 
> > > > It only support x86 for the moment, because a kind of generic timer interrupt
> > > > counter is needed on all archs to have it generic.
> > > > 
> > > > Signed-off-by: Frederic Weisbecker <fweisbec@gmail.com>
> > > 
> > > Hi Frederic,
> > > 
> > > This seems like a rewrite of the NMI lockup code. In my debugging, I 
> > > simply put ftrace_dump in the NMI lockup, which gives me a ftrace dump 
> > > as soon as NMI detects a lockup. I'm a bit confused at what this gives 
> > > us over that?
> > 
> > this is different from the NMI watchdog in a number of ways:
> > 
> >  - it works on all platforms and in all situations where the NMI watchdog 
> >    does not work.
> > 
> >  - in theory it can detect hard lockups in situations where the NMI 
> >    watchdog is disabled, such as suspend/resume or early bootup. 
> >    (especially early bootup lockups are nasty and the NMI watchdog is 
> >     enabled relatively late)
> > 
> >  - it could be extended to detect 'soft' lockups too - i.e. we could have 
> >    a one-stop facility to detect all kinds of "kernel does not seem to 
> >    progress" lockups.
> > 
> > But it's not as complete as the NMI watchdog: it relies on instrumented 
> > function calls rolling on and on during the lockup - that's not the case 
> > when we get a hard lockup due to a tight, infinite loop somewhere.
> 
> Ah, OK, the check is in the function tracer. Hmm, my logdev code had an 
> option to enable tracing at early bootup. Instead of using the normal 
> memory alloction for the ring buffer, it needed to use alloc_bootmem. I 
> wonder if it would be worth it to allow for a tracer to do the same if 
> it needs to be allocated early on (before memory is initialized)?

Yes, very much so!

Especially when using things like dump_trace options, this would be handy.

	Ingo

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] ftrace based hard lockup detector
  2009-01-19 13:18     ` Steven Rostedt
  2009-01-19 13:24       ` Ingo Molnar
@ 2009-01-19 13:25       ` Frédéric Weisbecker
  1 sibling, 0 replies; 5+ messages in thread
From: Frédéric Weisbecker @ 2009-01-19 13:25 UTC (permalink / raw)
  To: Steven Rostedt; +Cc: Ingo Molnar, linux-kernel

2009/1/19 Steven Rostedt <rostedt@goodmis.org>:
> Ah, OK, the check is in the function tracer. Hmm, my logdev code had an
> option to enable tracing at early bootup. Instead of using the normal
> memory alloction for the ring buffer, it needed to use alloc_bootmem. I
> wonder if it would be worth it to allow for a tracer to do the same if it
> needs to be allocated early on (before memory is initialized)?
>
> --Steve


Hi Steve,

Yeah I think that would be a good idea.


BTW, tote that for the moment, this hard lockup detector is basic, but
I could even make it relying
to the function graph tracer and then dump the whole stack of
functions called. I think that could
give a more precise backtrace.... (and more human readable).

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2009-01-19 13:26 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
     [not found] <4973cff5.05a0660a.650c.2f54@mx.google.com>
2009-01-19 12:59 ` [PATCH] ftrace based hard lockup detector Steven Rostedt
2009-01-19 13:04   ` Ingo Molnar
2009-01-19 13:18     ` Steven Rostedt
2009-01-19 13:24       ` Ingo Molnar
2009-01-19 13:25       ` Frédéric Weisbecker

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®