From: "Ray Lee" <ray-lk@madrabbit.org>
To: "John Sigler" <linux.kernel@free.fr>
Cc: linux-kernel@vger.kernel.org, linux-pci@atrey.karlin.mff.cuni.cz,
greg@kroah.com, grundler@parisc-linux.org
Subject: Re: How to debug complete kernel lock-ups
Date: Wed, 31 Oct 2007 14:28:11 -0700 [thread overview]
Message-ID: <2c0942db0710311428i7675a4b6saf3f79dc60a4f0be@mail.gmail.com> (raw)
In-Reply-To: <47284A21.1070901@free.fr>
On 10/31/07, John Sigler <linux.kernel@free.fr> wrote:
> "It seems that the PCI clock on this system has a rather large over- and
> undershoot and we suspect that the undershoot (of ~1V) is causing a drop
> in the core voltage of the on-board FPGA which results in lockup of the
> firmware. Both the under- and overshoot are well outside the allowed
> ranges (high=VCC+0.5V and low=-0.5V) of the PCI specification and a
> premature conclusion might be that the system does not comply to the PCI
> spec and that this is the cause of the lockup on this PC."
>
> This is waaay out of my league, as my area is software.
>
> Is it typical for voltage issues to hang hardware?
Yes, if the voltage is applied (or lacking) at the right place.
> Is it typical for one PCI board locking up to nail the entire system?
This doesn't appear to be a case of the *board* crashing, but rather
the board taking the pci bus and related hardware on-motherboard down
with it. Once that's down, anything that you need that goes through
the bus (on a PC, that's pretty much everything), is inaccessible.
> I don't understand why the lockup would only happen when I write to the
> 4 ports within a small time frame, and not when I only write to 2 ports
> (either one port on each card, or 2 ports on the same card). I suspected
> some kind of concurrency issue...
No, given the hardware guy's description, it's a power issue. Perhaps
when you're writing to a port, you're using more power on the card?
Four ports = 4 * the power draw. When the current load increases,
voltage drops, and if you underpower a chip, it's going to lose its
little head.
> I suppose the next logical step is to get the board's engineers
> and the system's engineers duke it out? :-)
Yes, all signs point to it being a pure hardware issue. You may be
able to work around it in software by initializing a 'counting
semaphore' to 2 to manage the maximum concurrency, so that you'll
never write more than 2 ports at a time until the hardware guys figure
it out.
Ray
prev parent reply other threads:[~2007-10-31 21:28 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2007-10-23 16:11 John Sigler
2007-10-24 9:17 ` John Sigler
2007-10-24 15:56 ` Greg KH
2007-10-24 16:19 ` Ray Lee
2007-10-25 4:06 ` Grant Grundler
2007-10-31 9:25 ` John Sigler
2007-10-31 21:28 ` Ray Lee [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2c0942db0710311428i7675a4b6saf3f79dc60a4f0be@mail.gmail.com \
--to=ray-lk@madrabbit.org \
--cc=greg@kroah.com \
--cc=grundler@parisc-linux.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@atrey.karlin.mff.cuni.cz \
--cc=linux.kernel@free.fr \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®