From: "jeff millar" <wa1hco@adelphia.net>
To: "Bill Davidsen" <davidsen@tmr.com>,
"Jakob Oestergaard" <jakob@unthought.net>
Cc: "Kernel mailing list" <linux-kernel@vger.kernel.org>
Subject: Re: RAID problems
Date: Tue, 30 Jul 2002 23:18:58 -0400 [thread overview]
Message-ID: <004501c23841$03265a30$6a01a8c0@wa1hco> (raw)
In-Reply-To: <Pine.LNX.3.96.1020730223102.6974A-100000@gatekeeper.tmr.com>
----- Original Message -----
From: "Bill Davidsen" <davidsen@tmr.com>
To: "Jakob Oestergaard" <jakob@unthought.net>
Cc: "Kernel mailing list" <linux-kernel@vger.kernel.org>
Sent: Tuesday, July 30, 2002 10:46 PM
Subject: Re: RAID problems
> On Tue, 30 Jul 2002, Jakob Oestergaard wrote:
>
> > On Tue, Jul 30, 2002 at 01:25:37PM -0400, Bill Davidsen wrote:
>
> > > Why does no one seem surprised that one drive failed and the system
marked
> > > four bad instead of using the spare in the first place? That's a more
> > > interesting question.
> >
> > As stated in the above URL (no shit, really!), this can happen for
> > example if some device hangs the SCSI bus.
>
> I think you misread my comment, it was not "why doesn't the documentation
> say this" but rather "why does software RAID have this problem?" I know
> this can happen in theory, but it seems that the docs imply that this
> isn't a surprise in practice. I've been running systems with SCSI RAID and
> hardware controllers since 1989 (or maybe 1990), I've got EMC and HDS
> boxes, aaraid and ipc controllers, Promise and CMC(??) on system boards,
> and a number of systems running software RAID. And of all the drive
> failures I've had, exactly one had more than one drive fail, and that was
> in a power {something} issue which took out multiple drives and system
> boards on many systems.
>
> I just surprised that the software RAID doesn't have better luck with
> this, I don't see any magic other than maybe a bus reset the firmware
> would be doing, and I'm wondering why this seems to be common with Linux.
> Or am I misreading the frequency with which it happens?
>
> > Did *anyone* read that section ?!? ;)
> >
> > If someone feels the explanation there could be better, please just send
> > the better explanation to me and it will get in. Really, this section
> > is one of the few sections that I did *not* update in the HOWTO, because
> > I really felt that it was still both adequate (since no-one has demanded
> > elaboration) and correct.
>
> Thye words are clear, I'm surprised at the behaviour. Yes, I know that's
> not your thing.
>
> --
> bill davidsen <davidsen@tmr.com>
> CTO, TMR Associates, Inc
> Doing interesting things with little computers since 1979.
In the 3 weeks since installing Linux software raid-5 (3 x 80 GB), I had to
reinitialize the raid devices twice (mkraid --force)...once due to an abrupt
shutdown and once due to some weird ATA/ATAPI/drive problem that caused a
disk to begin "clicking" spasmodically...and left the raid array all out of
whack..
Linux software raid seems very fragile and very scary to recover. I feel a
much stronger need for backup with raid than without it.
Raid needs an automatic way to maintain device synchronization. Why should
I have to...
manually examine the device data (lsraid)
find two devices that match
mark the others failed in /etc/raidtab
reinitialize the raid devices...putting all data at risk
hot add the "failed" device
wait for it to recover (hours)
change /etc/raidtab again
retest everything
This is 10 times worse that e2fsck and much more error prone. The file
system guru's worked hard on journalling to minimize this kind of risk.
jeff
next prev parent reply other threads:[~2002-07-31 3:15 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2002-07-25 11:54 Roy Sigurd Karlsbakk
2002-07-25 12:03 ` Neil Brown
2002-07-25 12:11 ` Jakob Oestergaard
2002-07-30 17:25 ` Bill Davidsen
2002-07-30 17:55 ` Jakob Oestergaard
2002-07-31 2:46 ` Bill Davidsen
2002-07-31 3:18 ` jeff millar [this message]
2002-07-31 3:32 ` Neil Brown
2002-07-31 14:54 ` Bill Davidsen
2002-07-31 13:35 ` Eyal Lebedinsky
2002-07-31 18:21 ` Bill Davidsen
2002-07-31 3:42 ` Jakob Oestergaard
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to='004501c23841$03265a30$6a01a8c0@wa1hco' \
--to=wa1hco@adelphia.net \
--cc=davidsen@tmr.com \
--cc=jakob@unthought.net \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®