mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "jeff millar" <wa1hco@adelphia.net>
To: "Bill Davidsen" <davidsen@tmr.com>,
	"Jakob Oestergaard" <jakob@unthought.net>
Cc: "Kernel mailing list" <linux-kernel@vger.kernel.org>
Subject: Re: RAID problems
Date: Tue, 30 Jul 2002 23:18:58 -0400	[thread overview]
Message-ID: <004501c23841$03265a30$6a01a8c0@wa1hco> (raw)
In-Reply-To: <Pine.LNX.3.96.1020730223102.6974A-100000@gatekeeper.tmr.com>

----- Original Message -----
From: "Bill Davidsen" <davidsen@tmr.com>
To: "Jakob Oestergaard" <jakob@unthought.net>
Cc: "Kernel mailing list" <linux-kernel@vger.kernel.org>
Sent: Tuesday, July 30, 2002 10:46 PM
Subject: Re: RAID problems


> On Tue, 30 Jul 2002, Jakob Oestergaard wrote:
>
> > On Tue, Jul 30, 2002 at 01:25:37PM -0400, Bill Davidsen wrote:
>
> > > Why does no one seem surprised that one drive failed and the system
marked
> > > four bad instead of using the spare in the first place? That's a more
> > > interesting question.
> >
> > As stated in the above URL (no shit, really!), this can happen for
> > example if some device hangs the SCSI bus.
>
> I think you misread my comment, it was not "why doesn't the documentation
> say this" but rather "why does software RAID have this problem?" I know
> this can happen in theory, but it seems that the docs imply that this
> isn't a surprise in practice. I've been running systems with SCSI RAID and
> hardware controllers since 1989 (or maybe 1990), I've got EMC and HDS
> boxes, aaraid and ipc controllers, Promise and CMC(??) on system boards,
> and a number of systems running software RAID. And of all the drive
> failures I've had, exactly one had more than one drive fail, and that was
> in a power {something} issue which took out multiple drives and system
> boards on many systems.
>
> I just surprised that the software RAID doesn't have better luck with
> this, I don't see any magic other than maybe a bus reset the firmware
> would be doing, and I'm wondering why this seems to be common with Linux.
> Or am I misreading the frequency with which it happens?
>
> > Did *anyone* read that section ?!?   ;)
> >
> > If someone feels the explanation there could be better, please just send
> > the better explanation to me and it will get in.  Really, this section
> > is one of the few sections that I did *not* update in the HOWTO, because
> > I really felt that it was still both adequate (since no-one has demanded
> > elaboration) and correct.
>
> Thye words are clear, I'm surprised at the behaviour. Yes, I know that's
> not your thing.
>
> --
> bill davidsen <davidsen@tmr.com>
>   CTO, TMR Associates, Inc
> Doing interesting things with little computers since 1979.

In the 3 weeks since installing Linux software raid-5 (3 x 80 GB), I had to
reinitialize the raid devices twice (mkraid --force)...once due to an abrupt
shutdown and once due to some weird ATA/ATAPI/drive problem that caused a
disk to begin "clicking" spasmodically...and left the raid array all out of
whack..

Linux software raid seems very fragile and very scary to recover.  I feel a
much stronger need for backup with raid than without it.

Raid needs an automatic way to maintain device synchronization.  Why should
I have to...
    manually examine the device data (lsraid)
    find two devices that match
    mark the others failed in /etc/raidtab
    reinitialize the raid devices...putting all data at risk
    hot add the "failed" device
    wait for it to recover (hours)
    change /etc/raidtab again
    retest everything

This is 10 times worse that e2fsck and much more error prone.  The file
system guru's worked hard on journalling to minimize this kind of risk.

jeff


  reply	other threads:[~2002-07-31  3:15 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2002-07-25 11:54 Roy Sigurd Karlsbakk
2002-07-25 12:03 ` Neil Brown
2002-07-25 12:11 ` Jakob Oestergaard
2002-07-30 17:25   ` Bill Davidsen
2002-07-30 17:55     ` Jakob Oestergaard
2002-07-31  2:46       ` Bill Davidsen
2002-07-31  3:18         ` jeff millar [this message]
2002-07-31  3:32           ` Neil Brown
2002-07-31 14:54             ` Bill Davidsen
2002-07-31 13:35           ` Eyal Lebedinsky
2002-07-31 18:21           ` Bill Davidsen
2002-07-31  3:42         ` Jakob Oestergaard

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to='004501c23841$03265a30$6a01a8c0@wa1hco' \
    --to=wa1hco@adelphia.net \
    --cc=davidsen@tmr.com \
    --cc=jakob@unthought.net \
    --cc=linux-kernel@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®