mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Linas Vepstas" <linasvepstas@gmail.com>
To: "Alan Cox" <alan@lxorguk.ukuu.org.uk>,
	"Martin K. Petersen" <martin.petersen@oracle.com>
Cc: "John Stoffel" <john@stoffel.org>,
	"Alistair John Strachan" <alistair@devzero.co.uk>,
	linux-kernel@vger.kernel.org
Subject: Re: amd64 sata_nv (massive) memory corruption
Date: Wed, 6 Aug 2008 16:33:04 -0500	[thread overview]
Message-ID: <3ae3aa420808061433i3d90c3dcgfb40d953da2941c8@mail.gmail.com> (raw)
In-Reply-To: <20080805182119.75913fa3@lxorguk.ukuu.org.uk>

2008/8/5 Alan Cox <alan@lxorguk.ukuu.org.uk>:

>> I'm game. Care to guide me through?  So: on every write, this
>> new device mapper module computes a checksum and stores
>> it somewhere. On every read, it computes a checksum and
>> compares to the stored value. Easy enough I guess.
>>
>> Several hard parts:
>> -- where to store the checksums?
>
> That is the million dollar question - plus you can argue it is the fs
> that should do it. There is stuff crawling through the standards world to
> provide a small per block additional info area on disk sectors.

My objection to fs-layer checksums (e.g. in some user-space
file system) is that it  doesn't leverage the extra info that RAID
has.  If a block is bad, RAID can probably fetch another one
that is good. You can't do this at the file-system level.

I assume I can layer device-mappers anywhere, right?
Layering one *underneath* md-raid would allow it to
reject/discard bad blocks, and then let the raid layer
try to find a good block somewhere else.

I assume that a device mapper can alter the number
of blocks-in to the number of blocks-out; that it doesn't
have to be 1-1. Then for every 10 sectors of data, it
would use 11 sectors of storage, one holding the
checksum.  I'm very naive about how the block layer
works, so I don't know what snags there might be.

The downside of this is that the disk wouldn't be
naively readable unless the specific mapper module
was in place -- so one would need a superblock of
some sort indicating the type of checksumming used,
etc.  Is there any "standardized" way of managing
superblocks for use by the device mapper?  I guess
the encrypting dm has to store meta-information
somewhere, too, specifying what kind of encryption
was used.  I'll look at that.

> Yes. If you can figure out where to keep the checksums without ruining
> performance

Heh. Unlikely. The act of checksumming will impact
performance. It should end up similar to the impact
from encryption (maybe not quite as bad), or comparable
to raid-5 (which computes various kinds of parity).

> (and of course if there isn't one lurking in device mapper
> world not yet submitted).

I'm googling, but I don't see anything.  However, I now see,
for the first time,   pending workd for 2.6.27 for a field in bio
called  "blk_integrity". I cannot figure out if this work requires
special-whiz-bang disk drives to be purchased.

Also, it seems to be limited to 8 bytes of checksums per 512
byte block? This is reasonable for checksumming, I guess,
but one could get even fancier and run ECC-type sums, if
one could store, say, an addtional 50 bytes for every 512
bytes. I'm cc'ing Martin Petersen, the developer, for
comments.


--linas

  reply	other threads:[~2008-08-06 21:39 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2008-08-01 17:30 Linas Vepstas
2008-08-01 20:51 ` John Stoffel
2008-08-02  3:06   ` Linas Vepstas
2008-08-01 22:19 ` Alistair John Strachan
2008-08-02  2:51   ` Linas Vepstas
2008-08-02 20:09     ` John Stoffel
2008-08-02 22:01       ` Linas Vepstas
2008-08-03  2:41         ` John Stoffel
2008-08-03 22:23           ` Linas Vepstas
2008-08-03 22:16             ` Alan Cox
2008-08-05 17:02               ` Linas Vepstas
2008-08-05 17:21                 ` Alan Cox
2008-08-06 21:33                   ` Linas Vepstas [this message]
2008-08-07  2:59                     ` Martin K. Petersen
2008-08-07  4:32                       ` Linas Vepstas
2008-08-07 16:42                         ` Martin K. Petersen
2008-08-07 17:23                           ` Linas Vepstas
2008-08-07 18:53                           ` John Stoffel
2008-08-07  7:45                     ` Pavel Machek
2008-08-02 21:55     ` Roger Heflin
     [not found] <fa.qB5d+HsAJ6G05jNoeU8Q9GV6Dow@ifi.uio.no>
     [not found] ` <fa.fxlDAHxOnGgcBiOH/EOauE67ZPc@ifi.uio.no>
     [not found]   ` <fa.1WYUmN6FHR5yW+sXoYRFN22Y8S8@ifi.uio.no>
     [not found]     ` <fa.LAUkvEUlYiF69V/F8F3wigxqH9w@ifi.uio.no>
     [not found]       ` <fa.mXeFXYNkfZfUYPQcGwzok0IOIfY@ifi.uio.no>
     [not found]         ` <fa.KjbvCGbUr2JeQTcwA1/sFGIIMik@ifi.uio.no>
2008-08-04  3:22           ` Robert Hancock
2008-08-05  5:29             ` Linas Vepstas
2008-08-05  6:36               ` Robert Hancock
2008-08-05 12:29               ` Alan Cox

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=3ae3aa420808061433i3d90c3dcgfb40d953da2941c8@mail.gmail.com \
    --to=linasvepstas@gmail.com \
    --cc=alan@lxorguk.ukuu.org.uk \
    --cc=alistair@devzero.co.uk \
    --cc=john@stoffel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=martin.petersen@oracle.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®