mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Linus Torvalds <torvalds@linux-foundation.org>
To: Robert Hancock <hancockr@shaw.ca>
Cc: Pekka Enberg <penberg@cs.helsinki.fi>,
	Neil Romig <neil@romig.demon.co.uk>,
	linux-kernel@vger.kernel.org, hyoshiok@miraclelinux.com,
	Andrew Morton <akpm@linux-foundation.org>
Subject: Re: File corruption when using kernels 2.6.18+
Date: Wed, 3 Oct 2007 20:39:41 -0700 (PDT)	[thread overview]
Message-ID: <alpine.LFD.0.999.0710032008190.3579@woody.linux-foundation.org> (raw)
In-Reply-To: <47045725.1070900@shaw.ca>



On Wed, 3 Oct 2007, Robert Hancock wrote:
> 
> Erratum 97: 128-Bit Streaming Stores May Cause Coherency Failure

The Intel-optimized memcpy doesn't use the SSE registers, just regular 
32-bit integer nontemporal stores (movnti). The reason is that the SSE 
state save is too expensive to be worth it.

So it's not that. Also, considering that it was a single-bit error in all 
the cases I saw, I wouldn't expect it to be a cache coherency problem, 
which I'd expect to corrupt a whole cacheline or possibly at least a whole 
access.

That said, bit corruption can be just about anything. It's certainly not 
impossible that it's a CPU bug.

But my first guess would be slightly dodgy motherboard, possibly coupled 
with a chipset that simply isn't very tolerant to any timing errors. If 
the motherboard traces to the DDR aren't impedance-matched, or if the 
traces don't have the same length, or if the capacitors that are supposed 
to handle spikes in burst current aren't up to snuff, you'll just get 
noisy lines.

And at some point, noisy lines means that you go from reliable operation 
to "oh, that bit didn't make it correctly".

Lowering the front-side bus frequency or altering the memory timings can 
help (ie doing things like running DDR-333 at DDR-266). Making sure that 
your power supply isn't even close to its limits is good. And choosing a 
motherboard and chipsets from a reliable manufacturer is more than a good 
idea.

The reason why it's interesting that the errors seemed to happen in the 
same byte-lane is that I think it's common policy to route data lines on 
the same layer, and matching trace length per group is very important, 
because you do signal clocking per-group, afaik. But on the other hand, 
multiple layers on the board are expensive, so people try to minimize 
them, and maybe you end up routing through a via to another layer - which 
then makes timing and capacitance harder.

Or there aren't ground lines close enough, or the data lines are too close 
to other lines and you get cross-talk etc etc.

No, I've not done board design, and I don't know what I'm talking about, 
but look at the interesting zig-zagging the data (and address) lines often 
do on the board. It often looks totally crazy ("why doesn't that line just 
go straight?"), but the thing is that the groups all need to have the same 
length, but the pins are all at different points, so you can't make the 
lines straight, or some of them would be much shorter than others.

And if something is border-line, it may work all of the time - *until* you 
hit specific patterns that cause lots of lines to wiggle around, and then 
a capacitor won't handle the extra current draw from switching, or 
cross-talk between lines hits you, and what used to work doesn't work any 
more.

I wish we all had ECC memory. That gets rid of a lot of worries.

			Linus

  reply	other threads:[~2007-10-04  3:40 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <fa.VCoGTwbB+qeNlkBJ24MiuXcGSRA@ifi.uio.no>
     [not found] ` <fa.R+FhT6IGmV/dDweUt/MdSeA5YbQ@ifi.uio.no>
     [not found]   ` <fa.7MxTIt/ik0uRFN/lMWUasYV/JyE@ifi.uio.no>
     [not found]     ` <fa.YJ4uCPzXT5TQElSosyz98cpuMSc@ifi.uio.no>
     [not found]       ` <fa.Tp9UUX9EYprQyLg0shgH1YG9DDM@ifi.uio.no>
     [not found]         ` <fa.PzpJEEoLXfC+eOQZjTWjdf9vdnE@ifi.uio.no>
2007-10-04  2:59           ` Robert Hancock
2007-10-04  3:39             ` Linus Torvalds [this message]
2007-09-30 15:40 Neil Romig
2007-09-30 16:29 ` Pekka Enberg
2007-10-02 21:05   ` Neil Romig
2007-10-03  5:18     ` Pekka Enberg
2007-10-03 18:42       ` Neil Romig
2007-10-03 18:48         ` Pekka Enberg
2007-10-03 19:22           ` Linus Torvalds
2007-10-03 19:35             ` Pekka Enberg
2007-10-03 19:54               ` Linus Torvalds
2007-10-04  1:11                 ` Hiro Yoshioka
2007-10-03 20:30               ` Alan Cox
2007-10-04 18:34                 ` Neil Romig
2007-10-02 22:30 ` Chuck Ebbert

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=alpine.LFD.0.999.0710032008190.3579@woody.linux-foundation.org \
    --to=torvalds@linux-foundation.org \
    --cc=akpm@linux-foundation.org \
    --cc=hancockr@shaw.ca \
    --cc=hyoshiok@miraclelinux.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=neil@romig.demon.co.uk \
    --cc=penberg@cs.helsinki.fi \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®