* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
@ 2009-04-30 14:39 Doug Thompson
0 siblings, 0 replies; 12+ messages in thread
From: Doug Thompson @ 2009-04-30 14:39 UTC (permalink / raw)
To: Andi Kleen; +Cc: akpm, greg, mingo, tglx, hpa, dougthompson, linux-kernel
--- On Thu, 4/30/09, Andi Kleen <andi@firstfloor.org> wrote:
> From: Andi Kleen <andi@firstfloor.org>
> Subject: Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
> To: "Doug Thompson" <norsk5@yahoo.com>
> Cc: "Andi Kleen" <andi@firstfloor.org>
> Date: Thursday, April 30, 2009, 1:05 AM
> > The problem we have had is once
> > an Uncorrected Error fires and dumps the address, mapping it
> > to the DIMM silk screen label is difficult, especially in
> > user space, in gaining access to the registers of the
> > controller.
>
> You can just do it either after reboot or in the crash
> kernel. I don't
> think it's required to put it all in kernel. Also you don't
> really
> need access to the registers;
Actually, according to AMD, their reference code for mapping from an error address to a memory slot does require access to the controller's registers. On page 67 of the BKDG for family F10 from their website is 2 and 1/2 pages of the code to perform that mapping. It takes into consideration interleaving of all kinds, etc. It is narly to say the least.
> SMBIOS provides this
> information and
> mcelog knows how to convert it.
As I undestand SMBIOS it provides a linear assignment of basic memory starts and lengths but does not provide the memory controller context as AMD's reference code takes into consideration
>
> Trying to add other consumers to mce.c will be likely very
> messy;
> there's really no generic way to do it. I hope you're not
> planning
> turning the nicely CPU independent code in mce.c into a
> mess
> of twisty CPU specific passages like the old 32bit code
> was.
>
> -Andi
No, not at all. Keeping the "clean" code is paramount, but we are seeking for an interface to accept the MCE error register structure and map that information to at least a DIMM label field, if not more.
The EDAC module would register for that interface upon loading and unregister upon module unload.
The MCE code would call a stub routine that either returns no mapping occurred OR call the EDAC mapper. MCE could then determine from that return code if a mapping occurred or not. If it did, then display the desired information, otherwise proceed as normal.
doug t
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-05-01 7:53 ` Borislav Petkov
@ 2009-05-03 0:32 ` Aristeu Rozanski
0 siblings, 0 replies; 12+ messages in thread
From: Aristeu Rozanski @ 2009-05-03 0:32 UTC (permalink / raw)
To: petkovbb, Andi Kleen, Borislav Petkov, akpm, greg, mingo, tglx,
hpa, dougthompson, linux-kernel, Ben Woodard,
Mauro Carvalho Chehab
> I thought I should summarize some points on the design direction, which,
> in my impression are what all the parties involved more or less would
> agree on:
>
> * mce interrupt delivers, if EDAC compiled in, the required batch of
> dumped MSRs to chipset/CPU specific EDAC driver. Aristeu, I guess you
> guys have some code on that, right?
it's in really early stage and like Andi said, it needs work.
> * EDAC module decodes MCE info, computes the DIMM label out of
> that. Additional hw setup stuff like memory hoisting/interleaving,
> ganged/unganged MC mode, testing infrastructure like DRAM error
> injection and all that pertaining to the specific hw is handled by the
> EDAC module.
>
> * If EDAC not enabled, mce operates as before.
agreed.
--
Aristeu
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 12:47 ` Andi Kleen
2009-04-30 14:48 ` Aristeu Rozanski
2009-04-30 18:37 ` Mauro Carvalho Chehab
@ 2009-05-01 12:39 ` Ingo Molnar
2 siblings, 0 replies; 12+ messages in thread
From: Ingo Molnar @ 2009-05-01 12:39 UTC (permalink / raw)
To: Andi Kleen
Cc: Borislav Petkov, akpm, greg, tglx, hpa, dougthompson, linux-kernel
* Andi Kleen <andi@firstfloor.org> wrote:
> > Kconfig, mce code delivers needed error info to edac which, in
> > turn, goes and decodes the error/does the mapping to DIMM
> > blocks/supplies DRAM error injection facility for testing
> > purposes and similar things. That way you have both and they
> > don't overlap in functionality.
>
> You can do that, but it's redundant because mcelog can do this
> this already. [...]
The thing is, when we took up x86 maintenance i had a good look at
the MCE situation, i checked both the kernel and the user-space
side.
The kernel side MCE code was in pretty bad shape to begin with, but
mcelog (the user-space tool) is a big stinking pile of poo on every
level.
It's one of the worst piece of kernel related code i ever saw. I
think you wrote all of it, and you should be ashamed of that code,
and you should be ashamed of the design and you should be ashamed of
the concept.
It even came with its own 'database' code: mcelog*/db.[ch] is 600+
lines of needless code instead of obvious library use. It's NIH and
self-serving complexity all over.
And the thing is, mcelog/mcedecode never really _did_ anything real
an useful, other than to:
1) Confuse kernel users who see a fatal MCE panic, with cryptic,
quirky codes, who write that down on paper, then run it through
the user-space tool - just to see a piece of information the
kernel could have provided already. (if they didnt make any
mistakes while writing down the codes)
2) Decode a quirky, binary MCE record and combine it with DMI data.
(which the kernel can and should do just fine.)
Yes, i know about tolerant=3 and certain people/companies opting to
ignore MCE fatality levels and live dangerously (and i also know
about non-fatal reporting and correction extensions in hw) - but for
99.999% of the Linux users the whole thing is just needless
complexity today, that does not offer anything valuable.
And that is really what happens when code is misdesigned and the
wrong pieces of code are pushed to user-space: a crappy, limited ABI
and an under-maintained, big pile of junk user-space kit.
The obvious truth is that hardware faults have to be caught, decoded
and optionally handled by the kernel.
The EDAC code at least has a sane design: it realizes that hardware
faults _must_ be fully known, decoded and potentially handled in the
kernel.
Piggyback-ing to user-space is plain idiotic and not defensible. So
if a piece of hardware capability is handled by the EDAC code, the
x86 MCE code will step aside and will stay the heck out of that
business. At least until the two concepts are merged into some sane
kernel hardware fault logging and handling framework.
And Andi, until you dont grasp such _basic_ design concepts, you
have no business writing such code really. You should stay the heck
away from it and you should stop 'advising' people who made the
right calls while you messed up. It is mcelog that is crap, not the
EDAC code.
Ingo
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 14:48 ` Aristeu Rozanski
@ 2009-05-01 7:53 ` Borislav Petkov
2009-05-03 0:32 ` Aristeu Rozanski
0 siblings, 1 reply; 12+ messages in thread
From: Borislav Petkov @ 2009-05-01 7:53 UTC (permalink / raw)
To: Aristeu Rozanski
Cc: Andi Kleen, Borislav Petkov, akpm, greg, mingo, tglx, hpa,
dougthompson, linux-kernel, Ben Woodard, Mauro Carvalho Chehab
Hi all,
I thought I should summarize some points on the design direction, which,
in my impression are what all the parties involved more or less would
agree on:
* mce interrupt delivers, if EDAC compiled in, the required batch of
dumped MSRs to chipset/CPU specific EDAC driver. Aristeu, I guess you
guys have some code on that, right?
* EDAC module decodes MCE info, computes the DIMM label out of
that. Additional hw setup stuff like memory hoisting/interleaving,
ganged/unganged MC mode, testing infrastructure like DRAM error
injection and all that pertaining to the specific hw is handled by the
EDAC module.
* If EDAC not enabled, mce operates as before.
Comments/suggestions?
Hit me! :)
--
Regards/Gruss,
Boris.
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 12:47 ` Andi Kleen
2009-04-30 14:48 ` Aristeu Rozanski
@ 2009-04-30 18:37 ` Mauro Carvalho Chehab
2009-05-01 12:39 ` Ingo Molnar
2 siblings, 0 replies; 12+ messages in thread
From: Mauro Carvalho Chehab @ 2009-04-30 18:37 UTC (permalink / raw)
To: Andi Kleen
Cc: Borislav Petkov, akpm, greg, mingo, tglx, hpa, dougthompson,
linux-kernel
On Thu, 30 Apr 2009, Andi Kleen wrote:
>> Kconfig, mce code delivers needed error info to edac which, in turn,
>> goes and decodes the error/does the mapping to DIMM blocks/supplies DRAM
>> error injection facility for testing purposes and similar things. That
>> way you have both and they don't overlap in functionality.
>
> You can do that, but it's redundant because mcelog can do this
> this already. I had some conversations with existing EDAC users
> recently and they seem to only care about the resulting output,
> so just querying from mcelog is fine.
> The only issue is that mcelog needs to get the DIMM data. In many
> cases it can do so from SMBIOS output, if not a suitable interface
> would need to be provided by the kernel.
>From what I've heard from the existing EDAC users, they have several
concerns that mcelog could be viable replacement to their EDAC usage, due
to performance issues, including the need of accessing SMBIOS in order to
get such information.
Also, EDAC interface is already stablished, and, as pointed by Doug, it is
very useful on cluster environments, where memory failures is a big issue
and need to be solved as soon as possible.
EDAC solves this issue very well and works on a wider range of designs
than mcelog. So, there's no reason to deprecate it or to reject patches
adding EDAC interfaces to other chips.
On the other hand, mcelog is also useful on different scenarios. So, they
are not competing technologies, but complementary ones.
So, assuming that both EDAC and mcelog are needed, the proper design for
those chipsets where the memory controller is integrated with other log
functions (like AMD64 and Nethalem) seem to build an unique kernel layer
that retrieves the error logs from the harware and allows access to the
same data via both mcelog and EDAC userspace API's.
Cheers,
Mauro
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 12:47 ` Andi Kleen
@ 2009-04-30 14:48 ` Aristeu Rozanski
2009-05-01 7:53 ` Borislav Petkov
2009-04-30 18:37 ` Mauro Carvalho Chehab
2009-05-01 12:39 ` Ingo Molnar
2 siblings, 1 reply; 12+ messages in thread
From: Aristeu Rozanski @ 2009-04-30 14:48 UTC (permalink / raw)
To: Andi Kleen
Cc: Borislav Petkov, akpm, greg, mingo, tglx, hpa, dougthompson,
linux-kernel, Ben Woodard, Mauro Carvalho Chehab
(adding Ben Woodard, Mauro C. Chehab to Cc list)
> > Kconfig, mce code delivers needed error info to edac which, in turn,
> > goes and decodes the error/does the mapping to DIMM blocks/supplies DRAM
> > error injection facility for testing purposes and similar things. That
> > way you have both and they don't overlap in functionality.
> You can do that, but it's redundant because mcelog can do this
> this already. I had some conversations with existing EDAC users
> recently and they seem to only care about the resulting output,
> so just querying from mcelog is fine.
what about using the same EDAC interface? for lots of memory controllers, even
in other architectures than x86, EDAC interface is available. sounds
inconsistent to force users to have to handle special cases on their scripts
just because _optional_ sharing of error information from mce code is not
available.
how about SW/HW scrubbing?
> The only issue is that mcelog needs to get the DIMM data. In many
> cases it can do so from SMBIOS output, if not a suitable interface
> would need to be provided by the kernel.
that can be done already in EDAC and in this driver
> > By the way, I think there's a similar attempt/proposal of letting mce
> > and edac talk to each other from Red Hat so I think this could be a
> There was a fairly dubious patch floating around I think, but it
> had a couple of problems.
and what if those problems are solved? a patch like that would make possible
to have EDAC support for both AMD64 and Nehalem and wouldn't hurt the
performance of people who choose not to use EDAC.
--
Aristeu
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
@ 2009-04-30 14:23 Doug Thompson
0 siblings, 0 replies; 12+ messages in thread
From: Doug Thompson @ 2009-04-30 14:23 UTC (permalink / raw)
To: Andi Kleen, Borislav Petkov
Cc: akpm, greg, mingo, tglx, hpa, dougthompson, linux-kernel
W1DUG
--- On Thu, 4/30/09, Borislav Petkov <borislav.petkov@amd.com> wrote:
> From: Borislav Petkov <borislav.petkov@amd.com>
> Subject: Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
> To: "Andi Kleen" <andi@firstfloor.org>
> Cc: akpm@linux-foundation.org, greg@kroah.com, mingo@elte.hu, tglx@linutronix.de, hpa@zytor.com, dougthompson@xmission.com, linux-kernel@vger.kernel.org
> Date: Thursday, April 30, 2009, 5:57 AM
> Hi,
>
> On Wed, Apr 29, 2009 at 09:30:31PM +0200, Andi Kleen
> wrote:
> > Borislav Petkov <borislav.petkov@amd.com>
> writes:
> >
> > > Hi,
> > >
> > > thanks to all reviewers of the previous
> submission, here is the second
> > > version of this series.
> >
> > The classic problem of the previous versions of these
> patches was that
> > they consume the same error registers (even if using
> pci config versus
> > msrs as access methods) as the kernel machine check
> poll/threshold
> > interrupt code.
Even the recommendation of AMD of having a polling thread for CORRECTABLE ERROR has a race issue to the same error registers due to the fact that a MCE is an exception and cannot be deferred or blocked off. In the middle of any poll cycle a MCE could fire and touch the same registers. small but present.
> > And with two logging agents
> racing on the same
> > registers you will always get junk results. Typically
> with threshold
> > enabled the mce code wins the race. I suspect this
> patchkit has
> > exactly the same fundamental design problem. EDAC
> really is not
> > particularly fitting for integrated memory controllers
> that report
> > their errors using standard machine check events.
>
> ok, how about we remove tha MSR/PCI cfg space reading bits
> and leave
> that task solely to the mce core. Then, iff you have edac
> turned on in
> Kconfig, mce code delivers needed error info to edac which,
> in turn,
> goes and decodes the error/does the mapping to DIMM
> blocks/supplies DRAM
> error injection facility for testing purposes and similar
> things. That
> way you have both and they don't overlap in functionality.
Adding the synchronization between the two is very doable. It is not yet in the current patch set, but a work in progress.
That is the solution we are pursuing, to have a mechanism to provide communication between MCE and EDAC providing the mapping operation to a DIMM label. The MCA exception fires retrieves the info and calls EDAC module for address mapping.
MCE polling handler calls the EDAC module for address mapping.
EDAC's basic model is a polling operation on the error registers at a 1 second (tunable) rate.
AMD's manual describes the UNCORRECTABLE MEMORY error handling via the MCE handler. It further recommends a polling thread to harvest CORRECTABLE MEMORY errors. Last time I checked the MCE poller was running on a 5 minute poll cycle.
That is where we have 2 different threads polling the same error registers without synchronization is problematic and where a "Listener" pattern can be created to provide callbacks for both or form into a single poller operation.
>
> By the way, I think there's a similar attempt/proposal of
> letting mce
> and edac talk to each other from Red Hat so I think this
> could be a
> viable thing to try.
Exactly
>
> > -Andi (who thinks all of this decoding should be in
> user space anyways)
>
> Think of a big data center with a thousands of 2,4,8 socket
> blades
> and the admin collecting mce output and running around
> decoding the
> errors on his workstation. Even worse, the blades have
> different DIMM
> configurations due to hw upgrades/newer machines. I'd much
> rather have
> the complete decoding done in kernel, where all the
> information needed
> for proper decoding is present and with the error landing
> in syslog or
> some other monitored buffer instead of reconstructing it in
> userspace.
>
> Thanks.
>
> --
> Regards/Gruss,
> Boris.
This model of clusters with thousands of multi-core nodes (5,000 in one case I think of) is used many times. The system console is tie to a serial port via a BIOS switch. The serial port is then attached to "conman" and all the consoles are funneled to a cluster controller which parses for a "bad memory" event.
In sites with EDAC deployed now the parser finds the node number, the CPU number on the node and extracts the EDAC DIMM label provided and generates a Repair Ticket. The technician proceeds to find the proper rack, blade and DIMM and takes that node out of service (for MCEs that are intermittent the node is reboot earlier). Then the bad DIMM is replaced - the one identified from EDAC - and the node quickly brought back online.
Without the DIMM Label provided by EDAC - or with just mce 'bad address' information - ALL the DIMMS are swapped out for off-line testing or all are return for warranty replacement. Getting the node back on line is the priority and reducing the time the technician spends on the rack floor.
Bare MCE information is logged on the cluster controller and no time is spent trying to retrieve the log and running a user space program. Cheaper (man hours) and faster to swapout out all the DIMMs. But that is frowned on, with EDAC solving the problem for themnow.
The requested feature from the customers is to provide the DIMM label WITH the MCE error information as well.
doug t
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 11:57 ` Borislav Petkov
2009-04-30 12:21 ` Ingo Molnar
@ 2009-04-30 12:47 ` Andi Kleen
2009-04-30 14:48 ` Aristeu Rozanski
` (2 more replies)
1 sibling, 3 replies; 12+ messages in thread
From: Andi Kleen @ 2009-04-30 12:47 UTC (permalink / raw)
To: Borislav Petkov
Cc: Andi Kleen, akpm, greg, mingo, tglx, hpa, dougthompson, linux-kernel
> ok, how about we remove tha MSR/PCI cfg space reading bits and leave
> that task solely to the mce core. Then, iff you have edac turned on in
That's the minimum fix, but even then the patchkit does a lot of
things, not necessarily all needs to be together.
> Kconfig, mce code delivers needed error info to edac which, in turn,
> goes and decodes the error/does the mapping to DIMM blocks/supplies DRAM
> error injection facility for testing purposes and similar things. That
> way you have both and they don't overlap in functionality.
You can do that, but it's redundant because mcelog can do this
this already. I had some conversations with existing EDAC users
recently and they seem to only care about the resulting output,
so just querying from mcelog is fine.
The only issue is that mcelog needs to get the DIMM data. In many
cases it can do so from SMBIOS output, if not a suitable interface
would need to be provided by the kernel.
> By the way, I think there's a similar attempt/proposal of letting mce
> and edac talk to each other from Red Hat so I think this could be a
There was a fairly dubious patch floating around I think, but it
had a couple of problems.
> > -Andi (who thinks all of this decoding should be in user space anyways)
>
> Think of a big data center with a thousands of 2,4,8 socket blades
> and the admin collecting mce output and running around decoding the
Nobody said anything about admins decoding on their workstation.
Corrected events (which are the 90+% case) get decoded in user space on the
same system. Uncorrected events get decoded after the reboot. Both happens
automatically and transparently.
-Andi
--
ak@linux.intel.com -- Speaking for myself only.
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-30 11:57 ` Borislav Petkov
@ 2009-04-30 12:21 ` Ingo Molnar
2009-04-30 12:47 ` Andi Kleen
1 sibling, 0 replies; 12+ messages in thread
From: Ingo Molnar @ 2009-04-30 12:21 UTC (permalink / raw)
To: Borislav Petkov
Cc: Andi Kleen, akpm, greg, tglx, hpa, dougthompson, linux-kernel
* Borislav Petkov <borislav.petkov@amd.com> wrote:
> > -Andi (who thinks all of this decoding should be in user space
> > anyways)
>
> Think of a big data center with a thousands of 2,4,8 socket blades
> and the admin collecting mce output and running around decoding
> the errors on his workstation. Even worse, the blades have
> different DIMM configurations due to hw upgrades/newer machines.
> I'd much rather have the complete decoding done in kernel, where
> all the information needed for proper decoding is present and with
> the error landing in syslog or some other monitored buffer instead
> of reconstructing it in userspace.
Yes, this aspect of the design is correct. The MCE code is seriously
mis-designed that way, lets not repeat it for EDAC :)
Ingo
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-29 19:30 ` Andi Kleen
@ 2009-04-30 11:57 ` Borislav Petkov
2009-04-30 12:21 ` Ingo Molnar
2009-04-30 12:47 ` Andi Kleen
0 siblings, 2 replies; 12+ messages in thread
From: Borislav Petkov @ 2009-04-30 11:57 UTC (permalink / raw)
To: Andi Kleen; +Cc: akpm, greg, mingo, tglx, hpa, dougthompson, linux-kernel
Hi,
On Wed, Apr 29, 2009 at 09:30:31PM +0200, Andi Kleen wrote:
> Borislav Petkov <borislav.petkov@amd.com> writes:
>
> > Hi,
> >
> > thanks to all reviewers of the previous submission, here is the second
> > version of this series.
>
> The classic problem of the previous versions of these patches was that
> they consume the same error registers (even if using pci config versus
> msrs as access methods) as the kernel machine check poll/threshold
> interrupt code. And with two logging agents racing on the same
> registers you will always get junk results. Typically with threshold
> enabled the mce code wins the race. I suspect this patchkit has
> exactly the same fundamental design problem. EDAC really is not
> particularly fitting for integrated memory controllers that report
> their errors using standard machine check events.
ok, how about we remove tha MSR/PCI cfg space reading bits and leave
that task solely to the mce core. Then, iff you have edac turned on in
Kconfig, mce code delivers needed error info to edac which, in turn,
goes and decodes the error/does the mapping to DIMM blocks/supplies DRAM
error injection facility for testing purposes and similar things. That
way you have both and they don't overlap in functionality.
By the way, I think there's a similar attempt/proposal of letting mce
and edac talk to each other from Red Hat so I think this could be a
viable thing to try.
> -Andi (who thinks all of this decoding should be in user space anyways)
Think of a big data center with a thousands of 2,4,8 socket blades
and the admin collecting mce output and running around decoding the
errors on his workstation. Even worse, the blades have different DIMM
configurations due to hw upgrades/newer machines. I'd much rather have
the complete decoding done in kernel, where all the information needed
for proper decoding is present and with the error landing in syslog or
some other monitored buffer instead of reconstructing it in userspace.
Thanks.
--
Regards/Gruss,
Boris.
Operating | Advanced Micro Devices GmbH
System | Karl-Hammerschmidt-Str. 34, 85609 Dornach b. München, Germany
Research | Geschäftsführer: Jochen Polster, Thomas M. McCoy, Giuliano Meroni
Center | Sitz: Dornach, Gemeinde Aschheim, Landkreis München
(OSRC) | Registergericht München, HRB Nr. 43632
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
2009-04-29 16:54 Borislav Petkov
@ 2009-04-29 19:30 ` Andi Kleen
2009-04-30 11:57 ` Borislav Petkov
0 siblings, 1 reply; 12+ messages in thread
From: Andi Kleen @ 2009-04-29 19:30 UTC (permalink / raw)
To: Borislav Petkov; +Cc: akpm, greg, mingo, tglx, hpa, dougthompson, linux-kernel
Borislav Petkov <borislav.petkov@amd.com> writes:
> Hi,
>
> thanks to all reviewers of the previous submission, here is the second
> version of this series.
The classic problem of the previous versions of these patches was that
they consume the same error registers (even if using pci config versus
msrs as access methods) as the kernel machine check poll/threshold
interrupt code. And with two logging agents racing on the same
registers you will always get junk results. Typically with threshold
enabled the mce code wins the race. I suspect this patchkit has
exactly the same fundamental design problem. EDAC really is not
particularly fitting for integrated memory controllers that report
their errors using standard machine check events.
-Andi (who thinks all of this decoding should be in user space anyways)
--
ak@linux.intel.com -- Speaking for myself only.
^ permalink raw reply [flat|nested] 12+ messages in thread
* [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64
@ 2009-04-29 16:54 Borislav Petkov
2009-04-29 19:30 ` Andi Kleen
0 siblings, 1 reply; 12+ messages in thread
From: Borislav Petkov @ 2009-04-29 16:54 UTC (permalink / raw)
To: akpm, greg
Cc: mingo, tglx, hpa, dougthompson, linux-kernel, Borislav Petkov,
Doug Thompson
Hi,
thanks to all reviewers of the previous submission, here is the second
version of this series.
Highlights are the addition of two helpers to read/write MSRs on several
CPUs, denoted by a cpumask and using an array of MSR values per-CPU, as
Peter suggested. Since IMHO they look generic enough I've added them to
arch/x86/lib/msr-on-cpu.c (now renamed to msr.c).
Moreover, I've addressed all the issues raised from the previous series.
Please let me know should there be anything else remaining.
Thanks,
Boris.
arch/x86/include/asm/msr.h | 11 +
arch/x86/lib/Makefile | 2 +-
arch/x86/lib/msr-on-cpu.c | 97 -
arch/x86/lib/msr.c | 151 ++
drivers/edac/Kconfig | 26 +
drivers/edac/Makefile | 1 +
drivers/edac/amd64_edac.c | 5385 ++++++++++++++++++++++++++++++++++++++++++++
7 files changed, 5575 insertions(+), 98 deletions(-)
^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2009-05-03 0:32 UTC | newest]
Thread overview: 12+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2009-04-30 14:39 [RFC PATCH 00/21 v2] amd64_edac: EDAC module for AMD64 Doug Thompson
-- strict thread matches above, loose matches on Subject: below --
2009-04-30 14:23 Doug Thompson
2009-04-29 16:54 Borislav Petkov
2009-04-29 19:30 ` Andi Kleen
2009-04-30 11:57 ` Borislav Petkov
2009-04-30 12:21 ` Ingo Molnar
2009-04-30 12:47 ` Andi Kleen
2009-04-30 14:48 ` Aristeu Rozanski
2009-05-01 7:53 ` Borislav Petkov
2009-05-03 0:32 ` Aristeu Rozanski
2009-04-30 18:37 ` Mauro Carvalho Chehab
2009-05-01 12:39 ` Ingo Molnar
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®