mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] x86_64: Fix misplaced `continue' in mce.c
@ 2007-06-22  0:13 Joshua Wise
  2007-06-22  4:57 ` Randy Dunlap
  0 siblings, 1 reply; 2+ messages in thread
From: Joshua Wise @ 2007-06-22  0:13 UTC (permalink / raw)
  To: linux-kernel; +Cc: ak, thockin, akpm

From: Joshua Wise <jwise@google.com>

Background:
  When a userspace application wants to know about machine check events, it
  opens /dev/mcelog and does a read(). Usually, we found that this interface
  works well, but in some cases, when the system was taking large numbers of
  machine check exceptions, the read() would hang. The system would output a
  soft-lockup warning, and the daemon reading from /dev/mcelog would suck up
  as much of a single CPU as it could spinning in system space.

Description:
  This patch fixes this bug. In particular, there was a "continue" inside a
  timeout loop that presumably was intended to break out of the outer loop,
  but instead caused the inner loop to continue. This patch also makes the
  condition for the break-out a little more evident by changing a
  !time_before to a time_after_eq.

Result:
  The read() no longer hangs in this test case.

Testing:
  On my system, I could replicate the bug with the following command:
    # for i in `seq 15000`; do ./inject_sbe.sh; done
  where inject_sbe.sh contains commands to inject a single-bit error into the
  next memory write transaction.

Patch:
  This patch is against git f1518a088bde6aea49e7c472ed6ab96178fcba3e.

Signed-off-by: Joshua Wise <jwise@google.com>
Signed-off-by: Tim Hockin <thockin@google.com>

--

diff --git a/arch/x86_64/kernel/mce.c b/arch/x86_64/kernel/mce.c
index a14375d..aa83023 100644
--- a/arch/x86_64/kernel/mce.c
+++ b/arch/x86_64/kernel/mce.c
@@ -497,15 +497,17 @@ static ssize_t mce_read(struct file *fil
  	for (i = 0; i < next; i++) {
  		unsigned long start = jiffies;
  		while (!mcelog.entry[i].finished) {
-			if (!time_before(jiffies, start + 2)) {
+			if (time_after_eq(jiffies, start + 2)) {
  				memset(mcelog.entry + i,0, sizeof(struct mce));
-				continue;
+				goto timeout;
  			}
  			cpu_relax();
  		}
  		smp_rmb();
  		err |= copy_to_user(buf, mcelog.entry + i, sizeof(struct mce));
  		buf += sizeof(struct mce); 
+	 timeout:
+	 	;
  	}

  	memset(mcelog.entry, 0, next * sizeof(struct mce));

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH] x86_64: Fix misplaced `continue' in mce.c
  2007-06-22  0:13 [PATCH] x86_64: Fix misplaced `continue' in mce.c Joshua Wise
@ 2007-06-22  4:57 ` Randy Dunlap
  0 siblings, 0 replies; 2+ messages in thread
From: Randy Dunlap @ 2007-06-22  4:57 UTC (permalink / raw)
  To: Joshua Wise; +Cc: linux-kernel, ak, thockin, akpm

On Thu, 21 Jun 2007 17:13:49 -0700 (PDT) Joshua Wise wrote:

> Testing:
>   On my system, I could replicate the bug with the following command:
>     # for i in `seq 15000`; do ./inject_sbe.sh; done
>   where inject_sbe.sh contains commands to inject a single-bit error into the
>   next memory write transaction.

Hi,

Does inject_sbe.sh require special hardware?  If so, what hardware
is it?

Thanks,
---
~Randy
*** Remember to use Documentation/SubmitChecklist when testing your code ***

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2007-06-22  4:57 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2007-06-22  0:13 [PATCH] x86_64: Fix misplaced `continue' in mce.c Joshua Wise
2007-06-22  4:57 ` Randy Dunlap

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®