mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* 2.6.0: kjournald 100% cpu + block errors(?)
@ 2004-01-23 20:30 JG
  0 siblings, 0 replies; only message in thread
From: JG @ 2004-01-23 20:30 UTC (permalink / raw)
  To: linux-kernel


[-- Attachment #1.1: Type: text/plain, Size: 2411 bytes --]

hi,

i just checked some stuff on my server via vnc when suddenly the cpu load went to 100% because of kjournald. i could only reboot my server with sysrq keys, because no shell was working/accepting input.
this is the dmesg output, which also scrolled down very fast on all consoles (not on vnc, but i also looked at the monitor of my server).

[...]
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
b_state=0x00000010, b_size=4096
block=4286578559, b_blocknr=18446744073701162879
[...]

i have no idea what could have caused this problem (the same problem, only other blocks, happend 11 days ago) and how i could debug this, since i have nothing else in the log files. when hitting the sysrq key to terminate the processes i got some other messages (which looked like oopses), but they were scrolling down so fast i couldn't read anything. is my OS hard disk (hda) dying (which is not older than 3 months)?

i also got this message on startup:
hda: set_drive_speed_status: status=0x58 { DriveReady SeekComplete DataRequest }
blk: queue e7dd1800, I/O limit 4095Mb (mask 0xffffffff)
hda: dma_timer_expiry: dma status == 0x60
hda: DMA timeout retry
hda: timeout waiting for DMA
blk: queue e7dd1400, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7c9b800, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7cbec00, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7cbe000, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7cafc00, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7caf400, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7caf000, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7ca6800, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7ca6400, I/O limit 4095Mb (mask 0xffffffff)
blk: queue e7c9bc00, I/O limit 4095Mb (mask 0xffffffff)
hda: status timeout: status=0xd0 { Busy }

hdb: DMA disabled
hda: drive not ready for command
ide0: reset: success
blk: queue e7dd1800, I/O limit 4095Mb (mask 0xffffffff)

i also attached the smartctl -a /dev/hda output, if it helps.

thanks very much for any info,
JG

[-- Attachment #1.2: smartctl --]
[-- Type: application/octet-stream, Size: 6250 bytes --]

smartctl version 5.26 Copyright (C) 2002-3 Bruce Allen
Home page is http://smartmontools.sourceforge.net/

=== START OF INFORMATION SECTION ===
Device Model:     HDS722512VLAT80
Serial Number:    VNR33EC3C1T2KK
Firmware Version: V33OA60A
Device is:        Not in smartctl database [for details use: -P showall]
ATA Version is:   6
ATA Standard is:  ATA/ATAPI-6 T13 1410D revision 3a
Local Time is:    Fri Jan 23 21:25:16 2004 CET
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x00)	Offline data collection activity was
					never started.
					Auto Offline Data Collection: Disabled.
Self-test execution status:      (  25)	The self-test routine was aborted by
					the host.
Total time to complete Offline 
data collection: 		 (2707) seconds.
Offline data collection
capabilities: 			 (0x1b) SMART execute Offline immediate.
					Auto Offline data collection on/off support.
					Suspend Offline collection upon new
					command.
					Offline surface scan supported.
					Self-test supported.
					No Conveyance Self-test supported.
					No Selective Self-test supported.
SMART capabilities:            (0x0003)	Saves SMART data before entering
					power-saving mode.
					Supports SMART auto save timer.
Error logging capability:        (0x01)	Error logging supported.
					General Purpose Logging supported.
Short self-test routine 
recommended polling time: 	 (   1) minutes.
Extended self-test routine
recommended polling time: 	 (  45) minutes.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x000b   096   096   060    Pre-fail  Always       -       9
  2 Throughput_Performance  0x0005   100   100   050    Pre-fail  Offline      -       0
  3 Spin_Up_Time            0x0007   100   100   024    Pre-fail  Always       -       195
  4 Start_Stop_Count        0x0012   100   100   000    Old_age   Always       -       6
  5 Reallocated_Sector_Ct   0x0033   100   100   005    Pre-fail  Always       -       16
  7 Seek_Error_Rate         0x000b   100   100   067    Pre-fail  Always       -       0
  8 Seek_Time_Performance   0x0005   100   100   020    Pre-fail  Offline      -       0
  9 Power_On_Hours          0x0012   100   100   000    Old_age   Always       -       1823
 10 Spin_Retry_Count        0x0013   100   100   060    Pre-fail  Always       -       0
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       6
192 Power-Off_Retract_Count 0x0032   100   100   050    Old_age   Always       -       80
193 Load_Cycle_Count        0x0012   100   100   050    Old_age   Always       -       80
194 Temperature_Celsius     0x0002   171   171   000    Old_age   Always       -       32 (Lifetime Min/Max 23/38)
196 Reallocated_Event_Count 0x0032   100   100   000    Old_age   Always       -       24
197 Current_Pending_Sector  0x0022   100   100   000    Old_age   Always       -       3
198 Offline_Uncorrectable   0x0008   100   100   000    Old_age   Offline      -       0
199 UDMA_CRC_Error_Count    0x000a   200   200   000    Old_age   Always       -       0

SMART Error Log Version: 1
ATA Error Count: 3
	CR = Command Register [HEX]
	FR = Features Register [HEX]
	SC = Sector Count Register [HEX]
	SN = Sector Number Register [HEX]
	CL = Cylinder Low Register [HEX]
	CH = Cylinder High Register [HEX]
	DH = Device/Head Register [HEX]
	DC = Device Command Register [HEX]
	ER = Error register [HEX]
	ST = Status register [HEX]
Timestamp = decimal seconds since the previous disk power-on.
Note: timestamp "wraps" after 2^32 msec = 49.710 days.

Error 3 occurred at disk power-on lifetime: 843 hours
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  40 51 70 64 ce 5c ee  Error: UNC

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Timestamp  Command/Feature_Name
  -- -- -- -- -- -- -- --   ---------  --------------------
  25 00 70 64 ce 5c e0 00 3030751.900  READ DMA EXT
  25 00 78 5c ce 5c e0 00 3030746.900  READ DMA EXT
  25 00 80 54 ce 5c e0 00 3030742.000  READ DMA EXT
  25 00 80 d4 34 14 e0 00 3030742.000  READ DMA EXT
  25 00 80 54 34 14 e0 00 3030742.000  READ DMA EXT

Error 2 occurred at disk power-on lifetime: 843 hours
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  40 51 77 5d ce 5c ee  Error: UNC

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Timestamp  Command/Feature_Name
  -- -- -- -- -- -- -- --   ---------  --------------------
  25 00 78 5c ce 5c e0 00 3030746.900  READ DMA EXT
  25 00 80 54 ce 5c e0 00 3030742.000  READ DMA EXT
  25 00 80 d4 34 14 e0 00 3030742.000  READ DMA EXT
  25 00 80 54 34 14 e0 00 3030742.000  READ DMA EXT
  25 00 80 d4 cd 5c e0 00 3030729.900  READ DMA EXT

Error 1 occurred at disk power-on lifetime: 843 hours
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  40 51 80 54 ce 5c ee  Error: UNC

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Timestamp  Command/Feature_Name
  -- -- -- -- -- -- -- --   ---------  --------------------
  25 00 80 54 ce 5c e0 00 3030742.000  READ DMA EXT
  25 00 80 d4 34 14 e0 00 3030742.000  READ DMA EXT
  25 00 80 54 34 14 e0 00 3030742.000  READ DMA EXT
  25 00 80 d4 cd 5c e0 00 3030729.900  READ DMA EXT
  25 00 80 54 cd 5c e0 00 3030729.900  READ DMA EXT

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Extended offline    Aborted by host               90%      1823         -


[-- Attachment #2: Type: application/pgp-signature, Size: 189 bytes --]

^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2004-01-23 20:30 UTC | newest]

Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2004-01-23 20:30 2.6.0: kjournald 100% cpu + block errors(?) JG

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome