* [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
@ 2005-09-15 17:40 Tomasz Kłoczko
2005-09-15 20:30 ` David S. Miller
0 siblings, 1 reply; 9+ messages in thread
From: Tomasz Kłoczko @ 2005-09-15 17:40 UTC (permalink / raw)
To: linux-kernel; +Cc: sparclinux, Aurora development, David S. Miller
[-- Attachment #1: Type: TEXT/PLAIN, Size: 587 bytes --]
I'm just catch series of kernel messages with soft lockup detected reports
(in attachment).
It occures during store big amout of data on NFS volume.
As NIC I use now Sun Swift gigabit eth (cassini driver). Probably this is
NFS related because I'm just browse yesterday logs and simillar was also
on Sun Happy Meal.
kloczek
--
-----------------------------------------------------------
*Ludzie nie mają problemów, tylko sobie sami je stwarzają*
-----------------------------------------------------------
Tomasz Kłoczko, sys adm @zie.pg.gda.pl|*e-mail: kloczek@rudy.mif.pg.gda.pl*
[-- Attachment #2: Type: TEXT/PLAIN, Size: 22089 bytes --]
BUG: soft lockup detected on CPU#0!
TSTATE: 0000009980009607 TPC: 00000000004775f4 TNPC: 00000000004775f8 Y: 00000000 Not tainted
TPC: <check_poison_obj+0x40/0x1bc>
g0: fffff80012b26030 g1: ffffffffffffffa5 g2: 000000000000006b g3: 0000000000180000
g4: fffff80004cb92a0 g5: fffff80000084000 g6: fffff80012b24000 g7: 00000000004c0bd9
o0: 00000000000000f0 o1: 000000000000ffff o2: ffffffffffff535e o3: ffffffffbb3f6458
o4: fffff80010fd8128 o5: 000000009913210b sp: fffff80012b258d1 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff800535f80f0 l1: 00000000000000f0 l2: 0000000000000000 l3: fffff80056b03c10
l4: fffff80043d5e628 l5: 0000000000000000 l6: 0000000000000014 l7: 00000000000005c8
i0: fffff8005fdf9b20 i1: fffff800535f80e8 i2: 0000000000000050 i3: 0000000000002c1c
i4: 0000000000002c1c i5: 00000000000001f8 i6: fffff80012b25991 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db3c TNPC: 000000000064db40 Y: 00000000 Not tainted
TPC: <unlock_kernel+0x58/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff8000a6506b8 o1: fffff8000a6506b8 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff8000a6506b8 i1: fffff8000a6507b0 i2: 0000000000000001 i3: 0000000000000000
i4: fffff8005d234470 i5: fffff8005d230000 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000009980009601 TPC: 00000000004775f4 TNPC: 00000000004775f8 Y: 00000000 Not tainted
TPC: <check_poison_obj+0x40/0x1bc>
g0: 0000000000000000 g1: ffffffffffffffa5 g2: 000000000000006b g3: 00000000006ac3b8
g4: fffff800382eae60 g5: fffff80000084000 g6: fffff8004fb90000 g7: 00000000000fd641
o0: 0000000000000800 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff800535f9dd0
o4: 0000000000000000 o5: 00000000006b5210 sp: fffff8004fb91d61 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff80036b2d8e0 l1: 0000000000000800 l2: 0000000000000000 l3: 000000000000009c
l4: 00000000000000b0 l5: fffff80043d5e628 l6: 0000000000000000 l7: fffff80056b03c10
i0: fffff800007adac0 i1: fffff80036b2d8d8 i2: 000000000000023f i3: 0000000000000000
i4: 0000000000000000 i5: 0000000000000094 i6: fffff8004fb91e21 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db38 TNPC: 000000000064db3c Y: 00000000 Not tainted
TPC: <unlock_kernel+0x54/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff8000cda66b8 o1: fffff8000cda66b8 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff8000cda66b8 i1: fffff8000cda67b0 i2: 0000000000000001 i3: 0000000000000000
i4: 0000000000000000 i5: 0000000000000008 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000000580009600 TPC: 000000000064db3c TNPC: 000000000064db40 Y: 00000000 Not tainted
TPC: <unlock_kernel+0x58/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800405289a0 g5: fffff80000084000 g6: fffff8000c61c000 g7: 0000018000000008
o0: fffff8000c61f698 o1: 0000000000000000 o2: 0000000000000000 o3: 0000000000000000
o4: 0000000000000000 o5: fffff80017f93be8 sp: fffff8000c61ede1 ret_pc: 00000000021c9480
RPC: <nfs_execute_write+0x20/0x50 [nfs]>
l0: fffff80044791138 l1: fffff8003f28a618 l2: fffff80038a79900 l3: 0000000000011250
l4: fffff8000c61f670 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff80017f939a0 i1: fffff80017f939a8 i2: 0000000000002000 i3: 0000000000000000
i4: 0000000000002000 i5: fffff80017f939a0 i6: fffff8000c61eeb1 i7: 00000000021c9904
I7: <nfs_flush_inode+0x454/0x538 [nfs]>
TSTATE: 0000009980009606 TPC: 00000000004560c0 TNPC: 00000000004560c4 Y: 00000000 Not tainted
TPC: <time_interpolator_get_offset+0x15c/0x170>
g0: fffff8005d2303a6 g1: 00000000000375db g2: 0000000000000010 g3: ffffffffffffffff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 000000000002800a
o0: 00000459d0aa1595 o1: 0000000000605668 o2: 0000000000000000 o3: 000000000000c0ac
o4: 000000000000c0ac o5: 00000000000001cc sp: fffff80057e72041 ret_pc: 0000000000456094
RPC: <time_interpolator_get_offset+0x130/0x170>
l0: 0000000000002e35 l1: 00000000006a0cd8 l2: 000000000079a800 l3: 000000000079a800
l4: 0000000000000000 l5: 000000000000c0ac l6: 0000000000000000 l7: 0000000000000000
i0: 0000000000000000 i1: fffff8005fe80ce0 i2: 0000000000000000 i3: fffffffff7ff1d04
i4: fffff80057e729d0 i5: 00000000ffffffff i6: fffff80057e72101 i7: 0000000000450254
I7: <do_gettimeofday+0x14/0x8c>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000004480009606 TPC: 0000000000477688 TNPC: 000000000047768c Y: 00000000 Not tainted
TPC: <check_poison_obj+0xd4/0x1bc>
g0: fffff80020f13ba0 g1: 000000000000006b g2: 000000000000006b g3: 0000000000000000
g4: fffff80044f75140 g5: fffff80000084000 g6: fffff800510a4000 g7: 000000000065687d
o0: 00000000000000f0 o1: fffff80043d5e628 o2: fffff800510a7330 o3: 0000000000000094
o4: 0000000000000000 o5: 00000000006b5210 sp: fffff800510a6001 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff80050fa0b40 l1: 00000000000000f0 l2: 0000000000000000 l3: 0000000000773c00
l4: 00000000006a0c00 l5: 0000000000000000 l6: 0000000000000000 l7: 00000001002cac74
i0: fffff8005fdf9b20 i1: fffff80050fa0b38 i2: 00000000000000a9 i3: 0000000000000010
i4: fffff80047ada064 i5: fffff80018a44c78 i6: fffff800510a60c1 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db38 TNPC: 000000000064db3c Y: 00000000 Not tainted
TPC: <unlock_kernel+0x54/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff800054de068 o1: fffff800054de068 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff800054de068 i1: fffff800054de160 i2: 0000000000000001 i3: 0000000000000000
i4: 000000000000ff00 i5: 000000000000004c i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000009980009606 TPC: 00000000004775f4 TNPC: 00000000004775f8 Y: 00000000 Not tainted
TPC: <check_poison_obj+0x40/0x1bc>
g0: fffff80020a3d570 g1: ffffffffffffffa5 g2: 000000000000006b g3: 00000000006ac3b8
g4: fffff800032ba2e0 g5: fffff80000084000 g6: fffff8004e7e4000 g7: 00000000002725ee
o0: 0000000000000800 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff800535f8828
o4: 0000000000000000 o5: 00000000006b5210 sp: fffff8004e7e6001 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff80013a7c910 l1: 0000000000000800 l2: 0000000000000000 l3: 0000000000773c00
l4: 0000000000704c00 l5: fffff8004e7e7140 l6: 0000000000000000 l7: 0000000000000008
i0: fffff800007adac0 i1: fffff80013a7c908 i2: 00000000000007b0 i3: 0000000000000010
i4: fffff8003ab20064 i5: fffff80018a44d80 i6: fffff8004e7e60c1 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db38 TNPC: 000000000064db3c Y: 00000000 Not tainted
TPC: <unlock_kernel+0x54/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff8003af03358 o1: fffff8003af03358 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff8003af03358 i1: fffff8003af03450 i2: 0000000000000001 i3: 0000000000000000
i4: 0000000000000000 i5: 0000000000000000 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000000080009602 TPC: 000000000050ae80 TNPC: 000000000050ae84 Y: 00000000 Not tainted
TPC: <radix_tree_tag_set+0x50/0xcc>
g0: fffff80047dd7930 g1: 0000000000000001 g2: 0000000000704f80 g3: fffff8003f28a9b0
g4: fffff80053865520 g5: fffff80000084000 g6: fffff8000e814000 g7: 0000000000000001
o0: fffff80004a13860 o1: 0000000000000000 o2: 0000080000000000 o3: 000000000000000c
o4: fffff800058d9900 o5: 0000000000000003 sp: fffff8000e816ca1 ret_pc: 000000000213d33c
RPC: <call_transmit+0x1f8/0x238 [sunrpc]>
l0: 0000000000000000 l1: 00000000021ccb98 l2: fffff80018a44528 l3: fffff80044791138
l4: fffff80038a790a8 l5: 0000000000011200 l6: 0000000000000003 l7: 0000000000000008
i0: fffff8004f082590 i1: 0000000000021cec i2: 0000000000000001 i3: 0000000000000001
i4: fffff8000e8175c0 i5: fffff80023569778 i6: fffff8000e816d61 i7: 00000000021c5ad8
I7: <nfs_set_page_writeback_locked+0x38/0x4c [nfs]>
TSTATE: 0000004480009604 TPC: 000000000064d94c TNPC: 000000000064d950 Y: 00000000 Not tainted
TPC: <_spin_trylock_bh+0x38/0x144>
g0: fffff80014ed7920 g1: 00000000000000ff g2: 0000000000000000 g3: fffff80000000000
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000000
o0: fffff8003f28a988 o1: 0000000000000003 o2: 0000000000000001 o3: fffff80057e73a30
o4: 0000000000000000 o5: fffff800055f04f8 sp: fffff80057e73191 ret_pc: 00000000021c8dd0
RPC: <nfs_mark_request_commit+0x1c/0x70 [nfs]>
l0: fffff8003f28aa20 l1: fffff8003f28a988 l2: fffff8003f28a888 l3: 0000000000000001
l4: 0000000000000040 l5: fffff80057e739a0 l6: 0000000000000000 l7: 0000000000000008
i0: fffff8000abfad70 i1: 0000000000000007 i2: 0000000000000000 i3: ffffffffb72b47f0
i4: fffff8000abfadc0 i5: fffff800007c3698 i6: fffff80057e73251 i7: 00000000021c9d74
I7: <nfs_writeback_done_full+0x148/0x190 [nfs]>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000004480009606 TPC: 0000000000477600 TNPC: 0000000000477684 Y: 00000000 Not tainted
TPC: <check_poison_obj+0x4c/0x1bc>
g0: 0000000000000000 g1: 000000000000006b g2: 000000000000006b g3: 00000000006ac3d0
g4: fffff80008392240 g5: fffff80000084000 g6: fffff8000d9c8000 g7: 000000000007ff2e
o0: 0000000000001000 o1: fffff8001c1e46b0 o2: fffff8000d9cb780 o3: 0000000000000000
o4: 0000000000000000 o5: 0000000000004c4a sp: fffff8000d9ca961 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff8002b8de000 l1: 0000000000001000 l2: 0000000000000000 l3: 0000000000000000
l4: 0000000000774000 l5: fffff8000d9cb160 l6: 0000000000000000 l7: 00000001002edc3b
i0: fffff8005fea00a0 i1: fffff8002b8de000 i2: 0000000000000017 i3: 00000001002edc65
i4: fffff800007bcbc0 i5: fffff800007bca00 i6: fffff8000d9caa21 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000009980009602 TPC: 000000000045609c TNPC: 00000000004560a0 Y: 00000000 Not tainted
TPC: <time_interpolator_get_offset+0x138/0x170>
g0: 00000000006d4568 g1: 0000049be7b30bb2 g2: 0000000000000001 g3: 000000004329aa59
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000034776968
o0: 0000049be7b30bb2 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff80030634078
o4: 0000000000000000 o5: 00000000000003f5 sp: fffff80057e72ba1 ret_pc: 0000000000456094
RPC: <time_interpolator_get_offset+0x130/0x170>
l0: 0000000000000200 l1: 00000000006a0cd8 l2: 0000000000000000 l3: 0000000000000000
l4: 0000000000001000 l5: 000000000000000c l6: fffff80057e70000 l7: 0000000000000008
i0: 0000000000000000 i1: fffff80030634070 i2: 000000000065fc00 i3: 0000000000000000
i4: 0000000000000000 i5: fffff800570240b0 i6: fffff80057e72c61 i7: 0000000000450254
I7: <do_gettimeofday+0x14/0x8c>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000000080009601 TPC: 00000000005162c8 TNPC: 00000000005162cc Y: 00000000 Not tainted
TPC: <__bzero+0x88/0xc0>
g0: 0000000000000004 g1: 0000000000000000 g2: 0000000000000800 g3: 00000000006ac3b8
g4: fffff8003101d520 g5: fffff80000084000 g6: fffff80001e14000 g7: 0000000000441dd7
o0: fffff80036b2de60 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff80036b2d8e0
o4: 0000000000000280 o5: 00000000006b5210 sp: fffff80001e15d61 ret_pc: 000000000047733c
RPC: <poison_obj+0x2c/0x44>
l0: 0000000000000800 l1: 0000000000000800 l2: 0000000000000000 l3: 000000000000009c
l4: 00000000000000b0 l5: fffff80043d5e628 l6: 0000000000000000 l7: fffff8000b6504f0
i0: fffff800007adac0 i1: fffff80036b2e0e0 i2: 000000000000005a i3: 000000005e680290
i4: fffff8005fdd1bc8 i5: fffff80020a3d250 i6: fffff80001e15e21 i7: 0000000000477c80
I7: <cache_alloc_debugcheck_after+0x30/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db38 TNPC: 000000000064db3c Y: 00000000 Not tainted
TPC: <unlock_kernel+0x54/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff8000ee4b9a8 o1: fffff8000ee4b9a8 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff8000ee4b9a8 i1: fffff8000ee4baa0 i2: 0000000000000001 i3: 0000000000000000
i4: 000000000005208c i5: fffff8000e9cbde0 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000004480009606 TPC: 0000000000477688 TNPC: 000000000047768c Y: 00000000 Not tainted
TPC: <check_poison_obj+0xd4/0x1bc>
g0: 0000000000000000 g1: 000000000000006b g2: 000000000000006b g3: 00000000006ac3d0
g4: fffff80004cb92a0 g5: fffff80000084000 g6: fffff800510ec000 g7: 00000000000a7f0d
o0: 0000000000001000 o1: fffff8001596a060 o2: fffff800510ef780 o3: 0000000000000008
o4: 0000000000000000 o5: 0000000000002d86 sp: fffff800510ee961 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff8003c95f000 l1: 0000000000001000 l2: 0000000000000000 l3: 0000000000704c00
l4: 0000000000704c00 l5: 0000000000000000 l6: 0000000000000000 l7: 00000001003005ee
i0: fffff8005fea00a0 i1: fffff8003c95f000 i2: 0000000000000cef i3: 0000000100300623
i4: fffff800007bca60 i5: fffff800007bca00 i6: fffff800510eea21 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000009980009602 TPC: 000000000041a7d0 TNPC: 000000000041a7d4 Y: 00000000 Not tainted
TPC: <tick_get_tick+0x0/0x18>
g0: 00000000000000b2 g1: 000000000041a7d0 g2: 0000000000000001 g3: 000000004329ab8a
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 000000002eed30f4
o0: 000004b84ccbc32c o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff80039cf06c0
o4: 0000000000000000 o5: 00000000000002c5 sp: fffff80057e72ba1 ret_pc: 0000000000456094
RPC: <time_interpolator_get_offset+0x130/0x170>
l0: 0000000000000200 l1: 00000000006a0cd8 l2: 0000000000000000 l3: 0000000000000000
l4: 0000000000001000 l5: 000000000000000c l6: fffff80057e70000 l7: 0000000000000008
i0: 0000000000000000 i1: fffff80039cf06b8 i2: 000000000065fc00 i3: 0000000000000000
i4: 0000000000000000 i5: fffff800570240b0 i6: fffff80057e72c61 i7: 0000000000450254
I7: <do_gettimeofday+0x14/0x8c>
BUG: soft lockup detected on CPU#0!
TSTATE: 0000000080009601 TPC: 0000000000476fbc TNPC: 0000000000476fc0 Y: 00000000 Not tainted
TPC: <dbg_redzone1+0x3c/0x44>
g0: 0000000000000800 g1: 0000000000000000 g2: 00000000000000f0 g3: 0000000000000010
g4: fffff80004cb92a0 g5: fffff80000084000 g6: fffff800510ec000 g7: 0000000000caf2cf
o0: fffff8002bd41488 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff8002bd41488
o4: 0000000000000000 o5: 00000000006b5210 sp: fffff800510edd61 ret_pc: 0000000000476fb0
RPC: <dbg_redzone1+0x30/0x44>
l0: 00000000000000f0 l1: 00000000000000f0 l2: 0000000000000000 l3: 000000000000009c
l4: 00000000000000b0 l5: fffff80043d5e628 l6: 0000000000000000 l7: fffff8000b6504f0
i0: fffff8005fdf9b20 i1: fffff8002bd41480 i2: 000000000000005a i3: 0000000000000000
i4: 0000000000000000 i5: 0000000000000094 i6: fffff800510ede21 i7: 0000000000477cbc
I7: <cache_alloc_debugcheck_after+0x6c/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db3c TNPC: 000000000064db40 Y: 00000000 Not tainted
TPC: <unlock_kernel+0x58/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff80036f059a8 o1: fffff80036f059a8 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff80036f059a8 i1: fffff80036f05aa0 i2: 0000000000000001 i3: 0000000000000000
i4: 0000000000000000 i5: fffff8002c5c1dc0 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
hw tcp v4 csum failed
BUG: soft lockup detected on CPU#0!
TSTATE: 0000000580009600 TPC: 000000000064db38 TNPC: 000000000064db3c Y: 00000000 Not tainted
TPC: <unlock_kernel+0x54/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff80043dbb160 g5: fffff80000084000 g6: fffff8004ed04000 g7: 0000018000000008
o0: fffff8004ed07698 o1: 0000000000000000 o2: 0000000000000000 o3: 0000000000000000
o4: 0000000000000000 o5: fffff80055b825d0 sp: fffff8004ed06de1 ret_pc: 00000000021c9480
RPC: <nfs_execute_write+0x20/0x50 [nfs]>
l0: fffff80044791138 l1: fffff8003f28b638 l2: fffff80038a79900 l3: 0000000000011250
l4: fffff8004ed07670 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff80055b82388 i1: fffff80055b82390 i2: 0000000000002000 i3: 0000000000000000
i4: 0000000000002000 i5: fffff80055b82388 i6: fffff8004ed06eb1 i7: 00000000021c9904
I7: <nfs_flush_inode+0x454/0x538 [nfs]>
TSTATE: 0000009980009607 TPC: 00000000004560b8 TNPC: 00000000004560bc Y: 00000000 Not tainted
TPC: <time_interpolator_get_offset+0x154/0x170>
g0: 0000000000000088 g1: 0000000000037d45 g2: 00000000006a0cd8 g3: ffffffffffffffff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 000000000002800a
o0: 000004d7f6f01758 o1: 000000000000ffff o2: ffffffffffff6f3f o3: 00000000285244c0
o4: fffff80031831ed0 o5: 0000000000000003 sp: fffff80057e71f81 ret_pc: 0000000000456094
RPC: <time_interpolator_get_offset+0x130/0x170>
l0: fffff8002bd41488 l1: 00000000006a0cd8 l2: 00000000000005dc l3: fffff8000b6504f0
l4: fffff80043d5e628 l5: 0000000000799000 l6: 0000000000000014 l7: 00000000000005c8
i0: 0000000000000000 i1: 0000000000605668 i2: 0000000000000000 i3: 000000000000ab43
i4: 000000000000ab43 i5: 00000000000005c8 i6: fffff80057e72041 i7: 0000000000450254
I7: <do_gettimeofday+0x14/0x8c>
process `named' is using obsolete setsockopt SO_BSDCOMPAT
BUG: soft lockup detected on CPU#0!
TSTATE: 0000004480009601 TPC: 0000000000477688 TNPC: 000000000047768c Y: 00000000 Not tainted
TPC: <check_poison_obj+0xd4/0x1bc>
g0: 0000000000000000 g1: 000000000000006b g2: 000000000000006b g3: 00000000006ac3b8
g4: fffff80053864ce0 g5: fffff80000084000 g6: fffff800047ec000 g7: 0000000000638b83
o0: 0000000000000800 o1: 0000000000000000 o2: 5a5a5a5a5a5a5a5a o3: fffff800535f9170
o4: 0000000000000000 o5: 00000000006b5210 sp: fffff800047edd61 ret_pc: 00000000004775c8
RPC: <check_poison_obj+0x14/0x1bc>
l0: fffff80043a1a558 l1: 0000000000000800 l2: 0000000000000000 l3: 000000000000009c
l4: 00000000000000b0 l5: fffff80043d5e628 l6: 0000000000000000 l7: fffff8000b650d30
i0: fffff800007adac0 i1: fffff80043a1a550 i2: 000000000000023d i3: 0000000000000000
i4: 0000000000000000 i5: 0000000000000094 i6: fffff800047ede21 i7: 0000000000477c70
I7: <cache_alloc_debugcheck_after+0x20/0x160>
TSTATE: 0000000580009601 TPC: 000000000064db3c TNPC: 000000000064db40 Y: 00000000 Not tainted
TPC: <unlock_kernel+0x58/0x64>
g0: 0000000000000000 g1: 0000000000704b80 g2: 0000000100000000 g3: 00000000000000ff
g4: fffff800584da6a0 g5: fffff8000008c000 g6: fffff80057e70000 g7: 0000000000000001
o0: fffff800381f5cd0 o1: fffff800381f5cd0 o2: 0000000000000001 o3: 0000000000000000
o4: 0000000000000000 o5: fffff800570240b0 sp: fffff80057e733e1 ret_pc: 0000000002141264
RPC: <__rpc_execute+0x98/0x280 [sunrpc]>
l0: 0000000000000000 l1: 000000000000000f l2: 000000000215f400 l3: 0000000000000400
l4: 0000000000000000 l5: 0000000000000000 l6: 0000000000000000 l7: 0000000000000008
i0: fffff800381f5cd0 i1: fffff800381f5dc8 i2: 0000000000000001 i3: 0000000000000000
i4: 0000000000000000 i5: 00000000000022b4 i6: fffff80057e734a1 i7: 000000000045ca4c
I7: <worker_thread+0x198/0x218>
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-15 17:40 [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0! Tomasz Kłoczko
@ 2005-09-15 20:30 ` David S. Miller
2005-09-15 20:36 ` Jima
2005-09-16 12:42 ` Tomasz Kłoczko
0 siblings, 2 replies; 9+ messages in thread
From: David S. Miller @ 2005-09-15 20:30 UTC (permalink / raw)
To: kloczek; +Cc: linux-kernel, sparclinux, aurora-sparc-devel, davem
From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
Date: Thu, 15 Sep 2005 19:40:27 +0200 (CEST)
> I'm just catch series of kernel messages with soft lockup detected reports
> (in attachment).
> It occures during store big amout of data on NFS volume.
>
> As NIC I use now Sun Swift gigabit eth (cassini driver). Probably this is
> NFS related because I'm just browse yesterday logs and simillar was also
> on Sun Happy Meal.
Interesting. Can you reproduce this with SLAB poisioning disabled?
That debugging feature is extremely expensive, although it shouldn't
make the CPU stop scheduling processes for more than 10 seconds.
I wonder if the NFS daemon code needs to have some limits put on
how much cpu it consumes handling requests before it gives up the
cpu. Perhaps, it has such throttling already, I don't know.
I'll also try to see if there can be some kind of sparc64 specific
issue which would cause this.
Where did you get that Cassini driver btw? It's not upstream,
although if it exists it should be.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-15 20:30 ` David S. Miller
@ 2005-09-15 20:36 ` Jima
2005-09-16 12:42 ` Tomasz Kłoczko
1 sibling, 0 replies; 9+ messages in thread
From: Jima @ 2005-09-15 20:36 UTC (permalink / raw)
To: David S. Miller
Cc: kloczek, linux-kernel, sparclinux, aurora-sparc-devel, davem
On Thu, 15 Sep 2005, David S. Miller wrote:
> Where did you get that Cassini driver btw? It's not upstream,
> although if it exists it should be.
Spot rolled it into the most recent Aurora kernel RPMs. He got the
driver from me; I in turn got it from Sun's web site. (It's tagged GPL.)
Neither Spot nor I care to 'own' the driver, due to both of us lacking
the hardware it supports, which makes maintenance difficult, at best.
Jima
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-15 20:30 ` David S. Miller
2005-09-15 20:36 ` Jima
@ 2005-09-16 12:42 ` Tomasz Kłoczko
2005-09-16 19:57 ` David S. Miller
2005-09-16 19:59 ` David S. Miller
1 sibling, 2 replies; 9+ messages in thread
From: Tomasz Kłoczko @ 2005-09-16 12:42 UTC (permalink / raw)
To: David S. Miller; +Cc: linux-kernel, sparclinux, aurora-sparc-devel, davem
[-- Attachment #1: Type: TEXT/PLAIN, Size: 2963 bytes --]
On Thu, 15 Sep 2005, David S. Miller wrote:
> From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
> Date: Thu, 15 Sep 2005 19:40:27 +0200 (CEST)
>
>> I'm just catch series of kernel messages with soft lockup detected reports
>> (in attachment).
>> It occures during store big amout of data on NFS volume.
>>
>> As NIC I use now Sun Swift gigabit eth (cassini driver). Probably this is
>> NFS related because I'm just browse yesterday logs and simillar was also
>> on Sun Happy Meal.
>
> Interesting. Can you reproduce this with SLAB poisioning disabled?
> That debugging feature is extremely expensive, although it shouldn't
> make the CPU stop scheduling processes for more than 10 seconds.
>
> I wonder if the NFS daemon code needs to have some limits put on
> how much cpu it consumes handling requests before it gives up the
> cpu. Perhaps, it has such throttling already, I don't know.
But this not case NFS server but NSF client. During this lookups I observe
rpciod takes 90-99% time of single processor. Load is between 10 and 20.
> I'll also try to see if there can be some kind of sparc64 specific
> issue which would cause this.
>
> Where did you get that Cassini driver btw? It's not upstream,
> although if it exists it should be.
It is avalaible on Sun pages (Copyright by Sun by it is GPLed):
http://www.sun.com/download/index.jsp?cat=Hardware%20Drivers&tab=3
Few days ago I'm talk about this driver with spot on #aurora channel and
it was included during prepare next kernel package.
But continue ..
Yesterday before testing NFS I was trying utilize NIC is using only
ftp/http/rcync/scp and all works correctly (without reporting lookups).
~23:00 CET after reported series of lookups system hangs completly. After
restart to now I see in dmesg some known messages:
[root@boss]# dmesg | sort | uniq -c | sort -n | tail -n 2
87 svc: bad direction 268435456, dropping request <==========
285 hw tcp v4 csum failed
Also I have question about second (hw tcp v4 csum failed):
[root@boss]# ifconfig eth0
eth0 Link encap:Ethernet HWaddr 00:03:BA:18:41:F9
inet addr:153.19.33.230 Bcast:153.19.33.255 Mask:255.255.255.0
inet6 addr: fe80::203:baff:fe18:41f9/64 Scope:Link
UP BROADCAST RUNNING MULTICAST MTU:1500 Metric:1
RX packets:2361199 errors:0 dropped:0 overruns:0 frame:0
TX packets:4060978 errors:0 dropped:0 overruns:0 carrier:0
collisions:0 txqueuelen:1000
RX bytes:352428791 (336.1 Mb) TX bytes:4884521951 (4658.2 Mb)
Interrupt:192
As you see in both errors is "0". Is this correct reporting broken packets
in kernel messages instead in RX errors ?
kloczek
--
-----------------------------------------------------------
*Ludzie nie mają problemów, tylko sobie sami je stwarzają*
-----------------------------------------------------------
Tomasz Kłoczko, sys adm @zie.pg.gda.pl|*e-mail: kloczek@rudy.mif.pg.gda.pl*
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-16 12:42 ` Tomasz Kłoczko
@ 2005-09-16 19:57 ` David S. Miller
2005-09-16 19:59 ` David S. Miller
1 sibling, 0 replies; 9+ messages in thread
From: David S. Miller @ 2005-09-16 19:57 UTC (permalink / raw)
To: kloczek; +Cc: davem, linux-kernel, sparclinux, aurora-sparc-devel
From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
Date: Fri, 16 Sep 2005 14:42:56 +0200 (CEST)
> Also I have question about second (hw tcp v4 csum failed):
...
> As you see in both errors is "0". Is this correct reporting broken packets
> in kernel messages instead in RX errors ?
It is protocol level error, therefore it is not counted
in raw device level statistics.
BTW, when you get the "hw tcp v4 csum failed", the kernel ignores
the hw checksum and rechecks it using software.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-16 12:42 ` Tomasz Kłoczko
2005-09-16 19:57 ` David S. Miller
@ 2005-09-16 19:59 ` David S. Miller
2005-09-17 15:22 ` Tomasz Kłoczko
1 sibling, 1 reply; 9+ messages in thread
From: David S. Miller @ 2005-09-16 19:59 UTC (permalink / raw)
To: kloczek; +Cc: davem, linux-kernel, sparclinux, aurora-sparc-devel
From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
Date: Fri, 16 Sep 2005 14:42:56 +0200 (CEST)
> On Thu, 15 Sep 2005, David S. Miller wrote:
>
> > I wonder if the NFS daemon code needs to have some limits put on
> > how much cpu it consumes handling requests before it gives up the
> > cpu. Perhaps, it has such throttling already, I don't know.
>
> But this not case NFS server but NSF client. During this lookups I observe
> rpciod takes 90-99% time of single processor. Load is between 10 and 20.
After studying some code yesterday, NFS client has the same
exact problem as NFS daemon, namely that if you give it enough
work it will never give up the cpu so that other tasks can
be scheduled.
This is a serious bug, and can easily trigger those soft lockup
messages. Based upon some other reports seen on linux-kernel
and elsewhere, things like the raid1 kernel daemon have a similar
issue as well.
I think you can help things _enormusly_ by turning off SLAB
poisioning, as I said that debugging feature is _VERY_ expensive.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-16 19:59 ` David S. Miller
@ 2005-09-17 15:22 ` Tomasz Kłoczko
2005-09-17 17:49 ` Tomasz Kłoczko
0 siblings, 1 reply; 9+ messages in thread
From: Tomasz Kłoczko @ 2005-09-17 15:22 UTC (permalink / raw)
To: David S. Miller; +Cc: davem, linux-kernel, sparclinux, aurora-sparc-devel
[-- Attachment #1: Type: TEXT/PLAIN, Size: 2125 bytes --]
On Fri, 16 Sep 2005, David S. Miller wrote:
> From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
> Date: Fri, 16 Sep 2005 14:42:56 +0200 (CEST)
>
>> On Thu, 15 Sep 2005, David S. Miller wrote:
>>
>>> I wonder if the NFS daemon code needs to have some limits put on
>>> how much cpu it consumes handling requests before it gives up the
>>> cpu. Perhaps, it has such throttling already, I don't know.
>>
>> But this not case NFS server but NSF client. During this lookups I observe
>> rpciod takes 90-99% time of single processor. Load is between 10 and 20.
>
> After studying some code yesterday, NFS client has the same
> exact problem as NFS daemon, namely that if you give it enough
> work it will never give up the cpu so that other tasks can
> be scheduled.
>
> This is a serious bug, and can easily trigger those soft lockup
> messages. Based upon some other reports seen on linux-kernel
> and elsewhere, things like the raid1 kernel daemon have a similar
> issue as well.
>
> I think you can help things _enormusly_ by turning off SLAB
> poisioning, as I said that debugging feature is _VERY_ expensive.
I'll try.
Now I have next call trace catched during reboot system after perform full
backup on NFS volume (during this as previouse was tons of lookups). This
case was during umounting NFS volumes (on system shutdown). Below call
trace may have some typos because I read tem from screen and write on
paper:
Badness in interruptible_sleep_on_timeout at kernel/shed.c: 3403 (Not tainted)
Call trace:
[00000000021c3cf8] nfs_kill_super+0x8c/0xbc [nfs]
[00000000004962d4] deactivate_super+0x58/0x7c
[00000000004ac5cc] sys_umount+0x2f8/0x308
[0000000000410fd4] linux_sparc_syscall32_0x34/0x40
[0000000000011ebc] 0x11ebc
Above may be is some resoult corrupting something during soft lookupus but
if not probably may be usefull on finding bugs.
kloczek
--
-----------------------------------------------------------
*Ludzie nie mają problemów, tylko sobie sami je stwarzają*
-----------------------------------------------------------
Tomasz Kłoczko, sys adm @zie.pg.gda.pl|*e-mail: kloczek@rudy.mif.pg.gda.pl*
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-17 15:22 ` Tomasz Kłoczko
@ 2005-09-17 17:49 ` Tomasz Kłoczko
2005-09-18 6:45 ` David S. Miller
0 siblings, 1 reply; 9+ messages in thread
From: Tomasz Kłoczko @ 2005-09-17 17:49 UTC (permalink / raw)
To: David S. Miller; +Cc: davem, linux-kernel, sparclinux, aurora-sparc-devel
[-- Attachment #1: Type: TEXT/PLAIN, Size: 698 bytes --]
On Sat, 17 Sep 2005, Tomasz Kłoczko wrote:
[..]
Something new.
I'm just finish rewrite backup procedure from dumping to NFS volume to
using piped dump|ssh|dd.
During this happens:
PSYCHO0: Uncorrectable Error, primary error type[DMA Write]
PSYCHO0: bytemask[0000] dword_offset[7] UPA_MID[1f] was_block(1)
PSYCHO0: UE AFAR [000000005c8bc040]
PSYCHO0: UE Secondary errors [(DMA Write)]
PSYCHO0: IOMMU Error, type[Protection Error]
kloczek
--
-----------------------------------------------------------
*Ludzie nie mają problemów, tylko sobie sami je stwarzają*
-----------------------------------------------------------
Tomasz Kłoczko, sys adm @zie.pg.gda.pl|*e-mail: kloczek@rudy.mif.pg.gda.pl*
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0!
2005-09-17 17:49 ` Tomasz Kłoczko
@ 2005-09-18 6:45 ` David S. Miller
0 siblings, 0 replies; 9+ messages in thread
From: David S. Miller @ 2005-09-18 6:45 UTC (permalink / raw)
To: kloczek; +Cc: davem, linux-kernel, sparclinux, aurora-sparc-devel
From: Tomasz Kłoczko <kloczek@rudy.mif.pg.gda.pl>
Date: Sat, 17 Sep 2005 19:49:17 +0200 (CEST)
> Something new.
> I'm just finish rewrite backup procedure from dumping to NFS volume to
> using piped dump|ssh|dd.
> During this happens:
>
> PSYCHO0: Uncorrectable Error, primary error type[DMA Write]
> PSYCHO0: bytemask[0000] dword_offset[7] UPA_MID[1f] was_block(1)
> PSYCHO0: UE AFAR [000000005c8bc040]
> PSYCHO0: UE Secondary errors [(DMA Write)]
> PSYCHO0: IOMMU Error, type[Protection Error]
A driver unmapped a DMA area while the device is still
accessing it. If you're still using the cassini card,
it's driver is the most likely suspect.
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2005-09-18 6:45 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2005-09-15 17:40 [2.6.14-rc1/sparc54]: BUG: soft lockup detected on CPU#0! Tomasz Kłoczko
2005-09-15 20:30 ` David S. Miller
2005-09-15 20:36 ` Jima
2005-09-16 12:42 ` Tomasz Kłoczko
2005-09-16 19:57 ` David S. Miller
2005-09-16 19:59 ` David S. Miller
2005-09-17 15:22 ` Tomasz Kłoczko
2005-09-17 17:49 ` Tomasz Kłoczko
2005-09-18 6:45 ` David S. Miller
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®