* Re: PROBLEM: page allocation or what in 2.6.8.1
@ 2004-08-25 14:51 Harry Edmon
2004-08-26 10:31 ` Andrew Morton
0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-25 14:51 UTC (permalink / raw)
To: linux-kernel
I have had another crash on the same system as my message of 23 August:
Unable to handle kernel paging request at virtual address cf0b4e1c
printing eip:
c0178616
*pde = 0003f067
*pte = 0f0b4000
Oops: 0000 [#1]
PREEMPT SMP DEBUG_PAGEALLOC
Modules linked in: af_packet nfs nfsd exportfs lockd sunrpc autofs capability co
mmoncap ipv6 eepro100 uhci_hcd usbcore pciehp shpchp pci_hotplug floppy pcspkr e
vdev e100 mii sd_mod 3w_xxxx scsi_mod ide_cd cdrom rtc reiserfs isofs ext2 ext3
jbd mbcache ide_generic siimage aec62xx trm290 alim15x3 hpt34x hpt366 ide_disk c
md64x piix rz1000 slc90e66 generic cs5530 cs5520 sc1200 triflex atiixp pdc202xx_
old pdc202xx_new opti621 ns87415 cy82c693 amd74xx sis5513 via82cxxx serverworks
ide_core raid1 md unix
CPU: 0
EIP: 0060:[<c0178616>] Not tainted
EFLAGS: 00010202 (2.6.8.1-debug)
EIP is at iput+0x4e/0x7c
eax: cf0b4df8 ebx: de835eb4 ecx: 00000001 edx: c01785af
esi: c5a03f68 edi: de835eb4 ebp: f7454000 esp: f7455ebc
ds: 007b es: 007b ss: 0068
Process kswapd0 (pid: 69, threadinfo=f7454000 task=f78b3a60)
Stack: de835ed0 c02d852c f7454000 c0174cce de835eb4 c035ef80 c5a03f68 f7454000
c90f4f68 c0175363 c5a03f68 d2614eb4 0000007b 00000080 f7454000 00000000
f7ffeb1c c017593d 00000080 c0147b2a 00000080 000000d0 00003e0e c201c7a0
Call Trace:
[<c0174cce>] dput+0x11a/0x25b
[<c0175363>] prune_dcache+0x1bb/0x247
[<c017593d>] shrink_dcache_memory+0x1f/0x45
[<c0147b2a>] shrink_slab+0x150/0x1a4
[<c014927c>] balance_pgdat+0x216/0x261
[<c0149397>] kswapd+0xd0/0xee
[<c011cf1e>] autoremove_wake_function+0x0/0x57
[<c0105fba>] ret_from_fork+0x6/0x14
[<c011cf1e>] autoremove_wake_function+0x0/0x57
[<c01492c7>] kswapd+0x0/0xee
[<c0104271>] kernel_thread_helper+0x5/0xb
Code: 8b 40 24 85 c0 74 08 8b 40 18 85 c0 0f 45 d0 89 1c 24 ff d2
<6>note: kswapd0[69] exited with preempt_count 1
Before the crash I see messages like the following:
oom-killer: gfp_mask=0xd0
DMA per-cpu:
cpu 0 hot: low 2, high 6, batch 1
cpu 0 cold: low 0, high 2, batch 1
cpu 1 hot: low 2, high 6, batch 1
cpu 1 cold: low 0, high 2, batch 1
Normal per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16
HighMem per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16
Free pages: 1112520kB (1106368kB HighMem)
Active:19525 inactive:9757 dirty:10 writeback:0 unstable:0 free:278130 slab:2076
51 mapped:8904 pagetables:308
DMA free:1904kB min:16kB low:32kB high:48kB active:0kB inactive:32kB present:163
84kB
protections[]: 8 476 732
Normal free:4248kB min:936kB low:1872kB high:2808kB active:25112kB inactive:2502
0kB present:901120kB
protections[]: 0 468 724
HighMem free:1106240kB min:512kB low:1024kB high:1536kB active:53004kB inactive:
14024kB present:1179584kB
protections[]: 0 0 256
DMA: 0*4kB 0*8kB 57*16kB 9*32kB 1*64kB 1*128kB 0*256kB 1*512kB 0*1024kB 0*2048kB
0*4096kB = 1904kB
Normal: 624*4kB 3*8kB 0*16kB 6*32kB 0*64kB 0*128kB 0*256kB 1*512kB 1*1024kB 0*20
48kB 0*4096kB = 4248kB
HighMem: 2242*4kB 4823*8kB 5072*16kB 4346*32kB 3379*64kB 2189*128kB 896*256kB 19
0*512kB 15*1024kB 0*2048kB 0*4096kB = 1106240kB
Swap cache: add 362198, delete 361774, find 247258/263510, race 0+3
Out of Memory: Killed process 6437 (apache2).
--
Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu
206-543-0547 harry@u.washington.edu
Dept of Atmospheric Sciences FAX: 206-543-0308
University of Washington, Box 351640, Seattle, WA 98195-1640
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: PROBLEM: page allocation or what in 2.6.8.1 2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon @ 2004-08-26 10:31 ` Andrew Morton 2004-08-31 16:36 ` Re[2]: " Harry Edmon 0 siblings, 1 reply; 7+ messages in thread From: Andrew Morton @ 2004-08-26 10:31 UTC (permalink / raw) To: Harry Edmon; +Cc: linux-kernel Harry Edmon <harry@atmos.washington.edu> wrote: > > I have had another crash on the same system as my message of 23 August: > > Unable to handle kernel paging request at virtual address cf0b4e1c hm. Is the hardware known to be good? > ... > > Before the crash I see messages like the following: > > oom-killer: gfp_mask=0xd0 That's because you've enabled CONFIG_DEBUG_PAGEALLOC. It enormously increases the size of slab objects, which seems to cause memory reclaim to blow up. (It shouldn't but it does. It's a low-priority problem though). It's unlikely that the oom-killing caused the oops, but it's possible I guess. There's supposed to be a dump_stack() in the out_of_memory() path, which would help in searching for bugs, but that seems to have got lost. ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re[2]: PROBLEM: page allocation or what in 2.6.8.1 2004-08-26 10:31 ` Andrew Morton @ 2004-08-31 16:36 ` Harry Edmon 2004-08-31 19:02 ` Andrew Morton 0 siblings, 1 reply; 7+ messages in thread From: Harry Edmon @ 2004-08-31 16:36 UTC (permalink / raw) To: akpm; +Cc: linux-kernel We believe the hardware is okay. We have run numerous memtests, all okay. This is a Tyan S2721-533 with dual 3.06 Xeons. We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially in nfsd. We have now gone to the 2.6.8.1-mm4 kernel and got the following crash: kfree_debugcheck: bad ptr f8c189fch. ------------[ cut here ]------------ kernel BUG at mm/slab.c:1833! invalid operand: 0000 [#1] PREEMPT SMP Modules linked in: nfs autofs4 nfsd exportfs lockd sunrpc capability commoncap i pv6 eepro100 joydev tsdev usbhid uhci_hcd usbcore evdev e100 mii sd_mod 3w_xxxx scsi_mod ide_cd cdrom rtc unix CPU: 0 EIP: 0060:[<c0143d0c>] Not tainted VLI EFLAGS: 00010086 (2.6.8.1-debug-mm) EIP is at kfree_debugcheck+0x4c/0x70 eax: 00000028 ebx: c1718300 ecx: c03729bc edx: c03729bc esi: f8c189fc edi: f8c189fc ebp: f2efbeac esp: f2efbe9c ds: 007b es: 007b ss: 0068 Process rpc.mountd (pid: 2529, threadinfo=f2efa000 task=f2ede0b0) Stack: c0339fa0 f8c189fc f8c189fc f7d5124c f2efbed0 c0144cb5 f8c189fc f2efbeb4 f7d51bcc 00000286 f7d5124c f7d5124c efdd925d f2efbee0 f8bc7a48 f8c189fc f7d51bcc f2efbf0c f8bc80d9 f7d5124c f8bddfc0 00000000 f2efa000 f8bdef28 Call Trace: [<c0106f69>] show_stack+0x80/0x96 [<c0107100>] show_registers+0x15f/0x1c3 [<c010730a>] die+0x10d/0x1a8 [<c010782b>] do_invalid_op+0x104/0x106 [<c0106b81>] error_code+0x2d/0x38 [<c0144cb5>] kfree+0x25/0x9f [<f8bc7a48>] ip_map_put+0x45/0x6e [sunrpc] [<f8bc80d9>] ip_map_lookup+0x2ce/0x3a9 [sunrpc] [<f8bc8215>] auth_unix_add_addr+0x61/0x9f [sunrpc] [<f8c02eb9>] exp_addclient+0xb2/0xbc [nfsd] [<f8bfa9ca>] nfsctl_transaction_write+0x6e/0x98 [nfsd] [<c01869cb>] sys_nfsservctl+0xc0/0x115 [<c01060a5>] sysenter_past_esp+0x52/0x71 Code: e3 05 03 1d 90 95 48 c0 8b 03 a9 80 00 00 00 74 0a 8b 5d f8 8b 75 fc 89 ec 5d c3 89 74 24 04 c7 04 24 a0 9f 33 c0 e8 f0 a6 fd ff <0f> 0b 29 07 d6 92 33 c0 eb dc 89 74 24 04 c7 04 24 e0 9f 33 c0 Andrew Morton <akpm@osdl.org> wrote: > Harry Edmon <harry@atmos.washington.edu> wrote: > > > > I have had another crash on the same system as my message of 23 August: > > > > Unable to handle kernel paging request at virtual address cf0b4e1c > > hm. Is the hardware known to be good? > > > ... > > > > Before the crash I see messages like the following: > > > > oom-killer: gfp_mask=0xd0 > > That's because you've enabled. It enormously > increases the size of slab objects, which seems to cause memory reclaim to > blow up. (It shouldn't but it does. It's a low-priority problem though). > > It's unlikely that the oom-killing caused the oops, but it's possible I > guess. There's supposed to be a dump_stack() in the out_of_memory() path, > which would help in searching for bugs, but that seems to have got lost. -- Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu 206-543-0547 harry@u.washington.edu Dept of Atmospheric Sciences FAX: 206-543-0308 University of Washington, Box 351640, Seattle, WA 98195-1640 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: Re[2]: PROBLEM: page allocation or what in 2.6.8.1 2004-08-31 16:36 ` Re[2]: " Harry Edmon @ 2004-08-31 19:02 ` Andrew Morton 2004-08-31 22:21 ` Re[4]: " Harry Edmon 0 siblings, 1 reply; 7+ messages in thread From: Andrew Morton @ 2004-08-31 19:02 UTC (permalink / raw) To: Harry Edmon; +Cc: linux-kernel Harry Edmon <harry@atmos.washington.edu> wrote: > > We believe the hardware is okay. We have run numerous memtests, all okay. This > is a Tyan S2721-533 with dual 3.06 Xeons. > > We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially > in nfsd. We have now gone to the 2.6.8.1-mm4 kernel and got the following > crash: > > kfree_debugcheck: bad ptr f8c189fch. ip_map_put() is doing kfree(garbage). That should be fixed by the below, which was merged subsequent to 2.6.8.1. diff -puN net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map net/sunrpc/svcauth_unix.c --- 25/net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map 2004-08-26 23:30:29.000000000 -0700 +++ 25-akpm/net/sunrpc/svcauth_unix.c 2004-08-26 23:30:29.061446776 -0700 @@ -90,7 +90,7 @@ static void svcauth_unix_domain_release( struct ip_map { struct cache_head h; - char *m_class; /* e.g. "nfsd" */ + char m_class[8]; /* e.g. "nfsd" */ struct in_addr m_addr; struct unix_domain *m_client; int m_add_change; @@ -104,7 +104,6 @@ void ip_map_put(struct cache_head *item, if (test_bit(CACHE_VALID, &item->flags) && !test_bit(CACHE_NEGATIVE, &item->flags)) auth_domain_put(&im->m_client->h); - kfree(im->m_class); kfree(im); } } @@ -121,8 +120,7 @@ static inline int ip_map_match(struct ip } static inline void ip_map_init(struct ip_map *new, struct ip_map *item) { - new->m_class = item->m_class; - item->m_class = NULL; + strcpy(new->m_class, item->m_class); new->m_addr.s_addr = item->m_addr.s_addr; } static inline void ip_map_update(struct ip_map *new, struct ip_map *item) @@ -171,6 +169,8 @@ static int ip_map_parse(struct cache_det /* class */ len = qword_get(&mesg, class, 50); if (len <= 0) return -EINVAL; + if (len >= sizeof(ipm.m_class)) + return -EINVAL; /* ip address */ len = qword_get(&mesg, buf, 50); @@ -194,9 +194,7 @@ static int ip_map_parse(struct cache_det } else dom = NULL; - ipm.m_class = strdup(class); - if (ipm.m_class == NULL) - return -ENOMEM; + strcpy(ipm.m_class, class); ipm.m_addr.s_addr = htonl((((((b1<<8)|b2)<<8)|b3)<<8)|b4); ipm.h.flags = 0; @@ -212,7 +210,6 @@ static int ip_map_parse(struct cache_det ip_map_put(&ipmp->h, &ip_map_cache); if (dom) auth_domain_put(dom); - if (ipm.m_class) kfree(ipm.m_class); if (!ipmp) return -ENOMEM; cache_flush(); @@ -272,9 +269,7 @@ int auth_unix_add_addr(struct in_addr ad if (dom->flavour != RPC_AUTH_UNIX) return -EINVAL; udom = container_of(dom, struct unix_domain, h); - ip.m_class = strdup("nfsd"); - if (!ip.m_class) - return -ENOMEM; + strcpy(ip.m_class, "nfsd"); ip.m_addr = addr; ip.m_client = udom; ip.m_add_change = udom->addr_changes+1; @@ -282,7 +277,7 @@ int auth_unix_add_addr(struct in_addr ad ip.h.expiry_time = NEVER; ipmp = ip_map_lookup(&ip, 1); - if (ip.m_class) kfree(ip.m_class); + if (ipmp) { ip_map_put(&ipmp->h, &ip_map_cache); return 0; @@ -306,7 +301,7 @@ struct auth_domain *auth_unix_lookup(str struct ip_map key, *ipm; struct auth_domain *rv; - key.m_class = "nfsd"; + strcpy(key.m_class, "nfsd"); key.m_addr = addr; ipm = ip_map_lookup(&key, 0); @@ -368,7 +363,7 @@ svcauth_null_accept(struct svc_rqst *rqs svc_putu32(resv, RPC_AUTH_NULL); svc_putu32(resv, 0); - key.m_class = rqstp->rq_server->sv_program->pg_class; + strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class); key.m_addr = rqstp->rq_addr.sin_addr; ipm = ip_map_lookup(&key, 0); @@ -464,7 +459,7 @@ svcauth_unix_accept(struct svc_rqst *rqs } - key.m_class = rqstp->rq_server->sv_program->pg_class; + strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class); key.m_addr = rqstp->rq_addr.sin_addr; _ ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re[4]: PROBLEM: page allocation or what in 2.6.8.1 2004-08-31 19:02 ` Andrew Morton @ 2004-08-31 22:21 ` Harry Edmon 2004-08-31 22:47 ` Andrew Morton 0 siblings, 1 reply; 7+ messages in thread From: Harry Edmon @ 2004-08-31 22:21 UTC (permalink / raw) To: akpm; +Cc: linux-kernel So far, no crash. But now I have NFS clients that from time to time are unable to access this server. The server has the following messages on it: Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted I can temporarily fix the problem by typing: exportfs -ar But eventually it happens again. -- Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu 206-543-0547 harry@u.washington.edu Dept of Atmospheric Sciences FAX: 206-543-0308 University of Washington, Box 351640, Seattle, WA 98195-1640 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: Re[4]: PROBLEM: page allocation or what in 2.6.8.1 2004-08-31 22:21 ` Re[4]: " Harry Edmon @ 2004-08-31 22:47 ` Andrew Morton 2004-09-16 21:01 ` Harry Edmon 0 siblings, 1 reply; 7+ messages in thread From: Andrew Morton @ 2004-08-31 22:47 UTC (permalink / raw) To: Harry Edmon; +Cc: linux-kernel Harry Edmon <harry@atmos.washington.edu> wrote: > > So far, no crash. But now I have NFS clients that from time to time are unable > to access this server. The server has the following messages on it: > > Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted > > I can temporarily fix the problem by typing: > > exportfs -ar > > But eventually it happens again. > Well there were a few other NFS fixes. Can you test the latest kernel from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ? ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: PROBLEM: page allocation or what in 2.6.8.1 2004-08-31 22:47 ` Andrew Morton @ 2004-09-16 21:01 ` Harry Edmon 0 siblings, 0 replies; 7+ messages in thread From: Harry Edmon @ 2004-09-16 21:01 UTC (permalink / raw) To: Andrew Morton; +Cc: linux-kernel Once I mounted "nfsd" and went to autofs4 all problems went away with 2.6.9-rc1-bk9. Thanks for the help. Andrew Morton wrote: >Harry Edmon <harry@atmos.washington.edu> wrote: > > >>So far, no crash. But now I have NFS clients that from time to time are unable >>to access this server. The server has the following messages on it: >> >>Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted >> >>I can temporarily fix the problem by typing: >> >>exportfs -ar >> >>But eventually it happens again. >> >> >> > >Well there were a few other NFS fixes. Can you test the latest kernel >from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ? > > ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2004-09-16 21:01 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon 2004-08-26 10:31 ` Andrew Morton 2004-08-31 16:36 ` Re[2]: " Harry Edmon 2004-08-31 19:02 ` Andrew Morton 2004-08-31 22:21 ` Re[4]: " Harry Edmon 2004-08-31 22:47 ` Andrew Morton 2004-09-16 21:01 ` Harry Edmon
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®