* Re: PROBLEM: page allocation or what in 2.6.8.1
@ 2004-08-25 14:51 Harry Edmon
2004-08-26 10:31 ` Andrew Morton
0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-25 14:51 UTC (permalink / raw)
To: linux-kernel
I have had another crash on the same system as my message of 23 August:
Unable to handle kernel paging request at virtual address cf0b4e1c
printing eip:
c0178616
*pde = 0003f067
*pte = 0f0b4000
Oops: 0000 [#1]
PREEMPT SMP DEBUG_PAGEALLOC
Modules linked in: af_packet nfs nfsd exportfs lockd sunrpc autofs capability co
mmoncap ipv6 eepro100 uhci_hcd usbcore pciehp shpchp pci_hotplug floppy pcspkr e
vdev e100 mii sd_mod 3w_xxxx scsi_mod ide_cd cdrom rtc reiserfs isofs ext2 ext3
jbd mbcache ide_generic siimage aec62xx trm290 alim15x3 hpt34x hpt366 ide_disk c
md64x piix rz1000 slc90e66 generic cs5530 cs5520 sc1200 triflex atiixp pdc202xx_
old pdc202xx_new opti621 ns87415 cy82c693 amd74xx sis5513 via82cxxx serverworks
ide_core raid1 md unix
CPU: 0
EIP: 0060:[<c0178616>] Not tainted
EFLAGS: 00010202 (2.6.8.1-debug)
EIP is at iput+0x4e/0x7c
eax: cf0b4df8 ebx: de835eb4 ecx: 00000001 edx: c01785af
esi: c5a03f68 edi: de835eb4 ebp: f7454000 esp: f7455ebc
ds: 007b es: 007b ss: 0068
Process kswapd0 (pid: 69, threadinfo=f7454000 task=f78b3a60)
Stack: de835ed0 c02d852c f7454000 c0174cce de835eb4 c035ef80 c5a03f68 f7454000
c90f4f68 c0175363 c5a03f68 d2614eb4 0000007b 00000080 f7454000 00000000
f7ffeb1c c017593d 00000080 c0147b2a 00000080 000000d0 00003e0e c201c7a0
Call Trace:
[<c0174cce>] dput+0x11a/0x25b
[<c0175363>] prune_dcache+0x1bb/0x247
[<c017593d>] shrink_dcache_memory+0x1f/0x45
[<c0147b2a>] shrink_slab+0x150/0x1a4
[<c014927c>] balance_pgdat+0x216/0x261
[<c0149397>] kswapd+0xd0/0xee
[<c011cf1e>] autoremove_wake_function+0x0/0x57
[<c0105fba>] ret_from_fork+0x6/0x14
[<c011cf1e>] autoremove_wake_function+0x0/0x57
[<c01492c7>] kswapd+0x0/0xee
[<c0104271>] kernel_thread_helper+0x5/0xb
Code: 8b 40 24 85 c0 74 08 8b 40 18 85 c0 0f 45 d0 89 1c 24 ff d2
<6>note: kswapd0[69] exited with preempt_count 1
Before the crash I see messages like the following:
oom-killer: gfp_mask=0xd0
DMA per-cpu:
cpu 0 hot: low 2, high 6, batch 1
cpu 0 cold: low 0, high 2, batch 1
cpu 1 hot: low 2, high 6, batch 1
cpu 1 cold: low 0, high 2, batch 1
Normal per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16
HighMem per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16
Free pages: 1112520kB (1106368kB HighMem)
Active:19525 inactive:9757 dirty:10 writeback:0 unstable:0 free:278130 slab:2076
51 mapped:8904 pagetables:308
DMA free:1904kB min:16kB low:32kB high:48kB active:0kB inactive:32kB present:163
84kB
protections[]: 8 476 732
Normal free:4248kB min:936kB low:1872kB high:2808kB active:25112kB inactive:2502
0kB present:901120kB
protections[]: 0 468 724
HighMem free:1106240kB min:512kB low:1024kB high:1536kB active:53004kB inactive:
14024kB present:1179584kB
protections[]: 0 0 256
DMA: 0*4kB 0*8kB 57*16kB 9*32kB 1*64kB 1*128kB 0*256kB 1*512kB 0*1024kB 0*2048kB
0*4096kB = 1904kB
Normal: 624*4kB 3*8kB 0*16kB 6*32kB 0*64kB 0*128kB 0*256kB 1*512kB 1*1024kB 0*20
48kB 0*4096kB = 4248kB
HighMem: 2242*4kB 4823*8kB 5072*16kB 4346*32kB 3379*64kB 2189*128kB 896*256kB 19
0*512kB 15*1024kB 0*2048kB 0*4096kB = 1106240kB
Swap cache: add 362198, delete 361774, find 247258/263510, race 0+3
Out of Memory: Killed process 6437 (apache2).
--
Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu
206-543-0547 harry@u.washington.edu
Dept of Atmospheric Sciences FAX: 206-543-0308
University of Washington, Box 351640, Seattle, WA 98195-1640
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: PROBLEM: page allocation or what in 2.6.8.1
2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon
@ 2004-08-26 10:31 ` Andrew Morton
2004-08-31 16:36 ` Re[2]: " Harry Edmon
0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-26 10:31 UTC (permalink / raw)
To: Harry Edmon; +Cc: linux-kernel
Harry Edmon <harry@atmos.washington.edu> wrote:
>
> I have had another crash on the same system as my message of 23 August:
>
> Unable to handle kernel paging request at virtual address cf0b4e1c
hm. Is the hardware known to be good?
> ...
>
> Before the crash I see messages like the following:
>
> oom-killer: gfp_mask=0xd0
That's because you've enabled CONFIG_DEBUG_PAGEALLOC. It enormously
increases the size of slab objects, which seems to cause memory reclaim to
blow up. (It shouldn't but it does. It's a low-priority problem though).
It's unlikely that the oom-killing caused the oops, but it's possible I
guess. There's supposed to be a dump_stack() in the out_of_memory() path,
which would help in searching for bugs, but that seems to have got lost.
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re[2]: PROBLEM: page allocation or what in 2.6.8.1
2004-08-26 10:31 ` Andrew Morton
@ 2004-08-31 16:36 ` Harry Edmon
2004-08-31 19:02 ` Andrew Morton
0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-31 16:36 UTC (permalink / raw)
To: akpm; +Cc: linux-kernel
We believe the hardware is okay. We have run numerous memtests, all okay. This
is a Tyan S2721-533 with dual 3.06 Xeons.
We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially
in nfsd. We have now gone to the 2.6.8.1-mm4 kernel and got the following
crash:
kfree_debugcheck: bad ptr f8c189fch.
------------[ cut here ]------------
kernel BUG at mm/slab.c:1833!
invalid operand: 0000 [#1]
PREEMPT SMP
Modules linked in: nfs autofs4 nfsd exportfs lockd sunrpc capability commoncap i
pv6 eepro100 joydev tsdev usbhid uhci_hcd usbcore evdev e100 mii sd_mod 3w_xxxx
scsi_mod ide_cd cdrom rtc unix
CPU: 0
EIP: 0060:[<c0143d0c>] Not tainted VLI
EFLAGS: 00010086 (2.6.8.1-debug-mm)
EIP is at kfree_debugcheck+0x4c/0x70
eax: 00000028 ebx: c1718300 ecx: c03729bc edx: c03729bc
esi: f8c189fc edi: f8c189fc ebp: f2efbeac esp: f2efbe9c
ds: 007b es: 007b ss: 0068
Process rpc.mountd (pid: 2529, threadinfo=f2efa000 task=f2ede0b0)
Stack: c0339fa0 f8c189fc f8c189fc f7d5124c f2efbed0 c0144cb5 f8c189fc f2efbeb4
f7d51bcc 00000286 f7d5124c f7d5124c efdd925d f2efbee0 f8bc7a48 f8c189fc
f7d51bcc f2efbf0c f8bc80d9 f7d5124c f8bddfc0 00000000 f2efa000 f8bdef28
Call Trace:
[<c0106f69>] show_stack+0x80/0x96
[<c0107100>] show_registers+0x15f/0x1c3
[<c010730a>] die+0x10d/0x1a8
[<c010782b>] do_invalid_op+0x104/0x106
[<c0106b81>] error_code+0x2d/0x38
[<c0144cb5>] kfree+0x25/0x9f
[<f8bc7a48>] ip_map_put+0x45/0x6e [sunrpc]
[<f8bc80d9>] ip_map_lookup+0x2ce/0x3a9 [sunrpc]
[<f8bc8215>] auth_unix_add_addr+0x61/0x9f [sunrpc]
[<f8c02eb9>] exp_addclient+0xb2/0xbc [nfsd]
[<f8bfa9ca>] nfsctl_transaction_write+0x6e/0x98 [nfsd]
[<c01869cb>] sys_nfsservctl+0xc0/0x115
[<c01060a5>] sysenter_past_esp+0x52/0x71
Code: e3 05 03 1d 90 95 48 c0 8b 03 a9 80 00 00 00 74 0a 8b 5d f8 8b 75 fc 89 ec
5d c3 89 74 24 04 c7 04 24 a0 9f 33 c0 e8 f0 a6 fd ff <0f> 0b 29 07 d6 92 33 c0
eb dc 89 74 24 04 c7 04 24 e0 9f 33 c0
Andrew Morton <akpm@osdl.org> wrote:
> Harry Edmon <harry@atmos.washington.edu> wrote:
> >
> > I have had another crash on the same system as my message of 23 August:
> >
> > Unable to handle kernel paging request at virtual address cf0b4e1c
>
> hm. Is the hardware known to be good?
>
> > ...
> >
> > Before the crash I see messages like the following:
> >
> > oom-killer: gfp_mask=0xd0
>
> That's because you've enabled. It enormously
> increases the size of slab objects, which seems to cause memory reclaim to
> blow up. (It shouldn't but it does. It's a low-priority problem though).
>
> It's unlikely that the oom-killing caused the oops, but it's possible I
> guess. There's supposed to be a dump_stack() in the out_of_memory() path,
> which would help in searching for bugs, but that seems to have got lost.
--
Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu
206-543-0547 harry@u.washington.edu
Dept of Atmospheric Sciences FAX: 206-543-0308
University of Washington, Box 351640, Seattle, WA 98195-1640
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: Re[2]: PROBLEM: page allocation or what in 2.6.8.1
2004-08-31 16:36 ` Re[2]: " Harry Edmon
@ 2004-08-31 19:02 ` Andrew Morton
2004-08-31 22:21 ` Re[4]: " Harry Edmon
0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-31 19:02 UTC (permalink / raw)
To: Harry Edmon; +Cc: linux-kernel
Harry Edmon <harry@atmos.washington.edu> wrote:
>
> We believe the hardware is okay. We have run numerous memtests, all okay. This
> is a Tyan S2721-533 with dual 3.06 Xeons.
>
> We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially
> in nfsd. We have now gone to the 2.6.8.1-mm4 kernel and got the following
> crash:
>
> kfree_debugcheck: bad ptr f8c189fch.
ip_map_put() is doing kfree(garbage). That should be fixed by the below,
which was merged subsequent to 2.6.8.1.
diff -puN net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map net/sunrpc/svcauth_unix.c
--- 25/net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map 2004-08-26 23:30:29.000000000 -0700
+++ 25-akpm/net/sunrpc/svcauth_unix.c 2004-08-26 23:30:29.061446776 -0700
@@ -90,7 +90,7 @@ static void svcauth_unix_domain_release(
struct ip_map {
struct cache_head h;
- char *m_class; /* e.g. "nfsd" */
+ char m_class[8]; /* e.g. "nfsd" */
struct in_addr m_addr;
struct unix_domain *m_client;
int m_add_change;
@@ -104,7 +104,6 @@ void ip_map_put(struct cache_head *item,
if (test_bit(CACHE_VALID, &item->flags) &&
!test_bit(CACHE_NEGATIVE, &item->flags))
auth_domain_put(&im->m_client->h);
- kfree(im->m_class);
kfree(im);
}
}
@@ -121,8 +120,7 @@ static inline int ip_map_match(struct ip
}
static inline void ip_map_init(struct ip_map *new, struct ip_map *item)
{
- new->m_class = item->m_class;
- item->m_class = NULL;
+ strcpy(new->m_class, item->m_class);
new->m_addr.s_addr = item->m_addr.s_addr;
}
static inline void ip_map_update(struct ip_map *new, struct ip_map *item)
@@ -171,6 +169,8 @@ static int ip_map_parse(struct cache_det
/* class */
len = qword_get(&mesg, class, 50);
if (len <= 0) return -EINVAL;
+ if (len >= sizeof(ipm.m_class))
+ return -EINVAL;
/* ip address */
len = qword_get(&mesg, buf, 50);
@@ -194,9 +194,7 @@ static int ip_map_parse(struct cache_det
} else
dom = NULL;
- ipm.m_class = strdup(class);
- if (ipm.m_class == NULL)
- return -ENOMEM;
+ strcpy(ipm.m_class, class);
ipm.m_addr.s_addr =
htonl((((((b1<<8)|b2)<<8)|b3)<<8)|b4);
ipm.h.flags = 0;
@@ -212,7 +210,6 @@ static int ip_map_parse(struct cache_det
ip_map_put(&ipmp->h, &ip_map_cache);
if (dom)
auth_domain_put(dom);
- if (ipm.m_class) kfree(ipm.m_class);
if (!ipmp)
return -ENOMEM;
cache_flush();
@@ -272,9 +269,7 @@ int auth_unix_add_addr(struct in_addr ad
if (dom->flavour != RPC_AUTH_UNIX)
return -EINVAL;
udom = container_of(dom, struct unix_domain, h);
- ip.m_class = strdup("nfsd");
- if (!ip.m_class)
- return -ENOMEM;
+ strcpy(ip.m_class, "nfsd");
ip.m_addr = addr;
ip.m_client = udom;
ip.m_add_change = udom->addr_changes+1;
@@ -282,7 +277,7 @@ int auth_unix_add_addr(struct in_addr ad
ip.h.expiry_time = NEVER;
ipmp = ip_map_lookup(&ip, 1);
- if (ip.m_class) kfree(ip.m_class);
+
if (ipmp) {
ip_map_put(&ipmp->h, &ip_map_cache);
return 0;
@@ -306,7 +301,7 @@ struct auth_domain *auth_unix_lookup(str
struct ip_map key, *ipm;
struct auth_domain *rv;
- key.m_class = "nfsd";
+ strcpy(key.m_class, "nfsd");
key.m_addr = addr;
ipm = ip_map_lookup(&key, 0);
@@ -368,7 +363,7 @@ svcauth_null_accept(struct svc_rqst *rqs
svc_putu32(resv, RPC_AUTH_NULL);
svc_putu32(resv, 0);
- key.m_class = rqstp->rq_server->sv_program->pg_class;
+ strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class);
key.m_addr = rqstp->rq_addr.sin_addr;
ipm = ip_map_lookup(&key, 0);
@@ -464,7 +459,7 @@ svcauth_unix_accept(struct svc_rqst *rqs
}
- key.m_class = rqstp->rq_server->sv_program->pg_class;
+ strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class);
key.m_addr = rqstp->rq_addr.sin_addr;
_
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re[4]: PROBLEM: page allocation or what in 2.6.8.1
2004-08-31 19:02 ` Andrew Morton
@ 2004-08-31 22:21 ` Harry Edmon
2004-08-31 22:47 ` Andrew Morton
0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-31 22:21 UTC (permalink / raw)
To: akpm; +Cc: linux-kernel
So far, no crash. But now I have NFS clients that from time to time are unable
to access this server. The server has the following messages on it:
Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted
I can temporarily fix the problem by typing:
exportfs -ar
But eventually it happens again.
--
Dr. Harry Edmon E-MAIL: harry@atmos.washington.edu
206-543-0547 harry@u.washington.edu
Dept of Atmospheric Sciences FAX: 206-543-0308
University of Washington, Box 351640, Seattle, WA 98195-1640
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: Re[4]: PROBLEM: page allocation or what in 2.6.8.1
2004-08-31 22:21 ` Re[4]: " Harry Edmon
@ 2004-08-31 22:47 ` Andrew Morton
2004-09-16 21:01 ` Harry Edmon
0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-31 22:47 UTC (permalink / raw)
To: Harry Edmon; +Cc: linux-kernel
Harry Edmon <harry@atmos.washington.edu> wrote:
>
> So far, no crash. But now I have NFS clients that from time to time are unable
> to access this server. The server has the following messages on it:
>
> Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted
>
> I can temporarily fix the problem by typing:
>
> exportfs -ar
>
> But eventually it happens again.
>
Well there were a few other NFS fixes. Can you test the latest kernel
from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ?
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: PROBLEM: page allocation or what in 2.6.8.1
2004-08-31 22:47 ` Andrew Morton
@ 2004-09-16 21:01 ` Harry Edmon
0 siblings, 0 replies; 7+ messages in thread
From: Harry Edmon @ 2004-09-16 21:01 UTC (permalink / raw)
To: Andrew Morton; +Cc: linux-kernel
Once I mounted "nfsd" and went to autofs4 all problems went away with
2.6.9-rc1-bk9. Thanks for the help.
Andrew Morton wrote:
>Harry Edmon <harry@atmos.washington.edu> wrote:
>
>
>>So far, no crash. But now I have NFS clients that from time to time are unable
>>to access this server. The server has the following messages on it:
>>
>>Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted
>>
>>I can temporarily fix the problem by typing:
>>
>>exportfs -ar
>>
>>But eventually it happens again.
>>
>>
>>
>
>Well there were a few other NFS fixes. Can you test the latest kernel
>from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ?
>
>
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2004-09-16 21:01 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon
2004-08-26 10:31 ` Andrew Morton
2004-08-31 16:36 ` Re[2]: " Harry Edmon
2004-08-31 19:02 ` Andrew Morton
2004-08-31 22:21 ` Re[4]: " Harry Edmon
2004-08-31 22:47 ` Andrew Morton
2004-09-16 21:01 ` Harry Edmon
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®