mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Re: PROBLEM: page allocation or what in 2.6.8.1
@ 2004-08-25 14:51 Harry Edmon
  2004-08-26 10:31 ` Andrew Morton
  0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-25 14:51 UTC (permalink / raw)
  To: linux-kernel

I have had another crash on the same system as my message of 23 August:

Unable to handle kernel paging request at virtual address cf0b4e1c

 printing eip:
c0178616
*pde = 0003f067
*pte = 0f0b4000
Oops: 0000 [#1]
PREEMPT SMP DEBUG_PAGEALLOC
Modules linked in: af_packet nfs nfsd exportfs lockd sunrpc autofs capability co
mmoncap ipv6 eepro100 uhci_hcd usbcore pciehp shpchp pci_hotplug floppy pcspkr e
vdev e100 mii sd_mod 3w_xxxx scsi_mod ide_cd cdrom rtc reiserfs isofs ext2 ext3 
jbd mbcache ide_generic siimage aec62xx trm290 alim15x3 hpt34x hpt366 ide_disk c
md64x piix rz1000 slc90e66 generic cs5530 cs5520 sc1200 triflex atiixp pdc202xx_
old pdc202xx_new opti621 ns87415 cy82c693 amd74xx sis5513 via82cxxx serverworks 
ide_core raid1 md unix
CPU:    0
EIP:    0060:[<c0178616>]    Not tainted
EFLAGS: 00010202   (2.6.8.1-debug) 
EIP is at iput+0x4e/0x7c
eax: cf0b4df8   ebx: de835eb4   ecx: 00000001   edx: c01785af
esi: c5a03f68   edi: de835eb4   ebp: f7454000   esp: f7455ebc
ds: 007b   es: 007b   ss: 0068
Process kswapd0 (pid: 69, threadinfo=f7454000 task=f78b3a60)
Stack: de835ed0 c02d852c f7454000 c0174cce de835eb4 c035ef80 c5a03f68 f7454000 
       c90f4f68 c0175363 c5a03f68 d2614eb4 0000007b 00000080 f7454000 00000000 
       f7ffeb1c c017593d 00000080 c0147b2a 00000080 000000d0 00003e0e c201c7a0 
Call Trace:
 [<c0174cce>] dput+0x11a/0x25b
 [<c0175363>] prune_dcache+0x1bb/0x247
 [<c017593d>] shrink_dcache_memory+0x1f/0x45
 [<c0147b2a>] shrink_slab+0x150/0x1a4
 [<c014927c>] balance_pgdat+0x216/0x261
 [<c0149397>] kswapd+0xd0/0xee
 [<c011cf1e>] autoremove_wake_function+0x0/0x57
 [<c0105fba>] ret_from_fork+0x6/0x14
 [<c011cf1e>] autoremove_wake_function+0x0/0x57
 [<c01492c7>] kswapd+0x0/0xee
 [<c0104271>] kernel_thread_helper+0x5/0xb
Code: 8b 40 24 85 c0 74 08 8b 40 18 85 c0 0f 45 d0 89 1c 24 ff d2 
 <6>note: kswapd0[69] exited with preempt_count 1

Before the crash I see messages like the following:

oom-killer: gfp_mask=0xd0
DMA per-cpu:
cpu 0 hot: low 2, high 6, batch 1
cpu 0 cold: low 0, high 2, batch 1
cpu 1 hot: low 2, high 6, batch 1
cpu 1 cold: low 0, high 2, batch 1
Normal per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16
HighMem per-cpu:
cpu 0 hot: low 32, high 96, batch 16
cpu 0 cold: low 0, high 32, batch 16
cpu 1 hot: low 32, high 96, batch 16
cpu 1 cold: low 0, high 32, batch 16

Free pages:     1112520kB (1106368kB HighMem)
Active:19525 inactive:9757 dirty:10 writeback:0 unstable:0 free:278130 slab:2076
51 mapped:8904 pagetables:308
DMA free:1904kB min:16kB low:32kB high:48kB active:0kB inactive:32kB present:163
84kB
protections[]: 8 476 732
Normal free:4248kB min:936kB low:1872kB high:2808kB active:25112kB inactive:2502
0kB present:901120kB
protections[]: 0 468 724
HighMem free:1106240kB min:512kB low:1024kB high:1536kB active:53004kB inactive:
14024kB present:1179584kB
protections[]: 0 0 256
DMA: 0*4kB 0*8kB 57*16kB 9*32kB 1*64kB 1*128kB 0*256kB 1*512kB 0*1024kB 0*2048kB
 0*4096kB = 1904kB
Normal: 624*4kB 3*8kB 0*16kB 6*32kB 0*64kB 0*128kB 0*256kB 1*512kB 1*1024kB 0*20
48kB 0*4096kB = 4248kB
HighMem: 2242*4kB 4823*8kB 5072*16kB 4346*32kB 3379*64kB 2189*128kB 896*256kB 19
0*512kB 15*1024kB 0*2048kB 0*4096kB = 1106240kB
Swap cache: add 362198, delete 361774, find 247258/263510, race 0+3
Out of Memory: Killed process 6437 (apache2).

-- 
 Dr. Harry Edmon			E-MAIL: harry@atmos.washington.edu
 206-543-0547				harry@u.washington.edu
 Dept of Atmospheric Sciences		FAX:	206-543-0308
 University of Washington, Box 351640, Seattle, WA 98195-1640

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon
@ 2004-08-26 10:31 ` Andrew Morton
  2004-08-31 16:36   ` Re[2]: " Harry Edmon
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-26 10:31 UTC (permalink / raw)
  To: Harry Edmon; +Cc: linux-kernel

Harry Edmon <harry@atmos.washington.edu> wrote:
>
> I have had another crash on the same system as my message of 23 August:
> 
> Unable to handle kernel paging request at virtual address cf0b4e1c

hm.  Is the hardware known to be good?

> ...
> 
> Before the crash I see messages like the following:
> 
> oom-killer: gfp_mask=0xd0

That's because you've enabled CONFIG_DEBUG_PAGEALLOC.  It enormously
increases the size of slab objects, which seems to cause memory reclaim to
blow up.  (It shouldn't but it does.  It's a low-priority problem though).

It's unlikely that the oom-killing caused the oops, but it's possible I
guess.  There's supposed to be a dump_stack() in the out_of_memory() path,
which would help in searching for bugs, but that seems to have got lost.


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re[2]: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-26 10:31 ` Andrew Morton
@ 2004-08-31 16:36   ` Harry Edmon
  2004-08-31 19:02     ` Andrew Morton
  0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-31 16:36 UTC (permalink / raw)
  To: akpm; +Cc: linux-kernel

We believe the hardware is okay.  We have run numerous memtests, all okay.  This
is a Tyan S2721-533 with dual 3.06 Xeons.

We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially
in nfsd.  We have now gone to the 2.6.8.1-mm4 kernel and got the following
crash:

kfree_debugcheck: bad ptr f8c189fch.

------------[ cut here ]------------

kernel BUG at mm/slab.c:1833!

invalid operand: 0000 [#1]

PREEMPT SMP 

Modules linked in: nfs autofs4 nfsd exportfs lockd sunrpc capability commoncap i
pv6 eepro100 joydev tsdev usbhid uhci_hcd usbcore evdev e100 mii sd_mod 3w_xxxx 
scsi_mod ide_cd cdrom rtc unix

CPU:    0

EIP:    0060:[<c0143d0c>]    Not tainted VLI

EFLAGS: 00010086   (2.6.8.1-debug-mm) 

EIP is at kfree_debugcheck+0x4c/0x70

eax: 00000028   ebx: c1718300   ecx: c03729bc   edx: c03729bc

esi: f8c189fc   edi: f8c189fc   ebp: f2efbeac   esp: f2efbe9c

ds: 007b   es: 007b   ss: 0068

Process rpc.mountd (pid: 2529, threadinfo=f2efa000 task=f2ede0b0)

Stack: c0339fa0 f8c189fc f8c189fc f7d5124c f2efbed0 c0144cb5 f8c189fc f2efbeb4 

       f7d51bcc 00000286 f7d5124c f7d5124c efdd925d f2efbee0 f8bc7a48 f8c189fc 

       f7d51bcc f2efbf0c f8bc80d9 f7d5124c f8bddfc0 00000000 f2efa000 f8bdef28 

Call Trace:

 [<c0106f69>] show_stack+0x80/0x96

 [<c0107100>] show_registers+0x15f/0x1c3

 [<c010730a>] die+0x10d/0x1a8

 [<c010782b>] do_invalid_op+0x104/0x106

 [<c0106b81>] error_code+0x2d/0x38

 [<c0144cb5>] kfree+0x25/0x9f

 [<f8bc7a48>] ip_map_put+0x45/0x6e [sunrpc]

 [<f8bc80d9>] ip_map_lookup+0x2ce/0x3a9 [sunrpc]

 [<f8bc8215>] auth_unix_add_addr+0x61/0x9f [sunrpc]

 [<f8c02eb9>] exp_addclient+0xb2/0xbc [nfsd]

 [<f8bfa9ca>] nfsctl_transaction_write+0x6e/0x98 [nfsd]

 [<c01869cb>] sys_nfsservctl+0xc0/0x115

 [<c01060a5>] sysenter_past_esp+0x52/0x71

Code: e3 05 03 1d 90 95 48 c0 8b 03 a9 80 00 00 00 74 0a 8b 5d f8 8b 75 fc 89 ec
 5d c3 89 74 24 04 c7 04 24 a0 9f 33 c0 e8 f0 a6 fd ff <0f> 0b 29 07 d6 92 33 c0
 eb dc 89 74 24 04 c7 04 24 e0 9f 33 c0 



Andrew Morton <akpm@osdl.org> wrote:
> Harry Edmon <harry@atmos.washington.edu> wrote:
> >
> > I have had another crash on the same system as my message of 23 August:
> > 
> > Unable to handle kernel paging request at virtual address cf0b4e1c
> 
> hm.  Is the hardware known to be good?
> 
> > ...
> > 
> > Before the crash I see messages like the following:
> > 
> > oom-killer: gfp_mask=0xd0
> 
> That's because you've enabled.  It enormously
> increases the size of slab objects, which seems to cause memory reclaim to
> blow up.  (It shouldn't but it does.  It's a low-priority problem though).
> 
> It's unlikely that the oom-killing caused the oops, but it's possible I
> guess.  There's supposed to be a dump_stack() in the out_of_memory() path,
> which would help in searching for bugs, but that seems to have got lost.

-- 
 Dr. Harry Edmon			E-MAIL: harry@atmos.washington.edu
 206-543-0547				harry@u.washington.edu
 Dept of Atmospheric Sciences		FAX:	206-543-0308
 University of Washington, Box 351640, Seattle, WA 98195-1640

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: Re[2]: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-31 16:36   ` Re[2]: " Harry Edmon
@ 2004-08-31 19:02     ` Andrew Morton
  2004-08-31 22:21       ` Re[4]: " Harry Edmon
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-31 19:02 UTC (permalink / raw)
  To: Harry Edmon; +Cc: linux-kernel

Harry Edmon <harry@atmos.washington.edu> wrote:
>
> We believe the hardware is okay.  We have run numerous memtests, all okay.  This
>  is a Tyan S2721-533 with dual 3.06 Xeons.
> 
>  We tried taking out CONFIG_DEBUG_PAGEALLOC and still we get crashes, especially
>  in nfsd.  We have now gone to the 2.6.8.1-mm4 kernel and got the following
>  crash:
> 
>  kfree_debugcheck: bad ptr f8c189fch.

ip_map_put() is doing kfree(garbage).  That should be fixed by the below,
which was merged subsequent to 2.6.8.1.


diff -puN net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map net/sunrpc/svcauth_unix.c
--- 25/net/sunrpc/svcauth_unix.c~use-fixed-size-buffer-instead-of-kmalloc-for-m_class-in-ip_map	2004-08-26 23:30:29.000000000 -0700
+++ 25-akpm/net/sunrpc/svcauth_unix.c	2004-08-26 23:30:29.061446776 -0700
@@ -90,7 +90,7 @@ static void svcauth_unix_domain_release(
 
 struct ip_map {
 	struct cache_head	h;
-	char			*m_class; /* e.g. "nfsd" */
+	char			m_class[8]; /* e.g. "nfsd" */
 	struct in_addr		m_addr;
 	struct unix_domain	*m_client;
 	int			m_add_change;
@@ -104,7 +104,6 @@ void ip_map_put(struct cache_head *item,
 		if (test_bit(CACHE_VALID, &item->flags) &&
 		    !test_bit(CACHE_NEGATIVE, &item->flags))
 			auth_domain_put(&im->m_client->h);
-		kfree(im->m_class);
 		kfree(im);
 	}
 }
@@ -121,8 +120,7 @@ static inline int ip_map_match(struct ip
 }
 static inline void ip_map_init(struct ip_map *new, struct ip_map *item)
 {
-	new->m_class = item->m_class;
-	item->m_class = NULL;
+	strcpy(new->m_class, item->m_class);
 	new->m_addr.s_addr = item->m_addr.s_addr;
 }
 static inline void ip_map_update(struct ip_map *new, struct ip_map *item)
@@ -171,6 +169,8 @@ static int ip_map_parse(struct cache_det
 	/* class */
 	len = qword_get(&mesg, class, 50);
 	if (len <= 0) return -EINVAL;
+	if (len >= sizeof(ipm.m_class))
+		return -EINVAL;
 
 	/* ip address */
 	len = qword_get(&mesg, buf, 50);
@@ -194,9 +194,7 @@ static int ip_map_parse(struct cache_det
 	} else
 		dom = NULL;
 
-	ipm.m_class = strdup(class);
-	if (ipm.m_class == NULL)
-		return -ENOMEM;
+	strcpy(ipm.m_class, class);
 	ipm.m_addr.s_addr =
 		htonl((((((b1<<8)|b2)<<8)|b3)<<8)|b4);
 	ipm.h.flags = 0;
@@ -212,7 +210,6 @@ static int ip_map_parse(struct cache_det
 		ip_map_put(&ipmp->h, &ip_map_cache);
 	if (dom)
 		auth_domain_put(dom);
-	if (ipm.m_class) kfree(ipm.m_class);
 	if (!ipmp)
 		return -ENOMEM;
 	cache_flush();
@@ -272,9 +269,7 @@ int auth_unix_add_addr(struct in_addr ad
 	if (dom->flavour != RPC_AUTH_UNIX)
 		return -EINVAL;
 	udom = container_of(dom, struct unix_domain, h);
-	ip.m_class = strdup("nfsd");
-	if (!ip.m_class)
-		return -ENOMEM;
+	strcpy(ip.m_class, "nfsd");
 	ip.m_addr = addr;
 	ip.m_client = udom;
 	ip.m_add_change = udom->addr_changes+1;
@@ -282,7 +277,7 @@ int auth_unix_add_addr(struct in_addr ad
 	ip.h.expiry_time = NEVER;
 	
 	ipmp = ip_map_lookup(&ip, 1);
-	if (ip.m_class) kfree(ip.m_class);
+
 	if (ipmp) {
 		ip_map_put(&ipmp->h, &ip_map_cache);
 		return 0;
@@ -306,7 +301,7 @@ struct auth_domain *auth_unix_lookup(str
 	struct ip_map key, *ipm;
 	struct auth_domain *rv;
 
-	key.m_class = "nfsd";
+	strcpy(key.m_class, "nfsd");
 	key.m_addr = addr;
 
 	ipm = ip_map_lookup(&key, 0);
@@ -368,7 +363,7 @@ svcauth_null_accept(struct svc_rqst *rqs
 	svc_putu32(resv, RPC_AUTH_NULL);
 	svc_putu32(resv, 0);
 
-	key.m_class = rqstp->rq_server->sv_program->pg_class;
+	strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class);
 	key.m_addr = rqstp->rq_addr.sin_addr;
 
 	ipm = ip_map_lookup(&key, 0);
@@ -464,7 +459,7 @@ svcauth_unix_accept(struct svc_rqst *rqs
 	}
 
 
-	key.m_class = rqstp->rq_server->sv_program->pg_class;
+	strcpy(key.m_class, rqstp->rq_server->sv_program->pg_class);
 	key.m_addr = rqstp->rq_addr.sin_addr;
 
 
_


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re[4]: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-31 19:02     ` Andrew Morton
@ 2004-08-31 22:21       ` Harry Edmon
  2004-08-31 22:47         ` Andrew Morton
  0 siblings, 1 reply; 7+ messages in thread
From: Harry Edmon @ 2004-08-31 22:21 UTC (permalink / raw)
  To: akpm; +Cc: linux-kernel

So far, no crash.  But now I have NFS clients that from time to time are unable
to access this server.  The server has the following messages on it:

Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted

I can temporarily fix the problem by typing:

exportfs -ar

But eventually it happens again.

-- 
 Dr. Harry Edmon			E-MAIL: harry@atmos.washington.edu
 206-543-0547				harry@u.washington.edu
 Dept of Atmospheric Sciences		FAX:	206-543-0308
 University of Washington, Box 351640, Seattle, WA 98195-1640

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: Re[4]: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-31 22:21       ` Re[4]: " Harry Edmon
@ 2004-08-31 22:47         ` Andrew Morton
  2004-09-16 21:01           ` Harry Edmon
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2004-08-31 22:47 UTC (permalink / raw)
  To: Harry Edmon; +Cc: linux-kernel

Harry Edmon <harry@atmos.washington.edu> wrote:
>
> So far, no crash.  But now I have NFS clients that from time to time are unable
> to access this server.  The server has the following messages on it:
> 
> Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted
> 
> I can temporarily fix the problem by typing:
> 
> exportfs -ar
> 
> But eventually it happens again.
> 

Well there were a few other NFS fixes.  Can you test the latest kernel
from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ?

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: PROBLEM: page allocation or what in 2.6.8.1
  2004-08-31 22:47         ` Andrew Morton
@ 2004-09-16 21:01           ` Harry Edmon
  0 siblings, 0 replies; 7+ messages in thread
From: Harry Edmon @ 2004-09-16 21:01 UTC (permalink / raw)
  To: Andrew Morton; +Cc: linux-kernel

Once I mounted "nfsd" and went to autofs4 all problems went away with 
2.6.9-rc1-bk9.  Thanks for the help.

Andrew Morton wrote:

>Harry Edmon <harry@atmos.washington.edu> wrote:
>  
>
>>So far, no crash.  But now I have NFS clients that from time to time are unable
>>to access this server.  The server has the following messages on it:
>>
>>Aug 31 15:16:43 funnel rpc.mountd: getfh failed: Operation not permitted
>>
>>I can temporarily fix the problem by typing:
>>
>>exportfs -ar
>>
>>But eventually it happens again.
>>
>>    
>>
>
>Well there were a few other NFS fixes.  Can you test the latest kernel
>from ftp://ftp.kernel.org/pub/linux/kernel/v2.6/snapshots/ ?
>  
>


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2004-09-16 21:01 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2004-08-25 14:51 PROBLEM: page allocation or what in 2.6.8.1 Harry Edmon
2004-08-26 10:31 ` Andrew Morton
2004-08-31 16:36   ` Re[2]: " Harry Edmon
2004-08-31 19:02     ` Andrew Morton
2004-08-31 22:21       ` Re[4]: " Harry Edmon
2004-08-31 22:47         ` Andrew Morton
2004-09-16 21:01           ` Harry Edmon

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®