From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S263671AbUECM5G (ORCPT ); Mon, 3 May 2004 08:57:06 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S263157AbUECM5G (ORCPT ); Mon, 3 May 2004 08:57:06 -0400 Received: from parcelfarce.linux.theplanet.co.uk ([195.92.249.252]:42209 "EHLO www.linux.org.uk") by vger.kernel.org with ESMTP id S263671AbUECM46 (ORCPT ); Mon, 3 May 2004 08:56:58 -0400 Date: Mon, 3 May 2004 09:58:29 -0300 From: Marcelo Tosatti To: Chris Stromsoe Cc: linux-kernel@vger.kernel.org, sct@redhat.com, davem@redhat.com Subject: Re: two lockups with 2.4.25 Message-ID: <20040503125829.GB29160@logos.cnet> References: Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.5.5.1i Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Tue, Apr 20, 2004 at 04:01:32PM -0700, Chris Stromsoe wrote: > I had two different machines lock up within hours of each other this > morning. They have the exact same hardware and had very similar uptimes > (within hours of each other). > > Both machines have a single SCSI disk on an onboard adaptec controller. > All partitions are ext3. There are 3 other disks arranged in a RAID5 > partition that are formatted JFS. > > The machines run as virus quarantine servers, writing about 50k messages > to disk per day. The files are written out into directories as > year/month/day/hour/filename. There are usually around 2k to 3k files per > directory. > > Both machines locked up hard. I was able to get sysrq+t output from one > machine, but nothing from the other (it has no serial console). I ran the > output through ksymoops. It's listed below. > > > > -Chris > > > ksymoops 2.4.5 on i686 2.4.25. Options used > -V (default) > -k /proc/ksyms (specified) > -l /proc/modules (specified) > -o /lib/modules/2.4.25/ (specified) > -m /boot/System.map-2.4.25 (specified) > > Pid: 26885, comm: sophie > EIP: 0010:[] CPU: 1 EFLAGS: 00000202 Not tainted > Using defaults from ksymoops -t elf32-i386 -a i386 > EAX: 00000001 EBX: 00000001 ECX: 00000202 EDX: 01000000 > ESI: f4762cc0 EDI: 081a1d28 EBP: e806fe94 DS: 0018 ES: 0018 > Warning (Oops_set_regs): garbage 'DS: 0018 ES: 0018' at end of register line ignored > CR0: 8005003b CR2: 081a1d28 CR3: 14c2c000 CR4: 00000690 > Call Trace: [] [] [] [] [] > [] [] [] > Warning (Oops_read): Code line not seen, dumping what data is available > > > >>EIP; c0110ed7 <===== > > >>EDX; 01000000 Before first symbol > >>ESI; f4762cc0 <_end+343c85cc/3896e90c> > >>EDI; 081a1d28 Before first symbol > >>EBP; e806fe94 <_end+27cd57a0/3896e90c> > > Trace; c011100f > Trace; c01259b7 > Trace; c01260de > Trace; c01132f9 > Trace; c0113158 > Trace; c0224ce1 <__kfree_skb+129/134> > Trace; c01145e3 > Trace; c0106fc4 The thing is flush_tlb_others can't block. It calls invlpg. > init R C3FFBF28 0 1 0 18517 (NOTLB) > Call Trace: [] [] [] [] [] > [] > sophie R current 0 26885 12367 26884 (NOTLB) > Call Trace: [] > mimedefang R DA151F28 4 26886 12405 26887 26879 (NOTLB) > Call Trace: [] [] [] [] [] > mimedefang R EF435F28 2408 26887 12405 26888 26886 (NOTLB) > Call Trace: [] [] [] [] [] > mimedefang R CC989F28 0 26888 12405 26889 26887 (NOTLB) > Call Trace: [] [] [] [] [] > mimedefang R E5A83F28 0 26889 12405 26890 26888 (NOTLB) > Call Trace: [] [] [] [] [] > mimedefang R F67E5F28 0 26890 12405 26889 (NOTLB) > Call Trace: [] [] [] [] [] > Warning (Oops_read): Code line not seen, dumping what data is available > > >>EIP; c3fa2000 <_end+3c0790c/3896e90c> <===== > > Trace; c0194c68 > Trace; c0105680 > Proc; jfsCommit > > >>EIP; f70c0cc0 <_end+36d265cc/3896e90c> <===== > > Trace; c019776b > Trace; c0197960 > Trace; c0105680 > Proc; jfsSync > > >>EIP; c0366180 <===== > > Trace; c0197ed3 > Trace; c0105680 > Proc; ahc_dv_0 This might be a problem in JFS? shaggy, can you take a look at this? > Trace; c0248577 > Trace; c0224b3f > Trace; c0224b56 > Trace; c0224ce1 <__kfree_skb+129/134> > Trace; c024c369 > Trace; c0131497 <__free_pages+1f/24> > Trace; c02539fe > Trace; c0223a93 > Trace; c0244711 > Trace; c014c8ec > Trace; c014b37b > Trace; c01218f6 > Trace; c010c605 > Trace; c0106ed3 > Proc; apache > > >>EIP; 00000001 Before first symbol <===== > > Trace; c019cd05 > Trace; f8d0c83d <[eepro100]speedo_start_xmit+16d/1f4> > Trace; c0231403 > Trace; c022648d > Trace; c0224b3f > Trace; c0224b56 > Trace; c0224ce1 <__kfree_skb+129/134> > Trace; c0253926 > Trace; c0223a93 > Trace; c0244711 > Trace; c014c8ec > Trace; c014b37b > Trace; c01218f6 > Trace; c010c605 > Trace; c0106ed3 > Proc; mimedefang There's also a lot of network activity.