From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261746AbUD1Tmf (ORCPT ); Wed, 28 Apr 2004 15:42:35 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261718AbUD1TlX (ORCPT ); Wed, 28 Apr 2004 15:41:23 -0400 Received: from mx1.redhat.com ([66.187.233.31]:52946 "EHLO mx1.redhat.com") by vger.kernel.org with ESMTP id S265040AbUD1SYv (ORCPT ); Wed, 28 Apr 2004 14:24:51 -0400 From: Jeff Moyer MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Transfer-Encoding: 7bit Message-ID: <16527.63123.869014.733258@segfault.boston.redhat.com> Date: Wed, 28 Apr 2004 14:23:15 -0400 To: Matt Mackall Cc: linux-kernel@vger.kernel.org Subject: Re: netconsole hangs w/ alt-sysrq-t In-Reply-To: <20040428142753.GE28459@waste.org> References: <16519.58589.773562.492935@segfault.boston.redhat.com> <20040425191543.GV28459@waste.org> <16527.42815.447695.474344@segfault.boston.redhat.com> <20040428140353.GC28459@waste.org> <16527.47765.286783.249944@segfault.boston.redhat.com> <20040428142753.GE28459@waste.org> X-Mailer: VM 7.14 under 21.4 (patch 13) "Rational FORTRAN" XEmacs Lucid Reply-To: jmoyer@redhat.com X-PGP-KeyID: 1F78E1B4 X-PGP-CertKey: F6FE 280D 8293 F72C 65FD 5A58 1FF8 A7CA 1F78 E1B4 X-PCLoadLetter: What the f**k does that mean? Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org ==> Regarding Re: netconsole hangs w/ alt-sysrq-t; Matt Mackall adds: [snip] mpm> Well process context defeats the purpose. Ok, I've more closely read mpm> your report and if I understand correctly, you're using the NAPI mpm> version of e100? There's some magic NAPI bits in netpoll_poll that mpm> might help here: >> Yes, sorry I didn't specify that earlier. >> mpm> if(trapped && np->dev->poll && test_bit(__LINK_STATE_RX_SCHED, mpm> &np->dev->state)) np-> dev->poll(np->dev, &budget); >> mpm> Perhaps we need to pull the trapped test out of there. Then with any mpm> luck, dev->hard_start_xmit will return non-zero in netpoll_send_skb, mpm> we'll call netpoll_poll to pump the card, and we'll be able to flush mpm> it. >> I don't think so. You can end up in code running in interrupt context >> that is not designed to (ip routing code, etc). I've been down that >> path already. I only defer to process context if irqs_disabled(). mpm> Fair enough. Turning on trapped basically short circuits the rest of mpm> the NAPI code so that stuff doesn't hit the stack when we call ->poll. mpm> Could you try doing a netpoll_set_trap(1)/(0) around the call to mpm> ->poll and see if that actually lets the thing work? Then we can try mpm> to figure out the right way to do this. Okay, I tried that before and I still ran into problems. Just to be sure, I grabbed 2.6.6-rc3 and tried your suggestion. The resulting stack trace is at the end of this message. I can't explain how we get from netif_receive_skb to anywhere up the stack, though. Perhaps you'll have better luck with it. Note that this is alt-sysrq-t output, so the first bits of stack you see are not relevant. What counts is everything after the Badness in local_bh_enable. mpm> My point about deferring still stands, as when we're dumping an oops, mpm> we can't really expect that we'll ever get to a process context. Sure, but you can explicitly flush the queue in such a circumstance, though this is still not an ideal solution. -Jeff pdflush S 00000000 0 9 4 10 8 (L-TLB) cfe03f7c 00000046 00000000 00000000 00000000 00000000 cfe03f48 c011fa9c 5ec54655 123c321a 00000000 00000002 cff725d0 cff736b0 cff736d0 c1244ce0 0000119b 4318ade1 00000006 cfe12650 cfe12808 00000046 00000000 00000003 Call Trace: [] recalc_task_prio+0x8c/0x1a0 [] pdflush+0x0/0x20 [] __pdflush+0xb8/0x320 [] pdflush+0x1a/0x20 [] pdflush+0x0/0x20 [] Badness in local_bh_enable at kernel/softirq.c:136 Call Trace: [] local_bh_enable+0x84/0x90 [] neigh_lookup+0x7e/0xb0 [] arp_process+0x200/0x530 [] rebalance_tick+0x3b/0xf0 [] netif_receive_skb+0x1d3/0x280 [] e100_poll+0x6e7/0x760 [e100] [] e100_xmit_frame+0x284/0x400 [e100] [] e100_poll+0x0/0x760 [e100] [] netpoll_poll+0x5e/0x60 [] netpoll_send_skb+0xc3/0x120 [] write_msg+0x58/0x60 [netconsole] [] write_msg+0x0/0x60 [netconsole] [] __call_console_drivers+0x56/0x60 [] release_console_sem+0x77/0x140 [] printk+0x1c0/0x270 [] pdflush+0x0/0x20 [] kthread+0x9c/0xb0 [] show_trace+0xa5/0xc0 [] kthread+0x9c/0xb0 [] show_stack+0x7b/0x90 [] show_state+0x5c/0x90 [] __handle_sysrq_nolock+0x6e/0xe6 [] handle_sysrq+0x40/0x60 [] kbd_event+0x32/0x70 [] input_event+0xf5/0x400 [] atkbd_report_key+0x2f/0x80 [] atkbd_interrupt+0x223/0x530 [] serio_interrupt+0x66/0x70 [] i8042_interrupt+0xf4/0x190 [] handle_IRQ_event+0x30/0x60 [] do_IRQ+0x121/0x290 ======================= [] common_interrupt+0x18/0x20 [] transmeta_identify+0x2b/0x60 [] apm_bios_call_simple+0x9f/0xe0 [] apm_do_idle+0x18/0x70 [] apm_cpu_idle+0x7c/0x150 [] default_idle+0x0/0x40 [] cpu_idle+0x46/0x50 [] start_kernel+0x19f/0x200 [] unknown_bootoption+0x0/0x130