From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1760847AbYGJTKw (ORCPT ); Thu, 10 Jul 2008 15:10:52 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S1754750AbYGJTKo (ORCPT ); Thu, 10 Jul 2008 15:10:44 -0400 Received: from qb-out-0506.google.com ([72.14.204.225]:36796 "EHLO qb-out-0506.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1754629AbYGJTKn (ORCPT ); Thu, 10 Jul 2008 15:10:43 -0400 DomainKey-Signature: a=rsa-sha1; c=nofws; d=gmail.com; s=gamma; h=message-id:date:from:to:subject:cc:in-reply-to:mime-version :content-type:content-transfer-encoding:content-disposition :references; b=O5EjDCZ0Zs+uh85uXe41ZLHg/DAfvgB+egJ5AWful1XaPfFsj9a9f02lvzPLTMHK4b 2a/nHeR0sTMbnNZNGri+vq3vpsURmnZdJMkrD5YmXZO2ZMMAllWLY6dpSxkZEGkIaIq3 O+ei45zFCppBerobP0sEqwwq55GfpBdqP3iao= Message-ID: <19f34abd0807101210t4043de1en59cd2e64b13d259a@mail.gmail.com> Date: Thu, 10 Jul 2008 21:10:41 +0200 From: "Vegard Nossum" To: "Zhang, Yanmin" Subject: Re: v2.6.26-rc7: BUG: unable to handle kernel NULL pointer dereference Cc: "Rusty Russell" , "Mike Travis" , "Adrian Bunk" , "Srivatsa Vaddagiri" , linux-kernel@vger.kernel.org, "Gautham R Shenoy" , "Rafael J. Wysocki" , "Zhang, Yanmin" , "Heiko Carstens" In-Reply-To: <1214294783.25608.75.camel@ymzhang> MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Disposition: inline References: <20080622125633.GA8166@damson.getinternet.no> <200806231326.11328.rusty@rustcorp.com.au> <485FD644.80208@sgi.com> <200806241136.52430.rusty@rustcorp.com.au> <1214294783.25608.75.camel@ymzhang> Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Content-Transfer-Encoding: 8bit X-MIME-Autoconverted: from base64 to 8bit by alpha.home.local id m6AJAufq004339 On Tue, Jun 24, 2008 at 10:06 AM, Zhang, Yanmin wrote:> In function _cpu_up, the panic happens when calling __raw_notifier_call_chain> at the second time. Kernel doesn't panic when calling it at the first time. If> just say because of nr_cpu_ids, that's not right.>> By checking source codes, I find function do_boot_cpu is the culprit.> Consider below call chain:> _cpu_up=>__cpu_up=>smp_ops.cpu_up=>native_cpu_up=>do_boot_cpu.>> So do_boot_cpu is called in the end. In do_boot_cpu, if boot_error==true,> cpu_clear(cpu, cpu_possible_map) is executed. So later on, when _cpu_up> calls __raw_notifier_call_chain at the second time to report CPU_UP_CANCELED,> because this cpu is already cleared from cpu_possible_map, get_cpu_sysdev returns> NULL.>> Many resources are related to cpu_possible_map, so it's better not to change it.>> Below patch against 2.6.26-rc7 fixes it by removing the bit clearing in cpu_possible_map.>> Vegard, would you like to help test it? Yay! I just hit this again with your patch applied Inquiring remote APIC #1...... APIC #1 ID: failed... APIC #1 VERSION: failed... APIC #1 SPIV: failed and it works correctly, no NULL pointer error this time :-) I know it's applied to mainline already, actually with this tag too,even though I wasn't able to confirm it before: Tested-by: Vegard Nossum Thanks :-) Vegard -- "The animistic metaphor of the bug that maliciously sneaked in whilethe programmer was not looking is intellectually dishonest as itdisguises that the error is the programmer's own creation." -- E. W. Dijkstra, EWD1036{.n++%ݶw{.n+{G{ayʇڙ,jfhz_(階ݢj"mG?&~iOzv^m ?I