From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1757866AbYJQVs7 (ORCPT ); Fri, 17 Oct 2008 17:48:59 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S1756999AbYJQVsU (ORCPT ); Fri, 17 Oct 2008 17:48:20 -0400 Received: from wf-out-1314.google.com ([209.85.200.171]:1998 "EHLO wf-out-1314.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1756944AbYJQVsT (ORCPT ); Fri, 17 Oct 2008 17:48:19 -0400 Message-ID: <13c67e2c0810171448o6858827ei1ccc9e0ddf487f8@mail.gmail.com> Date: Fri, 17 Oct 2008 14:48:18 -0700 From: "Ani Sinha" To: linux-kernel@vger.kernel.org, torvalds@linux-foundation.org Subject: panic() logic Cc: kernel@anirban.org MIME-Version: 1.0 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: 7bit Content-Disposition: inline Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi All: I noticed an issue with the panic() firing on a back core in SMP lately. We are mostly working on mips architectures but it might effect other archs as well. Therefore, I am putting forward my thoughts and comments to the whole linux community. In the following, by front core I mean core#0 and by back core I mean other cores. The panic() call does a smp_send_stop() pretty early in the call process. There is an issue with this: smp_send_stop basically marks all the other cores as 'down' and updates the cpu bitmap. One implication of this is that you can not do an IPI later on to other cores (smp_send_function() does a 'for_earch_online_cpu'). This makes sense since you should not be allowed to do anything on a down cpu. But what if a particular architecture had logic to do specific things for the front core and other things on the back cores as a part of 'graceful reboot' process? For example, the arch dependent emergency_restart() might try to do send IPI to the front core to do a graceful reboot and simply place the others on an infinite loop? As a concrete example, mips sibyte processor has logic in arch dependent restart code to call cfe_exit() on front core and an infinite loop on the back cores. It does this through an IPI which obviously does not succeed because of the early smp_send_stop(). So, my proposal is, can we just run panic on the front core and put the back cores in loop right away. I mean, can we do this in pseudo code? panic() { /* do stuff that must be done on the current core */ on_each_cpu(continue_panic, ...); } continue_panic() { if (!smp_processor_id()) { smp_send_stop(); /* rest of existing panic() logic */ }else