From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f54.google.com (mail-wm1-f54.google.com [209.85.128.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 676722BD035 for ; Mon, 12 Jan 2026 16:24:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768235079; cv=none; b=dKgql75NeiGwpD364qBcmjdKACqJG/EKtUc5T4C+vZj3VRjO2ERFZsHT5DnPC/jBoYIKF1kjbu4uRVous9UsItWJ6ZybHFYFN0UGT2wnutxp9uf5Vyp/sa2Kt8ePAR4CI3PpI6+54ynracSkMX2lY5v8Y4B9fh7bkOrnRzKLeAQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768235079; c=relaxed/simple; bh=yQc7xxiFEzNbsrFcwFRVgPOJ3KERRxk+2YYAVH5H/40=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Tk0sdmd93kaAii5/MvBRfgafm9itP6P1o2goAxZBV9Pbvq/M85qyL9rNx61sXx0r33NOMVMwSK67CqOyiOpTosIAPV+E8ivHLkJ0C+uwj6d9gXF3+p2+y8KHrJpde87eCMrztFFqbNAL8Rj8nr2oQSVh38V923LMTBySR1bGoVM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=D7XRrDhe; arc=none smtp.client-ip=209.85.128.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="D7XRrDhe" Received: by mail-wm1-f54.google.com with SMTP id 5b1f17b1804b1-47795f6f5c0so40747495e9.1 for ; Mon, 12 Jan 2026 08:24:37 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1768235076; x=1768839876; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=2444t0YQQiBZl36bNcGtq7Feua62iZIWvVRaQJE+FwM=; b=D7XRrDheLuaSXODXC4aBdD901bYvTF2bcaRWLNQKqjHTob8N+YN19y1M3hx+8YK0l3 zNq8TQGHvJ9LXnoUMktzi+bc9/jdn5YPGBv1G6ZJrOnAvNhk+elWLVFRwRXgsQl1Tg8Y j4Lmq64bpTluQVA2ag8aov6JVFTWHO3Ck1q+KjdC75FU56FAemnDYQT/0GrF3SY92MfN 3AQqY5p8Gi8JdTrwJI14gf3t5PewF8mQkGhl6QAkm5NGFMBPt25ckFkO3sebRvLzuTro Ii0wmg5d3AQ90JVdMs/kpebL8KzJikQQoGrukmE8/iucA3XkDnfDAFdJr9FtBMmOT2WO tEEg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768235076; x=1768839876; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=2444t0YQQiBZl36bNcGtq7Feua62iZIWvVRaQJE+FwM=; b=ZfuofQ0Ry+lmpS/NstyrMGR1yWdfJsH7HQvMQe3sZ6Lw7toWjG8zgBK4evBO6/bQzz rm3Up+lKVCWFY4nPp3p7LVWfGLp9iA6vW5sw7jbhidNXA5QDolch9IPRw1Sbr448NgSB ByvsUBOf3PcpFIHr4SOInobRxvgI4+Wm9VssvM97DoIWTO42+gGhKtLl/dFXlC+bxkiS DmGEsnY3aw+G1EdIoxSD5dIZKJx8mwas8+S6fyirtLcEammF+PKAgn6vLErXffJ15Rq9 QIzZGN0/pBYDYxjaXmPoROJBWciCTUAvCUMxZABYWcVsmrl/v0Q9l1T8SDbbjPYvDb8M BlYA== X-Forwarded-Encrypted: i=1; AJvYcCVUb3YlIaHYS/nKSUdnWAe/aPAPFN0AzemkKEuj656oxqoxxNk8XvMZ1IkZGe+cWBZ3u84XuoglRZTKWD8=@vger.kernel.org X-Gm-Message-State: AOJu0YxOtsUrpBbxR/vlEwZoowEy5nJOxrAkGqO68Wu/tUp20tknS7St rMBE64Q3uTMaP5pwAFWKCgUzZcSt5udTfP948NtY4TbRETgAlyFo1h7LmHIUOuGQShU= X-Gm-Gg: AY/fxX5kW8GdMnOMK5ATzDmTEX/E3cTzSVjcoycH1TeEuVQCOpsKaYUBqsRiRopS+r8 K7SCkT00TLx5LSsIXJuOO1oca1hCdlAQ4fc5gn4GsFFo8yjGkv0H1R4B8FL4jq2AENHaRsjEUHS vGyNQHTZG9PI4kPaefwaMRiJ+4+jywbW2c2ZGhM1QEwJwRToeoxZnEI47PvFdPe5FHUWl3TqBS6 SntgPmKk/PPeymhbCUybuqc1NTeJxCcJt4V78rsNBoKqRO2jJHqG31geASHkTNDYssCKSEx4dS8 45O4qwxJo4xZMOp2nTHX/HcBeNkiKMP5Lg/I/hq5Inchi21BLY+vyPUfQqLOA+EptM/hf0TiLqw oE94DXN02mxe5++R/4a2Ajt/M0T24BtMXR1s+5XNgUnTgapIbIxjs2Q7NdyBMeCSHlltfxeCL1X RF7gCqQ4lI0i+S9w== X-Google-Smtp-Source: AGHT+IGmAFpcGJILBJZUBMdCfBa1w5TDMkDkuphhV6DuxPBJUPospmk89YHeurbAbCB6le/o0PHoqw== X-Received: by 2002:a05:600c:4fc6:b0:477:7b9a:bb0a with SMTP id 5b1f17b1804b1-47d84b54c52mr213113025e9.21.1768235075555; Mon, 12 Jan 2026 08:24:35 -0800 (PST) Received: from pathway.suse.cz ([176.114.240.130]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-432bd0e6784sm40182186f8f.19.2026.01.12.08.24.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 12 Jan 2026 08:24:35 -0800 (PST) Date: Mon, 12 Jan 2026 17:24:28 +0100 From: Petr Mladek To: Pnina Feder Cc: akpm@linux-foundation.org, bhe@redhat.com, linux-kernel@vger.kernel.org, lkp@intel.com, mgorman@suse.de, mingo@redhat.com, peterz@infradead.org, rostedt@goodmis.org, senozhatsky@chromium.org, tglx@linutronix.de, vkondra@mobileye.com Subject: Re: [PATCH v6] panic: add panic_force_cpu= parameter to redirect panic to a specific CPU Message-ID: References: <20260111123656.1563887-1-pnina.feder@mobileye.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260111123656.1563887-1-pnina.feder@mobileye.com> On Sun 2026-01-11 14:36:56, Pnina Feder wrote: > Some platforms require panic handling to execute on a specific CPU for > crash dump to work reliably. This can be due to firmware limitations, > interrupt routing constraints, or platform-specific requirements where > only a single CPU is able to safely enter the crash kernel. > > Add the panic_force_cpu= kernel command-line parameter to redirect panic > execution to a designated CPU. When the parameter is provided, the CPU > that initially triggers panic forwards the panic context to the target > CPU via IPI, which then proceeds with the normal panic and kexec flow. > > The IPI delivery is implemented as a weak function (panic_smp_redirect_cpu) > so architectures with NMI support can override it for more reliable delivery. > > If the specified CPU is invalid, offline, or a panic is already in > progress on another CPU, the redirection is skipped and panic continues > on the current CPU. > > --- a/kernel/panic.c > +++ b/kernel/panic.c > @@ -300,6 +300,121 @@ void __weak crash_smp_send_stop(void) > > atomic_t panic_cpu = ATOMIC_INIT(PANIC_CPU_INVALID); > > +#if defined(CONFIG_SMP) && defined(CONFIG_CRASH_DUMP) > +/* CPU to redirect panic to, or -1 if disabled */ > +static int panic_force_cpu = -1; > + > +static int __init panic_force_cpu_setup(char *str) > +{ > + int cpu; > + > + if (!str) > + return -EINVAL; > + > + if (kstrtoint(str, 0, &cpu) || cpu < 0) { > + pr_warn("panic_force_cpu: invalid value '%s'\n", str); > + return -EINVAL; > + } > + > + panic_force_cpu = cpu; > + return 0; > +} > +early_param("panic_force_cpu", panic_force_cpu_setup); > + > +static void do_panic_on_target_cpu(void *info) > +{ > + panic("%s", (char *)info); > +} > + > +/** > + * panic_smp_redirect_cpu - Redirect panic to target CPU > + * @target_cpu: CPU that should handle the panic > + * @msg: formatted panic message > + * > + * Default implementation uses IPI. Architectures with NMI support > + * can override this for more reliable delivery. > + * > + * Return: 0 on success, negative errno on failure > + */ > +int __weak panic_smp_redirect_cpu(int target_cpu, void *msg) > +{ > + static call_single_data_t panic_csd; > + > + panic_csd.func = do_panic_on_target_cpu; > + panic_csd.info = msg; > + > + return smp_call_function_single_async(target_cpu, &panic_csd); > +} > + > +/** > + * panic_force_target_cpu - Redirect panic to a specific CPU for crash kernel > + * @buf: buffer to format the panic message into > + * @buf_size: size of the buffer > + * @fmt: panic message format string > + * @args: arguments for format string > + * > + * Some platforms require panic handling to occur on a specific CPU > + * for the crash kernel to function correctly. This function redirects > + * panic handling to the CPU specified via the panic_redirect_cpu= boot parameter. > + * > + * Returns true if panic should proceed on current CPU. > + * Returns false (never returns) if panic was redirected. > + */ > +__printf(3, 0) > +static bool panic_force_target_cpu(char *buf, int buf_size, const char *fmt, va_list args) > +{ > + int cpu = raw_smp_processor_id(); > + int target_cpu = panic_force_cpu; What is the reason to read the value into a local variable? If the reason was to avoid a race then READ_ONCE() should be used. Otherwise, it looks a bit misleading to use so different names. Maybe, rename the function to panic_try_force_cpu() use the global variable. Also, please invert the logic. The function should return "false" when it was not redirected (logical failure). > + /* Feature not enabled via boot parameter */ > + if (target_cpu < 0) > + return true; > + > + /* Already on target CPU - proceed normally */ > + if (cpu == target_cpu) > + return true; > + > + /* Target CPU is offline, can't redirect */ > + if (!cpu_online(target_cpu)) > + return true; > + > + /* Another panic already in progress */ > + if (panic_in_progress()) > + return true; > + > + vsnprintf(buf, buf_size, fmt, args); This is using a global buffer without any serialization. More CPUs might call panic()/panic_force_target_cpu() in parallel. The buffer might contain a mess as a result. I am afraid that we need a separate buffer. And only one CPU can be allowed to use it. We would need similar synchronization as with @panic_cpu for @panic_redirect_cpu. > + > + console_verbose(); > + bust_spinlocks(1); > + > + pr_emerg("panic: Redirecting from CPU %d to CPU %d for crash kernel\n", > + cpu, target_cpu); > + > + /* Dump original CPU's stack before redirecting */ > + if (test_taint(TAINT_DIE) || oops_in_progress > 1) { > + panic_this_cpu_backtrace_printed = true; > + } else if (IS_ENABLED(CONFIG_DEBUG_BUGVERBOSE)) { > + dump_stack(); > + panic_this_cpu_backtrace_printed = true; > + } The "panic_this_cpu_backtrace_printed" variable is checked in panic_trigger_all_cpu_backtrace() to see whether we want to print backtrace for this CPU or not. panic_smp_redirect_cpu() is going to call panic() on another CPU. Do we want to print backtrace from the other CPU? I guess, not. We should make the other panic() aware that it was redirected from here. Maybe, using the @panic_redirect_cpu variable which I suggested above to synchronize the access to the helper buffer. And panic() should do something like: if (panic_redirect_cpu >= 0 && panic_force_cpu == raw_smp_processor_id()) { /* Backtrace was printed on the original CPU. */ pr_emerg("panic: Redirected from CPU %d to CPU %d\n", panic_redirect_cpu, panic_force_cpu); } else if (test_taint(TAINT_DIE) || oops_in_progress > 1) { panic_this_cpu_backtrace_printed = true; } else if (IS_ENABLED(CONFIG_DEBUG_BUGVERBOSE)) { dump_stack(); panic_this_cpu_backtrace_printed = true; } Also we might need to check @panic_regirect_cpu in panic_trigger_all_cpu_backtrace() and skip this particular CPU there. > + > + printk_legacy_allow_panic_sync(); > + console_flush_on_panic(CONSOLE_FLUSH_PENDING); > + > + if (panic_smp_redirect_cpu(target_cpu, buf) != 0) > + return true; > + > + /* IPI/NMI sent, this CPU should stop */ > + return false; > +} > +#else > +__printf(3, 0) > +static inline bool panic_force_target_cpu(char *buf, int buf_size, const char *fmt, va_list args) > +{ > + return true; > +} > +#endif /* CONFIG_SMP && CONFIG_CRASH_DUMP */ > + > bool panic_try_start(void) > { > int old_cpu, this_cpu; > @@ -451,6 +566,13 @@ void vpanic(const char *fmt, va_list args) > local_irq_disable(); > preempt_disable_notrace(); > > + /* > + * Redirect panic to target CPU if configured via panic_force_cpu=. > + * Returns false and never returns if panic was redirected. The 2nd sentence is confusing. IMHO, panic_smp_self_stop() always returns. The point is that this CPU should stop itself when panic() was redirected. But wait! The panic_cpu will eventually do smp_send_stop(). On x86_64, it would call native_stop_other_cpus(). It woult wait until this CPU clears the related bit in cpus_stop_mask(). But it would never when when this CPU already spins in panic_smp_self_stop(). Or do I miss anything, please? IMHO, panic_smp_self_stop() can't be used here. Or we need to make stop_other_cpus() aware that this one is already stopped. Sigh, it is getting complicated. > + */ > + if (!panic_force_target_cpu(buf, sizeof(buf), fmt, args)) > + panic_smp_self_stop(); > + > /* > * It's possible to come here directly from a panic-assertion and > * not have preempt disabled. Some functions called from here want Best Regards, Petr