From: Stephen Brennan <stephen.s.brennan@oracle.com>
To: "Guilherme G. Piccoli" <gpiccoli@igalia.com>,
Zack Rusin <zack.rusin@broadcom.com>,
Petr Mladek <pmladek@suse.com>
Cc: Borislav Petkov <bp@alien8.de>,
Ajay Kaher <ajay.kaher@broadcom.com>,
Alexey Makhalov <alexey.makhalov@broadcom.com>,
x86@kernel.org, Joel Granados <joel.granados@kernel.org>,
Baoquan He <baoquan.he@linux.dev>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Dave Hansen <dave.hansen@linux.intel.com>,
"H . Peter Anvin" <hpa@zytor.com>,
virtualization@lists.linux.dev,
bcm-kernel-feedback-list@broadcom.com,
linux-kernel@vger.kernel.org,
John Ogness <john.ogness@linutronix.de>,
Steven Rostedt <rostedt@goodmis.org>,
Sergey Senozhatsky <senozhatsky@chromium.org>,
Kees Cook <kees@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Mike Rapoport <rppt@kernel.org>,
Pasha Tatashin <pasha.tatashin@soleen.com>,
Pratyush Yadav <pratyush@kernel.org>,
Dave Young <ruirui.yang@linux.dev>,
Jonathan Corbet <corbet@lwn.net>, Bo Gan <bo.gan@broadcom.com>,
Brennan Lamoreaux <brennan.lamoreaux@broadcom.com>,
kexec@lists.infradead.org, linux-doc@vger.kernel.org
Subject: Re: [PATCH v1 4/4] x86/vmware: Run panic diagnostics before kdump by default
Date: Mon, 28 Sep 2026 10:35:10 -0700 [thread overview]
Message-ID: <87v77p9t4h.fsf@oracle.com> (raw)
In-Reply-To: <bc0bdd14-5e6a-0676-cd04-14fb0bb93cd5@igalia.com>
"Guilherme G. Piccoli" <gpiccoli@igalia.com> writes:
> This Message Is From an External Sender
> This message came from outside your organization.
> Report Suspicious
>
> Hi Petr, Zack - thanks for CCing me!
> Some comments below:
>
>
> On 18/09/2026 00:23, Zack Rusin wrote:
>>> [...]
>>> Maybe, we should start with something simple, and introduce
>>> one more panic notifier as a start. It might be called either:
>>>
>>> + "panic_hypervisor_list" because "crash_kexec_post_notifiers = true"
>>> seems to be primary set on hypervisors.
>>>
>>> But I would rather make it more generic and call it
>>>
>>> + panic_pre_crash_kexec or panic_pre_kdump because there might be
>>> more notifiers which are either 100% safe and useful or are worth
>>> the risk before calling crash dump.
>>>
>>> We could put there x86/vmware notifiers as a start. And we could later
>>> move there other important notifiers.
>>>
>>> How does that sound, please?
>>
>
> It's a good idea, IMO. We could start with this, Zach commented some
> implementation details below...and after it gets merged, we could move
> other hypervisors that currently set "crash_kexec_post_notifiers" to
> this list and eventually, unexport this symbol. We should avoid having
> code forcing this parameter, as Petr said, many notifiers are executed
> if that is set.
>
> (I'm CCing Stephen Brennan here, I recall he had problems with this
> being auto-set, we talked about that in the panic notifiers big
> discussions in the past heh)
Hi Guilherme, thanks for CCing me! You're right, I'm interested too :D
Your email arrived on the first day of my vacation, so I'm sorry for the
delayed response.
I like the above idea, starting simple with a second list for the
notifiers which are straightforward and have an urgent need to run.
My previous complaints were with code overriding
crash_kexec_post_notifiers, because that really interferes with your
mental model when debugging panic/kdump issues. Having two notifier
lists is a reasonable mental model, or at least more understandable than
silently and unconditionally modifying crash_kexec_post_notifiers.
I like the name "panic_pre_kdump" for my own tastes, though n==1 on
that.
I do think we're going to have trouble articulating exactly what are the
criteria for a "pre_kdump" notifier vs a "panic" notifier? Anyone
submitting a patch believes their notifier is important, so of course
they believe it belongs on the pre_kdump list! How will we draw the line
in the future?
- Is it safety? As in "the notifier doesn't take any locks"? I feel
this should be part of the criteria. (Though I have never been
successful in predicting what code would cause a kdump/panic failure)
- Is it urgency? As in "without this, the expected behavior of the
system during a panic will not occur".
- Is it the fact that the notifier is communicating with the hypervisor?
From my perspective: informing the hypervisor of a panic is both safe
and urgent. If the hypervisor is configured to immediately halt and do
its own panic handling, great.
Providing the last 4KB of logs seems less safe and less urgent (though
obviously still immensely helpful). The kernel provides a huge variety
of consoles through which the hypervisor could have already received
these logs. If the system is configured not to use them... it seems the
tradeoff to not be debuggable has already been made :P
To me, that logic belongs in a standard panic notifier.
Plenty of people will likely disagree with me on that, but that's why it
would be helpful to have clarity on what separates pre_kdump from
standard panic notifiers.
> The only thing I'd like to suggest: I think we should have a parameter
> that disables running this list, which would be the opposite of
> "crash_kexec_post_notifiers".
>
> I would implement it as something like: "postpone_pre_kexec_notifiers"
> or something like that. The parameter would basically "move" this list
> execution to the same time as the current notifiers, gating them to
> "crash_kexec_post_notifiers". This way, we'd allow users to debug kexec
> failures maybe related to the "early" notifiers. WDYT?
Agreed for precisely the reason I said above: it's difficult to predict
which notifier is going to cause a problem. Having an escape hatch to
change this behavior without rebuilding a production kernel is critical,
IMO.
Thanks,
Stephen
>
>> [...]
>> I think that without that default though, x86 oops_end() can enter
>> crash_kexec(regs) before reaching panic(), for example with
>> panic_on_oops=1. To cover that path too, I'd call the chain from
>> __crash_kexec() after the image check and register capture, under the
>> existing kexec lock. A second call in vpanic(), immediately before
>> kmsg_dump_desc(), would cover the fallback path. And I think a
>> set-once guard would prevent duplicate or recursive dispatch.
>>
>
> Regarding this, 2 things:
>
> a) I think you could change kexec_should_crash() to "return 0" also in
> case the new list is set to run, the same is done currently for
> "crash_kexec_post_notifiers". Makes sense?
>
> b) Well, does this whole panic diag thing you're implementing here aims
> only at x86 guests ? Or would it be possible to run, for example, arm64
> guests? Asking this because in x86 and some other architectures (but not
> arm64[0]), it's possible to override machine_crash_shutdown() handler,
> and run things prior to a kexec. Take a look on how Hyper-V does that on
> arch/x86 - this could be just what you need, except if you plan to have
> it for all architectures heh
>
> Finally, if possible please keep me looped in the following patches, I'm
> very interested on that =)
> Cheers,
>
>
> Guilherme
>
>
> [0]
> https://lore.kernel.org/r/427a8277-49f0-4317-d6c3-4a15d7070e55@igalia.com/
next prev parent reply other threads:[~2026-09-28 17:36 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-08 18:07 [PATCH v1 0/4] x86/vmware: Preserve panic diagnostics in vmware.log Zack Rusin
2026-09-08 18:07 ` [PATCH v1 1/4] x86/vmware: Add a bounded panic log sender Zack Rusin
2026-09-21 13:59 ` Michael Kelley
2026-09-08 18:07 ` [PATCH v1 2/4] x86/vmware: Add the vmware_record_panic_msg sysctl Zack Rusin
2026-09-10 8:31 ` Joel Granados
2026-09-10 11:44 ` Zack Rusin
2026-09-16 15:46 ` Zack Rusin
2026-09-08 18:07 ` [PATCH v1 3/4] x86/vmware: Report guest crashes after kmsg dumpers Zack Rusin
2026-09-08 18:07 ` [PATCH v1 4/4] x86/vmware: Run panic diagnostics before kdump by default Zack Rusin
2026-09-17 9:07 ` Petr Mladek
2026-09-18 3:23 ` Zack Rusin
2026-09-18 14:21 ` Guilherme G. Piccoli
2026-09-18 23:00 ` Zack Rusin
2026-09-20 21:15 ` Guilherme G. Piccoli
2026-09-21 8:23 ` Petr Mladek
2026-09-28 17:35 ` Stephen Brennan [this message]
2026-09-17 23:07 ` [PATCH v1 0/4] x86/vmware: Preserve panic diagnostics in vmware.log Maaz Mombasawala
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=87v77p9t4h.fsf@oracle.com \
--to=stephen.s.brennan@oracle.com \
--cc=ajay.kaher@broadcom.com \
--cc=akpm@linux-foundation.org \
--cc=alexey.makhalov@broadcom.com \
--cc=baoquan.he@linux.dev \
--cc=bcm-kernel-feedback-list@broadcom.com \
--cc=bo.gan@broadcom.com \
--cc=bp@alien8.de \
--cc=brennan.lamoreaux@broadcom.com \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=gpiccoli@igalia.com \
--cc=hpa@zytor.com \
--cc=joel.granados@kernel.org \
--cc=john.ogness@linutronix.de \
--cc=kees@kernel.org \
--cc=kexec@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=pasha.tatashin@soleen.com \
--cc=pmladek@suse.com \
--cc=pratyush@kernel.org \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ruirui.yang@linux.dev \
--cc=senozhatsky@chromium.org \
--cc=tglx@kernel.org \
--cc=virtualization@lists.linux.dev \
--cc=x86@kernel.org \
--cc=zack.rusin@broadcom.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®