From: James Morse <james.morse@arm.com>
To: Borislav Petkov <bp@alien8.de>, David Arcari <darcari@redhat.com>
Cc: "Rafael J. Wysocki" <rjw@rjwysocki.net>,
Linux ACPI <linux-acpi@vger.kernel.org>,
Lenny Szubowicz <lszubowi@redhat.com>,
Len Brown <lenb@kernel.org>, Tony Luck <tony.luck@intel.com>,
"Eric W. Biederman" <ebiederm@xmission.com>,
Alexandru Gagniuc <mr.nuke.me@gmail.com>,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH] ACPI/APEI: Clear GHES block_status before panic()
Date: Fri, 21 Dec 2018 18:52:20 +0000 [thread overview]
Message-ID: <0bb80989-4fe5-c320-8ffc-0f39502110c9@arm.com> (raw)
In-Reply-To: <1691682.ckC9UJvgEr@aspire.rjw.lan>
On 21/12/2018 11:17, Rafael J. Wysocki wrote:
> On Thursday, December 20, 2018 8:24:47 PM CET Borislav Petkov wrote:
>> + James.
Thanks,
>> On Wed, Dec 19, 2018 at 11:50:52AM -0500, David Arcari wrote:
>>> From: Lenny Szubowicz <lszubowi@redhat.com>
>>>
>>> In __ghes_panic() clear the block status in the APEI generic
>>> error status block for that generic hardware error source before
>>> calling panic() to prevent a second panic() in the crash kernel
>>> for exactly the same fatal error.
>>>
>>> Otherwise ghes_probe(), running in the crash kernel, would see
>>> an unhandled error in the APEI generic error status block and
>>> panic again, thereby precluding any crash dump.
I bet that was fun to watch!
>>> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
>>> index 02c6fd9..f008ba7 100644
>>> --- a/drivers/acpi/apei/ghes.c
>>> +++ b/drivers/acpi/apei/ghes.c
>>> @@ -691,6 +691,8 @@ static void __ghes_panic(struct ghes *ghes)
>>> {
>>> __ghes_print_estatus(KERN_EMERG, ghes->generic, ghes->estatus);
>>>
>>> + ghes_clear_estatus(ghes);
>>> +
>>> /* reboot to log the error! */
>>> if (!panic_timeout)
>>> panic_timeout = ghes_panic_timeout;
>>
>> Acked-by: Borislav Petkov <bp@suse.de>
>
> Patch applied, thanks!
Great!
Do we need to ghes_ack_error() too?
With the location cleared the new kernel will never find the records, and
firmware can never re-use that location because it wasn't ack'd. The upshot is
RAS records can't be generated for the kdump kernel. The acpi spec talks about
use of the memory, so I don't think its fair for it to use this to disarm a
watchdog.
I think we can live with this as the kdump kernel isn't going to handle RAS
errors for the bulk of memory anyway.
Thanks,
James
next prev parent reply other threads:[~2018-12-21 18:52 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2018-12-19 16:50 David Arcari
2018-12-20 18:28 ` Tyler Baicar
2018-12-20 19:24 ` Borislav Petkov
2018-12-21 11:17 ` Rafael J. Wysocki
2018-12-21 18:52 ` James Morse [this message]
2018-12-21 18:59 ` Borislav Petkov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=0bb80989-4fe5-c320-8ffc-0f39502110c9@arm.com \
--to=james.morse@arm.com \
--cc=bp@alien8.de \
--cc=darcari@redhat.com \
--cc=ebiederm@xmission.com \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lszubowi@redhat.com \
--cc=mr.nuke.me@gmail.com \
--cc=rjw@rjwysocki.net \
--cc=tony.luck@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®