From: "Alexandre Courbot" <acourbot@nvidia.com>
To: "John Hubbard" <jhubbard@nvidia.com>
Cc: "Danilo Krummrich" <dakr@kernel.org>,
"Timur Tabi" <ttabi@nvidia.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Eliot Courtney" <ecourtney@nvidia.com>,
"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Miguel Ojeda" <ojeda@kernel.org>,
"Alex Gaynor" <alex.gaynor@gmail.com>,
"Boqun Feng" <boqun.feng@gmail.com>,
"Gary Guo" <gary@garyguo.net>,
"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
"Benno Lossin" <lossin@kernel.org>,
"Andreas Hindborg" <a.hindborg@kernel.org>,
"Alice Ryhl" <aliceryhl@google.com>,
"Trevor Gross" <tmgross@umich.edu>,
nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing
Date: Mon, 31 Aug 2026 14:35:58 +0900 [thread overview]
Message-ID: <DL2VVZMCMKFJ.3VSCVJRWO5HKR@nvidia.com> (raw)
In-Reply-To: <20260829013324.499542-15-jhubbard@nvidia.com>
On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
<...>
> @@ -748,11 +756,13 @@ fn send_command<M>(&mut self, bar: Bar0<'_>, command: M) -> Result<u32>
> /// # Errors
> ///
> /// - `ETIMEDOUT` if `timeout` has elapsed before any message becomes available.
> - /// - `EIO` if there was some inconsistency (e.g. message shorter than advertised) on the
> - /// message queue.
> - ///
> - /// Error codes returned by the message constructor are propagated as-is.
> + /// - `EIO` if the framing or the checksum is invalid, or the queue was already poisoned by an
> + /// earlier such failure. Either failure poisons the queue, so recovery requires a reset.
> fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
> + if self.poisoned.get() {
> + return Err(EIO);
> + }
> +
> // Wait for a message to arrive from the GSP.
> let (slice_1, slice_2) = read_poll_timeout(
> || Ok(self.gsp_mem.driver_read_area()),
> @@ -763,7 +773,10 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
> .map(|(slice_1, slice_2)| (slice_1.as_flattened(), slice_2.as_flattened()))?;
>
> // Extract the `GspMsgElement`.
> - let (header, slice_1) = GspMsgElement::from_bytes_prefix(slice_1).ok_or(EIO)?;
> + let Some((header, slice_1)) = GspMsgElement::from_bytes_prefix(slice_1) else {
> + self.poisoned.set(true);
We probably want to `dev_err` some error here, or the queue will just
stop working without explanation.
> + return Err(EIO);
> + };
>
> dev_dbg!(
> &self.dev,
> @@ -777,6 +790,7 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
>
> // Check that the driver read area is large enough for the message.
> if slice_1.len() + slice_2.len() < payload_length {
> + self.poisoned.set(true);
Same here. Maybe we can have a `fn poison(&self, reason:
fmt::Arguments<'_>) -> Error` that performs the log and returns `EIO`
for convenience?
next prev parent reply other threads:[~2026-08-31 5:36 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-29 1:22 [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29 1:22 ` [PATCH v2 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-08-31 1:10 ` Alexandre Courbot
2026-08-29 1:22 ` [PATCH v2 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-08-31 1:10 ` Alexandre Courbot
2026-08-29 1:22 ` [PATCH v2 03/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-08-29 1:22 ` [PATCH v2 04/15] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-08-31 14:24 ` Alexandre Courbot
2026-09-01 13:16 ` Alexandre Courbot
2026-08-29 1:22 ` [PATCH v2 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-01 1:15 ` Alexandre Courbot
2026-08-29 1:25 ` [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29 1:35 ` John Hubbard
2026-08-29 1:33 ` [PATCH v2 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-01 7:03 ` Alexandre Courbot
2026-08-29 1:33 ` [PATCH v2 07/15] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-01 12:52 ` Alexandre Courbot
2026-08-29 1:33 ` [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them John Hubbard
2026-08-31 5:06 ` Alexandre Courbot
2026-08-29 1:33 ` [PATCH v2 09/15] gpu: nova-core: match GSP RPC replies by sequence, not just function John Hubbard
2026-08-31 1:09 ` Alexandre Courbot
2026-08-31 4:33 ` John Hubbard
2026-08-31 22:18 ` John Hubbard
2026-08-31 22:46 ` John Hubbard
2026-08-29 1:33 ` [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-08-31 5:35 ` Alexandre Courbot [this message]
2026-08-29 1:33 ` [PATCH v2 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-08-31 6:04 ` Alexandre Courbot
2026-08-29 1:33 ` [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-01 14:54 ` Alexandre Courbot
2026-09-01 15:08 ` Danilo Krummrich
2026-09-02 14:33 ` Alexandre Courbot
2026-09-03 3:06 ` John Hubbard
2026-08-29 1:33 ` [PATCH v2 13/15] gpu: nova-core: retrigger the GSP falcon and clear every latched cause John Hubbard
2026-09-02 15:00 ` Alexandre Courbot
2026-08-29 1:33 ` [PATCH v2 14/15] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-08-29 1:33 ` [PATCH v2 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
2026-09-02 15:07 ` Alexandre Courbot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DL2VVZMCMKFJ.3VSCVJRWO5HKR@nvidia.com \
--to=acourbot@nvidia.com \
--cc=a.hindborg@kernel.org \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=aliceryhl@google.com \
--cc=apopple@nvidia.com \
--cc=bhelgaas@google.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=dakr@kernel.org \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=jhubbard@nvidia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=tmgross@umich.edu \
--cc=ttabi@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®