From: John Hubbard <jhubbard@nvidia.com>
To: Danilo Krummrich <dakr@kernel.org>,
Alexandre Courbot <acourbot@nvidia.com>
Cc: "Timur Tabi" <ttabi@nvidia.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Eliot Courtney" <ecourtney@nvidia.com>,
"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Miguel Ojeda" <ojeda@kernel.org>,
"Alex Gaynor" <alex.gaynor@gmail.com>,
"Boqun Feng" <boqun.feng@gmail.com>,
"Gary Guo" <gary@garyguo.net>,
"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
"Benno Lossin" <lossin@kernel.org>,
"Andreas Hindborg" <a.hindborg@kernel.org>,
"Alice Ryhl" <aliceryhl@google.com>,
"Trevor Gross" <tmgross@umich.edu>,
nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
"John Hubbard" <jhubbard@nvidia.com>
Subject: [PATCH v4 12/17] gpu: nova-core: bound a GSP wait by a single deadline
Date: Fri, 11 Sep 2026 21:43:55 -0700 [thread overview]
Message-ID: <20260912044400.677097-13-jhubbard@nvidia.com> (raw)
In-Reply-To: <20260912044400.677097-1-jhubbard@nvidia.com>
The GSP posts unsolicited events on the same queue as command replies,
so a caller waiting for one message consumes whatever arrives first and
reads again.
Every read started a fresh five-second timeout, so a steady stream of
events extended the wait without bound. The two boot-time waits for an
unsolicited event also released the queue mutex between reads, so a
command sent from another thread could consume the event and leave the
waiter to time out.
Compute one deadline when the wait begins and pass the time remaining to
each read, and hold the queue mutex across the whole wait. Put the loop
in one helper that the command reply wait and both boot-time event waits
share.
Assisted-by: LLM
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
---
drivers/gpu/nova-core/gsp/cmdq.rs | 73 ++++++++++++++++++++------
drivers/gpu/nova-core/gsp/commands.rs | 8 +--
drivers/gpu/nova-core/gsp/sequencer.rs | 8 +--
3 files changed, 60 insertions(+), 29 deletions(-)
diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
index 45c3c3aea8f9..4595aa2176e5 100644
--- a/drivers/gpu/nova-core/gsp/cmdq.rs
+++ b/drivers/gpu/nova-core/gsp/cmdq.rs
@@ -32,7 +32,11 @@
},
Mutex, //
},
- time::Delta,
+ time::{
+ Delta,
+ Instant,
+ Monotonic, //
+ },
transmute::{
AsBytes,
FromBytes, //
@@ -134,7 +138,9 @@ fn size(&self) -> usize {
/// Trait representing messages received from the GSP.
///
-/// This trait tells [`Cmdq::receive_msg`] how it can receive a given type of message.
+/// A reply that [`Cmdq::send_command`] waits for, or an event that [`Cmdq::await_msg`] waits for.
+/// The receiver matches a message's function code against [`Self::FUNCTION`] and decodes the
+/// message with [`Self::read`].
pub(crate) trait MessageFromGsp: Sized {
/// Function identifying this message from the GSP.
const FUNCTION: MsgFunction;
@@ -569,8 +575,9 @@ fn notify_gsp(bar: Bar0<'_>) {
///
/// # Errors
///
- /// - `ETIMEDOUT` if space does not become available to send the command, or if the reply is
- /// not received within the timeout.
+ /// - `ETIMEDOUT` if space does not become available to send the command, or if the reply does
+ /// not arrive within [`Self::RECEIVE_TIMEOUT`] of the send, however many events arrive
+ /// while waiting.
/// - `EIO` if the variable payload requested by the command has not been entirely
/// written to by its [`CommandToGsp::init_variable_payload`] method.
///
@@ -585,13 +592,7 @@ pub(crate) fn send_command<M>(&self, bar: Bar0<'_>, command: M) -> Result<M::Rep
let mut inner = self.inner.lock();
inner.send_command(bar, command)?;
- loop {
- match inner.receive_msg::<M::Reply>(Self::RECEIVE_TIMEOUT) {
- Ok(reply) => break Ok(reply),
- Err(ENOMSG) => continue,
- Err(e) => break Err(e),
- }
- }
+ inner.await_msg()
}
/// Sends `command` to the GSP without waiting for a reply.
@@ -611,15 +612,25 @@ pub(crate) fn send_command_no_wait<M>(&self, bar: Bar0<'_>, command: M) -> Resul
self.inner.lock().send_command(bar, command)
}
- /// Receive a message from the GSP.
+ /// Waits for an unsolicited GSP event of type `M`. Events that arrive before it are logged and
+ /// consumed.
+ ///
+ /// The queue mutex is held for the whole wait, up to [`Self::RECEIVE_TIMEOUT`], so no other
+ /// caller can send a command or consume an event meanwhile.
///
- /// See [`CmdqInner::receive_msg`] for details.
- pub(crate) fn receive_msg<M: MessageFromGsp>(&self, timeout: Delta) -> Result<M>
+ /// # Errors
+ ///
+ /// - `ETIMEDOUT` if the event does not arrive within [`Self::RECEIVE_TIMEOUT`] of the call,
+ /// however many other events arrive while waiting.
+ /// - `EIO` if the queue is poisoned, or if a message fails framing or checksum validation.
+ ///
+ /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
+ pub(crate) fn await_msg<M: MessageFromGsp>(&self) -> Result<M>
where
// This allows all error types, including `Infallible`, to be used for `M::InitError`.
Error: From<M::InitError>,
{
- self.inner.lock().receive_msg(timeout)
+ self.inner.lock().await_msg()
}
}
@@ -901,6 +912,38 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
result
}
+ /// Receives a message of type `M`, waiting up to [`Cmdq::RECEIVE_TIMEOUT`] from the call.
+ ///
+ /// Any other message that arrives first is logged as an event and does not extend the
+ /// deadline.
+ ///
+ /// # Errors
+ ///
+ /// - `ETIMEDOUT` if no message of type `M` arrives before the deadline, however many other
+ /// messages arrive while waiting.
+ /// - `EIO` if the queue is poisoned or a message fails framing or checksum validation (see
+ /// [`Self::wait_for_msg`]).
+ ///
+ /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
+ fn await_msg<M: MessageFromGsp>(&mut self) -> Result<M>
+ where
+ // This allows all error types, including `Infallible`, to be used for `M::InitError`.
+ Error: From<M::InitError>,
+ {
+ let deadline = Instant::<Monotonic>::now() + Cmdq::RECEIVE_TIMEOUT;
+ loop {
+ let remaining = deadline - Instant::<Monotonic>::now();
+ if remaining.is_negative() {
+ break Err(ETIMEDOUT);
+ }
+ match self.receive_msg::<M>(remaining) {
+ Ok(msg) => break Ok(msg),
+ Err(ENOMSG) => continue,
+ Err(e) => break Err(e),
+ }
+ }
+ }
+
/// Logs an event, meaning a message that no caller was waiting for.
///
/// An OS error or robust-channel record is logged at error level and an unknown function code
diff --git a/drivers/gpu/nova-core/gsp/commands.rs b/drivers/gpu/nova-core/gsp/commands.rs
index d1c80cf3c452..a01ab14299c6 100644
--- a/drivers/gpu/nova-core/gsp/commands.rs
+++ b/drivers/gpu/nova-core/gsp/commands.rs
@@ -188,13 +188,7 @@ fn read(
/// Waits for GSP initialization to complete.
pub(crate) fn wait_gsp_init_done(cmdq: &Cmdq<'_>) -> Result {
- loop {
- match cmdq.receive_msg::<GspInitDone>(Cmdq::RECEIVE_TIMEOUT) {
- Ok(_) => break Ok(()),
- Err(ENOMSG) => continue,
- Err(e) => break Err(e),
- }
- }
+ cmdq.await_msg::<GspInitDone>().map(|_| ())
}
/// The `GetGspStaticInfo` command.
diff --git a/drivers/gpu/nova-core/gsp/sequencer.rs b/drivers/gpu/nova-core/gsp/sequencer.rs
index 1782ed7d7ca6..250adc9fe74f 100644
--- a/drivers/gpu/nova-core/gsp/sequencer.rs
+++ b/drivers/gpu/nova-core/gsp/sequencer.rs
@@ -343,13 +343,7 @@ pub(crate) fn run(
libos: &'a Coherent<'a, [LibosMemoryRegionInitArgument]>,
bootloader_app_version: u32,
) -> Result {
- let seq_info = loop {
- match cmdq.receive_msg::<GspSequence>(Cmdq::RECEIVE_TIMEOUT) {
- Ok(seq_info) => break seq_info,
- Err(ENOMSG) => continue,
- Err(e) => return Err(e),
- }
- };
+ let seq_info = cmdq.await_msg::<GspSequence>()?;
let sequencer = GspSequencer {
bar: ctx.bar,
--
2.55.0
next prev parent reply other threads:[~2026-09-12 4:44 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 4:43 [PATCH v4 00/17] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-12 4:43 ` [PATCH v4 01/17] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-12 4:43 ` [PATCH v4 02/17] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-12 4:43 ` [PATCH v4 03/17] gpu: nova-core: add the GIN vector, leaf and subtree types John Hubbard
2026-09-12 4:43 ` [PATCH v4 04/17] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-12 4:43 ` [PATCH v4 05/17] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-12 4:43 ` [PATCH v4 06/17] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-12 4:43 ` [PATCH v4 07/17] gpu: nova-core: wait for GFW boot in probe, not in the Gpu constructor John Hubbard
2026-09-12 4:43 ` [PATCH v4 08/17] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-12 4:43 ` [PATCH v4 09/17] gpu: nova-core: log GSP events instead of discarding them John Hubbard
2026-09-12 4:43 ` [PATCH v4 10/17] gpu: nova-core: stop re-parsing a bad GSP message John Hubbard
2026-09-12 4:43 ` [PATCH v4 11/17] gpu: nova-core: return ENOMSG for an unmatched " John Hubbard
2026-09-12 4:43 ` John Hubbard [this message]
2026-09-12 4:43 ` [PATCH v4 13/17] gpu: nova-core: add a GSP message queue drain John Hubbard
2026-09-12 4:43 ` [PATCH v4 14/17] gpu: nova-core: add the falcon interrupt registers and their HAL John Hubbard
2026-09-12 4:43 ` [PATCH v4 15/17] gpu: nova-core: service GSP events from the SWGEN0 interrupt John Hubbard
2026-09-12 4:43 ` [PATCH v4 16/17] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-09-12 4:44 ` [PATCH v4 17/17] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260912044400.677097-13-jhubbard@nvidia.com \
--to=jhubbard@nvidia.com \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=aliceryhl@google.com \
--cc=apopple@nvidia.com \
--cc=bhelgaas@google.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=dakr@kernel.org \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=tmgross@umich.edu \
--cc=ttabi@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®