From: Zhi Wang <zhiw@nvidia.com>
To: <dakr@kernel.org>, <acourbot@nvidia.com>
Cc: <alex@shazbot.org>, <jgg@nvidia.com>, <yishaih@nvidia.com>,
<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
<airlied@gmail.com>, <simona@ffwll.ch>, <ojeda@kernel.org>,
<alex.gaynor@gmail.com>, <boqun.feng@gmail.com>,
<gary@garyguo.net>, <bjorn3_gh@protonmail.com>,
<lossin@kernel.org>, <a.hindborg@kernel.org>,
<aliceryhl@google.com>, <tmgross@umich.edu>,
<jhubbard@nvidia.com>, <ecourtney@nvidia.com>, <cjia@nvidia.com>,
<smitra@nvidia.com>, <kjaju@nvidia.com>, <alkumar@nvidia.com>,
<ankita@nvidia.com>, <aniketa@nvidia.com>, <kwankhede@nvidia.com>,
<targupta@nvidia.com>, <nova-gpu@lists.linux.dev>,
<linux-kernel@vger.kernel.org>, <zhiwang@kernel.org>,
Zhi Wang <zhiw@nvidia.com>
Subject: [PATCH v3 12/31] gpu: nova-core: gsp: add synchronous GMC transactions
Date: Mon, 28 Sep 2026 13:28:15 +0300 [thread overview]
Message-ID: <aee97e22c6a87984e449c5d4e89bcde4e48a77dc.1790580105.git.zhiw@nvidia.com> (raw)
In-Reply-To: <cover.1790580105.git.zhiw@nvidia.com>
vGPU management queries need an owned response payload after the receive
queue releases its slot. Sending and receiving must share one queue
guard so another receiver cannot consume the reply.
Match the response flag, command ID and request sequence, then copy the
payload and return the raw firmware status. Support the default receive
timeout and a caller-selected timeout without converting firmware
rejection into a transport error.
Share the receive deadline and retry loop in CmdqInner with the existing
GMC boot response wait. Preserve each caller's matching and status
policy, consume every valid element once even when its callback fails,
and keep a single deadline after sending. Document unmatched-message
consumption and distinguish a zero response payload capacity from a
command that sends no reply.
Co-developed-by: Alok Kumar <alkumar@nvidia.com>
Signed-off-by: Alok Kumar <alkumar@nvidia.com>
Signed-off-by: Zhi Wang <zhiw@nvidia.com>
---
Documentation/gpu/nova/core/interrupts.rst | 17 +-
drivers/gpu/nova-core/gsp/cmdq.rs | 174 +++++++++++++++++----
drivers/gpu/nova-core/gsp/fw.rs | 6 +-
3 files changed, 161 insertions(+), 36 deletions(-)
diff --git a/Documentation/gpu/nova/core/interrupts.rst b/Documentation/gpu/nova/core/interrupts.rst
index fa124bc8ca05..45ba788bfe27 100644
--- a/Documentation/gpu/nova/core/interrupts.rst
+++ b/Documentation/gpu/nova/core/interrupts.rst
@@ -576,9 +576,12 @@ command reply or an unsolicited event, and the two differ in the function code.
records, lifecycle notices) need no action and get no line of their own,
because the receive trace at debug level already records every message's
arrival with its sequence number, function code, and length.
-* A GMC message carries a command id in place of a function code. Only the
- ``GSP_INIT`` wait during boot claims GMC messages, so one that arrives
- anywhere else is logged at warning level and dropped.
+* A GMC message carries a command id in place of a function code. The
+ ``GSP_INIT`` wait and synchronous GMC transactions claim responses with the
+ expected command id and sequence number. These waits consume interleaved RPC
+ messages as events. Unmatched GMC messages are handled by the wait's callback
+ or logged and consumed; a queue drain with no waiting caller logs them at
+ warning level and drops them.
A command's reply must carry the RPC sequence number that nova-core wrote into
the command, as well as its function code. A message with the awaited function
@@ -600,9 +603,11 @@ the device is reset.
The polling path and the IRQ thread both read the queue under the command-queue
mutex. Replies and events share one queue and one read pointer, so one lock is
held across the whole drain. A thread waiting for a reply logs each event that
-arrives before the reply and keeps waiting. One deadline of 5 seconds applies
-to the whole wait, rather than a fresh timeout after each message, and the
-thread holds the mutex for the whole wait, so no other caller consumes the
+arrives before the reply and keeps waiting. One receive deadline applies to the
+whole wait, rather than a fresh timeout after each message. The default timeout
+is 5 seconds; GMC transactions can specify another timeout. For send-and-wait operations the receive deadline starts after sending,
+so it does not bound waiting for the mutex or for command-queue space. The thread
+holds the mutex from sending through receiving, so no other caller consumes the
message it waits for.
With one lock, a drain waits for an in-flight command's receive to finish or
diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
index c63cd450df41..df3d45a13a13 100644
--- a/drivers/gpu/nova-core/gsp/cmdq.rs
+++ b/drivers/gpu/nova-core/gsp/cmdq.rs
@@ -548,6 +548,16 @@ fn payload_length(&self) -> Option<usize> {
}
}
+/// Response from a GMC API command.
+pub(crate) struct GmcResponse {
+ /// Response status (`NV_STATUS` code). Zero means success.
+ #[expect(dead_code)]
+ pub(crate) status: u32,
+ /// Response payload copied out of the message queue.
+ #[expect(dead_code)]
+ pub(crate) payload: KVec<u8>,
+}
+
/// GSP command queue.
///
/// Provides the ability to send commands and receive messages from the GSP using a shared memory
@@ -688,6 +698,99 @@ pub(crate) fn send_gmc_no_wait(
.send_gmc(command_id, payload, max_response_size)
}
+ /// Sends a GMC API command and waits for its matching response.
+ ///
+ /// Uses [`Self::RECEIVE_TIMEOUT`] as the receive timeout. Firmware rejection is returned in
+ /// [`GmcResponse::status`], not as an error.
+ ///
+ /// See [`Self::send_gmc_and_receive_timeout`] for message consumption and locking behavior.
+ ///
+ /// # Errors
+ ///
+ /// Returns the same errors as [`Self::send_gmc_and_receive_timeout`].
+ #[expect(dead_code)]
+ pub(crate) fn send_gmc_and_receive(
+ &self,
+ command_id: u32,
+ payload: &[u8],
+ max_response_size: u32,
+ ) -> Result<GmcResponse> {
+ self.send_gmc_and_receive_timeout(
+ command_id,
+ payload,
+ max_response_size,
+ Self::RECEIVE_TIMEOUT,
+ )
+ }
+
+ /// Sends a GMC API command and waits up to `timeout` for its matching response.
+ ///
+ /// Matches both the command ID and the request's sequence number. A nonzero firmware status
+ /// is returned in [`GmcResponse::status`]; the caller decides how to handle rejection.
+ /// `max_response_size` advertises the response payload capacity in bytes, so zero still
+ /// permits a status-only response.
+ ///
+ /// Unmatched GMC responses and events are debug-logged and consumed. Interleaved RPC messages
+ /// are logged as events and consumed. The matching response is also consumed if copying its
+ /// payload fails.
+ ///
+ /// The queue stays locked from sending through receiving. One receive deadline starts after
+ /// sending and is not extended by other messages. It does not bound waiting for the mutex or
+ /// for space to send the request.
+ ///
+ /// # Errors
+ ///
+ /// - `EMSGSIZE` if the request exceeds the command queue's maximum element size.
+ /// - `ETIMEDOUT` if space does not become available to send the request, or if the matching
+ /// response does not arrive before the receive deadline.
+ /// - `EIO` if the command queue slot cannot hold the request headers, the receive queue is
+ /// poisoned, or a received element fails framing validation.
+ /// - `ENOMEM` if the response payload cannot be allocated.
+ ///
+ /// Errors from initializing the request headers are propagated as-is.
+ pub(crate) fn send_gmc_and_receive_timeout(
+ &self,
+ command_id: u32,
+ payload: &[u8],
+ max_response_size: u32,
+ timeout: Delta,
+ ) -> Result<GmcResponse> {
+ let mut inner = self.inner.lock();
+ let expected_sequence = inner.send_gmc(command_id, payload, max_response_size)?;
+ let dev = inner.dev;
+
+ let deadline = Instant::<Monotonic>::now() + timeout;
+ inner.await_gmc(deadline, |header, payload_0, payload_1| {
+ let header = &header.gmc;
+ if !header.is_response_to(command_id, expected_sequence) {
+ let kind = if header.is_response() {
+ "response"
+ } else {
+ "event"
+ };
+ dev_dbg!(
+ dev,
+ "GSP GMC: skip {} seq {} cmd {:#x}; want response seq {} cmd {:#x}\n",
+ kind,
+ header.sequence,
+ header.command_id(),
+ expected_sequence,
+ command_id,
+ );
+ return Ok(None);
+ }
+
+ // Each byte slice is at most `isize::MAX` bytes, so their sum fits in `usize`.
+ let mut payload = KVec::with_capacity(payload_0.len() + payload_1.len(), GFP_KERNEL)?;
+ payload.extend_from_slice(payload_0, GFP_KERNEL)?;
+ payload.extend_from_slice(payload_1, GFP_KERNEL)?;
+ Ok(Some(GmcResponse {
+ status: header.status(),
+ payload,
+ }))
+ })
+ }
+
/// Sends a GMC API request that GSP-RM does not answer.
///
/// # Errors
@@ -1395,6 +1498,35 @@ fn receive_gmc_and_dispatch<R>(
})
}
+ /// Waits until `handler` returns a value for a GMC element, using one receive deadline.
+ ///
+ /// Unclaimed elements do not extend the deadline. Every valid element is consumed, including
+ /// one for which `handler` returns an error; RPC elements are logged as events and consumed.
+ /// The caller retains the queue guard throughout the wait and the handler calls.
+ ///
+ /// # Errors
+ ///
+ /// - `ETIMEDOUT` if no GMC element satisfies `handler` before `deadline`.
+ /// - `EIO` if the queue is poisoned or an element fails framing validation.
+ ///
+ /// Errors from `handler` are propagated as-is.
+ fn await_gmc<R>(
+ &mut self,
+ deadline: Instant<Monotonic>,
+ mut handler: impl FnMut(&GspGmcMsgElement, &[u8], &[u8]) -> Result<Option<R>>,
+ ) -> Result<R> {
+ loop {
+ let remaining = deadline - Instant::<Monotonic>::now();
+ if remaining.is_negative() {
+ return Err(ETIMEDOUT);
+ }
+
+ if let Some(value) = self.receive_gmc_and_dispatch(remaining, &mut handler)? {
+ return Ok(value);
+ }
+ }
+ }
+
/// Waits for the response to the GMC request with command id `command_id` and RPC sequence
/// number `sequence`, up to [`Cmdq::RECEIVE_TIMEOUT`] from the call.
///
@@ -1420,35 +1552,23 @@ fn await_gmc_response<R>(
) -> Result<R> {
let dev = self.dev;
let deadline = Instant::<Monotonic>::now() + Cmdq::RECEIVE_TIMEOUT;
- loop {
- let remaining = deadline - Instant::<Monotonic>::now();
- if remaining.is_negative() {
- break Err(ETIMEDOUT);
+ self.await_gmc(deadline, |header, payload_0, payload_1| {
+ if !header.gmc.is_response_to(command_id, sequence) {
+ return on_other(header, payload_0, payload_1).map(|()| None);
}
- let response =
- self.receive_gmc_and_dispatch(remaining, |header, payload_0, payload_1| {
- if !header.gmc.is_response_to(command_id, sequence) {
- return on_other(header, payload_0, payload_1).map(|()| None);
- }
-
- let status = header.gmc.status();
- if status != 0 {
- dev_err!(
- dev,
- "GSP GMC: command 0x{:x} failed, status={:#x}\n",
- command_id,
- status
- );
- return Err(EIO);
- }
-
- decode(payload_0, payload_1).map(Some)
- })?;
-
- if let Some(response) = response {
- break Ok(response);
+ let status = header.gmc.status();
+ if status != 0 {
+ dev_err!(
+ dev,
+ "GSP GMC: command 0x{:x} failed, status={:#x}\n",
+ command_id,
+ status
+ );
+ return Err(EIO);
}
- }
+
+ decode(payload_0, payload_1).map(Some)
+ })
}
}
diff --git a/drivers/gpu/nova-core/gsp/fw.rs b/drivers/gpu/nova-core/gsp/fw.rs
index 90bb324666e5..71732418e7a5 100644
--- a/drivers/gpu/nova-core/gsp/fw.rs
+++ b/drivers/gpu/nova-core/gsp/fw.rs
@@ -814,7 +814,7 @@ pub(crate) fn status(&self) -> u32 {
/// sequence number `sequence`.
pub(crate) fn is_response_to(&self, command_id: u32, sequence: u32) -> bool {
self.is_response()
- && self.command_id() == command_id
+ && self.command_id() == (command_id & GMCAPI_COMMAND_ID_MASK)
&& self.sequence == u64::from(sequence)
}
}
@@ -911,8 +911,8 @@ impl GspGmcMsgElement {
/// Creates the queue element header and the GMC API header of a request that carries
/// `payload_size` bytes of payload under the RPC sequence number `sequence`.
///
- /// `max_response_size` is the largest response that the sender accepts, and zero for a request
- /// that GSP-RM does not answer.
+ /// `max_response_size` is the response payload capacity in bytes. Zero permits a status-only
+ /// response; whether GSP-RM sends a response at all is determined by the command's protocol.
///
/// # Errors
///
next prev parent reply other threads:[~2026-09-28 10:31 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 10:28 [PATCH v3 00/31] Introduce NVIDIA vGPU manager Zhi Wang
2026-09-28 10:28 ` [PATCH v3 01/31] gpu: nova-core: gsp: pass boot context through setup helpers Zhi Wang
2026-09-28 10:28 ` [PATCH v3 02/31] gpu: nova-core: gsp: decouple boot context from VgpuManager Zhi Wang
2026-09-28 10:28 ` [PATCH v3 03/31] gpu: nova-core: vgpu: detect boot state independently Zhi Wang
2026-09-28 10:28 ` [PATCH v3 04/31] gpu: nova-core: gpu: add a channel ID pool for vGPU Zhi Wang
2026-09-28 10:28 ` [PATCH v3 05/31] gpu: nova-core: gsp: decode the FIFO engine table Zhi Wang
2026-09-28 10:28 ` [PATCH v3 06/31] gpu: nova-core: vgpu: initialize runtime parameters after GSP boot Zhi Wang
2026-09-28 10:28 ` [PATCH v3 07/31] gpu: nova-core: vgpu: reserve the 48-VM WPR2 heap Zhi Wang
2026-09-28 10:28 ` [PATCH v3 08/31] gpu: nova-core: mm: borrow BarUser for temporary BAR1 access Zhi Wang
2026-09-28 10:28 ` [PATCH v3 09/31] gpu: nova-core: mm: add VramBlock Zhi Wang
2026-09-28 10:28 ` [PATCH v3 10/31] gpu: nova-core: mm: add VramRegion Zhi Wang
2026-09-28 10:28 ` [PATCH v3 11/31] gpu: nova-core: mm: add BarMapping Zhi Wang
2026-09-28 10:28 ` Zhi Wang [this message]
2026-09-28 10:28 ` [PATCH v3 13/31] gpu: nova-core: gsp: wait for GMC completion events Zhi Wang
2026-09-28 10:28 ` [PATCH v3 14/31] gpu: nova-core: vgpu: add r000 plugin bindings Zhi Wang
2026-09-28 10:28 ` [PATCH v3 15/31] gpu: nova-core: vgpu: add VRAM slot allocator Zhi Wang
2026-09-28 10:28 ` [PATCH v3 16/31] gpu: nova-core: gsp: factor out NVKV payload conversion Zhi Wang
2026-09-28 10:28 ` [PATCH v3 17/31] gpu: nova-core: vgpu: query VF assignments and properties Zhi Wang
2026-09-28 10:28 ` [PATCH v3 18/31] gpu: nova-core: vgpu: add instance create/destroy Zhi Wang
2026-09-28 10:28 ` [PATCH v3 19/31] gpu: nova-core: vgpu: encode vGPU boot requests Zhi Wang
2026-09-28 10:28 ` [PATCH v3 20/31] gpu: nova-core: vgpu: add GSP plugin communication buffers Zhi Wang
2026-09-28 10:28 ` [PATCH v3 21/31] gpu: nova-core: vgpu: add instance boot Zhi Wang
2026-09-28 10:28 ` [PATCH v3 22/31] gpu: nova-core: vgpu: add instance shutdown Zhi Wang
2026-09-28 10:28 ` [PATCH v3 23/31] gpu: nova-core: vgpu: initialize GSP plugin RPC buffers Zhi Wang
2026-09-28 10:28 ` [PATCH v3 24/31] gpu: nova-core: vgpu: add GSP plugin RPC transactions Zhi Wang
2026-09-28 10:28 ` [PATCH v3 25/31] gpu: nova-core: vgpu: negotiate the GSP plugin RPC version Zhi Wang
2026-09-28 10:28 ` [PATCH v3 26/31] gpu: nova-core: vgpu: send GSP plugin configuration parameters Zhi Wang
2026-09-28 10:28 ` [PATCH v3 27/31] gpu: nova-core: vgpu: update the GSP plugin BME state Zhi Wang
2026-09-28 10:28 ` [PATCH v3 28/31] gpu: nova-core: vgpu: add CeUtils commands Zhi Wang
2026-09-28 10:28 ` [PATCH v3 29/31] gpu: nova-core: vgpu: scrub guest VRAM with CeUtils Zhi Wang
2026-09-28 10:28 ` [PATCH v3 30/31] gpu: nova-core: vgpu: export plugin log buffers via debugfs Zhi Wang
2026-09-28 10:28 ` [PATCH v3 31/31] gpu: nova-core: vgpu: introduce SR-IOV PF APIs Zhi Wang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aee97e22c6a87984e449c5d4e89bcde4e48a77dc.1790580105.git.zhiw@nvidia.com \
--to=zhiw@nvidia.com \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=alex@shazbot.org \
--cc=aliceryhl@google.com \
--cc=alkumar@nvidia.com \
--cc=aniketa@nvidia.com \
--cc=ankita@nvidia.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=cjia@nvidia.com \
--cc=dakr@kernel.org \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=jgg@nvidia.com \
--cc=jhubbard@nvidia.com \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kwankhede@nvidia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=skolothumtho@nvidia.com \
--cc=smitra@nvidia.com \
--cc=targupta@nvidia.com \
--cc=tmgross@umich.edu \
--cc=yishaih@nvidia.com \
--cc=zhiwang@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®