[PATCH v3 13/31] gpu: nova-core: gsp: wait for GMC completion events

From: Zhi Wang

Date: Mon Sep 28 2026 - 06:34:41 EST


GSP plugin shutdown completes through a separate event carrying the
GFID. Releasing the queue guard after sending would let another receiver
consume that completion.

Keep the guard from sending through the matching event and use the
shared Inner receive deadline. Offer GMC events to the completion
predicate, pass unmatched events to the caller's handler, and consume
interleaved responses and RPC messages. Represent completion with a
value and all continuing cases with None.

Propagate callback errors after consuming the current valid element.
Document callback non-reentrancy and that the receive timeout starts
after sending, excluding mutex acquisition and command-queue space
waits.

Co-developed-by: Alok Kumar <alkumar@xxxxxxxxxx>
Signed-off-by: Alok Kumar <alkumar@xxxxxxxxxx>
Signed-off-by: Zhi Wang <zhiw@xxxxxxxxxx>
---
Documentation/gpu/nova/core/interrupts.rst | 6 ++-
drivers/gpu/nova-core/gsp/cmdq.rs | 55 ++++++++++++++++++++++
drivers/gpu/nova-core/gsp/fw.rs | 7 +++
3 files changed, 66 insertions(+), 2 deletions(-)

diff --git a/Documentation/gpu/nova/core/interrupts.rst b/Documentation/gpu/nova/core/interrupts.rst
index 45ba788bfe27..f78b1294905b 100644
--- a/Documentation/gpu/nova/core/interrupts.rst
+++ b/Documentation/gpu/nova/core/interrupts.rst
@@ -578,7 +578,8 @@ command reply or an unsolicited event, and the two differ in the function code.
arrival with its sequence number, function code, and length.
* A GMC message carries a command id in place of a function code. The
``GSP_INIT`` wait and synchronous GMC transactions claim responses with the
- expected command id and sequence number. These waits consume interleaved RPC
+ expected command id and sequence number. A GMC completion-event wait uses its
+ caller's predicate to select an event. These waits consume interleaved RPC
messages as events. Unmatched GMC messages are handled by the wait's callback
or logged and consumed; a queue drain with no waiting caller logs them at
warning level and drops them.
@@ -605,7 +606,8 @@ mutex. Replies and events share one queue and one read pointer, so one lock is
held across the whole drain. A thread waiting for a reply logs each event that
arrives before the reply and keeps waiting. One receive deadline applies to the
whole wait, rather than a fresh timeout after each message. The default timeout
-is 5 seconds; GMC transactions can specify another timeout. For send-and-wait operations the receive deadline starts after sending,
+is 5 seconds; GMC transactions and completion-event waits can specify another
+timeout. For send-and-wait operations the receive deadline starts after sending,
so it does not bound waiting for the mutex or for command-queue space. The thread
holds the mutex from sending through receiving, so no other caller consumes the
message it waits for.
diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
index df3d45a13a13..43038c7fe823 100644
--- a/drivers/gpu/nova-core/gsp/cmdq.rs
+++ b/drivers/gpu/nova-core/gsp/cmdq.rs
@@ -791,6 +791,61 @@ pub(crate) fn send_gmc_and_receive_timeout(
})
}

+ /// Sends an asynchronous GMC command and waits atomically for its event.
+ ///
+ /// The command queue remains locked from the send through the matching
+ /// event, preventing another transaction from consuming its completion.
+ /// GMC events are passed to `predicate`, then to `handler` if the predicate returns false.
+ /// GMC responses are debug-logged and consumed without invoking either callback. Interleaved
+ /// RPC messages are logged as events and consumed.
+ ///
+ /// One receive deadline starts after sending and is not extended by other messages. It does
+ /// not bound waiting for the mutex or for space to send the request.
+ ///
+ /// Both callbacks run with the queue locked and must not reenter this queue or reset it.
+ /// Their second argument is the raw `max_resp_or_status` word of the event header.
+ ///
+ /// # Errors
+ ///
+ /// - `EMSGSIZE` if the request exceeds the command queue's maximum element size.
+ /// - `ETIMEDOUT` if space does not become available to send the request, or if no event
+ /// satisfies `predicate` before the receive deadline.
+ /// - `EIO` if the command queue slot cannot hold the request headers, the receive queue is
+ /// poisoned, or a received element fails framing validation.
+ ///
+ /// Errors from initializing the request headers and from either callback are propagated
+ /// as-is. The current valid element is consumed before a callback error is returned.
+ #[expect(dead_code)]
+ pub(crate) fn send_gmc_and_wait_event(
+ &self,
+ command_id: u32,
+ payload: &[u8],
+ timeout: Delta,
+ mut predicate: impl FnMut(u32, u32, u64, &[u8], &[u8]) -> Result<bool>,
+ mut handler: impl FnMut(u32, u32, u64, &[u8], &[u8]) -> Result,
+ ) -> Result {
+ let mut inner = self.inner.lock();
+ inner.send_gmc(command_id, payload, 0)?;
+ let deadline = Instant::<Monotonic>::now() + timeout;
+
+ inner.await_gmc(deadline, |header, payload_0, payload_1| {
+ let header = &header.gmc;
+ if header.is_response() {
+ return Ok(None);
+ }
+
+ let command = header.command_id();
+ let max_resp_or_status = header.raw_status_word();
+ let sequence = header.sequence;
+ if predicate(command, max_resp_or_status, sequence, payload_0, payload_1)? {
+ return Ok(Some(()));
+ }
+
+ handler(command, max_resp_or_status, sequence, payload_0, payload_1)?;
+ Ok(None)
+ })
+ }
+
/// Sends a GMC API request that GSP-RM does not answer.
///
/// # Errors
diff --git a/drivers/gpu/nova-core/gsp/fw.rs b/drivers/gpu/nova-core/gsp/fw.rs
index 71732418e7a5..a416826737bb 100644
--- a/drivers/gpu/nova-core/gsp/fw.rs
+++ b/drivers/gpu/nova-core/gsp/fw.rs
@@ -802,6 +802,13 @@ pub(crate) fn sequence_number(&self) -> u64 {
self.sequence & !GMC_EVENT_SEQUENCE_BASE
}

+ /// Returns the raw `max_resp_or_status` word carried by a GMC header.
+ ///
+ /// Its meaning for an event is defined by that event's command.
+ pub(crate) fn raw_status_word(&self) -> u32 {
+ self.max_resp_or_status
+ }
+
/// Returns the `NV_STATUS` that a response carries.
///
/// The value is meaningful only when [`Self::is_response`] is `true`. In a request, the same