Re: [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them
From: Alexandre Courbot
Date: Mon Aug 31 2026 - 01:07:17 EST
On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
> The GSP posts unsolicited messages onto the same queue that carries
> command replies: logs, OS error and robust-channel records, and
> lifecycle notices.
>
> Anything that was not the reply a caller awaited was discarded, and an
> unrecognized function code aborted the in-flight command, so the GSP's
> error reports never reached the log.
>
> Route every non-reply message to a dispatcher, which logs the error
> records and leaves the in-flight command waiting for its reply. The
> dispatch runs on the existing command and wait loops, so events are
> handled during normal operation before any interrupt exists. Event
> payloads, such as XID numbers and log contents, are not decoded.
The naming looks a bit misleading to me - `dispatch_event` doesn't
really dispatch anything, it logs or ignores what is passed to it. If we
plan on implementing a real dispatch mechanism in the future I'm ok with
keeping the name, but we should document that intent in its doccomment.
>
> Assisted-by: Cursor:claude-opus-5
> Signed-off-by: John Hubbard <jhubbard@xxxxxxxxxx>
> ---
> drivers/gpu/nova-core/gsp/cmdq.rs | 64 ++++++++++++++++++++++++-------
> 1 file changed, 51 insertions(+), 13 deletions(-)
>
> diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
> index f0f28b6ded7a..0df52df1da89 100644
> --- a/drivers/gpu/nova-core/gsp/cmdq.rs
> +++ b/drivers/gpu/nova-core/gsp/cmdq.rs
> @@ -547,11 +547,11 @@ fn notify_gsp(bar: Bar0<'_>) {
>
> /// Sends `command` to the GSP and waits for the reply.
> ///
> - /// Messages with non-matching function codes are silently consumed until the expected reply
> - /// arrives.
> + /// A message read while waiting that is not the reply goes to
> + /// [`CmdqInner::dispatch_event`].
`CmdqInner` is an internal private type and an implementation detail, we
shouldn't reference it in public documentation. Let's say something like
"A message read while waiting that is logged if it is an error record,
and ignored otherwise".
> ///
> - /// The queue is locked for the entire send+receive cycle to ensure that no other command can
> - /// be interleaved.
> + /// The queue is locked for the entire send+receive cycle, so no other command can be
> + /// interleaved.
nit: is this hunk necessary?
> ///
> /// # Errors
> ///
> @@ -805,8 +805,10 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
>
> /// Receive a message from the GSP.
> ///
> - /// The expected message type is specified using the `M` generic parameter. If the pending
> - /// message has a different function code, `ERANGE` is returned and the message is consumed.
> + /// The expected message type is specified using the `M` generic parameter. A message whose
> + /// function code matches is decoded and returned. Any other message, whether its function code
> + /// is a different one or is unrecognized, goes to [`Self::dispatch_event`] and `ERANGE` is
> + /// returned.
> ///
> /// The read pointer is always advanced past the message, regardless of whether it matched.
> ///
> @@ -815,8 +817,7 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
> /// - `ETIMEDOUT` if `timeout` has elapsed before any message becomes available.
> /// - `EIO` if there was some inconsistency (e.g. message shorter than advertised) on the
> /// message queue.
> - /// - `EINVAL` if the function code of the message was not recognized.
> - /// - `ERANGE` if the message had a recognized but non-matching function code.
> + /// - `ERANGE` if the message was not the awaited reply.
> ///
> /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
> fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
> @@ -825,11 +826,13 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
> Error: From<M::InitError>,
> {
> let message = self.wait_for_msg(timeout)?;
> - let function = message.header.function().map_err(|_| EINVAL)?;
> + let function = message.header.function();
> + let seq = message.header.sequence();
> + let matched = matches!(function, Ok(f) if f == M::FUNCTION);
>
> - // Extract the message. Store the result as we want to advance the read pointer even in
> - // case of failure.
> - let result = if function == M::FUNCTION {
> + // Bind the result rather than returning early. The read pointer must advance past this
> + // message on every path.
Oooh nice catch, the queue was irreversibly unusable after an unknown
message is received on the previous code.
> + let result = if matched {
> let (cmd, contents_1) = M::Message::from_bytes_prefix(message.contents.0).ok_or(EIO)?;
> let mut sbuffer = SBufferIter::new_reader([contents_1, message.contents.1]);
>
> @@ -840,7 +843,7 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
> dev_warn!(
> &self.dev,
> "GSP message {:?} has unprocessed data\n",
> - function
> + M::FUNCTION
> );
> }
> })
> @@ -853,6 +856,41 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
> message.header.length().div_ceil(GSP_PAGE_SIZE),
> )?);
>
> + if !matched {
> + self.dispatch_event(function, seq);
> + }
We already have an `else` branch for the `if matched` statement above,
you can move this block there and avoid testing again.
It's also arguably (slightly) better since the message is now logged
before the fallible `u32::try_from` is called.
> +
> result
> }
> +
> + /// Routes a GSP message that is not the reply a caller is waiting for.
> + ///
> + /// GSP-reported errors are logged at error level and unrecognized function codes at warning
> + /// level. Every other known function code is consumed without a log line, because the RPC
> + /// receive trace in [`Self::wait_for_msg`] already records its arrival.
> + fn dispatch_event(&self, function: Result<MsgFunction, u32>, seq: u32) {
That's where we could describe the intent to turn this into an actual
dispatcher if we plan to do so in the future. Otherwise, something like
`classify_event` is probably a more fit name.