Re: [PATCH v2 1/9] gpu: nova-core: gsp: cmdq: validate checksum earlier on receive
From: Alexandre Courbot
Date: Mon Sep 28 2026 - 02:26:05 EST
On Mon Sep 28, 2026 at 1:25 PM JST, Eliot Courtney wrote:
> On Sun Sep 27, 2026 at 10:46 PM JST, Alexandre Courbot wrote:
>> The checksum validation error message of `wait_for_msg` prints the
>> message's sequence number before validating the checksum.
>>
>> But the checksum is part of the transport layer, and it not validating
>> indicates a corruption that could very well be in the RPC message
>> header, meaning the value of this field cannot be trusted.
>>
>> Furthermore, the transport layer is not supposed to know the kind of
>> message it transports, and it accessing the RPC header is a layering
>> violation.
>>
>> Thus, move the checksum validation before we start looking at the
>> message header, and drop that information from the error message.
>>
>> Signed-off-by: Alexandre Courbot <acourbot@xxxxxxxxxx>
>> ---
>
> N.B. this means that checking the length of the message (which uses the
> payload length) almost always gets replaced with a bad checksum error.
> But since we can't distinguish between corrupted data and something
> wrong with the lengths, we can't do anything about it anyway so lgtm.
> It also means we don't know which message had the bad checksum, but
> again I think this fine too.
Thankfully r000 gets rid of the checksum (which indeed is not very
critical when messages are sent through shared memory...), so that part
will eventually disappear. I just wanted to get the layering right to
make the split in the next patches easier.