RE: [EXT] [PATCH v2] crypto: caam/qi2 - lower the algorithm priority
From: Sahil Malhotra (OSS)
Date: Wed Sep 30 2026 - 06:22:57 EST
Hi Vincent,
Got your point, Regarding the numbers on LS1088 and LS2088 I need to fetch them.
For this patch, I am ok.
Reviewed-by: Sahil Malhotra <sahil.malhotra@xxxxxxx>
Regards,
Sahil Malhotra
NXP Confidential
> -----Original Message-----
> From: Vincent Jardin <vjardin@xxxxxxx>
> Sent: 30 September 2026 14:39
> To: Sahil Malhotra (OSS) <sahil.malhotra@xxxxxxxxxxx>
> Cc: Horia Geanta <horia.geanta@xxxxxxx>; Pankaj Gupta
> <pankaj.gupta@xxxxxxx>; Herbert Xu <herbert@xxxxxxxxxxxxxxxxxxx>; David S.
> Miller <davem@xxxxxxxxxxxxx>; Gaurav Jain <gaurav.jain@xxxxxxx>; Eric
> Biggers <ebiggers@xxxxxxxxxx>; linux-crypto@xxxxxxxxxxxxxxx; linux-
> kernel@xxxxxxxxxxxxxxx
> Subject: Re: [EXT] [PATCH v2] crypto: caam/qi2 - lower the algorithm priority
>
> Caution: This is an external email. Please take care when clicking links or opening
> attachments. When in doubt, report the message using the 'Report this email'
> button
>
>
> Hi Sahil,
>
> Le 30/09/26 07:21, Sahil Malhotra (OSS) a écrit :
> > Hi Vincent,
> >
> > I'm not sure I understand the need for this change.
> > Is it now expected that hardware accelerators should have a lower default
> priority than ARM CE?
> > Users who want to use ARM CE can already choose it at runtime, so changing
> the default priority does not seem necessary from my perspective.
>
> Not as a general rule. The default priority should select what is best for most
> users with the platform as it ships, and on the LX2160A this is the CE.
>
> Measured on an LX2160A (16x Cortex-A72 at 2.2 GHz) with tcrypt, AES-128-GCM
> encryption, 4 KiB requests:
>
> gcm-aes-ce 1 core 1107 MB/s
> gcm-aes-caam-qi2, 1 in flight 128 MB/s
> gcm-aes-caam-qi2, 32 in flight 1.4 cores 418 MB/s
>
> It is the same for cbc, ctr, xts, rfc4106 and sha256/512, at every size I measured:
> the SEC is slower, and its driver path costs more CPU per byte than doing the
> crypto on the core.
>
> The SEC can win, but only in a narrow case. The MC DPC and DPL have to be
> tuned for it (one DPIO and one DPSECI queue pair per CPU, and the SEC
> coherency setting in the DPC), and the load has to be large buffers over many
> keys. I did set such scenario for my internal working cases, but it they are not the
> default cases.
>
> This is the case of QAT in commit
> 8024774190a5 ("crypto: qat - lower priority for skcipher and aead algorithms")
> most users call the crypto API synchronously on small buffers and do not benefit
> from the accelerator.
>
> > Users who want to use ARM CE can already choose it at runtime, so
> > changing the default priority does not seem necessary from my
> > perspective.
>
> In practice it works the other way round. IPsec, dm-crypt or kTLS ask for an
> algorithm by name and get the highest priority, they do not choose a driver. The
> users who gain from the SEC are the ones who also tune the DPC and DPL for it,
> and they can raise its priority with crypto_user, or ask for the driver by name.
>
> So the CPU should be the 1st gear, and the SEC should be the next one, for those
> who set the platform up for it. It is not a statement that hardware crypto is slower
> in general.
>
> I only measured the LX2160A, not the LS1088A or LS2088A. If you have numbers
> there with the SEC ahead, I'm happy to look at them.
>
> Best regards,
> Vincent