Re: [PATCH 1/2] crypto: pcrypt - Remove pcrypt

From: Hendrik Donner

Date: Fri Jul 24 2026 - 14:22:53 EST



Hallo,

On 7/22/26 21:40, Eric Biggers wrote:
On Wed, Jul 22, 2026 at 06:12:16PM +0200, Hendrik Donner wrote:
Hello,

On 7/21/26 21:50, Eric Biggers wrote:
On Tue, Jul 21, 2026 at 08:59:21PM +0200, Hendrik Donner wrote:
Hello,

On 7/14/26 00:32, Eric Biggers wrote:
pcrypt was originally intended to improve IPsec performance. However,
it's no longer useful for that. Reports from the rare cases that anyone
has actually tried to use it over the years indicate that it actually
reduces IPsec performance, e.g.:

* https://github.com/libreswan/libreswan/wiki/Internals:-Cryptographic-Acceleration#obsoleted-ipsec-accelerations
* https://users.strongswan.narkive.com/liqTaTq8/strongswan-problem-with-pcrypt
* https://unix.stackexchange.com/questions/594336/ipsec-multithreading-via-pcrypt-worse-than-single-thread

It's also undocumented and quite difficult to actually use. Its design
is also broken, in that any unprivileged program can enable pcrypt
systemwide at any time (by instantiating it using AF_ALG).

Meanwhile, pcrypt has been a regular source of bugs, including at least
four that have received CVEs.

Let's just remove it. No one seems to care about it anymore other than
people looking for vulnerabilities.


my company is a user. We have a hardware platform based on an IMX6 SoC
using IPSec and configure pcrypt using crconf. Current performance
difference:

iperf3 -c <IP> --time 60 -R

pcrypt:
Download: 107 Mbits/sec

No pcrypt:
Download: 59.3 Mbits/sec

iperf3 -c <IP> --time 60

pcrypt:
Upload: 65.9 Mbits/sec

No pcrypt:
Upload: 52.0 Mbits/sec

The relevant crypto templates are configured in early userspace and
since i got curious, that has been the case since 2017.

Mostly using

pcrypt(gcm_base(ctr-aes-neonbs,ghash-generic))

nowadays, AES-CBC in the past/as a fallback option.

So at least on some platforms there is still a significant performance
boots, at least for downloads in this case.

Thanks for bringing up your use case.

Have you looked into alternative solutions such as Receive Side Scaling
(https://docs.kernel.org/networking/scaling.html#rss-receive-side-scaling)?
AFAIK it's not just the crypto performance that makes pcrypt unnecessary
these days, but also the design of the networking layer.

I'm looking into this more, the IMX.6 is a bit limited with IRQ handling and
queue distribution.

Thanks! Maybe Steffen and the other IPsec folks would have some advice
too.


so i'm now on 7.1.4 with

PCI: imx6: Keep i.MX6 Root Port MSI/MSI-X Capabilities with iMSI-RX to work around hardware bug

on top to be able to tune queue settings. And to have a working ethernet
in the first place, without the patch the NETDEV WATCHDOG resets the
card all the time due to queues stalling. But now more than 1 CPU are
serving IRQs.

With pcrypt
(seqiv(rfc4106(pcrypt(gcm_base(ctr-aes-neonbs,ghash-lib))))):

Upload:
[ 4] 0.00-60.00 sec 901 MBytes 126 Mbits/sec

Download:
[ 4] 0.00-60.00 sec 1.23 GBytes 177 Mbits/sec

Without pcrypt
(seqiv(rfc4106(gcm_base(ctr-aes-neonbs,ghash-lib)))):

Upload:
[ 4] 0.00-60.00 sec 679 MBytes 94.9 Mbits/sec

Download:
[ 4] 0.00-60.00 sec 674 MBytes 94.3 Mbits/sec

So counterintuitively pcrypt matters more again. I repeated the tests a
few times, those numbers are fairly representative. Every run is over a
60 sec window.

Regards,
Hendrik

I understand that i.MX6 doesn't have the ARMv8 crypto extensions.
However, surely you could at least use the NEON-optimized GHASH code?
Is there a reason you're not using it?

I think it was historically not working well for us, retested:

NEON GHASH with pcrypt:

Download: 116 Mbits/sec
Upload: 74.3 Mbits/sec

NEON GHASH baseline:

Download: 90.7 Mbits/sec
Upload: 60.8 Mbits/sec

Looks better, still a ~15 Mbit improvement with pcrypt.

I want to point out that without IPSec our baseline is:

Download: 942 Mbits/sec
Upload: 942 Mbits/sec

Intel IGB ethernet.

That already shows that about two-thirds of the improvement you were
getting from pcrypt can be gotten just by using the correct GHASH
implementation for the platform (59.3 => 90.7 download, vs 59.3 => 107;
and 52.0 => 60.8 upload, vs 52.0 => 65.9).

And with that new baseline, for downloads, pcrypt adds just 28% more
throughput (90.7 => 116) rather than 80% as it did before (59.3 => 107).

It seems clear that the usefulness of pcrypt rapidly decreases as the
actual crypto gets faster.

Note: we've been enabling crypto optimizations by default in recent
kernels, so that people can no longer use the generic code by accident.
For example in v7.1 and later, GHASH optimizations are always enabled.

I'm working on further optimizations to the AES-GCM code as well,
specifically implementing AES-GCM directly on all platforms without the
inefficient gcm_base template that is being used here.

So while maybe pcrypt does still help a bit on this platform for now,
the approach does seem quite dated and largely a workaround for
inefficiencies elsewhere in the stack (including systems where the
optimized crypto code is accidentally not enabled).

- Eric