Re: [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+

From: David Woodhouse

Date: Fri Oct 09 2026 - 06:04:11 EST


On Mon, 2026-09-07 at 17:40 +0100, Will Deacon wrote:
> Hi folks,
>
> This series implements CONFIG_HOTPLUG_PARALLEL for arm64 by utilising
> the argument to CPU_ON introduced in PSCI v0.2. This supersedes the
> previous series from Jinjie [1], and I'm grateful to him for his support
> getting this alternative version in shape for posting.
>
> I previously spoke at KVM forum about this work last year:
>
>   https://www.youtube.com/watch?v=Q6kOshnnQuE
>
> There are three major benefits realised by these changes:
>
> 1. We ditch our home-brew secondary CPU synchronisation in favour of
>    constructs in the core code.
>
> 2. They offer a performance advantage on some platforms, where getting a
>    CPU into the kernel can take time. I'm hoping Jinjie can provide
>    numbers in this case.
>
> 3. When used inside a confidential guest, they provide protection
>    against a malicious VMM that puts secondary CPUs into the guest
>    kernel after the onlining operation has timed out.
>
> The first four patches are small changes to the generic code to make it
> a little more "Arm-shaped" but the rest of the series is largely
> confined to the arm64 architecture and PSCI driver code.
>
> The series is based on -rc2, as it otherwise conflicts with Fuad's GMID
> fix that was recently merged upstream.
>
> All feedback welcome,

Thanks both of you for working on this. It was always on my list to
come back and do it for arm64, but that list is long. Also on my list
FWIW is parallelising the *later* stages of CPU hotplug¹, not just the
early path into start_secondary()/secondary_start_kernel(). IIRC there
was more win to be had, but we had to prove that a lot of once-
serialized code could safely run concurrently. Ideally without just
naïvely adding locking and serializing it all again.

I gave your series a quick spin across a range of EC2 systems, and it
gives a ~40% win on systems with 192 cores — at least, as a
microbenchmark of the CPU onlining. For these systems, PSCI is fairly
fast to bring the CPUs online and it isn't a huge proportion of the
*overall* kexec time.

However, your code in serial mode (cpuhp.parallel=0) is up to 40%
*slower* than before. A/B testing results, all values in milliseconds:

┌───────────┬──────┬───────────┬────────────────┬─────────────────┐
│ platform │ ncpu │ A 7.3-rc2 │ A' Will-serial │ B Will-parallel │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ a1-metal │ 16 │ 5.7 │ 6.7 (+16.8%) │ 5.2 (−8.6%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ a1-virt │ 16 │ 13.6 │ 15.7 (+15.4%) │ 10.5 (−22.5%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ m6g-metal │ 64 │ 21.7 │ 29.8 (+37.3%) │ 15.5 (−28.8%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ m6g-virt │ 64 │ 27.0 │ 30.9 (+14.5%) │ 23.0 (−14.6%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ c7g-metal │ 64 │ 20.6 │ 26.0 (+26.3%) │ 11.9 (−41.9%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ c7g-virt │ 64 │ 26.8 │ 31.0 (+15.4%) │ 21.7 (−19.1%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ c8g-metal │ 192 │ 81.7 │ 111.2 (+36.0%) │ 53.0 (−35.2%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ c8g-virt │ 192 │ 110.7 │ 120.4 (+8.8%) │ 96.4 (−12.9%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ m9g-metal │ 192 │ 69.0 │ 97.2 (+40.9%) │ 43.9 (−36.3%) │
├───────────┼──────┼───────────┼────────────────┼─────────────────┤
│ m9g-virt │ 192 │ 128.6 │ 138.9 (+8.0%) │ 112.2 (−12.7%) │
└───────────┴──────┴───────────┴────────────────┴─────────────────┘

We also tried it on another system where the firmware takes about 9ms
to bring each CPU up (after CPU_ON returns fairly quickly). Fanning out
the CPU_ON calls didn't make any difference either; the CPUs came
online, one at a time, about 9ms apart. The parallel onlining saved
only about 20ms out of 820ms (→800ms) here.

The real answer for such platforms (at least for kexec) is *not* to do
the CPU_OFF/CPU_ON thing at all. Pasha's Caretaker work² still does so
for the reclaim at hotplug time in the next kernel; I'm experimenting
with eliding that, which should give the biggest improvement on such
platforms. So during kexec the APs just spin and wait to be asked to
come back, instead of going completely offline.

Once I have that working in its upstreamable Caretaker-compatible form,
I'll post the full 3x2 matrix of results for each platform.

¹ https://lore.kernel.org/all/ff1f4bcd728c3500f80259d1bd9319f6d4cabef2.camel@xxxxxxxxxxxxx/t/#u
² https://lore.kernel.org/all/20260920193650.3373435-1-pasha.tatashin@xxxxxxxxxx/

Attachment: smime.p7s
Description: S/MIME cryptographic signature