Re: [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+
From: David Woodhouse
Date: Fri Oct 09 2026 - 11:07:00 EST
On Fri, 2026-10-09 at 12:08 +0100, Will Deacon wrote:
> Hi David,
>
> Thanks for replying with so much data! It looks like I sent my v2 out
> at the same time.
Yeah. I'm redoing the full set of tests on that, along with the
spinwait variant.
> On Fri, Oct 09, 2026 at 11:00:08AM +0100, David Woodhouse wrote:
> > Thanks both of you for working on this. It was always on my list to
> > come back and do it for arm64, but that list is long. Also on my list
> > FWIW is parallelising the *later* stages of CPU hotplug¹, not just the
> > early path into start_secondary()/secondary_start_kernel(). IIRC there
> > was more win to be had, but we had to prove that a lot of once-
> > serialized code could safely run concurrently. Ideally without just
> > naïvely adding locking and serializing it all again.
>
> That gets _really_ hard because you need RCU up and running pretty early.
RCU should be OK, as long as a CPU which is registered doesn't *then*
block in a spin-wait, like the traditional CPU hotplug machinery does.
> > I gave your series a quick spin across a range of EC2 systems, and it
> > gives a ~40% win on systems with 192 cores — at least, as a
> > microbenchmark of the CPU onlining. For these systems, PSCI is fairly
> > fast to bring the CPUs online and it isn't a huge proportion of the
> > *overall* kexec time.
> >
> > However, your code in serial mode (cpuhp.parallel=0) is up to 40%
> > *slower* than before. A/B testing results, all values in milliseconds:
>
> Oh, that's unexpected.
> ...
> Would it be possible for you to pick one of these platforms and bisect
> the serial regression, please? I'm not really sure where to look, but
> serial boot should work all the way through the series so if you can
> identify the point at which it regresses (and presumably doesn't recover)
> then that would hopefully point me in the right direction. One possibility
> is that the extra atomics in the new state machine logic are slowing things
> down. Another possibility is the timeout logic in
> cpuhp_wait_for_sync_state() ends up sleeping in your tests.
Looks like the latter, falling back to usleep_range(1000, 2000). Goes
away if you make it atomic_cond_read_relaxed():
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=cba8bdecff74
Of course, if we weren't synchronously *waiting* for the APs and they
just turned up later, the world would be a better place. And the ones
with slow firmware wouldn't matter either.
Unless we're actually doing parallel initcalls, we don't even *need*
the secondary CPUs until much later, do we?
> Probably worth using my v2 just in case one of the fixes there helps,
> but I'm not hopeful.
>
> > We also tried it on another system where the firmware takes about 9ms
> > to bring each CPU up (after CPU_ON returns fairly quickly). Fanning out
> > the CPU_ON calls didn't make any difference either; the CPUs came
> > online, one at a time, about 9ms apart. The parallel onlining saved
> > only about 20ms out of 820ms (→800ms) here.
>
> Damn, I guess the firmware has some serialisation in that case?
$DEITY knows, but I guess so. We don't get to write the firmware for
those ones, AIUI.
> > The real answer for such platforms (at least for kexec) is *not* to do
> > the CPU_OFF/CPU_ON thing at all. Pasha's Caretaker work² still does so
> > for the reclaim at hotplug time in the next kernel; I'm experimenting
> > with eliding that, which should give the biggest improvement on such
> > platforms. So during kexec the APs just spin and wait to be asked to
> > come back, instead of going completely offline.
>
> I was talking to Tarun about that at LPC. I was envisaging a form of
> CPU_OFF that would leave the MMU enabled so we could make use of the
> existing EFI boot logic, but I hadn't thought about it beyond that.
I think it ties closely into the caretaker thing. We don't take the CPU
offline at all (no firmware, no CPU_OFF of any kind). Instead, we just
leave it parked somewhere.
The first stage is leaving it parked and just spinning waiting to come
back. Which is all we care about here.
The next stages involve making it do useful *work* instead of just
spinning. Like Pasha's KVM vCPU caretaker, and the thing I'm working on
which runs a minimal C runtime driving an e1000 through VFIO,
transitioning onto bare metal on the "offline" CPUs across kexec, then
back to running in a Linux thread after kexec. All in QEMU while being
flood-pinged from the host across the kexec.
Attachment:
smime.p7s
Description: S/MIME cryptographic signature