Re: [PATCH] arm64: smp: distinguish secondary CPUs that hang after reaching head.S

From: Naman Jain

Date: Mon Jul 27 2026 - 01:38:50 EST




On 7/26/2026 7:16 PM, Will Deacon wrote:
On Fri, Jul 24, 2026 at 10:16:03AM +0530, Naman Jain wrote:
On 7/23/2026 11:26 AM, Anshuman Khandual wrote:
On 23/07/26 8:19 AM, Jinjie Ruan wrote:
在 2026/7/22 19:30, Naman Jain 写道:
When a secondary CPU fails to come online, __cpu_up() falls back to
__early_cpu_boot_status, but boot status 0x0 is ambiguous: it cannot
distinguish a CPU that never executed head.S (firmware/hypervisor never
dispatched it, so it never ran a single instruction) from one that
entered head.S, started executing, and then got stuck somewhere in kernel
bring-up. Add a change to let us tell those two cases apart, which
narrows down where to look when a CPU goes missing during boot.
I previously encountered this issue when debugging the parallel startup
of ARM64 secondary cores. It is difficult for the kernel to determine
whether the secondary core is hung in the firmware or whether it has not
executed a single instruction. So I think this motive is reasonable.

Why should kernel determine the difference here ? Would not the firmware
know if it has started any secondary CPU for the kernel which must have
come inside head.S ? If the cpu gets hung inside firmware while starting
up then the debug responsibilities belong there instead.

Still wondering what's the rationale for this change.

Hello Anshuman,
This sounds fair to me. Let me elaborate the problem, beyond the scope of
this patch. In production, we occasionally see these crashes where one of
the CPU fails to bring up online, with 0x0 status code. Hypervisor may be
missing the telemetry, but the problem is that we don't know if the
secondary CPU ever started executing the instructions or is stuck somewhere
between the start of head.S and marking itself online at the end of
secondary_start_kernel().
There are couple of places, where we get those other status codes, but not
everywhere. If the issue is not easily reproducible, experiments on local
setups do not yield anything. That's where I am attempting to add some more
information in kernel to debug these issues.

I think this is a game of diminishing returns. There's a lot of stuff
that the firmware/hypervisor can get wrong here and trying to detect or
handle that in Linux is going to be a real mess. For example, if it
enters the kernel at the wrong address, or in the wrong mode, or with
the MMU enabled etc. It sounds like you don't have much idea about
what happens in the failure case, so it might not even execute the code
that you're adding correctly.

That is true. I agree. While we cannot and should not worry about adding logs for each of these firmware failure points in kernel, do you see any merit in adding any of this information to the kernel to at least narrow down the problem?

Or, can I safely consider 0x0 unknown error in CPU bring-up to be definitely a result of firmware/Hypervisor issues?

Regards,
Naman


It's also going to conflict heavily with the ongoing parallel bringup
work, which significantly reworks this code.

I think you need to add the debug to your hypervisor, rather than the
kernel.

Will