Re: glymur: hard reset on the first system-domain idle entry (SS3, 0x0200c354) after a heavy load - Lenovo Yoga Slim 7x Gen 11

From: Joonhoe Kim

Date: Sun Oct 04 2026 - 05:37:24 EST


Hi,

A data point from Kaanapali (SM8850), which has the same domain states
(cluster 0x01000054, system 0x0200c354): Lenovo Legion Tab Y700 Gen 5,
v7.3-rc4, PSCI OSI mode, with the CPU PM domains split in two cluster
domains (CPU0-5, CPU6-7) under power-domain-system. With the single
cluster domain of upstream kaanapali.dtsi the firmware rejects almost
every domain state, so this does not show there. In short: we see
silent resets from plain idle, and here they follow the cluster state
rather than SS3.

Symptom: a silent reset after minutes to an hour of idle with the
display off. Nothing in the printk ring or pstore dmesg; the console
ramoops only has "watchdog: CPU7: Watchdog detected hard LOCKUP on cpu 0", and
the watchdog bites before the hardlockup panic is printed.

Keeping SS3 out of runtime idle first seemed to help, but the resets
came back with SS3 never entered at runtime.

An idle-entry recorder (per-CPU ring, records cleaned to PoC around the
PSCI call), read from a RAM dump, shows the lost CPUs entering CPU
retention (0x4) around a cluster 0 power-down (0x01000054) and never
returning from that PSCI call, with IPIs and expired hrtimers pending.
In one case the last CPU, the one requesting the cluster state, did not
return either. Most cluster cycles in the same window were fine, so it
looks like a race between the cluster state and a CPU entering
retention.

Refusing the cluster domain state at runtime: 3 h idle without a reset
(562k cluster-off requests refused). Same kernel with the state allowed:
2 resets in 78 min. Screen-off idle power did not change measurably.

Is the cluster state (and SS3) meant to be used from runtime idle on
these SoCs, or is there a firmware prerequisite we miss?

Not tested on linux-next with Ulf's CPU PM domain series yet. I can run
patches or collect dumps.

Thanks,
Joonhoe Kim