[Question] Userspace throttling + "sched/fair: Combine detach into dequeue when migrating task" causes guest boot hang

From: Chen Jinghuang

Date: Tue Aug 25 2026 - 08:21:53 EST


Hi, I'm seeing a VM boot hang on mainline, and I'd like to understand the
interaction between userspace throttling and the scheduler patch
"e1f078f50478 sched/fair: Combine detach into dequeue when migrating
task".

Host:
aarch64, 96 CPUs (0-95), 4 NUMA nodes:
node0: 0-23, node1: 24-47, node2: 48-71, node3: 72-95
Mainline kernel tag: 7.2-rc1.

Guest(libvirt/KVM) - described in words:
An aarch64 (virt-6.2) UEFI VM launched with `virsh create`; key config:

- 128 vCPUs (statically placed, oversubscribed — the host has only 96
physical CPUs).
- host-passthrough CPU model; GICv3; 64 GiB RAM.
- <cputune> has <global_quota> set to 400000; all <vcpupin> and
<emulatorpin> entries are commented out, so there is no vCPU pinning.
- Storage: qcow2 on virtio-scsi (cache=none, io=native). HPET disabled.

Userspace throttling:
The VM runs under a CPU-quota cap applied on the host. The actual values
from the cgroup controller are:

cpu.cfs_period_us = 100000
cpu.cfs_quota_us = 400000

I also found that if I set cpu.cfs_quota_us to -1, or enlarge it beyond a
certain point, the guest boots fine.

Symptom:
The guest hangs at some command early in boot and never reaches the login
prompt.

Observations:
Only reverting both of the following together makes it boot (neither one
alone suffices):

1. The kernel patch for userspace throttling.
2. The scheduler patch:
e1f078f50478 ("sched/fair: Combine detach into dequeue when migrating
task").

Reverting only one of them still hangs; reverting both together boots fine.

Question:
I don't fully understand how these two interact. My rough guess: e1f078f50478
("sched/fair: Combine detach into dequeue when migrating task") affects the
PELT accounting, and the userspace throttling also has logic that affects PELT
accounting. When both are combined, load balancing and subsequent scheduling
behavior may end up misbehaving, stalling the guest.

This looks like a real regression on mainline in the 128-vCPU oversubscribed
VM on a 96-core/4-NUMA host scenario. Any pointer to the correct mechanism or
a fix direction would be very much appreciated.

Thanks,
Chen Jinghuang