Re: [Question] Userspace throttling + "sched/fair: Combine detach into dequeue when migrating task" causes guest boot hang
From: Aaron Lu
Date: Wed Aug 26 2026 - 06:20:19 EST
On Tue, Aug 25, 2026 at 12:06:29PM +0000, Chen Jinghuang wrote:
> Hi, I'm seeing a VM boot hang on mainline, and I'd like to understand the
> interaction between userspace throttling and the scheduler patch
> "e1f078f50478 sched/fair: Combine detach into dequeue when migrating
> task".
>
> Host:
> aarch64, 96 CPUs (0-95), 4 NUMA nodes:
> node0: 0-23, node1: 24-47, node2: 48-71, node3: 72-95
> Mainline kernel tag: 7.2-rc1.
>
> Guest(libvirt/KVM) - described in words:
> An aarch64 (virt-6.2) UEFI VM launched with `virsh create`; key config:
>
> - 128 vCPUs (statically placed, oversubscribed — the host has only 96
> physical CPUs).
> - host-passthrough CPU model; GICv3; 64 GiB RAM.
> - <cputune> has <global_quota> set to 400000; all <vcpupin> and
> <emulatorpin> entries are commented out, so there is no vCPU pinning.
> - Storage: qcow2 on virtio-scsi (cache=none, io=native). HPET disabled.
>
> Userspace throttling:
> The VM runs under a CPU-quota cap applied on the host. The actual values
> from the cgroup controller are:
>
> cpu.cfs_period_us = 100000
> cpu.cfs_quota_us = 400000
>
> I also found that if I set cpu.cfs_quota_us to -1, or enlarge it beyond a
> certain point, the guest boots fine.
>
> Symptom:
> The guest hangs at some command early in boot and never reaches the login
> prompt.
>
I tried this on an x86 machine with v7.2-rc1 kernel and with quota set
to 4 cpus, the VM booted fine; when I further reduced quota to 1 cpu, the
guest kernel would dump a ton of soft lockups during boot. I also tried
running an old 5.10 kernel(which doesn't have per-task throttle) and it
behaved the same as v7.2-rc1.
The x86 machine has 64cores/128cpus and the VM I created has 128cpus and
128G memory.
> Observations:
> Only reverting both of the following together makes it boot (neither one
> alone suffices):
>
> 1. The kernel patch for userspace throttling.
> 2. The scheduler patch:
> e1f078f50478 ("sched/fair: Combine detach into dequeue when migrating
> task")
I'm curious how you found e1f078f50478, just because it touched pelt?
>
> Reverting only one of them still hangs; reverting both together boots fine.
>
On top of v7.2-rc1, right?
> Question:
> I don't fully understand how these two interact. My rough guess: e1f078f50478
> ("sched/fair: Combine detach into dequeue when migrating task") affects the
> PELT accounting, and the userspace throttling also has logic that affects PELT
> accounting. When both are combined, load balancing and subsequent scheduling
> behavior may end up misbehaving, stalling the guest.
Is the host busy? If the host has many idle cpus, even the pelt is
wrecked(which I doubt), it should not cause the qemu task being starved.
The PELT accounting matters when tasks have to compet the same CPU, but
if your host system has many idle cpus, that should not happen.
And from the log you posted for the cpu usage, it appears that task
group is getting cpu time.