Re: [Question] Userspace throttling + "sched/fair: Combine detach into dequeue when migrating task" causes guest boot hang

From: Aaron Lu

Date: Thu Aug 27 2026 - 07:18:06 EST


On Thu, Aug 27, 2026 at 11:36:33AM +0800, chenjinghuang wrote:
> On 8/26/2026 6:00 PM, Aaron Lu wrote:
> > On Tue, Aug 25, 2026 at 12:06:29PM +0000, Chen Jinghuang wrote:
> >> Hi, I'm seeing a VM boot hang on mainline, and I'd like to understand the
> >> interaction between userspace throttling and the scheduler patch
> >> "e1f078f50478 sched/fair: Combine detach into dequeue when migrating
> >> task".
> >>
> >> Host:
> >> aarch64, 96 CPUs (0-95), 4 NUMA nodes:
> >> node0: 0-23, node1: 24-47, node2: 48-71, node3: 72-95
> >> Mainline kernel tag: 7.2-rc1.
> >>
> >> Guest(libvirt/KVM) - described in words:
> >> An aarch64 (virt-6.2) UEFI VM launched with `virsh create`; key config:
> >>
> >> - 128 vCPUs (statically placed, oversubscribed — the host has only 96
> >> physical CPUs).
> >> - host-passthrough CPU model; GICv3; 64 GiB RAM.
> >> - <cputune> has <global_quota> set to 400000; all <vcpupin> and
> >> <emulatorpin> entries are commented out, so there is no vCPU pinning.
> >> - Storage: qcow2 on virtio-scsi (cache=none, io=native). HPET disabled.
> >>
> >> Userspace throttling:
> >> The VM runs under a CPU-quota cap applied on the host. The actual values
> >> from the cgroup controller are:
> >>
> >> cpu.cfs_period_us = 100000
> >> cpu.cfs_quota_us = 400000
> >>
> >> I also found that if I set cpu.cfs_quota_us to -1, or enlarge it beyond a
> >> certain point, the guest boots fine.
> >>
> >> Symptom:
> >> The guest hangs at some command early in boot and never reaches the login
> >> prompt.
> >>
> >
> > I tried this on an x86 machine with v7.2-rc1 kernel and with quota set
> > to 4 cpus, the VM booted fine; when I further reduced quota to 1 cpu, the
> > guest kernel would dump a ton of soft lockups during boot. I also tried
> > running an old 5.10 kernel(which doesn't have per-task throttle) and it
> > behaved the same as v7.2-rc1.
> >
> > The x86 machine has 64cores/128cpus and the VM I created has 128cpus and
> > 128G memory.
> >
> My host is an ARM64 machine without SMT, so 96 physical cores correspond to
> 96 logical CPUs. The VM is configured with 128 vCPUs, which is indeed a typical
> CPU oversubscription scenario. In my machine, if I configure the VM with 96 vCPUs,
> the issue don't occur either.
>

OK, so this looks like it has something to do with the task number. More
vCPUs translated to more qemu threads and that caused your guest boot
hang. I think there should be softlockup dumps in your case, is it that
the softlockup detector is not enabled in your guest kernel config?

With this said, I increased vCPU to 256 on this 64core/128cpus Intel host
to see if I can reproduce this. With 4 cpus quota, it managed to boot
the guest. There are several softlockups during boot though and that is
an indication the qemu task is in need of cpu time. I can imagine if I
increase the vCPU number even more or reduce the quota further, it will
eventually hang the guest with tons of softlockup msgs.