Re: [Question] Userspace throttling + "sched/fair: Combine detach into dequeue when migrating task" causes guest boot hang

From: chenjinghuang

Date: Wed Aug 26 2026 - 04:28:52 EST


On 8/25/2026 9:15 PM, K Prateek Nayak wrote:
> Hello Chen,
>
> Thank you for your report.
>
Thanks for your response, Here are the updated test results for your questions:
> On 8/25/2026 5:36 PM, Chen Jinghuang wrote:
>> Hi, I'm seeing a VM boot hang on mainline, and I'd like to understand the
>> interaction between userspace throttling and the scheduler patch
>> "e1f078f50478 sched/fair: Combine detach into dequeue when migrating
>> task".
>>
>> Host:
>> aarch64, 96 CPUs (0-95), 4 NUMA nodes:
>> node0: 0-23, node1: 24-47, node2: 48-71, node3: 72-95
>> Mainline kernel tag: 7.2-rc1.
>>
>> Guest(libvirt/KVM) - described in words:
>> An aarch64 (virt-6.2) UEFI VM launched with `virsh create`; key config:
>>
>> - 128 vCPUs (statically placed, oversubscribed — the host has only 96
>> physical CPUs).
>> - host-passthrough CPU model; GICv3; 64 GiB RAM.
>> - <cputune> has <global_quota> set to 400000; all <vcpupin> and
>> <emulatorpin> entries are commented out, so there is no vCPU pinning.
>> - Storage: qcow2 on virtio-scsi (cache=none, io=native). HPET disabled.
>>
>> Userspace throttling:
>> The VM runs under a CPU-quota cap applied on the host. The actual values
>> from the cgroup controller are:
>>
>> cpu.cfs_period_us = 100000
>> cpu.cfs_quota_us = 400000
>
> 400ms across 96 CPUs per 100ms seems awfully low. Let me go see if I can
> reproduce this.
>
>>
>> I also found that if I set cpu.cfs_quota_us to -1, or enlarge it beyond a
>> certain point, the guest boots fine.
>
> Sounds a lot like guest side lock-holder preemption stalling the guest.
> If you give it enough time, does the guest progress?
>
Even if given plenty of time, the guest fails to complete booting. During
boot, it first pauses for a while at the early console log:
[ 0.003862][ T0] printk: console [tty0] enabled
[ 0.004709][ T0] printk: bootconsole [pl11] disabled

Then it progresses a bit and hangs again here:
[ 129.095680][ T1] systemd[1]: Finished Create List of Static Device Nodes.
[ 129.097121][ T1] systemd[1]: sysinit.target: starting held back, waiting for: systemd-sysctl.service

Softlockups are occasionally triggered inside the guest, but it never reaches
the login prompt.
>>
>> Symptom:
>> The guest hangs at some command early in boot and never reaches the login
>> prompt.
>
> What happens if you allow it to boot and then enforce the more
> the aggressive limits later? Do you see RCU stalls / lockups?
>
If I allow the guest to boot normally first by setting echo -1 > cpu.cfs_quota_us,
and then echo 400000 > cpu.cfs_quota_us after boot completes, the guest can be
operated normally without any RCU stalls.
>>
>> Observations:
>> Only reverting both of the following together makes it boot (neither one
>> alone suffices):
>>
>> 1. The kernel patch for userspace throttling.
>
> Are these Aaron's patches too or just the recent rework that I did?
> Could you please paste a log of all the reverts.
>
The series of userspace throttling patches includes:
- sched/fair: Add related data structure for task based throttle
- sched/fair: Implement throttle task work and related helpers
- sched/fair: Switch to task based throttle model
- sched/fair: Task based throttle time accounting
- sched/fair: Get rid of throttled_lb_pair()
- sched/fair: Propagate load for throttled cfs_rq
- sched/fair: update_cfs_group() for throttled cfs_rqs
- sched/fair: Do not balance task to a throttled cfs_rq
- sched/fair: Start a cfs_rq on throttled hierarchy with PELT clock throttled
- sched/fair: Prevent cfs_rq from being unthrottled with zero runtime_remaining

Even with e1f078f50478 reverted, switching to commit sched/fair: Switch to task based throttle model
reproduces the hang, whereas switching to the commit sched/fair: Implement throttle task work and related helpers
works fine.

>> 2. The scheduler patch:
>> e1f078f50478 ("sched/fair: Combine detach into dequeue when migrating
>> task").
>>
>> Reverting only one of them still hangs; reverting both together boots fine.
>
> Can you check your cgroup stats to see how much time the vCPUs are getting
> before and after the revert? Very surprising that e1f078f50478 has some
> effect here.
>
Here are the cgroup statistics sampled 10 seconds apart:

before revert:
cpu.stat:
=== T0 ===
nr_periods 1150
nr_throttled 701
throttled_time 5843715989650
nr_bursts 0
burst_time 0
=== T1 (10s later) ===
nr_periods 1250
nr_throttled 801
throttled_time 6732356889110
nr_bursts 0
burst_time 0

cpuacct.usage:
=== T0 ===
541942826770
=== T1 (10s later) ===
581940183840

after revert:
cpu.stat:
=== T0 ===
nr_periods 2899
nr_throttled 719
throttled_time 3892756326930
nr_bursts 0
burst_time 0
adjust_runtime 0
=== T1 (10s later) ===
nr_periods 2999
nr_throttled 808
throttled_time 4390035054220
nr_bursts 0
burst_time 0
adjust_runtime 0

cpuacct.usage:
=== T0 ===
1016649387110
=== T1 (10s later) ===
1051395917140
>>
>> Question:
>> I don't fully understand how these two interact. My rough guess: e1f078f50478
>> ("sched/fair: Combine detach into dequeue when migrating task") affects the
>> PELT accounting, and the userspace throttling also has logic that affects PELT
>> accounting. When both are combined, load balancing and subsequent scheduling
>> behavior may end up misbehaving, stalling the guest.
>>
>> This looks like a real regression on mainline in the 128-vCPU oversubscribed
>> VM on a 96-core/4-NUMA host scenario. Any pointer to the correct mechanism or
>> a fix direction would be very much appreciated.
>
> Both, with exit-to-user throttling, and the legacy method, we would have
> preempted the vCPU in xfer_to_guest_mode_work():
>
> if (ti_work & (_TIF_NEED_RESCHED | _TIF_NEED_RESCHED_LAZY))
> schedule();
>
> if (ti_work & _TIF_NOTIFY_RESUME)
> resume_user_mode_work(NULL);
>
> Previously, task would have taken the schedule() route out, and now it
> is done via resume_user_mode_work() -> schedule() / preempt_schedule()
>
> Since e1f078f50478 only takes effect at migration, does 1:1 pinning
> help progress the boot?
>
With 1:1 static vCPU pinning, the issue no longer occurs.
Additionally, if I restrict all vCPU tasks to CPU range 0-3 (e.g. by setting cpuset="0-3"),
the issue also disappears.

> Are there any splats in your dmesg?
>
>
There are no errors, warnings, or splats in the host dmesg.