Re: [PATCH 0/2] Batch register access for live migration optimization
From: Yize Wang
Date: Thu Oct 08 2026 - 04:41:44 EST
在 2026/10/3 22:53, Marc Zyngier 写道:
On Sun, 20 Sep 2026 10:15:25 +0100,
Yize Wang <wangyize7@xxxxxxxxxx> wrote:
在 2026/9/18 20:08, Marc Zyngier 写道:All these locks should be non-blocking by the time you save anything
On Fri, 18 Sep 2026 09:18:13 +0100,
Yize Wang<wangyize7@xxxxxxxxxx> wrote:
This series adds batch register access support to KVM/arm64 to reduceQuestions:
syscall overhead during VM live migration.
Currently, QEMU issues one ioctl per register when saving/restoring VGIC
state. On large VM configurations this means tens of thousands of syscalls,
where lock acquisition and context switch overhead dominates migration
downtime. Thus, we provide a batch register method to allow userspace
read/write multiple distributor and redistributor registers in a single call.
In this way, we can significantly reduce syscalls and migration downtime.
Test the VM migration time under pressure conditions.
The VM specifications for migration are as follows:
- VM use 4-K page;
- the number of VCPU is 160;
- the total memory is 320Gigabit;
- use 'Redis SET-benchmark' to pressurize VM;
Performance results (3-run average, ms):
| Metric | Without patch | With patch | Improvement |
|---------------------|---------------|------------|-------------|
| Migration downtime | 536 | 321 | 40% |
| Source (total) | 344 | 230 | 33% |
| - VGIC put | 158 | 40 | 75% |
| - VGIC get | 120 | 19 | 84% |
| Destination (total) | 192 | 91 | 53% |
| - VGIC put | 132 | 27 | 80% |
Yize Wang (2):
KVM: arm64: Add batch group constant and data structure to UAPI header
KVM: arm64: Add VGIC v3 batch register access implementation
- Why only the MMIO registers?
- Why not the sysregs?
- Why only the GIC?
- Why not all of the state?
- Where is the corresponding userspace code?
More importantly, since this is about batching system calls:
- Why can't this be done with io_uring instead?
M.
Hi, Marc! Thank you for the review.
These patches focus on optimizing GICv3 register access during live
migration. We found that there are a large number of locks (kvm->lock,
vcpus, config_lock) in the GIC, these lock operations wil cost large
time waste. The batches of sysreg for vcpu optimization will come in
follow as a separate series. And let me address these questions one by
one.
related to the GIC, because:
- none of the vcpu can be running
- this must be a single threaded operation
So if you are seeing anything contended, this is either the sign of a
bad KVM bug, or an indication that you are violating the above
requirements.
M.
Hi, Marc! Thanks for your review.
You are right that when saving the GIC state during live migration process, all vCPUs are not running and it is a single threaded operation. Our patches process the packaged register states from the userspace under the same lock. These operations only aim to reduce the lock opreation downtime and do not involve lock contention.
In addition, these patches focus on reducing the syscalls of ioctl. Now, every register requires an ioctl, resulting in (num_cpu * registers_per_cpu) system calls. This will cost lots of time during the large-scale VM migration. Our patches batch registers by category in qemu and then send them to kernel in batches. This operations significantly reduce the number of ioctl syscalls and save a great amount of downtime.
The comparison is as follow:
BEFORE:
QEMU (Userspace)Kernel (Kernelspace)
+--------------------------++------------------------------+
||||
|kvm_arm_gicv3_get()||vgic_v3_attr_regs_access()|
||||
|for ncpu in num_cpu:|||
|+--------------+|||
|| GICR_CTLR|----|--- ioctl ------> |mutex_lock(&kvm->lock)|
|+--------------+ <--|--- val --------- |kvm_trylock_all_vcpus()|
|+--------------+||mutex_lock(&config_lock)|
|| GICR_STATUSR |----|--- ioctl ------> ||
|+--------------+ <--|--- val --------- |read/write single register|
|+--------------+|||
|| GICR_WAKER|----|--- ioctl ------> |mutex_unlock(&config_lock)|
|+--------------+ <--|--- val --------- |kvm_unlock_all_vcpus()|
|...||mutex_unlock(&kvm->lock)|
|+--------------+|||
|| ICC_SRE_EL1|----|--- ioctl ------> |(repeat lock -> read -> unlock
|+--------------+ <--|--- val --------- |for each register)|
|...|||
||||
+--------------------------++------------------------------+
AFTER:
QEMU (Userspace)Kernel (Kernelspace)
+----------------------------++--------------------------------+
||||
|kvm_arm_gicv3_get()||vgic_v3_batch_access()|
||| |
|+- Collect Phase ----+|||
||| |||
|| batch_add(GICR_CTLR)| |||
|| batch_add(GICR_STATUSR) |||
|| batch_add(GICR_WAKER)||||
|| batch_add(GICR_IGROUPR0)| ||
|| ...||||
|| batch_add(ICC_SRE_EL1) |||
|| ...||||
|+---------------------+ |||
||1 ioctl|copy_from_user(entries)|
|+- Commit Phase -----+|||
||| --- ioctl(BATCH)---> |mutex_lock(&kvm->lock)|
||batch_commit()| ||kvm_trylock_all_vcpus()|
|||||mutex_lock(&config_lock)|
||||||
|+---------------------+ ||+-- Process Each Entry --+|
||||||
|+- Unpack Phase -----+|||entry[0]: read GICR_CTLR| |
||| <---batch result ---||entry[1]: read GICR_STATUSR|
||Write back to QEMU | | ||entry[2]: read GICR_WAKER| |
||state||||...||
||(post-process:| | ||entry[N]: read ICC_SRE | |
||unshuffle, split| | ||| |
||bytes, etc.)| |||* FLAG_LOCKED: skip||
||| | ||locking (already held)| |
|+---------------------+ ||+-------------------------+ |
||| |
+----------------------------++--------------------------------+