[PATCH v5 00/15] KVM: arm64: FEAT_HDBSS support for stage-2 dirty tracking
From: Tian Zheng
Date: Tue Sep 29 2026 - 06:56:30 EST
This series implements hardware-assisted stage-2 dirty tracking on
arm64 using FEAT_HDBSS. It combines Leonardo Bras' HAFDBS descriptor
rework [1] with the HDBSS buffer support built on top of it, as
requested during the review of [1].
Patches 1-4 are from Leonardo's RFC, reworked per the review; patch 5
adds the folio-account harvest from the HDBSS validation work. The
descriptor encoding moves the stage-2 write permission from S2AP[1]
to DBM and reuses S2AP[1] as the dirty state:
RO (DBM=0, S2AP[1]=0) read-only, write -> permission fault
WC (DBM=1, S2AP[1]=0) writable-clean, hw promotes on write
WD (DBM=1, S2AP[1]=1) writable-dirty
The remaining patches implement the HDBSS buffer machinery: the
buffer size is configured before the first vCPU is created, the
buffers are allocated at vCPU creation, a full buffer raises a fault
that forces an exit, and every VM exit flushes the entries into the
dirty bitmap or the dirty ring. The size ioctl only applies to
dirty-bitmap mode, as dirty-ring mode pins the buffer size to
PAGE_SIZE. A single derived hardware-dirty mode selects the
configuration:
migration with HDBSS -> HD|HA|HDBSS
migration without -> HD off, write-protect faults
no migration -> HD|HA, only written pages go dirty
The HDBSS registers are programmed on vCPU load (patch 7): HDBSS can
be enabled while a vCPU is mid-KVM_RUN, and hardware appends dirty
entries without any fault, so HDBSSBR_EL2 must already point at the
running vCPU's buffer.
HAFDBS is not gated on nested virtualization: shadow stage-2 MMUs
build their own VTCR without HD via kvm_get_vtcr(). HDBSS stays
gated, as the nested exit-flush and harvest paths are unaudited.
This series was only tested on non-nested guests, so reviewer
attention on the nested paths is appreciated.
The KVM_CAP_ARM_HDBSS_BUFFER_SIZE interface is exercised by a new
selftest (patch 15), including the contract that dirty-ring mode
pins the buffer size to the default.
Comments welcome, especially on:
- allocating the HDBSS buffer at vCPU creation and programming the
HDBSSBR_EL2 and HDBSSPROD_EL2 registers on vCPU load (patch 7),
rather than at each mode switch,
- keeping HDBSS NV-gated while HAFDBS is not (patch 12),
- using the target MMU's live HD state (kvm_hw_dirty_enabled())
instead of the hardware capability (patch 12), so shadow MMUs
keep installing writable-dirty entries as before.
Relative to Leonardo's RFC [1]:
- Patch 1 keeps reading stage-2 writability from S2AP[1] in the
nested walker, as DBM is RES0 from L1's perspective.
- Patch 2 splits the two dirty ledgers on the fault paths: the
host folio account is marked speculatively on PROT_W while the
KVM dirty bitmap is only marked on PROT_DIRTY.
- Patch 5's HAFDBS toggle becomes the derived mode (patch 12),
which computes the full VTCR_EL2 dirty configuration (off,
HAFDBS or HDBSS) from the capabilities and the logging state.
- The folio-account harvest, the dirty-ring reservation and the
buffer-size UAPI are new.
Changes since v4 [2]:
- Rebased onto Leonardo's descriptor rework [1]: write permission
moves to DBM, S2AP[1] becomes the pure dirty state. Replaces v4's
auto-DBM patch and drops the eager-splitting dependency, as the
walker clears DBM on blocks so lazy splitting keeps working.
- New: harvest of the stage-2 dirty state into the host folio
account at unmap/write-protect time.
- Buffer lifetime tied to the vCPU: allocated at creation, freed at
destruction, registers programmed on ownership. Closes the
use-after-free window on a concurrent mode switch.
- Auto enable/disable replaced by a single derived mode: no illegal
intermediate VTCR_EL2 state, no locking.
- Outside migration, HAFDBS is now enabled (Leonardo's RFC [1]):
read faults install writable-clean pages, hardware promotes them
on write, and only pages actually written to become dirty.
- Flush and HDBSS fault handling split into separate patches, with
the flush unified at VM exit so that entries pushed to the dirty
ring are accounted for before the vCPU re-enters the guest.
- New: dirty-ring support (ring reserves room for a full flush,
buffer pinned to PAGE_SIZE) and the KVM_CAP_ARM_HDBSS_BUFFER_SIZE
UAPI with documentation and a selftest.
Changes since v3 [3]:
- Merge sysreg definitions into the FEAT_HDBSS detection patch (was a
separate patch in v3).
- Add auto DBM (Dirty Bit Modifier) support as a new patch, suggested
by Leonardo Bras. DBM is now controlled as a page-table level flag
(KVM_PGTABLE_S2_DBM) rather than per-PTE. Note that DBM is injected
at stage-2 MMU creation time, not lazily on first dirty access. This
means the first write to a dirty-logged page does not generate a
page fault, which is a key reason for the mandatory dependency on
Leonardo's eager hugepage splitting patch.
- Split the v3 "Enable HDBSS support and handle HDBSSF events" patch
into three patches: per-vCPU buffer management, fault handling and
buffer flush, and auto enable/disable on dirty logging change. This
implements kernel-managed automatic HDBSS enable/disable.
- Remove the KVM_CAP_ARM_HW_DIRTY_STATE_TRACK ioctl for manual HDBSS
on/off. HDBSS is now automatically enabled/disabled based on dirty
logging state via kvm_arch_commit_memory_region().
- Change HDBSS buffer flush triggers to vcpu_put, check_vcpu_requests,
and kvm_handle_guest_abort.
- Store hdbss_order at VM level (kvm->arch.hdbss_order) instead of
per-vCPU, since all vCPUs share the same order.
- Document patch is not included in this version; will be sent in a
follow-up series.
Changes since v2 [4]:
- Remove the ARM64_HDBSS configuration option and ensure this feature
is only enabled in VHE mode.
- Move HDBSS-related variables to the arch-independent portion of the
kvm structure.
- Remove error messages during HDBSS enable/disable operations.
- Change HDBSS buffer flushing from handle_exit to vcpu_put,
check_vcpu_requests, and kvm_handle_guest_abort.
- Add fault handling for HDBSS including buffer full, external abort,
and general protection fault (GPF).
- Add support for a 4KB HDBSS buffer size, mapped to the value 0b0000.
- Add a second argument to the ioctl to turn HDBSS on or off.
Changes since v1 [5]:
- Removed redundant macro definitions and switched to tool-generated.
- Split HDBSS interface and implementation into separate patches.
- Integrate system_supports_hdbss() into ARM feature initialization.
- Refactored HDBSS data structure to store meaningful values instead
of raw register contents.
- Fixed permission checks when applying DBM bits in page tables to
prevent potential memory corruption.
- Removed unnecessary dsb instructions.
- Drop the debugging printks.
- Merged the two patches "using ioctl to enable/disable the HDBSS
feature" and "support to handle the HDBSSF event" into one.
[1] https://lore.kernel.org/all/20260901171558.2674031-1-leo.bras@xxxxxxx/
[2] https://lore.kernel.org/all/20260709104026.2612599-1-zhengtian10@xxxxxxxxxx/
[3] https://lore.kernel.org/all/20260225040421.2683931-1-zhengtian10@xxxxxxxxxx/
[4] https://lore.kernel.org/all/20251121092342.3393318-1-zhengtian10@xxxxxxxxxx/
[5] https://lore.kernel.org/all/20250311040321.1460-1-yezhenyu2@xxxxxxxxxx/
Signed-off-by: Tian Zheng <zhengtian10@xxxxxxxxxx>
Eillon (3):
KVM: arm64: Add HDBSS per-vCPU buffer management
KVM: arm64: Flush the HDBSS buffer on VM exit
KVM: arm64: Handle HDBSS faults
Leonardo Bras (4):
KVM: arm64: pgtables: Change write bit from S2AP_W to DBM
KVM: arm64: Add KVM_PGTABLE_PROT_DIRTY
KVM: arm64: Introduce a dedicated walker for stage2 write-protect
KVM: arm64: Add KVM_REQ_RELOAD_STAGE2
Tian Zheng (8):
KVM: arm64: Harvest stage-2 dirty state into the host folio account
KVM: arm64: Add support for FEAT_HDBSS
KVM: Add kvm_arch_dirty_ring_size_updated() hook
KVM: arm64: Reserve dirty ring space for the HDBSS buffer
KVM: arm64: Derive the VM hardware dirty mode from dirty logging
KVM: arm64: Add HDBSS buffer size ioctl for dirty-bitmap mode
KVM: arm64: Document HDBSS buffer size ioctl
KVM: arm64: selftests: Add HDBSS buffer size ioctl interface test
Documentation/virt/kvm/api.rst | 28 +++
arch/arm64/include/asm/cpufeature.h | 5 +
arch/arm64/include/asm/esr.h | 5 +
arch/arm64/include/asm/kvm_dirty_bit.h | 44 ++++
arch/arm64/include/asm/kvm_host.h | 15 ++
arch/arm64/include/asm/kvm_mmu.h | 18 ++
arch/arm64/include/asm/kvm_nested.h | 9 +-
arch/arm64/include/asm/kvm_pgtable.h | 14 +-
arch/arm64/include/asm/sysreg.h | 9 +
arch/arm64/kernel/cpufeature.c | 12 +
arch/arm64/kvm/Makefile | 1 +
arch/arm64/kvm/arm.c | 84 ++++++-
arch/arm64/kvm/dirty_bit.c | 129 ++++++++++
arch/arm64/kvm/hyp/pgtable.c | 74 +++++-
arch/arm64/kvm/hyp/vhe/switch.c | 17 ++
arch/arm64/kvm/mmu.c | 118 ++++++++-
arch/arm64/kvm/nested.c | 5 +
arch/arm64/kvm/ptdump.c | 10 +-
arch/arm64/kvm/reset.c | 3 +
arch/arm64/tools/cpucaps | 1 +
include/linux/kvm_dirty_ring.h | 1 +
include/uapi/linux/kvm.h | 1 +
tools/testing/selftests/kvm/Makefile.kvm | 1 +
.../testing/selftests/kvm/arm64/hdbss_test.c | 224 ++++++++++++++++++
virt/kvm/dirty_ring.c | 4 +
virt/kvm/kvm_main.c | 1 +
26 files changed, 803 insertions(+), 30 deletions(-)
create mode 100644 arch/arm64/include/asm/kvm_dirty_bit.h
create mode 100644 arch/arm64/kvm/dirty_bit.c
create mode 100644 tools/testing/selftests/kvm/arm64/hdbss_test.c
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.43.0