Re: [PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking

From: Tian Zheng

Date: Mon Aug 31 2026 - 08:41:08 EST




On 8/10/2026 7:01 PM, Leonardo Bras wrote:
On Wed, Aug 05, 2026 at 11:41:51AM +0800, Tian Zheng wrote:


On 8/4/2026 7:10 PM, Leonardo Bras wrote:
On Tue, Aug 04, 2026 at 12:54:16PM +0800, Tian Zheng wrote:


On 8/4/2026 12:32 AM, Leonardo Bras wrote:
On Mon, Aug 03, 2026 at 09:57:46PM +0800, Tian Zheng wrote:


On 8/3/2026 6:21 PM, Leonardo Bras wrote:
On Mon, Aug 03, 2026 at 12:04:24PM +0800, Tian Zheng wrote:


On 8/3/2026 9:33 AM, Tian Zheng wrote:
09, 2026 at 06:40:23PM +0800, Tian Zheng wrote:
-    if (prot & KVM_PGTABLE_PROT_W)
+    if (prot & KVM_PGTABLE_PROT_W) {
            set |= KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W;

+        /*
+         * No DEVICE filter needed here:
relax_perms is only called
+         * on FSC_PERM faults. Device pages
always get full RW from
+         * initial mapping and are never write-protected during
+         * migration, so they never trigger a permission fault.
+         */
+        if (pgt->flags & KVM_PGTABLE_S2_DBM)
+            set |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
+    } else {
+        /*
+         * Clear DBM on W→RO downgrade to prevent hardware from
+         * silently upgrading RO+DBM back to W+dirty, which would
+         * bypass KVM's write tracking and cause data corruption.
+         */
+        clr |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
+    }
+
This block makes it pretty evident that the DBM bit really *is* the
write permission bit. I'd much rather we
introduce the concept of dirty
state to the page table library and migrate the abstract write
permission to the DBM field, even if we don't have FEAT_HAFDBS.


Ohh, that's an amazing idea!

Thinking about that again...
If we adopt the encoding with DBM being the write-permission
bit, and all
PTEs have it since the start, how can we have lazy-splitting happening?

Only way I think of is removing both DBM and S2_S2AP_W bit
from writable
PTEs during dirty-track enable, and re-adding them during
the first write
fault. If we don't remove the DBM bit, systems with HDBSS
would just dirty
it by hardware, without causing a fault.

DBM=0 would need to happen only in the first write-protect (only on
lazy-splitting). All other write-protecting would just clean
the S2_S2AP_W
bit, as everything is already split.

Is that what was intended?

Thanks!
Leo

Hi Leo,

I think the cleanest way to handle this is to simply avoid setting DBM
on block mappings. If we only set DBM on page-level PTEs, then block
mappings will naturally stay DBM=0 and trigger a write fault on first
access — exactly what we need for lazy splitting.

When the fault occurs, the block gets split into page-level PTEs, and at
that point we can set DBM=1 on the resulting leaf entries. This way:

1. Lazy split works naturally (fault -> split -> set DBM=1)

2. No need to clear DBM globally at dirty-track enable

3. No special handling for block mappings

So I think global DBM is still viable — we just need to filter out block
mappings when setting the DBM bit. That way the lazy split path
is preserved
without extra complexity.

Hi Tian,

Humm, but would not that be contrary to what Oliver suggested:
changing the
encoding from the PTE for all entries?

(Like, if the PTE is writable, it has to have DBM set)

IIUC what you said, on first faulting of the page in the VM:
- If the entry is a page (level-3 leaf) and writable, add DBM
- If it's a block entry (leaf but not a level-3), don't add DBM

So after we enable dirty-logging:
- a level-3 entry would not fault, using HDBSS, and
- a block entry would fault, do the splitting, and add DBM to level-3
   entries during the split.

If I got that correct, that would be clean indeed.

But then we would have a different encoding for block entries and page
entries. In page entries, DBM could be used to say if the page is
writable,
but on block entries one would have to look at the 'dirty-bit'.

Would that be ok?

Thanks!
Leo

Hi Leo,

My initial concern was that clearing all DBM bits at the start of
migration would be too expensive, so I thought distinguishing between
level-3 entries and block entries would be better.


I think we expect it to be expensive, but since we already clean the
dirty-bit (ro/rw) bit, we can have both happening in the same write :)

(since we only mark the DBM bit when we fault the memory on lazy-splitting,
we are expecting to have the same amount of writes to pagetable as we have
before HDBSS, both on faulting and 1st iteration cleaning)

Hi, Leo

Actually, I have thought about this approach too, but if we clear DBM in
kvm_pgtable_stage2_wrprotect(), then during the first round of
migration, we will fault and release RO -> W, and then add DBM.

Yeah, that's only for lazy-splitting, though.


But next time, when we migrate the dirty pages in round two, we will run
kvm_pgtable_stage2_wrprotect() again, which will clear DBM again. And
finally, HDBSS will be useless during migration.

Right, on lazy splitting, we have to clean the DBM bit on the
write-protect only if it's a block entry (hugepage).

Once it faults for the first time, it will lazy-split, and we don't need to
clean the DBM bit.



However, I ran a quick test on a 400GB VM (4 vCPUs), and the overhead
turned out to be around 30ns — which I think is acceptable.

Just a quick correction — I misstated the unit in my previous email. The
overhead for clearing DBM on the 400GB VM (4 vCPUs) was around 32 µs, not 30
ns.


Oh, that seems more likely :)

Question: is tha above amount of memory initially in Level-1 blocks,
level-2 blocks or level-3 pages? (aka: were you using explicit/transparent
hugepages?)


I'm using transparent hugepages. However, if we were to use level-3 stage-2
pages with -mem-prealloc enabled in QEMU, I believe the time cost would be
extremely high — potentially out of our control.


Yeah, that's the issue.
For this not to explode like this, we need to mark as RO only when the
entries are blocks AND we are doing lazy splitting.

We have:
Mode DBM Dirty bit
RO 0 X
WC 1 0
WD 1 1

On write-protect:
- Lazy splitting + block entry (hugepage, level 2-) -> RO
- Otherwise -> WC

On first fault, the block entry will be lazy-splitten, and we can set DBM=1
in every new page.

That way we guarantee that we are not faulting level-3 pages unecessarily,
nor need to go through the whole tree setting DBM=1 or DBM=0 on level-3
pages.

How does that sound?

Thanks!
Leo



Hi Leo,

I've also been thinking about this approach: clear
KVM_PTE_LEAF_ATTR_HI_S2_DBM when kvm_pgtable_stage2_wrprotect() calls
stage2_update_leaf_attrs(). And we can check whether a page is a block
page during the page walk, right?

Hi Tian,
That was what I was thinking :)


So we can check the page level in the walker callback
stage2_attr_walker(), filter there, clear DBM for block pages and
preserve DBM on level-3 pages. Something like this:

```
pte &= ~data->attr_clr; // wrprotect: clears S2AP_W only
pte |= data->attr_set;
if (ctx->level < KVM_PGTABLE_LAST_LEVEL)

Only on lazy splitting, right?

Or maybe we get the DBM bit on during eager splitting...

Hi, Leo

No, it works for both. Whether eager or lazy split, wrprotect always runs
before split, so the block DBM is cleared first.

Hi Tian,

On eager splitting, why should we ever strip DBM?

Eager splitting means we won't have to fault to do the splitting, so we can
have HAFDBS/HDBSS handle every fault, including the first one that we use
for lazy-splitting.


For eager split: wrprotect clears W (W=1->W=0) and strips DBM from the
block. Then the block is split into level-3 pages, but since W=0 and DBM
depends on W, DBM stays 0. DBM is only set back to 1 during the first write
fault (relax_perms: W=0->W=1, which also sets DBM=1).

I suggest that, on splitting, set all writable level-3 pages with DBM=1.
If lazy-splitting, that will happen naturally on the first fault's split.
If eager-splitting that will happen during the split as well.

So total solution looks like:
- Dirty-logging on, walk the memslot pagetable
- On block: mark as read-only
- On page: mark as writable-clean
- On splitting: Mark writable block's new pages as writable-clean
- On fault: after the split, mark the faulting page as writable-dirty

This should take care of everything, including lazy/eager splitting
differences, as well as set the groundwork for HDBSS to work properly.

What do you think?

Thanks!
Leo



Hi Leo,

Yes, that's exactly the approach in v5.

During write-protect, stage2_attr_walker() strips DBM from blocks (RO) but keeps it on level-3 pages (WC). After a block is split, the new level-3 pages are marked writable-clean. On the first write fault, relax_perms() upgrades the faulting page to writable-dirty.

So all three cases you listed are already covered:

- Blocks -> RO
- Pages -> WC
- Fault -> WD

Thanks,
Tian