Re: [PATCH v6 08/25] KVM: arm64: iommu: Shadow host stage-2 page table

From: Mostafa Saleh

Date: Tue Jul 14 2026 - 04:34:57 EST


On Tue, Jul 14, 2026 at 09:16:29AM +0100, Vincent Donnefort wrote:
> On Mon, Jul 13, 2026 at 06:19:24PM +0000, Mostafa Saleh wrote:
> >
> > That would walk the table twice, first time with a massive map, then
> > to unmap the pages that was just mapped, issuing TLB invalidations
> > for all of those, also as the IOMMU page table code never frees tables
> > that means we would immediately exahust the IOMMU pool.
>
> Well the massive map is pretty quick we would mostly just allocate the PGD? And
> then it would be just for the subsequent unmaps to split the blocks. So nothing
> to be freed really and no risk to exhaust the pool. But if you put the calls to
> the host_stage2_* functions that's pretty transparent.

Not for the IOMMU, everything is mapped with a leaf mapping, because
the pgtable code does not support splitting blocks, this is something
I plan to add support for in a follow up series after this one is
merged along side other optimizations.

>
> Just to make sure I am clear. Here's what it could look like:
>
> in iommu.c:
>
> int kvm_iommu_stage2_idmap_init(struct kvm_s2_mmu *host_mmu)
> {
> u64 addr = 0;
>
> for (i = 0; i < hyp_memblock_nr; i++) {
> kvm_iommu_stage2_idmap(addr, reg->start, IOMMU_MMIO);
>
> addr = reg->start + reg->size;
> kvm_iommu_stage2_idmap(reg->start, addr, IOMMU_READ | IOMMU_WRITE);
> }
>
> kvm_iommu_stage2_idmap(addr, kvm_phys_size(host_mmu), IOMMU_MMIO);
> }
>
> in mem_protect.c:
>
> int kvm_host_prepare_stage2(void *pgt_pool_base)
> {
> ...
>
> return kvm_iommu_stage2_idmap_init(mmu);
>
> /*
> * From here, all the changes to the host stage-2 will be reported to
> * the iommu driver
> */
> }
>
> On the pro side: no custom walker and we really couple the iommu idmap to the
> host stage-2 map.
>
> On the con side: As you've pointed out, we'd need indeed the iommu to be ready
> early enough.
>
> If you want to stick to the snapshot function, could we have the walker to
> first walk the hyp_memblock? the .arg could just pass if we are walking a MMIO
> region? That would avoid that find_mem_range_from() and I believe would
> massively simplify the snapshot walker.
>

I see, let me try that.


Unrelated to this, but while looking into option#2 I see that
__pkvm_guest_share_host() does not call either of the
host_stage2_set_owner* so that requires another explicit
pkvm_iommu_host_stage2_idmap() call.

I still plan to use option#2 in v7.

Thanks,
Mostafa