Re: [PATCH v4 02/21] iommufd: Add iommufd_sw_map_msi()

From: Jason Gunthorpe

Date: Sat Aug 22 2026 - 10:06:41 EST


On Sat, Aug 22, 2026 at 03:50:05PM +0200, Andrew Jones wrote:

> > The viommu can set the msi_addr_pattern
>
> msi_addr_pattern should be under the control of the hypervisor since
> it, and its companion msi_addr_mask, represent a block of GPAs. It
> will necessarily have to match the vIMSIC topology described by the
> VMM to work anyway.

I mean the iommufd viommu, in the hypervisor.

> > It is unfortunate you can't learn the vCPU the MSI is targetting from
> > the MSI descriptor, in terms of linux that's a pretty difficult choice
> > to implement.
>
> With the RISC-V IOMMU MSI table and irqbypass support in KVM, we can
> write the guest's MSI messages directly to the device without
> interpreting the IOVAs.

Every architecture can work this way. Nobody has impleented Linux
support for it, and if we do, it must be arch generic.

Not being able to discover the vCPU from the MSI descriptor means you
can't use any of the existing less-optimal Linux flows and are forced
to implement this hard thing..

> Guest changes to S1 or IRQ affinity therefore do not require MSI table
> updates. And, when a vCPU moves to another pCPU, or switches between a
> VS-file and an MRIF, the hypervisor only updates the MSI table entry for
> that vIMSIC GPA. The device message and S1 mapping are unaffected.

I'm not sure how you will end up controlling the msi table..

> As I understand it, Nicolin's series addresses the ARM split between
> SMMU translation and ITS interrupt remapping.

ARM is basically the same, the ITS page goes through the S1 and S2, so
the goal is to get a valid ITS page into the S2, tell the guest to use
it and create a S1 pointing at it then feed the MSI descriptor from
the guest unmodified to the HW.

Exactly the same as what you want.

> The guest-selected IOVA must be preserved for S1 while ITS routing
> is managed separately. RISC-V performs MSI remapping in the IOMMU
> after S1, so guest MSI target changes do not require corresponding
> per-vector coordination with a separate interrupt controller.

In ARM you'd manage vCPU mapping in the GIC since it is already doing
translation lookups, this happens without involving the iommu. It
doesn't need to change the descriptor to do this.

But conceptually it is the same thing, there is a part along the MSI
chain that maps from vCPU to pCPU. Either in the GIC HW tables or in
the IOMMU MSI tables, doesn't matter much, IMHO.

I think the two x86's also have a very similar thing where the iommu
interrupt remapping tables can do the vCPU to pCPU translation.

Again nobody has built this and if we do it must somehow be able to
work generically for all arches, so it will take a while to get it
sorted out I think. Since riscv's HW design boxed themselves into
doing all this work I guess you have no choice.

I strongly recommend you keep it seperate from this initial series
(And remove the irq domains from this series) and come with a proposal
that everyone can understand how an arch could hook into all of this
machinery.

Jason