Re: [PATCH v5 13/24] iommu/amd: Program IOMMU DTE with the private IPA domain
From: Suthikulpanit, Suravee
Date: Wed Sep 30 2026 - 02:10:50 EST
On 9/19/2026 10:26 PM, guanghuifeng@xxxxxxxxxxxxxxxxx wrote:
@@ -2770,6 +2810,51 @@ static spinlock_t *amd_iommu_get_top_lock(struct pt_iommu *iommupt)
return &pdom->lock;
}
+#if IS_ENABLED(CONFIG_AMD_IOMMU_IOMMUFD)
+/*
+ * The vIOMMU private IPA domain programs a synthetic DTE for the
+ * IOMMU's own requester ID so hardware can DMA to backing store.
+ * That object is not on pdom->dev_list: it is not an IOMMU-API
+ * attach, and walkers such as rlookup and clone_aliases assume a
+ * real struct device.
+ *
+ * amd_iommu_change_top() therefore misses it. Early maps fit under
+ * the initial page-table top (PT_FEAT_DYNAMIC_TOP). Later DevID and
+ * DomID maps at high IPA call increase_top(); the old root stays
+ * live as a child of the new one, but the self DTE still holds the
+ * old MODE and would not translate those IOVAs.
+ *
+ * Walk iommu_array (from amd_iommu_pdom_bind_iommu()) and rewrite
+ * iommu->viommu_dev_data when this domain is that IOMMU's private
+ * IPA table. set_dte_entry() skips clone_aliases() because the
+ * synthetic DTE has no struct device.
+ */
+static void update_viommu_self_dte(struct protection_domain *pdom,
+ phys_addr_t top_paddr,
+ unsigned int top_level)
+{
+ struct pdom_iommu_info *pdom_iommu_info;
+ unsigned long i;
+
+ lockdep_assert_held(&pdom->lock);
+
+ xa_for_each(&pdom->iommu_array, i, pdom_iommu_info) {
+ struct amd_iommu *iommu = pdom_iommu_info->iommu;
+
+ if (iommu->viommu_pdom != pdom || !iommu->viommu_dev_data)
+ continue;
+ set_dte_entry(iommu, iommu->viommu_dev_data, top_paddr,
+ top_level);
+ }
+}
Self DTE Mode maybe does not comply with spec requirement of Mode=100b
At initialization time, only the 8MB General Backing Storage region (IPA 0x0 - 0x800000) is mapped, so the page table top is likely level 1 or 2, resulting in Mode = 010b or 011b.
However, spec Section 2.10.1 states explicitly:
"The DTE for the IOMMU's DeviceID must be set with V=1, TV=1, GV=0, Mode=100b."
This is a hard requirement ("must"), not a recommendation. Mode=100b means 4-level page table (48-bit address space), which is necessary because the private IPA address map extends up to 48'h0050_0000_0000 (~5TB), well beyond the 39-bit limit of a 3-level table.
For GstBufferTRPMode=0, the largest IOMMU Private Address = 48'h0050_0000_0000 (39-bit address)
For GstBufferTRPMode=1, the largest IPA = 48'h0028_0000_0000 (38-bit address)
In both cases, DTE[Mode]=011b (39-bit GPA space) should be sufficient. I have discussed this with AMD IOMMU hardware designer, and has been confirmed that 011b should be sufficient. I have also experimented with DTE[Mode]=011b with the highest possible GID (i.e. 0x7FFF) and do not see issues.
Please note also that:
* Linux default to 3-level page table initially, and grow the table level automatically as higher IOVA are mapped.
* Linux supports GstBufferTRPMode=1 only.
The AMD IOMMU spec should restrict to DTE[Mode]=100b. The spec will be updated in the next revision. I am also adding comment to describe in patch series v6.
Thanks,
Suravee