Re: [PATCH v5 13/18] iommu/vt-d: Preserve PASID table of preserved device

From: Samiullah Khawaja

Date: Fri Oct 09 2026 - 22:32:18 EST


On Fri, Oct 09, 2026 at 11:38:57AM +0800, Baolu Lu wrote:
On 9/21/2026 8:48 AM, Samiullah Khawaja wrote:
In scalable mode the PASID table is used to fetch the io page tables.
Preserve and restore the PASID table of the preserved devices.

Signed-off-by: Samiullah Khawaja <skhawaja@xxxxxxxxxx>
---
drivers/iommu/intel/liveupdate.c | 141 +++++++++++++++++++++++++++++--
drivers/iommu/intel/pasid.c | 10 ++-
drivers/iommu/intel/pasid.h | 8 ++
include/linux/kho/abi/iommu.h | 17 ++++
4 files changed, 169 insertions(+), 7 deletions(-)

Should this patch come before patch 12/18, which restores the device’s
domain attachment?

Agreed. I will move it before the patch that restores the device's
domain attachment.



[snip]

diff --git a/drivers/iommu/intel/pasid.c b/drivers/iommu/intel/pasid.c
index e4f24d3f19a6..59cc69383799 100644
--- a/drivers/iommu/intel/pasid.c
+++ b/drivers/iommu/intel/pasid.c
@@ -13,6 +13,7 @@
#include <linux/cpufeature.h>
#include <linux/dmar.h>
#include <linux/iommu.h>
+#include <linux/iommu-liveupdate.h>
#include <linux/memory.h>
#include <linux/pci.h>
#include <linux/pci-ats.h>
@@ -60,8 +61,13 @@ int intel_pasid_alloc_table(struct device *dev)
size = max_pasid >> (PASID_PDE_SHIFT - 3);
order = size ? get_order(size) : 0;
- dir = iommu_alloc_pages_node_sz(info->iommu->node, GFP_KERNEL,
- 1 << (order + PAGE_SHIFT));
+
+ max_pasid = 1 << (order + PAGE_SHIFT + 3);
+ if (dev_iommu_restored_state(dev))
+ dir = intel_pasid_restore_table(dev, max_pasid);
+ else
+ dir = iommu_alloc_pages_node_sz(info->iommu->node, GFP_KERNEL,
+ 1 << (order + PAGE_SHIFT));

This restores only the PASID table pages. The PASID table is eventually
installed in the context entry during probe_device:

if (sm_supported(iommu) && !dev_is_real_dma_subdevice(dev)) {
ret = intel_pasid_alloc_table(dev);
if (ret) {
dev_err(dev, "PASID table allocation failed\n");
goto clear_rbtree;
}

if (!context_copied(iommu, info->bus, info->devfn)) {
ret = intel_pasid_setup_sm_context(dev);
if (ret)
goto free_table;
}
}

For a restored PASID table, the call to intel_pasid_setup_sm_context()
should be skipped. Instead, it should check whether the preserved pasid
table is compatible with the new kernel environment.

I will skip the setup call here, but the compatibility check is done in
the intel_pasid_restore_table(). The command line configuration and
other things can be verified during iommu unit restore as you pointed
out in the other patch.

if (!dir) {
kfree(pasid_table);
return -ENOMEM;
diff --git a/include/linux/kho/abi/iommu.h b/include/linux/kho/abi/iommu.h

[snip]

index 5aaa29da6832..308cd83fd3e3 100644
--- a/include/linux/kho/abi/iommu.h
+++ b/include/linux/kho/abi/iommu.h
@@ -129,6 +129,19 @@ struct iommu_dev_map_ser {
u64 iommu_phys;
} __packed;
+/**
+ * struct iommu_device_intel_ser - Intel specific state of serialized device
+ * @restored: Whether the device state is restored
+ * @pasid_table: Physical address of pasid table
+ * @max_pasid: Maximum supported pasid
+ */
+struct iommu_device_intel_ser {
+ u8 restored;
+ u8 padding[7];
+ u64 pasid_table;
+ u64 max_pasid;
+} __packed;
+
/**
* struct iommu_device_ser - Serialized state of a device
* @hdr: Common object header
@@ -136,6 +149,7 @@ struct iommu_dev_map_ser {
* @pci_domain_nr: PCI domain number
* @dma_owner_token: Token to identify the DMA owner of this device
* @domain_iommu_ser: Domain and IOMMU mapping
+ * @intel: Intel specific serialization data
*/
struct iommu_device_ser {
struct iommu_hdr_ser hdr;
@@ -143,6 +157,9 @@ struct iommu_device_ser {
u32 pci_domain_nr;
u64 dma_owner_token;
struct iommu_dev_map_ser domain_iommu_ser;
+ union {
+ struct iommu_device_intel_ser intel;
+ };
} __packed;
/* There are maximum 256 buses, so maximum 512 context tables */

Thanks,
baolu

Thanks,
Sami