Re: [PATCH v5 11/18] iommu: Restore and reattach preserved domains to devices

From: Samiullah Khawaja

Date: Fri Oct 09 2026 - 23:51:24 EST


On Wed, Oct 07, 2026 at 12:44:31PM -0700, Nicolin Chen wrote:
On Mon, Sep 21, 2026 at 12:48:27AM +0000, Samiullah Khawaja wrote:
@@ -694,7 +700,8 @@ static int __iommu_probe_device(struct device *dev, struct list_head *group_list
}

for_each_group_device(group, gdev2) {
- if (dev_iommu_preserved_state(gdev2->dev)) {
+ if (dev_iommu_preserved_state(gdev2->dev) ||
+ dev_iommu_restored_state(gdev2->dev)) {
ret = -EBUSY;
goto err_free_gdev;

Maybe it should -EBUSY on a group that already has a device so it
wouldn't end up with a multi-device group.

Agreed. This should make sure that a preserved device does not get added
into a group that already has a device. Will update.

@@ -2211,6 +2242,7 @@ static int __iommu_attach_device(struct iommu_domain *domain,
ret = domain->ops->attach_dev(domain, dev, old);
if (ret)
return ret;
+
dev->iommu->attach_deferred = 0;
trace_attach_device_to_domain(dev);
return 0;

Unnecessary change.

Will remove.

@@ -3175,6 +3207,62 @@ int iommu_fwspec_add_ids(struct device *dev, const u32 *ids, int num_ids)
}
EXPORT_SYMBOL_GPL(iommu_fwspec_add_ids);

+static struct device *__iommu_group_restored_device(struct iommu_group *group)
+{
+ struct group_device *gdev;
+
+ lockdep_assert_held(&group->mutex);
+ for_each_group_device(group, gdev) {
+ if (!dev_is_pci(gdev->dev))
+ continue;
+
+ if (dev_iommu_restored_state(gdev->dev))
+ return gdev->dev;

list_first_entry instead of for_each_group_device since there's a
singleton enforcement.

Agreed. Will Update.

+ }
+
+ return NULL;
+}
+
+static int __iommu_group_restore_domain(struct iommu_group *group)
+{
+ struct iommu_device_ser *device_ser;
+ struct iommu_domain *domain;
+ struct device *dev;
+ void *owner;
+ int ret;
+
+ lockdep_assert_held(&group->mutex);
+ if (group->domain)
+ return -EBUSY;
+
+ dev = __iommu_group_restored_device(group);
+ device_ser = dev_iommu_restored_state(dev);
+ if (!device_ser)
+ return -ENOENT;
+
+ ret = __iommu_group_alloc_blocking_domain(group);
+ if (ret)
+ return ret;
+
+ domain = iommu_restore_domain(dev, device_ser, &owner);
+ if (WARN_ON(IS_ERR(domain)))
+ return PTR_ERR(domain);
+
+ /* The restored domain is attached with the restored device. */
+ ret = __iommu_group_set_domain(group, domain);
+ if (ret)
+ return ret;

If (ret), how about the restored domain by iommu_restore_domain()?

The restored domain is not leaked and it remains restored and associated
with the preserved device in FLB and will be reused later if there is
another rescan.

+ /*
+ * Ownership of groups with preserved devices is set during boot. These
+ * will be reclaimed later by the entity (iommufd) that preserved them.
+ */
+ WARN_ON(group->owner);
+ group->owner = owner;
+ group->owner_cnt = 1;
+ return ret;
+}
+
/**
* iommu_setup_default_domain - Set the default_domain for the group
* @group: Group to change
@@ -3233,6 +3321,16 @@ static int iommu_setup_default_domain(struct iommu_group *group,

/* We must set default_domain early for __iommu_device_set_domain */
group->default_domain = dom;
+
+ /* Preserved devices need to be attached to the restore domain */
+ if (__iommu_group_restored_device(group)) {
+ ret = __iommu_group_restore_domain(group);

__iommu_group_restored_device is called twice: here (outside) and
inside __iommu_group_restore_domain.

Perhaps change to:
dev = __iommu_group_restored_device(group);
if (dev) {
ret = __iommu_device_restore_domain(dev);
...
?

I will have to get the group again inside the
__iommu_device_restore_domain(), but I think that is fine. Will update
this.

+void iommu_init_device_preserved_data(struct device *dev)
+{
+ struct iommu_device_ser *device_ser = NULL;

"= NULL" doesn't seem necessary.

Will remove.

+ struct iommu_device_array_ser *array;
+ struct iommu_flb_obj *flb_obj;
+ int ret, idx;
+
+ if (!dev_is_pci(dev))
+ return;
+
+ ret = iommu_liveupdate_flb_get_incoming(&flb_obj);
+ if (ret)
+ return;
+
+ mutex_lock(&flb_obj->lock);
+ array = phys_to_virt(flb_obj->ser->device_array_phys);
+ iommu_liveupdate_for_each_arr(array) {
+ iommu_liveupdate_for_each_obj(array, device_ser, idx) {
+ if (match_device_ser(device_ser, to_pci_dev(dev))) {
+ device_ser->hdr.flags |= IOMMU_SER_FLAG_INCOMING;
+ goto out;
+ }
+ }
+ }
+
+ device_ser = NULL;
+out:
+ WRITE_ONCE(dev->iommu->device_ser, device_ser);

dev->iommu->device_ser is NULL after kzalloc.

So, maybe drop "device_ser = NULL" and move WRITE_ONCE() into the
loop (under match_device_ser)?

Will update.

+ mutex_unlock(&flb_obj->lock);
+ liveupdate_flb_put_incoming(&iommu_flb);

Hmm, you might want to check the lifecycle of this flb thing.

dev->iommu->device_ser points to something inside the flb, which
might be freed somewhere?

The lifecycle is bound to the FD that is preserved into LUO. The
incoming FLB is only freed when that FD is finished, and the device_ser
will be cleared before finish. That logic is not part of phase 1, as it
doesn't do the iommufd restore. In phase 1, iommufd's can_finish()
always returns false, so the FLB is never freed.

+struct iommu_domain *iommu_restore_domain(struct device *dev,
[...]
+ domain_ser = phys_to_virt(ser->domain_iommu_ser.domain_phys);
+ if (domain_ser->restored_domain) {
+ *owner = ser;
+ domain = domain_ser->restored_domain;
+ goto out;
+ }
+
+ domain_ser->hdr.flags |= IOMMU_SER_FLAG_INCOMING;

Drop the extra space before "IOMMU".

Will update.

+/**
+ * dev_iommu_restore_did() - Get restored domain ID for a device
+ * @dev: Target device
+ * @domain: Target domain
+ *
+ * Fetches the domain ID preserved for @dev and @domain across Live Update.
+ *
+ * Return: Domain ID or -1 on error.
+ */
+static inline int dev_iommu_restore_did(struct device *dev, struct iommu_domain *domain)

"did" is an intel thing..

+{
+ struct iommu_device_ser *ser = dev_iommu_restored_state(dev);
+
+ if (ser && iommu_domain_restored_state(domain))
+ return ser->domain_iommu_ser.attachment_id;

... so, it could be just dev_iommu_restored_attachment_id()?

Yes, this looks good. I will update it.

Nicolin

Thanks,
Sami