Re: [RFC PATCH v6 08/11] iommufd: Add vIOMMU provider support
From: Jason Gunthorpe
Date: Tue Sep 29 2026 - 09:15:21 EST
On Tue, Sep 29, 2026 at 06:15:39PM +0530, Aneesh Kumar K.V wrote:
> Jason Gunthorpe <jgg@xxxxxxxxxx> writes:
>
> > On Tue, Sep 29, 2026 at 11:44:28AM +0530, Aneesh Kumar K.V wrote:
> >
> >> tsm_viommu_get_ops() runs with pci_tsm_rwsem held for read and takes a
> >> temporary reference on the backend module (the CCA module). This keeps
> >> the selected ops callable until vIOMMU initialization completes. iommufd
> >> then drops the reference. This does not pin a particular TSM
> >> registration or prevent tsm_unregister().
> >
> > That seems over complicated. Maybe we can't get to the sane locking I
> > suggested earlier where TSM module is stable while a driver is bound,
> > but we absolutely must have sane locking where we can "pin" the tsm
> > for a pdev and it cannot be unregistered for long periods of time,
> > such as while a viommu/vdev exists.
> >
> > That period should start right before getting the ops and continue to
> > until the viommu is destroyed.
> >
> > No hot unplug of tsm modules while things are active.
> >
>
>
> That is essentially how it works. I decided to take the module reference
> here and the tsm_dev reference in viommu_init() to keep the rest of
> viommu_alloc() cleaner. That is, we have:
>
> struct module *owner = NULL;
>
> ops = tsm_viommu_get_ops(idev->dev, cmd->type, &owner);
>
> if (!ops) {
> ops = iommu_dev->ops->get_viommu_ops(idev->dev, cmd->type);
>
> rc = ops->viommu_init(viommu, idev->dev,
> if (rc)
> goto out_put_hwpt;
>
> out_put_idev:
> module_put(owner);
if we have a get_ops we need a put_ops()..
> The tsm_dev and module details are needed only by tsm_viommu, not by a
> generic SMMU driver. For viommu_init() to take ownership of the resources
> acquired by get_ops(), I would either need to add a viommu_info argument
> to viommu_init(), affecting all IOMMU driver implementations, or make the
> error handling conditional and awkward.
Just don't, iommufd can hold the tsm ops if it knows it created the
viommu through tsm.
> >> CCA vIOMMU initialization takes a tsm_dev reference, keeping the TSM
> >> object and its PCI/TSM resources alive until the vIOMMU is
> >> destroyed.
> >
> > iommufd should do this, so long as the viommu object exists the tsm
> > for it exists. It should not be inside tsm drivers.
> >
>
> That would expose more TSM details to iommufd. The reference must also
> be acquired under pci_tsm_rwsem.
That seems like overkill.
> /* Pin the current TSM and revalidate the selected vIOMMU operations. */
> struct tsm_dev * pci_tsm_viommu_get_tsm_dev(struct pci_dev *pdev,
> enum iommu_viommu_type type,
> const struct iommufd_viommu_ops *expected_ops)
> {
> const struct pci_tsm_ops *ops;
> const struct iommufd_viommu_ops *viommu_ops;
> struct tsm_dev *tsm_dev;
> int ret = -ENODEV;
>
> {
> guard(rwsem_read)(&pci_tsm_rwsem);
> if (!pdev->tsm)
> return ERR_PTR(-ENODEV);
> if (pdev->tsm->tsm_dev->unregistering)
> return ERR_PTR(-ENODEV);
> ops = to_pci_tsm_ops(pdev->tsm);
> if (!ops->viommu_get_ops)
> return ERR_PTR(-ENODEV);
> tsm_dev = pdev->tsm->tsm_dev;
> /* Unlike get_device(), this also keeps PCI/TSM resources active. */
> if (!tsm_try_get(tsm_dev))
> return ERR_PTR(-ENODEV);
This is way too complicated for what should be a very simple scheme :(
rcu_read_lock()
tsm = rcu_derference(pdev->tsm);
if (!tsm)
return NULL;
/* Prevent the module from unloading which must be the only way to
trigger unregister */
if (!try_module_get(tsm->ops->module))
return NULL;
rcu_read_unlock()
And this should be exposed to iommufd..
viommu_create:
tsm = tsm_get_device(pdev)
tsm_get_viommu_ops(tsm,...)
[..]
viommu_destroy:
tsm_put_device(tsm)
No lock, no registering FSM, just RCU free the pdev->tsm memory and do
that module unload rcu synchronize during module unload. No hand of of
the lifecylce to other layers.
Hold the module get in iommufd inside the iommufd viommu object.
It is simple and easy to understand.
> viommu_ops = ops->viommu_get_ops(&pdev->dev, type);
> if (viommu_ops == expected_ops)
> return tsm_dev;
> if (IS_ERR(viommu_ops))
> ret = PTR_ERR(viommu_ops);
> }
This should just be try_module_get and a touch of RCU.
> > ideally the module refcount handles this and the only way to trigger a
> > tsm remove is through module unload, with no sysfs path?
>
> I agree. The existing code takes extra care to allow tsm_unregister().
> However, if unloading arm-cca-host.ko is the only way to trigger
> unregistration, as it currently is, we can avoid this complexity.
I've been badly traumitized by hot unplug races, bugs and deadlock,
let's not introduce anything like this here, there is no need.
Jason