Re: [PATCH 12/13] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver
From: Danilo Krummrich
Date: Wed Sep 16 2026 - 11:57:24 EST
On Wed Sep 16, 2026 at 4:17 PM CEST, Jason Gunthorpe wrote:
> On Tue, Sep 15, 2026 at 11:19:01PM +0200, Danilo Krummrich wrote:
>
>> The driver_data pointer in struct device is defined to be a pointer where a
>> driver can store *arbitrary* data for the duration the driver is bound to this
>> device.
>
> There are places in the kernel where the drvdata of the bound device
> ends up owned by the subsystem, not the end driver. Yes, an ideal
> driver + subsystem should never even need drvdatab beyond remove. Yet,
> things are not perfect..
>
> The fundamental issue is some kernel API surfaces that the subystem
> needs to work with only provide a struct device in their callbacks and
> the subystem has no option but to use the drvdata for its own purpose
> to recover the subsystem specific data.
>
> For example VFIO hooks into this nasty API:
>
> ret = vga_client_register(pdev, vfio_pci_set_decode);
> if (ret)
> return ret;
>
> Which doesn't provide a void * token to pass the core code's
> struct.
>
> Another is all the PCI callbacks which assume the op behind them uses
> drvdata to get its data.
This is where the class (i.e. vfio-pci) should instead take a driver callback,
as only the driver really knowns about the layering details.
Usually, class device implementations can't make assumptions of the underlying
bus, because they have to work for any bus. I.e. there's no other way than
providing helpers and letting drivers do the glue code between the bus and the
class device.
> It is not necessarily easy to fix. Rrouting all those PCI callbacks
> through trampolines in every single driver is really not an appealing
> design. There are alot of VFIO drivers.
The reason this seems undesirable from a vfio-pci perspective is that it is
special in the sense that it is a class device that is specifically built to sit
on top of a spcific bus device (i.e. struct pci_dev).
Since this is a rare (maybe even unique?) edge case, the driver core has indeed
no infrastructure in place to represent this.
> Maybe it needs a dev->drvdata and dev->subsystem_data, maybe it needs
> some PCI thing where the pm ops can get a void *, IDK.
It still makes me think that there should be some closer integration of vfio-pci
with the PCI core, as it is specifically built for this bus.
In fact, what you say above is the generalization of my "hack" [1], which, the
more I think about this, seems actually less of a hack. :)
I would object a bit to the generalization with dev->subsystem_data as it
screams for abuse, but the thing in [1] seems more and more reasonable to me.
(Another option would be to really split it up and have a virtual vfio-pci bus
on top of PCI, which would also allow for custom match logic, but that also
seems pretty overkill.)
>> As mentoined above, there are people volunteering now. Without starting it, it
>> can't scale further than that. :)
>
> There are lots of other vfio patches that need attention too, and it
> seems we are short of that more than anything. Now you need to do a
> bunch of C refactoring patches as well just to get things ready to
> show a bunch of rust code. It is a lot of work.
It's only the drvdata thing, what else is missing?
> I don't really understand in a nutshell why we should do this for nova
> the mails were so long... Can we not just ignore the lifetime
> imperfection for this?
I mentioned some points in the first two paragraphs of [2]. Besides that, I
don't see a reason why we should spend time and effort for working out the
inferior solution, where the better alternative is even less effort, contributes
to better quality and stability of the whole driver project and also offers a
chance for the vfio subsystem to gain new contributors and gather experience
with the language that has proven itself in many areas already.
I mean, there'd still be the option to make it an experiment and say let's add
the abstractions and the nvidia-vgpu driver and see how it works out for a while
before allowing more Rust drivers. And if it really turns out to be bad, it
should also be easy to rip it out, replace nvidia-vgpu with a C driver and throw
in a crappy FFI layer. :)
>> However, I don't really see the use-case; you can't load nvidia-vgpu without
>> nova-core in the first place, so it would require to unbind nova-core through
>> sysfs force unbind, no?
>
> nvidia gpu is a more unique scenario, if you are building a general
> bindings it has to support the flows like this.
Sure, as mentioned, the code I sketched up should already be able to do this.
> I've wanted to rework the way the common ops are shimmed in for a
> while, you'd probably want to do that before rust bindings.
TBH, I don't think it makes a difference; having Rust abstractions is pretty
much the same as having another driver. I.e. it would be equivalent to saying
"before we accept another pci-vfio driver we need to do some rework".
Actually, I even think it can be an advantage, as you could think of the Rust
abstractions like a driver with built-in correctness checks, so it can help to
validate the changes.
Of course, this requires the responsibles of the Rust code to help with that and
I think we have this commitment.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/dakr/linux.git/commit/?id=76b3bfd6386a01f338780e1a37b5bad3f5a48d31
[2] https://lore.kernel.org/nova-gpu/DLCRZLO06SIO.LS7TWQXIPZSQ@xxxxxxxxxx/