Re: [PATCH v9 00/13] PCI: liveupdate: PCI core support for Live Update

From: David Matlack

Date: Thu Sep 24 2026 - 18:03:35 EST


On 2026-09-23 04:27 PM, Zhu Yanjun wrote:
> 在 2026/9/22 11:53, David Matlack 写道:
> > On Tue, Sep 22, 2026 at 11:36 AM Zhu Yanjun<yanjun.zhu@xxxxxxxxx> wrote:
> > > 在 2026/9/18 13:06, David Matlack 写道:
> > > > Future Work
> > > > -----------
> > > >
> > > > Following this series, we expect to make further improvements to the PCI
> > > > core support for Live Update:
> > > >
> > > > - Allow P2P across Live Update by avoiding resizing or moving
> > > > preserved device BARs and preserving all upstream bridge windows.
> > > >
> > > > - Support preserving Virtual Functions by preserving SR-IOV
> > > > configuration on PFs and enumerating VFs after Live Update.
> > > Preserving the PCIe topology, bus numbers, ACS, and Bus Mastering across
> > > kexec is a foundational step for minimizing downtime.
> > >
> > > As we look toward complete end-to-end support for DMA preservation
> > > across Live Update—especially for VFIO device passthrough and dma-buf
> > > sharing scenarios—IOMMU table/domain preservation becomes crucial to
> > > prevent IOMMU page faults when devices continue performing DMA during kexec.
> > >
> > > I would like to ask about the current status and roadmap regarding IOMMU
> > > Live Update / KHO (Kexec Handover) support:
> > >
> > > Is there an ongoing effort or RFC series for IOMMU handover / page-table
> > > preservation currently in development or under discussion?
> > >
> > > How is the coordination between the PCI core Live Update mechanisms and
> > > the IOMMU subsystem being envisioned for preserving IOVA mappings (e.g.,
> > > restoring domains or handing over root tables)?
> > >
> > > Any pointers to active discussion threads, RFCs, or future plans
> > > regarding IOMMU participation in Live Update would be greatly appreciated.
> > The first IOMMU series to support Live update, can be found here:
> >
> > https://lore.kernel.org/linux-iommu/20260921004834.2601285-1-skhawaja@xxxxxxxxxx/
>
> Thanks.
>
> I have a question regarding the restoration sequencing when module
> dependencies are involved during a Live Update reboot.
>
> For the standard hardware/driver stack, the sequence seems to naturally
> follow the kernel's early initcalls and device probing (e.g., IOMMU early
> hardware handover -> PCI bus topology/BME preservation -> IOMMU domain
> attach & DMA ownership claim -> VFIO/iommufd cdev binding).
>
> However, if there is a custom kernel module or subsystem (let's call it
> Module A) that is not part of the standard PCI/IOMMU device probe callback
> chain, but strictly depends on the fully restored state of PCI, IOMMU, and
> VFIO/iommufd:
>
> 1.
>
> What is the recommended or standardized way in the Live Update
> architecture to guarantee that Module A's restoration happens
> *after* all its underlying dependencies (PCI / IOMMU / VFIO) have
> completely finished their restore processes?
>
> 2.
>
> Is the expectation to rely on LUO (Live Update Orchestrator) phase
> notification callbacks (e.g., late restore notifiers), Driver Core
> mechanisms like |-EPROBE_DEFER| / |device_link|, or something else?
>
> Any guidance on how cross-subsystem restoration order and async probe
> dependencies should be handled in the Live Update framework would be greatly
> appreciated.

Note: Some of what I write below is not yet merged into the Live Update
tree so may change. Samiullah Khawaja will be giving a talk about file
dependencies at LPC where some of these topics will be discussed.

LUO does not have a global "restore phase" or late-restore notifier.
Restoration is on-demand and driven by dependencies, and userspace does
the orchestration.

There are two types of objects that LUO manages:

1. FLB (File-Lifecycle-Bound) data, for shared/global state.

An FLB is retrieved lazily the first time someone calls
liveupdate_flb_get_incoming(). The PCI core does that from
pci_setup_device() during enumeration, and the IOMMU driver does it
when it initializes. So the order in which this global state is
restored is simply the normal boot/initcall/probe order of the
subsystems that use it. LUO does not impose any extra order.

2. Files in sessions, for per-object state (vfio cdevs, iommufds,
memfds, ...).

Files are retrieved either by userspace (LIVEUPDATE_SESSION_RETRIEVE_FD)
or by kernel code (liveupdate_get_file_incoming()). Retrieval can
happen in any order and is idempotent. When one file depends on
another, the dependency is expressed between the two files, not
through a global phase.

VFIO -> iommufd is an example to follow (see Samiullah's IOMMU series
[1] on top of Vipin's VFIO series [2]):

- Outgoing: when a vfio cdev is preserved, iommufd_device_preserve()
calls liveupdate_get_token_outgoing() on the iommufd file the device
is attached to. That fails unless userspace has already preserved
the iommufd in the same session, so the dependency is enforced at
preserve time. The iommufd token is then recorded in the preserved
device state.

- Incoming: the recorded token is what connects the device back to its
iommufd after kexec. Retrieving and re-attaching to the restored
iommufd is the next phase of the IOMMU work, so that part is not in
[1] yet. The building blocks are there, though:
liveupdate_get_file_incoming() lets one file handler pull in a file
it depends on by token (it is idempotent, so the order userspace
retrieves in doesn't matter), and ->can_finish() lets a handler
block LIVEUPDATE_SESSION_FINISH until everything it depends on is in
a consistent state.

- Device binding: vfio-pci's ->retrieve() simply fails (-ENODEV) if the
preserved device is not bound to vfio-pci yet. There is no probe
deferral inside LUO. Userspace is expected to make sure the device
is bound (e.g. wait for udev) before retrieving the file. In the
meantime the device is protected: the PCI core keeps bus mastering
and BDFs stable, and the IOMMU core reattaches the preserved domain
and claims DMA ownership, so no other driver can bind to it.

If you can share more about what Module A is and what state it needs to
carry across the update, I can try to go into more detail. But
generically I would recommend:

- If Module A has state that must survive the Live Update, model it as
a LUO file handler. If it depends on specific VFIO/iommufd files,
record their tokens with liveupdate_get_token_outgoing() in your
->preserve(). Then in your ->retrieve(), get them back with
liveupdate_get_file_incoming(). If something isn't ready yet (e.g.
the device hasn't probed), fail ->retrieve() and let userspace retry,
and use ->can_finish() to keep the session from finishing too early.

- If Module A is a driver that binds to a device, the normal driver
core mechanisms (-EPROBE_DEFER, device links) still apply for
probe-time dependencies. Live Update doesn't replace them. But
restoring the preserved state itself should probably go through LUO
as above.

- If Module A doesn't need to preserve anything itself but just
consumes the restored VFIO/iommufd objects, the simplest option is to
let userspace sequence it. Userspace already knows when the devices
are bound and the FDs have been retrieved, so it can hand them to
Module A at that point.

[1] https://lore.kernel.org/linux-iommu/20260921004834.2601285-1-skhawaja@xxxxxxxxxx/
[2] https://lore.kernel.org/kvm/20260714151505.3466855-1-vipinsh@xxxxxxxxxx/