Re: [PATCH v3] iommu/amd: Make PerfOpt compulsory

From: Mario Limonciello

Date: Thu Oct 01 2026 - 13:14:21 EST


On 10/1/26 06:34, Vasant Hegde wrote:


On 10/1/2026 4:10 AM, Jason Gunthorpe wrote:
[ ... 336 lines skipped ... ]
@@ -2370,7 +2370,7 @@ static int attach_device(struct device *dev,
if (ret)
goto out;
- if (dev_data->perfopt)
+ if (dev_data->perfopt && pdom_is_in_pt_mode(domain))
goto skip_caps;

This should also disable PASID support since without a gcr3 table
that's broken. Set the max_pasid to zero during probe device.

I think we can add it inside amd_iommu_probe_device().

OK, I'll rev the patch to do this.



The flow is weird like this because the driver still hasn't cleaned up
the identity/blocking domain flows. There is no need to track them in
lists and things because they never need invalidation..

Right. attach_device path cleanup is in my list. Now that _probe() path is
settled (thanks to Pranjal), this is probably next item to pick up.



This is much nicer if the above could be written in side
amd_iommu_identity_attach().

I'd say for now its fine to keep it in attach_device() as it gets called anyway.
When I rework, will make sure its moved under amd_iommu_identity_attach().


I'll leave these two items to Vasant.



Also, I'm a little confused, I thought this had to be toggled on and
off so it is only set while in identity, how does the DTE influence
what happens when in this special mode? I was sort of expecting it was
ignored entirely in HW. This version seems to lock it to always on,
but still permits a blocking domain to attach.


Yeah; AFAICT all the skips of features are mostly symbolic and for consistency when this is enabled.

I suppose really we should not permit a blocking domain to attach to make sure it's a clear message what's going on.


Mario, Can you check what happens if we move device to BLOCKED domain (DTE[v]=0) ?

-Vasant


I tested this on an APU that applies this.

Short answer: with PerfOpt enabled, moving the GPU towards a blocked DTE
has no effect -- the device keeps doing DMA as if nothing changed. This
matches Jason's expectation that the DTE is ignored entirely in HW for the
device once it is in the PerfOpt/identity fast path.

Longer answer below (I had an agent help do this check).

Method
------
To keep the normal driver bound and the GPU doing real work, I added a
small debugfs knob that flips the live DTE in place and back, reusing the
existing paths:

clear_dte_entry() -> V=1, TV=0, IR=IW=0 (same DTE as blocking domain)
set_dte_entry() -> restore passthrough

Note the AMD "blocked" DTE is V=1/TV=0, not V=0 -- make_clear_dte() keeps
V set, and TV=0 is what is supposed to block DMA. I confirmed the change
via /sys/kernel/debug/iommu/amd/devtbl:

passthrough: QWORD[0] = 0x6000000000000003 (V=1 TV=1 IR=1 IW=1)
blocked: QWORD[0] = 0x0000000000000001 (V=1 TV=0 IR=0 IW=0)

The write goes through update_dte256() + iommu_flush_dte_sync() + a
completion wait, so the IOMMU re-fetched the entry.

For a guaranteed IOMMU-routed DMA source I used amdgpu_benchmark test 1
(SDMA VRAM<->GTT copies). The GTT side is system memory, so for an
identity device it has to traverse the IOMMU DTE.

Result
------
With the DTE blocked (TV=0), SDMA still moved 2GB to and from GTT at full
speed:

dma 1024 bo moves of 1024 kB from 2 to 4 ... 5890 MB/s (GTT->VRAM)
dma 1024 bo moves of 1024 kB from 4 to 2 ... 10922 MB/s (VRAM->GTT)

(vs ~10-11 GB/s unblocked, i.e. no change). During the blocked window:

- mem_target_abort (amd_iommu PMU) = 0
- no IO_PAGE_FAULT in the event log
- no SDMA error / ring timeout / GPU reset

I also tried to witness the GPU's DMA directly via the amd_iommu PMU
filtered by devid. The filter itself works (NVMe, in a translating
domain, shows ~120M mem_trans_total under load), but the GPU reads zero
on every counter -- mem_trans_total, mem_pass_untrans, mem_target_abort --
even under the 2GB benchmark. The GPU is the only identity device on this
system, so I had no non-PerfOpt passthrough device to validate the
passthrough counters against; hence I am relying on the functional
benchmark result above rather than the zero counts.

Conclusion
----------
Once PerfOpt is enabled and the device is in the identity/pt fast path,
the DTE is effectively ignored by HW: TV=0 does not block DMA. So a
blocking-domain attach would not actually stop the device while it is in
this state.