Re: [PATCH] PCI: Reserve prefetchable window headroom for Resizable BARs

From: Christian König

Date: Mon Aug 31 2026 - 17:43:41 EST


On 8/31/26 19:49, Geramy Loveless wrote:
> On 8/31/26 9:58 AM, Ilpo Järvinen wrote:
>> On Mon, 31 Aug 2026, Christian König wrote:
>>
>>> On 8/28/26 23:37, Geramy Loveless wrote:
>>>> Firmware typically sizes prefetchable bridge windows for the boot-time
>>>> BAR size. Behind a fixed (non-hotplug) PCIe switch fabric, there is then
>>>> no room for a driver to grow a Resizable BAR afterwards: every window
>>>> from the leaf up to the root is sized for the small BAR, so
>>>> pci_resize_resource() fails with -ENOSPC.
>>>>
>>>> Furthermore, a small prefetchable BAR (e.g., a 2 MiB doorbell) sharing a
>>>> bridge's single prefetchable window with a much larger one (e.g., a 32 GiB
>>>> VRAM BAR) pushes the required window size past the large BAR's alignment.
>>>> Because bridge windows round up to a power of two, this forces a massive
>>>> alignment waste (e.g., 32 GiB + 2 MiB rounds up to a 64 GiB window,
>>>> wasting ~32 GiB per GPU).
>>>>
>>>> This patch solves both issues to enable ReBAR on cascaded switch fabrics:
>>>>
>>>> 1. Reserve Headroom:
>>>> Reserve prefetchable window headroom for the maximum size of each
>>>> downstream Resizable BAR during the bridge sizing pass. The device BAR
>>>> and hardware ReBAR are left at their boot size to prevent tearing down
>>>> firmware-loaded state (e.g., AMD R9700 PSP) before the driver binds.
>>> Yeah that was suggested before but that is clearly not something you can
>>> do in common code.
>>>
>>> The ReBAR fields often doesn't reflect the actual needed space but
>>> rather the maximum the HW address logic can resolve.
>>>
>>> So what you end up with is allocating multiple TiB for a window which
>>> just needs few GiB, sometimes even completely overflowing the 64bit
>>> address space made available by the root complex.
>
> Yeah I could imagine that would be bad, hence I tried to compact the 64-bit space as well. As far as I understand PCI/PCIe standards its making this patch difficult.
> Christian I'm not sure if you could share these reference documents but that probably would be a better start for me to look at before I update the patch or make changes, also I need to wait for a review on the last patch I submitted too, so that leaves me with some time to review some standards if you know specifically where to look, if not thats fine too.
>
>> Yes. A naive approach to (only) go to the max ReBAR allows just doesn't
>> work well enough to be usable in general case.
> I did not know the max would report above the amount is actually needed because of vendors decisions in the cards, that's interesting.

Yeah as I said before to actually figure out how large your BAR needs to be for an AMD GPU you have to load quite a bit of FW and probe the connected RAM chips.

If the GPU is your initial display device that is not much of a problem because that happens quite early with the VBIOS load, but if you have non display HW the driver actually needs to do this ASIC init/reset.

The bus interface just gives out the possible BAR sizes which could potentially make sense on the board.

At least AMD has also produced quite a number of GPUs with bugs in the ReBAR registers, so for example Navi 1x reports possible BAR sizes of only between 1 - 4 MiB while it should be 8GiB at least.

I clearly don't recommend using that register as source for the resource estimation for upstream bridges.

Regards,
Christian.