Re: [RFC PATCH] PCI: Tolerate non-prefetchable 64-bit BARs in prefetchable windows
From: Ard Biesheuvel
Date: Mon Sep 14 2026 - 09:17:07 EST
On Mon, 14 Sep 2026, at 13:38, Ilpo Järvinen wrote:
> On Fri, 11 Sep 2026, Ard Biesheuvel wrote:
>> On Fri, 11 Sep 2026, at 11:27, Ilpo Järvinen wrote:
>> > On Thu, 10 Sep 2026, Ard Biesheuvel wrote:
>> >> On Thu, 10 Sep 2026, at 19:31, Ilpo Järvinen wrote:
...
>> >> >
>> >> > I was more thinking along the lines of always setting IORESOURCE_PREFETCH
>> >> > for 64-bit BARs on PCIe devices but I've not had time to look at that/test
>> >> > how many things would break as a result. ...It would seem much simpler
>> >> > solution to differentiate PCI from PCIe while keeping the existing logic
>> >> > without adding similar pci_is_pcie() checks everywhere.
>> >> >
>> >>
>> >> I agree that the code changes would be much simpler.
>> >>
>> >> However, would this impact the PCI metadata observed by all consumers,
>> >> including userspace,
>> >
>> > Yes, it will impact userspace. But quoting you from above:
>> >
>> > "is being abused to inform memory mapping attributes and other device/BAR
>> > level properties that it was never intended for."
>> >
>> > What does the userspace then do with the information? Does it qualify
>> > under "it was never intended for"?
>> >
>>
>> Perhaps, but that does not mean we are allowed to break it now.
>>
>> More specifically, would the output of 'lspci' change as a result?
>
> I'm not sure and I'm under impression there are multiple ways lspci can
> derive its information.
>
>> >> and drivers that may expect a certain BAR layout,
>> >
>> > ??? Would that even be spec compliant??
>>
>> Again, maybe not, but if something that works fine today stops
>> working because the PCI subsystem started lying to the driver about
>> how the PCIe device describes itself, we'll be on the hook to fix it.
>
> Understood.
>
> So I suppose you're not wanting to do this for the actual placement
> algorithm then because of the same risk but only cover the case where FW
> placed 64-bit non-pref BAR into prefetchable window like this patch
> currently does?
>
Basically. The firmware knows the platform better than the OS, so if it
produces a resource allocation with non-prefetchable PCIe BARs in
prefetchable bridge windows, there is a good chance this was deliberate.
> ...And limiting to that case only likely implies remove + rescan cycle may
> fail because resources can no longer be placed into the same windows
> which will surprise user (arguably, not the most common use case but
> definitely surprising for the user if the kernel cannot place the
> resources the same way they were after boot => another way to get problem
> reports).
>
> (Unrelated to this change, I'm going to open that BAR placement can of
> worms myself in a week or two because of resource placement changes I'm
> preparing. And I already hit one such problem where a particular BAR
> address results in probe failure I'll probably have to quirk around. And
> I didn't even have to move things into another window to trigger that.)
>
Awesome :-)
>> >> and/or base decisions about memory attributes on this?
>> >
>> > The point is to consider them 64-bit window eligible so yes, kernel would
>> > definitely be basing decision on that but that's intentional.
>>
>> That would mean that a driver may decide to use ioremap_wc() rather than
>> ioremap() to map a non-prefetchable BAR that we decided to misrepresent
>> as a prefetchable one. Even if the (pseudo-)PCI-PCI bridge will not do
>> any readahead, the CPU or interconnect may behave very differently as a
>> result, and touch BAR regions that the driver never accessed explicitly.
>>
>> >> In particular, I am concerned about non-prefetchable BARs that actually
>> >> have side effects on read, being mapped with WC (or Normal-NC on arm64)
>> >> semantics, where the interconnect may widen, combine or reorder accesses.
>> >
>> > So on a more concrete terms, you're referring to the check in
>> > proc_bus_pci_mmap()? And the one in __pci_resource_attr_is_visible() +
>> > pci_dev_resource_wc_is_visible()? I suppose that wouldn't work then.
>>
>> No, I am referring to the hundreds of ioremap() and ioremap_wc() calls
>> under drivers. Maybe none of them are affected but who knows.
>
> I'm left to wonder how many of those are based on a IORESOURCE_PREFETCH
> check, I strongly suspect none. Somehow I feel I'm the only one running
> git grep and I fail to locate any examples of the problem you mention. No
> offense meant, I just feel we're dicussion code that is hypothetical and
> doesn't exist for real, and definitely not on large scale.
>
I can assure you you are not the only one running 'git grep'.
But I feel that presenting non-prefetchable BARs as prefetchable in every
respect just to simplify our logic wrt bridge window placement is not the
right approach. We must assume that non-prefetchable BARs may have side
effects on read - we are just relying on the fact that PCIe ports modeled
as PPBs will never perform readahead through their prefetchable windows.
For instance, such non-prefetchable BARs will end up having resourceN_wc
sysfs nodes created for them. They will also become mmap()'able with
WC attributes through /proc/bus/pci (via proc_bus_pci_mmap()).
> That being said, I don't have problem in accepting adding
> IORESOURCE_PREFETCH to all PCIe resources was not a good idea so there's
> no point in continuing discussion towards this direction.
>
Ok.
>> > So if just setting IORESOURCE_PREFETCH is not workable, how about adding a
>> > getter for res->flags which adds IORESOURCE_PREFETCH into the returned
>> > flags if it's PCIe device and (in the end) use the raw value only in those
>> > places that actually care about wc distinction. What I don't want to see
>> > us adding that pci_is_pcie() everywhere.
>>
>> Maybe add another IORESOURCE_PREFETCH_xxx flag that indicates that the
>> resource may be placed in a prefetchable bridge window? We'd only have
>> to set it in a single place (when probing the BAR), and we can add
>> support for it piecemeal in the validation and allocation logic.
>
> Unfortunately, I've earlier discovered there's no "single place" unless
> we add some gross res->flags fixup hack into PCI core. See e.g., the
> commit bdb32359eab9 ("sparc/PCI: Correct 64-bit non-pref -> pref BAR
> resources"), which we could hopefully revert after your change! So I'm
> afraid if you go to the new flag approach, besides drivers/pci/ you'd have
> to hunt down these from under arch/, and likely miss a few in the process.
>
> Also, res->flags is currently full (for 32-bit) so you'd need to make the
> field (and therefore struct resource) larger.
>
Shame.
> So I still suggest having a flags getter for this purpose would be better.
> In the complete solution covering also resource fitting and assignment
> algorithm, the main complication with that approach comes from the few
> cases that don't have the struct pci_dev readily available. Some can
> easily be handled by passing it as a param but there might be cases that
> are given pci_bus as parameter that could be somewhat trickier if there is
> no resource nor pci_dev (in case of a root bus), but I'm not immediately
> sure if there are actually any cases for real that fall into the latter
> category.
>
Yeah this is getting complicated very quickly :-)
So I take it you prefer to address this as a single feature, rather than
starting out by being more permissive when preserving the existing resource
allocation?