Re: [PATCH 5/5] PCI/quirks: Avoid certain BAR 0 address with igb

From: David Laight

Date: Mon Sep 28 2026 - 10:09:28 EST


On Mon, 28 Sep 2026 15:20:13 +0300 (EEST)
Ilpo Järvinen <ilpo.jarvinen@xxxxxxxxxxxxxxx> wrote:

> On Thu, 24 Sep 2026, Bjorn Helgaas wrote:
>
> > [+cc Rafael, ACPI resource question]
> >
> > On Wed, Sep 23, 2026 at 04:17:55PM +0300, Ilpo Järvinen wrote:
> > > While testing the resource placement changes, my tests hit a case where
> > > igb fails to probe when BAR 0 is placed at 0x9c000000:
> > >
> > > 90000000-9cffffff : PCI Bus 0000:a0
> > > - 90000000-902fffff : PCI Bus 0000:a1
> > > - 90000000-900fffff : 0000:a1:00.0
> > > - 90000000-900fffff : igb
> > > - 90100000-901fffff : 0000:a1:00.0
> > > - 90200000-90203fff : 0000:a1:00.0
> > > - 90200000-90203fff : igb
> > > + 9be00000-9c0fffff : PCI Bus 0000:a1

Is that valid?
I'm no expect but I wouldn't expect an address range to cross a power of 2 boundary.

David

> > > + 9be00000-9befffff : 0000:a1:00.0
> > > + 9bf00000-9bf03fff : 0000:a1:00.0
> > > + 9c000000-9c0fffff : 0000:a1:00.0
> > > 9c100000-9c17ffff : amd_iommu
> > > 9c180000-9c1803ff : IOAPIC 8
> > >
> > > - Region 0: Memory at 90000000 (32-bit, non-prefetchable) [size=1M]
> > > - Region 3: Memory at 90200000 (32-bit, non-prefetchable) [size=16K]
> > > - Expansion ROM at 90100000 [disabled] [size=1M]
> > > + Region 0: Memory at 9c000000 (32-bit, non-prefetchable) [size=1M]
> > > + Region 3: Memory at 9bf00000 (32-bit, non-prefetchable) [size=16K]
> > > + Expansion ROM at 9be00000 [disabled] [size=1M]
> > >
> > > igb 0000:a1:00.0 0000:a1:00.0 (uninitialized): PCIe link lost
> > > ------------[ cut here ]------------
> > > igb: Failed to read reg 0x18!
> > > WARNING: drivers/net/ethernet/intel/igb/igb_main.c:724 at igb_rd32.cold+0x3c/0x4f [igb], CPU#32: kworker/32:1/706
> > > ...
> > > igb_get_invariants_82575+0xff/0xf00 [igb]
> > > igb_probe+0x3c8/0x1190 [igb]
> > > local_pci_probe+0x3b/0x80
> > >
> > > Apparently, the igb driver bails out, after its initial sanity check
> > > detects an unexpected ~0 read. Hacking around the sanity check just
> > > results in more failures down the road so the sanity check itself is not
> > > the cause for the failure.
> > >
> > > The resource placement looks valid so the actual placement patches seem
> > > to work normally.
> > >
> > > All other possible 1M address I could test (with a hack patch) did work.
> >
> > Super weird. Is it possible there's some other device there? It's
> > conceivable ACPI might have a _CRS method describing it. I think
> > there are ACPI devices for which we don't reserve space mentioned in
> > _CRS. Maybe Rafael knows a debug option to log everything in _CRS?
>
> Now that you mentioned it, there certainly something going on with
> that address:
>
> [ 0.000000] BIOS-e820: [mem 0x0000000070000000-0x000000008fffffff] device reserved
> [ 0.000000] BIOS-e820: [gap 0x0000000090000000-0x000000009bffffff]
> [ 0.000000] BIOS-e820: [mem 0x000000009c000000-0x000000009cffffff] device reserved
> [ 0.000000] BIOS-e820: [gap 0x000000009d000000-0x00000000a8ffffff]
> [ 0.000000] BIOS-e820: [mem 0x00000000a9000000-0x00000000a9ffffff] device reserved
> ...
> [ 0.000000] efi: Remove mem48: MMIO range=[0x80000000-0x8fffffff] (256MB) from e820 map
> [ 0.000000] e820: remove [mem 0x80000000-0x8fffffff] device reserved
> [ 0.000000] efi: Remove mem49: MMIO range=[0x9c000000-0x9cffffff] (16MB) from e820 map
> [ 0.000000] e820: remove [mem 0x9c000000-0x9cffffff] device reserved
> [ 0.000000] efi: Remove mem50: MMIO range=[0xa9000000-0xa9ffffff] (16MB) from e820 map
> [ 0.000000] e820: remove [mem 0xa9000000-0xa9ffffff] device reserved
>
> But given the comment above efi_remove_e820_mmio() it sounds like this is
> a red herring.
>
> In any case, the address is inside the provided root bus resource:
>
> [ 4.854840] ACPI: PCI Root Bridge [PC05] (domain 0000 [bus a0-bf])
> ...
> [ 4.854848] PCI host bridge to bus 0000:a0
> [ 4.854848] pci_bus 0000:a0: root bus resource [io 0x6000-0x6fff window]
> [ 4.854848] pci_bus 0000:a0: root bus resource [mem 0x90000000-0x9cffffff window]
> [ 4.854848] pci_bus 0000:a0: root bus resource [mem 0x71d60000000-0x7fcffffffff window]
> [ 4.854848] pci_bus 0000:a0: root bus resource [bus a0-bf]
>
> > Does igb seem sensitive about this exact address on a variety of
> > machines? If so I would expect some kind of igb hardware erratum for
> > it.
>
> I've not heard anything to that effect.
>
> But existance of the sanity check itself in the igb driver looks almost
> like a smoking gun so I don't know what to think of it.
>
> I'm also inclined to think that the placement approaches out there so far
> might not have covered that many addresses but used the left edge of the
> window. So this series, when it often moves resources to right edge of the
> window, goes to what might not be on well-charted territory.
>
> > Do other non-igb devices work at that address?
>
> Unfortunately there are not other devices underneath the RP so I might not
> be able to test this.
>
> > > Add quirk to reshuffle igb resources, use BAR 3 to block the problematic
> > > address.
> > >
> > > Signed-off-by: Ilpo Järvinen <ilpo.jarvinen@xxxxxxxxxxxxxxx>
> > > ---
> > >
> > > I know this is ugly and I don't like it either but do not know better
> > > way to avoid the regression.
> > >
> > > I've tried with iommu=off and that did not resolve the issue.
> > >
> > > I also managed to prove igb works with the same resource layout in
> > > another system. So identifying the case should probably be tightened
> > > by matching with more devices than the one used by igb. This is open
> > > to discussion.
> >
> > I guess this answers one of my questions above.
>
> Yeah.
>
> But then there's the sanity check in igb which hints otherwise.
>
> > > ---
> > > drivers/pci/quirks.c | 56 ++++++++++++++++++++++++++++++++++++++++++++
> > > 1 file changed, 56 insertions(+)
> > >
> > > diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c
> > > index de9bbccda21f..e6f3e2ab1fd4 100644
> > > --- a/drivers/pci/quirks.c
> > > +++ b/drivers/pci/quirks.c
> > > @@ -6288,6 +6288,62 @@ DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1536, rom_bar_overlap_defect);
> > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1537, rom_bar_overlap_defect);
> > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1538, rom_bar_overlap_defect);
> > >
> > > +/*
> > > + * The igb driver probe (due to reads returning ~0 unexpected) when BAR 0
> > > + * appears at 0x9c000000. The cause is unknown.
> >
> > I suppose this is missing "fails"? "igb driver probe fails"?
>
> Obviously, thanks.
>