On Thu, 24 Sep 2026, Bjorn Helgaas wrote: > [+cc Rafael, ACPI resource question] > > On Wed, Sep 23, 2026 at 04:17:55PM +0300, Ilpo Järvinen wrote: > > While testing the resource placement changes, my tests hit a case where > > igb fails to probe when BAR 0 is placed at 0x9c000000: > > > > 90000000-9cffffff : PCI Bus 0000:a0 > > - 90000000-902fffff : PCI Bus 0000:a1 > > - 90000000-900fffff : 0000:a1:00.0 > > - 90000000-900fffff : igb > > - 90100000-901fffff : 0000:a1:00.0 > > - 90200000-90203fff : 0000:a1:00.0 > > - 90200000-90203fff : igb > > + 9be00000-9c0fffff : PCI Bus 0000:a1 > > + 9be00000-9befffff : 0000:a1:00.0 > > + 9bf00000-9bf03fff : 0000:a1:00.0 > > + 9c000000-9c0fffff : 0000:a1:00.0 > > 9c100000-9c17ffff : amd_iommu > > 9c180000-9c1803ff : IOAPIC 8 > > > > - Region 0: Memory at 90000000 (32-bit, non-prefetchable) [size=1M] > > - Region 3: Memory at 90200000 (32-bit, non-prefetchable) [size=16K] > > - Expansion ROM at 90100000 [disabled] [size=1M] > > + Region 0: Memory at 9c000000 (32-bit, non-prefetchable) [size=1M] > > + Region 3: Memory at 9bf00000 (32-bit, non-prefetchable) [size=16K] > > + Expansion ROM at 9be00000 [disabled] [size=1M] > > > > igb 0000:a1:00.0 0000:a1:00.0 (uninitialized): PCIe link lost > > ------------[ cut here ]------------ > > igb: Failed to read reg 0x18! > > WARNING: drivers/net/ethernet/intel/igb/igb_main.c:724 at igb_rd32.cold+0x3c/0x4f [igb], CPU#32: kworker/32:1/706 > > ... > > igb_get_invariants_82575+0xff/0xf00 [igb] > > igb_probe+0x3c8/0x1190 [igb] > > local_pci_probe+0x3b/0x80 > > > > Apparently, the igb driver bails out, after its initial sanity check > > detects an unexpected ~0 read. Hacking around the sanity check just > > results in more failures down the road so the sanity check itself is not > > the cause for the failure. > > > > The resource placement looks valid so the actual placement patches seem > > to work normally. > > > > All other possible 1M address I could test (with a hack patch) did work. > > Super weird. Is it possible there's some other device there? It's > conceivable ACPI might have a _CRS method describing it. I think > there are ACPI devices for which we don't reserve space mentioned in > _CRS. Maybe Rafael knows a debug option to log everything in _CRS? Now that you mentioned it, there certainly something going on with that address: [ 0.000000] BIOS-e820: [mem 0x0000000070000000-0x000000008fffffff] device reserved [ 0.000000] BIOS-e820: [gap 0x0000000090000000-0x000000009bffffff] [ 0.000000] BIOS-e820: [mem 0x000000009c000000-0x000000009cffffff] device reserved [ 0.000000] BIOS-e820: [gap 0x000000009d000000-0x00000000a8ffffff] [ 0.000000] BIOS-e820: [mem 0x00000000a9000000-0x00000000a9ffffff] device reserved ... [ 0.000000] efi: Remove mem48: MMIO range=[0x80000000-0x8fffffff] (256MB) from e820 map [ 0.000000] e820: remove [mem 0x80000000-0x8fffffff] device reserved [ 0.000000] efi: Remove mem49: MMIO range=[0x9c000000-0x9cffffff] (16MB) from e820 map [ 0.000000] e820: remove [mem 0x9c000000-0x9cffffff] device reserved [ 0.000000] efi: Remove mem50: MMIO range=[0xa9000000-0xa9ffffff] (16MB) from e820 map [ 0.000000] e820: remove [mem 0xa9000000-0xa9ffffff] device reserved But given the comment above efi_remove_e820_mmio() it sounds like this is a red herring. In any case, the address is inside the provided root bus resource: [ 4.854840] ACPI: PCI Root Bridge [PC05] (domain 0000 [bus a0-bf]) ... [ 4.854848] PCI host bridge to bus 0000:a0 [ 4.854848] pci_bus 0000:a0: root bus resource [io 0x6000-0x6fff window] [ 4.854848] pci_bus 0000:a0: root bus resource [mem 0x90000000-0x9cffffff window] [ 4.854848] pci_bus 0000:a0: root bus resource [mem 0x71d60000000-0x7fcffffffff window] [ 4.854848] pci_bus 0000:a0: root bus resource [bus a0-bf] > Does igb seem sensitive about this exact address on a variety of > machines? If so I would expect some kind of igb hardware erratum for > it. I've not heard anything to that effect. But existance of the sanity check itself in the igb driver looks almost like a smoking gun so I don't know what to think of it. I'm also inclined to think that the placement approaches out there so far might not have covered that many addresses but used the left edge of the window. So this series, when it often moves resources to right edge of the window, goes to what might not be on well-charted territory. > Do other non-igb devices work at that address? Unfortunately there are not other devices underneath the RP so I might not be able to test this. > > Add quirk to reshuffle igb resources, use BAR 3 to block the problematic > > address. > > > > Signed-off-by: Ilpo Järvinen > > --- > > > > I know this is ugly and I don't like it either but do not know better > > way to avoid the regression. > > > > I've tried with iommu=off and that did not resolve the issue. > > > > I also managed to prove igb works with the same resource layout in > > another system. So identifying the case should probably be tightened > > by matching with more devices than the one used by igb. This is open > > to discussion. > > I guess this answers one of my questions above. Yeah. But then there's the sanity check in igb which hints otherwise. > > --- > > drivers/pci/quirks.c | 56 ++++++++++++++++++++++++++++++++++++++++++++ > > 1 file changed, 56 insertions(+) > > > > diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c > > index de9bbccda21f..e6f3e2ab1fd4 100644 > > --- a/drivers/pci/quirks.c > > +++ b/drivers/pci/quirks.c > > @@ -6288,6 +6288,62 @@ DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1536, rom_bar_overlap_defect); > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1537, rom_bar_overlap_defect); > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1538, rom_bar_overlap_defect); > > > > +/* > > + * The igb driver probe (due to reads returning ~0 unexpected) when BAR 0 > > + * appears at 0x9c000000. The cause is unknown. > > I suppose this is missing "fails"? "igb driver probe fails"? Obviously, thanks. -- i. > > + * Use BAR 3 to block 0x9c000000 address. > > + */ > > +static void bar0_address_breakage(struct pci_dev *dev) > > +{ > > + struct resource *bar0 = pci_resource_n(dev, 0); > > + struct resource *bar3 = pci_resource_n(dev, 3); > > + resource_size_t broken_addr = 0x9c000000; > > + struct resource *res; > > + int i, ret; > > + > > + if (bar0->start != broken_addr) > > + return; > > + > > + /* > > + * HW BAR sizes seems to vary. Exclude non-1M BAR 0 case and > > + * sanity check BAR 0 & 3 before attempting this quirk. > > + */ > > + if (resource_type(bar0) != IORESOURCE_MEM || > > + resource_size(bar0) != SZ_1M || > > + resource_type(bar3) != IORESOURCE_MEM) > > + return; > > + > > + pci_info(dev, "%pR: relocating BAR\n", bar0); > > + > > + pci_dev_for_each_resource(dev, res, i) { > > + if (!resource_assigned(res) || > > + resource_type(res) != IORESOURCE_MEM) > > + continue; > > + > > + pci_release_resource(dev, i); > > + } > > + > > + resource_set_range(bar3, broken_addr, resource_size(bar3)); > > + bar3->flags &= ~IORESOURCE_UNSET; > > + pci_claim_resource(dev, 3); > > + if (!resource_assigned(bar3)) { > > + bar3->flags |= IORESOURCE_UNSET; > > + pci_warn(dev, "resource relocation failed\n"); > > + } > > + > > + pci_dev_for_each_resource(dev, res, i) { > > + if (resource_assigned(res) || > > + resource_type(res) != IORESOURCE_MEM) > > + continue; > > + > > + ret = pci_assign_resource(dev, i); > > + if (ret) > > + pci_warn(dev, "resource relocation failed\n"); > > + } > > +} > > +DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_INTEL, 0x1533, bar0_address_breakage); > > + > > #ifdef CONFIG_PCIEASPM > > /* > > * Several Intel DG2 graphics devices advertise that they can only tolerate > > -- > > 2.47.3 > > >