On Mon, 5 Oct 2026, Bjorn Helgaas wrote: > [+cc Boris] > > On Fri, Oct 02, 2026 at 02:33:15PM +0300, Ilpo Järvinen wrote: > > While testing the resource placement changes, tests hit a case where > > igb fails to probe when BAR 0 is placed at 0x9c000000: > > > > 90000000-9cffffff : PCI Bus 0000:a0 > > - 90000000-902fffff : PCI Bus 0000:a1 > > - 90000000-900fffff : 0000:a1:00.0 > > - 90000000-900fffff : igb > > - 90100000-901fffff : 0000:a1:00.0 > > - 90200000-90203fff : 0000:a1:00.0 > > - 90200000-90203fff : igb > > + 9be00000-9c0fffff : PCI Bus 0000:a1 > > + 9be00000-9befffff : 0000:a1:00.0 > > + 9bf00000-9bf03fff : 0000:a1:00.0 > > + 9c000000-9c0fffff : 0000:a1:00.0 > > 9c100000-9c17ffff : amd_iommu > > 9c180000-9c1803ff : IOAPIC 8 > > > > - Region 0: Memory at 90000000 (32-bit, non-prefetchable) [size=1M] > > - Region 3: Memory at 90200000 (32-bit, non-prefetchable) [size=16K] > > - Expansion ROM at 90100000 [disabled] [size=1M] > > + Region 0: Memory at 9c000000 (32-bit, non-prefetchable) [size=1M] > > + Region 3: Memory at 9bf00000 (32-bit, non-prefetchable) [size=16K] > > + Expansion ROM at 9be00000 [disabled] [size=1M] > > > > igb 0000:a1:00.0 0000:a1:00.0 (uninitialized): PCIe link lost > > ------------[ cut here ]------------ > > igb: Failed to read reg 0x18! > > WARNING: drivers/net/ethernet/intel/igb/igb_main.c:724 at igb_rd32.cold+0x3c/0x4f [igb], CPU#32: kworker/32:1/706 > > ... > > igb_get_invariants_82575+0xff/0xf00 [igb] > > igb_probe+0x3c8/0x1190 [igb] > > local_pci_probe+0x3b/0x80 > > > > It turns out there is a 64kB iomem black hole at 9c000000 that returns > > ~0 and this is where igb's BAR 0 resides. If another BAR of the same > > card is placed into that address, it is similarly black holed. > > > > Mark the problematic 64kB range reserved using a quirk bound to the > > bridge found in the problematic system. > > I think this is my fault, or at least it looks like it's related to > 07eab0901ede ("efi/x86: Remove EfiMemoryMappedIO from E820 map"). > > This platform describes the [0x9c000000-0x9cffffff] range as > E820_TYPE_RESERVED in the E820 table and as EFI_MEMORY_MAPPED_IO in > the EFI memory map (from the dmesg at > https://bugzilla.kernel.org/show_bug.cgi?id=222074): > > BIOS-e820: [gap 0x0000000090000000-0x000000009bffffff] > BIOS-e820: [mem 0x000000009c000000-0x000000009cffffff] device reserved > efi: Remove mem49: MMIO range=[0x9c000000-0x9cffffff] (16MB) from e820 map > e820: remove [mem 0x9c000000-0x9cffffff] device reserved > pci_bus 0000:a0: root bus resource [mem 0x90000000-0x9cffffff window] > > /proc/iomem: > > 90000000-9cffffff : PCI Bus 0000:a0 > 9c100000-9c17ffff : amd_iommu > 9c180000-9c1803ff : IOAPIC 8 > > I haven't worked out all the details, but I bet that if 07eab0901ede > had not removed [0x9c000000-0x9cffffff] from the E820 table, it would > show up in /proc/iomem as "Reserved" and would not be available for > use by a BAR. amd_iommu and IOAPIC 8 occupy some of that space, and I > suspect there are other devices in there that we don't know about. > > It looks like the last 16MB of every 32-bit PCI host bridge window is > EFI_MEMORY_MAPPED_IO, so if you can move the igb device to a different > host bridge, I suspect the same problem would happen if you put the > BAR at 16MB below the end. Perhaps things could break but even in this range I've found addresses above 9c000000 that do work. And the problem at 9c000000 is a black hole (returning ~0), not something that seems a meaningful device. With 07eab0901ede ("efi/x86: Remove EfiMemoryMappedIO from E820 map") reverted, I get this: 90000000-9cffffff : PCI Bus 0000:a0 9bd00000-9bffffff : PCI Bus 0000:a1 9bd00000-9bdfffff : 0000:a1:00.0 9bd00000-9bdfffff : igb 9befc000-9befffff : 0000:a1:00.0 9befc000-9befffff : igb 9bf00000-9bffffff : 0000:a1:00.0 9c000000-9cffffff : Reserved 9c100000-9c17ffff : amd_iommu 9c180000-9c1803ff : IOAPIC 8 -- i. > > Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222074 > > Suggested-by: Mario Limonciello > > Signed-off-by: Ilpo Järvinen > > --- > > > > Mario suggested the quirk to be based on the bridge instead of the > > endpoint device which certainly looks better and cleaner than the > > approach used in v1. > > > > The current plan is to try a different card in the same slot but it is > > a bit hard for me to predictable on what timescale that can be done. > > > > --- > > drivers/pci/quirks.c | 35 +++++++++++++++++++++++++++++++++++ > > 1 file changed, 35 insertions(+) > > > > diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c > > index de9bbccda21f..5483b47d8d54 100644 > > --- a/drivers/pci/quirks.c > > +++ b/drivers/pci/quirks.c > > @@ -6288,6 +6288,41 @@ DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1536, rom_bar_overlap_defect); > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1537, rom_bar_overlap_defect); > > DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x1538, rom_bar_overlap_defect); > > > > +/* > > + * TODO: Remove when/if root cause is found. > > + * > > + * Genoa appears to have a 64kB iomem black hole starting at 0x9c000000 > > + * address for which all reads return ~0 blocking an overlapping BAR from > > + * working. > > + * > > + * Work around the problem by reserving the space prior to making any iomem > > + * allocations that could overlap with the black hole. > > + */ > > +static struct resource black_hole_res = DEFINE_RES_MEM_NAMED(0x9c000000, SZ_64K, > > + "reserved"); > > + > > +static void genoa_iomem_black_hole(struct pci_dev *dev) > > +{ > > + struct resource *r; > > + > > + pci_bus_for_each_resource(dev->bus, r) { > > + if (!r || !r->flags || !resource_assigned(r)) > > + continue; > > + > > + if (!__resource_contains_unbound(r, &black_hole_res)) > > + continue; > > + > > + if (resource_assigned(&black_hole_res)) > > + pci_dbg(dev, "iomem black hole workaround already applied\n"); > > + else if (!insert_resource(r, &black_hole_res)) > > + pci_info(dev, "iomem black hole workaround enabled\n"); > > + else > > + pci_dbg(dev, "iomem black hole workaround add failed\n"); > > + return; > > + } > > +} > > +DECLARE_PCI_FIXUP_HEADER(PCI_VENDOR_ID_AMD, 0x14ab, genoa_iomem_black_hole); > > + > > #ifdef CONFIG_PCIEASPM > > /* > > * Several Intel DG2 graphics devices advertise that they can only tolerate > > -- > > 2.47.3 > > > -- i.