mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>
To: Ard Biesheuvel <ardb@kernel.org>,
	 Lorenzo Pieralisi <lpieralisi@kernel.org>
Cc: Ard Biesheuvel <ardb+git@google.com>,
	linux-pci@vger.kernel.org,  LKML <linux-kernel@vger.kernel.org>,
	Bjorn Helgaas <bhelgaas@google.com>
Subject: Re: [RFC PATCH] PCI: Tolerate non-prefetchable 64-bit BARs in prefetchable windows
Date: Mon, 14 Sep 2026 14:38:11 +0300 (EEST)	[thread overview]
Message-ID: <d2a561a9-e175-e102-3fa2-0be02342774a@linux.intel.com> (raw)
In-Reply-To: <48bb0fb6-10dc-46a7-84f6-3464d7c7bdfe@app.fastmail.com>

[-- Attachment #1: Type: text/plain, Size: 9058 bytes --]

On Fri, 11 Sep 2026, Ard Biesheuvel wrote:
> On Fri, 11 Sep 2026, at 11:27, Ilpo Järvinen wrote:
> > On Thu, 10 Sep 2026, Ard Biesheuvel wrote:
> >> On Thu, 10 Sep 2026, at 19:31, Ilpo Järvinen wrote:
> >> > On Thu, 10 Sep 2026, Ard Biesheuvel wrote:
> >> >
> >> >> From: Ard Biesheuvel <ardb@kernel.org>
> >> >> 
> >> >> The prefetchable vs. non-prefetchable distinction is a relic of
> >> >> conventional PCI, to denote from which regions PCI-PCI bridges were
> >> >> permitted to perform speculative readahead.
> >> >> 
> >> >> For software compatibility reasons, PCI Express inherited the Type 1
> >> >> header and models PCIe root ports as PCI-PCI bridges. However, this
> >> >> readahead behavior does not exist in PCIe, and so this distinction has
> >> >> mostly become meaningless on the bridge level.
> >> >> 
> >> >> As per the PCIe r6.3 ECN "Removing Prefetchable Terminology", the
> >> >> 'prefetchable' designation has been removed from the specification
> >> >> entirely, on the basis that it is obsolete, and is being abused to
> >> >> inform memory mapping attributes and other device/BAR level properties
> >> >> that it was never intended for.
> >> >> 
> >> >> Given the limited range for non-prefetchable windows in the Type 1
> >> >> header, and the fact that the distinction no longer exists for PCIe,
> >> >> resource allocation performed by firmware may result in non-prefetchable
> >> >> 64-bit BARs being allocated inside prefetchable bridge windows.
> >> >> 
> >> >> Linux rejects such allocations ("can't claim; no compatible bridge
> >> >> window") when it encounters them, but will usually fail to produce an
> >> >> alternative allocation, given that firmware wouldn't have placed them
> >> >> there in the first place if there was sufficient space in the
> >> >> non-prefetchable window.
> >> >> 
> >> >> So at the very least, let's not reject such allocations when they were
> >> >> made by the firmware.
> >> >> 
> >> >> Cc: Bjorn Helgaas <bhelgaas@google.com>
> >> >> Cc: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>
> >> >> Signed-off-by: Ard Biesheuvel <ardb@kernel.org>
> >> >> ---
> >> >> Link: https://github.com/tianocore/edk2/issues/13104
> >> >> 
> >> >>  drivers/pci/pci.c   | 3 ++-
> >> >>  include/linux/pci.h | 4 ++--
> >> >>  2 files changed, 4 insertions(+), 3 deletions(-)
> >> >> 
> >> >> diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
> >> >> index b2879a6be5f8..e33eb9f3a139 100644
> >> >> --- a/drivers/pci/pci.c
> >> >> +++ b/drivers/pci/pci.c
> >> >> @@ -761,7 +761,8 @@ struct resource *pci_find_parent_resource(const struct pci_dev *dev,
> >> >>  			 * not, the allocator made a mistake.
> >> >>  			 */
> >> >>  			if (r->flags & IORESOURCE_PREFETCH &&
> >> >> -			    !(res->flags & IORESOURCE_PREFETCH))
> >> >> +			    !(res->flags & IORESOURCE_PREFETCH) &&
> >> >> +			    !pci_is_pcie(dev))
> >> >>  				return NULL;
> >> >>  
> >> >>  			/*
> >> >
> >> > I was more thinking along the lines of always setting IORESOURCE_PREFETCH 
> >> > for 64-bit BARs on PCIe devices but I've not had time to look at that/test 
> >> > how many things would break as a result. ...It would seem much simpler 
> >> > solution to differentiate PCI from PCIe while keeping the existing logic 
> >> > without adding similar pci_is_pcie() checks everywhere.
> >> >
> >> 
> >> I agree that the code changes would be much simpler.
> >> 
> >> However, would this impact the PCI metadata observed by all consumers,
> >> including userspace,
> >
> > Yes, it will impact userspace. But quoting you from above:
> >
> > "is being abused to inform memory mapping attributes and other device/BAR 
> > level properties that it was never intended for."
> >
> > What does the userspace then do with the information? Does it qualify 
> > under "it was never intended for"?
> >
> 
> Perhaps, but that does not mean we are allowed to break it now.
> 
> More specifically, would the output of 'lspci' change as a result?

I'm not sure and I'm under impression there are multiple ways lspci can 
derive its information.

> >> and drivers that may expect a certain BAR layout,
> >
> > ??? Would that even be spec compliant??
> 
> Again, maybe not, but if something that works fine today stops
> working because the PCI subsystem started lying to the driver about
> how the PCIe device describes itself, we'll be on the hook to fix it.

Understood.

So I suppose you're not wanting to do this for the actual placement 
algorithm then because of the same risk but only cover the case where FW 
placed 64-bit non-pref BAR into prefetchable window like this patch 
currently does?

...And limiting to that case only likely implies remove + rescan cycle may 
fail because resources can no longer be placed into the same windows 
which will surprise user (arguably, not the most common use case but 
definitely surprising for the user if the kernel cannot place the 
resources the same way they were after boot => another way to get problem 
reports).

(Unrelated to this change, I'm going to open that BAR placement can of 
worms myself in a week or two because of resource placement changes I'm 
preparing. And I already hit one such problem where a particular BAR 
address results in probe failure I'll probably have to quirk around. And 
I didn't even have to move things into another window to trigger that.)

> >> and/or base decisions about memory attributes on this?
> >
> > The point is to consider them 64-bit window eligible so yes, kernel would 
> > definitely be basing decision on that but that's intentional.
> 
> That would mean that a driver may decide to use ioremap_wc() rather than
> ioremap() to map a non-prefetchable BAR that we decided to misrepresent
> as a prefetchable one. Even if the (pseudo-)PCI-PCI bridge will not do
> any readahead, the CPU or interconnect may behave very differently as a
> result, and touch BAR regions that the driver never accessed explicitly.
> 
> >> In particular, I am concerned about non-prefetchable BARs that actually
> >> have side effects on read, being mapped with WC (or Normal-NC on arm64)
> >> semantics, where the interconnect may widen, combine or reorder accesses.
> >
> > So on a more concrete terms, you're referring to the check in 
> > proc_bus_pci_mmap()? And the one in __pci_resource_attr_is_visible() + 
> > pci_dev_resource_wc_is_visible()? I suppose that wouldn't work then.
> 
> No, I am referring to the hundreds of ioremap() and ioremap_wc() calls
> under drivers. Maybe none of them are affected but who knows.

I'm left to wonder how many of those are based on a IORESOURCE_PREFETCH 
check, I strongly suspect none. Somehow I feel I'm the only one running 
git grep and I fail to locate any examples of the problem you mention. No 
offense meant, I just feel we're dicussion code that is hypothetical and 
doesn't exist for real, and definitely not on large scale.

That being said, I don't have problem in accepting adding 
IORESOURCE_PREFETCH to all PCIe resources was not a good idea so there's 
no point in continuing discussion towards this direction.

> > So if just setting IORESOURCE_PREFETCH is not workable, how about adding a 
> > getter for res->flags which adds IORESOURCE_PREFETCH into the returned 
> > flags if it's PCIe device and (in the end) use the raw value only in those 
> > places that actually care about wc distinction. What I don't want to see 
> > us adding that pci_is_pcie() everywhere.
> 
> Maybe add another IORESOURCE_PREFETCH_xxx flag that indicates that the
> resource may be placed in a prefetchable bridge window? We'd only have
> to set it in a single place (when probing the BAR), and we can add
> support for it piecemeal in the validation and allocation logic.

Unfortunately, I've earlier discovered there's no "single place" unless 
we add some gross res->flags fixup hack into PCI core. See e.g., the 
commit bdb32359eab9 ("sparc/PCI: Correct 64-bit non-pref -> pref BAR 
resources"), which we could hopefully revert after your change! So I'm 
afraid if you go to the new flag approach, besides drivers/pci/ you'd have 
to hunt down these from under arch/, and likely miss a few in the process.

Also, res->flags is currently full (for 32-bit) so you'd need to make the 
field (and therefore struct resource) larger.

So I still suggest having a flags getter for this purpose would be better.
In the complete solution covering also resource fitting and assignment 
algorithm, the main complication with that approach comes from the few 
cases that don't have the struct pci_dev readily available. Some can 
easily be handled by passing it as a param but there might be cases that 
are given pci_bus as parameter that could be somewhat trickier if there is 
no resource nor pci_dev (in case of a root bus), but I'm not immediately 
sure if there are actually any cases for real that fall into the latter 
category.


-- 
 i.

  reply	other threads:[~2026-09-14 11:38 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 14:34 Ard Biesheuvel
2026-09-10 17:31 ` Ilpo Järvinen
2026-09-10 21:04   ` Ard Biesheuvel
2026-09-11  9:27     ` Ilpo Järvinen
2026-09-11 10:01       ` Ard Biesheuvel
2026-09-14 11:38         ` Ilpo Järvinen [this message]
2026-09-14 12:50           ` Ard Biesheuvel
2026-09-11 10:14       ` Lorenzo Pieralisi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d2a561a9-e175-e102-3fa2-0be02342774a@linux.intel.com \
    --to=ilpo.jarvinen@linux.intel.com \
    --cc=ardb+git@google.com \
    --cc=ardb@kernel.org \
    --cc=bhelgaas@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=lpieralisi@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®