From: Qiang Yu <qiang.yu@oss.qualcomm.com>
To: Bjorn Helgaas <helgaas@kernel.org>
Cc: "Manivannan Sadhasivam" <mani@kernel.org>,
"Lorenzo Pieralisi" <lpieralisi@kernel.org>,
"Krzysztof Wilczyński" <kwilczynski@kernel.org>,
"Rob Herring" <robh@kernel.org>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Konrad Dybcio" <konrad.dybcio@oss.qualcomm.com>,
linux-pci@vger.kernel.org, linux-arm-msm@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2] PCI: qcom: Block accesses to downstream devices on link down
Date: Sun, 13 Sep 2026 23:27:21 -0700 [thread overview]
Message-ID: <aqeTyRnmsEcChW8G@hu-qianyu-lv.qualcomm.com> (raw)
In-Reply-To: <20260911171807.GA402207@bhelgaas>
On Fri, Sep 11, 2026 at 12:18:07PM -0500, Bjorn Helgaas wrote:
> On Fri, Sep 11, 2026 at 08:17:29AM +0200, Manivannan Sadhasivam wrote:
> > On Tue, Sep 08, 2026 at 06:00:30PM -0500, Bjorn Helgaas wrote:
> > > On Wed, Aug 19, 2026 at 11:36:54PM -0700, Qiang Yu wrote:
> > > > After a PCIe link goes down, software may still access the BAR
> > > > (MMIO) space or configuration space of devices behind that link
> > > > before recovery has run. As the link is down, these accesses
> > > > never complete, resulting in a storm of Completion Timeout AERs.
> > >
> > > What is special about qcom here? It seems like the Completion
> > > Timeouts and AER interrupts should happen with every PCIe
> > > controller.
> >
> > The special behavior which is common across many (not all) ARM SoCs
> > is that they don't synthesize all-one response for completion
> > timeouts, unlike RCs in x86 machines. Rather, they return AXI error
> > response, resulting in CPU treating them as SError, in-addition to
> > AER storm.
> >
> > Commit message missed mentioning SError though.
>
> I don't know how SError works, but this sounds like a pretty big open
> issue with respect to RAS. I don't think we want a kernel panic
> because a device failed to respond to a config or MMIO access, e.g.,
> if a card or Thunderbolt cable got unplugged.
>
> > > Is this mitigating an issue that will still happen on other
> > > controllers and should be solved elsewhere, e.g., by changing the
> > > software that accesses the BAR to look for the error responses it
> > > gets when the Completion Timeout happens?
> >
> > That would be too late as we don't get a proper error response.
> > That's why this patch is used.
> >
> > > The patch refers to the ECAM blocker (which I assume affects
> > > config accesses) and doesn't mention MMIO. Is the SLV_AXI stuff
> > > for MMIO?
> >
> > ECAM blocker is a Qcom's custom implementation which prevents the
> > config access to go out of the link and terminate the request
> > properly. And yes, it doesn't affect MMIO like BAR.
>
> The commit log implies that the patch blocks both config and MMIO
> accesses: "software may access MMIO or config space ... Use ECAM
> blocker to drop these accesses." If the ECAM blocker only affects
> config accesses, we need to reword that description.
>
This patch can block both config space and MMIO accesses.
> If this patch doesn't affect MMIO accesses, I assume they will still
> fail and lead to SError? I'm a little dubious about the utility of
> this because I assume most runtime driver accesses will be to MMIO
> space, not config space.
>
> I assume this is also racy: any config access that happens in the
> window between the link going down and the ECAM blocker being enabled
> will still fail and result in SError, right?
>
Yes, there is a window between link going down and the ECAM blocker being
enabled, this is why I mentioned this patch can *reduce* the completion
timeout errors in commit msg.
- Qiang Yu
> > > Assume the root port leads to a switch, and the switch has two
> > > links to downstream devices. If one of the switch downstream
> > > links goes down, is this going to kill the device on the other
> > > link that is still up?
> >
> > ECAM blocker is only enabled when the PCIe link to the Root Port
> > goes down. We don't get LDn interrupt when the switch downstream
> > link goes down. In that case, the switch should synthesize the error
> > response if the CPU tries to access the unavailable device.
>
> For config or MMIO reads targeting a device below a switch, if the
> link to the device is down, I think the switch will return an
> Unsupported Request completion without data (PCIe r7.0, sec 2.3.1.1).
>
> The PCIe spec doesn't prescribe whether RC synthesizes data to
> complete the read, but x86 RCs typically synthesize ~0 data. I guess
> some arm64 RCs assert SError instead of completing the read.
>
> IIUC, if the link from the RP goes down, we'll take the LDn interrupt,
> enable ECAM blocker, and future config reads would return ~0. But if
> a link below a switch goes down, there's no LDn interrupt and it
> sounds like future config reads would result in a UR Cpl and SError.
next prev parent reply other threads:[~2026-09-14 6:27 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-20 6:36 Qiang Yu
2026-08-20 11:29 ` Konrad Dybcio
2026-09-01 12:35 ` Manivannan Sadhasivam
2026-09-01 12:39 ` Konrad Dybcio
2026-09-01 14:06 ` Xilin Wu
2026-09-03 7:50 ` Krishna Chaitanya Chundru
2026-09-08 23:00 ` Bjorn Helgaas
2026-09-11 6:17 ` Manivannan Sadhasivam
2026-09-11 17:18 ` Bjorn Helgaas
2026-09-14 6:27 ` Qiang Yu [this message]
2026-09-30 16:51 ` Mayank Rana
2026-10-01 14:03 ` Qiang Yu
2026-09-15 17:16 ` Manivannan Sadhasivam
2026-09-15 23:14 ` Bjorn Helgaas
2026-09-16 8:03 ` Marc Zyngier
2026-09-16 16:36 ` Bjorn Helgaas
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqeTyRnmsEcChW8G@hu-qianyu-lv.qualcomm.com \
--to=qiang.yu@oss.qualcomm.com \
--cc=bhelgaas@google.com \
--cc=helgaas@kernel.org \
--cc=konrad.dybcio@oss.qualcomm.com \
--cc=kwilczynski@kernel.org \
--cc=linux-arm-msm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=mani@kernel.org \
--cc=robh@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®