From: "Bowman, Terry" <terry.bowman@amd.com>
To: "Cheatham, Benjamin" <benjamin.cheatham@amd.com>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Bjorn Helgaas <bhelgaas@google.com>,
Dan Williams <djbw@kernel.org>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
linux-cxl@vger.kernel.org
Cc: Tony Luck <tony.luck@intel.com>, Borislav Petkov <bp@alien8.de>,
Hanjun Guo <guohanjun@huawei.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Richard Cheng <icheng@nvidia.com>,
Robert Richter <rrichter@amd.com>, Lukas Wunner <lukas@wunner.de>,
linux-pci@vger.kernel.org, linux-acpi@vger.kernel.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow
Date: Thu, 10 Sep 2026 11:55:26 -0500 [thread overview]
Message-ID: <90b550f6-bca9-4bb5-b38d-b4577b03ba1d@amd.com> (raw)
In-Reply-To: <a8ea024b-a608-4eea-bbdc-cba37b2bf2af@amd.com>
On 9/2/2026 3:57 PM, Cheatham, Benjamin wrote:
> On 9/2/2026 8:39 AM, Terry Bowman wrote:
>> Establish a single CXL protocol error path shared by CXL Virtual
>> Hierarchy (VH) and Restricted CXL Host (RCH) topologies. AER dispatch in
>> handle_error_source() routes CXL protocol errors, gated by
>> is_cxl_error(), through the AER-CXL kfifo to a cxl_core consumer for
>> logging and recovery. Producer and consumer go live together so no CXL
>> error is silently dropped across a bisect.
>>
>> is_cxl_error() expands from Endpoint-only to also cover Root Port,
>> Upstream Port, and Downstream Port. RCDs report on behalf of an upstream
>> RCH Downstream Port and instead reach the kfifo via
>> cxl_rch_handle_error().
>>
>> For uncorrectable errors, cxl_proto_err_wait_for_empty() drains the CXL
>> plane (RAS read, panic policy, state clear) before pci_aer_handle_error()
>> drives PCIe recovery, so recovery does not tear down RAS iomaps while the
>> consumer is still reading them. Correctable errors run asynchronously.
>>
>> Panic policy: cxl_do_recovery() panics on a confirmed UCE, and also when
>> the RAS registers cannot be mapped -- an unconfirmable UCE is treated
>> conservatively as fatal since CXL.mem coherency may be lost. A
>> mapped-but-clear status is logged as spurious with no panic.
>>
>> to_ras_base() centralizes RAS base lookup (dport->regs.ras for
>> Root/Downstream Ports, port->regs.ras otherwise) and provides an
>> injection point for RAS status simulation during testing. The
>> cxl_cor_error_detected() AER callback is removed; correctable Endpoint
>> errors now route through the kfifo like every other CXL protocol error.
>>
>> Update cxl_handle_rdport_errors() with locking to prevent dport from
>> being freed and RAS from being unmapped.
>>
>> At this step cxl_handle_rdport_errors() still dispatches a single
>> severity per pass (matching the pre-series baseline). The following
>> patch, "cxl/ras: Handle RCH correctable and uncorrectable errors in one
>> pass", processes a simultaneously signalled CE and UCE together.
>>
>> Co-developed-by: Dan Williams <djbw@kernel.org>
>> Signed-off-by: Dan Williams <djbw@kernel.org>
>> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>>
>> ---
>>
>
> One small nit, but otherwise LGTM:
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
>
> ...
>
>> pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
>> pci_channel_state_t state)
>> {
>> - struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
>> - struct cxl_memdev *cxlmd = cxlds->cxlmd;
>> - struct device *dev = &cxlmd->dev;
>> - bool ue;
>> + struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
>> + bool ue = false;
>> +
>> + if (!port)
>> + return PCI_ERS_RESULT_DISCONNECT;
>> +
>> + if (is_cxl_restricted(pdev))
>> + cxl_handle_rdport_errors(pdev);
>>
>> - scoped_guard(device, dev) {
>> - if (!dev->driver) {
>> + scoped_guard(device, &port->dev) {
>> + if (!port->dev.driver) {
>> dev_warn(&pdev->dev,
>> - "%s: memdev disabled, abort error handling\n",
>> - dev_name(dev));
>> + "%s: port disabled, abort error handling\n",
>> + dev_name(&port->dev));
>> return PCI_ERS_RESULT_DISCONNECT;
>> }
>>
>> - if (cxlds->rcd)
>> - cxl_handle_rdport_errors(cxlds);
>> /*
>> - * A frozen channel indicates an impending reset which is fatal to
>> - * CXL.mem operation, and will likely crash the system. On the off
>> - * chance the situation is recoverable dump the status of the RAS
>> - * capability registers and bounce the active state of the memdev.
>> + * The CXL RAS read is unconditional regardless of channel
>> + * state. Any uncorrectable error bit set in the CXL RAS
>> + * status register triggers a panic below because CXL.mem
>> + * cache coherency is already lost; continuing risks silent
>> + * data corruption.
>> */
>> - ue = cxl_handle_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
>> + ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
>> }
>>
>> + /*
>> + * CXL.mem UCE means cache coherency is lost. Continuing risks
>> + * silent data corruption.
>> + */
>
> Don't need this comment and the last sentence in the comment above.
I'll keep the panic-site comment (it's the "why" at the point of consequence) and will
remove the duplicated coherency/corruption sentence from the block comment above (done
in 2/9) so there's no duplication at any point in the series.
-Terry
>> + if (ue)
>> + panic("CXL cachemem error");
>> +
>> switch (state) {
>> case pci_channel_io_normal:
>> - if (ue) {
>> - device_release_driver(dev);
>> - return PCI_ERS_RESULT_NEED_RESET;
>> - }
>> return PCI_ERS_RESULT_CAN_RECOVER;
>> case pci_channel_io_frozen:
>> dev_warn(&pdev->dev,
>> "%s: frozen state error detected, disable CXL.mem\n",
>> - dev_name(dev));
>> - device_release_driver(dev);
>> + dev_name(port->uport_dev));
>> + device_release_driver(port->uport_dev);
>> return PCI_ERS_RESULT_NEED_RESET;
>> case pci_channel_io_perm_failure:
>> dev_warn(&pdev->dev,
>> @@ -335,3 +371,82 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
>> return PCI_ERS_RESULT_NEED_RESET;
>> }
>> EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
next prev parent reply other threads:[~2026-09-10 16:55 UTC|newest]
Thread overview: 33+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-08 0:51 ` Jonathan Cameron
2026-09-09 15:38 ` Bowman, Terry
2026-09-09 22:02 ` Jonathan Cameron
2026-09-10 14:57 ` Bowman, Terry
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-10 16:55 ` Bowman, Terry [this message]
2026-09-08 0:57 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-08 1:06 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-09 14:42 ` Bowman, Terry
2026-09-09 15:21 ` Bowman, Terry
2026-09-08 17:47 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-09 15:16 ` Lukas Wunner
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-09-08 18:13 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-09-08 18:39 ` Jonathan Cameron
2026-09-10 15:19 ` Bowman, Terry
2026-09-09 16:03 ` [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Lukas Wunner
2026-09-09 20:31 ` Bowman, Terry
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=90b550f6-bca9-4bb5-b38d-b4577b03ba1d@amd.com \
--to=terry.bowman@amd.com \
--cc=alison.schofield@intel.com \
--cc=benjamin.cheatham@amd.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=djbw@kernel.org \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=lukas@wunner.de \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®