mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Bowman, Terry" <terry.bowman@amd.com>
To: "Cheatham, Benjamin" <benjamin.cheatham@amd.com>,
	Jonathan Cameron <jic23@kernel.org>,
	Dave Jiang <dave.jiang@intel.com>,
	Alison Schofield <alison.schofield@intel.com>,
	Vishal Verma <vishal.l.verma@intel.com>,
	Davidlohr Bueso <dave@stgolabs.net>,
	Bjorn Helgaas <bhelgaas@google.com>,
	Dan Williams <djbw@kernel.org>,
	"Rafael J . Wysocki" <rafael@kernel.org>,
	Jonathan Corbet <corbet@lwn.net>,
	linux-cxl@vger.kernel.org
Cc: Tony Luck <tony.luck@intel.com>, Borislav Petkov <bp@alien8.de>,
	Hanjun Guo <guohanjun@huawei.com>,
	Mauro Carvalho Chehab <mchehab@kernel.org>,
	Shuai Xue <xueshuai@linux.alibaba.com>,
	Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
	Li Ming <ming.li@zohomail.com>,
	Shuah Khan <skhan@linuxfoundation.org>,
	Richard Cheng <icheng@nvidia.com>,
	Robert Richter <rrichter@amd.com>, Lukas Wunner <lukas@wunner.de>,
	linux-pci@vger.kernel.org, linux-acpi@vger.kernel.org,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow
Date: Thu, 10 Sep 2026 11:55:26 -0500	[thread overview]
Message-ID: <90b550f6-bca9-4bb5-b38d-b4577b03ba1d@amd.com> (raw)
In-Reply-To: <a8ea024b-a608-4eea-bbdc-cba37b2bf2af@amd.com>

On 9/2/2026 3:57 PM, Cheatham, Benjamin wrote:
> On 9/2/2026 8:39 AM, Terry Bowman wrote:
>> Establish a single CXL protocol error path shared by CXL Virtual
>> Hierarchy (VH) and Restricted CXL Host (RCH) topologies. AER dispatch in
>> handle_error_source() routes CXL protocol errors, gated by
>> is_cxl_error(), through the AER-CXL kfifo to a cxl_core consumer for
>> logging and recovery. Producer and consumer go live together so no CXL
>> error is silently dropped across a bisect.
>>
>> is_cxl_error() expands from Endpoint-only to also cover Root Port,
>> Upstream Port, and Downstream Port. RCDs report on behalf of an upstream
>> RCH Downstream Port and instead reach the kfifo via
>> cxl_rch_handle_error().
>>
>> For uncorrectable errors, cxl_proto_err_wait_for_empty() drains the CXL
>> plane (RAS read, panic policy, state clear) before pci_aer_handle_error()
>> drives PCIe recovery, so recovery does not tear down RAS iomaps while the
>> consumer is still reading them. Correctable errors run asynchronously.
>>
>> Panic policy: cxl_do_recovery() panics on a confirmed UCE, and also when
>> the RAS registers cannot be mapped -- an unconfirmable UCE is treated
>> conservatively as fatal since CXL.mem coherency may be lost. A
>> mapped-but-clear status is logged as spurious with no panic.
>>
>> to_ras_base() centralizes RAS base lookup (dport->regs.ras for
>> Root/Downstream Ports, port->regs.ras otherwise) and provides an
>> injection point for RAS status simulation during testing. The
>> cxl_cor_error_detected() AER callback is removed; correctable Endpoint
>> errors now route through the kfifo like every other CXL protocol error.
>>
>> Update cxl_handle_rdport_errors() with locking to prevent dport from
>> being freed and RAS from being unmapped.
>>
>> At this step cxl_handle_rdport_errors() still dispatches a single
>> severity per pass (matching the pre-series baseline). The following
>> patch, "cxl/ras: Handle RCH correctable and uncorrectable errors in one
>> pass", processes a simultaneously signalled CE and UCE together.
>>
>> Co-developed-by: Dan Williams <djbw@kernel.org>
>> Signed-off-by: Dan Williams <djbw@kernel.org>
>> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>>
>> ---
>>
> 
> One small nit, but otherwise LGTM:
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> 
> ...
> 
>>  pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
>>  				    pci_channel_state_t state)
>>  {
>> -	struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
>> -	struct cxl_memdev *cxlmd = cxlds->cxlmd;
>> -	struct device *dev = &cxlmd->dev;
>> -	bool ue;
>> +	struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
>> +	bool ue = false;
>> +
>> +	if (!port)
>> +		return PCI_ERS_RESULT_DISCONNECT;
>> +
>> +	if (is_cxl_restricted(pdev))
>> +		cxl_handle_rdport_errors(pdev);
>>  
>> -	scoped_guard(device, dev) {
>> -		if (!dev->driver) {
>> +	scoped_guard(device, &port->dev) {
>> +		if (!port->dev.driver) {
>>  			dev_warn(&pdev->dev,
>> -				 "%s: memdev disabled, abort error handling\n",
>> -				 dev_name(dev));
>> +				 "%s: port disabled, abort error handling\n",
>> +				 dev_name(&port->dev));
>>  			return PCI_ERS_RESULT_DISCONNECT;
>>  		}
>>  
>> -		if (cxlds->rcd)
>> -			cxl_handle_rdport_errors(cxlds);
>>  		/*
>> -		 * A frozen channel indicates an impending reset which is fatal to
>> -		 * CXL.mem operation, and will likely crash the system. On the off
>> -		 * chance the situation is recoverable dump the status of the RAS
>> -		 * capability registers and bounce the active state of the memdev.
>> +		 * The CXL RAS read is unconditional regardless of channel
>> +		 * state. Any uncorrectable error bit set in the CXL RAS
>> +		 * status register triggers a panic below because CXL.mem
>> +		 * cache coherency is already lost; continuing risks silent
>> +		 * data corruption.
>>  		 */
>> -		ue = cxl_handle_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
>> +		ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
>>  	}
>>  
>> +	/*
>> +	 * CXL.mem UCE means cache coherency is lost. Continuing risks
>> +	 * silent data corruption.
>> +	 */
> 
> Don't need this comment and the last sentence in the comment above.

I'll keep the panic-site comment (it's the "why" at the point of consequence) and will 
remove the duplicated coherency/corruption sentence from the block comment above (done 
in 2/9) so there's no duplication at any point in the series.

-Terry

>> +	if (ue)
>> +		panic("CXL cachemem error");
>> +
>>  	switch (state) {
>>  	case pci_channel_io_normal:
>> -		if (ue) {
>> -			device_release_driver(dev);
>> -			return PCI_ERS_RESULT_NEED_RESET;
>> -		}
>>  		return PCI_ERS_RESULT_CAN_RECOVER;
>>  	case pci_channel_io_frozen:
>>  		dev_warn(&pdev->dev,
>>  			 "%s: frozen state error detected, disable CXL.mem\n",
>> -			 dev_name(dev));
>> -		device_release_driver(dev);
>> +			 dev_name(port->uport_dev));
>> +		device_release_driver(port->uport_dev);
>>  		return PCI_ERS_RESULT_NEED_RESET;
>>  	case pci_channel_io_perm_failure:
>>  		dev_warn(&pdev->dev,
>> @@ -335,3 +371,82 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
>>  	return PCI_ERS_RESULT_NEED_RESET;
>>  }
>>  EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");


  reply	other threads:[~2026-09-10 16:55 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-08  0:51   ` Jonathan Cameron
2026-09-09 15:38     ` Bowman, Terry
2026-09-09 22:02       ` Jonathan Cameron
2026-09-10 14:57         ` Bowman, Terry
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-10 16:55     ` Bowman, Terry [this message]
2026-09-08  0:57   ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-08  1:06   ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-09 14:42     ` Bowman, Terry
2026-09-09 15:21     ` Bowman, Terry
2026-09-08 17:47   ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-09 15:16   ` Lukas Wunner
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-09-08 18:13   ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-09-02 20:57   ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-09-08 18:39   ` Jonathan Cameron
2026-09-10 15:19     ` Bowman, Terry
2026-09-09 16:03 ` [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Lukas Wunner
2026-09-09 20:31   ` Bowman, Terry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=90b550f6-bca9-4bb5-b38d-b4577b03ba1d@amd.com \
    --to=terry.bowman@amd.com \
    --cc=alison.schofield@intel.com \
    --cc=benjamin.cheatham@amd.com \
    --cc=bhelgaas@google.com \
    --cc=bp@alien8.de \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=djbw@kernel.org \
    --cc=guohanjun@huawei.com \
    --cc=icheng@nvidia.com \
    --cc=iweiny@kernel.org \
    --cc=jic23@kernel.org \
    --cc=lenb@kernel.org \
    --cc=linux-acpi@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=lukas@wunner.de \
    --cc=mchehab@kernel.org \
    --cc=ming.li@zohomail.com \
    --cc=rafael@kernel.org \
    --cc=rrichter@amd.com \
    --cc=skhan@linuxfoundation.org \
    --cc=tony.luck@intel.com \
    --cc=vishal.l.verma@intel.com \
    --cc=xueshuai@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®