From: Srirangan Madhavan <smadhavan@nvidia.com>
To: Jonathan Cameron <jic23@kernel.org>
Cc: Alison Schofield <alison.schofield@intel.com>,
Bjorn Helgaas <bhelgaas@google.com>,
Dave Jiang <dave.jiang@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Ira Weiny <ira.weiny@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
linux-cxl@vger.kernel.org, linux-pci@vger.kernel.org,
linux-kernel@vger.kernel.org,
Alex Williamson <alex.williamson@redhat.com>,
vsethi@nvidia.com, alwilliamson@nvidia.com,
Sai Yashwanth Reddy Kancherla <skancherla@nvidia.com>,
Vishal Aslot <vaslot@nvidia.com>,
Manish Honap <mhonap@nvidia.com>, Jiandi An <jan@nvidia.com>,
Richard Cheng <icheng@nvidia.com>,
linux-tegra@vger.kernel.org
Subject: Re: [PATCH v12 07/12] cxl: Validate HDM ranges before CXL reset
Date: Tue, 22 Sep 2026 17:08:21 -0700 [thread overview]
Message-ID: <354b7192-401e-43b3-84eb-8488a0fd89b1@nvidia.com> (raw)
In-Reply-To: <20260912023344.4cdeaf7a@jic23-hlaptop>
On 9/11/26 6:33 PM, Jonathan Cameron wrote:
> External email: Use caution opening links or attachments
>
>
> On Thu, 10 Sep 2026 07:08:03 +0000
> Srirangan Madhavan <smadhavan@nvidia.com> wrote:
>
>> Before reset, require cached HDM decoder state, collect enabled decoder
>> ranges, and reserve them with request_mem_region(). This rejects reset
>> while affected CXL memory is busy and keeps the validation stable
>> through reset.
>>
>> If CPU cache invalidation support is available, invalidate the affected
>> ranges before reset.
>
> If it is not what happens? I'm thinking dead system as coherency just broke
> but maybe not. My gut is no reset without cache invalidation support.
>
> In general document your cache invalidation flow in this commit message. I'd
> suspicious you don't have a pair of invalidations, one before to ensure
> all caches are clean and one after to evict prefetched lines if your
> device was mean enough to serve garbage during the reset
> (which many will do!)
>
> If more comes in later patches, but a breadcrumb in here to say so.
>
>> If the runtime backend is unavailable, continue
>> after the range reservation succeeds.
>>
>> Reject CXL Reset when no cached HDM decoder state is available. The reset
>> path needs the cached address map to validate affected ranges and perform
>> CPU cache invalidation. Also reject normalized-addressing decoders for
>> now because the cached decoder range is not a system physical address.
>>
>> Signed-off-by: Srirangan Madhavan <smadhavan@nvidia.com>
>
>
>
>> +static int cxl_hdm_range_flush_cache(struct cxl_hdm_range *range)
>> +{
>> + struct pci_dev *pdev = range->pdev;
>> + const struct range *hpa_range = &range->hpa_range;
>> + int rc;
>> +
>> + rc = cpu_cache_invalidate_memregion(hpa_range->start, range->len);
>> + if (rc)
>> + pci_err(pdev,
>> + "failed to invalidate CPU cache [%#llx-%#llx]: %d\n",
>> + hpa_range->start, hpa_range->end, rc);
> What does the wrapper really bring?
>
>> +
>> + return rc;
>> +}
>> +
>> +static int cxl_hdm_ranges_flush_cpu_caches(struct cxl_hdm_range_context *ctx,
>> + struct pci_dev *pdev)
>> +{
>> + struct cxl_hdm_range *range;
>> + int rc;
>> +
>> + if (list_empty(&ctx->ranges))
>> + return 0;
>> +
>> + if (!cpu_cache_has_invalidate_memregion()) {
>> + pci_warn(pdev,
>> + "CPU cache synchronization unavailable; continuing without cache invalidation\n");
>
> boom. Unless you have very strong reasonsing for me this is a no.
> This introduces extremely subtle cache coherency corruption into your
> system and no one wants to debug results of that.
>
>> + return 0;
>> + }
>> +
>> + list_for_each_entry(range, &ctx->ranges, list) {
>> + rc = cxl_hdm_range_flush_cache(range);
>> + if (rc)
>> + return rc;
>> + }
>> +
>> + return 0;
>> +}
>
These go into the v13 patch 11 now. I have:
- reset is rejected when active HDM ranges exist and CPU cache
invalidation is unavailable;
- the affected ranges are invalidated before reset and again after
state restoration;
- range reservations and IOMMU exclusion remain active through the
second invalidation; and
- added description for invalidation sequence in the commit message.
The previous wrapper around one invalidation call is now a range-context
helper used by both invalidation passes.
https://lore.kernel.org/linux-cxl/20260922083924.2451158-12-smadhavan@nvidia.com/
--
Regards,
Srirangan
next prev parent reply other threads:[~2026-09-23 0:08 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 7:07 [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
2026-09-10 7:07 ` [PATCH v12 01/12] cxl: Move HDM decoder programming helpers Srirangan Madhavan
2026-09-11 23:30 ` Jonathan Cameron
2026-09-22 23:28 ` Srirangan Madhavan
2026-09-10 7:07 ` [PATCH v12 02/12] cxl: Make HDM commit helpers available to reset code Srirangan Madhavan
2026-09-10 7:07 ` [PATCH v12 03/12] cxl: Share HDM decoder decode logic Srirangan Madhavan
2026-09-12 0:07 ` Jonathan Cameron
2026-09-22 23:52 ` Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 04/12] cxl: Cache decoder settings on PCI devices Srirangan Madhavan
2026-09-12 0:22 ` Jonathan Cameron
2026-09-22 23:59 ` Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 05/12] cxl: Cache endpoint decoder settings during PCI enumeration Srirangan Madhavan
2026-09-12 1:03 ` Jonathan Cameron
2026-09-23 0:02 ` Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 06/12] cxl: Add CXL Device Reset helper Srirangan Madhavan
2026-09-12 1:26 ` Jonathan Cameron
2026-09-15 13:56 ` Lucero Palau, Alejandro
2026-09-23 0:51 ` Srirangan Madhavan
2026-09-23 0:05 ` Srirangan Madhavan
2026-09-23 0:20 ` Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 07/12] cxl: Validate HDM ranges before CXL reset Srirangan Madhavan
2026-09-12 1:33 ` Jonathan Cameron
2026-09-23 0:08 ` Srirangan Madhavan [this message]
2026-09-10 7:08 ` [PATCH v12 08/12] PCI/CXL: Reject CXL Reset on multifunction devices Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 09/12] cxl: Restore CXL state after PCI reset Srirangan Madhavan
2026-09-12 1:43 ` Jonathan Cameron
2026-09-23 0:10 ` Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 10/12] PCI/CXL: Expose CXL Reset as a PCI reset method Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 11/12] Documentation/ABI: Document CXL Reset " Srirangan Madhavan
2026-09-10 7:08 ` [PATCH v12 12/12] PCI/CXL: Restore CXL state after CXL bus reset Srirangan Madhavan
2026-09-10 7:31 ` [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=354b7192-401e-43b3-84eb-8488a0fd89b1@nvidia.com \
--to=smadhavan@nvidia.com \
--cc=alex.williamson@redhat.com \
--cc=alison.schofield@intel.com \
--cc=alwilliamson@nvidia.com \
--cc=bhelgaas@google.com \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=icheng@nvidia.com \
--cc=ira.weiny@intel.com \
--cc=jan@nvidia.com \
--cc=jic23@kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linux-tegra@vger.kernel.org \
--cc=mhonap@nvidia.com \
--cc=skancherla@nvidia.com \
--cc=vaslot@nvidia.com \
--cc=vishal.l.verma@intel.com \
--cc=vsethi@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®