mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Srirangan Madhavan <smadhavan@nvidia.com>
To: Jonathan Cameron <jic23@kernel.org>
Cc: Alison Schofield <alison.schofield@intel.com>,
	Bjorn Helgaas <bhelgaas@google.com>,
	Dave Jiang <dave.jiang@intel.com>,
	Davidlohr Bueso <dave@stgolabs.net>,
	Ira Weiny <ira.weiny@intel.com>,
	Vishal Verma <vishal.l.verma@intel.com>,
	linux-cxl@vger.kernel.org, linux-pci@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	Alex Williamson <alex.williamson@redhat.com>,
	vsethi@nvidia.com, alwilliamson@nvidia.com,
	Sai Yashwanth Reddy Kancherla <skancherla@nvidia.com>,
	Vishal Aslot <vaslot@nvidia.com>,
	Manish Honap <mhonap@nvidia.com>, Jiandi An <jan@nvidia.com>,
	Richard Cheng <icheng@nvidia.com>,
	linux-tegra@vger.kernel.org
Subject: Re: [PATCH v12 07/12] cxl: Validate HDM ranges before CXL reset
Date: Tue, 22 Sep 2026 17:08:21 -0700	[thread overview]
Message-ID: <354b7192-401e-43b3-84eb-8488a0fd89b1@nvidia.com> (raw)
In-Reply-To: <20260912023344.4cdeaf7a@jic23-hlaptop>

On 9/11/26 6:33 PM, Jonathan Cameron wrote:
> External email: Use caution opening links or attachments
> 
> 
> On Thu, 10 Sep 2026 07:08:03 +0000
> Srirangan Madhavan <smadhavan@nvidia.com> wrote:
> 
>> Before reset, require cached HDM decoder state, collect enabled decoder
>> ranges, and reserve them with request_mem_region(). This rejects reset
>> while affected CXL memory is busy and keeps the validation stable
>> through reset.
>>
>> If CPU cache invalidation support is available, invalidate the affected
>> ranges before reset.
> 
> If it is not what happens? I'm thinking dead system as coherency just broke
> but maybe not. My gut is no reset without cache invalidation support.
> 
> In general document your cache invalidation flow in this commit message. I'd
> suspicious you don't have a pair of invalidations, one before to ensure
> all caches are clean and one after to evict prefetched lines if your
> device was mean enough to serve garbage during the reset
> (which many will do!)
> 
> If more comes in later patches, but a breadcrumb in here to say so.
> 
>> If the runtime backend is unavailable, continue
>> after the range reservation succeeds.
>>
>> Reject CXL Reset when no cached HDM decoder state is available. The reset
>> path needs the cached address map to validate affected ranges and perform
>> CPU cache invalidation. Also reject normalized-addressing decoders for
>> now because the cached decoder range is not a system physical address.
>>
>> Signed-off-by: Srirangan Madhavan <smadhavan@nvidia.com>
> 
> 
> 
>> +static int cxl_hdm_range_flush_cache(struct cxl_hdm_range *range)
>> +{
>> +     struct pci_dev *pdev = range->pdev;
>> +     const struct range *hpa_range = &range->hpa_range;
>> +     int rc;
>> +
>> +     rc = cpu_cache_invalidate_memregion(hpa_range->start, range->len);
>> +     if (rc)
>> +             pci_err(pdev,
>> +                     "failed to invalidate CPU cache [%#llx-%#llx]: %d\n",
>> +                     hpa_range->start, hpa_range->end, rc);
> What does the wrapper really bring?
> 
>> +
>> +     return rc;
>> +}
>> +
>> +static int cxl_hdm_ranges_flush_cpu_caches(struct cxl_hdm_range_context *ctx,
>> +                                        struct pci_dev *pdev)
>> +{
>> +     struct cxl_hdm_range *range;
>> +     int rc;
>> +
>> +     if (list_empty(&ctx->ranges))
>> +             return 0;
>> +
>> +     if (!cpu_cache_has_invalidate_memregion()) {
>> +             pci_warn(pdev,
>> +                      "CPU cache synchronization unavailable; continuing without cache invalidation\n");
> 
> boom.  Unless you have very strong reasonsing for me this is a no.
> This introduces extremely subtle cache coherency corruption into your
> system and no one wants to debug results of that.
> 
>> +             return 0;
>> +     }
>> +
>> +     list_for_each_entry(range, &ctx->ranges, list) {
>> +             rc = cxl_hdm_range_flush_cache(range);
>> +             if (rc)
>> +                     return rc;
>> +     }
>> +
>> +     return 0;
>> +}
> 

These go into the v13 patch 11 now. I have:
   - reset is rejected when active HDM ranges exist and CPU cache
     invalidation is unavailable;
   - the affected ranges are invalidated before reset and again after 
state restoration;
   - range reservations and IOMMU exclusion remain active through the
     second invalidation; and
   - added description  for invalidation sequence in the commit message.

The previous wrapper around one invalidation call is now a range-context 
helper used by both invalidation passes.

https://lore.kernel.org/linux-cxl/20260922083924.2451158-12-smadhavan@nvidia.com/

-- 
Regards,
Srirangan

  reply	other threads:[~2026-09-23  0:08 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10  7:07 [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 01/12] cxl: Move HDM decoder programming helpers Srirangan Madhavan
2026-09-11 23:30   ` Jonathan Cameron
2026-09-22 23:28     ` Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 02/12] cxl: Make HDM commit helpers available to reset code Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 03/12] cxl: Share HDM decoder decode logic Srirangan Madhavan
2026-09-12  0:07   ` Jonathan Cameron
2026-09-22 23:52     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 04/12] cxl: Cache decoder settings on PCI devices Srirangan Madhavan
2026-09-12  0:22   ` Jonathan Cameron
2026-09-22 23:59     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 05/12] cxl: Cache endpoint decoder settings during PCI enumeration Srirangan Madhavan
2026-09-12  1:03   ` Jonathan Cameron
2026-09-23  0:02     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 06/12] cxl: Add CXL Device Reset helper Srirangan Madhavan
2026-09-12  1:26   ` Jonathan Cameron
2026-09-15 13:56     ` Lucero Palau, Alejandro
2026-09-23  0:51       ` Srirangan Madhavan
2026-09-23  0:05     ` Srirangan Madhavan
2026-09-23  0:20     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 07/12] cxl: Validate HDM ranges before CXL reset Srirangan Madhavan
2026-09-12  1:33   ` Jonathan Cameron
2026-09-23  0:08     ` Srirangan Madhavan [this message]
2026-09-10  7:08 ` [PATCH v12 08/12] PCI/CXL: Reject CXL Reset on multifunction devices Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 09/12] cxl: Restore CXL state after PCI reset Srirangan Madhavan
2026-09-12  1:43   ` Jonathan Cameron
2026-09-23  0:10     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 10/12] PCI/CXL: Expose CXL Reset as a PCI reset method Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 11/12] Documentation/ABI: Document CXL Reset " Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 12/12] PCI/CXL: Restore CXL state after CXL bus reset Srirangan Madhavan
2026-09-10  7:31 ` [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=354b7192-401e-43b3-84eb-8488a0fd89b1@nvidia.com \
    --to=smadhavan@nvidia.com \
    --cc=alex.williamson@redhat.com \
    --cc=alison.schofield@intel.com \
    --cc=alwilliamson@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=icheng@nvidia.com \
    --cc=ira.weiny@intel.com \
    --cc=jan@nvidia.com \
    --cc=jic23@kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-tegra@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=skancherla@nvidia.com \
    --cc=vaslot@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®