mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Srirangan Madhavan <smadhavan@nvidia.com>
To: Jonathan Cameron <jic23@kernel.org>
Cc: Alison Schofield <alison.schofield@intel.com>,
	Bjorn Helgaas <bhelgaas@google.com>,
	Dave Jiang <dave.jiang@intel.com>,
	Davidlohr Bueso <dave@stgolabs.net>,
	Ira Weiny <ira.weiny@intel.com>,
	Vishal Verma <vishal.l.verma@intel.com>,
	linux-cxl@vger.kernel.org, linux-pci@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	Alex Williamson <alex.williamson@redhat.com>,
	vsethi@nvidia.com, alwilliamson@nvidia.com,
	Sai Yashwanth Reddy Kancherla <skancherla@nvidia.com>,
	Vishal Aslot <vaslot@nvidia.com>,
	Manish Honap <mhonap@nvidia.com>, Jiandi An <jan@nvidia.com>,
	Richard Cheng <icheng@nvidia.com>,
	linux-tegra@vger.kernel.org
Subject: Re: [PATCH v12 05/12] cxl: Cache endpoint decoder settings during PCI enumeration
Date: Tue, 22 Sep 2026 17:02:26 -0700	[thread overview]
Message-ID: <23266146-f92c-4b97-bf1b-99e4973b7dfc@nvidia.com> (raw)
In-Reply-To: <20260912020305.64278c3e@jic23-hlaptop>

On 9/11/26 6:03 PM, Jonathan Cameron wrote:
> External email: Use caution opening links or attachments
> 
> 
> On Thu, 10 Sep 2026 07:08:01 +0000
> Srirangan Madhavan <smadhavan@nvidia.com> wrote:
> 
>> Populate pci_dev->hdm for CXL.mem functions from pci_bus_add_device(),
>> after final PCI fixups and state save but before driver binding. This
>> gives driver-free reset paths an early HDM snapshot while avoiding the
>> pre-resource-assignment window in PCI capability initialization.
>>
>> Use the CXL Register Locator BAR Indicator to find the component register
>> BAR, reject unassigned, disabled, or zero memory BAR resources before
>> temporarily enabling Memory Space, and restore the original PCI_COMMAND
>> value before returning. Cache the CXL Device DVSEC control register with
>> the HDM state for reset recovery before a driver can alter it.
>>
>> CXL core refreshes the cache as decoders are committed or reset, and keeps
>> the cached DVSEC control synchronized when CXL.mem is enabled or disabled.
>> Move the register helpers into the built-in CONFIG_CXL_RESET set so the
>> early cache path is available without cxl_core, and keep the cxl-test mock
>> core from building a duplicate regs.o.
>>
>> Signed-off-by: Srirangan Madhavan <smadhavan@nvidia.com>
> Hi Srirangan,
> 
> Unless I'm reading this wrong, this has evolved to the point that it
> needs a step back and a rethink.  There is complexity in here I
> don't think you need at all.  I may well be missing something
> though given it's Friday evening!
> 
> Jonathan
> 
> 
>> diff --git a/drivers/cxl/core/resource.c b/drivers/cxl/core/resource.c
>> index e6aa55079c76..6d9f8fe14b16 100644
>> --- a/drivers/cxl/core/resource.c
>> +++ b/drivers/cxl/core/resource.c
>> @@ -2,9 +2,17 @@
>>   /* Copyright (c) 2026 NVIDIA Corporation & Affiliates */
>>   #include <linux/delay.h>
>>   #include <linux/bug.h>
>> +#include <linux/bitfield.h>
>> +#include <linux/cleanup.h>
>>   #include <linux/errno.h>
>>   #include <linux/export.h>
>> +#include <linux/io.h>
>> +#include <linux/ioport.h>
>>   #include <linux/kernel.h>
>> +#include <linux/pci.h>
>> +#include <linux/slab.h>
>> +
>> +#include <cxl/pci.h>
>>
>>   #include "cxl.h"
>>   #include "core.h"
>> @@ -156,3 +164,331 @@ int cxl_hdm_decode_decoder(struct cxl_decoder_settings *settings, int id,
>>                                  &settings->interleave_granularity);
>>   }
>>   EXPORT_SYMBOL_FOR_MODULES(cxl_hdm_decode_decoder, "cxl_core");
>> +
>> +struct cxl_hdm_decoder_state {
>> +     u32 ctrl;
>> +     u32 base_low;
>> +     u32 base_high;
>> +     u32 size_low;
>> +     u32 size_high;
>> +     u32 target_low;
>> +     u32 target_high;
>> +};
>> +
>> +static void cxl_pci_hdm_info_free(struct cxl_hdm_info *info)
>> +{
>> +     if (!info)
>> +             return;
>> +
>> +     kfree(info->decoder_state);
>> +     kfree(info);
>> +}
>> +
>> +DEFINE_FREE(cxl_pci_hdm_info, struct cxl_hdm_info *,
>> +         cxl_pci_hdm_info_free(_T))
>> +
>> +void pci_cxl_hdm_release(struct pci_dev *pdev)
>> +{
>> +     struct cxl_hdm_info *info = pdev->hdm;
>> +
>> +     pdev->hdm = NULL;
> 
> If the order here matters, you need a barrier or WRITE_ONCE() might
> do it.
> 
>> +     cxl_pci_hdm_info_free(info);
>> +}
>> +
>> +static bool cxl_pci_bar_usable(struct pci_dev *pdev, int bar)
> 
> This doesn't seem to have anything CXL specific about it.  Maybe
> give it a more generic name and move it to pci.c?
> However, see later - I'm not sure you need this.
> 
>> +{
>> +     struct resource *res = &pdev->resource[bar];
>> +
>> +     if (!pci_resource_len(pdev, bar))
>> +             return false;
>> +     if (res->flags & (IORESOURCE_UNSET | IORESOURCE_DISABLED))
>> +             return false;
>> +     if (resource_type(res) != IORESOURCE_MEM)
>> +             return false;
>> +     if (!res->start || !res->end)
>> +             return false;
>> +
>> +     return true;
>> +}
>> +
>> +static int cxl_pci_hdm_find_bar(struct pci_dev *pdev, resource_size_t hdm_start,
>> +                             resource_size_t hdm_size, int *bar,
>> +                             resource_size_t *offset)
>> +{
>> +     resource_size_t hdm_end;
>> +
>> +     if (!hdm_size)
>> +             return -EINVAL;
>> +
>> +     hdm_end = hdm_start + hdm_size - 1;
>> +     if (hdm_end < hdm_start)
>> +             return -EINVAL;
>> +
>> +     for (int i = 0; i < PCI_STD_NUM_BARS; i++) {
> 
> This feels unduly painful and overly specific to this case.
> It is just looking for a resource to bar and offset.
> 
> It is not CXL specific so maybe ask Bjorn if such a helper might go
> in pci.c (assuming not discussed and dismissed in earlier rounds of
> review!)  I'd still like it to mention hdm or anything even if
> local to here.
> static int cxl_pci_find_resource_bar(struct pci_dev *pdev,
>                                       resource_size_t start, resource_size_t size,
>                                       int *bar, resource_size_t *offset)
> 
> Noted later, you can skip this entirely as this is going in a circle.
> 
>> +             struct resource *res = &pdev->resource[i];
>> +
>> +             if (!cxl_pci_bar_usable(pdev, i))
>> +                     continue;
>> +             if (hdm_start < res->start || hdm_end > res->end)
>> +                     continue;
>> +
>> +             if (bar)
>> +                     *bar = i;
>> +             if (offset)
>> +                     *offset = hdm_start - res->start;
>> +             return 0;
>> +     }
>> +
>> +     return -ENODEV;
>> +}
> 
>> +
>> +static void cxl_pci_hdm_read_decoder_state(struct cxl_hdm_decoder_state *state,
>> +                                        void __iomem *hdm, int id)
>> +{
>> +     state->ctrl = readl(hdm + CXL_HDM_DECODER0_CTRL_OFFSET(id));
>> +     state->base_low = readl(hdm + CXL_HDM_DECODER0_BASE_LOW_OFFSET(id));
>> +     state->base_high = readl(hdm + CXL_HDM_DECODER0_BASE_HIGH_OFFSET(id));
>> +     state->size_low = readl(hdm + CXL_HDM_DECODER0_SIZE_LOW_OFFSET(id));
>> +     state->size_high = readl(hdm + CXL_HDM_DECODER0_SIZE_HIGH_OFFSET(id));
>> +     state->target_low = readl(hdm + CXL_HDM_DECODER0_TL_LOW(id));
>> +     state->target_high = readl(hdm + CXL_HDM_DECODER0_TL_HIGH(id));
>> +}
>> +
>> +static int cxl_pci_hdm_read_decoder(struct pci_dev *pdev,
>> +                                 struct cxl_hdm_decoder_state *state,
>> +                                 struct cxl_decoder_settings *settings,
>> +                                 void __iomem *hdm, int id)
>> +{
>> +     u64 target_or_skip, base, size;
>> +     int rc;
>> +
>> +     cxl_pci_hdm_read_decoder_state(state, hdm, id);
>> +
>> +     base = ((u64)state->base_high << 32) | state->base_low;
>> +     size = ((u64)state->size_high << 32) | state->size_low;
>> +     target_or_skip = ((u64)state->target_high << 32) | state->target_low;
> 
> I'm not sure I get why we cache the registers and the stuff derived from them.
> Why isn't one source of info enough? Or do the have different lifetimes?
> If they do then add a comment to structure definition on that.
> 
>> +
>> +     rc = cxl_hdm_decode_decoder(settings, id, state->ctrl, base, size,
>> +                                 target_or_skip, NULL);
>> +     if (rc) {
>> +             pci_err(pdev, "CXL HDM decoder %d has invalid configuration: %d\n",
>> +                     id, rc);
>> +             return rc;
>> +     }
>> +     return 0;
>> +}
> 
>> +
>> +static int __cxl_pci_hdm_read_info(struct pci_dev *pdev,
>> +                                struct cxl_register_map *map,
>> +                                struct cxl_hdm_info *info)
>> +{
>> +     struct cxl_decoder_settings *settings;
>> +     int decoder_count;
>> +     int rc;
>> +
>> +     rc = cxl_setup_regs(map);
>> +     if (rc)
>> +             return rc;
>> +
>> +     if (!map->component_map.hdm_decoder.valid)
>> +             return -ENODEV;
>> +
>> +     void __iomem *hdm __free(cxl_hdm_iounmap) =
>> +             cxl_pci_hdm_map(pdev, map, info);
>> +     if (IS_ERR(hdm))
>> +             return PTR_ERR(no_free_ptr(hdm));
>> +
>> +     decoder_count = cxl_hdm_decoder_count(readl(hdm +
>> +                                                 CXL_HDM_DECODER_CAP_OFFSET));
> 
> Go long on lines like this.  As long as you stay only a bit over 80 no one will
> mind.
> 
>> +     if (decoder_count < 0)
>> +             return decoder_count;
>> +
>> +     if (decoder_count > ARRAY_SIZE(info->settings))
>> +             return -ENXIO;
>> +
>> +     if (CXL_HDM_DECODER0_CTRL_OFFSET(decoder_count - 1) + 0x10 >
> 
> That 0x10 needs to be a define or other useful code. I have no idea what
> it is...
> 
>> +         info->hdm_size) {
>> +             pci_err(pdev,
>> +                     "CXL HDM decoder count exceeds mapped register block\n");
>> +             return -ENXIO;
>> +     }
>> +
>> +     info->decoder_count = decoder_count;
>> +     info->global_ctrl = readl(hdm + CXL_HDM_DECODER_CTRL_OFFSET);
>> +     info->decoder_state = kcalloc(decoder_count,
>> +                                   sizeof(*info->decoder_state),
>> +                                   GFP_KERNEL);
> 
> That's same size as settings, so put the two of them a struct at the end of
> cxl_hdm_info with a short name (e.g. decoder[].state, decoder[].settings) and
> use a struct_size() allocation for them both.
> 
>> +     if (!info->decoder_state)
>> +             return -ENOMEM;
>> +
>> +     settings = info->settings;
>> +     for (int i = 0; i < info->decoder_count; i++) {
>> +             rc = cxl_pci_hdm_read_decoder(pdev, &info->decoder_state[i],
>> +                                           &settings[i], hdm, i);
>> +             if (rc)
>> +                     return rc;
>> +     }
>> +
>> +     return 0;
>> +}
>> +
>> +static int cxl_pci_hdm_read_info(struct pci_dev *pdev,
>> +                              struct cxl_register_map *map,
>> +                              struct cxl_hdm_info *info)
>> +{
>> +     bool restore_command;
>> +     u16 command;
>> +     int rc, rc2;
>> +
>> +     guard(pci_dev)(pdev);
>> +
>> +     rc = pci_read_config_word(pdev, PCI_COMMAND, &command);
>> +     if (rc)
>> +             return pcibios_err_to_errno(rc);
>> +
>> +     restore_command = !(command & PCI_COMMAND_MEMORY);
>> +     if (restore_command) {
>> +             rc = pci_write_config_word(pdev, PCI_COMMAND,
>> +                                        command | PCI_COMMAND_MEMORY);
>> +             if (rc)
>> +                     return pcibios_err_to_errno(rc);
>> +     }
>> +
>> +     rc = __cxl_pci_hdm_read_info(pdev, map, info);
>> +
>> +     if (!restore_command)
>> +             return rc;
>> +
>> +     rc2 = pci_write_config_word(pdev, PCI_COMMAND, command);
>> +     if (rc2) {
>> +             rc2 = pcibios_err_to_errno(rc2);
>> +             pci_err(pdev,
>> +                     "failed to restore PCI_COMMAND after CXL HDM cache init: %d\n",
>> +                     rc2);
> 
>> +             if (!rc)
>> +                     rc = rc2;
>> +     }
> Dance is more complex to read than just duplicating a little.
> 
>          if (rc)
>                  goto reset_command_reg
> 
>          return pci_write_config_word(pdev, PCI_COMMAND, command);
> 
> reset_command_reg:
>          if (pci_write_config_word(pdev, PCI_COMMAND, command) !=
>              PCIBIOS_SUCCESSFUL)
>                  pci_err(pdev, ...);
> 
> I doubt we care about what return of that is given we are on fire.
> 
>          return rc;
> 
>> +
>> +     return rc;
>> +}
>> +
>> +static int __pci_cxl_hdm_init(struct pci_dev *pdev)
>> +{
>> +     struct cxl_register_map map = { 0 };
> 
> The 0 doesn't add anything = { };
> 
>> +     int dvsec;
>> +     int rc;
>> +
>> +     if (!cxl_pci_hdm_capable(pdev))
>> +             return -ENOTTY;
>> +
>> +     rc = cxl_find_regblock(pdev, CXL_REGLOC_RBI_COMPONENT, &map);
>> +     if (rc)
>> +             return rc;
> So this reads the regblock locator to get where the component regs are
> by decodeing the bar and offset then finding the resource form that...
>> +
>> +     rc = cxl_pci_hdm_find_bar(pdev, map.resource, map.max_size, NULL, NULL);
>> +     if (rc)
>> +             return rc;
> This takes the resource and finds the bar and offset?
> 
> Going in circles. Can't you pull a helper out of the start of
> cxl_decode_regblock() and get the bar and offset directly.
> Even if you need the sanity checks along the way, I think you can
> do them more directly and only go in one direction.
> 
>> +
>> +     struct cxl_hdm_info *info __free(cxl_pci_hdm_info) =
>> +             kzalloc_obj(*info, GFP_KERNEL);
>> +     if (!info)
>> +             return -ENOMEM;
> 
> Why allocate here when so much can still fail?  We don't need
> it for a few more calls.
> 
>> +
>> +     dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
>> +                                       PCI_DVSEC_CXL_DEVICE);
>> +     if (!dvsec)
>> +             return -ENOTTY;
>> +
>> +     rc = pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CTRL,
>> +                               &info->dvsec_ctrl);
>> +     if (rc)
>> +             return pcibios_err_to_errno(rc);
>> +     info->dvsec_ctrl_valid = true;
>> +
> 
> I'd make this read and allocate - having found the size so we don't
> allocate more than necessary.  If any crazy device changes number of
> decoders on reset then we can just clamp it and print a rude message.
> 
>> +     rc = cxl_pci_hdm_read_info(pdev, &map, info);
>> +     if (rc)
>> +             return rc;
>> +
>> +     pdev->hdm = no_free_ptr(info);
>> +
>> +     return 0;
>> +}
>> +
>> +void pci_cxl_hdm_init(struct pci_dev *pdev)
>> +{
>> +     int rc;
>> +
>> +     rc = __pci_cxl_hdm_init(pdev);
>> +     if (rc && rc != -ENOTTY && rc != -ENODEV)
>> +             pci_dbg(pdev, "CXL HDM cache init failed: %d\n", rc);
>> +}
> 

These are now getting addressed in v13 patch 9:

   - removed the resource-to-BAR reverse lookup and now obtain the BAR 
and offset directly from the Register Locator path;
   - removed the duplicate raw-register decoder cache and retain one 
decoded settings representation;
   - combined the per-decoder state into the flexible array and allocate 
it with struct_size() after discovering the decoder count;
   - replaced the unexplained 0x10 size check.
   - protect cache publication and removal with cxl_rwsem.dpa; and
   - simplified PCI_COMMAND restoration and its error path.

I chose to reject a changed decoder count rather than clamp it because a 
partial snapshot would not describe all current hardware state. Added a 
comment for this.

https://lore.kernel.org/linux-cxl/20260922083924.2451158-10-smadhavan@nvidia.com/

-- 
Regards,
Srirangan

  reply	other threads:[~2026-09-23  0:02 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10  7:07 [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 01/12] cxl: Move HDM decoder programming helpers Srirangan Madhavan
2026-09-11 23:30   ` Jonathan Cameron
2026-09-22 23:28     ` Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 02/12] cxl: Make HDM commit helpers available to reset code Srirangan Madhavan
2026-09-10  7:07 ` [PATCH v12 03/12] cxl: Share HDM decoder decode logic Srirangan Madhavan
2026-09-12  0:07   ` Jonathan Cameron
2026-09-22 23:52     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 04/12] cxl: Cache decoder settings on PCI devices Srirangan Madhavan
2026-09-12  0:22   ` Jonathan Cameron
2026-09-22 23:59     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 05/12] cxl: Cache endpoint decoder settings during PCI enumeration Srirangan Madhavan
2026-09-12  1:03   ` Jonathan Cameron
2026-09-23  0:02     ` Srirangan Madhavan [this message]
2026-09-10  7:08 ` [PATCH v12 06/12] cxl: Add CXL Device Reset helper Srirangan Madhavan
2026-09-12  1:26   ` Jonathan Cameron
2026-09-15 13:56     ` Lucero Palau, Alejandro
2026-09-23  0:51       ` Srirangan Madhavan
2026-09-23  0:05     ` Srirangan Madhavan
2026-09-23  0:20     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 07/12] cxl: Validate HDM ranges before CXL reset Srirangan Madhavan
2026-09-12  1:33   ` Jonathan Cameron
2026-09-23  0:08     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 08/12] PCI/CXL: Reject CXL Reset on multifunction devices Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 09/12] cxl: Restore CXL state after PCI reset Srirangan Madhavan
2026-09-12  1:43   ` Jonathan Cameron
2026-09-23  0:10     ` Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 10/12] PCI/CXL: Expose CXL Reset as a PCI reset method Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 11/12] Documentation/ABI: Document CXL Reset " Srirangan Madhavan
2026-09-10  7:08 ` [PATCH v12 12/12] PCI/CXL: Restore CXL state after CXL bus reset Srirangan Madhavan
2026-09-10  7:31 ` [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=23266146-f92c-4b97-bf1b-99e4973b7dfc@nvidia.com \
    --to=smadhavan@nvidia.com \
    --cc=alex.williamson@redhat.com \
    --cc=alison.schofield@intel.com \
    --cc=alwilliamson@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=icheng@nvidia.com \
    --cc=ira.weiny@intel.com \
    --cc=jan@nvidia.com \
    --cc=jic23@kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-tegra@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=skancherla@nvidia.com \
    --cc=vaslot@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®