mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jonathan Cameron <jic23@kernel.org>
To: Manish Honap <mhonap@nvidia.com>
Cc: Gregory Price <gourry@gourry.net>,
	"alex@shazbot.org" <alex@shazbot.org>,
	"jgg@ziepe.ca" <jgg@ziepe.ca>, Ankit Agrawal <ankita@nvidia.com>,
	"dave.jiang@intel.com" <dave.jiang@intel.com>,
	"alejandro.lucero-palau@amd.com" <alejandro.lucero-palau@amd.com>,
	Srirangan Madhavan <smadhavan@nvidia.com>,
	"corbet@lwn.net" <corbet@lwn.net>,
	"skhan@linuxfoundation.org" <skhan@linuxfoundation.org>,
	"dave@stgolabs.net" <dave@stgolabs.net>,
	"alison.schofield@intel.com" <alison.schofield@intel.com>,
	"vishal.l.verma@intel.com" <vishal.l.verma@intel.com>,
	"iweiny@kernel.org" <iweiny@kernel.org>,
	"ming.li@zohomail.com" <ming.li@zohomail.com>,
	Yishai Hadas <yishaih@nvidia.com>,
	Shameer Kolothum Thodi <skolothumtho@nvidia.com>,
	"kevin.tian@intel.com" <kevin.tian@intel.com>,
	"bhelgaas@google.com" <bhelgaas@google.com>,
	"dmatlack@google.com" <dmatlack@google.com>,
	"kees@kernel.org" <kees@kernel.org>,
	"gustavoars@kernel.org" <gustavoars@kernel.org>,
	Neo Jia <cjia@nvidia.com>, Krishnakant Jaju <kjaju@nvidia.com>,
	Vikram Sethi <vsethi@nvidia.com>, Zhi Wang <zhiw@nvidia.com>,
	"linux-doc@vger.kernel.org" <linux-doc@vger.kernel.org>,
	"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
	"kvm@vger.kernel.org" <kvm@vger.kernel.org>,
	"linux-cxl@vger.kernel.org" <linux-cxl@vger.kernel.org>,
	"linux-pci@vger.kernel.org" <linux-pci@vger.kernel.org>,
	"linux-kselftest@vger.kernel.org"
	<linux-kselftest@vger.kernel.org>,
	"linux-hardening@vger.kernel.org"
	<linux-hardening@vger.kernel.org>
Subject: Re: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough
Date: Thu, 24 Sep 2026 01:54:09 +0100	[thread overview]
Message-ID: <20260924015409.3681ad6d@jic23-hlaptop> (raw)
In-Reply-To: <IA1PR12MB9030771093FA13DCB01C481EBD842@IA1PR12MB9030.namprd12.prod.outlook.com>

On Mon, 21 Sep 2026 09:43:34 +0000
Manish Honap <mhonap@nvidia.com> wrote:

> Hello Gregory,
> 
> Thank you for the careful read of the documentation. I will rework the
> sections you flagged as below:
> 

Hi Manish,

It is much more helpful to reply inline with the relevant sections so we can
see the context wrt to what Gregory was asking.  That saves others
reading the discussion from having to carefully build up a global model
of what was said to understand what the replies are about.  Also I'll hazzard
a guess that Gregory may have read enough code since he sent the review
that he has already forgotten some of his own comments! (maybe that's just
me!)  I appreciate that doesn't always work if you are going to address
the comment via more radical restructuring.  Or in short, please don't
top post!

Thanks,

Jonathan


> - Address model
> 
>   I will update the wording to mention that VMM does know the GPA: it builds
>   the guest's CFMWS window and so chooses the guest-physical range the device
>   can inhabit. What the host owns is the HPA and the placement within it, not
>   knowledge of the GPA.
>   I will add the two models you described, fixed placement for an accelerator
>   that needs exact physical placement versus no fixed placement for pooled or
>   compressed memory, with a small HPA/GPA example, and note that the current
>   series under discussion implements the fixed-offset case.
> 
> - Virtual decoders
> 
>   I will rename the The "Guest decoder and commit" section to "Virtual
>   decoders" and add a short flow showing a guest vdecoder write being absorbed
>   and a live read returning COMMITTED from the host-locked physical decoder.
> 
> - "The kernel virtualizes ..."
> 
>   I will update it to say vfio-pci-core (its config-space permission hooks).
> 
> - guest IOAS, stage-2 mapping, and struct-page-less coherent memory
> 
>   I will define these terms where first used and include that the range is
>   handed to the device whole and never onlined as system RAM. I will also add
>   details why a stale stage-2 mapping must not outlive the HDM window, and how
>   the VMM rebuilds it.
> 
> - Reset
> 
>   I will reword this section. FLRs are virtualized so a guest reset can't
>   inflict physical CXL.mem/decoder effects on the host rather than loose
>   wording "must not take an FLR"
> 
> I can share the revised text ahead of the v6 posting if that is easier to
> review.
> 
> Thanks,
> Manish
> 
> > -----Original Message-----
> > From: Gregory Price <gourry@gourry.net>
> > Sent: Thursday, September 17, 2026 1:03 AM
> > To: Manish Honap <mhonap@nvidia.com>
> > Cc: alex@shazbot.org; jgg@ziepe.ca; Ankit Agrawal <ankita@nvidia.com>;
> > jic23@kernel.org; dave.jiang@intel.com; alejandro.lucero-palau@amd.com;
> > Srirangan Madhavan <smadhavan@nvidia.com>; corbet@lwn.net;
> > skhan@linuxfoundation.org; dave@stgolabs.net; alison.schofield@intel.com;
> > vishal.l.verma@intel.com; iweiny@kernel.org; ming.li@zohomail.com; Yishai
> > Hadas <yishaih@nvidia.com>; Shameer Kolothum Thodi
> > <skolothumtho@nvidia.com>; kevin.tian@intel.com; bhelgaas@google.com;
> > dmatlack@google.com; kees@kernel.org; gustavoars@kernel.org; Neo Jia
> > <cjia@nvidia.com>; Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi
> > <vsethi@nvidia.com>; Zhi Wang <zhiw@nvidia.com>; linux-
> > doc@vger.kernel.org; linux-kernel@vger.kernel.org; kvm@vger.kernel.org;
> > linux-cxl@vger.kernel.org; linux-pci@vger.kernel.org; linux-
> > kselftest@vger.kernel.org; linux-hardening@vger.kernel.org
> > Subject: Re: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-
> > 2 device passthrough
> > 
> > External email: Use caution opening links or attachments
> > 
> > 
> > On Thu, Sep 17, 2026 at 12:05:39AM +0530, mhonap@nvidia.com wrote:  
> > > From: Manish Honap <mhonap@nvidia.com>
> > >  
> > 
> > 1) Thank you so much for writing documentation, i truly appreciate this.
> > 
> > 2) I apologize in advance for my terseness, I know writing is hard,
> >    please do not interpret this as disliking your writing or series.
> >   
> > > +Address model
> > > +=============
> > > +
> > > +The HDM memory is a coherent host physical range (HPA). The host
> > > +kernel resolves that range before the guest sees the device, and owns
> > > +it for the bind lifetime. The guest only chooses where the memory
> > > +appears in its own physical address space (GPA), by programming a
> > > +virtual endpoint HDM decoder. The guest never reprograms the physical  
> > decoder.  
> > > +
> > > +The kernel holds the HPA and does not see the GPA. The guest programs
> > > +a GPA and does not see the HPA. The VMM holds the device fd, reads
> > > +the committed base from the decoder-register region described below,
> > > +and maps the HPA-backed HDM region at the GPA the guest committed.
> > > +The base the guest reads back is the GPA, not the HPA.
> > > +  
> > 
> > I think this must be slightly inaccurate / imprecise wording.
> > 
> > The host *must* provide some form of physical memory window to the guest
> > at initialization time, otherwise the guest has no way to know - at boot time -
> > that there's even a window of memory it can use.
> > 
> > That's what the CFMWS is.  This is initialized by the hypervisor - which is
> > controlled by the host.
> > 
> > So the host (at least the VMM) must know, for the region the entire device
> > *could* inhabit, what that GPA is - because it's the one that makes the CFMWS
> > for the guest.
> > 
> > If this is not the case, then something is missing from this documentation to
> > explain why.
> > 
> > 
> > If you're actually trying to say is that the GPA's programmed into the virtual
> > decoders are largely symbolic - this at best feels a bit inaccurate and simply an
> > implementation detail.
> > 
> > The host's virtio device could enforce ....:
> > 
> > Host Range:
> >    CFMWS HPA - [0x10000, 0x20000]
> >                    |         |
> > Guest Range:       |         |
> >    CFMWS GPA - [0x50000, 0x60000]
> > 
> > In that case, you'd get the following translation...
> >    vdecoder0.0 - [0x58000, 0x60000]
> >    CFMWS GPA   - [0x58000, 0x60000]
> >    CFMWS HPA   - [0x18000, 0x20000]
> > 
> > Or the virtio device could not enforce that and let the host page-fault just hand
> > it a random page from the actual CXL device.
> > 
> >    vdecoder0.0 - [0x58000, 0x60000]
> >    CFMWS GPA   - [0x58000, 0x60000]
> >                          |
> >               No discrete host mapping
> > 
> > 
> > The former makes sense if the device (accelerator) requires exact physical
> > placement to do its accelerator nonsense.
> > 
> > The latter makes sense if the device (accelerator) doesn't care about placement
> > (compressed memory).
> > 
> > This is not saying we need support both out of the box, but we shouldn't lock
> > ourselves into the former unless there's some reason why the latter is not
> > reasonable.
> > 
> > Can you please help document what the actual expected behavior is with
> > examples in the Address model section so it's easier to understand the intent?
> > That will help quite a bit.
> >   
> > > +Guest decoder and commit
> > > +========================
> > > +
> > > +The guest programs its virtual endpoint decoder through the trapped
> > > +region: it writes a base (a GPA), a size, and then the COMMIT bit.
> > > +The host already resolved and committed the physical placement before
> > > +the guest ran,  
> > 
> > So the host does know GPA, just not exact placement.
> >   
> > > so a live read of the decoder always shows COMMITTED and the
> > > +guest's commit poll completes. The physical decoder is never
> > > +rewritten; the guest's writes are absorbed.
> > > +  
> > 
> > Rather clunky, round-about way to say "The guest decoders are
> > virtualized".   If possible, it would be nice to formalize this concept
> > into "Virtual Decoders" - since that's what this is.
> > 
> > With that concept i think you can probably generate some nice diagrams that
> > show how the guest vdecoder's interact with the host drivers.
> >   
> > > +The VMM observes the commit, reads the committed base, and maps the
> > > +HDM region at that GPA.
> > > +  
> > 
> > So the host does know the GPA.
> >   
> > > +CXL Device DVSEC
> > > +================
> > > +
> > > +The kernel virtualizes the CXL Device DVSEC body through the
> > > +config-space permission hooks. Reads and writes inside the DVSEC body
> > > +use a per-open shadow; a guest write stays in the shadow and does not  
> > reach hardware.  
> > > +Accesses outside the DVSEC body go to the device as usual.
> > > +  
> > 
> > "The kernel" - what part? vfio-pci ? the vmm ?
> > 
> >   
> > > +DMA and iommufd
> > > +===============
> > > +
> > > +A Type-2 accelerator issues ATS-translated DMA to addresses inside
> > > +its own HDM window, so that range must be present in the guest IOAS
> > > +that backs the nested stage-2 translation.  
> > 
> > Type-2, ATS, DMA, HDM window, guest IOAS, stage-2 translation
> > 
> > I think the only thing i don't know in this sentence is "guest IOAS" and it's still
> > hurting my brain to read.
> > 
> > Are all accelerators expect to have this particular interaction, or just yours?
> >   
> > > The HDM range is struct-page-less coherent
> > > +memory, which a userspace-VA ``IOMMU_IOAS_MAP`` cannot pin.
> > > +  
> > 
> > The hardest part about writing about virtualization is keeping a consistent
> > mental model from section to section.
> > 
> > which userspace? guest? host?  (i presume guest here)
> > 
> > `struct-page-less coherent memory`
> >    e.g. the host never hotplugs this, it hands the entire region
> >    directly to the VFIO device, right?
> > 
> >    I think this would be nice to spell out somewhere.
> >   
> > > +The HDM memory region is therefore exportable as a dma-buf:
> > > +``VFIO_DEVICE_FEATURE_DMA_BUF`` on that region returns an fd that
> > > +iommufd maps with ``IOMMU_IOAS_MAP_FILE``, mapping the physical  
> > range  
> > > +without a VA or a page pin. The dma-buf is revoked whenever the
> > > +mapping is torn down (reset, power transition, teardown), so a stale
> > > +stage-2 mapping cannot outlive the HDM window.
> > > +  
> > 
> > For the sake of readers, I think either a little bit more information on this
> > "stage-2 mapping" concept is needed to make sense of what's going on here
> > and why it mustn't outlive the HDM window.
> >   
> > > +Reset
> > > +=====
> > > +
> > > +A CXL Type-2 function must not take a Function Level Reset: an FLR
> > > +resets the coherent CXL.mem state and the HDM decoder. The PCI core
> > > +reflects this by preferring the CXL reset over FLR, so a function
> > > +reset of a CXL device runs the CXL DVSEC reset sequence, which resets
> > > +the function and then restores the HDM decoder and the PCI config state.
> > > +  
> > 
> > I think what you're trying to say is that FLRs are never passed to the device
> > because it can cause physical device effects that defeat the purpose of the
> > virtualization, yes?
> > 
> > So we virtualize FLRs...
> >   
> > > +A guest requests a reset by writing Initiate CXL Reset in the DVSEC.
> > > +That write only stamps completion in the shadow. The real reset runs
> > > +at the vfio reset points (the reset ioctl and a virtualized FLR through config  
> > space):
> > 
> > As you describe here.
> > 
> > So it's not that an accelerator "must not take an FLR" - it's that FLRs are
> > virtualized to prevent deleterious effects on the host / hardware.
> > 
> > Am I misunderstanding this?
> >   
> > > +the kernel zaps the HDM mapping and revokes the dma-buf, then runs
> > > +the CXL reset, which always clears the device memory, and restores
> > > +and re-samples the decoder afterwards. A CXL port masks Secondary Bus
> > > +Reset by default, so a ``VFIO_DEVICE_PCI_HOT_RESET`` does not reach
> > > +the endpoint and the HDM state is untouched. If the port has SBR
> > > +unmasked the reset can decommit the decoder without restoring it, so
> > > +the reset_done handler gates HDM access; a ``VFIO_DEVICE_RESET`` then
> > > +runs the CXL reset sequence and restores it.
> > > +
> > > +The decoder register region is served by live reads of the hardware
> > > +decoder with guest writes absorbed: the decoder is committed and
> > > +locked by the host, so a guest can neither decommit nor reprogram it,
> > > +and the kernel keeps no shadow of the decoder state.  
> > 
> > This is basically what I said at the beginning - it must either be that the host
> > provides locked auto-decoders at boot, or it must provide proper virtualization
> > so that the decoders settings are fully virtualized.
> > 
> > Seems it's the former, and that makes sense.  Please correct me if i'm
> > misunderstanding.
> >   
> > > After a reset the kernel restores and
> > > +re-samples the firmware-committed decoder, so the geometry the guest
> > > +reads back is unchanged. A VMM that dropped its HDM mapping, for
> > > +example across a reset or a D3hot->D0 transition, must rescan the
> > > +decoder and rebuild its
> > > +stage-2 mapping before it resumes HDM access.
> > > +  
> > 
> > Yeah i think we need a bit more information about this stage-2 mapping
> > rebuild to make sense of this.  Maybe I'm just not read-up enough on this
> > particular setup - is there another part of the docs you can link to that talk
> > about this, or are you able to share some details as to what this rebuild
> > process looks like?
> > 
> > ~Gregory  


  reply	other threads:[~2026-09-24  0:54 UTC|newest]

Thread overview: 52+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 18:35 [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-09-16 18:35 ` [PATCH v5 01/27] cxl/regs: Split the BAR block request and ioremap helpers mhonap
2026-09-22  1:28   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 02/27] cxl/regs: Let a BAR-owning driver own the component register block mhonap
2026-09-22  1:36   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-09-22  1:43   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 04/27] cxl: Add cxl_reset_dvsec_sequence() for vfio-pci mhonap
2026-09-16 18:35 ` [PATCH v5 05/27] vfio/pci: Add the CXL provider ops registration interface mhonap
2026-09-16 18:35 ` [PATCH v5 06/27] vfio/pci: Detect CXL devices and load the CXL provider on demand mhonap
2026-09-17  8:48   ` Richard Cheng
2026-09-21 10:05     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 07/27] vfio/pci: Honor -EPROBE_DEFER from CXL provider probe mhonap
2026-09-16 18:35 ` [PATCH v5 08/27] vfio/pci: Fall back to plain vfio-pci when CXL init fails mhonap
2026-09-22  2:15   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 09/27] vfio/pci: Add a generic excluded-range list mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 10/27] vfio/pci: Migrate MSI-X exclusion onto the " mhonap
2026-09-16 18:35 ` [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 12/27] vfio/pci: Call the CXL open and close hooks around device use mhonap
2026-09-16 18:35 ` [PATCH v5 13/27] vfio/pci: Bracket PCI resets with the CXL reset hooks mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 14/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-09-16 18:35 ` [PATCH v5 15/27] vfio/cxl: Add the vfio-cxl provider module skeleton mhonap
2026-09-16 18:35 ` [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 17/27] vfio/cxl: Own the whole component register BAR mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 18/27] vfio/cxl: Expose the HDM memory region to the guest mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 19/27] vfio/cxl: Contain HDM memory errors with memory_failure() mhonap
2026-09-16 18:35 ` [PATCH v5 20/27] vfio/cxl: Expose the HDM decoder registers read-only to the guest mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 21/27] vfio/cxl: Exclude the HDM decoder registers from direct BAR access mhonap
2026-09-17  7:28   ` Richard Cheng
2026-09-21  9:52     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 22/27] vfio/cxl: Clear the HDM access gate after a hot reset mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 23/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-09-16 18:35 ` [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf mhonap
2026-09-17  7:55   ` Richard Cheng
2026-09-21  9:58     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 25/27] vfio/cxl: Run the CXL reset at the vfio reset points mhonap
2026-09-17  8:11   ` Richard Cheng
2026-09-21 10:01     ` Manish Honap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-09-16 19:33   ` Gregory Price
2026-09-21  9:43     ` Manish Honap
2026-09-24  0:54       ` Jonathan Cameron [this message]
2026-09-16 18:35 ` [PATCH v5 27/27] selftests/vfio: Add CXL Type-2 passthrough tests mhonap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260924015409.3681ad6d@jic23-hlaptop \
    --to=jic23@kernel.org \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alex@shazbot.org \
    --cc=alison.schofield@intel.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=cjia@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=dmatlack@google.com \
    --cc=gourry@gourry.net \
    --cc=gustavoars@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=kjaju@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=ming.li@zohomail.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skolothumtho@nvidia.com \
    --cc=smadhavan@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    --cc=yishaih@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®