From: "Danilo Krummrich" <dakr@kernel.org>
To: "Alex Williamson" <alex@shazbot.org>
Cc: "Jason Gunthorpe" <jgg@nvidia.com>, "Zhi Wang" <zhiw@nvidia.com>,
<acourbot@nvidia.com>, <yishaih@nvidia.com>,
<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
<airlied@gmail.com>, <simona@ffwll.ch>, <ojeda@kernel.org>,
<alex.gaynor@gmail.com>, <boqun.feng@gmail.com>,
<gary@garyguo.net>, <bjorn3_gh@protonmail.com>,
<lossin@kernel.org>, <a.hindborg@kernel.org>,
<aliceryhl@google.com>, <tmgross@umich.edu>,
<jhubbard@nvidia.com>, <ecourtney@nvidia.com>, <cjia@nvidia.com>,
<smitra@nvidia.com>, <kjaju@nvidia.com>, <alkumar@nvidia.com>,
<ankita@nvidia.com>, <aniketa@nvidia.com>, <kwankhede@nvidia.com>,
<targupta@nvidia.com>, <nova-gpu@lists.linux.dev>,
<linux-kernel@vger.kernel.org>, <zhiwang@kernel.org>,
<kvm@vger.kernel.org>
Subject: Re: [PATCH 12/13] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver
Date: Mon, 14 Sep 2026 23:36:24 +0200 [thread overview]
Message-ID: <DLFD2ZDSK9YQ.3A4R66G8UJMD8@kernel.org> (raw)
In-Reply-To: <20260914121217.70fa0d93@shazbot.org>
On Mon Sep 14, 2026 at 8:12 PM CEST, Alex Williamson wrote:
> Hi Danilo,
>
> On Fri, 11 Sep 2026 22:39:16 +0200
> "Danilo Krummrich" <dakr@kernel.org> wrote:
>
>> Hi Alex, Jason, Zhi,
>>
>> On Sat Sep 5, 2026 at 10:11 AM CEST, Zhi Wang wrote:
>> > NVIDIA vGPU VFs require their open, reset, and close lifecycle to be
>> > coordinated with the PF-side nova-core driver.
>>
>> [...]
>>
>> > drivers/vfio/pci/nvidia-vgpu/main.c | 253 ++++++++++++++++++++++++++
>>
>> This is going to be a longer response; sorry about this in advance.
>>
>> Looking at the FFI boundary introduced in the previous patch, I'm concerned that
>> it translates the driver model relationships we've expressed through Rust's
>> ownership and lifetime model back into raw pointers and lifetime assumptions
>> that callers must uphold. It also introduces manual lifecycle management across
>> the boundary, rather than preserving nova-core's RAII-based ownership model.
>>
>> I think implementing the NVIDIA vGPU driver in Rust would let us preserve those
>> relationships across the interface, make lifecycle management less error-prone,
>> and fit naturally alongside nova-core and nova-drm.
> [snip]
>>
>> If you've made it this far, thanks for reading through this long write-up. I
>> hope you find it useful. Please let me know if you have any questions or
>> thoughts.
>
> I can't really say I made it this far with comprehension, but thanks
> for the effort ;)
>
> The one piece here that I can actually review is [5], where
> dev_get_drvdata() is replaced with a vfio-pci-core struct pointer
> embedded in the struct pci_dev, which is a non-starter as far as having
> a common PCI-core shared by various drivers.
Well, that was just a quick hack to get it out of the way. :)
I think there are a couple of options.
(1) Make the PM helpers take a struct vfio_pci_core_device * in the first
place and let the driver forward to the helpers in its own PM callbacks.
(2) Provide an (optional?) driver callback that translates a struct pci_dev to
struct vfio_pci_core_device.
(3) Provide a macro for drivers to define PM ops, letting drivers provide the
function that translates struct pci_dev to struct vfio_pci_core_device.
(4) Give struct vfio_pci_core_device its own PM domain (which is probably a
bit overkill :).
I understand that the idea is to hide the PM handling in the vfio-pci framwork,
but I think the existing implementation is a bit of a layering violation, since
class device implementations shouldn't impose requirements on the bus device
private data layout.
I also think that the approach to fully hide it in the framework is only really
worth if it doesn't otherwise impose subtle requirements on the driver (such as
the layout requirement of the bus device private data).
Thus, I'd personally just go with (1) as it is the most honest approach in terms
of driver layering. But I think (2) is a good alternative that is not more
invasive than asking drivers to set the bus device private data to
struct vfio_pci_core_device *.
> My concerns are of course who is going to review the Rust vfio-pci
> variant drivers from a vfio perspective, not just a drm driver
> viewpoint.
I did put some considerations about this; I think Zhi would be a great candidate
for this. :) And as mentioned I can also take some of the responsibility, but I
also want to be honest about the fact that I have a lot on my plate already.
> Who is going to be responsive when the interfaces break and
> monitor vfio proactively to prevent such breakages, and how do we avoid
> derailing feature development in the core code base.
In practice we shouldn't see any breakages with the Rust code that wouldn't also
break the C VFIO drivers. If this would be the case it would mean that the Rust
code relies on guarantees that none of the C drivers rely on, which would likely
mean that is was never a guarantee that was actually promised by the VFIO core
code.
There may be cases where Rust code breaks, and C code won't, but those should be
mechanical things, such as a type mismatch where Rust is more careful, e.g. if
we'd change a CPP define to an enum, etc.
> Can a Rust vfio-pci variant driver be self-contained, or to what extent does
> it impose on the framework, such as the drvdata idiom.
As mentioned above, I think this one is more of a layering violation in the
vfio-pci-core; class devices shouldn't impose layout requirements on bus device
private data.
The reason C drivers can get away with it more easily is e.g. that C relies on
procedural cleanup and that all responsibility for managing lifetimes sits on
the drivers themselves, so it is easier for them to adjust. But in general, it
wouldn't work out if all class device registrations or other core primitives
would have the same expectation.
For Rust specifically it is that the driver core controls the lifetime of the
bus device private data, which is a fundamental requirement to e.g. represent
registrations as RAII types. E.g. the vfio::pci::Registration has to be stored
in the bus device private data, such that it is guaranteed to be correctly
destroyed on driver unbind.
To get back to your question, a Rust vfio-pci variant driver should be
self-contained. The interface sits in the abstraction that translates the C
driver API to a Rust driver API. It sometimes can help quite significantly (e.g.
in terms of how complex the Rust code needs to get in order to actually be safe)
if the C code does a minor adjustment, but it shouldn't be necessary.
For instance, the only addition to the driver core we have is an additional
callback in struct device_driver, and in the future an additional pointer in
struct device_private; none of those couldn't be worked around in some way
though.
On the other hand there are examples where the Rust introduction motivated
design improvements on the C side, or bug fixes for issues that were caught
while writing a safe abstraction. For instance, we recently had some fixes
around dyn IDs in the PCI and USB core, which both were motivated by Rust code.
> FWIW, AI can only go so far to support reviews. Having the code
> insight to ask the right questions is essential. A human in the loop
> is a requirement.
>
> Additionally, if we can't narrow the device matching to only the
> Nova-core supported VFs,
Agreed, and I think that should be possible. Is there a particular case you
think of where this wouldn't hold?
next prev parent reply other threads:[~2026-09-14 21:36 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-05 8:11 [PATCH 00/13] Introduce NVIDIA vGPU manager and " Zhi Wang
2026-09-05 8:11 ` [PATCH 01/13] gpu: nova-core: vgpu: add post-GSP-boot vGPU initialization Zhi Wang
2026-09-11 6:51 ` Alexandre Courbot
2026-09-05 8:11 ` [PATCH 02/13] gpu: nova-core: mm: add VramBlock and Bar1Map Zhi Wang
2026-09-11 5:01 ` Alistair Popple
2026-09-05 8:11 ` [PATCH 03/13] gpu: nova-core: vgpu: add VRAM slot allocator Zhi Wang
2026-09-05 8:11 ` [PATCH 04/13] gpu: nova-core: vgpu: add r000 plugin bindings Zhi Wang
2026-09-05 8:11 ` [PATCH 05/13] gpu: nova-core: vgpu: add instance create/destroy Zhi Wang
2026-09-05 8:11 ` [PATCH 06/13] gpu: nova-core: gsp: add GMC transaction helpers Zhi Wang
2026-09-05 8:11 ` [PATCH 07/13] gpu: nova-core: vgpu: add vGPU bootload Zhi Wang
2026-09-05 8:11 ` [PATCH 08/13] gpu: nova-core: vgpu: implement PluginRpc channel and config params Zhi Wang
2026-09-05 8:11 ` [PATCH 09/13] gpu: nova-core: vgpu: scrub guest framebuffer memory with CeUtils Zhi Wang
2026-09-05 8:11 ` [PATCH 10/13] gpu: nova-core: vgpu: export plugin log buffers via debugfs Zhi Wang
2026-09-05 8:11 ` [PATCH 11/13] gpu: nova-core: vgpu: export lifecycle operations to VFIO Zhi Wang
2026-09-05 8:11 ` [PATCH 12/13] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver Zhi Wang
2026-09-09 3:00 ` Alex Williamson
2026-09-11 20:39 ` Danilo Krummrich
2026-09-14 18:12 ` Alex Williamson
2026-09-14 21:36 ` Danilo Krummrich [this message]
2026-09-05 8:11 ` [PATCH 13/13] gpu: nova-core: reserve the 48-VM WPR2 heap Zhi Wang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLFD2ZDSK9YQ.3A4R66G8UJMD8@kernel.org \
--to=dakr@kernel.org \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=alex@shazbot.org \
--cc=aliceryhl@google.com \
--cc=alkumar@nvidia.com \
--cc=aniketa@nvidia.com \
--cc=ankita@nvidia.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=cjia@nvidia.com \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=jgg@nvidia.com \
--cc=jhubbard@nvidia.com \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kvm@vger.kernel.org \
--cc=kwankhede@nvidia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=skolothumtho@nvidia.com \
--cc=smitra@nvidia.com \
--cc=targupta@nvidia.com \
--cc=tmgross@umich.edu \
--cc=yishaih@nvidia.com \
--cc=zhiw@nvidia.com \
--cc=zhiwang@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®