From: "Danilo Krummrich" <dakr@kernel.org>
To: "Jason Gunthorpe" <jgg@nvidia.com>
Cc: "Alex Williamson" <alex@shazbot.org>,
"Zhi Wang" <zhiw@nvidia.com>, <acourbot@nvidia.com>,
<yishaih@nvidia.com>, <skolothumtho@nvidia.com>,
<kevin.tian@intel.com>, <airlied@gmail.com>, <simona@ffwll.ch>,
<ojeda@kernel.org>, <alex.gaynor@gmail.com>,
<boqun.feng@gmail.com>, <gary@garyguo.net>,
<bjorn3_gh@protonmail.com>, <lossin@kernel.org>,
<a.hindborg@kernel.org>, <aliceryhl@google.com>,
<tmgross@umich.edu>, <jhubbard@nvidia.com>,
<ecourtney@nvidia.com>, <cjia@nvidia.com>, <smitra@nvidia.com>,
<kjaju@nvidia.com>, <alkumar@nvidia.com>, <ankita@nvidia.com>,
<aniketa@nvidia.com>, <kwankhede@nvidia.com>,
<targupta@nvidia.com>, <nova-gpu@lists.linux.dev>,
<linux-kernel@vger.kernel.org>, <zhiwang@kernel.org>,
<kvm@vger.kernel.org>
Subject: Re: [PATCH 12/13] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver
Date: Wed, 16 Sep 2026 17:38:21 +0200 [thread overview]
Message-ID: <DLGUPX9KPFKF.12CU304YAST79@kernel.org> (raw)
In-Reply-To: <20260916141718.GU3968357@nvidia.com>
On Wed Sep 16, 2026 at 4:17 PM CEST, Jason Gunthorpe wrote:
> On Tue, Sep 15, 2026 at 11:19:01PM +0200, Danilo Krummrich wrote:
>
>> The driver_data pointer in struct device is defined to be a pointer where a
>> driver can store *arbitrary* data for the duration the driver is bound to this
>> device.
>
> There are places in the kernel where the drvdata of the bound device
> ends up owned by the subsystem, not the end driver. Yes, an ideal
> driver + subsystem should never even need drvdatab beyond remove. Yet,
> things are not perfect..
>
> The fundamental issue is some kernel API surfaces that the subystem
> needs to work with only provide a struct device in their callbacks and
> the subystem has no option but to use the drvdata for its own purpose
> to recover the subsystem specific data.
>
> For example VFIO hooks into this nasty API:
>
> ret = vga_client_register(pdev, vfio_pci_set_decode);
> if (ret)
> return ret;
>
> Which doesn't provide a void * token to pass the core code's
> struct.
>
> Another is all the PCI callbacks which assume the op behind them uses
> drvdata to get its data.
This is where the class (i.e. vfio-pci) should instead take a driver callback,
as only the driver really knowns about the layering details.
Usually, class device implementations can't make assumptions of the underlying
bus, because they have to work for any bus. I.e. there's no other way than
providing helpers and letting drivers do the glue code between the bus and the
class device.
> It is not necessarily easy to fix. Rrouting all those PCI callbacks
> through trampolines in every single driver is really not an appealing
> design. There are alot of VFIO drivers.
The reason this seems undesirable from a vfio-pci perspective is that it is
special in the sense that it is a class device that is specifically built to sit
on top of a spcific bus device (i.e. struct pci_dev).
Since this is a rare (maybe even unique?) edge case, the driver core has indeed
no infrastructure in place to represent this.
> Maybe it needs a dev->drvdata and dev->subsystem_data, maybe it needs
> some PCI thing where the pm ops can get a void *, IDK.
It still makes me think that there should be some closer integration of vfio-pci
with the PCI core, as it is specifically built for this bus.
In fact, what you say above is the generalization of my "hack" [1], which, the
more I think about this, seems actually less of a hack. :)
I would object a bit to the generalization with dev->subsystem_data as it
screams for abuse, but the thing in [1] seems more and more reasonable to me.
(Another option would be to really split it up and have a virtual vfio-pci bus
on top of PCI, which would also allow for custom match logic, but that also
seems pretty overkill.)
>> As mentoined above, there are people volunteering now. Without starting it, it
>> can't scale further than that. :)
>
> There are lots of other vfio patches that need attention too, and it
> seems we are short of that more than anything. Now you need to do a
> bunch of C refactoring patches as well just to get things ready to
> show a bunch of rust code. It is a lot of work.
It's only the drvdata thing, what else is missing?
> I don't really understand in a nutshell why we should do this for nova
> the mails were so long... Can we not just ignore the lifetime
> imperfection for this?
I mentioned some points in the first two paragraphs of [2]. Besides that, I
don't see a reason why we should spend time and effort for working out the
inferior solution, where the better alternative is even less effort, contributes
to better quality and stability of the whole driver project and also offers a
chance for the vfio subsystem to gain new contributors and gather experience
with the language that has proven itself in many areas already.
I mean, there'd still be the option to make it an experiment and say let's add
the abstractions and the nvidia-vgpu driver and see how it works out for a while
before allowing more Rust drivers. And if it really turns out to be bad, it
should also be easy to rip it out, replace nvidia-vgpu with a C driver and throw
in a crappy FFI layer. :)
>> However, I don't really see the use-case; you can't load nvidia-vgpu without
>> nova-core in the first place, so it would require to unbind nova-core through
>> sysfs force unbind, no?
>
> nvidia gpu is a more unique scenario, if you are building a general
> bindings it has to support the flows like this.
Sure, as mentioned, the code I sketched up should already be able to do this.
> I've wanted to rework the way the common ops are shimmed in for a
> while, you'd probably want to do that before rust bindings.
TBH, I don't think it makes a difference; having Rust abstractions is pretty
much the same as having another driver. I.e. it would be equivalent to saying
"before we accept another pci-vfio driver we need to do some rework".
Actually, I even think it can be an advantage, as you could think of the Rust
abstractions like a driver with built-in correctness checks, so it can help to
validate the changes.
Of course, this requires the responsibles of the Rust code to help with that and
I think we have this commitment.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/dakr/linux.git/commit/?id=76b3bfd6386a01f338780e1a37b5bad3f5a48d31
[2] https://lore.kernel.org/nova-gpu/DLCRZLO06SIO.LS7TWQXIPZSQ@kernel.org/
next prev parent reply other threads:[~2026-09-16 15:38 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-05 8:11 [PATCH 00/13] Introduce NVIDIA vGPU manager and " Zhi Wang
2026-09-05 8:11 ` [PATCH 01/13] gpu: nova-core: vgpu: add post-GSP-boot vGPU initialization Zhi Wang
2026-09-11 6:51 ` Alexandre Courbot
2026-09-05 8:11 ` [PATCH 02/13] gpu: nova-core: mm: add VramBlock and Bar1Map Zhi Wang
2026-09-11 5:01 ` Alistair Popple
2026-09-05 8:11 ` [PATCH 03/13] gpu: nova-core: vgpu: add VRAM slot allocator Zhi Wang
2026-09-05 8:11 ` [PATCH 04/13] gpu: nova-core: vgpu: add r000 plugin bindings Zhi Wang
2026-09-05 8:11 ` [PATCH 05/13] gpu: nova-core: vgpu: add instance create/destroy Zhi Wang
2026-09-05 8:11 ` [PATCH 06/13] gpu: nova-core: gsp: add GMC transaction helpers Zhi Wang
2026-09-05 8:11 ` [PATCH 07/13] gpu: nova-core: vgpu: add vGPU bootload Zhi Wang
2026-09-05 8:11 ` [PATCH 08/13] gpu: nova-core: vgpu: implement PluginRpc channel and config params Zhi Wang
2026-09-05 8:11 ` [PATCH 09/13] gpu: nova-core: vgpu: scrub guest framebuffer memory with CeUtils Zhi Wang
2026-09-05 8:11 ` [PATCH 10/13] gpu: nova-core: vgpu: export plugin log buffers via debugfs Zhi Wang
2026-09-05 8:11 ` [PATCH 11/13] gpu: nova-core: vgpu: export lifecycle operations to VFIO Zhi Wang
2026-09-05 8:11 ` [PATCH 12/13] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver Zhi Wang
2026-09-09 3:00 ` Alex Williamson
2026-09-11 20:39 ` Danilo Krummrich
2026-09-14 18:12 ` Alex Williamson
2026-09-14 21:36 ` Danilo Krummrich
2026-09-15 18:01 ` Alex Williamson
2026-09-15 21:19 ` Danilo Krummrich
2026-09-15 22:55 ` Zhi Wang
2026-09-16 14:17 ` Jason Gunthorpe
2026-09-16 15:38 ` Danilo Krummrich [this message]
2026-09-16 16:28 ` Jason Gunthorpe
2026-09-16 18:02 ` Danilo Krummrich
2026-09-17 1:49 ` Dave Airlie
2026-09-15 23:44 ` Dave Airlie
2026-09-16 14:25 ` Jason Gunthorpe
2026-09-05 8:11 ` [PATCH 13/13] gpu: nova-core: reserve the 48-VM WPR2 heap Zhi Wang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLGUPX9KPFKF.12CU304YAST79@kernel.org \
--to=dakr@kernel.org \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=alex@shazbot.org \
--cc=aliceryhl@google.com \
--cc=alkumar@nvidia.com \
--cc=aniketa@nvidia.com \
--cc=ankita@nvidia.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=cjia@nvidia.com \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=jgg@nvidia.com \
--cc=jhubbard@nvidia.com \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kvm@vger.kernel.org \
--cc=kwankhede@nvidia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=skolothumtho@nvidia.com \
--cc=smitra@nvidia.com \
--cc=targupta@nvidia.com \
--cc=tmgross@umich.edu \
--cc=yishaih@nvidia.com \
--cc=zhiw@nvidia.com \
--cc=zhiwang@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®