mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Alexandre Courbot" <acourbot@nvidia.com>
To: "John Hubbard" <jhubbard@nvidia.com>
Cc: "Danilo Krummrich" <dakr@kernel.org>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Gary Guo" <gary@garyguo.net>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
	"Will Pierce" <wpierce@nvidia.com>
Subject: Re: [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt
Date: Wed, 02 Sep 2026 23:33:45 +0900	[thread overview]
Message-ID: <DL4WKU4TOA0C.2FGP44BOVFHLV@nvidia.com> (raw)
In-Reply-To: <20260829013324.499542-17-jhubbard@nvidia.com>

On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
<...>
> @@ -79,13 +97,42 @@ fn probe<'bound>(
>              pdev.set_master();
>  
>              Ok(try_pin_init!(NovaCore {
> +                vectors: crate::irq::alloc_vectors(pdev, crate::irq::gsp::GSP_SUBTREE.into())?,
> +                // SAFETY: `vectors` is initialized above, lives at a pinned stable address, and
> +                // is dropped after all fields that use `vectors_ref` (struct field drop order).
> +                vectors_ref: unsafe { &*core::ptr::from_ref(vectors.as_ref().get_ref()) },
>                  bar: pdev.iomap_region_sized::<BAR0_SIZE>(0, c"nova-core/bar0")?,
>                  // TODO: Use `&bar` self-referential pin-init syntax once available.
>                  //
>                  // SAFETY: `bar` is initialized before this expression is evaluated
> -                // (`try_pin_init!()` initializes fields in declaration order), lives at a pinned
> -                // stable address, and is dropped after `gpu` (struct field drop order).
> -                gpu <- Gpu::new(pdev, unsafe { &*core::ptr::from_ref(bar) }),
> +                // (`try_pin_init!()` initializes fields in the order they appear here), lives at a
> +                // pinned stable address, and is dropped after `gpu` (struct field drop order).
> +                gpu <- Gpu::new(pdev, unsafe { &*core::ptr::from_ref(bar) }, vectors_ref),
> +                // Quiesce the interrupt tree before registering the handler below.
> +                _: {
> +                    // SAFETY: as for the `bar` borrow above.
> +                    let bar = unsafe { &*core::ptr::from_ref(bar) };
> +                    crate::irq::gsp::quiesce(bar, gpu.chipset(), vectors_ref.irq_type());
> +                },
> +                // Register the permanent GSP SWGEN0 handler, which enables the interrupt.
> +                //
> +                // SAFETY: `bar` and `vectors` are initialized and pinned (see above). `_gsp_irq`
> +                // is declared before `vectors` in the struct, so it is dropped first, ensuring
> +                // `free_irq` runs before the vectors are freed. The registration is stored in
> +                // `NovaCore` and never leaked.
> +                _gsp_irq <- unsafe {
> +                    GspIrq::new(
> +                        pdev,
> +                        vectors_ref,
> +                        &*core::ptr::from_ref(bar),
> +                        gpu.cmdq(),
> +                        gpu.chipset(),
> +                    )
> +                },
> +                // Drain the messages the GSP posted during boot, before relying on the interrupt.
> +                _: {
> +                    gpu.cmdq().drain()?;
> +                },

Can't these last 3 blocks (quiesce, _gsp_irq, and cmdq drain) be moved
inside `Gpu`? It seems to make sense in terms of ownership, as the `Gpu`
would own its interrupt handler, and things would be much cleaner as
well: there would be no need for the unsafe `bar` lifetime conversion,
`vectors_ref`, or the `chipset` and `cmdq` accessor methods.

Just gave it a quick try locally and it builds fine (with a net -20
LoCs), and AFAICT drop order is also preserved.

Pushing a bit further I could also put `vectors` into `Gpu`, which again
makes sense to me ownership-wise (because the set of interrupts we want
to serve might depend on e.g. the GPU architecture). The only drawback
is that I had to reintroduce `vectors_ref`, but that's a small and
temporary hack. I'd say this belongs in `gpu.rs` as well.

(after looking some more at the code) Ok, I'm now completely convinced
this belongs here. We could put these blocks right after `gsp_resources`
(which boots the GSP), with the benefit that the GSP interrupts will be
working to build `gsp_static_info`, which is obtained by sending a
regular GSP message! Right now we are still polling to build it, but
with the IRQ handler ready we could just wait for the signal to read the
reply. I am not saying this should be done for this series (let's do it
as a follow-up), but this is to illustrate that this is where the IRQ
setup should be done, not in `driver.rs`.

<...>
> +/// Clears the interrupt state that GSP boot left behind.
> +///
> +/// Disables every vector in every implemented leaf, clears the falcon's SWGEN0 latch, clears the
> +/// tree's pending bits, and rearms PCI interrupt delivery. On return no vector is enabled, so the
> +/// tree delivers nothing.
> +pub(crate) fn quiesce(bar: Bar0<'_>, chipset: Chipset, irq_type: pci::IrqType) {
> +    let tree = Tree::new(bar, chipset, irq_type, GSP_SUBTREE.into());
> +    tree.disable_all_leaves();
> +    // GSP boot consumes its notifications by polling the queue, which leaves SWGEN0 latched.
> +    // Clear it before the tree drain below, so the drain clears the tree state the clear sets.
> +    // Messages already posted raise no interrupt of their own, and the caller's queue drain
> +    // covers them.
> +    GspFalcon::clear_swgen0_intr(bar);
> +    tree.drain();
> +    // The `TOP_EN` cycle in `drain` is the rearm for the two enable-cycle methods, but pre-Hopper
> +    // MSI rearms through a configuration-space write instead. An interrupt delivered before probe
> +    // leaves delivery un-armed on that path, with no handler to have rearmed it.
> +    tree.rearm_pci_irq(GSP_SUBTREE);
> +}

The `clear_swgen0_intr` bit is interesting - note that we already do it
in `Gpu::new` (I had no idea why, now I understand! :)), so that
vestigial one can be removed (which should also simplify `falcon/gsp.rs`
a bit).

It also looks like this function could become a method of
`SubtreeVectors` that resets the tree covered by the allocation.

> +
> +/// Threaded IRQ handler for the GSP SWGEN0 event.
> +///
> +/// The top half clears the GIN leaf and reads the falcon SWGEN0 latch. The IRQ thread drains the
> +/// GSP-to-CPU message queue, which takes the command-queue lock.
> +#[pin_data]
> +pub(crate) struct GspInterrupt<'a> {

This type doesn't need `#[pin_data]`. Its constructor can also return
just `Self` - you will just need to wrap its call as a parameter of
`irq::ThreadedRegistration::new` into an `Ok(...)` to make it happy, but
it's simpler overall.

I'll stop here for this revision - I suppose there are more minor
things, but it will be easier to discover them with the bigger cleanups
applied.

  parent reply	other threads:[~2026-09-02 14:33 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-29  1:22 [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:22 ` [PATCH v2 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 03/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-08-29  1:22 ` [PATCH v2 04/15] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-08-31 14:24   ` Alexandre Courbot
2026-09-01 13:16   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-01  1:15   ` Alexandre Courbot
2026-08-29  1:25 ` [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:35   ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-01  7:03   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 07/15] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-01 12:52   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them John Hubbard
2026-08-31  5:06   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 09/15] gpu: nova-core: match GSP RPC replies by sequence, not just function John Hubbard
2026-08-31  1:09   ` Alexandre Courbot
2026-08-31  4:33     ` John Hubbard
2026-08-31 22:18       ` John Hubbard
2026-08-31 22:46         ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-08-31  5:35   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-08-31  6:04   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-01 14:54   ` Alexandre Courbot
2026-09-01 15:08     ` Danilo Krummrich
2026-09-02 14:33   ` Alexandre Courbot [this message]
2026-09-03  3:06     ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 13/15] gpu: nova-core: retrigger the GSP falcon and clear every latched cause John Hubbard
2026-09-02 15:00   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 14/15] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-08-29  1:33 ` [PATCH v2 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
2026-09-02 15:07   ` Alexandre Courbot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DL4WKU4TOA0C.2FGP44BOVFHLV@nvidia.com \
    --to=acourbot@nvidia.com \
    --cc=a.hindborg@kernel.org \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=wpierce@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®