mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "John Hubbard" <jhubbard@nvidia.com>,
	"Danilo Krummrich" <dakr@kernel.org>,
	"Alexandre Courbot" <acourbot@nvidia.com>
Cc: "Alistair Popple" <apopple@nvidia.com>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Benno Lossin" <lossin@kernel.org>, "Zhi Wang" <zhiw@nvidia.com>,
	"David Airlie" <airlied@gmail.com>,
	"Alexandre Courbot" <acourbot@nvidia.com>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Will Pierce" <wpierce@nvidia.com>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Trevor Gross" <tmgross@umich.edu>, "Gary Guo" <gary@garyguo.net>,
	nova-gpu@lists.linux.dev,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Alice Ryhl" <aliceryhl@google.com>,
	LKML <linux-kernel@vger.kernel.org>,
	"Danilo Krummrich" <dakr@kernel.org>,
	"Miguel Ojeda" <ojeda@kernel.org>
Subject: Re: [PATCH v3 12/14] gpu: nova-core: drive GSP events with the SWGEN0 interrupt
Date: Thu, 03 Sep 2026 03:28:14 +0000	[thread overview]
Message-ID: <20260903032815.6BDA71F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260903031514.1515905-13-jhubbard@nvidia.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Unserviceable GSP falcon interrupts cause an infinite interrupt storm and CPU soft lockup.
--

commit e6648318ceeb9312972b220805b21b9370c7409e
Author: John Hubbard <jhubbard@nvidia.com>

gpu: nova-core: drive GSP events with the SWGEN0 interrupt
    
This commit implements threaded IRQ handling for the GSP SWGEN0 event in
nova-core. It replaces polling with an interrupt handler that drains the
GSP-to-CPU message queue and clears the GIN leaf in hardirq context.

> diff --git a/drivers/gpu/nova-core/irq/gsp.rs b/drivers/gpu/nova-core/irq/gsp.rs
> new file mode 100644
> index 0000000000000..174488565f790
> --- /dev/null
> +++ b/drivers/gpu/nova-core/irq/gsp.rs

[ ... ]

> +impl irq::ThreadedHandler for GspInterrupt<'_> {
> +    /// Top half: clears the GIN leaf, takes every falcon cause pending for the host, and rearms
> +    /// PCI interrupt delivery.
> +    fn handle(&self) -> irq::ThreadedIrqReturn {
> +        let bar = self.bar;
> +
> +        // Only service our own vector: require the GSP bit in the leaf and clear just that bit, so
> +        // a co-pending vector in the same leaf stays pending for whoever services it. The subtree
> +        // stays enabled, so there is no whole-tree disable and enable.
> +        let leaf = self.tree.read_pending(GSP_INTR_0_VECTOR.leaf_index());
> +        if !leaf.vectors().contains(GSP_INTR_0_VECTOR.leaf_mask()) {
> +            // Nothing to service, but nova-core is the only consumer of this PCI interrupt, so
> +            // skipping the rearm here would silence every later interrupt as well.
> +            self.tree.rearm_pci_irq(GSP_SUBTREE);
> +            return irq::ThreadedIrqReturn::None;
> +        }
> +        leaf.clear_vectors(GSP_INTR_0_VECTOR.leaf_mask());
> +
> +        let status = GspFalcon::take_host_intr(bar, self.chipset);
> +
> +        // A cause left latched holds the falcon's host-routed set non-empty, and the falcon
> +        // signals the tree only on a transition of that set, so no later SWGEN0 would signal.
> +        let unserviceable = status.with_swgen0(false);
> +        if unserviceable.into_raw() != 0 {
> +            // nova-core has no recovery path for a cause other than a posted message, for example
> +            // a HALT from a GSP crash, so report it rather than discarding it.
> +            dev_err!(
> +                &self.dev,
> +                "unserviceable GSP falcon interrupt, IRQSTAT {:#x}\n",
> +                status.into_raw()
> +            );
> +            GspFalcon::clear_intr(bar, unserviceable);
> +        }
> +
> +        // The leaf clear above consumed the tree's record of this interrupt, and the falcon signals
> +        // the tree only on a transition of its host-routed causes, so a cause that arrived while
> +        // this handler ran would never reach the CPU. Re-emit to supply that transition.
> +        GspFalcon::retrigger_intr(bar, self.chipset);
> +
> +        // Delivery resumes only after this, so it must happen on every path that services the
> +        // vector, including the fault path above.
> +        self.tree.rearm_pci_irq(GSP_SUBTREE);

[Severity: High]
Does this lead to an interrupt storm if the GSP remains in a fault state like a
HALT or WDT timeout?

If there is no recovery path to reset the GSP, wouldn't the unserviceable
interrupt immediately re-latch after GspFalcon::clear_intr() is called?

The subsequent calls to GspFalcon::retrigger_intr() and self.tree.rearm_pci_irq()
would then force a new edge to the GIN tree and re-enable PCI delivery,
potentially trapping the CPU in an endless loop servicing the same unserviceable
interrupt.

Could the GIN leaf source be disabled here instead when encountering an
unrecoverable fault?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260903031514.1515905-1-jhubbard@nvidia.com?part=12

  reply	other threads:[~2026-09-03  3:28 UTC|newest]

Thread overview: 30+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03  3:14 [PATCH v3 00/14] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-03  3:15 ` [PATCH v3 01/14] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-03  3:15 ` [PATCH v3 02/14] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-03  3:15 ` [PATCH v3 03/14] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-09-05  1:39   ` Alexandre Courbot
2026-09-03  3:15 ` [PATCH v3 04/14] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-03  3:15 ` [PATCH v3 05/14] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-05  6:11   ` Alexandre Courbot
2026-09-03  3:15 ` [PATCH v3 06/14] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-05 13:55   ` Alexandre Courbot
2026-09-03  3:15 ` [PATCH v3 07/14] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-03  3:29   ` sashiko-bot
2026-09-03  3:57     ` John Hubbard
2026-09-03  3:15 ` [PATCH v3 08/14] gpu: nova-core: log GSP events instead of discarding them John Hubbard
2026-09-03  3:15 ` [PATCH v3 09/14] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-09-04 10:53   ` Alexandre Courbot
2026-09-04 11:17     ` Gary Guo
2026-09-04 13:45       ` Alexandre Courbot
2026-09-03  3:15 ` [PATCH v3 10/14] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-09-04 11:13   ` Alexandre Courbot
2026-09-04 11:26     ` Gary Guo
2026-09-04 13:32       ` Alexandre Courbot
2026-09-04 13:41         ` Gary Guo
2026-09-03  3:15 ` [PATCH v3 11/14] gpu: nova-core: add the falcon interrupt status and routing registers John Hubbard
2026-09-03  3:15 ` [PATCH v3 12/14] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-03  3:28   ` sashiko-bot [this message]
2026-09-03  3:55     ` John Hubbard
2026-09-04  1:53       ` John Hubbard
2026-09-03  3:15 ` [PATCH v3 13/14] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-09-03  3:15 ` [PATCH v3 14/14] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260903032815.6BDA71F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=a.hindborg@kernel.org \
    --cc=acourbot@nvidia.com \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=wpierce@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®