mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: John Hubbard <jhubbard@nvidia.com>
To: Danilo Krummrich <dakr@kernel.org>,
	Alexandre Courbot <acourbot@nvidia.com>
Cc: "Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Gary Guo" <gary@garyguo.net>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
	"John Hubbard" <jhubbard@nvidia.com>,
	"Joel Fernandes" <joelagnelf@nvidia.com>
Subject: [PATCH v5 08/15] gpu: nova-core: add an interrupt delivery self-test
Date: Tue, 29 Sep 2026 20:41:41 -0700	[thread overview]
Message-ID: <20260930034148.590687-9-jhubbard@nvidia.com> (raw)
In-Reply-To: <20260930034148.590687-1-jhubbard@nvidia.com>

A GPU interrupt can be lost in the MSI or MSI-X allocation, in the GIN
tree's enables, or in the rearm, and every one of those failures looks
the same: no interrupt arrives, and no register or log line says which
one broke.

Add a probe-time self-test, built under NOVA_CORE_SELFTESTS, that
latches the CPU doorbell vector through the GIN software trigger and
waits for a registered handler to service it. One delivery would pass
with a broken rearm, because the first message-signaled interrupt
arrives whether or not the driver rearms, so the test triggers twice and
waits for the first handler to finish before the second trigger. It runs
after GFW boot and before GSP boot, on a quiesced tree, and fails probe
unless both deliveries arrive, each finds only the doorbell pending, and
the leaf ends clear.

The doorbell has the same vector on every supported GPU, so the test
names it without asking GSP-RM. It allocates the PCI vectors for the
doorbell's subtree and releases them before returning, so under MSI-X
the delivery also exercises that subtree's table entry.

Assisted-by: LLM
Co-developed-by: Joel Fernandes <joelagnelf@nvidia.com>
Signed-off-by: Joel Fernandes <joelagnelf@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
---
 drivers/gpu/nova-core/Kconfig               |   5 +
 drivers/gpu/nova-core/driver.rs             |   5 +
 drivers/gpu/nova-core/irq.rs                |   2 +
 drivers/gpu/nova-core/irq/doorbell_test.rs  | 257 ++++++++++++++++++++
 drivers/gpu/nova-core/irq/interrupt_tree.rs |  15 ++
 drivers/gpu/nova-core/nova_core.rs          |   2 +-
 6 files changed, 285 insertions(+), 1 deletion(-)
 create mode 100644 drivers/gpu/nova-core/irq/doorbell_test.rs

diff --git a/drivers/gpu/nova-core/Kconfig b/drivers/gpu/nova-core/Kconfig
index 1934f17baa8b..2e11e46c99c7 100644
--- a/drivers/gpu/nova-core/Kconfig
+++ b/drivers/gpu/nova-core/Kconfig
@@ -24,4 +24,9 @@ config NOVA_CORE_SELFTESTS
 	help
 	  Build the driver self-tests and run them when the GPU is probed.
 
+	  If the interrupt delivery test fails, the probe fails and the driver
+	  does not bind to the GPU. A broken interrupt path would otherwise
+	  show up later as a hang, far from its cause. Every other self-test
+	  logs its failure and lets the probe continue.
+
 	  If unsure, say N.
diff --git a/drivers/gpu/nova-core/driver.rs b/drivers/gpu/nova-core/driver.rs
index fc321c6a10b0..6d45fec6d7cc 100644
--- a/drivers/gpu/nova-core/driver.rs
+++ b/drivers/gpu/nova-core/driver.rs
@@ -123,6 +123,11 @@ fn probe<'bound>(
                 // We must wait for GFW_BOOT completion before doing any significant setup on
                 // the GPU.
                 gpu::wait_gfw_boot_completion(pdev.as_ref(), bar, spec.chipset)?;
+
+                // The self-test disables and drains the whole tree, so it has to run before
+                // `Gpu::new` boots the GSP.
+                #[cfg(CONFIG_NOVA_CORE_SELFTESTS)]
+                crate::irq::doorbell_test::run_selftest(pdev, bar, spec.chipset)?;
             },
 
             // TODO: Use self-referential pin-init syntax once available.
diff --git a/drivers/gpu/nova-core/irq.rs b/drivers/gpu/nova-core/irq.rs
index b2cfe73af114..ff8b00a82442 100644
--- a/drivers/gpu/nova-core/irq.rs
+++ b/drivers/gpu/nova-core/irq.rs
@@ -9,6 +9,8 @@
 //!
 //! See `Documentation/gpu/nova/core/interrupts.rst`.
 
+#[cfg(CONFIG_NOVA_CORE_SELFTESTS)]
+pub(crate) mod doorbell_test;
 mod hal;
 pub(crate) mod interrupt_tree;
 mod regs;
diff --git a/drivers/gpu/nova-core/irq/doorbell_test.rs b/drivers/gpu/nova-core/irq/doorbell_test.rs
new file mode 100644
index 000000000000..aef9f4c7f2e8
--- /dev/null
+++ b/drivers/gpu/nova-core/irq/doorbell_test.rs
@@ -0,0 +1,257 @@
+// SPDX-License-Identifier: GPL-2.0
+// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+
+//! Interrupt delivery self-test.
+//!
+//! The test triggers the CPU doorbell vector from software, twice, and checks that each trigger
+//! reaches a registered handler. It runs during probe under `CONFIG_NOVA_CORE_SELFTESTS`.
+//!
+//! See "Self-test" in `Documentation/gpu/nova/core/interrupts.rst`.
+
+use core::pin::Pin;
+
+use kernel::{
+    device::Bound,
+    irq,
+    pci,
+    prelude::*,
+    sync::{
+        atomic::{
+            Atomic,
+            Relaxed, //
+        },
+        Completion, //
+    },
+    time, //
+};
+
+use super::interrupt_tree::{
+    GinVector,
+    LeafEnableGuard,
+    LeafMask,
+    Subtree,
+    TopEnableGuard,
+    Tree, //
+};
+
+use crate::{
+    driver::Bar0,
+    gpu::Chipset,
+    selftest_assert,
+    selftest_assert_eq, //
+};
+
+/// The CPU doorbell vector. Every supported GPU uses this number, so the test needs nothing from
+/// GSP-RM, which is not running yet.
+const DOORBELL_VECTOR: GinVector = GinVector::new::<129>();
+
+/// The only subtree that this test services.
+const DOORBELL_SUBTREE: Subtree = DOORBELL_VECTOR.subtree();
+
+/// Time allowed for each delivery to arrive.
+const DELIVERY_TIMEOUT_MS: time::Msecs = 1000;
+
+/// The self-test's interrupt handler.
+///
+/// It clears only the doorbell's bit, rearms delivery, and never walks the tree. A missing rearm
+/// shows up as a timeout on the second delivery.
+#[pin_data]
+struct DoorbellTestHandler<'a> {
+    tree: &'a Tree<'a>,
+    /// Completed by the first delivery.
+    #[pin]
+    first: Completion,
+    /// Completed by the second delivery.
+    #[pin]
+    second: Completion,
+    /// Deliveries that found the doorbell bit set.
+    irq_count: Atomic<u32>,
+    /// The doorbell leaf's pending bits, as read by the first delivery.
+    first_pending: Atomic<u32>,
+    /// The doorbell leaf's pending bits, as read by the second delivery.
+    second_pending: Atomic<u32>,
+}
+
+impl irq::Handler for DoorbellTestHandler<'_> {
+    fn handle(&self) -> irq::IrqReturn {
+        let leaf = self.tree.read_pending(DOORBELL_VECTOR.leaf_index());
+        let pending = leaf.vectors();
+        if !pending.contains(DOORBELL_VECTOR.leaf_mask()) {
+            self.tree.rearm_pci_irq(DOORBELL_SUBTREE);
+            return irq::IrqReturn::None;
+        }
+        leaf.clear_vectors(DOORBELL_VECTOR.leaf_mask());
+        // Rearm before completing, since the waiting thread triggers the next doorbell as soon as
+        // it wakes.
+        self.tree.rearm_pci_irq(DOORBELL_SUBTREE);
+
+        match self.irq_count.fetch_add(1, Relaxed) {
+            0 => {
+                self.first_pending.store(pending.into_raw(), Relaxed);
+                self.first.complete_all();
+            }
+            1 => {
+                self.second_pending.store(pending.into_raw(), Relaxed);
+                self.second.complete_all();
+            }
+            _ => (),
+        }
+
+        irq::IrqReturn::Handled
+    }
+}
+
+/// The self-test's handler registration and the enables that deliver to it.
+///
+/// Drops in the order that "Enabling the GSP event" in
+/// `Documentation/gpu/nova/core/interrupts.rst` requires: the vector is disabled, then the
+/// handler is freed, then the subtree is disabled.
+struct SelftestResources<'a, 'r> {
+    _leaf_guard: LeafEnableGuard<'a>,
+    reg: Pin<KBox<irq::Registration<'r, DoorbellTestHandler<'a>>>>,
+    _top_guard: TopEnableGuard<'a>,
+}
+
+impl<'a> SelftestResources<'a, '_> {
+    fn handler(&self) -> &DoorbellTestHandler<'a> {
+        self.reg.handler()
+    }
+
+    /// Disables the doorbell vector and waits for a handler in flight on another CPU to finish.
+    ///
+    /// The handler's counters and the leaf's pending bits are final on return.
+    fn quiesce_source(&self) {
+        self.handler()
+            .tree
+            .disable_leaf(DOORBELL_VECTOR.leaf_index(), DOORBELL_VECTOR.leaf_mask());
+        self.reg.synchronize();
+    }
+}
+
+/// Runs the interrupt delivery self-test.
+///
+/// Call this only during probe, before GSP boot: it disables every vector in the tree and clears
+/// every pending bit. On return, the doorbell's subtree is disabled at `TOP`, and the test's PCI
+/// vectors and handler are released.
+///
+/// # Errors
+///
+/// `EINVAL` if `chipset` does not implement the doorbell's subtree. `ETIMEDOUT` if a delivery
+/// does not arrive within [`DELIVERY_TIMEOUT_MS`]. `EIO` if a self-test assertion fails.
+/// Otherwise the error from allocating the PCI vectors or registering the handler.
+pub(crate) fn run_selftest(pdev: &pci::Device<Bound>, bar: Bar0<'_>, chipset: Chipset) -> Result {
+    let tree = Tree::new(pdev, bar, chipset, DOORBELL_SUBTREE.into())?;
+    let tree_ref = &tree;
+    let request = tree.request_for(DOORBELL_SUBTREE)?;
+    let doorbell = DOORBELL_VECTOR.leaf_index();
+    let doorbell_mask = DOORBELL_VECTOR.leaf_mask();
+
+    dev_info!(
+        pdev,
+        "interrupt self-test: starting on vector {}, subtree {}, with {:?}\n",
+        DOORBELL_VECTOR.into_raw(),
+        DOORBELL_SUBTREE.index(),
+        tree.msi_type,
+    );
+
+    // GFW boot can leave vectors enabled and pending. Registering a handler unmasks the PCI
+    // interrupt, and they would be delivered to a handler that services only the doorbell.
+    tree.reset();
+
+    // A delivery proves nothing unless the doorbell bit starts out clear.
+    let pre_pending = tree.read_pending(doorbell).vectors();
+    selftest_assert!(
+        pdev,
+        !pre_pending.contains(doorbell_mask),
+        "vector {} already pending, leaf[{}] is {:#x}",
+        DOORBELL_VECTOR.into_raw(),
+        doorbell.get(),
+        pre_pending.into_raw()
+    );
+
+    let handler_init = try_pin_init!(DoorbellTestHandler {
+        tree: tree_ref,
+        first <- Completion::new(),
+        second <- Completion::new(),
+        irq_count: Atomic::new(0),
+        first_pending: Atomic::new(0),
+        second_pending: Atomic::new(0),
+    }? Error);
+
+    // Registration must precede any enable, or a delivery reaches no handler.
+    let reg = KBox::pin_init(
+        // SAFETY: this registration is dropped before the enclosing function returns, so its
+        // `Drop`, which calls `free_irq()`, always runs.
+        unsafe {
+            irq::Registration::new(
+                request,
+                irq::Flags::TRIGGER_NONE,
+                c"nova-core-selftest",
+                handler_init,
+            )
+        },
+        GFP_KERNEL,
+    )?;
+
+    let resources = SelftestResources {
+        _leaf_guard: tree.enable_leaf_guarded(doorbell, doorbell_mask),
+        _top_guard: tree.enable_top_guarded(),
+        reg,
+    };
+    let handler = resources.handler();
+
+    tree.trigger(DOORBELL_VECTOR)?;
+    let mut completed = handler
+        .first
+        .wait_for_completion_timeout(time::msecs_to_jiffies(DELIVERY_TIMEOUT_MS))
+        .is_some();
+
+    // The second trigger waits for the first delivery, or the two could coalesce.
+    if completed {
+        tree.trigger(DOORBELL_VECTOR)?;
+        completed = handler
+            .second
+            .wait_for_completion_timeout(time::msecs_to_jiffies(DELIVERY_TIMEOUT_MS))
+            .is_some();
+    }
+
+    resources.quiesce_source();
+
+    let count = handler.irq_count.load(Relaxed);
+    let first_pending = LeafMask::from_raw(handler.first_pending.load(Relaxed));
+    let second_pending = LeafMask::from_raw(handler.second_pending.load(Relaxed));
+    let residual = tree.read_pending(doorbell).vectors();
+
+    if !completed {
+        dev_err!(
+            pdev,
+            "interrupt self-test: only {} of 2 deliveries arrived within {} ms\n",
+            count,
+            DELIVERY_TIMEOUT_MS,
+        );
+        return Err(ETIMEDOUT);
+    }
+
+    selftest_assert_eq!(pdev, count, 2, "delivery count");
+
+    // Every other vector in the leaf is disabled and was drained, so require the exact mask.
+    selftest_assert_eq!(pdev, first_pending, doorbell_mask, "first delivery");
+    selftest_assert_eq!(pdev, second_pending, doorbell_mask, "second delivery");
+    selftest_assert!(
+        pdev,
+        !residual.contains(doorbell_mask),
+        "vector {} still pending, leaf[{}] is {:#x}",
+        DOORBELL_VECTOR.into_raw(),
+        doorbell.get(),
+        residual.into_raw()
+    );
+
+    dev_info!(
+        pdev,
+        "interrupt self-test: passed, subtree {}, {} deliveries\n",
+        DOORBELL_SUBTREE.index(),
+        count,
+    );
+
+    Ok(())
+}
diff --git a/drivers/gpu/nova-core/irq/interrupt_tree.rs b/drivers/gpu/nova-core/irq/interrupt_tree.rs
index 6e97828bbd42..4bb27cc8b6cd 100644
--- a/drivers/gpu/nova-core/irq/interrupt_tree.rs
+++ b/drivers/gpu/nova-core/irq/interrupt_tree.rs
@@ -444,6 +444,21 @@ pub(super) fn read_pending(&self, leaf: LeafIndex) -> LeafPending<'a> {
         }
     }
 
+    /// Sets the pending bit of `vector` in its leaf, as if the vector's source had raised it.
+    ///
+    /// # Errors
+    ///
+    /// `EINVAL` if this tree does not implement `vector`.
+    #[cfg_attr(not(CONFIG_NOVA_CORE_SELFTESTS), expect(dead_code))]
+    pub(super) fn trigger(&self, vector: GinVector) -> Result {
+        vector.validate(self.hal.leaf_count())?;
+        self.bar.write_reg(
+            NV_VIRTUAL_FUNCTION_PRIV_CPU_INTR_LEAF_TRIGGER::zeroed().with_vector(vector),
+        );
+
+        Ok(())
+    }
+
     /// Disables every vector in every implemented leaf, including the subtrees that nova-core does
     /// not service. Call this only during probe.
     pub(super) fn disable_all_leaves(&self) {
diff --git a/drivers/gpu/nova-core/nova_core.rs b/drivers/gpu/nova-core/nova_core.rs
index 8202c4982efa..8c0761dedda4 100644
--- a/drivers/gpu/nova-core/nova_core.rs
+++ b/drivers/gpu/nova-core/nova_core.rs
@@ -18,7 +18,7 @@
 mod fsp;
 mod gpu;
 mod gsp;
-#[expect(dead_code)]
+#[cfg_attr(not(CONFIG_NOVA_CORE_SELFTESTS), expect(dead_code))]
 mod irq;
 mod mctp;
 mod mm;
-- 
2.55.0


  parent reply	other threads:[~2026-09-30  3:43 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  3:41 [PATCH v5 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-30  3:41 ` [PATCH v5 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-30  3:41 ` [PATCH v5 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-30  3:59   ` sashiko-bot
2026-09-30  3:41 ` [PATCH v5 03/15] gpu: nova-core: add the GIN vector, leaf and subtree types John Hubbard
2026-09-30  3:41 ` [PATCH v5 04/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-30  3:41 ` [PATCH v5 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-30  3:41 ` [PATCH v5 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-30  3:41 ` [PATCH v5 07/15] gpu: nova-core: wait for GFW boot in probe, not in the Gpu constructor John Hubbard
2026-09-30  3:41 ` John Hubbard [this message]
2026-09-30  3:41 ` [PATCH v5 09/15] gpu: nova-core: log GSP events instead of discarding them John Hubbard
2026-09-30  3:41 ` [PATCH v5 10/15] gpu: nova-core: return ENOMSG for an unmatched GSP message John Hubbard
2026-09-30  3:41 ` [PATCH v5 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-09-30  3:41 ` [PATCH v5 12/15] gpu: nova-core: add the falcon interrupt registers and HAL methods John Hubbard
2026-09-30  3:41 ` [PATCH v5 13/15] gpu: nova-core: service GSP events from the SWGEN0 interrupt John Hubbard
2026-09-30  3:41 ` [PATCH v5 14/15] gpu: nova-core: add KUnit tests for the interrupt tree John Hubbard
2026-09-30  3:41 ` [PATCH v5 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930034148.590687-9-jhubbard@nvidia.com \
    --to=jhubbard@nvidia.com \
    --cc=a.hindborg@kernel.org \
    --cc=acourbot@nvidia.com \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=joelagnelf@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®