mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
@ 2026-10-09  5:05 Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
                   ` (3 more replies)
  0 siblings, 4 replies; 11+ messages in thread
From: Naman Jain @ 2026-10-09  5:05 UTC (permalink / raw)
  To: linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Michael Kelley, linux-hyperv, linux-kernel

Hi,

On systems with several fast NVMe controllers, completion interrupts can
keep returning to the same CPUs faster than scheduled work can run. Each
handler may drain only a small number of completions, but the combined
interrupt stream can still prevent scheduler and watchdog progress.

This series keeps the existing hardirq completion path for normal I/O.
For an I/O queue with a dedicated MSI-X vector, if an interrupt arrives
with a completion pending while need_resched() is set, mask that vector
and hand the queue to irq_poll. The poller drains bounded batches and
returns the queue to interrupt mode as soon as it catches up.

The completion-path transitions are:

                         +-----------------+
                         | normal IRQ mode |
                         +-----------------+
                                  |
                   CQE pending and need_resched()
                                  |
                                  v
                     mask vector, schedule irq_poll
                                  |
                                  v
                +--------------------------------+
          +---->| drain up to one poll budget    |
          |     +--------------------------------+
          |              |                 |
          |         full budget        short batch
          |              |                 |
          |              v                 v
          |     remain on poll list   irq_poll_complete()
          |              |           mark handoff allowance
          |     generic irq_poll      clear poll ownership
          |     invokes next pass     enable vector
          |              |                 |
          +--------------+                 v
                                  +-----------------+
                                  | normal IRQ mode |
                                  +-----------------+

A full-budget return deliberately stays on the irq_poll list. The generic
irq_poll core invokes another bounded pass. Eventually a pass consumes
less than its budget, takes the short-batch path, and restores interrupt
mode.

An MSI-X message may have been latched while the vector was masked, but
not every MSI-X hierarchy can report pending state. The handoff allowance
therefore does not claim that an interrupt is known to be stale. It lets
exactly one empty interrupt after unmasking return IRQ_HANDLED. If a new
completion is present, the normal completion path handles it.

The core idea is simple: when a hardirq sees need_resched(), it masks the
interrupt and offloads completion processing to irq_poll. The interrupt is
unmasked once the CQ catches up. Everything else handles queue and
controller state changes while irq_poll is active, because its asynchronous
callback can outlive the hardirq that scheduled it.

The lifecycle transitions are:

  normal IRQ mode / irq_poll
              |
              | set PAUSED, join IRQ and irq_poll, mask vector
              v
          +--------+
          | paused |---- temporary ----> resume after controller is ready
          +--------+                      clear PAUSED, restore IRQ mode
              |
              | permanent stop
              | clear INIT, balance owned mask, synchronize replay
              v
        free registered IRQ

INIT means the queue has an initialized irq_poll instance and a dedicated
MSI-X vector. IRQ_POLL owns completion processing, while MASKED records the
matching disable_irq_nosync(). PAUSED blocks IRQ handling and new admission.
HANDOFF allows one empty interrupt after unmasking. A temporary pause keeps
MASKED set for resume; permanent stop keeps PAUSED set through IRQ removal.

Admin, shared, legacy, and threaded interrupt paths are unchanged. Explicit
polling never enters irq_poll; it only disables bottom halves while holding
the shared CQ polling lock. The series adds no timer, rate sampler,
workqueue, sysfs knob, module parameter, or IRQ thread.

Why not use nvme.use_threaded_interrupts=1 instead?

Threaded mode installs nvme_irq_check() as the primary handler and
nvme_irq() as the IRQ thread. When the primary handler finds a pending
completion, it returns IRQ_WAKE_THREAD, so completion queue draining runs
through a schedulable task rather than the normal hardirq fast path.

That remains a useful opt-in mode and this series does not change it. On
the affected VM it prevented the lockup, but it also reduced peak
throughput. This series has a narrower goal: retain direct hardirq
completion while the CPU is keeping up, and use bounded irq_poll work
only after the scheduler has already asserted need_resched().

The distinction is not that irq_poll is always faster than an IRQ thread.
It is that the normal path is left alone, while the fallback creates a
scheduling boundary only after a reschedule request is observed.

need_resched() is used here as a best-effort pressure signal, not as an
interrupt-flood detector or an elapsed-time guarantee. It is transient,
and irq_poll may execute one bounded callback before the scheduler runs.
The narrower claim is that, once an NVMe hardirq observes an already
pending reschedule request, it stops unbounded CQ draining in hardirq.

Previous NVMe attempts made the fallback the normal high-load path. An
all-irq_poll RFC in 2016 reported an 8-10% single-core IOPS loss and a
measurable low-queue-depth latency cost. A hardirq-first, threaded
overflow series in 2019 fixed the Azure lockup, but Azure testing
reported a substantial throughput loss.

Relevant discussions from the past around this problem:

  https://lore.kernel.org/linux-nvme/1475660534-16681-1-git-send-email-sagi@grimberg.me/
  https://lore.kernel.org/lkml/1566281669-48212-1-git-send-email-longli@linuxonhyperv.com/
  https://lore.kernel.org/linux-nvme/20191209175622.1964-1-kbusch@kernel.org/

The first patch selects IRQ_POLL. The second makes existing task-context
CQ polling safe against the new softirq user. The third adds the NVMe
queue handoff and lifecycle handling.

The RFC starts with a per-queue weight of 256, matching the existing
global irq_poll budget. This is a starting point, not a claim that 256 is
the optimal NVMe value.

Testing included:

  - full x86_64 kernel build and strict checkpatch
  - four-controller QEMU fio, full-disable S3, controller reset, and
    post-reset I/O
  - QEMU lockdep, prove-locking, and atomic-context validation
  - ARM64 Hyper-V VM with four 14-queue NVMe data controllers
  - natural irq_poll activation with threaded interrupts disabled
  - matching default-hardirq, IRQ-poll fallback, and threaded-IRQ comparison
  - four controller resets followed by successful reads

Read-only 4 KiB random-read results follow. "Default hardirq" uses the
unpatched base. The other modes use the same kernel built from this series:
"IRQ-poll fallback" uses the default nvme.use_threaded_interrupts=0, while
"threaded IRQs" uses nvme.use_threaded_interrupts=1. Low-load values are
medians of two 30-second runs. High-load fallback and threaded values are
medians of two and three 90-second runs respectively; the default-hardirq
run locked up before it could complete.

                              default       IRQ-poll      threaded
                              hardirq       fallback      IRQs
  1 disk, QD1
    IOPS                      24.42K        24.96K        21.24K
    mean completion latency  38.94 us      38.17 us      44.60 us
    p99 completion latency   90.6 us       90.1 us       94.2 us
    p99.9 completion latency 94.7 us       94.7 us      103.9 us

  4 disks, 128 jobs, QD256
    status                    lockup (26s)  pass           pass
    IOPS                      N/A           4.986M         3.165M
    bandwidth                 N/A          19.0 GiB/s     12.1 GiB/s
    mean completion latency   N/A           6.544 ms      10.332 ms
    p99 completion latency    N/A          14.483 ms      22.151 ms
    p99.9 completion latency  N/A          22.544 ms      24.773 ms

IRQ-poll fallback prevents hardirq starvation while avoiding the 15-37%
IOPS loss and higher latency observed with threaded interrupts, by
activating only under scheduler pressure.

Naman Jain (3):
  nvme-pci: select IRQ_POLL
  nvme-pci: make completion queue polling softirq-safe
  nvme-pci: defer completions when rescheduling is needed

 drivers/nvme/host/Kconfig |   1 +
 drivers/nvme/host/pci.c   | 295 ++++++++++++++++++++++++++++++++++++++++++++--
 2 files changed, 283 insertions(+), 13 deletions(-)

base-commit: eea3fef32a9cf36abcb5975a5a594e4135a6b026
-- 
2.43.0

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [RFC PATCH 1/3] nvme-pci: select IRQ_POLL
  2026-10-09  5:05 [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Naman Jain
@ 2026-10-09  5:05 ` Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe Naman Jain
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 11+ messages in thread
From: Naman Jain @ 2026-10-09  5:05 UTC (permalink / raw)
  To: linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Michael Kelley, linux-hyperv, linux-kernel

A later patch uses irq_poll to process NVMe completions in bounded
batches. Select IRQ_POLL so that support is available when the NVMe PCI
driver is built into the kernel or as a module.

Assisted-by: LLM
Signed-off-by: Naman Jain <namjain@linux.microsoft.com>
---
 drivers/nvme/host/Kconfig | 1 +
 1 file changed, 1 insertion(+)

diff --git a/drivers/nvme/host/Kconfig b/drivers/nvme/host/Kconfig
index 31974c7dd20c9..22164b901da85 100644
--- a/drivers/nvme/host/Kconfig
+++ b/drivers/nvme/host/Kconfig
@@ -5,6 +5,7 @@ config NVME_CORE
 config BLK_DEV_NVME
 	tristate "NVM Express block device"
 	depends on PCI && BLOCK
+	select IRQ_POLL
 	select NVME_CORE
 	help
 	  The NVM Express driver is for solid state drives directly

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe
  2026-10-09  5:05 [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
@ 2026-10-09  5:05 ` Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
  2026-10-09  5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
  3 siblings, 0 replies; 11+ messages in thread
From: Naman Jain @ 2026-10-09  5:05 UTC (permalink / raw)
  To: linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Michael Kelley, linux-hyperv, linux-kernel

A following patch processes interrupt-driven completion queues from
IRQ_POLL_SOFTIRQ. Ensure every task-context user disables bottom halves
while holding cq_poll_lock so the softirq cannot interrupt the local
owner.

This includes explicit blk-mq polling. Those queues do not use irq_poll,
but all cq_poll_lock instances share one lockdep class because they are
initialized at the same call site. Keep that class consistently
softirq-safe rather than reclassifying reused queue objects across reset.

Assisted-by: LLM
Signed-off-by: Naman Jain <namjain@linux.microsoft.com>
---
 drivers/nvme/host/pci.c | 12 ++++++------
 1 file changed, 6 insertions(+), 6 deletions(-)

diff --git a/drivers/nvme/host/pci.c b/drivers/nvme/host/pci.c
index dbcfc7ddb78aa..bd4a6461b620f 100644
--- a/drivers/nvme/host/pci.c
+++ b/drivers/nvme/host/pci.c
@@ -1673,9 +1673,9 @@ static void nvme_poll_irqdisable(struct nvme_queue *nvmeq)
 
 	irq = pci_irq_vector(pdev, nvmeq->cq_vector);
 	disable_irq(irq);
-	spin_lock(&nvmeq->cq_poll_lock);
+	spin_lock_bh(&nvmeq->cq_poll_lock);
 	nvme_poll_cq(nvmeq, NULL);
-	spin_unlock(&nvmeq->cq_poll_lock);
+	spin_unlock_bh(&nvmeq->cq_poll_lock);
 	enable_irq(irq);
 }
 
@@ -1688,9 +1688,9 @@ static int nvme_poll(struct blk_mq_hw_ctx *hctx, struct io_comp_batch *iob)
 	    !nvme_cqe_pending(nvmeq))
 		return 0;
 
-	spin_lock(&nvmeq->cq_poll_lock);
+	spin_lock_bh(&nvmeq->cq_poll_lock);
 	found = nvme_poll_cq(nvmeq, iob);
-	spin_unlock(&nvmeq->cq_poll_lock);
+	spin_unlock_bh(&nvmeq->cq_poll_lock);
 
 	return found;
 }
@@ -2086,9 +2086,9 @@ static void nvme_reap_pending_cqes(struct nvme_dev *dev)
 	int i;
 
 	for (i = dev->ctrl.queue_count - 1; i > 0; i--) {
-		spin_lock(&dev->queues[i].cq_poll_lock);
+		spin_lock_bh(&dev->queues[i].cq_poll_lock);
 		nvme_poll_cq(&dev->queues[i], NULL);
-		spin_unlock(&dev->queues[i].cq_poll_lock);
+		spin_unlock_bh(&dev->queues[i].cq_poll_lock);
 	}
 }
 

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed
  2026-10-09  5:05 [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
  2026-10-09  5:05 ` [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe Naman Jain
@ 2026-10-09  5:05 ` Naman Jain
  2026-10-09 15:30   ` Keith Busch
  2026-10-09  5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
  3 siblings, 1 reply; 11+ messages in thread
From: Naman Jain @ 2026-10-09  5:05 UTC (permalink / raw)
  To: linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Michael Kelley, linux-hyperv, linux-kernel

Under sustained I/O, a CPU can finish one NVMe interrupt and immediately
enter another. This can keep the scheduler from running even when each
interrupt completes only a small amount of work.

For dedicated MSI-X I/O queues, check need_resched() before draining the
completion queue. If the scheduler needs the CPU, mask the vector and move
completion processing to irq_poll. The poller processes bounded batches
and returns the queue to interrupt mode after it catches up.

Stop and join irq_poll before deleting I/O queues or making the controller
inaccessible during suspend or PCI error recovery. Resume it only after
the controller is operational again. This prevents deferred completion
processing from using deleted queue state or inaccessible registers.

Admin queues, shared and legacy interrupts, explicit polling queues, and
threaded interrupts continue to use their existing paths.

Assisted-by: LLM
Signed-off-by: Naman Jain <namjain@linux.microsoft.com>
---
 drivers/nvme/host/pci.c | 283 +++++++++++++++++++++++++++++++++++++++-
 1 file changed, 276 insertions(+), 7 deletions(-)

diff --git a/drivers/nvme/host/pci.c b/drivers/nvme/host/pci.c
index bd4a6461b620f..dc1865d59686c 100644
--- a/drivers/nvme/host/pci.c
+++ b/drivers/nvme/host/pci.c
@@ -13,6 +13,7 @@
 #include <linux/init.h>
 #include <linux/interrupt.h>
 #include <linux/io.h>
+#include <linux/irq_poll.h>
 #include <linux/kstrtox.h>
 #include <linux/memremap.h>
 #include <linux/mm.h>
@@ -99,6 +100,8 @@ MODULE_PARM_DESC(sgl_threshold,
 
 #define NVME_PCI_MIN_QUEUE_SIZE 2
 #define NVME_PCI_MAX_QUEUE_SIZE 4095
+#define NVME_IRQ_POLL_BUDGET	256
+
 static int io_queue_depth_set(const char *val, const struct kernel_param *kp);
 static const struct kernel_param_ops io_queue_depth_ops = {
 	.set = io_queue_depth_set,
@@ -283,6 +286,8 @@ struct nvme_queue;
 static void nvme_dev_disable(struct nvme_dev *dev, bool shutdown);
 static void nvme_delete_io_queues(struct nvme_dev *dev);
 static void nvme_update_attrs(struct nvme_dev *dev);
+static void nvme_pause_io_irq_poll(struct nvme_dev *dev);
+static void nvme_resume_io_irq_poll(struct nvme_dev *dev);
 
 struct nvme_descriptor_pools {
 	struct dma_pool *large;
@@ -369,6 +374,10 @@ struct nvme_queue {
 	spinlock_t sq_lock;
 	void *sq_cmds
 		__guarded_by(&sq_lock);
+	struct irq_poll iopoll;
+	int irq;
+	/* Serializes irq_poll pause, resume, and permanent teardown. */
+	struct mutex irq_poll_lock;
 	 /* only used for poll queues: */
 	spinlock_t cq_poll_lock ____cacheline_aligned_in_smp;
 	struct nvme_completion *cqes;
@@ -390,6 +399,29 @@ struct nvme_queue {
 #define NVMEQ_SQ_CMB		1
 #define NVMEQ_DELETE_ERROR	2
 #define NVMEQ_POLLED		3
+/*
+ * IRQ-poll state:
+ *
+ * NVMEQ_IRQ_POLL:
+ *	irq_poll owns completion processing for this queue.
+ * NVMEQ_IRQ_POLL_INIT:
+ *	The queue has an initialized irq_poll and dedicated MSI-X.
+ * NVMEQ_IRQ_POLL_HANDOFF:
+ *	Allow one empty interrupt after the vector is unmasked.
+ * NVMEQ_IRQ_POLL_PAUSED:
+ *	Block IRQ handling and new irq_poll admission.
+ * NVMEQ_IRQ_POLL_MASKED:
+ *	This state machine owns one disable_irq_nosync().
+ *
+ * A temporary pause keeps NVMEQ_IRQ_POLL_MASKED set for resume. Permanent
+ * teardown clears NVMEQ_IRQ_POLL_INIT before joining irq_poll and keeps
+ * NVMEQ_IRQ_POLL_PAUSED set through IRQ removal.
+ */
+#define NVMEQ_IRQ_POLL		4
+#define NVMEQ_IRQ_POLL_INIT	5
+#define NVMEQ_IRQ_POLL_HANDOFF	6
+#define NVMEQ_IRQ_POLL_PAUSED	7
+#define NVMEQ_IRQ_POLL_MASKED	8
 	__le32 *dbbuf_sq_db;
 	__le32 *dbbuf_cq_db;
 	__le32 *dbbuf_sq_ei;
@@ -1638,11 +1670,104 @@ static inline bool nvme_poll_cq(struct nvme_queue *nvmeq,
 	return found;
 }
 
+static int nvme_poll_cq_bounded(struct nvme_queue *nvmeq,
+				struct io_comp_batch *iob, int budget)
+{
+	int found = 0;
+
+	while (found < budget && nvme_cqe_pending(nvmeq)) {
+		dma_rmb();
+		nvme_handle_cqe(nvmeq, iob, nvmeq->cq_head);
+		nvme_update_cq_head(nvmeq);
+		found++;
+	}
+
+	return found;
+}
+
+/*
+ * A hardirq that observes need_resched() masks the vector and hands the
+ * queue to irq_poll. A full-budget pass remains scheduled. A short pass
+ * completes polling and returns the queue to interrupt mode.
+ */
+static int nvme_irq_poll(struct irq_poll *iop, int budget)
+{
+	struct nvme_queue *nvmeq =
+		container_of(iop, struct nvme_queue, iopoll);
+	struct pci_dev *pdev = to_pci_dev(nvmeq->dev->dev);
+	DEFINE_IO_COMP_BATCH(iob);
+	int found;
+
+	spin_lock(&nvmeq->cq_poll_lock);
+	if (unlikely(test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) ||
+		     READ_ONCE(pdev->error_state) != pci_channel_io_normal)) {
+		irq_poll_complete(iop);
+		spin_unlock(&nvmeq->cq_poll_lock);
+		return 0;
+	}
+	found = nvme_poll_cq_bounded(nvmeq, &iob, budget);
+	/* AER may freeze the channel while this callback is draining. */
+	if (found && READ_ONCE(pdev->error_state) == pci_channel_io_normal)
+		nvme_ring_cq_doorbell(nvmeq);
+	spin_unlock(&nvmeq->cq_poll_lock);
+
+	if (!rq_list_empty(&iob.req_list))
+		nvme_pci_complete_batch(&iob);
+
+	if (found < budget) {
+		spin_lock(&nvmeq->cq_poll_lock);
+		if (unlikely(test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) ||
+			     READ_ONCE(pdev->error_state) !=
+					pci_channel_io_normal)) {
+			irq_poll_complete(iop);
+			spin_unlock(&nvmeq->cq_poll_lock);
+			return found;
+		}
+		irq_poll_complete(iop);
+		/*
+		 * An interrupt may have been latched while the vector was
+		 * masked. Allow one empty interrupt after handing the queue
+		 * back without claiming that one is known to be pending.
+		 */
+		set_bit(NVMEQ_IRQ_POLL_HANDOFF, &nvmeq->flags);
+		clear_bit_unlock(NVMEQ_IRQ_POLL, &nvmeq->flags);
+		if (test_and_clear_bit(NVMEQ_IRQ_POLL_MASKED,
+				       &nvmeq->flags))
+			enable_irq(nvmeq->irq);
+		spin_unlock(&nvmeq->cq_poll_lock);
+	}
+
+	return found;
+}
+
 static irqreturn_t nvme_irq(int irq, void *data)
 {
 	struct nvme_queue *nvmeq = data;
 	DEFINE_IO_COMP_BATCH(iob);
 
+	if (unlikely(test_and_clear_bit(NVMEQ_IRQ_POLL_HANDOFF,
+					&nvmeq->flags))) {
+		if (!nvme_cqe_pending(nvmeq))
+			return IRQ_HANDLED;
+	}
+	if (unlikely(test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) &&
+		     (!test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags) ||
+		      test_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags) ||
+		      READ_ONCE(to_pci_dev(nvmeq->dev->dev)->error_state) !=
+				pci_channel_io_normal)))
+		return IRQ_HANDLED;
+
+	if (unlikely(in_hardirq() &&
+		     test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags) &&
+		     !test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) &&
+		     need_resched() && nvme_cqe_pending(nvmeq) &&
+		     !test_and_set_bit_lock(NVMEQ_IRQ_POLL, &nvmeq->flags))) {
+		set_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
+		disable_irq_nosync(irq);
+		irq_poll_sched(&nvmeq->iopoll);
+		return IRQ_HANDLED;
+	}
+
 	if (nvme_poll_cq(nvmeq, &iob)) {
 		if (!rq_list_empty(&iob.req_list))
 			nvme_pci_complete_batch(&iob);
@@ -1713,6 +1838,7 @@ static void nvme_pci_submit_async_event(struct nvme_ctrl *ctrl)
 static int nvme_pci_subsystem_reset(struct nvme_ctrl *ctrl)
 {
 	struct nvme_dev *dev = to_nvme_dev(ctrl);
+	bool paused = false;
 	int ret = 0;
 
 	/*
@@ -1732,6 +1858,8 @@ static int nvme_pci_subsystem_reset(struct nvme_ctrl *ctrl)
 		goto unlock;
 	}
 
+	nvme_pause_io_irq_poll(dev);
+	paused = true;
 	writel(NVME_SUBSYS_RESET, dev->bar + NVME_REG_NSSR);
 
 	if (!nvme_change_ctrl_state(ctrl, NVME_CTRL_CONNECTING) ||
@@ -1744,6 +1872,8 @@ static int nvme_pci_subsystem_reset(struct nvme_ctrl *ctrl)
 	 */
 	readl(dev->bar + NVME_REG_CSTS);
 unlock:
+	if (paused && nvme_ctrl_state(ctrl) == NVME_CTRL_LIVE)
+		nvme_resume_io_irq_poll(dev);
 	mutex_unlock(&dev->shutdown_lock);
 	return ret;
 }
@@ -2050,6 +2180,98 @@ static void nvme_free_queues(struct nvme_dev *dev, int lowest)
 	}
 }
 
+static void __nvme_pause_irq_poll(struct nvme_queue *nvmeq)
+	__must_hold(&nvmeq->irq_poll_lock)
+{
+	if (test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags))
+		return;
+
+	set_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags);
+	synchronize_irq(nvmeq->irq);
+	irq_poll_disable(&nvmeq->iopoll);
+	spin_lock_bh(&nvmeq->cq_poll_lock);
+	clear_bit(NVMEQ_IRQ_POLL, &nvmeq->flags);
+	spin_unlock_bh(&nvmeq->cq_poll_lock);
+
+	if (!test_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags) &&
+	    READ_ONCE(to_pci_dev(nvmeq->dev->dev)->error_state) ==
+			pci_channel_io_normal) {
+		/*
+		 * A PAUSED-but-unmasked handler drains normally until
+		 * disable_irq() has joined it.
+		 */
+		disable_irq(nvmeq->irq);
+		set_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
+	}
+	/* Resume or permanent teardown balances any mask we still own. */
+}
+
+static void nvme_pause_irq_poll(struct nvme_queue *nvmeq)
+{
+	mutex_lock(&nvmeq->irq_poll_lock);
+	if (test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags))
+		__nvme_pause_irq_poll(nvmeq);
+	mutex_unlock(&nvmeq->irq_poll_lock);
+}
+
+static void nvme_resume_irq_poll(struct nvme_queue *nvmeq)
+{
+	bool masked;
+
+	mutex_lock(&nvmeq->irq_poll_lock);
+	if (!test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags) ||
+	    !test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) ||
+	    READ_ONCE(to_pci_dev(nvmeq->dev->dev)->error_state) !=
+			pci_channel_io_normal)
+		goto out_unlock;
+
+	/* Consume saved mask ownership before reopening hardirq admission. */
+	masked = test_and_clear_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
+	if (masked)
+		set_bit(NVMEQ_IRQ_POLL_HANDOFF, &nvmeq->flags);
+	irq_poll_enable(&nvmeq->iopoll);
+	clear_bit_unlock(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags);
+	if (masked)
+		enable_irq(nvmeq->irq);
+out_unlock:
+	mutex_unlock(&nvmeq->irq_poll_lock);
+}
+
+static void nvme_stop_irq_poll(struct nvme_queue *nvmeq)
+{
+	mutex_lock(&nvmeq->irq_poll_lock);
+	if (!test_and_clear_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags))
+		goto out_unlock;
+
+	__nvme_pause_irq_poll(nvmeq);
+	if (READ_ONCE(to_pci_dev(nvmeq->dev->dev)->error_state) ==
+			pci_channel_io_normal &&
+	    test_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags)) {
+		/* PAUSED makes any replay harmless; join it before IRQ removal. */
+		enable_irq(nvmeq->irq);
+		synchronize_irq(nvmeq->irq);
+		clear_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
+	}
+out_unlock:
+	mutex_unlock(&nvmeq->irq_poll_lock);
+}
+
+static void nvme_pause_io_irq_poll(struct nvme_dev *dev)
+{
+	int i;
+
+	for (i = dev->ctrl.queue_count - 1; i > 0; i--)
+		nvme_pause_irq_poll(&dev->queues[i]);
+}
+
+static void nvme_resume_io_irq_poll(struct nvme_dev *dev)
+{
+	int i;
+
+	for (i = 1; i < dev->ctrl.queue_count; i++)
+		nvme_resume_irq_poll(&dev->queues[i]);
+}
+
 static void nvme_suspend_queue(struct nvme_dev *dev, unsigned int qid)
 {
 	struct nvme_queue *nvmeq = &dev->queues[qid];
@@ -2063,8 +2285,11 @@ static void nvme_suspend_queue(struct nvme_dev *dev, unsigned int qid)
 	nvmeq->dev->online_queues--;
 	if (!nvmeq->qid && nvmeq->dev->ctrl.admin_q)
 		nvme_quiesce_admin_queue(&nvmeq->dev->ctrl);
-	if (!test_and_clear_bit(NVMEQ_POLLED, &nvmeq->flags))
+	if (!test_and_clear_bit(NVMEQ_POLLED, &nvmeq->flags)) {
+		nvme_stop_irq_poll(nvmeq);
 		pci_free_irq(to_pci_dev(dev->dev), nvmeq->cq_vector, nvmeq);
+		clear_bit(NVMEQ_IRQ_POLL_HANDOFF, &nvmeq->flags);
+	}
 }
 
 static void nvme_suspend_io_queues(struct nvme_dev *dev)
@@ -2087,7 +2312,8 @@ static void nvme_reap_pending_cqes(struct nvme_dev *dev)
 
 	for (i = dev->ctrl.queue_count - 1; i > 0; i--) {
 		spin_lock_bh(&dev->queues[i].cq_poll_lock);
-		nvme_poll_cq(&dev->queues[i], NULL);
+		/* Reap CQ memory without ringing a disabled controller. */
+		nvme_poll_cq_bounded(&dev->queues[i], NULL, INT_MAX);
 		spin_unlock_bh(&dev->queues[i].cq_poll_lock);
 	}
 }
@@ -2164,6 +2390,7 @@ static int nvme_alloc_queue(struct nvme_dev *dev, int qid, int depth)
 	nvmeq->dev = dev;
 	spin_lock_init(&nvmeq->sq_lock);
 	spin_lock_init(&nvmeq->cq_poll_lock);
+	mutex_init(&nvmeq->irq_poll_lock);
 	nvmeq->cq_head = 0;
 	nvmeq->cq_phase = 1;
 	nvmeq->q_db = &dev->dbs[qid * 2 * dev->db_stride];
@@ -2183,14 +2410,36 @@ static int queue_request_irq(struct nvme_queue *nvmeq)
 {
 	struct pci_dev *pdev = to_pci_dev(nvmeq->dev->dev);
 	int nr = nvmeq->dev->ctrl.instance;
+	int ret;
+
+	/* The queue may be reused with a different IRQ layout after reset. */
+	clear_bit(NVMEQ_IRQ_POLL, &nvmeq->flags);
+	clear_bit(NVMEQ_IRQ_POLL_HANDOFF, &nvmeq->flags);
+	clear_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags);
+	clear_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
 
 	if (use_threaded_interrupts) {
 		return pci_request_irq(pdev, nvmeq->cq_vector, nvme_irq_check,
 				nvme_irq, nvmeq, "nvme%dq%d", nr, nvmeq->qid);
-	} else {
+	}
+	/*
+	 * Queue IDs map one-to-one to vectors here.  Restrict IRQ masking to
+	 * dedicated MSI-X I/O vectors; shared MSI and legacy IRQs stay on the
+	 * regular handler.
+	 */
+	if (!nvmeq->qid || !pdev->msix_enabled ||
+	    nvmeq->dev->num_vecs <= nvmeq->cq_vector ||
+	    nvmeq->cq_vector != nvmeq->qid)
 		return pci_request_irq(pdev, nvmeq->cq_vector, nvme_irq,
 				NULL, nvmeq, "nvme%dq%d", nr, nvmeq->qid);
-	}
+
+	nvmeq->irq = pci_irq_vector(pdev, nvmeq->cq_vector);
+	irq_poll_init(&nvmeq->iopoll, NVME_IRQ_POLL_BUDGET, nvme_irq_poll);
+	ret = pci_request_irq(pdev, nvmeq->cq_vector, nvme_irq,
+			      NULL, nvmeq, "nvme%dq%d", nr, nvmeq->qid);
+	if (!ret)
+		set_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags);
+	return ret;
 }
 
 static void nvme_init_queue(struct nvme_queue *nvmeq, u16 qid)
@@ -3073,6 +3322,7 @@ static int nvme_setup_io_queues(struct nvme_dev *dev)
 
 	if (dev->online_queues - 1 < dev->max_qid) {
 		nr_io_queues = dev->online_queues - 1;
+		nvme_pause_io_irq_poll(dev);
 		nvme_delete_io_queues(dev);
 		result = nvme_setup_io_queues_trylock(dev);
 		if (result)
@@ -3330,6 +3580,7 @@ static void nvme_dev_disable(struct nvme_dev *dev, bool shutdown)
 	}
 
 	nvme_quiesce_io_queues(&dev->ctrl);
+	nvme_pause_io_irq_poll(dev);
 
 	if (!dead && dev->ctrl.queue_count > 0) {
 		nvme_delete_io_queues(dev);
@@ -3948,9 +4199,14 @@ static int nvme_resume(struct device *dev)
 	if (ctrl->hmpre && nvme_setup_host_mem(ndev))
 		goto reset;
 
+	mutex_lock(&ndev->shutdown_lock);
+	nvme_resume_io_irq_poll(ndev);
+	mutex_unlock(&ndev->shutdown_lock);
 	return 0;
 reset:
-	return nvme_try_sched_reset(ctrl);
+	if (nvme_ctrl_state(ctrl) == NVME_CTRL_RESETTING)
+		return nvme_try_sched_reset(ctrl);
+	return nvme_reset_ctrl(ctrl);
 }
 
 static int nvme_suspend(struct device *dev)
@@ -3958,6 +4214,7 @@ static int nvme_suspend(struct device *dev)
 	struct pci_dev *pdev = to_pci_dev(dev);
 	struct nvme_dev *ndev = pci_get_drvdata(pdev);
 	struct nvme_ctrl *ctrl = &ndev->ctrl;
+	bool irq_poll_paused = false;
 	int ret = -EBUSY;
 
 	ndev->last_ps = U32_MAX;
@@ -3985,8 +4242,14 @@ static int nvme_suspend(struct device *dev)
 	nvme_wait_freeze(ctrl);
 	nvme_sync_queues(ctrl);
 
-	if (nvme_ctrl_state(ctrl) != NVME_CTRL_LIVE)
+	mutex_lock(&ndev->shutdown_lock);
+	if (nvme_ctrl_state(ctrl) != NVME_CTRL_LIVE) {
+		mutex_unlock(&ndev->shutdown_lock);
 		goto unfreeze;
+	}
+	nvme_pause_io_irq_poll(ndev);
+	irq_poll_paused = true;
+	mutex_unlock(&ndev->shutdown_lock);
 
 	/*
 	 * Host memory access may not be successful in a system suspend state,
@@ -4022,10 +4285,16 @@ static int nvme_suspend(struct device *dev)
 		 * Clearing npss forces a controller reset on resume. The
 		 * correct value will be rediscovered then.
 		 */
-		ret = nvme_disable_prepare_reset(ndev, true);
 		ctrl->npss = 0;
+		ret = nvme_disable_prepare_reset(ndev, true);
+		goto unfreeze;
 	}
 unfreeze:
+	if (ret && irq_poll_paused) {
+		mutex_lock(&ndev->shutdown_lock);
+		nvme_resume_io_irq_poll(ndev);
+		mutex_unlock(&ndev->shutdown_lock);
+	}
 	nvme_unfreeze(ctrl);
 	return ret;
 }

^ permalink raw reply	[flat|nested] 11+ messages in thread

* RE: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09  5:05 [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Naman Jain
                   ` (2 preceding siblings ...)
  2026-10-09  5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
@ 2026-10-09  5:54 ` Michael Kelley
  2026-10-09  6:51   ` Naman Jain
  3 siblings, 1 reply; 11+ messages in thread
From: Michael Kelley @ 2026-10-09  5:54 UTC (permalink / raw)
  To: Naman Jain, linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	linux-hyperv, linux-kernel

From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October 8, 2026 10:06 PM
>
> On systems with several fast NVMe controllers, completion interrupts can
> keep returning to the same CPUs faster than scheduled work can run. Each
> handler may drain only a small number of completions, but the combined
> interrupt stream can still prevent scheduler and watchdog progress.

See this recent proposal [1] that sounds like it is addressing the same or a
similar issue. And there is this [2] more global approach. It's worthwhile to read
through the discussion on both threads. I haven't done a detailed comparison
of either vs. your proposal.

Michael

[1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@bytedance.com/
[2] https://lore.kernel.org/lkml/20260819124341.4185621-1-lrizzo@google.com/

>
> This series keeps the existing hardirq completion path for normal I/O.
> For an I/O queue with a dedicated MSI-X vector, if an interrupt arrives
> with a completion pending while need_resched() is set, mask that vector
> and hand the queue to irq_poll. The poller drains bounded batches and
> returns the queue to interrupt mode as soon as it catches up.
>
> The completion-path transitions are:
>
>                          +-----------------+
>                          | normal IRQ mode |
>                          +-----------------+
>                                   |
>                    CQE pending and need_resched()
>                                   |
>                                   v
>                      mask vector, schedule irq_poll
>                                   |
>                                   v
>                 +--------------------------------+
>           +---->| drain up to one poll budget    |
>           |     +--------------------------------+
>           |              |                 |
>           |         full budget        short batch
>           |              |                 |
>           |              v                 v
>           |     remain on poll list   irq_poll_complete()
>           |              |           mark handoff allowance
>           |     generic irq_poll      clear poll ownership
>           |     invokes next pass     enable vector
>           |              |                 |
>           +--------------+                 v
>                                   +-----------------+
>                                   | normal IRQ mode |
>                                   +-----------------+
>
> A full-budget return deliberately stays on the irq_poll list. The generic
> irq_poll core invokes another bounded pass. Eventually a pass consumes
> less than its budget, takes the short-batch path, and restores interrupt
> mode.
>
> An MSI-X message may have been latched while the vector was masked, but
> not every MSI-X hierarchy can report pending state. The handoff allowance
> therefore does not claim that an interrupt is known to be stale. It lets
> exactly one empty interrupt after unmasking return IRQ_HANDLED. If a new
> completion is present, the normal completion path handles it.
>
> The core idea is simple: when a hardirq sees need_resched(), it masks the
> interrupt and offloads completion processing to irq_poll. The interrupt is
> unmasked once the CQ catches up. Everything else handles queue and
> controller state changes while irq_poll is active, because its asynchronous
> callback can outlive the hardirq that scheduled it.
>
> The lifecycle transitions are:
>
>   normal IRQ mode / irq_poll
>               |
>               | set PAUSED, join IRQ and irq_poll, mask vector
>               v
>           +--------+
>           | paused |---- temporary ----> resume after controller is ready
>           +--------+                      clear PAUSED, restore IRQ mode
>               |
>               | permanent stop
>               | clear INIT, balance owned mask, synchronize replay
>               v
>         free registered IRQ
>
> INIT means the queue has an initialized irq_poll instance and a dedicated
> MSI-X vector. IRQ_POLL owns completion processing, while MASKED records the
> matching disable_irq_nosync(). PAUSED blocks IRQ handling and new admission.
> HANDOFF allows one empty interrupt after unmasking. A temporary pause keeps
> MASKED set for resume; permanent stop keeps PAUSED set through IRQ removal.
>
> Admin, shared, legacy, and threaded interrupt paths are unchanged. Explicit
> polling never enters irq_poll; it only disables bottom halves while holding
> the shared CQ polling lock. The series adds no timer, rate sampler,
> workqueue, sysfs knob, module parameter, or IRQ thread.
>
> Why not use nvme.use_threaded_interrupts=1 instead?
>
> Threaded mode installs nvme_irq_check() as the primary handler and
> nvme_irq() as the IRQ thread. When the primary handler finds a pending
> completion, it returns IRQ_WAKE_THREAD, so completion queue draining runs
> through a schedulable task rather than the normal hardirq fast path.
>
> That remains a useful opt-in mode and this series does not change it. On
> the affected VM it prevented the lockup, but it also reduced peak
> throughput. This series has a narrower goal: retain direct hardirq
> completion while the CPU is keeping up, and use bounded irq_poll work
> only after the scheduler has already asserted need_resched().
>
> The distinction is not that irq_poll is always faster than an IRQ thread.
> It is that the normal path is left alone, while the fallback creates a
> scheduling boundary only after a reschedule request is observed.
>
> need_resched() is used here as a best-effort pressure signal, not as an
> interrupt-flood detector or an elapsed-time guarantee. It is transient,
> and irq_poll may execute one bounded callback before the scheduler runs.
> The narrower claim is that, once an NVMe hardirq observes an already
> pending reschedule request, it stops unbounded CQ draining in hardirq.
>
> Previous NVMe attempts made the fallback the normal high-load path. An
> all-irq_poll RFC in 2016 reported an 8-10% single-core IOPS loss and a
> measurable low-queue-depth latency cost. A hardirq-first, threaded
> overflow series in 2019 fixed the Azure lockup, but Azure testing
> reported a substantial throughput loss.
>
> Relevant discussions from the past around this problem:
>
>
> https://lore.kernel.org/
> %2Flinux-nvme%2F1475660534-16681-1-git-send-email-
> sagi%40grimberg.me%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808df25c30
> 4c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764118549
> %7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwM
> CIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=1
> mVDHNgaLqX48N6F4S5ny1KC76E4Yv6CM4k%2BOI7V3ks%3D&reserved=0
>
> https://lore.kernel.org/
> %2Flkml%2F1566281669-48212-1-git-send-email-
> longli%40linuxonhyperv.com%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808
> df25c304c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764
> 147973%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAu
> MDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C
> &sdata=8LgU81UBs5FTTkkc7zbQo7Gx828IjfJO6Un9aebcdkM%3D&reserved=0
>
> https://lore.kernel.org/
> %2Flinux-nvme%2F20191209175622.1964-1-
> kbusch%40kernel.org%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808df25c3
> 04c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764169164
> %7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwM
> CIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=N
> IMpebg%2B3xIyJvnIk7FbUJANUxhUSFmljodgIAqswTY%3D&reserved=0
>
> The first patch selects IRQ_POLL. The second makes existing task-context
> CQ polling safe against the new softirq user. The third adds the NVMe
> queue handoff and lifecycle handling.
>
> The RFC starts with a per-queue weight of 256, matching the existing
> global irq_poll budget. This is a starting point, not a claim that 256 is
> the optimal NVMe value.
>
> Testing included:
>
>   - full x86_64 kernel build and strict checkpatch
>   - four-controller QEMU fio, full-disable S3, controller reset, and
>     post-reset I/O
>   - QEMU lockdep, prove-locking, and atomic-context validation
>   - ARM64 Hyper-V VM with four 14-queue NVMe data controllers
>   - natural irq_poll activation with threaded interrupts disabled
>   - matching default-hardirq, IRQ-poll fallback, and threaded-IRQ comparison
>   - four controller resets followed by successful reads
>
> Read-only 4 KiB random-read results follow. "Default hardirq" uses the
> unpatched base. The other modes use the same kernel built from this series:
> "IRQ-poll fallback" uses the default nvme.use_threaded_interrupts=0, while
> "threaded IRQs" uses nvme.use_threaded_interrupts=1. Low-load values are
> medians of two 30-second runs. High-load fallback and threaded values are
> medians of two and three 90-second runs respectively; the default-hardirq
> run locked up before it could complete.
>
>                               default       IRQ-poll      threaded
>                               hardirq       fallback      IRQs
>   1 disk, QD1
>     IOPS                      24.42K        24.96K        21.24K
>     mean completion latency  38.94 us      38.17 us      44.60 us
>     p99 completion latency   90.6 us       90.1 us       94.2 us
>     p99.9 completion latency 94.7 us       94.7 us      103.9 us
>
>   4 disks, 128 jobs, QD256
>     status                    lockup (26s)  pass           pass
>     IOPS                      N/A           4.986M         3.165M
>     bandwidth                 N/A          19.0 GiB/s     12.1 GiB/s
>     mean completion latency   N/A           6.544 ms      10.332 ms
>     p99 completion latency    N/A          14.483 ms      22.151 ms
>     p99.9 completion latency  N/A          22.544 ms      24.773 ms
>
> IRQ-poll fallback prevents hardirq starvation while avoiding the 15-37%
> IOPS loss and higher latency observed with threaded interrupts, by
> activating only under scheduler pressure.
>
> Naman Jain (3):
>   nvme-pci: select IRQ_POLL
>   nvme-pci: make completion queue polling softirq-safe
>   nvme-pci: defer completions when rescheduling is needed
>
>  drivers/nvme/host/Kconfig |   1 +
>  drivers/nvme/host/pci.c   | 295
> ++++++++++++++++++++++++++++++++++++++++++++--
>  2 files changed, 283 insertions(+), 13 deletions(-)
>
> base-commit: eea3fef32a9cf36abcb5975a5a594e4135a6b026
> --
> 2.43.0

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09  5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
@ 2026-10-09  6:51   ` Naman Jain
  2026-10-09 10:16     ` Naman Jain
  0 siblings, 1 reply; 11+ messages in thread
From: Naman Jain @ 2026-10-09  6:51 UTC (permalink / raw)
  To: Michael Kelley, linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	linux-hyperv, linux-kernel



On 10/9/2026 11:24 AM, Michael Kelley wrote:
> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October 8, 2026 10:06 PM
>>
>> On systems with several fast NVMe controllers, completion interrupts can
>> keep returning to the same CPUs faster than scheduled work can run. Each
>> handler may drain only a small number of completions, but the combined
>> interrupt stream can still prevent scheduler and watchdog progress.
> 
> See this recent proposal [1] that sounds like it is addressing the same or a
> similar issue. And there is this [2] more global approach. It's worthwhile to read
> through the discussion on both threads. I haven't done a detailed comparison
> of either vs. your proposal.
> 
> Michael
> 
> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@bytedance.com/
> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1-lrizzo@google.com/
> 


Thanks for sharing these Michael. I'll check more on these, and try it out.

Regards,
Naman

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09  6:51   ` Naman Jain
@ 2026-10-09 10:16     ` Naman Jain
  2026-10-09 11:27       ` Luigi Rizzo
  2026-10-09 11:36       ` Fengnan
  0 siblings, 2 replies; 11+ messages in thread
From: Naman Jain @ 2026-10-09 10:16 UTC (permalink / raw)
  To: Michael Kelley, linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	linux-hyperv, linux-kernel, changfengnan, Luigi Rizzo



On 10/9/2026 12:21 PM, Naman Jain wrote:
> 
> 
> On 10/9/2026 11:24 AM, Michael Kelley wrote:
>> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October 
>> 8, 2026 10:06 PM
>>>
>>> On systems with several fast NVMe controllers, completion interrupts can
>>> keep returning to the same CPUs faster than scheduled work can run. Each
>>> handler may drain only a small number of completions, but the combined
>>> interrupt stream can still prevent scheduler and watchdog progress.
>>
>> See this recent proposal [1] that sounds like it is addressing the 
>> same or a
>> similar issue. And there is this [2] more global approach. It's 
>> worthwhile to read
>> through the discussion on both threads. I haven't done a detailed 
>> comparison
>> of either vs. your proposal.
>>
>> Michael
>>
>> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1- 
>> changfengnan@bytedance.com/
>> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1- 
>> lrizzo@google.com/
>>
> 
> 
> Thanks for sharing these Michael. I'll check more on these, and try it out.
> 
> Regards,
> Naman

++ authors of these two series, for awareness and if there is some 
configuration in their patches I should be trying to fix these lockup 
issues.

I tested both GSIM v5 and the NVMe adaptive interrupt polling patch on
the ARM64 Azure system where the NVMe hardirq soft lockup is reproducible.

Test system
===========

The VM has:

   - 128 Arm Neoverse-V2 vCPUs
   - two 64-CPU sockets / NUMA nodes
   - approximately 862 GiB RAM
   - four 3.5-TB Microsoft NVMe Direct Disk v2 data devices
   - one NVMe OS device and one additional accelerator-facing NVMe
     controller
   - Hyper-V vPCI with MSI-X
   - 14 I/O queues per data controller

The four data controllers independently map their I/O vectors to the
same 14 CPUs:

   0, 10, 19, 28, 37, 46, 55, 64,
   74, 83, 92, 101, 110, 119

Thus each of these CPUs handles corresponding queues from all four data
controllers.

The kernel base for both experiments was:

   next-20261006
   eea3fef32a9cf36abcb5975a5a594e4135a6b026
   7.3.0-rc6-next-20261006

The main lockup workload was read-only:

   - 4-KiB random reads
   - libaio
   - O_DIRECT
   - four data devices
   - 128 jobs
   - iodepth 256
   - 90-second nominal runtime

NVMe interrupt coalescing was disabled (FID 0x08 = 0), and
nvme.use_threaded_interrupts was zero.

Without a mitigation, the watchdog reports soft lockups after about
26 seconds, normally on all 14 CPUs listed above.

Conclusion
==========

On this VM:

   - GSIM v5 did not prevent the lockup with any tested adaptive or fixed
     setting, including the documented benchmark settings and the maximum
     allowed delay.
   - NVMe adaptive polling can improve throughput for sufficiently dense
     individual queues, but it did not prevent the production lockup
     because the aggregate cross-controller load was spread across enough
     queues that most queues did not meet the fixed 10-us admission
     threshold.

The production failure is triggered by aggregate scheduler starvation
from many queues and controllers sharing the same IRQ CPUs. Neither
generic interrupt-rate moderation nor per-queue throughput-based
adaptation directly observes that condition.


More details:

GSIM v5
=======

I tested the complete seven-patch v5 series:

https://lore.kernel.org/all/20260819124341.4185621-1-lrizzo@google.com/

The kernel was built with:

   CONFIG_IRQ_SW_MODERATION=y
   CONFIG_IRQ_TIME_ACCOUNTING=y

GSIM is runtime-disabled and per-IRQ opt-in by default. I identified the
56 dedicated I/O vectors belonging to the four data controllers. All 56
were eligible and exposed allow_sw_moderation, and only those vectors
were enabled.

I first verified the same GSIM-patched kernel with runtime moderation
disabled. It reproduced the soft lockup, as expected.

I then tested the suggested adaptive configuration from the cover letter:

   delay_us=100
   target_intr_rate=1000000
   hardirq_percent=70
   update_ms=5

I also tested the exact adaptive configuration used in patch 6's
benchmark section:

   delay_us=200
   target_intr_rate=1000000
   hardirq_percent=70
   update_ms=5

In addition, I tested fixed moderation at:

   10, 25, 50, 75, 100, 200 and 500 us

The 500-us value is the maximum allowed by the implementation.

GSIM was definitely active. For the adaptive 100-us test, on the 14
affected CPUs it selected delays between about 63 and 100 us, set
323,811 moderation timers, enqueued 506,557 IRQs, and recorded 24,332
hardirq-over-threshold events.

However, every full-duration GSIM-only configuration still soft-locked.

A summary is:

                          QD1 IOPS       QD256 outcome

   GSIM off                 24.2K        soft lockup
   adaptive 100 us           4.0K        soft lockup
   adaptive 200 us           3.3K        soft lockup / timeout
   fixed 10 us              22.1K        soft lockup
   fixed 200 us              3.2K        soft lockup
   fixed 500 us              1.7K        soft lockup / timeout

My interpretation is that GSIM reduces how often the NVMe interrupt
handler runs, but it does not bound how much work the handler performs
once entered. Delaying an interrupt permits more CQEs to accumulate, and
the NVMe hardirq still drains the CQ until empty. On this topology that
produces fewer, larger, still-unbounded hardirq executions.

This does not contradict the reported GSIM benefits for systems limited
by aggregate MSI-X traffic or PCIe/SoC backpressure. It means that on
this VM the limiting issue is scheduler fairness inside the NVMe
completion handler rather than interrupt-delivery overhead.

NVMe adaptive interrupt polling
===============================

I also tested:

https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@bytedance.com/

The posted revision also has the irq_poll full-budget bookkeeping issue
reported in the review thread: it can call irq_poll_complete() and still
return the full budget. Since the author acknowledged this and said it
would be fixed in the next revision, I added only the corresponding
one-line fix before boot testing. I did not boot the known-buggy state.

The tested adaptive algorithm uses fixed constants:

   - 10-us poll period and admission threshold
   - 8,192 CQEs per sample/trial window
   - irq_poll budget of 64 CQEs
   - 64 successful windows per polling episode
   - two immediate failed trials followed by a 524,288-CQE backoff

I tested the same patched kernel with:

   1. adaptive policy off
   2. adaptive policy enabled at runtime through
      /sys/class/nvme/nvmeX/adaptive_irq_polling
   3. nvme.use_adaptive_irq_polling=1 at boot

The runtime policy was enabled only for the four data controllers.

For the production lockup workload, both runtime-on and boot-default-on
still soft-locked at about 26 seconds.

IRQ_POLL increased by only about 520 callbacks during the entire
runtime-on test and by about 539 during the boot-default-on test.

The reason appears to be the fixed per-queue admission threshold. With
about 5M aggregate IOPS spread across 56 I/O queues:

   5M / 56 ~= 89K CQEs/s per queue
   average interval ~= 11.2 us

The adaptive patch attempts polling only when the measured per-queue
completion interval is at most 10 us. The workload is intense in
aggregate, but the individual queues are just below the polling
admission threshold.

Boot-time enablement produced the same result, so runtime sysfs
switching was not the issue.

I then tested workloads intended to match the patch's expected dense
per-queue case.

With one SSD, one job:

               IRQ mode       adaptive mode

   QD32        328K IOPS      328K IOPS
   QD64        335K IOPS      443K IOPS
   QD128       406K IOPS      232K IOPS

At QD64, adaptive polling was clearly active (about 1.83M IRQ_POLL
callbacks) and improved IOPS by about 32% while reducing latency. This
matches the direction reported by the author.

The results were not stable across repeats, however. A repeated QD64
run was approximately neutral. QD128 produced both a regression and an
improvement depending on run order/device state. This appears consistent
with the author's comment that some benchmark results were still
affected by drive-state variability.

I also tested one SSD with 16 deep jobs pinned to one CPU. Adaptive
polling activated, but averaged about 3% fewer IOPS than IRQ mode.

Thanks,
Naman

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09 10:16     ` Naman Jain
@ 2026-10-09 11:27       ` Luigi Rizzo
  2026-10-09 14:58         ` Luigi Rizzo
  2026-10-09 11:36       ` Fengnan
  1 sibling, 1 reply; 11+ messages in thread
From: Luigi Rizzo @ 2026-10-09 11:27 UTC (permalink / raw)
  To: Naman Jain
  Cc: Michael Kelley, linux-nvme, Keith Busch, Jens Axboe,
	Christoph Hellwig, Sagi Grimberg, linux-hyperv, linux-kernel,
	changfengnan

On Fri, Oct 9, 2026 at 12:17 PM Naman Jain <namjain@linux.microsoft.com> wrote:
>
>
>
> On 10/9/2026 12:21 PM, Naman Jain wrote:
> >
> >
> > On 10/9/2026 11:24 AM, Michael Kelley wrote:
> >> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October
> >> 8, 2026 10:06 PM
> >>>
> >>> On systems with several fast NVMe controllers, completion interrupts can
> >>> keep returning to the same CPUs faster than scheduled work can run. Each
> >>> handler may drain only a small number of completions, but the combined
> >>> interrupt stream can still prevent scheduler and watchdog progress.
> >>
> >> See this recent proposal [1] that sounds like it is addressing the
> >> same or a
> >> similar issue. And there is this [2] more global approach. It's
> >> worthwhile to read
> >> through the discussion on both threads. I haven't done a detailed
> >> comparison
> >> of either vs. your proposal.
> >>
> >> Michael
> >>
> >> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-
> >> changfengnan@bytedance.com/
> >> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1-
> >> lrizzo@google.com/
> >>
> >
> >
> > Thanks for sharing these Michael. I'll check more on these, and try it out.
> >
> > Regards,
> > Naman
>
> ++ authors of these two series, for awareness and if there is some
> configuration in their patches I should be trying to fix these lockup
> issues.

I see your thorough analysis below, thanks for the details.
Is there any reason why you used nvme.use_threaded_interrupts=0 ?

Setting to 1 moves the work to a kernel thread and seems to be the
best way to prevent too much work in the hardirq and the soft lockup.

I am surprised that GSIM fails to prevent the soft lockup, though:
under high load the sequence of events (starting from unmoderated)
should be the following:

1. HW sends MSIx interrupt, not blocked by anything
2. SW calls handle_fasteoi_irq() --> handle_irq_event() --> ... -->
action->handler() which starts processing
3. HW possibly sends another MSIx interrupt
4. SW handle_irq_event() completes
5. SW irq_start_moderation() starts the timer, __disable_irq() and
sets IRQD_IRQ_INPROGRESS | IRQD_MODERATED
6. if #3 happened, or another HW interrupt comes before the timer expires:
  SW handle_fasteoi_irq() finds IRQD_IRQ_INPROGRESS | IRQD_MODERATED both set,
  and does not call handle_irq_event(), postponing the call
7. when the timer expires, the pending interrupt is reinjected.

The above suggests that we should see some pauses between runs of
action->handler(),
at least for a single queue per CPU.

Do you have any way to look with perf and bpftrace at the CPU
processing interrupts,
to see what keeps it busy (the pause between 6 and 7 should keep it
well below 100%)
and try to verify whether the calls to hndle_irq_event() are actually spaced
as the sequence above describes?

Having multiple interrupts on the same CPU mean that the gap
between interrupts for one can be used by others, and that may
definitely make the CPU 100% busy in hardirq.

Again to address these cases, nvme.use_threaded_interrupts=1 would
greatly reduce the time spent in hardirq, moving the bulk of the work to softirq
hence at least becoming interruptible.

cheers
luigi

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09 10:16     ` Naman Jain
  2026-10-09 11:27       ` Luigi Rizzo
@ 2026-10-09 11:36       ` Fengnan
  1 sibling, 0 replies; 11+ messages in thread
From: Fengnan @ 2026-10-09 11:36 UTC (permalink / raw)
  To: Naman Jain, Michael Kelley, linux-nvme
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	linux-hyperv, linux-kernel, Luigi Rizzo

Hi Naman:

This is an issue I hadn’t noticed before. 
I'll take a closer look at this issue later.

Regarding testing, there are a few points to note:
1. My patch has been updated to V2, https://lore.kernel.org/linux-nvme/20260826061546.56006-1-changfengnan@bytedance.com/
2. In my experience, when do performance tests, it’s best to do precondition
first and then execute an ABBA test; this ensures reliable results.
3. IIRC, disabling interrupts in the QEMU kernel does not actually disable hardware
interrupts; QEMU still handles them, which has a significant impact on performance testing.
Perhaps you could test this on a physical machine.

Also, I’d like to ask: does this softlockup issue only occur on QEMU+ARM?
or does it also occur on x86 physical machines?
Have you tried enabling the “Posted Interrupt” feature in VM? Maybe this will solve
this problem in VM.

Thanks.

在 2026/10/9 18:16, Naman Jain 写道:
> 
> 
> On 10/9/2026 12:21 PM, Naman Jain wrote:
>>
>>
>> On 10/9/2026 11:24 AM, Michael Kelley wrote:
>>> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October 8, 2026 10:06 PM
>>>>
>>>> On systems with several fast NVMe controllers, completion interrupts can
>>>> keep returning to the same CPUs faster than scheduled work can run. Each
>>>> handler may drain only a small number of completions, but the combined
>>>> interrupt stream can still prevent scheduler and watchdog progress.
>>>
>>> See this recent proposal [1] that sounds like it is addressing the same or a
>>> similar issue. And there is this [2] more global approach. It's worthwhile to read
>>> through the discussion on both threads. I haven't done a detailed comparison
>>> of either vs. your proposal.
>>>
>>> Michael
>>>
>>> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1- changfengnan@bytedance.com/
>>> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1- lrizzo@google.com/
>>>
>>
>>
>> Thanks for sharing these Michael. I'll check more on these, and try it out.
>>
>> Regards,
>> Naman
> 
> ++ authors of these two series, for awareness and if there is some configuration in their patches I should be trying to fix these lockup issues.
> 
> I tested both GSIM v5 and the NVMe adaptive interrupt polling patch on
> the ARM64 Azure system where the NVMe hardirq soft lockup is reproducible.
> 
> Test system
> ===========
> 
> The VM has:
> 
>   - 128 Arm Neoverse-V2 vCPUs
>   - two 64-CPU sockets / NUMA nodes
>   - approximately 862 GiB RAM
>   - four 3.5-TB Microsoft NVMe Direct Disk v2 data devices
>   - one NVMe OS device and one additional accelerator-facing NVMe
>     controller
>   - Hyper-V vPCI with MSI-X
>   - 14 I/O queues per data controller
> 
> The four data controllers independently map their I/O vectors to the
> same 14 CPUs:
> 
>   0, 10, 19, 28, 37, 46, 55, 64,
>   74, 83, 92, 101, 110, 119
> 
> Thus each of these CPUs handles corresponding queues from all four data
> controllers.
> 
> The kernel base for both experiments was:
> 
>   next-20261006
>   eea3fef32a9cf36abcb5975a5a594e4135a6b026
>   7.3.0-rc6-next-20261006
> 
> The main lockup workload was read-only:
> 
>   - 4-KiB random reads
>   - libaio
>   - O_DIRECT
>   - four data devices
>   - 128 jobs
>   - iodepth 256
>   - 90-second nominal runtime
> 
> NVMe interrupt coalescing was disabled (FID 0x08 = 0), and
> nvme.use_threaded_interrupts was zero.
> 
> Without a mitigation, the watchdog reports soft lockups after about
> 26 seconds, normally on all 14 CPUs listed above.
> 
> Conclusion
> ==========
> 
> On this VM:
> 
>   - GSIM v5 did not prevent the lockup with any tested adaptive or fixed
>     setting, including the documented benchmark settings and the maximum
>     allowed delay.
>   - NVMe adaptive polling can improve throughput for sufficiently dense
>     individual queues, but it did not prevent the production lockup
>     because the aggregate cross-controller load was spread across enough
>     queues that most queues did not meet the fixed 10-us admission
>     threshold.
> 
> The production failure is triggered by aggregate scheduler starvation
> from many queues and controllers sharing the same IRQ CPUs. Neither
> generic interrupt-rate moderation nor per-queue throughput-based
> adaptation directly observes that condition.
> 
> 
> More details:
> 
> GSIM v5
> =======
> 
> I tested the complete seven-patch v5 series:
> 
> https://lore.kernel.org/all/20260819124341.4185621-1-lrizzo@google.com/
> 
> The kernel was built with:
> 
>   CONFIG_IRQ_SW_MODERATION=y
>   CONFIG_IRQ_TIME_ACCOUNTING=y
> 
> GSIM is runtime-disabled and per-IRQ opt-in by default. I identified the
> 56 dedicated I/O vectors belonging to the four data controllers. All 56
> were eligible and exposed allow_sw_moderation, and only those vectors
> were enabled.
> 
> I first verified the same GSIM-patched kernel with runtime moderation
> disabled. It reproduced the soft lockup, as expected.
> 
> I then tested the suggested adaptive configuration from the cover letter:
> 
>   delay_us=100
>   target_intr_rate=1000000
>   hardirq_percent=70
>   update_ms=5
> 
> I also tested the exact adaptive configuration used in patch 6's
> benchmark section:
> 
>   delay_us=200
>   target_intr_rate=1000000
>   hardirq_percent=70
>   update_ms=5
> 
> In addition, I tested fixed moderation at:
> 
>   10, 25, 50, 75, 100, 200 and 500 us
> 
> The 500-us value is the maximum allowed by the implementation.
> 
> GSIM was definitely active. For the adaptive 100-us test, on the 14
> affected CPUs it selected delays between about 63 and 100 us, set
> 323,811 moderation timers, enqueued 506,557 IRQs, and recorded 24,332
> hardirq-over-threshold events.
> 
> However, every full-duration GSIM-only configuration still soft-locked.
> 
> A summary is:
> 
>                          QD1 IOPS       QD256 outcome
> 
>   GSIM off                 24.2K        soft lockup
>   adaptive 100 us           4.0K        soft lockup
>   adaptive 200 us           3.3K        soft lockup / timeout
>   fixed 10 us              22.1K        soft lockup
>   fixed 200 us              3.2K        soft lockup
>   fixed 500 us              1.7K        soft lockup / timeout
> 
> My interpretation is that GSIM reduces how often the NVMe interrupt
> handler runs, but it does not bound how much work the handler performs
> once entered. Delaying an interrupt permits more CQEs to accumulate, and
> the NVMe hardirq still drains the CQ until empty. On this topology that
> produces fewer, larger, still-unbounded hardirq executions.
> 
> This does not contradict the reported GSIM benefits for systems limited
> by aggregate MSI-X traffic or PCIe/SoC backpressure. It means that on
> this VM the limiting issue is scheduler fairness inside the NVMe
> completion handler rather than interrupt-delivery overhead.
> 
> NVMe adaptive interrupt polling
> ===============================
> 
> I also tested:
> 
> https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@bytedance.com/
> 
> The posted revision also has the irq_poll full-budget bookkeeping issue
> reported in the review thread: it can call irq_poll_complete() and still
> return the full budget. Since the author acknowledged this and said it
> would be fixed in the next revision, I added only the corresponding
> one-line fix before boot testing. I did not boot the known-buggy state.
> 
> The tested adaptive algorithm uses fixed constants:
> 
>   - 10-us poll period and admission threshold
>   - 8,192 CQEs per sample/trial window
>   - irq_poll budget of 64 CQEs
>   - 64 successful windows per polling episode
>   - two immediate failed trials followed by a 524,288-CQE backoff
> 
> I tested the same patched kernel with:
> 
>   1. adaptive policy off
>   2. adaptive policy enabled at runtime through
>      /sys/class/nvme/nvmeX/adaptive_irq_polling
>   3. nvme.use_adaptive_irq_polling=1 at boot
> 
> The runtime policy was enabled only for the four data controllers.
> 
> For the production lockup workload, both runtime-on and boot-default-on
> still soft-locked at about 26 seconds.
> 
> IRQ_POLL increased by only about 520 callbacks during the entire
> runtime-on test and by about 539 during the boot-default-on test.
> 
> The reason appears to be the fixed per-queue admission threshold. With
> about 5M aggregate IOPS spread across 56 I/O queues:
> 
>   5M / 56 ~= 89K CQEs/s per queue
>   average interval ~= 11.2 us
> 
> The adaptive patch attempts polling only when the measured per-queue
> completion interval is at most 10 us. The workload is intense in
> aggregate, but the individual queues are just below the polling
> admission threshold.
> 
> Boot-time enablement produced the same result, so runtime sysfs
> switching was not the issue.
> 
> I then tested workloads intended to match the patch's expected dense
> per-queue case.
> 
> With one SSD, one job:
> 
>               IRQ mode       adaptive mode
> 
>   QD32        328K IOPS      328K IOPS
>   QD64        335K IOPS      443K IOPS
>   QD128       406K IOPS      232K IOPS
> 
> At QD64, adaptive polling was clearly active (about 1.83M IRQ_POLL
> callbacks) and improved IOPS by about 32% while reducing latency. This
> matches the direction reported by the author.
> 
> The results were not stable across repeats, however. A repeated QD64
> run was approximately neutral. QD128 produced both a regression and an
> improvement depending on run order/device state. This appears consistent
> with the author's comment that some benchmark results were still
> affected by drive-state variability.
> 
> I also tested one SSD with 16 deep jobs pinned to one CPU. Adaptive
> polling activated, but averaged about 3% fewer IOPS than IRQ mode.
> 
> Thanks,
> Naman

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
  2026-10-09 11:27       ` Luigi Rizzo
@ 2026-10-09 14:58         ` Luigi Rizzo
  0 siblings, 0 replies; 11+ messages in thread
From: Luigi Rizzo @ 2026-10-09 14:58 UTC (permalink / raw)
  To: Naman Jain
  Cc: Michael Kelley, linux-nvme, Keith Busch, Jens Axboe,
	Christoph Hellwig, Sagi Grimberg, linux-hyperv, linux-kernel,
	changfengnan

On Fri, Oct 9, 2026 at 1:27 PM Luigi Rizzo <lrizzo@google.com> wrote:
>
> On Fri, Oct 9, 2026 at 12:17 PM Naman Jain <namjain@linux.microsoft.com> wrote:
> >
> >
> >
> > On 10/9/2026 12:21 PM, Naman Jain wrote:
> > >
> > >
> > > On 10/9/2026 11:24 AM, Michael Kelley wrote:
> > >> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October
> > >> 8, 2026 10:06 PM
> > >>>
> > >>> On systems with several fast NVMe controllers, completion interrupts can
> > >>> keep returning to the same CPUs faster than scheduled work can run. Each
> > >>> handler may drain only a small number of completions, but the combined
> > >>> interrupt stream can still prevent scheduler and watchdog progress.
> > >>
> > >> See this recent proposal [1] that sounds like it is addressing the
> > >> same or a
> > >> similar issue. And there is this [2] more global approach. It's
> > >> worthwhile to read
> > >> through the discussion on both threads. I haven't done a detailed
> > >> comparison
> > >> of either vs. your proposal.
> > >>
> > >> Michael
> > >>
> > >> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-
> > >> changfengnan@bytedance.com/
> > >> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1-
> > >> lrizzo@google.com/
> > >>
> > >
> > >
> > > Thanks for sharing these Michael. I'll check more on these, and try it out.
> > >
> > > Regards,
> > > Naman
> >
> > ++ authors of these two series, for awareness and if there is some
> > configuration in their patches I should be trying to fix these lockup
> > issues.
>
> I see your thorough analysis below, thanks for the details.
> Is there any reason why you used nvme.use_threaded_interrupts=0 ?
>
> Setting to 1 moves the work to a kernel thread and seems to be the
> best way to prevent too much work in the hardirq and the soft lockup.
>
> I am surprised that GSIM fails to prevent the soft lockup, though:
> under high load the sequence of events (starting from unmoderated)
> should be the following:
>
> 1. HW sends MSIx interrupt, not blocked by anything
> 2. SW calls handle_fasteoi_irq() --> handle_irq_event() --> ... -->
> action->handler() which starts processing
> 3. HW possibly sends another MSIx interrupt
> 4. SW handle_irq_event() completes
> 5. SW irq_start_moderation() starts the timer, __disable_irq() and
> sets IRQD_IRQ_INPROGRESS | IRQD_MODERATED
> 6. if #3 happened, or another HW interrupt comes before the timer expires:
>   SW handle_fasteoi_irq() finds IRQD_IRQ_INPROGRESS | IRQD_MODERATED both set,
>   and does not call handle_irq_event(), postponing the call
> 7. when the timer expires, the pending interrupt is reinjected.
>
> The above suggests that we should see some pauses between runs of
> action->handler(),
> at least for a single queue per CPU.
>
> Do you have any way to look with perf and bpftrace at the CPU
> processing interrupts,
> to see what keeps it busy (the pause between 6 and 7 should keep it
> well below 100%)
> and try to verify whether the calls to hndle_irq_event() are actually spaced
> as the sequence above describes?
>
> Having multiple interrupts on the same CPU mean that the gap
> between interrupts for one can be used by others, and that may
> definitely make the CPU 100% busy in hardirq.

Following up on the multiple interrupt sources on one CPU, you could
try something like this:
every time an interrupt source comes while the timer is active,
increment a counter.
When the timer expires and the counter is > 0, add as many "gaps" as
there were sources.
This should restore the invariant that every interrupt source gets one
run per moderation interval,
and the adaptive control loop should adjust the value to hit the target

--- a/kernel/irq/irq_moderation.h
    +++ b/kernel/irq/irq_moderation.h

    @@  struct irq_mod_state {
    +    unsigned int        kicked;
         unsigned int        sleep_ns;
    --- a/kernel/irq/irq_moderation.c
    +++ b/kernel/irq/irq_moderation.c
    @@ bool irq_moderation_do_start(struct irq_desc *desc, struct
irq_mod_state *m)
             /* We need moderation, start the timer. */
             m->timer_set++;
    +        m->kicked = 0;
             hrtimer_start_range_ns(&m->timer, ns_to_ktime(m->sleep_ns),
                            slack_ns, HRTIMER_MODE_REL_PINNED_HARD);
    +    } else if (READ_ONCE(irq_mod_params.timer_push)) {
    +        /* Timer armed: record the kick, timer_callback() will
push it forward. */
    +        m->kicked++;
         }
    @@ static enum hrtimer_restart timer_callback(struct hrtimer *timer)
         lockdep_assert_irqs_disabled();

    +    /* Push the timer forward by sleep_ns for each kick received
while armed. */
    +    if (m->kicked) {
    +        hrtimer_forward_now(timer, ns_to_ktime((u64)m->kicked *
m->sleep_ns));
    +        m->kicked = 0;
    +        return HRTIMER_RESTART;
    +    }
    +

cheers
luigi

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed
  2026-10-09  5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
@ 2026-10-09 15:30   ` Keith Busch
  0 siblings, 0 replies; 11+ messages in thread
From: Keith Busch @ 2026-10-09 15:30 UTC (permalink / raw)
  To: Naman Jain
  Cc: linux-nvme, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Michael Kelley, linux-hyperv, linux-kernel

On Fri, Oct 09, 2026 at 05:05:56AM +0000, Naman Jain wrote:
>  static irqreturn_t nvme_irq(int irq, void *data)
>  {
>  	struct nvme_queue *nvmeq = data;
>  	DEFINE_IO_COMP_BATCH(iob);
>  
> +	if (unlikely(test_and_clear_bit(NVMEQ_IRQ_POLL_HANDOFF,
> +					&nvmeq->flags))) {
> +		if (!nvme_cqe_pending(nvmeq))
> +			return IRQ_HANDLED;
> +	}
> +	if (unlikely(test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) &&
> +		     (!test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags) ||
> +		      test_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags) ||
> +		      READ_ONCE(to_pci_dev(nvmeq->dev->dev)->error_state) !=
> +				pci_channel_io_normal)))
> +		return IRQ_HANDLED;
> +
> +	if (unlikely(in_hardirq() &&
> +		     test_bit(NVMEQ_IRQ_POLL_INIT, &nvmeq->flags) &&
> +		     !test_bit(NVMEQ_IRQ_POLL_PAUSED, &nvmeq->flags) &&
> +		     need_resched() && nvme_cqe_pending(nvmeq) &&
> +		     !test_and_set_bit_lock(NVMEQ_IRQ_POLL, &nvmeq->flags))) {
> +		set_bit(NVMEQ_IRQ_POLL_MASKED, &nvmeq->flags);
> +		disable_irq_nosync(irq);
> +		irq_poll_sched(&nvmeq->iopoll);
> +		return IRQ_HANDLED;
> +	}

The disable_irq_nosync still does non-posted transactions for msi and
msi-x. NVMe has a shortcut for msi, but there's currently nothing for
msi-x. For that case, I have a prep patch I consulted with Thomas, and
the result is below. I'm delayed sending it out while catching up from
travel, but I think this will be even more efficient.

---
diff --git a/drivers/irqchip/irq-msi-lib.c b/drivers/irqchip/irq-msi-lib.c
index 45e0ed3134ce1..bff622b4839aa 100644
--- a/drivers/irqchip/irq-msi-lib.c
+++ b/drivers/irqchip/irq-msi-lib.c
@@ -126,6 +126,7 @@ bool msi_lib_init_dev_msi_info(struct device *dev, struct irq_domain *domain,
 	 */
 	if (info->flags & MSI_FLAG_PCI_MSI_MASK_PARENT) {
 		chip->irq_mask		= irq_chip_mask_parent;
+		chip->irq_mask_nowait	= NULL;
 		chip->irq_unmask	= irq_chip_unmask_parent;
 	}
 
diff --git a/drivers/pci/msi/irqdomain.c b/drivers/pci/msi/irqdomain.c
index 6e65f0f44112e..92a53c2772659 100644
--- a/drivers/pci/msi/irqdomain.c
+++ b/drivers/pci/msi/irqdomain.c
@@ -162,6 +162,11 @@ static void pci_irq_mask_msix(struct irq_data *data)
 	pci_msix_mask(irq_data_get_msi_desc(data));
 }
 
+static void pci_irq_mask_nowait_msix(struct irq_data *data)
+{
+	pci_msix_mask_nowait(irq_data_get_msi_desc(data));
+}
+
 static void pci_irq_unmask_msix(struct irq_data *data)
 {
 	pci_msix_unmask(irq_data_get_msi_desc(data));
@@ -182,6 +187,7 @@ static const struct msi_domain_template pci_msix_template = {
 		.irq_startup		= pci_irq_startup_msix,
 		.irq_shutdown		= pci_irq_shutdown_msix,
 		.irq_mask		= pci_irq_mask_msix,
+		.irq_mask_nowait	= pci_irq_mask_nowait_msix,
 		.irq_unmask		= pci_irq_unmask_msix,
 		.irq_write_msi_msg	= pci_msi_domain_write_msg,
 		.flags			= IRQCHIP_ONESHOT_SAFE,
diff --git a/drivers/pci/msi/msi.h b/drivers/pci/msi/msi.h
index 0b420b319f50f..337ff9d9da0ed 100644
--- a/drivers/pci/msi/msi.h
+++ b/drivers/pci/msi/msi.h
@@ -40,10 +40,16 @@ static inline void pci_msix_write_vector_ctrl(struct msi_desc *desc, u32 ctrl)
 		writel(ctrl, desc_addr + PCI_MSIX_ENTRY_VECTOR_CTRL);
 }
 
-static inline void pci_msix_mask(struct msi_desc *desc)
+/* Posted write only: the device may still send an interrupt after return */
+static inline void pci_msix_mask_nowait(struct msi_desc *desc)
 {
 	desc->pci.msix_ctrl |= PCI_MSIX_ENTRY_CTRL_MASKBIT;
 	pci_msix_write_vector_ctrl(desc, desc->pci.msix_ctrl);
+}
+
+static inline void pci_msix_mask(struct msi_desc *desc)
+{
+	pci_msix_mask_nowait(desc);
 	/* Flush write to device */
 	readl(desc->pci.mask_base);
 }
diff --git a/include/linux/interrupt.h b/include/linux/interrupt.h
index 3bf969ad8fe07..d40c992ddb8a1 100644
--- a/include/linux/interrupt.h
+++ b/include/linux/interrupt.h
@@ -231,6 +231,7 @@ extern void devm_free_irq(struct device *dev, unsigned int irq, void *dev_id);
 
 bool irq_has_action(unsigned int irq);
 extern void disable_irq_nosync(unsigned int irq);
+extern void disable_irq_nowait(unsigned int irq);
 extern bool disable_hardirq(unsigned int irq);
 extern void disable_irq(unsigned int irq);
 extern void disable_percpu_irq(unsigned int irq);
diff --git a/include/linux/irq.h b/include/linux/irq.h
index f485369b1b4f7..de095299f20ed 100644
--- a/include/linux/irq.h
+++ b/include/linux/irq.h
@@ -455,6 +455,9 @@ static inline irq_hw_number_t irqd_to_hwirq(struct irq_data *d)
  * @irq_disable:	disable the interrupt
  * @irq_ack:		start of a new interrupt
  * @irq_mask:		mask an interrupt source
+ * @irq_mask_nowait:	mask an interrupt source without waiting for the
+ *			hardware to acknowledge it (optional, falls back to
+ *			@irq_mask)
  * @irq_mask_ack:	ack and mask an interrupt source
  * @irq_unmask:		unmask an interrupt source
  * @irq_eoi:		end of interrupt
@@ -505,6 +508,7 @@ struct irq_chip {
 
 	void		(*irq_ack)(struct irq_data *data);
 	void		(*irq_mask)(struct irq_data *data);
+	void		(*irq_mask_nowait)(struct irq_data *data);
 	void		(*irq_mask_ack)(struct irq_data *data);
 	void		(*irq_unmask)(struct irq_data *data);
 	void		(*irq_eoi)(struct irq_data *data);
diff --git a/kernel/irq/chip.c b/kernel/irq/chip.c
index de754db414d1d..4e76f025b5f68 100644
--- a/kernel/irq/chip.c
+++ b/kernel/irq/chip.c
@@ -399,6 +399,37 @@ void irq_disable(struct irq_desc *desc)
 	__irq_disable(desc, irq_settings_disable_unlazy(desc));
 }
 
+static void mask_irq_nowait(struct irq_desc *desc)
+{
+	if (irqd_irq_masked(&desc->irq_data))
+		return;
+
+	if (desc->irq_data.chip->irq_mask_nowait) {
+		desc->irq_data.chip->irq_mask_nowait(&desc->irq_data);
+		irq_state_set_masked(desc);
+	} else {
+		mask_irq(desc);
+	}
+}
+
+/**
+ * irq_disable_nowait - Mark interrupt disabled and mask it without waiting
+ * @desc:	irq descriptor which should be disabled
+ *
+ * Like irq_disable() with lazy disable turned off, but the hardware may
+ * still deliver an interrupt after this returns. The flow handler deals
+ * with that the same way it does for a lazily disabled interrupt.
+ */
+void irq_disable_nowait(struct irq_desc *desc)
+{
+	if (desc->irq_data.chip->irq_disable) {
+		__irq_disable(desc, true);
+		return;
+	}
+	irq_state_set_disabled(desc);
+	mask_irq_nowait(desc);
+}
+
 void irq_percpu_enable(struct irq_desc *desc, unsigned int cpu)
 {
 	if (desc->irq_data.chip->irq_enable)
diff --git a/kernel/irq/internals.h b/kernel/irq/internals.h
index 0ce21dd454047..3ba13932ba831 100644
--- a/kernel/irq/internals.h
+++ b/kernel/irq/internals.h
@@ -98,6 +98,7 @@ extern void irq_startup_managed(struct irq_desc *desc);
 extern void irq_shutdown(struct irq_desc *desc);
 extern void irq_shutdown_and_deactivate(struct irq_desc *desc);
 extern void irq_disable(struct irq_desc *desc);
+extern void irq_disable_nowait(struct irq_desc *desc);
 extern void irq_percpu_enable(struct irq_desc *desc, unsigned int cpu);
 extern void irq_percpu_disable(struct irq_desc *desc, unsigned int cpu);
 extern void mask_irq(struct irq_desc *desc);
diff --git a/kernel/irq/manage.c b/kernel/irq/manage.c
index 2fbff2618a1e2..ecb4d2144b9d5 100644
--- a/kernel/irq/manage.c
+++ b/kernel/irq/manage.c
@@ -705,6 +705,27 @@ void disable_irq_nosync(unsigned int irq)
 }
 EXPORT_SYMBOL(disable_irq_nosync);
 
+/**
+ * disable_irq_nowait - mask an irq without waiting for the hardware
+ * @irq: Interrupt to disable
+ *
+ * Like disable_irq_nosync(), but masks the interrupt at the chip right away
+ * instead of lazily, and does not wait for the hardware to acknowledge the
+ * mask. An interrupt that still arrives is held pending and replayed by
+ * enable_irq(). Use this when an extra interrupt is cheaper than flushing
+ * the mask, e.g. switching from interrupts to polling.
+ *
+ * This function may be called from IRQ context.
+ */
+void disable_irq_nowait(unsigned int irq)
+{
+	scoped_irqdesc_get_and_buslock(irq, IRQ_GET_DESC_CHECK_GLOBAL) {
+		if (!scoped_irqdesc->depth++)
+			irq_disable_nowait(scoped_irqdesc);
+	}
+}
+EXPORT_SYMBOL_GPL(disable_irq_nowait);
+
 /**
  * disable_irq - disable an irq and wait for completion
  * @irq: Interrupt to disable
--

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-10-09 15:30 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09  5:05 [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Naman Jain
2026-10-09  5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
2026-10-09  5:05 ` [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe Naman Jain
2026-10-09  5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
2026-10-09 15:30   ` Keith Busch
2026-10-09  5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
2026-10-09  6:51   ` Naman Jain
2026-10-09 10:16     ` Naman Jain
2026-10-09 11:27       ` Luigi Rizzo
2026-10-09 14:58         ` Luigi Rizzo
2026-10-09 11:36       ` Fengnan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®