mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Naman Jain <namjain@linux.microsoft.com>
To: linux-nvme@lists.infradead.org
Cc: Keith Busch <kbusch@kernel.org>, Jens Axboe <axboe@kernel.dk>,
	Christoph Hellwig <hch@lst.de>, Sagi Grimberg <sagi@grimberg.me>,
	Michael Kelley <mhklinux@outlook.com>,
	linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
Date: Fri,  9 Oct 2026 05:05:53 +0000	[thread overview]
Message-ID: <20261009050556.2817978-1-namjain@linux.microsoft.com> (raw)

Hi,

On systems with several fast NVMe controllers, completion interrupts can
keep returning to the same CPUs faster than scheduled work can run. Each
handler may drain only a small number of completions, but the combined
interrupt stream can still prevent scheduler and watchdog progress.

This series keeps the existing hardirq completion path for normal I/O.
For an I/O queue with a dedicated MSI-X vector, if an interrupt arrives
with a completion pending while need_resched() is set, mask that vector
and hand the queue to irq_poll. The poller drains bounded batches and
returns the queue to interrupt mode as soon as it catches up.

The completion-path transitions are:

                         +-----------------+
                         | normal IRQ mode |
                         +-----------------+
                                  |
                   CQE pending and need_resched()
                                  |
                                  v
                     mask vector, schedule irq_poll
                                  |
                                  v
                +--------------------------------+
          +---->| drain up to one poll budget    |
          |     +--------------------------------+
          |              |                 |
          |         full budget        short batch
          |              |                 |
          |              v                 v
          |     remain on poll list   irq_poll_complete()
          |              |           mark handoff allowance
          |     generic irq_poll      clear poll ownership
          |     invokes next pass     enable vector
          |              |                 |
          +--------------+                 v
                                  +-----------------+
                                  | normal IRQ mode |
                                  +-----------------+

A full-budget return deliberately stays on the irq_poll list. The generic
irq_poll core invokes another bounded pass. Eventually a pass consumes
less than its budget, takes the short-batch path, and restores interrupt
mode.

An MSI-X message may have been latched while the vector was masked, but
not every MSI-X hierarchy can report pending state. The handoff allowance
therefore does not claim that an interrupt is known to be stale. It lets
exactly one empty interrupt after unmasking return IRQ_HANDLED. If a new
completion is present, the normal completion path handles it.

The core idea is simple: when a hardirq sees need_resched(), it masks the
interrupt and offloads completion processing to irq_poll. The interrupt is
unmasked once the CQ catches up. Everything else handles queue and
controller state changes while irq_poll is active, because its asynchronous
callback can outlive the hardirq that scheduled it.

The lifecycle transitions are:

  normal IRQ mode / irq_poll
              |
              | set PAUSED, join IRQ and irq_poll, mask vector
              v
          +--------+
          | paused |---- temporary ----> resume after controller is ready
          +--------+                      clear PAUSED, restore IRQ mode
              |
              | permanent stop
              | clear INIT, balance owned mask, synchronize replay
              v
        free registered IRQ

INIT means the queue has an initialized irq_poll instance and a dedicated
MSI-X vector. IRQ_POLL owns completion processing, while MASKED records the
matching disable_irq_nosync(). PAUSED blocks IRQ handling and new admission.
HANDOFF allows one empty interrupt after unmasking. A temporary pause keeps
MASKED set for resume; permanent stop keeps PAUSED set through IRQ removal.

Admin, shared, legacy, and threaded interrupt paths are unchanged. Explicit
polling never enters irq_poll; it only disables bottom halves while holding
the shared CQ polling lock. The series adds no timer, rate sampler,
workqueue, sysfs knob, module parameter, or IRQ thread.

Why not use nvme.use_threaded_interrupts=1 instead?

Threaded mode installs nvme_irq_check() as the primary handler and
nvme_irq() as the IRQ thread. When the primary handler finds a pending
completion, it returns IRQ_WAKE_THREAD, so completion queue draining runs
through a schedulable task rather than the normal hardirq fast path.

That remains a useful opt-in mode and this series does not change it. On
the affected VM it prevented the lockup, but it also reduced peak
throughput. This series has a narrower goal: retain direct hardirq
completion while the CPU is keeping up, and use bounded irq_poll work
only after the scheduler has already asserted need_resched().

The distinction is not that irq_poll is always faster than an IRQ thread.
It is that the normal path is left alone, while the fallback creates a
scheduling boundary only after a reschedule request is observed.

need_resched() is used here as a best-effort pressure signal, not as an
interrupt-flood detector or an elapsed-time guarantee. It is transient,
and irq_poll may execute one bounded callback before the scheduler runs.
The narrower claim is that, once an NVMe hardirq observes an already
pending reschedule request, it stops unbounded CQ draining in hardirq.

Previous NVMe attempts made the fallback the normal high-load path. An
all-irq_poll RFC in 2016 reported an 8-10% single-core IOPS loss and a
measurable low-queue-depth latency cost. A hardirq-first, threaded
overflow series in 2019 fixed the Azure lockup, but Azure testing
reported a substantial throughput loss.

Relevant discussions from the past around this problem:

  https://lore.kernel.org/linux-nvme/1475660534-16681-1-git-send-email-sagi@grimberg.me/
  https://lore.kernel.org/lkml/1566281669-48212-1-git-send-email-longli@linuxonhyperv.com/
  https://lore.kernel.org/linux-nvme/20191209175622.1964-1-kbusch@kernel.org/

The first patch selects IRQ_POLL. The second makes existing task-context
CQ polling safe against the new softirq user. The third adds the NVMe
queue handoff and lifecycle handling.

The RFC starts with a per-queue weight of 256, matching the existing
global irq_poll budget. This is a starting point, not a claim that 256 is
the optimal NVMe value.

Testing included:

  - full x86_64 kernel build and strict checkpatch
  - four-controller QEMU fio, full-disable S3, controller reset, and
    post-reset I/O
  - QEMU lockdep, prove-locking, and atomic-context validation
  - ARM64 Hyper-V VM with four 14-queue NVMe data controllers
  - natural irq_poll activation with threaded interrupts disabled
  - matching default-hardirq, IRQ-poll fallback, and threaded-IRQ comparison
  - four controller resets followed by successful reads

Read-only 4 KiB random-read results follow. "Default hardirq" uses the
unpatched base. The other modes use the same kernel built from this series:
"IRQ-poll fallback" uses the default nvme.use_threaded_interrupts=0, while
"threaded IRQs" uses nvme.use_threaded_interrupts=1. Low-load values are
medians of two 30-second runs. High-load fallback and threaded values are
medians of two and three 90-second runs respectively; the default-hardirq
run locked up before it could complete.

                              default       IRQ-poll      threaded
                              hardirq       fallback      IRQs
  1 disk, QD1
    IOPS                      24.42K        24.96K        21.24K
    mean completion latency  38.94 us      38.17 us      44.60 us
    p99 completion latency   90.6 us       90.1 us       94.2 us
    p99.9 completion latency 94.7 us       94.7 us      103.9 us

  4 disks, 128 jobs, QD256
    status                    lockup (26s)  pass           pass
    IOPS                      N/A           4.986M         3.165M
    bandwidth                 N/A          19.0 GiB/s     12.1 GiB/s
    mean completion latency   N/A           6.544 ms      10.332 ms
    p99 completion latency    N/A          14.483 ms      22.151 ms
    p99.9 completion latency  N/A          22.544 ms      24.773 ms

IRQ-poll fallback prevents hardirq starvation while avoiding the 15-37%
IOPS loss and higher latency observed with threaded interrupts, by
activating only under scheduler pressure.

Naman Jain (3):
  nvme-pci: select IRQ_POLL
  nvme-pci: make completion queue polling softirq-safe
  nvme-pci: defer completions when rescheduling is needed

 drivers/nvme/host/Kconfig |   1 +
 drivers/nvme/host/pci.c   | 295 ++++++++++++++++++++++++++++++++++++++++++++--
 2 files changed, 283 insertions(+), 13 deletions(-)

base-commit: eea3fef32a9cf36abcb5975a5a594e4135a6b026
-- 
2.43.0

             reply	other threads:[~2026-10-09  5:06 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09  5:05 Naman Jain [this message]
2026-10-09  5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
2026-10-09  5:05 ` [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe Naman Jain
2026-10-09  5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
2026-10-09 15:30   ` Keith Busch
2026-10-09  5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
2026-10-09  6:51   ` Naman Jain
2026-10-09 10:16     ` Naman Jain
2026-10-09 11:27       ` Luigi Rizzo
2026-10-09 14:58         ` Luigi Rizzo
2026-10-09 11:36       ` Fengnan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009050556.2817978-1-namjain@linux.microsoft.com \
    --to=namjain@linux.microsoft.com \
    --cc=axboe@kernel.dk \
    --cc=hch@lst.de \
    --cc=kbusch@kernel.org \
    --cc=linux-hyperv@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-nvme@lists.infradead.org \
    --cc=mhklinux@outlook.com \
    --cc=sagi@grimberg.me \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®