From: Naman Jain <namjain@linux.microsoft.com>
To: Michael Kelley <mhklinux@outlook.com>,
"linux-nvme@lists.infradead.org" <linux-nvme@lists.infradead.org>
Cc: Keith Busch <kbusch@kernel.org>, Jens Axboe <axboe@kernel.dk>,
Christoph Hellwig <hch@lst.de>, Sagi Grimberg <sagi@grimberg.me>,
"linux-hyperv@vger.kernel.org" <linux-hyperv@vger.kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
changfengnan@bytedance.com, Luigi Rizzo <lrizzo@google.com>
Subject: Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure
Date: Fri, 9 Oct 2026 15:46:55 +0530 [thread overview]
Message-ID: <3f4a11ea-bbc0-4dc5-9720-3447346739cf@linux.microsoft.com> (raw)
In-Reply-To: <04d88297-1732-4761-ae82-d68c441a7156@linux.microsoft.com>
On 10/9/2026 12:21 PM, Naman Jain wrote:
>
>
> On 10/9/2026 11:24 AM, Michael Kelley wrote:
>> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, October
>> 8, 2026 10:06 PM
>>>
>>> On systems with several fast NVMe controllers, completion interrupts can
>>> keep returning to the same CPUs faster than scheduled work can run. Each
>>> handler may drain only a small number of completions, but the combined
>>> interrupt stream can still prevent scheduler and watchdog progress.
>>
>> See this recent proposal [1] that sounds like it is addressing the
>> same or a
>> similar issue. And there is this [2] more global approach. It's
>> worthwhile to read
>> through the discussion on both threads. I haven't done a detailed
>> comparison
>> of either vs. your proposal.
>>
>> Michael
>>
>> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-
>> changfengnan@bytedance.com/
>> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1-
>> lrizzo@google.com/
>>
>
>
> Thanks for sharing these Michael. I'll check more on these, and try it out.
>
> Regards,
> Naman
++ authors of these two series, for awareness and if there is some
configuration in their patches I should be trying to fix these lockup
issues.
I tested both GSIM v5 and the NVMe adaptive interrupt polling patch on
the ARM64 Azure system where the NVMe hardirq soft lockup is reproducible.
Test system
===========
The VM has:
- 128 Arm Neoverse-V2 vCPUs
- two 64-CPU sockets / NUMA nodes
- approximately 862 GiB RAM
- four 3.5-TB Microsoft NVMe Direct Disk v2 data devices
- one NVMe OS device and one additional accelerator-facing NVMe
controller
- Hyper-V vPCI with MSI-X
- 14 I/O queues per data controller
The four data controllers independently map their I/O vectors to the
same 14 CPUs:
0, 10, 19, 28, 37, 46, 55, 64,
74, 83, 92, 101, 110, 119
Thus each of these CPUs handles corresponding queues from all four data
controllers.
The kernel base for both experiments was:
next-20261006
eea3fef32a9cf36abcb5975a5a594e4135a6b026
7.3.0-rc6-next-20261006
The main lockup workload was read-only:
- 4-KiB random reads
- libaio
- O_DIRECT
- four data devices
- 128 jobs
- iodepth 256
- 90-second nominal runtime
NVMe interrupt coalescing was disabled (FID 0x08 = 0), and
nvme.use_threaded_interrupts was zero.
Without a mitigation, the watchdog reports soft lockups after about
26 seconds, normally on all 14 CPUs listed above.
Conclusion
==========
On this VM:
- GSIM v5 did not prevent the lockup with any tested adaptive or fixed
setting, including the documented benchmark settings and the maximum
allowed delay.
- NVMe adaptive polling can improve throughput for sufficiently dense
individual queues, but it did not prevent the production lockup
because the aggregate cross-controller load was spread across enough
queues that most queues did not meet the fixed 10-us admission
threshold.
The production failure is triggered by aggregate scheduler starvation
from many queues and controllers sharing the same IRQ CPUs. Neither
generic interrupt-rate moderation nor per-queue throughput-based
adaptation directly observes that condition.
More details:
GSIM v5
=======
I tested the complete seven-patch v5 series:
https://lore.kernel.org/all/20260819124341.4185621-1-lrizzo@google.com/
The kernel was built with:
CONFIG_IRQ_SW_MODERATION=y
CONFIG_IRQ_TIME_ACCOUNTING=y
GSIM is runtime-disabled and per-IRQ opt-in by default. I identified the
56 dedicated I/O vectors belonging to the four data controllers. All 56
were eligible and exposed allow_sw_moderation, and only those vectors
were enabled.
I first verified the same GSIM-patched kernel with runtime moderation
disabled. It reproduced the soft lockup, as expected.
I then tested the suggested adaptive configuration from the cover letter:
delay_us=100
target_intr_rate=1000000
hardirq_percent=70
update_ms=5
I also tested the exact adaptive configuration used in patch 6's
benchmark section:
delay_us=200
target_intr_rate=1000000
hardirq_percent=70
update_ms=5
In addition, I tested fixed moderation at:
10, 25, 50, 75, 100, 200 and 500 us
The 500-us value is the maximum allowed by the implementation.
GSIM was definitely active. For the adaptive 100-us test, on the 14
affected CPUs it selected delays between about 63 and 100 us, set
323,811 moderation timers, enqueued 506,557 IRQs, and recorded 24,332
hardirq-over-threshold events.
However, every full-duration GSIM-only configuration still soft-locked.
A summary is:
QD1 IOPS QD256 outcome
GSIM off 24.2K soft lockup
adaptive 100 us 4.0K soft lockup
adaptive 200 us 3.3K soft lockup / timeout
fixed 10 us 22.1K soft lockup
fixed 200 us 3.2K soft lockup
fixed 500 us 1.7K soft lockup / timeout
My interpretation is that GSIM reduces how often the NVMe interrupt
handler runs, but it does not bound how much work the handler performs
once entered. Delaying an interrupt permits more CQEs to accumulate, and
the NVMe hardirq still drains the CQ until empty. On this topology that
produces fewer, larger, still-unbounded hardirq executions.
This does not contradict the reported GSIM benefits for systems limited
by aggregate MSI-X traffic or PCIe/SoC backpressure. It means that on
this VM the limiting issue is scheduler fairness inside the NVMe
completion handler rather than interrupt-delivery overhead.
NVMe adaptive interrupt polling
===============================
I also tested:
https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@bytedance.com/
The posted revision also has the irq_poll full-budget bookkeeping issue
reported in the review thread: it can call irq_poll_complete() and still
return the full budget. Since the author acknowledged this and said it
would be fixed in the next revision, I added only the corresponding
one-line fix before boot testing. I did not boot the known-buggy state.
The tested adaptive algorithm uses fixed constants:
- 10-us poll period and admission threshold
- 8,192 CQEs per sample/trial window
- irq_poll budget of 64 CQEs
- 64 successful windows per polling episode
- two immediate failed trials followed by a 524,288-CQE backoff
I tested the same patched kernel with:
1. adaptive policy off
2. adaptive policy enabled at runtime through
/sys/class/nvme/nvmeX/adaptive_irq_polling
3. nvme.use_adaptive_irq_polling=1 at boot
The runtime policy was enabled only for the four data controllers.
For the production lockup workload, both runtime-on and boot-default-on
still soft-locked at about 26 seconds.
IRQ_POLL increased by only about 520 callbacks during the entire
runtime-on test and by about 539 during the boot-default-on test.
The reason appears to be the fixed per-queue admission threshold. With
about 5M aggregate IOPS spread across 56 I/O queues:
5M / 56 ~= 89K CQEs/s per queue
average interval ~= 11.2 us
The adaptive patch attempts polling only when the measured per-queue
completion interval is at most 10 us. The workload is intense in
aggregate, but the individual queues are just below the polling
admission threshold.
Boot-time enablement produced the same result, so runtime sysfs
switching was not the issue.
I then tested workloads intended to match the patch's expected dense
per-queue case.
With one SSD, one job:
IRQ mode adaptive mode
QD32 328K IOPS 328K IOPS
QD64 335K IOPS 443K IOPS
QD128 406K IOPS 232K IOPS
At QD64, adaptive polling was clearly active (about 1.83M IRQ_POLL
callbacks) and improved IOPS by about 32% while reducing latency. This
matches the direction reported by the author.
The results were not stable across repeats, however. A repeated QD64
run was approximately neutral. QD128 produced both a regression and an
improvement depending on run order/device state. This appears consistent
with the author's comment that some benchmark results were still
affected by drive-state variability.
I also tested one SSD with 16 deep jobs pinned to one CPU. Adaptive
polling activated, but averaged about 3% fewer IOPS than IRQ mode.
Thanks,
Naman
next prev parent reply other threads:[~2026-10-09 10:17 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 5:05 Naman Jain
2026-10-09 5:05 ` [RFC PATCH 1/3] nvme-pci: select IRQ_POLL Naman Jain
2026-10-09 5:05 ` [RFC PATCH 2/3] nvme-pci: make completion queue polling softirq-safe Naman Jain
2026-10-09 5:05 ` [RFC PATCH 3/3] nvme-pci: defer completions when rescheduling is needed Naman Jain
2026-10-09 15:30 ` Keith Busch
2026-10-09 5:54 ` [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure Michael Kelley
2026-10-09 6:51 ` Naman Jain
2026-10-09 10:16 ` Naman Jain [this message]
2026-10-09 11:27 ` Luigi Rizzo
2026-10-09 14:58 ` Luigi Rizzo
2026-10-09 11:36 ` Fengnan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3f4a11ea-bbc0-4dc5-9720-3447346739cf@linux.microsoft.com \
--to=namjain@linux.microsoft.com \
--cc=axboe@kernel.dk \
--cc=changfengnan@bytedance.com \
--cc=hch@lst.de \
--cc=kbusch@kernel.org \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-nvme@lists.infradead.org \
--cc=lrizzo@google.com \
--cc=mhklinux@outlook.com \
--cc=sagi@grimberg.me \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®