mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v4 0/2] nvme-tcp: allow setting per-queue io_cpu through sysfs
@ 2026-09-27  7:09 Saravanan D
  2026-09-27  7:09 ` [PATCH v4 1/2] nvme: add per-queue sysfs directories Saravanan D
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Saravanan D @ 2026-09-27  7:09 UTC (permalink / raw)
  To: linux-nvme
  Cc: kbusch, hch, sagi, axboe, nilay, dwagner, linux-kernel,
	iyamahata, kiyer, ganbalagane, sj, Saravanan D

On hosts that partition their CPUs after connect time, the connect
time io_cpu selection can land a queue's socket work on CPUs owned by
a different workload. This series lets a control plane set each
queue's io_cpu directly through sysfs.

The first patch adds a small common queue info that a transport embeds
in its queue structure, backing a transport neutral

  /sys/class/nvme/nvmeX/queues/<qid>/

directory under the controller device. The second patch has nvme-tcp
expose io_cpu and managed attributes there.

Changes since v3 [1]:
- Split into two patches and introduced the common queue info that
  Sagi proposed [2], named nvme_queue_info since the PCI driver
  already uses nvme_queue, as Nilay pointed out.
- Renamed the tcp_queues directory to queues so the ABI does not bake
  in the transport, as suggested by Nilay [3].
- Added the read-only managed attribute, 1 while io_cpu is driver
  managed and 0 once the user assigned it, as suggested by Nilay.
- io_cpu now accepts -1 or the string "unbound" to leave the socket
  work unbound, instead of restoring the connect time selection, as
  suggested by Sagi.
- A user assignment is tracked in the common queue info flags and
  still persists across reconnects, following Sagi's earlier guidance
  that io_cpu should not change on reconnect.
- The queue directories register once per controller lifetime and
  persist across reconnects, and the io_cpu accounting moved from the
  queue lock to a controller level lock since a write can now arrive
  while a queue is torn down.
- The io_cpu accounting now also covers a pin to a queue whose connect
  time pick did not bind a cpu, and the queue teardown clears any
  accounting left by a write racing the teardown.

[1] https://lore.kernel.org/linux-nvme/20260909211432.6741-1-saravanand@crusoe.ai/
[2] https://lore.kernel.org/linux-nvme/13c160d0-0657-47ce-9a16-58d348267ab0@grimberg.me/
[3] https://lore.kernel.org/linux-nvme/8dfee10f-89f3-463b-8d16-587873c3394c@linux.ibm.com/

Saravanan D (2):
  nvme: add per-queue sysfs directories
  nvme-tcp: allow setting per-queue io_cpu through sysfs

 Documentation/ABI/stable/sysfs-nvme |  21 ++++
 drivers/nvme/host/core.c            |   2 +
 drivers/nvme/host/nvme.h            |  23 +++++
 drivers/nvme/host/sysfs.c           |  42 ++++++++
 drivers/nvme/host/tcp.c             | 153 +++++++++++++++++++++++++++-
 5 files changed, 240 insertions(+), 1 deletion(-)


base-commit: 2ee54f01f07c0307deaf90ca8691a4643ae0357b
-- 
2.55.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-09  2:01 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27  7:09 [PATCH v4 0/2] nvme-tcp: allow setting per-queue io_cpu through sysfs Saravanan D
2026-09-27  7:09 ` [PATCH v4 1/2] nvme: add per-queue sysfs directories Saravanan D
2026-10-08  9:59   ` Nilay Shroff
2026-09-27  7:09 ` [PATCH v4 2/2] nvme-tcp: allow setting per-queue io_cpu through sysfs Saravanan D
2026-10-08  9:45   ` Nilay Shroff
2026-10-09  2:01     ` Saravanan D
2026-10-08  5:39 ` [PATCH v4 0/2] " Saravanan D

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®