mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support
@ 2026-10-03 12:40 Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
                   ` (3 more replies)
  0 siblings, 4 replies; 7+ messages in thread
From: Tao Cui @ 2026-10-03 12:40 UTC (permalink / raw)
  To: tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov

From: Tao Cui <cuitao@kylinos.cn>

This is v10 of the RFC.  A BPF program can now take over the cost
model of one device: attaching an iocost_model_ops struct_ops creates
the ioc if needed and switches the device to the model, io.cost.model
selects between the attached and the builtin model, and detaching or
device removal restores the builtin one.  The attachment, the cgroup
callback pairing and the removal paths are serialized with what blkg
creation and destruction actually use.

Why a pluggable model at all
----------------------------

When iocost landed in 2019, its commit message already promised that
"a later patch will also allow using bpf progs for cost models", and
the code has carried the split for it ever since: calc_vtime_cost()
is a dispatcher whose only implementation is calc_vtime_cost_builtin().
Seven years later the builtin linear model is still the only one.
This series fills that slot: builtin algorithms remain the default
while new ones can be prototyped in BPF.

The measured problems
---------------------

The builtin model prices each IO with a binary sequential/random base
picked by a single per-cgroup cursor and a 16MB seek threshold, plus
a per-page cost.  On a virtio-blk device with the HDD autop profile,
a 4k IO costs ~24us when judged sequential and ~2.7ms when judged
random, a 112x spread, so a wrong judgement becomes a wrong price.
Three classes of mispricing, all measured:

 1. Heuristic rigidity.  Two legitimate sequential readers in one
    cgroup (a database with multiple tablespaces, a threaded backup)
    ping-pong the single cursor and are all priced random: a
    measured 89x overcharge collapses throughput under the same
    weight.  Random IO within a hot window smaller than the 16MB
    threshold is priced sequential: measured 107x undercharge, an
    accounting escape for hotspot workloads.  No setting of the six
    builtin parameters seems able to fix this: telling the streams
    apart requires per-IO state tracking, which looks like logic
    rather than coefficients.

 2. Device nonlinearity.  SLC-cache phases, SMR band placement and
    shared controllers (multiple NVMe namespaces multiplexing one
    device) make the real cost of an identical IO vary by an order
    of magnitude over time or across namespaces.  A static
    6-parameter linear model has no way to express that.

 3. Unpriced operations.  Flush and zone append fall through to a
    cost of zero and bypass throttling entirely, and the same pattern
    extends to device quirks the builtin model was never taught.

Mispricing feeds directly into the control loop: vtime budgets,
surplus donation and the vrate feedback all consume the model's
output, so a wrong model can skew the whole controller.

How
---

Attachment follows the hid_bpf_ops model: the target device is set
in the dev member of the struct_ops from userspace before load; the
device is looked up with blkdev_get_no_open() and disk_live() is
checked under rq_qos_mutex, like blkg_conf_open_bdev().  Attaching
creates the ioc if needed, like an io.cost.model write does, and
switches the model under the same freeze and quiesce those writes
go through; enabling stays with io.cost.qos and wbt stays with it
too.  When the disk goes away, the model is ejected completely.

Attachment is independent of the model selection and of the
controller state: while a model is attached, "model" selects between
it and the builtin model - "model=bpf" switches to the attached
model, "model=linear" switches back to the builtin model, and
neither detaches the struct_ops; only detaching removes the model,
after which "model=bpf" fails.  "ctrl" keeps describing the builtin
coefficients, which are kept while the BPF model is in use and take
effect again when switched back.  Enabling and disabling the
controller, and with it the wbt handover, remains with io.cost.qos.

A bound model in use owns pricing for every charged IO on the
device: it is called from the bio charging path and prices every
operation including flushes.  The completion-time request sizing
which feeds the latency met/missed accounting uses the transfer cost
coefficients the struct_ops carries (vtime per page for reads and
writes), so the builtin latency tracking and vrate adjustment
follow the model's pricing while it is in use; extending the
model to the QoS side itself is left for a later interface.

The cgroup callbacks are bound to the iocg policy lifetime and pair
up: iocg_init() is delivered to every cgroup which already has a
blkg on the device when the model is attached, and to each one
appearing afterwards; iocg_free() is delivered to the cgroups which
still exist when the model is detached.  The existing iocgs are
walked over q->blkg_list under q->blkcg_mutex, the same walk the
blkcg policy teardown uses; cgroups without a blkg on the device
yet are not missed, as their callbacks are delivered at blkg
creation and destruction.

    u64 calc_cost(struct bio *bio, u64 model_flags)

The model reads whatever it needs from the bio itself: the
operation flags (the operation must be extracted with a mask, and
the REQ_* flag bits, including PREFLUSH/FUA, are part of it), the
size, the start sector and the issuing cgroup through
bio->bi_blkg.  model_flags carries iocost-specific metadata which
is not a property of the bio, such as whether the cost calculation
is for a merged request; the return value is vtime, clamped to 1
second of device time per IO.  Per-cgroup state can be stored in
BPF_MAP_TYPE_CGRP_STORAGE keyed by the cgroup of bi_blkg; it
follows the cgroup lifetime.  Sleepable models are rejected at
verification, since calc_cost() runs under RCU read lock.

Patch overview:

 1/4: the BPF struct_ops cost model support: Kconfig, ops
      definition, per-device attachment, unified dispatch and
      verifier checks
 2/4: selftests with two example models (the full builtin linear
      HDD formula at double cost, and a multi-stream sequentiality
      model) plus a runner and the selftest kernel config entries
 3/4: add an iocost_ioc_tick tracepoint emitting the per-period
      controller state, so model quality can be evaluated without
      drgn (existing events are state-change driven and silent in
      steady state)
 4/4: document the attachment in cgroup-v2.rst

Does it work
------------

Mechanism, verified functionally on a virtio-blk device (HDD
profile, sequential-read workload from a 1%-weight cgroup, builtin
vs the 2x example model):

 - per-IO charge: 2882us -> 5722us, a factor of 1.985x; the
   completed IO count halves and total cost.usage is conserved,
   i.e. the model output drives both charging and budgeting
 - attachment: model=bpf readback while in use (ctrl keeps
   describing the builtin coefficients), EBUSY for a second model
   on the same device, a second device attaches an independent
   model, "model=linear"/"model=bpf" switch between the attached
   model and the builtin without detaching, detaching removes the
   model and "model=bpf" fails afterwards
 - edge cases: writing an unknown device fails and nothing is
   applied; the selftest runner checks the write error and errno
   of every step, including the restoration

Workload-shape verification (same setup, 4k IOs at weight 1000,
builtin vs the 2x example model):

 - flush-heavy workload (read/write/fsync alternating): priced
   1.99x the builtin, i.e. flushes no longer reset the cursor
   and misjudge the following IO as random
 - non-page-multiple IO (6 KiB): priced ~2x, matching the
   builtin's truncating page count
 - first IO from a high LBA (past 16 MiB): priced 2.01x, i.e.
   a fresh cgroup's zero cursor no longer misjudges the first IO
   as random

Payoff, demonstrated with the multi-stream example model (2/4) on
the same setup, 4k IOs at weight 1000, builtin vs the model:

 - two sequential readers in one cgroup: priced 1961us/op by builtin
   (both judged random by the single cursor) and 23us/op by the
   model (each stream keeps its own slot); the completed IO count
   rises by two orders of magnitude
 - random IO inside an 8M window: priced 24us/op by builtin
   (undercharge, an accounting escape) and 2607us/op by the model
 - single-stream sequential and whole-disk random pricing are
   unchanged, so the model fixes both directions of mispricing
   without introducing a new one

Non-interference, measured on enterprise NVMe: no measurable
overhead when the BPF model is not attached.

Changes in v9:
- the attach holds rq_qos_mutex from the disk_live() check through
  the unfreeze, like an io.cost.model write does, instead of
  releasing it between creating the ioc and switching the model;
  the queue reference, the second q_to_ioc() and the lockless
  traversal it needed are gone
- the attachment is published, cleared and the blkg list is walked
  under blkcg_mutex and queue_lock: blkg creation runs ioc_pd_init()
  and adds to q->blkg_list under queue_lock only, so the previous
  blkcg_mutex-only serialization could double-deliver iocg_init() or
  deliver iocg_free() without iocg_init().  The init walk now goes
  parents first, like blkcg_activate_policy()
- the device removal ejection reads ops->bdev before clearing ops->q:
  once q is NULL .unreg returns without a lock and the map can be
  freed, so touching ops afterwards was a use-after-free window.  The
  ejection moved to an ioc_bpf_eject() helper
- .unreg only detaches when the closing link is the one which owns
  the attachment, so a map re-attached to a new device through a
  newer link is not torn down by an old link; the v8 ops->q re-entry
  guard at .reg is gone, a second attach of an attached map now
  fails with -EBUSY like any other model would
- the ioc->attached and ioc->model pointers are unconditional and
  blk_get_queue_rcu() is declared in block/blk.h without an export,
  which drops the local prototypes and the remaining #ifdefs
- the unused struct_ops .init name lookup, the bdev/q cases in
  .init_member (the core already rejects nonzero non-function members
  it does not claim) and the unused write_cost_model() buffers are
  gone; the -EOPNOTSUPP selftest paths call test__skip()
- the selftest writes the dev member through the skeleton's typed
  struct_ops access instead of assuming it is the first member, and
  the maps are ".struct_ops.link" so closing the fd detaches; the
  example model's seek judgement is gated on a non-zero IO size so a
  dataless flush is priced sequentially, and write_cost_model()
  returns the negated errno ASSERT_ERR() expects
- the tick tracepoint's running field reports the active iocg list
  instead of ioc->running, which has not transitioned to IOC_IDLE at
  the emit point, so the final tick reads active=0 running=0
- the attach and detach take the queue freeze before rq_qos_mutex,
  like ioc_qos_write() does; taking the mutex around the freeze
  instead formed a lockdep cycle with the io.cost.qos write path,
  and .unreg's link recheck moved under the mutex inside
  ioc_bpf_detach()
- the io.cost.model documentation states that detaching restores the
  builtin model and that ctrl only accepts "auto" and "user" and
  never selects the model, says attaching rather than loading binds
  the model, and that removing the device removes the model as well,
  with the struct_ops link still having to be closed afterwards

Changes in v10:
- one map binds one device at a time: the attach rejects a second
  attach while the ops is still bound, which the v9 owning-link
  rework had dropped.  Without it, attaching the same map to a
  second device clobbered the single device slot, leaked the first
  bdev reference and left the first device pointing at the ops
- the detach takes the queue the caller holds a reference on as an
  argument instead of re-reading ops->q, fetches the ioc under
  rq_qos_mutex, and re-checks the owning link under the mutex, so
  closing an old link after an ejection and a re-attach no longer
  tears down the newer attachment
- iocg_free() is only delivered to a cgroup which received
  iocg_init(), tracked per iocg, so a blkg appearing between the
  attach walk and its own pd_init no longer gets an unpaired
  iocg_free(); a blkg whose radix_tree_insert() failed may still
  keep an iocg_init() without iocg_free(), which is now documented
  in the interface comment


Link: https://lore.kernel.org/r/20260908100143.47598-1-cui.tao@linux.dev # v1
Link: https://lore.kernel.org/r/20260910125817.223354-1-cui.tao@linux.dev # v2
Link: https://lore.kernel.org/r/20260914073356.791518-1-cui.tao@linux.dev # v3
Link: https://lore.kernel.org/r/20260916072302.1068871-1-cui.tao@linux.dev # v4
Link: https://lore.kernel.org/r/20260918031751.1255420-1-cui.tao@linux.dev # v5
Link: https://lore.kernel.org/r/20260918055001.1273840-1-cui.tao@linux.dev # v6
Link: https://lore.kernel.org/r/20260924054549.2271705-1-cui.tao@linux.dev # v7
Link: https://lore.kernel.org/r/20260930075154.189958-1-cui.tao@linux.dev # v8
Tao Cui (4):
  blk-iocost: add BPF struct_ops cost model support
  selftests/bpf: add iocost cost model test
  blk-iocost: add iocost_ioc_tick tracepoint for per-period device
    summary
  docs: cgroup-v2: document the iocost BPF cost model attachment
Link: https://lore.kernel.org/r/20261003013033.149288-1-cui.tao@linux.dev # v9

^ permalink raw reply	[flat|nested] 7+ messages in thread

* [RFC PATCH v10 1/4] blk-iocost: add BPF struct_ops cost model support
  2026-10-03 12:40 [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
@ 2026-10-03 12:40 ` Tao Cui
  2026-10-03 13:21   ` bot+bpf-ci
  2026-10-03 20:45   ` Alexei Starovoitov
  2026-10-03 12:40 ` [RFC PATCH v10 2/4] selftests/bpf: add iocost cost model test Tao Cui
                   ` (2 subsequent siblings)
  3 siblings, 2 replies; 7+ messages in thread
From: Tao Cui @ 2026-10-03 12:40 UTC (permalink / raw)
  To: tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov

iocost prices IO with a linear model derived from a handful of
device parameters.  The model is reasonable for common device classes,
but the real cost of an IO depends on device internals the linear
formula cannot capture (write amplification, GC, compression, cache
behavior), and the right formula is device specific.  This lets a BPF
program take over pricing for a device, so the formula can be adapted
to the device and updated without rebuilding the kernel.

An iocost_model_ops struct_ops is attached to one device: the target
device is set in the dev member from userspace before load and .reg
attaches the model to it, creating the ioc if needed, like an
io.cost.model write does, and switching the device to the model under
the same queue freeze and quiesce those writes go through.
io.cost.model selects between the attached model and the builtin one:
"model=bpf" switches to the attached model, "model=linear" back to
the builtin coefficients; neither detaches the struct_ops, and only
.unreg or device removal does.  Enabling and disabling the controller
stays with io.cost.qos.

While the BPF model is the one in use, it prices every charged IO:
calc_cost() is called from the IO submission path with RCU read lock
held, receives the bio itself and returns the cost in vtime units,
clamped to one second of device time per IO.  IOs which are never
charged, root cgroup IOs and IOs while the controller is disabled,
do not reach it.  The struct_ops also carries the transfer cost
coefficients, vtime per page for reads and writes, which the
completion-time request sizing of READ and WRITE requests uses, so
the latency tracking and vrate adjustment follow the model's
pricing for those requests.

The cgroup callbacks are bound to the iocg policy lifetime, one
(cgroup, device) pair per invocation: iocg_init() is delivered on
attach to every cgroup which already has a blkg on the device and to
each one appearing afterwards, iocg_free() on detach to every cgroup
still existing then, and at policy deactivation time for the rest, so
init and free pair up, except for a blkg whose radix_tree_insert()
failed inside blkg_create(): it had iocg_init() delivered but never
reaches the blkg list, so the detach walk misses it and iocg_free()
may not be delivered.  The attach and detach paths walk the
blkg list under blkcg_mutex and queue_lock, which is what blkg
creation and destruction synchronize on; device removal only clears
the attachment, as the policy deactivation which precedes it has
already delivered the remaining iocg_free() calls.

Device removal ejects the model completely, and .unreg only detaches
when the link which owns the attachment is closed, so a map
re-attached to a new device through a newer link is not torn down by
an old one.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 block/Kconfig              |  10 +
 block/Makefile             |   1 +
 block/blk-core.c           |  16 ++
 block/blk-iocost-bpf.c     | 184 +++++++++++++++
 block/blk-iocost.c         | 459 ++++++++++++++++++++++++++++++++++++-
 block/blk.h                |   1 +
 include/linux/blk-iocost.h | 126 ++++++++++
 7 files changed, 789 insertions(+), 8 deletions(-)
 create mode 100644 block/blk-iocost-bpf.c
 create mode 100644 include/linux/blk-iocost.h

diff --git a/block/Kconfig b/block/Kconfig
index 70e4a66d941ff..1cafc1bd0dda3 100644
--- a/block/Kconfig
+++ b/block/Kconfig
@@ -231,4 +231,14 @@ config BLK_ERROR_INJECTION
 
 source "block/Kconfig.iosched"
 
+config BLK_CGROUP_IOCOST_BPF
+	bool "Enable BPF pluggable cost model support for the cost IO controller"
+	depends on BLK_CGROUP_IOCOST && BPF_SYSCALL && BPF_JIT && DEBUG_INFO_BTF
+	help
+	 Enabling this option registers the "iocost_model_ops" BPF
+	 struct_ops type, which allows a BPF program to fully replace
+	 the builtin linear cost model on the device it is attached
+	 to.  The struct_ops is attached per device, following the
+	 hid_bpf_ops model.
+
 endif # BLOCK
diff --git a/block/Makefile b/block/Makefile
index e7bd320e3d697..ee5cebeea006f 100644
--- a/block/Makefile
+++ b/block/Makefile
@@ -39,3 +39,4 @@ obj-$(CONFIG_BLK_INLINE_ENCRYPTION)	+= blk-crypto.o blk-crypto-profile.o \
 					   blk-crypto-sysfs.o
 obj-$(CONFIG_BLK_INLINE_ENCRYPTION_FALLBACK)	+= blk-crypto-fallback.o
 obj-$(CONFIG_BLOCK_HOLDER_DEPRECATED)	+= holder.o
+obj-$(CONFIG_BLK_CGROUP_IOCOST_BPF)	+= blk-iocost-bpf.o
diff --git a/block/blk-core.c b/block/blk-core.c
index 13dc70e8f55d9..a902736d381a5 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -536,6 +536,22 @@ bool blk_get_queue(struct request_queue *q)
 }
 EXPORT_SYMBOL(blk_get_queue);
 
+/**
+ * blk_get_queue_rcu - get a queue reference regardless of the dying flag
+ * @q: the request_queue to reference
+ *
+ * Unlike blk_get_queue(), this succeeds on a dying queue, so a caller
+ * which holds the queue only through RCU (e.g. a detach path which
+ * read the pointer locklessly) can still take a reference and touch
+ * the queue under its own lifetime.  The caller must hold
+ * rcu_read_lock() so the memory is valid.  Fails only when the
+ * refcount already dropped to zero.
+ */
+bool blk_get_queue_rcu(struct request_queue *q)
+{
+	return refcount_inc_not_zero(&q->refs);
+}
+
 #ifdef CONFIG_FAIL_MAKE_REQUEST
 
 static DECLARE_FAULT_ATTR(fail_make_request);
diff --git a/block/blk-iocost-bpf.c b/block/blk-iocost-bpf.c
new file mode 100644
index 0000000000000..56959c5fc3937
--- /dev/null
+++ b/block/blk-iocost-bpf.c
@@ -0,0 +1,184 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * blk-iocost: BPF struct_ops plumbing for pluggable cost models.
+ *
+ * Registers the "iocost_model_ops" struct_ops type.  Attachment is
+ * per-device and follows the hid_bpf_ops model: the target device is
+ * set in the ops from userspace before load, .reg attaches the model
+ * to that device, creating the ioc if needed and switching to the
+ * model under the same queue freeze and quiesce as io.cost.model
+ * writes, and .unreg detaches it; enabling and disabling the
+ * controller stays with io.cost.qos.  The struct_ops core owns the
+ * program lifetime.
+ */
+#include <linux/init.h>
+#include <linux/kernel.h>
+#include <linux/module.h>
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/blk-iocost.h>
+#include <linux/blk-mq.h>
+#include <linux/mutex.h>
+#include "blk.h"
+
+/* nothing to look up: the core finds the struct by name before calling */
+static int bpf_iocost_model_init(struct btf *btf)
+{
+	return 0;
+}
+
+static bool bpf_iocost_is_valid_access(int off, int size,
+				       enum bpf_access_type type,
+				       const struct bpf_prog *prog,
+				       struct bpf_insn_access_aux *info)
+{
+	return bpf_tracing_btf_ctx_access(off, size, type, prog, info);
+}
+
+/*
+ * No iocost-specific helpers; bpf_base_func_proto already covers the
+ * cgroup storage helpers under CONFIG_CGROUPS.
+ */
+static const struct bpf_func_proto *
+bpf_iocost_get_func_proto(enum bpf_func_id func_id,
+			  const struct bpf_prog *prog)
+{
+	return bpf_base_func_proto(func_id, prog);
+}
+
+static int bpf_iocost_check_member(const struct btf_type *t,
+				   const struct btf_member *member,
+				   const struct bpf_prog *prog)
+{
+	/* every callback runs under RCU read lock or a spinlock */
+	if (prog->sleepable)
+		return -EINVAL;
+	return 0;
+}
+
+static int bpf_iocost_init_member(const struct btf_type *t,
+				  const struct btf_member *member,
+				  void *kdata, const void *udata)
+{
+	struct iocost_model_ops *ops = kdata;
+	const struct iocost_model_ops *uops = udata;
+	u32 moff = __btf_member_bit_offset(t, member) / 8;
+
+	switch (moff) {
+	/*
+	 * bdev, q and link are kernel-private and start zeroed; the
+	 * struct_ops core rejects a nonzero userspace value for
+	 * non-function members this callback does not claim.
+	 */
+	case offsetof(struct iocost_model_ops, dev):
+		ops->dev = uops->dev;
+		return 1;
+	case offsetof(struct iocost_model_ops, read_vtime_per_page):
+		ops->read_vtime_per_page = uops->read_vtime_per_page;
+		return 1;
+	case offsetof(struct iocost_model_ops, write_vtime_per_page):
+		ops->write_vtime_per_page = uops->write_vtime_per_page;
+		return 1;
+	}
+
+	return 0;
+}
+
+/*
+ * kvalue is zeroed at map allocation and function members are only
+ * written when the BPF side provides a prog, so a model which did
+ * not implement calc_cost leaves it NULL.  The dispatch would call
+ * it on every charged IO, so reject it here.
+ */
+static int bpf_iocost_validate(void *kdata)
+{
+	struct iocost_model_ops *ops = kdata;
+
+	return ops->calc_cost ? 0 : -EINVAL;
+}
+
+static int bpf_iocost_reg(void *kdata, struct bpf_link *link)
+{
+	struct iocost_model_ops *ops = kdata;
+
+	if (!ops->dev)
+		return -EINVAL;
+
+	return ioc_bpf_attach(ops, link);
+}
+
+static void bpf_iocost_unreg(void *kdata, struct bpf_link *link)
+{
+	struct iocost_model_ops *ops = kdata;
+	struct request_queue *q;
+
+	/*
+	 * ops->q may already have been cleared by the removal ejection;
+	 * take a queue reference under RCU before entering the queue,
+	 * as the queue may be dying and its memory is only guaranteed
+	 * under rcu_read_lock()
+	 */
+	rcu_read_lock();
+	/* pairs with the smp_store_release() in the attach: a non-NULL
+	 * q guarantees the owning link store is visible */
+	q = smp_load_acquire(&ops->q);
+	if (!q || !blk_get_queue_rcu(q)) {
+		rcu_read_unlock();
+		return;
+	}
+	rcu_read_unlock();
+
+	/*
+	 * detach only if this link still owns the attachment: after a
+	 * device removal the same map may have been attached to a new
+	 * device through another link, and closing the old link must
+	 * not tear that down.  The link is re-checked under
+	 * rq_qos_mutex inside ioc_bpf_detach().
+	 */
+	ioc_bpf_detach(ops, link, q);
+
+	blk_put_queue(q);
+}
+
+static const struct bpf_verifier_ops bpf_iocost_verifier_ops = {
+	.get_func_proto = bpf_iocost_get_func_proto,
+	.is_valid_access = bpf_iocost_is_valid_access,
+};
+
+static u64 bpf_iocost_calc_cost_stub(struct bio *bio, u64 flags)
+{
+	return 0;
+}
+
+static void bpf_iocost_iocg_init_stub(struct blkcg *blkcg,
+				      struct request_queue *q)
+{ }
+static void bpf_iocost_iocg_free_stub(struct blkcg *blkcg,
+				      struct request_queue *q)
+{ }
+
+static struct iocost_model_ops __bpf_ops_iocost_model_ops = {
+	.calc_cost = bpf_iocost_calc_cost_stub,
+	.iocg_init = bpf_iocost_iocg_init_stub,
+	.iocg_free = bpf_iocost_iocg_free_stub,
+};
+
+static struct bpf_struct_ops bpf_iocost_model_ops = {
+	.verifier_ops = &bpf_iocost_verifier_ops,
+	.init = bpf_iocost_model_init,
+	.check_member = bpf_iocost_check_member,
+	.init_member = bpf_iocost_init_member,
+	.validate = bpf_iocost_validate,
+	.reg = bpf_iocost_reg,
+	.unreg = bpf_iocost_unreg,
+	.name = "iocost_model_ops",
+	.cfi_stubs = &__bpf_ops_iocost_model_ops,
+	.owner = THIS_MODULE,
+};
+
+static int __init bpf_iocost_init(void)
+{
+	return register_bpf_struct_ops(&bpf_iocost_model_ops, iocost_model_ops);
+}
+late_initcall(bpf_iocost_init);
diff --git a/block/blk-iocost.c b/block/blk-iocost.c
index 2745bffcd5eef..509ae36f99edf 100644
--- a/block/blk-iocost.c
+++ b/block/blk-iocost.c
@@ -177,6 +177,7 @@
 #include <linux/timer.h>
 #include <linux/time64.h>
 #include <linux/parser.h>
+#include <linux/blk-iocost.h>
 #include <linux/sched/signal.h>
 #include <asm/local.h>
 #include <asm/local64.h>
@@ -445,6 +446,11 @@ struct ioc {
 	int				autop_idx;
 	bool				user_qos_params:1;
 	bool				user_cost_model:1;
+
+	/* the struct_ops attached to this device, NULL = none */
+	const struct iocost_model_ops	__rcu *attached;
+	/* the cost model in use, NULL = builtin linear model */
+	const struct iocost_model_ops	__rcu *model;
 };
 
 struct iocg_pcpu_stat {
@@ -462,6 +468,8 @@ struct iocg_stat {
 struct ioc_gq {
 	struct blkg_policy_data		pd;
 	struct ioc			*ioc;
+	/* whether the attached model's iocg_init() was delivered */
+	bool				bpf_iocg_inited;
 
 	/*
 	 * A iocg can get its weight from two sources - an explicit
@@ -780,6 +788,33 @@ static void ioc_refresh_period_us(struct ioc *ioc)
 	ioc_refresh_margins(ioc);
 }
 
+/*
+ * The BPF model in use, or NULL when the builtin linear model prices
+ * this device; with !CONFIG_BLK_CGROUP_IOCOST_BPF the field is
+ * always NULL, so callers need no #ifdefs.
+ */
+static const struct iocost_model_ops *ioc_model_in_use(struct ioc *ioc)
+{
+	return rcu_dereference(ioc->model);
+}
+
+/*
+ * The struct_ops attached to this device, or NULL.  While attachment
+ * and model selection are independent, the cgroup callbacks follow
+ * the attachment, not the selection.
+ */
+static const struct iocost_model_ops *ioc_attached_or_null(struct ioc *ioc)
+{
+	return rcu_dereference(ioc->attached);
+}
+
+static const struct iocost_model_ops *
+ioc_model_in_use_locked(struct ioc *ioc)
+{
+	return rcu_dereference_protected(ioc->model,
+					 lockdep_is_held(&ioc->lock));
+}
+
 /*
  *  ioc->rqos.disk isn't initialized when this function is called from
  *  the init path.
@@ -803,8 +838,13 @@ static int ioc_autop_idx(struct ioc *ioc, struct gendisk *disk)
 	if (idx < AUTOP_SSD_DFL)
 		return AUTOP_SSD_DFL;
 
-	/* if user is overriding anything, maintain what was there */
-	if (ioc->user_qos_params || ioc->user_cost_model)
+	/*
+	 * if user is overriding anything, or a BPF model is in use,
+	 * maintain what was there: the builtin coefficients are inert
+	 * then, so stepping the profile is pointless
+	 */
+	if (ioc->user_qos_params || ioc->user_cost_model ||
+	    ioc_model_in_use_locked(ioc))
 		return idx;
 
 	/* step up/down based on the vrate */
@@ -2571,8 +2611,19 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg,
 
 static u64 calc_vtime_cost(struct bio *bio, struct ioc_gq *iocg, bool is_merge)
 {
+	const struct iocost_model_ops *model;
 	u64 cost;
 
+	rcu_read_lock();
+	model = ioc_model_in_use(iocg->ioc);
+	if (model) {
+		cost = model->calc_cost(bio,
+				is_merge ? IOCOST_COST_F_MERGE : 0);
+		rcu_read_unlock();
+		return min(cost, VTIME_PER_SEC);
+	}
+	rcu_read_unlock();
+
 	calc_vtime_cost_builtin(bio, iocg, is_merge, &cost);
 	return cost;
 }
@@ -2594,10 +2645,32 @@ static void calc_size_vtime_cost_builtin(struct request *rq, struct ioc *ioc,
 	}
 }
 
+/*
+ * Called from the request completion path, where no ioc->lock is
+ * held; the model pointer is read under RCU, matching the bio-side
+ * calc_vtime_cost().
+ */
 static u64 calc_size_vtime_cost(struct request *rq, struct ioc *ioc)
 {
+	const struct iocost_model_ops *model;
 	u64 cost;
 
+	rcu_read_lock();
+	model = ioc_model_in_use(ioc);
+	if (model && (req_op(rq) == REQ_OP_READ ||
+		      req_op(rq) == REQ_OP_WRITE)) {
+		unsigned int pages =
+			blk_rq_stats_sectors(rq) >> IOC_SECT_TO_PAGE_SHIFT;
+		u64 coeff = req_op(rq) == REQ_OP_READ ?
+			model->read_vtime_per_page :
+			model->write_vtime_per_page;
+
+		cost = pages * coeff;
+		rcu_read_unlock();
+		return cost;
+	}
+	rcu_read_unlock();
+
 	calc_size_vtime_cost_builtin(rq, ioc, &cost);
 	return cost;
 }
@@ -2888,6 +2961,51 @@ static void ioc_rqos_queue_depth_changed(struct rq_qos *rqos)
 	spin_unlock_irq(&ioc->lock);
 }
 
+#ifdef CONFIG_BLK_CGROUP_IOCOST_BPF
+/*
+ * Eject the model on device removal.  The caller holds rq_qos_mutex
+ * and blkcg_deactivate_policy() has already run, so every existing
+ * blkg got iocg_free() from ioc_pd_free() and no walk is needed.
+ *
+ * bdev is read into a local before ops->q is cleared: once q is
+ * NULL, .unreg returns without taking any lock and the map holding
+ * ops can be freed, so ops must not be touched afterwards.
+ */
+static void ioc_bpf_eject(struct ioc *ioc)
+{
+	struct request_queue *q = ioc->rqos.disk->queue;
+	struct iocost_model_ops *ops;
+	struct block_device *bdev;
+
+	spin_lock_irq(&q->queue_lock);
+	spin_lock(&ioc->lock);
+	ops = (struct iocost_model_ops *)rcu_dereference_protected(
+			ioc->attached, lockdep_is_held(&ioc->lock));
+	if (!ops) {
+		spin_unlock(&ioc->lock);
+		spin_unlock_irq(&q->queue_lock);
+		return;
+	}
+	/*
+	 * ioc->attached and ioc->model are left set: the ioc is freed
+	 * right below, so clearing them buys nothing, and any
+	 * ioc_pd_free() which still runs finds the attachment and
+	 * delivers its iocg_free() through the struct_ops image grace
+	 * period instead of silently losing it.
+	 */
+	bdev = ops->bdev;
+	spin_unlock(&ioc->lock);
+	WRITE_ONCE(ops->link, NULL);
+	ops->bdev = NULL;
+	smp_store_release(&ops->q, NULL);
+	spin_unlock_irq(&q->queue_lock);
+
+	blkdev_put_no_open(bdev);
+}
+#else
+static inline void ioc_bpf_eject(struct ioc *ioc) { }
+#endif
+
 static void ioc_rqos_exit(struct rq_qos *rqos)
 {
 	struct ioc *ioc = rqos_to_ioc(rqos);
@@ -2900,6 +3018,10 @@ static void ioc_rqos_exit(struct rq_qos *rqos)
 
 	timer_shutdown_sync(&ioc->timer);
 	free_percpu(ioc->pcpu_stat);
+
+	/* eject the model on device removal */
+	ioc_bpf_eject(ioc);
+
 	kfree(ioc);
 }
 
@@ -3022,6 +3144,7 @@ static void ioc_pd_init(struct blkg_policy_data *pd)
 	struct ioc_now now;
 	struct blkcg_gq *tblkg;
 	unsigned long flags;
+	const struct iocost_model_ops *model;
 
 	ioc_now(ioc, &now);
 
@@ -3048,6 +3171,20 @@ static void ioc_pd_init(struct blkg_policy_data *pd)
 	spin_lock_irqsave(&ioc->lock, flags);
 	weight_updated(iocg, &now);
 	spin_unlock_irqrestore(&ioc->lock, flags);
+
+	/*
+	 * the attached model is RCU-protected: a concurrent detach
+	 * publishes NULL and the struct_ops image survives it by a
+	 * grace period, so the callback is safe inside the read-side
+	 * critical section
+	 */
+	rcu_read_lock();
+	model = ioc_attached_or_null(ioc);
+	if (model && model->iocg_init) {
+		model->iocg_init(blkg->blkcg, ioc->rqos.disk->queue);
+		iocg->bpf_iocg_inited = true;
+	}
+	rcu_read_unlock();
 }
 
 static void iocg_release(struct rcu_head *rcu)
@@ -3066,8 +3203,31 @@ static void ioc_pd_free(struct blkg_policy_data *pd)
 	struct blkcg_gq *blkg = pd_to_blkg(pd);
 	struct ioc *ioc = iocg->ioc;
 	unsigned long flags;
+	const struct iocost_model_ops *model;
 
 	if (ioc) {
+		/*
+		 * gate the free on whether init was delivered, so a
+		 * blkg which appeared between the attach walk and its
+		 * own pd_init does not get an unpaired iocg_free().
+		 * The reverse can still happen to a blkg whose
+		 * radix_tree_insert() failed after pd_init delivered
+		 * iocg_init(): it never reaches q->blkg_list, the
+		 * detach walk misses it and the model may be gone by
+		 * the time this runs.  That window requires the GFP_NOWAIT
+		 * radix insert to fail and is documented in the interface
+		 * comment.
+		 */
+		if (iocg->bpf_iocg_inited) {
+			rcu_read_lock();
+			model = ioc_attached_or_null(ioc);
+			if (model && model->iocg_free)
+				model->iocg_free(blkg->blkcg,
+						 ioc->rqos.disk->queue);
+			rcu_read_unlock();
+			iocg->bpf_iocg_inited = false;
+		}
+
 		spin_lock_irqsave(&ioc->lock, flags);
 
 		if (!list_empty(&iocg->active_list)) {
@@ -3433,17 +3593,21 @@ static u64 ioc_cost_model_prfill(struct seq_file *sf,
 	const char *dname = blkg_dev_name(pd->blkg);
 	struct ioc *ioc = pd_to_iocg(pd)->ioc;
 	u64 *u = ioc->params.i_lcoefs;
+	const struct iocost_model_ops *model;
 
 	if (!dname)
 		return 0;
 
 	spin_lock_irq(&ioc->lock);
-	seq_printf(sf, "%s ctrl=%s model=linear "
+	model = ioc_model_in_use_locked(ioc);
+	seq_printf(sf, "%s ctrl=%s model=%s "
 		   "rbps=%llu rseqiops=%llu rrandiops=%llu "
 		   "wbps=%llu wseqiops=%llu wrandiops=%llu\n",
 		   dname, ioc->user_cost_model ? "user" : "auto",
-		   u[I_LCOEF_RBPS], u[I_LCOEF_RSEQIOPS], u[I_LCOEF_RRANDIOPS],
-		   u[I_LCOEF_WBPS], u[I_LCOEF_WSEQIOPS], u[I_LCOEF_WRANDIOPS]);
+		   model ? "bpf" : "linear",
+		   u[I_LCOEF_RBPS], u[I_LCOEF_RSEQIOPS],
+		   u[I_LCOEF_RRANDIOPS], u[I_LCOEF_WBPS],
+		   u[I_LCOEF_WSEQIOPS], u[I_LCOEF_WRANDIOPS]);
 	spin_unlock_irq(&ioc->lock);
 	return 0;
 }
@@ -3457,6 +3621,265 @@ static int ioc_cost_model_show(struct seq_file *sf, void *v)
 	return 0;
 }
 
+#ifdef CONFIG_BLK_CGROUP_IOCOST_BPF
+/*
+ * Deliver iocg_init()/iocg_free() to the cgroups which already have a
+ * blkg on the queue, the same q->blkg_list walk the blkcg policy
+ * teardown uses.  The queue is frozen and quiesced and blkcg_mutex
+ * serializes against blkg creation and destruction.  Cgroups without
+ * a blkg on the device yet are not missed: their blkg is created
+ * later and ioc_pd_init()/ioc_pd_free() deliver the callbacks then.
+ */
+static void ioc_bpf_walk_iocgs(struct ioc *ioc,
+			       const struct iocost_model_ops *ops, bool init)
+{
+	struct request_queue *q = ioc->rqos.disk->queue;
+	struct blkcg_gq *blkg;
+
+	/*
+	 * blkg_create() holds blkcg_mutex across ioc_pd_init() and the
+	 * list_add(), and blkg_free_workfn() holds blkcg_mutex across
+	 * pd_free and delays the list deletion until after it, so
+	 * holding blkcg_mutex makes the list stable and delivers each
+	 * callback exactly once; queue_lock additionally guards the
+	 * ioc->attached pointer the callbacks are gated on.
+	 */
+	lockdep_assert_held(&q->blkcg_mutex);
+	lockdep_assert_held(&q->queue_lock);
+
+	/*
+	 * the cgroup callbacks are optional; validate() only requires
+	 * calc_cost.  The flag keeps the pairing per cgroup across
+	 * models: an init is only delivered when the previous model's
+	 * free was, and a free is only delivered to a cgroup which
+	 * received an init, from whichever model delivered it
+	 */
+	if (init) {
+		/* parents first, like blkcg_activate_policy() */
+		list_for_each_entry_reverse(blkg, &q->blkg_list, q_node) {
+			if (blkg_to_iocg(blkg) && ops->iocg_init &&
+			    !blkg_to_iocg(blkg)->bpf_iocg_inited) {
+				ops->iocg_init(blkg->blkcg, q);
+				blkg_to_iocg(blkg)->bpf_iocg_inited = true;
+			}
+		}
+	} else {
+		list_for_each_entry(blkg, &q->blkg_list, q_node) {
+			if (!blkg_to_iocg(blkg) ||
+			    !blkg_to_iocg(blkg)->bpf_iocg_inited)
+				continue;
+			if (ops->iocg_free)
+				ops->iocg_free(blkg->blkcg, q);
+			blkg_to_iocg(blkg)->bpf_iocg_inited = false;
+		}
+	}
+}
+
+int ioc_bpf_attach(struct iocost_model_ops *ops, struct bpf_link *link)
+{
+	struct block_device *bdev;
+	struct request_queue *q;
+	struct gendisk *disk;
+	struct ioc *ioc;
+	unsigned int memflags;
+	int ret;
+
+	bdev = blkdev_get_no_open(new_decode_dev(ops->dev), false);
+	if (!bdev)
+		return -ENODEV;
+	q = bdev->bd_queue;
+	disk = bdev->bd_disk;
+
+	if (bdev_is_partition(bdev)) {
+		ret = -EINVAL;
+		goto put;
+	}
+	if (!queue_is_mq(q)) {
+		ret = -EOPNOTSUPP;
+		goto put;
+	}
+
+	/*
+	 * the freeze is taken before rq_qos_mutex; taking the mutex
+	 * around the freeze instead cycles against the io.cost.qos
+	 * write path, which freezes with the mutex held.  Holding
+	 * rq_qos_mutex from the disk_live() check through the model
+	 * switch keeps the ioc from being freed under us.
+	 */
+	memflags = blk_mq_freeze_queue(q);
+
+	mutex_lock(&q->rq_qos_mutex);
+
+	/*
+	 * one map binds one device at a time: the ops carries a single
+	 * device slot, so a second attach while it is bound elsewhere
+	 * would clobber it.  Re-attaching after a detach or the device
+	 * removal ejection is fine, both clear ops->q, and the
+	 * release/acquire protocol below orders that clearing against
+	 * this check even when the ejection runs on another device's
+	 * mutex.
+	 *
+	 * ops->q is the release flag of the ops state: the detach and
+	 * the ejection write ops->link first and ops->q last with
+	 * smp_store_release(), and this acquire read of a NULL q
+	 * guarantees their link clearing is already visible, so the
+	 * link store below cannot be clobbered by a concurrent ejection
+	 * on another device, which no lock here serializes against.
+	 */
+	if (smp_load_acquire(&ops->q)) {
+		mutex_unlock(&q->rq_qos_mutex);
+		blk_mq_unfreeze_queue(q, memflags);
+		ret = -EBUSY;
+		goto put;
+	}
+
+	if (!disk_live(disk)) {
+		mutex_unlock(&q->rq_qos_mutex);
+		blk_mq_unfreeze_queue(q, memflags);
+		ret = -ENODEV;
+		goto put;
+	}
+	ioc = q_to_ioc(q);
+	if (!ioc) {
+		ret = blk_iocost_init(disk);
+		if (ret) {
+			mutex_unlock(&q->rq_qos_mutex);
+			blk_mq_unfreeze_queue(q, memflags);
+			goto put;
+		}
+		ioc = q_to_ioc(q);
+	}
+
+	blk_mq_quiesce_queue(q);
+
+	mutex_lock(&q->blkcg_mutex);
+	spin_lock_irq(&q->queue_lock);
+
+	/*
+	 * the model pointers are published under ioc->lock, nested in
+	 * queue_lock like ioc_pd_init() does, so the ioc->lock readers
+	 * (the timer's autop check, the io.cost.model staging) stay
+	 * synchronized with the writers
+	 */
+	spin_lock(&ioc->lock);
+
+	if (rcu_dereference_protected(ioc->attached,
+				      lockdep_is_held(&ioc->lock))) {
+		spin_unlock(&ioc->lock);
+		spin_unlock_irq(&q->queue_lock);
+		mutex_unlock(&q->blkcg_mutex);
+		blk_mq_unquiesce_queue(q);
+		blk_mq_unfreeze_queue(q, memflags);
+		mutex_unlock(&q->rq_qos_mutex);
+		ret = -EBUSY;
+		goto put;
+	}
+
+	rcu_assign_pointer(ioc->attached, ops);
+	rcu_assign_pointer(ioc->model, ops);
+	spin_unlock(&ioc->lock);
+
+	ops->bdev = bdev;
+	WRITE_ONCE(ops->link, link);
+	smp_store_release(&ops->q, q);
+
+	/* pair iocg_init() with the cgroups which already exist */
+	ioc_bpf_walk_iocgs(ioc, ops, true);
+
+	spin_unlock_irq(&q->queue_lock);
+	mutex_unlock(&q->blkcg_mutex);
+
+	blk_mq_unquiesce_queue(q);
+	blk_mq_unfreeze_queue(q, memflags);
+	mutex_unlock(&q->rq_qos_mutex);
+	return 0;
+
+put:
+	blkdev_put_no_open(bdev);
+	return ret;
+}
+
+/*
+ * Detach a model: switch back to the builtin model when the attached
+ * model is in use, clear the attachment, and deliver iocg_free() to
+ * the cgroups which still exist, under the same freeze and quiesce as
+ * the attach.  The caller has already checked ops->q.
+ */
+void ioc_bpf_detach(struct iocost_model_ops *ops, struct bpf_link *link,
+		   struct request_queue *q)
+{
+	struct block_device *bdev;
+	struct ioc *ioc;
+	unsigned int memflags;
+
+	if (!q)
+		return;
+
+	/* eject may clear ops->bdev on a dying queue concurrently */
+	bdev = READ_ONCE(ops->bdev);
+
+	/*
+	 * the freeze is taken before rq_qos_mutex like in the attach;
+	 * the attachment and the link are re-checked under the mutex
+	 * below, as the ejection or a newer attachment through another
+	 * link may have won the race meanwhile
+	 */
+	memflags = blk_mq_freeze_queue(q);
+	blk_mq_quiesce_queue(q);
+
+	mutex_lock(&q->rq_qos_mutex);
+
+	ioc = q_to_ioc(q);
+	if (!ioc) {
+		mutex_unlock(&q->rq_qos_mutex);
+		blk_mq_unquiesce_queue(q);
+		blk_mq_unfreeze_queue(q, memflags);
+		return;
+	}
+
+	mutex_lock(&q->blkcg_mutex);
+	spin_lock_irq(&q->queue_lock);
+
+	spin_lock(&ioc->lock);
+	if (rcu_dereference_protected(ioc->attached,
+				      lockdep_is_held(&ioc->lock)) != ops ||
+	    READ_ONCE(ops->link) != link) {
+		/* the ejection or a newer link won; nothing to detach */
+		spin_unlock(&ioc->lock);
+		spin_unlock_irq(&q->queue_lock);
+		mutex_unlock(&q->blkcg_mutex);
+		mutex_unlock(&q->rq_qos_mutex);
+		blk_mq_unquiesce_queue(q);
+		blk_mq_unfreeze_queue(q, memflags);
+		return;
+	}
+	if (rcu_dereference_protected(ioc->model,
+				      lockdep_is_held(&ioc->lock)) == ops)
+		rcu_assign_pointer(ioc->model, NULL);
+	rcu_assign_pointer(ioc->attached, NULL);
+	spin_unlock(&ioc->lock);
+	WRITE_ONCE(ops->link, NULL);
+	ops->bdev = NULL;
+	smp_store_release(&ops->q, NULL);
+
+	ioc_bpf_walk_iocgs(ioc, ops, false);
+
+	spin_unlock_irq(&q->queue_lock);
+	mutex_unlock(&q->blkcg_mutex);
+
+	mutex_unlock(&q->rq_qos_mutex);
+
+	blk_mq_unquiesce_queue(q);
+	blk_mq_unfreeze_queue(q, memflags);
+
+	/* drop the attach reference: the detacher owns it */
+	if (bdev)
+		blkdev_put_no_open(bdev);
+}
+
+
+#endif
+
 static const match_table_t cost_ctrl_tokens = {
 	{ COST_CTRL,		"ctrl=%s"	},
 	{ COST_MODEL,		"model=%s"	},
@@ -3484,6 +3907,8 @@ static ssize_t ioc_cost_model_write(struct kernfs_open_file *of, char *input,
 	bool user;
 	char *body, *p;
 	int ret;
+	const struct iocost_model_ops *new_model = NULL;
+	bool model_write = false;
 
 	blkg_conf_init(&ctx, input);
 
@@ -3536,9 +3961,25 @@ static ssize_t ioc_cost_model_write(struct kernfs_open_file *of, char *input,
 			continue;
 		case COST_MODEL:
 			match_strlcpy(buf, &args[0], sizeof(buf));
-			if (strcmp(buf, "linear"))
-				goto unlock;
-			continue;
+			if (!strcmp(buf, "linear")) {
+				/* staged and committed below, so a parse
+				 * error later in the same write leaves
+				 * the model selection untouched
+				 */
+				new_model = NULL;
+				model_write = true;
+				continue;
+			}
+			if (!strcmp(buf, "bpf")) {
+				new_model = rcu_dereference_protected(
+						ioc->attached,
+						lockdep_is_held(&ioc->lock));
+				if (!new_model)
+					goto unlock;
+				model_write = true;
+				continue;
+			}
+			goto unlock;
 		}
 
 		tok = match_token(p, i_lcoef_tokens, args);
@@ -3556,6 +3997,8 @@ static ssize_t ioc_cost_model_write(struct kernfs_open_file *of, char *input,
 	} else {
 		ioc->user_cost_model = false;
 	}
+	if (model_write)
+		rcu_assign_pointer(ioc->model, new_model);
 	ioc_refresh_params(ioc, true);
 
 	ret = 0;
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c532..2ef75df930d6f 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -112,6 +112,7 @@ static inline void blk_wait_io(struct completion *done)
 }
 
 struct block_device *blkdev_get_no_open(dev_t dev, bool autoload);
+bool blk_get_queue_rcu(struct request_queue *q);
 void blkdev_put_no_open(struct block_device *bdev);
 
 bool bvec_try_merge_hw_page(struct request_queue *q, struct bio_vec *bv,
diff --git a/include/linux/blk-iocost.h b/include/linux/blk-iocost.h
new file mode 100644
index 0000000000000..9bcdbdc94a535
--- /dev/null
+++ b/include/linux/blk-iocost.h
@@ -0,0 +1,126 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_BLK_IOCOST_H
+#define _LINUX_BLK_IOCOST_H
+
+#include <linux/types.h>
+#include <linux/blk_types.h>
+#include <linux/blkdev.h>
+
+struct bio;
+struct blkcg;
+struct bpf_link;
+struct request_queue;
+
+/*
+ * Pluggable cost model interface for blk-iocost.
+ *
+ * A BPF struct_ops implementation is attached to one device, identified
+ * by the dev member set from userspace before load, following the
+ * hid_bpf_ops model: the struct_ops core owns the lifetime of the
+ * program.  Attaching creates the ioc if needed, like an io.cost.model
+ * write does, switches the device to the BPF model under the same
+ * queue freeze and quiesce, and delivers iocg_init() to the cgroups
+ * which already exist on the device.  Enabling and disabling the
+ * controller stays with io.cost.qos.  While a model is attached,
+ * io.cost.model selects between it and the builtin model: "model=bpf"
+ * switches to the attached model, "model=linear" switches back to the
+ * builtin model, and neither detaches the struct_ops; only detaching
+ * removes the model.  The builtin linear coefficients are kept while
+ * the BPF model is in use and take effect again when switched back.
+ *
+ * While the BPF model is the one in use, it owns pricing for every
+ * charged IO on the device: it prices all operations, including
+ * flushes, from the bio charging path.
+ *
+ * calc_cost() is called from the IO submission path with RCU read lock
+ * held and must not sleep.  It receives the bio itself so the model
+ * can read whatever it needs (operation flags, size, sector, the
+ * issuing cgroup through bio->bi_blkg).  It returns the cost of the
+ * IO in vtime units, where 1 second of device time equals
+ * VTIME_PER_SEC (2^37, available to BPF programs through vmlinux.h).
+ * The returned value is clamped to 1 second of device time per IO.
+ *
+ * The struct_ops also carries the transfer cost coefficients, vtime
+ * per page for reads and writes: while the BPF model is in use,
+ * these replace the builtin linear coefficients in the
+ * completion-time request sizing of READ and WRITE requests.  Letting
+ * a model take over the QoS side entirely (latency tracking, vrate
+ * control) is left for a later extension.
+ *
+ * The cgroup callbacks are bound to the iocg policy lifetime, one
+ * (cgroup, device) pair per invocation: iocg_init() is delivered on
+ * attach to every cgroup which already exists on the device and to
+ * each one appearing afterwards; iocg_free() is delivered on detach to
+ * every cgroup still existing then, and at policy deactivation time
+ * for the rest, so init and free pair up, with one exception: a
+ * blkg whose radix_tree_insert() failed inside blkg_create() has
+ * had iocg_init() delivered but never reaches the blkg list, so
+ * the detach walk misses it and iocg_free() may not be delivered.
+ * Models which keep per-cgroup state should tolerate this, e.g.
+ * by using cgroup storage which is freed with the cgroup.  Both
+ * callbacks are called under the queue lock or the blkcg lock, or
+ * inside an RCU read-side critical section, and must not sleep.
+ */
+
+/*
+ * iocost-specific call metadata for calc_cost()'s model_flags
+ * argument; the merge indicator is not a property of the bio.
+ * An enum so the value is exported through BTF and BPF models can
+ * use it from vmlinux.h.
+ */
+enum {
+	IOCOST_COST_F_MERGE	= 1 << 0,	/* called from merge path */
+};
+
+struct iocost_model_ops {
+	/*
+	 * target device (major:minor), set from userspace before load
+	 * through the struct_ops map's initial value, in the userspace
+	 * dev_t encoding new_decode_dev() accepts
+	 */
+	dev_t dev;
+
+	/* vtime per page, used by the builtin sizing and vrate logic */
+	u64 read_vtime_per_page;
+	u64 write_vtime_per_page;
+
+	u64 (*calc_cost)(struct bio *bio, u64 model_flags);
+	void (*iocg_init)(struct blkcg *blkcg, struct request_queue *q);
+	void (*iocg_free)(struct blkcg *blkcg, struct request_queue *q);
+
+	/* private: */
+
+	/*
+	 * bdev reference held while attached; dropped by the removal
+	 * ejection or .unreg, whichever detaches the model first
+	 */
+	struct block_device	*bdev;
+	/* queue of the attached device, NULL = not attached */
+	struct request_queue	*q;
+	/*
+	 * the link which owns the attachment; .unreg only detaches
+	 * when it matches, so a map re-attached through a newer link
+	 * is not torn down by closing an old one
+	 */
+	struct bpf_link		*link;
+};
+
+#ifdef CONFIG_BLK_CGROUP_IOCOST_BPF
+
+int ioc_bpf_attach(struct iocost_model_ops *ops, struct bpf_link *link);
+void ioc_bpf_detach(struct iocost_model_ops *ops,
+			 struct bpf_link *link, struct request_queue *q);
+
+#else	/* CONFIG_BLK_CGROUP_IOCOST_BPF */
+
+static inline int ioc_bpf_attach(struct iocost_model_ops *ops,
+				 struct bpf_link *link)
+{
+	return -EOPNOTSUPP;
+}
+static inline void ioc_bpf_detach(struct iocost_model_ops *ops,
+				   struct bpf_link *link,
+				   struct request_queue *q) { }
+
+#endif	/* CONFIG_BLK_CGROUP_IOCOST_BPF */
+#endif	/* _LINUX_BLK_IOCOST_H */
-- 
2.43.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [RFC PATCH v10 2/4] selftests/bpf: add iocost cost model test
  2026-10-03 12:40 [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
@ 2026-10-03 12:40 ` Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 3/4] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 4/4] docs: cgroup-v2: document the iocost BPF cost model attachment Tao Cui
  3 siblings, 0 replies; 7+ messages in thread
From: Tao Cui @ 2026-10-03 12:40 UTC (permalink / raw)
  To: tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov

Add two example iocost cost models and tests which attach them to
one device, given as major:minor in $IOCOST_TEST_DEV.

The first model implements the builtin linear HDD formula at double
cost.  The second implements the multi-stream sequentiality
detection: it keeps a per-cgroup table of stream slots instead of
the single builtin cursor, so
concurrent sequential streams in one cgroup are priced sequentially.
Both set the target device through the struct_ops member before
load, and attaching the struct_ops binds the model to the device;
the maps are declared ".struct_ops.link" so that a test dying in
between does not leave the model attached.

The test verifies the model=bpf readback while attached, that
ctrl=bpf is rejected, that a second model on the same device fails
with -EBUSY, that model=linear and model=bpf switch between the
attached and builtin models without detaching, and that detaching
restores the builtin model.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 tools/testing/selftests/bpf/config            |   2 +
 .../selftests/bpf/prog_tests/iocost_model.c   | 233 ++++++++++++++++++
 .../selftests/bpf/progs/iocost_model.c        | 141 +++++++++++
 tools/testing/selftests/bpf/progs/iocost_ms.c | 163 ++++++++++++
 4 files changed, 539 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/iocost_model.c
 create mode 100644 tools/testing/selftests/bpf/progs/iocost_model.c
 create mode 100644 tools/testing/selftests/bpf/progs/iocost_ms.c

diff --git a/tools/testing/selftests/bpf/config b/tools/testing/selftests/bpf/config
index 2b883b388f90c..c5634ee91e690 100644
--- a/tools/testing/selftests/bpf/config
+++ b/tools/testing/selftests/bpf/config
@@ -140,3 +140,5 @@ CONFIG_SMC_HS_CTRL_BPF=y
 CONFIG_DIBS=y
 CONFIG_DIBS_LO=y
 CONFIG_PM_WAKELOCKS=y
+CONFIG_BLK_CGROUP_IOCOST=y
+CONFIG_BLK_CGROUP_IOCOST_BPF=y
diff --git a/tools/testing/selftests/bpf/prog_tests/iocost_model.c b/tools/testing/selftests/bpf/prog_tests/iocost_model.c
new file mode 100644
index 0000000000000..fa3ba82a840d5
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/iocost_model.c
@@ -0,0 +1,233 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+#include <fcntl.h>
+#include <sys/sysmacros.h>
+#include <unistd.h>
+#include "iocost_model.skel.h"
+#include "iocost_ms.skel.h"
+
+/*
+ * Read back the io.cost.model line of dev and copy the model= value
+ * into @model.  Returns 0 on success.
+ */
+static int readback_model(const char *dev, char *model, size_t model_sz)
+{
+	char line[256], word[256], *m, *end;
+	FILE *fp;
+	int found = 0;
+
+	fp = fopen("/sys/fs/cgroup/io.cost.model", "r");
+	if (!fp)
+		return -1;
+	while (fgets(line, sizeof(line), fp)) {
+		if (sscanf(line, "%255s", word) == 1 && !strcmp(word, dev)) {
+			found = 1;
+			break;
+		}
+	}
+	fclose(fp);
+	if (!found)
+		return -1;
+
+	m = strstr(line, "model=");
+	if (!m)
+		return -1;
+	m += strlen("model=");
+	end = m;
+	while (*end && *end != ' ')
+		end++;
+	snprintf(model, model_sz, "%.*s", (int)(end - m), m);
+	return 0;
+}
+
+/*
+ * Write a line to io.cost.model for dev and return the errno of the
+ * write negated (0 on success), as ASSERT_ERR() expects.
+ */
+static int write_cost_model(const char *dev, const char *what)
+{
+	char buf[128];
+	FILE *fp;
+	int err = 0;
+
+	snprintf(buf, sizeof(buf), "%s %s", dev, what);
+	fp = fopen("/sys/fs/cgroup/io.cost.model", "w");
+	if (!fp)
+		return -EACCES;
+	if (fwrite(buf, 1, strlen(buf), fp) != strlen(buf))
+		err = ferror(fp) ? -errno : -EIO;
+	if (fclose(fp) && !err)
+		err = -errno;
+	return err;
+}
+
+/*
+ * Attach the example model to one device, given as major:minor in
+ * $IOCOST_TEST_DEV: the dev member is written through the struct_ops
+ * map's initial value before load, as hid_bpf_ops does with hid_id,
+ * and attaching the struct_ops attaches the model to the device.
+ * Detaching the struct_ops restores the builtin model.
+ *
+ * Requires root, cgroup v2 and a device with iocost support.
+ */
+void serial_test_iocost_model(void)
+{
+	struct iocost_model *skel, *second;
+	unsigned int maj, min;
+	int err;
+	char model[32], *dev;
+
+	dev = getenv("IOCOST_TEST_DEV");
+	if (!dev || geteuid() != 0 || sscanf(dev, "%u:%u", &maj, &min) != 2) {
+		printf("%s:SKIP:needs root and IOCOST_TEST_DEV=<maj>:<min>\n",
+		       __func__);
+		test__skip();
+		return;
+	}
+
+	skel = iocost_model__open();
+	if (!ASSERT_OK_PTR(skel, "skel_open"))
+		return;
+
+	skel->struct_ops.iocost_2x->dev = (__u32)makedev(maj, min);
+
+	err = iocost_model__load(skel);
+	if (!ASSERT_OK(err, "skel_load")) {
+		iocost_model__destroy(skel);
+		return;
+	}
+
+	err = iocost_model__attach(skel);
+	if (err == -EOPNOTSUPP || err == -ENODEV) {
+		printf("%s:SKIP:device is not blk-mq, has no iocost or does not exist\n",
+		       __func__);
+		iocost_model__destroy(skel);
+		test__skip();
+		return;
+	}
+	if (ASSERT_OK(err, "attach")) {
+		/*
+		 * attached: the read path reports model=bpf until the
+		 * struct_ops is detached; ctrl keeps describing the
+		 * builtin coefficients
+		 */
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback"))
+			ASSERT_EQ(strcmp(model, "bpf"), 0, "model_bpf");
+
+		/* a second model on the same device fails with -EBUSY */
+		second = iocost_model__open();
+		if (ASSERT_OK_PTR(second, "second_open")) {
+			second->struct_ops.iocost_2x->dev =
+				(__u32)makedev(maj, min);
+			err = iocost_model__load(second);
+			if (ASSERT_OK(err, "second_load")) {
+				struct bpf_link *l2;
+
+				/*
+				 * the kernel rejects attaching a second
+				 * model to the device with EBUSY
+				 */
+				l2 = bpf_map__attach_struct_ops(
+						second->maps.iocost_2x);
+				if (!ASSERT_ERR_PTR(l2, "second_ebusy"))
+					bpf_link__destroy(l2);
+				else
+					ASSERT_EQ(libbpf_get_error(l2), -EBUSY,
+						  "second_ebusy_errno");
+			}
+			iocost_model__destroy(second);
+		}
+
+		/*
+		 * attachment and model selection are independent:
+		 * ctrl=bpf is never accepted, model=linear switches
+		 * back to the builtin without detaching, model=bpf
+		 * switches to the attached model again
+		 */
+		err = write_cost_model(dev, "ctrl=bpf");
+		ASSERT_ERR(err, "ctrl_bpf_rejected");
+
+		err = write_cost_model(dev, "model=linear");
+		ASSERT_OK(err, "model_linear_write");
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback_linear"))
+			ASSERT_EQ(strcmp(model, "linear"), 0, "model_linear");
+
+		/* still attached: switching back to the model works */
+		err = write_cost_model(dev, "model=bpf");
+		ASSERT_OK(err, "model_bpf_write");
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback_bpf"))
+			ASSERT_EQ(strcmp(model, "bpf"), 0, "model_bpf");
+
+		bpf_link__destroy(skel->links.iocost_2x);
+		/* cleared, or the skeleton destroy below detaches again */
+		skel->links.iocost_2x = NULL;
+
+		/* after detach, model=bpf fails: nothing is attached */
+		err = write_cost_model(dev, "model=bpf");
+		ASSERT_ERR(err, "model_bpf_after_detach");
+
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback_after_detach"))
+			ASSERT_EQ(strcmp(model, "linear"), 0, "model_linear");
+	}
+
+	iocost_model__destroy(skel);
+}
+/*
+ * Same check for the multi-stream example model.  Only one model can
+ * be attached to a device at a time; both tests attach and detach, so
+ * they are serial.
+ */
+void serial_test_iocost_model_streams(void)
+{
+	struct iocost_ms *skel;
+	unsigned int maj, min;
+	int err;
+	char model[32], *dev;
+
+	dev = getenv("IOCOST_TEST_DEV");
+	if (!dev || geteuid() != 0 || sscanf(dev, "%u:%u", &maj, &min) != 2) {
+		printf("%s:SKIP:needs root and IOCOST_TEST_DEV=<maj>:<min>\n",
+		       __func__);
+		test__skip();
+		return;
+	}
+
+	skel = iocost_ms__open();
+	if (!ASSERT_OK_PTR(skel, "skel_open"))
+		return;
+
+	skel->struct_ops.iocost_ms->dev = (__u32)makedev(maj, min);
+
+	err = iocost_ms__load(skel);
+	if (!ASSERT_OK(err, "skel_load")) {
+		iocost_ms__destroy(skel);
+		return;
+	}
+
+	err = iocost_ms__attach(skel);
+	if (err == -EOPNOTSUPP || err == -ENODEV) {
+		printf("%s:SKIP:device is not blk-mq, has no iocost or does not exist\n",
+		       __func__);
+		iocost_ms__destroy(skel);
+		test__skip();
+		return;
+	}
+	if (ASSERT_OK(err, "attach")) {
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback"))
+			ASSERT_EQ(strcmp(model, "bpf"), 0, "model_bpf");
+
+		bpf_link__destroy(skel->links.iocost_ms);
+		skel->links.iocost_ms = NULL;
+
+		err = readback_model(dev, model, sizeof(model));
+		if (ASSERT_OK(err, "readback_after_detach"))
+			ASSERT_EQ(strcmp(model, "linear"), 0, "model_linear");
+	}
+
+	iocost_ms__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/iocost_model.c b/tools/testing/selftests/bpf/progs/iocost_model.c
new file mode 100644
index 0000000000000..e4414d4802e2b
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/iocost_model.c
@@ -0,0 +1,141 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Example iocost cost model: the builtin linear HDD formula with all
+ * costs doubled, for one device given by the dev member of the
+ * struct_ops.
+ *
+ * The constants mirror what calc_lcoefs() derives from the AUTOP_HDD
+ * defaults (rbps=174019176 rseqiops=41708 rrandiops=370, w-side
+ * analog) in vtime units where 1s == 2^37.  On a rotational device
+ * still on ctrl=auto, a device with this model attached charges
+ * twice the builtin model under the same read/write workload
+ * (the builtin prices flushes as zero; this model prices them
+ * as one-page writes), so the
+ * doubled cost is a direct check that accounting goes through the
+ * BPF path.  On a non-rotational device, or one with user-pinned
+ * coefficients, the ratio to the builtin model is arbitrary.
+ *
+ * The cursor handling follows the builtin: a zero cursor means "no
+ * previous IO", and the cursor is advanced for bios the builtin
+ * prices (READ/WRITE with a non-zero size; the builtin skips the
+ * whole update when the cost is zero), so flushes and discards
+ * leave it alone and long merged streams do not drift past the
+ * 16MB seek threshold.
+ *
+ * The model implements the full linear formula itself, including
+ * flushes: there is no fallback to the builtin model, which prices
+ * a dataless WRITE|REQ_PREFLUSH as zero; this model prices it as
+ * a one-page sequential write.
+ */
+
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+/*
+ * VTIME_PER_SEC, IOC_PAGE_SIZE/SHIFT, IOC_SECT_TO_PAGE_SHIFT and
+ * IOCOST_COST_F_MERGE come from vmlinux.h (BTF enum constants)
+ */
+#define LCOEF_RANDIO_PAGES	4096	/* 16MB seek threshold */
+#define IOCOST_REQ_OP_MASK	0xff		/* REQ_OP_MASK, not in BTF */
+
+/*
+ * DIV64_U64_ROUND_UP / DIV_ROUND_UP_ULL equivalents, folded at
+ * compile time
+ */
+#define RU(x, y)		((x) / (y) + (((x) % (y)) ? 1 : 0))
+
+#define RBPS	174019176ULL
+#define RSEQIOPS	41708ULL
+#define RRANDIOPS	370ULL
+#define WBPS	178075866ULL
+#define WSEQIOPS	42705ULL
+#define WRANDIOPS	378ULL
+
+#define RPAGE	(RU(VTIME_PER_SEC, RU(RBPS, IOC_PAGE_SIZE)))
+#define RSEQIO	(RU(VTIME_PER_SEC, RSEQIOPS) - RPAGE)
+#define RRANDIO	(RU(VTIME_PER_SEC, RRANDIOPS) - RPAGE)
+#define WPAGE	(RU(VTIME_PER_SEC, RU(WBPS, IOC_PAGE_SIZE)))
+#define WSEQIO	(RU(VTIME_PER_SEC, WSEQIOPS) - WPAGE)
+#define WRANDIO	(RU(VTIME_PER_SEC, WRANDIOPS) - WPAGE)
+
+/*
+ * per-cgroup cursor storage: keyed by the cgroup and freed with it,
+ * so a cgroup that comes back starts with no cursor
+ */
+struct {
+	__uint(type, BPF_MAP_TYPE_CGRP_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, __u64);
+} cursor_store SEC(".maps");
+
+SEC("struct_ops")
+u64 BPF_PROG(iocost_2x_calc_cost, struct bio *bio, u64 model_flags)
+{
+	u64 opf = bio->bi_opf, nbytes = bio->bi_iter.bi_size;
+	u64 sector = bio->bi_iter.bi_sector;
+	struct blkcg *blkcg = bio->bi_blkg->blkcg;
+	u64 pages, seek_pages = 0, base, coef_page, randio, cost;
+	__u64 *cursor, cur;
+	int priced;
+
+	/* builtin truncates: max(sectors >> IOC_SECT_TO_PAGE_SHIFT, 1) */
+	pages = nbytes >> IOC_PAGE_SHIFT;
+	if (!pages)
+		pages = 1;
+
+	if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ) {
+		base = RSEQIO; coef_page = RPAGE; randio = RRANDIO;
+	} else if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE) {
+		base = WSEQIO; coef_page = WPAGE; randio = WRANDIO;
+	} else {
+		/*
+		 * a fully owning model must price every op; unknown
+		 * ops are priced per page at the write coefficient
+		 */
+		base = 0; coef_page = WPAGE; randio = 0;
+	}
+
+	/* mirror the builtin cursor semantics described above */
+	priced = (opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ ||
+		 (opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE;
+	cursor = bpf_cgrp_storage_get(&cursor_store,
+				      blkcg->css.cgroup, NULL,
+				      BPF_LOCAL_STORAGE_GET_F_CREATE);
+	if (!cursor) {
+		if (model_flags & IOCOST_COST_F_MERGE)
+			base = 0;
+		return 2 * (base + pages * coef_page);
+	}
+	cur = *cursor;
+	/*
+	 * a size-zero IO never moves the cursor, so it must not be
+	 * judged against it either; the builtin skips it entirely,
+	 * this model prices it as a one-page sequential write
+	 */
+	if (cur && priced && nbytes) {
+		seek_pages = sector > cur ? sector - cur
+					   : cur - sector;
+		seek_pages >>= IOC_SECT_TO_PAGE_SHIFT;
+		if (seek_pages > LCOEF_RANDIO_PAGES)
+			base = randio;
+	}
+	if (priced && nbytes)
+		*cursor = sector + (nbytes >> 9);
+
+	if (model_flags & IOCOST_COST_F_MERGE)
+		base = 0;
+
+	cost = 2 * (base + pages * coef_page);
+	return cost;
+}
+
+SEC(".struct_ops.link") /* detach on close */
+struct iocost_model_ops iocost_2x = {
+	.read_vtime_per_page = 2 * RPAGE,
+	.write_vtime_per_page = 2 * WPAGE,
+	.calc_cost = (void *)iocost_2x_calc_cost,
+};
+
+char LICENSE[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/iocost_ms.c b/tools/testing/selftests/bpf/progs/iocost_ms.c
new file mode 100644
index 0000000000000..c4bd5138a95ed
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/iocost_ms.c
@@ -0,0 +1,163 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Example multi-stream sequentiality detection cost model.
+ *
+ * The builtin model keeps a single cursor per cgroup, so two
+ * interleaved sequential readers in one cgroup are all priced random
+ * (measured 89x overcharge, 12.9x throughput collapse), while random
+ * IO inside a hot window smaller than the 16MB seek threshold is
+ * priced sequential (measured 107x undercharge).  This model replaces
+ * the single cursor with a per-cgroup table of stream slots: an IO is
+ * sequential iff its sector matches the expected next sector of any
+ * tracked stream.  Interleaved streams keep their own slots, and
+ * windowed random IO rarely matches a moving expectation.
+ *
+ * Stream state lives in a CGRP_STORAGE map, so it is created and
+ * freed with the cgroup.  The model implements the full builtin
+ * linear formula itself, including flush pricing.
+ */
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+/*
+ * VTIME_PER_SEC, IOC_PAGE_SIZE/SHIFT, IOC_SECT_TO_PAGE_SHIFT and
+ * IOCOST_COST_F_MERGE come from vmlinux.h (BTF enum constants)
+ */
+#define IOCOST_REQ_OP_MASK	0xff		/* REQ_OP_MASK, not in BTF */
+
+/*
+ * DIV64_U64_ROUND_UP / DIV_ROUND_UP_ULL equivalents, folded at
+ * compile time
+ */
+#define RU(x, y)		((x) / (y) + (((x) % (y)) ? 1 : 0))
+
+#define RBPS	174019176ULL
+#define RSEQIOPS	41708ULL
+#define RRANDIOPS	370ULL
+#define WBPS	178075866ULL
+#define WSEQIOPS	42705ULL
+#define WRANDIOPS	378ULL
+
+#define RPAGE	(RU(VTIME_PER_SEC, RU(RBPS, IOC_PAGE_SIZE)))
+#define RSEQIO	(RU(VTIME_PER_SEC, RSEQIOPS) - RPAGE)
+#define RRANDIO	(RU(VTIME_PER_SEC, RRANDIOPS) - RPAGE)
+#define WPAGE	(RU(VTIME_PER_SEC, RU(WBPS, IOC_PAGE_SIZE)))
+#define WSEQIO	(RU(VTIME_PER_SEC, WSEQIOPS) - WPAGE)
+#define WRANDIO	(RU(VTIME_PER_SEC, WRANDIOPS) - WPAGE)
+
+#define NSLOTS	4
+
+struct streams {
+	__u64 expected[NSLOTS];	/* next expected sector, per stream */
+	__u64 stamp[NSLOTS];	/* LRU stamp, 0 = empty */
+};
+
+/*
+ * per-cgroup stream table: keyed by the cgroup, freed with it
+ */
+struct {
+	__uint(type, BPF_MAP_TYPE_CGRP_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, struct streams);
+} stream_tab SEC(".maps");
+
+SEC("struct_ops")
+u64 BPF_PROG(iocost_ms_calc_cost, struct bio *bio, u64 model_flags)
+{
+	u64 opf = bio->bi_opf, nbytes = bio->bi_iter.bi_size;
+	u64 sector = bio->bi_iter.bi_sector;
+	struct blkcg *blkcg = bio->bi_blkg->blkcg;
+	struct streams *s;
+	u64 pages, base, coef_page, randio, advance, now;
+	u32 i, victim = 0, found = 0xFFFFFFFF;
+
+	if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ) {
+		base = RSEQIO; coef_page = RPAGE; randio = RRANDIO;
+	} else if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE) {
+		base = WSEQIO; coef_page = WPAGE; randio = WRANDIO;
+	} else {
+		/*
+		 * a fully owning model must price every op; unknown
+		 * ops are priced as per-page writes
+		 */
+		base = 0; coef_page = WPAGE; randio = 0;
+	}
+	advance = nbytes >> 9;	/* sectors, truncated like the builtin */
+
+	/*
+	 * only bios the builtin prices participate in stream tracking;
+	 * the early return below skips the merge discount, which only
+	 * applies to tracked data bios
+	 */
+	if (!(((opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ ||
+	       (opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE) && nbytes)) {
+		pages = nbytes >> IOC_PAGE_SHIFT;
+		if (!pages)
+			pages = 1;
+		return base + pages * coef_page;
+	}
+
+	s = bpf_cgrp_storage_get(&stream_tab, blkcg->css.cgroup, NULL,
+				 BPF_LOCAL_STORAGE_GET_F_CREATE);
+	if (!s) {
+		/* no storage: price per page, truncating like the builtin */
+		pages = nbytes >> IOC_PAGE_SHIFT;
+		if (!pages)
+			pages = 1;
+		return base + pages * coef_page;
+	}
+
+	/*
+	 * Slot access is lockless, mirroring the builtin single-cursor
+	 * update in ioc_rqos_throttle(): concurrent CPUs submitting for
+	 * the same cgroup can race on slot updates; mispricing is
+	 * bounded and acceptable for an example model.
+	 */
+	now = bpf_ktime_get_ns();
+	for (i = 0; i < NSLOTS; i++) {
+		if (s->expected[i] == sector && s->stamp[i]) {
+			found = i;
+			break;
+		}
+	}
+	if (found != 0xFFFFFFFF) {
+		/* sequential: keep the seq base from the op branch */
+		s->expected[found] = sector + advance;
+		s->stamp[found] = now;
+	} else {
+		base = randio;
+		for (i = 1; i < NSLOTS; i++) {
+			if (s->stamp[i] < s->stamp[victim])
+				victim = i;
+		}
+		s->expected[victim] = sector + advance;
+		s->stamp[victim] = now;
+	}
+
+	/* builtin truncates: max(sectors >> IOC_SECT_TO_PAGE_SHIFT, 1) */
+	pages = nbytes >> IOC_PAGE_SHIFT;
+	if (!pages)
+		pages = 1;
+	if (model_flags & IOCOST_COST_F_MERGE) {
+		/*
+		 * merged bios skip the base cost but still advance
+		 * the stream position above, so a merge at the
+		 * expected sector does not make the following new IO
+		 * look random
+		 */
+		base = 0;
+	}
+
+	return base + pages * coef_page;
+}
+
+SEC(".struct_ops.link") /* detach on close */
+struct iocost_model_ops iocost_ms = {
+	.read_vtime_per_page = RPAGE,
+	.write_vtime_per_page = WPAGE,
+	.calc_cost = (void *)iocost_ms_calc_cost,
+};
+
+char LICENSE[] SEC("license") = "GPL";
-- 
2.43.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [RFC PATCH v10 3/4] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary
  2026-10-03 12:40 [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 2/4] selftests/bpf: add iocost cost model test Tao Cui
@ 2026-10-03 12:40 ` Tao Cui
  2026-10-03 12:40 ` [RFC PATCH v10 4/4] docs: cgroup-v2: document the iocost BPF cost model attachment Tao Cui
  3 siblings, 0 replies; 7+ messages in thread
From: Tao Cui @ 2026-10-03 12:40 UTC (permalink / raw)
  To: tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov

Add iocost_ioc_tick, emitted once per period from ioc_timer_fn()
with the overall controller state: period_us, vrate, busy_level,
active iocg count, usage percentage and whether the controller
still has active iocgs.  It reads the values the period ran in:
the event is emitted before the vrate adjustment and the period
transition, so each tick reports the completed period.  ioc->running
has not transitioned to IOC_IDLE at the emit point, so the running
field reports the active iocg list instead.

It fires every period the controller is running, including steady
states, plus one final tick before the controller goes idle, which
reads active=0 running=0, making the transition to dormancy (e.g. a
device saturated entirely by uncharged IO) directly visible.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 block/blk-iocost.c            | 14 +++++++++++
 include/trace/events/iocost.h | 45 +++++++++++++++++++++++++++++++++++
 2 files changed, 59 insertions(+)

diff --git a/block/blk-iocost.c b/block/blk-iocost.c
index 509ae36f99edf..a565892b1be07 100644
--- a/block/blk-iocost.c
+++ b/block/blk-iocost.c
@@ -2278,6 +2278,7 @@ static void ioc_timer_fn(struct timer_list *timer)
 	struct ioc_now now;
 	LIST_HEAD(surpluses);
 	int nr_debtors, nr_shortages = 0, nr_lagging = 0;
+	int nr_active = 0;
 	u64 usage_us_sum = 0;
 	u32 ppm_rthr;
 	u32 ppm_wthr;
@@ -2314,6 +2315,8 @@ static void ioc_timer_fn(struct timer_list *timer)
 		u64 vdone, vtime, usage_us;
 		u32 hw_active, hw_inuse;
 
+		nr_active++;
+
 		/*
 		 * Collect unused and wind vtime closer to vnow to prevent
 		 * iocgs from accumulating a large amount of budget.
@@ -2475,6 +2478,17 @@ static void ioc_timer_fn(struct timer_list *timer)
 
 	ioc->busy_level = clamp(ioc->busy_level, -1000, 1000);
 
+	/*
+	 * Everything the tick reports is final here: busy_level was just
+	 * computed, cur_period hasn't changed, nr_active and usage_us_sum
+	 * are complete, and vrate and period_us still hold the values
+	 * this period ran in.  Emit before the refresh below so the
+	 * event reads the completed period directly; ioc->running has
+	 * not transitioned to IOC_IDLE yet, so the running field reports
+	 * the active iocg list instead.
+	 */
+	trace_iocost_ioc_tick(ioc, &now, nr_active, usage_us_sum);
+
 	ioc_adjust_base_vrate(ioc, rq_wait_pct, nr_lagging, nr_shortages,
 			      prev_busy_level, missed_ppm);
 
diff --git a/include/trace/events/iocost.h b/include/trace/events/iocost.h
index e772b1bc60d60..dc14c574894de 100644
--- a/include/trace/events/iocost.h
+++ b/include/trace/events/iocost.h
@@ -178,6 +178,51 @@ TRACE_EVENT(iocost_ioc_vrate_adj,
 	)
 );
 
+/*
+ * Periodic per-device summary, emitted once per period from the tail of
+ * ioc_timer_fn().  Unlike the state-change events above, this fires every
+ * period the controller is running, including steady states, and carries
+ * the overall controller state so basic monitoring doesn't require drgn.
+ */
+TRACE_EVENT(iocost_ioc_tick,
+
+	TP_PROTO(struct ioc *ioc, const struct ioc_now *now,
+		 int nr_active, u64 usage_us_sum),
+
+	TP_ARGS(ioc, now, nr_active, usage_us_sum),
+
+	TP_STRUCT__entry (
+		__string(devname, ioc_name(ioc))
+		__field(u64, cur_period)
+		__field(u32, period_us)
+		__field(u64, vrate)
+		__field(int, busy_level)
+		__field(int, nr_active)
+		__field(u32, usage_pct)
+		/* whether the controller still has active iocgs */
+		__field(int, running)
+	),
+
+	TP_fast_assign(
+		__assign_str(devname);
+		__entry->cur_period = atomic64_read(&ioc->cur_period);
+		__entry->period_us = ioc->period_us;
+		__entry->vrate = ioc->vtime_base_rate;
+		__entry->busy_level = ioc->busy_level;
+		__entry->nr_active = nr_active;
+		__entry->usage_pct = now->now > ioc->period_at ?
+			div64_u64(usage_us_sum * 100,
+				  now->now - ioc->period_at) : 0;
+		__entry->running = !list_empty(&ioc->active_iocgs);
+	),
+
+	TP_printk("[%s] period=%llu:%uus vrate=%llu busy=%d active=%d usage=%u%% running=%d",
+		__get_str(devname), __entry->cur_period, __entry->period_us,
+		__entry->vrate, __entry->busy_level, __entry->nr_active,
+		__entry->usage_pct, __entry->running
+	)
+);
+
 TRACE_EVENT(iocost_iocg_forgive_debt,
 
 	TP_PROTO(struct ioc_gq *iocg, const char *path, struct ioc_now *now,
-- 
2.43.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [RFC PATCH v10 4/4] docs: cgroup-v2: document the iocost BPF cost model attachment
  2026-10-03 12:40 [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
                   ` (2 preceding siblings ...)
  2026-10-03 12:40 ` [RFC PATCH v10 3/4] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary Tao Cui
@ 2026-10-03 12:40 ` Tao Cui
  3 siblings, 0 replies; 7+ messages in thread
From: Tao Cui @ 2026-10-03 12:40 UTC (permalink / raw)
  To: tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov

Document the BPF cost model attachment in the io.cost.model section
of the cgroup v2 documentation: attaching an iocost_model_ops
struct_ops to a device by its major:minor, the model=bpf readback
while attached (ctrl keeps describing the coefficients), that
detaching restores the builtin model, and that ctrl= writes never
select a model.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 Documentation/admin-guide/cgroup-v2.rst | 22 ++++++++++++++++++++++
 1 file changed, 22 insertions(+)

diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-guide/cgroup-v2.rst
index 7fe950425216c..f883d1c9e67d7 100644
--- a/Documentation/admin-guide/cgroup-v2.rst
+++ b/Documentation/admin-guide/cgroup-v2.rst
@@ -2117,8 +2117,30 @@ IO Interface Files
 	  =====		================================
 	  ctrl		"auto" or "user"
 	  model		The cost model in use - "linear"
+			or "bpf" while the BPF model
+			is in use
 	  =====		================================
 
+	When CONFIG_BLK_CGROUP_IOCOST_BPF is enabled, a BPF cost model
+	can be bound to a device by attaching an "iocost_model_ops"
+	struct_ops carrying the whole disk's major:minor in its "dev"
+	member (a partition's major:minor is rejected).  Attaching
+	creates the controller if needed, like an io.cost.model write
+	does, and
+	switches the device's pricing to the model; enabling and
+	disabling the controller stays with io.cost.qos.  While a model
+	is attached, "model" selects between it and the builtin model:
+	"model=bpf" switches to the attached model and "model=linear"
+	switches back to the builtin model, and neither detaches the
+	struct_ops; only detaching removes the model, after which
+	"model=bpf" fails and the device prices with the builtin model
+	again.  Device removal ejects the attached model as well, while
+	the struct_ops link remains until userspace destroys it.  "ctrl"
+	accepts only "auto" and "user" and never selects a model; it
+	keeps describing the builtin coefficients, which take effect
+	again when switched back, and the automatic profile stepping
+	does not switch profiles while the BPF model is in use.
+
 	When "ctrl" is "auto", the kernel may change all parameters
 	dynamically.  When "ctrl" is set to "user" or any other
 	parameters are written to, "ctrl" become "user" and the
-- 
2.43.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [RFC PATCH v10 1/4] blk-iocost: add BPF struct_ops cost model support
  2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
@ 2026-10-03 13:21   ` bot+bpf-ci
  2026-10-03 20:45   ` Alexei Starovoitov
  1 sibling, 0 replies; 7+ messages in thread
From: bot+bpf-ci @ 2026-10-03 13:21 UTC (permalink / raw)
  To: cui.tao, tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, ast,
	daniel, linux-kselftest, cui.tao, cuitao, ameryhung,
	alexei.starovoitov, ast, andrii, daniel, martin.lau, eddyz87,
	yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2184 bytes --]

> blk-iocost: add BPF struct_ops cost model support

TL;DR: Documentation issue only: cgroup-v2.rst isn't updated for io.cost.model
accepting "model=bpf", and writing "model=linear" now switches a device off an
attached BPF model, which may surprise existing tools.

> diff --git a/block/blk-iocost.c b/block/blk-iocost.c
> index 2745bffcd5ee..509ae36f99ed 100644
> --- a/block/blk-iocost.c
> +++ b/block/blk-iocost.c

[ ... ]

> @@ -3536,9 +3961,25 @@ static ssize_t ioc_cost_model_write(struct kernfs_open_file *of, char *input,
>  			continue;
>  		case COST_MODEL:
>  			match_strlcpy(buf, &args[0], sizeof(buf));
> -			if (strcmp(buf, "linear"))
> -				goto unlock;
> -			continue;
> +			if (!strcmp(buf, "linear")) {
> +				/* staged and committed below, so a parse
> +				 * error later in the same write leaves
> +				 * the model selection untouched
> +				 */
> +				new_model = NULL;
> +				model_write = true;
> +				continue;
> +			}
> +			if (!strcmp(buf, "bpf")) {
> +				new_model = rcu_dereference_protected(
> +						ioc->attached,
> +						lockdep_is_held(&ioc->lock));
> +				if (!new_model)
> +					goto unlock;
> +				model_write = true;
> +				continue;
> +			}
> +			goto unlock;
>  		}

Does this change to the io.cost.model interface need a matching update to
Documentation/admin-guide/cgroup-v2.rst?

ioc_cost_model_write() now accepts "model=bpf", and "model=linear" is no
longer only a value check.  It now clears ioc->model and moves the device
off an attached BPF model.  ioc_cost_model_prfill() can also print
"model=bpf".

The documentation still describes the key as:

  model		The cost model in use - "linear"

and says nothing about "bpf" or the new effect of writing "linear".

Existing tools that write lines like "MAJ:MIN ctrl=user model=linear
rbps=..." will now silently switch a device away from an attached BPF
model.  Could cgroup-v2.rst be updated to cover the new value and the
changed semantics?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37124306614

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [RFC PATCH v10 1/4] blk-iocost: add BPF struct_ops cost model support
  2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
  2026-10-03 13:21   ` bot+bpf-ci
@ 2026-10-03 20:45   ` Alexei Starovoitov
  1 sibling, 0 replies; 7+ messages in thread
From: Alexei Starovoitov @ 2026-10-03 20:45 UTC (permalink / raw)
  To: Tao Cui, tj, josef, axboe
  Cc: cgroups, linux-block, linux-kernel, bpf, andrii, eddyz87, daniel,
	linux-kselftest, cuitao, ameryhung

On Sat, Oct 03, 2026 at 08:40 PM Tao Cui <cui.tao@linux.dev> wrote:

v10 was sent 11 hours after v9.
Patch 1 grew from 712 to 789 lines and device removal is still
not tested.
Pls slow down. Wait for Tejun before sending v11.

pw-bot: cr

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-03 20:45 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-03 12:40 [RFC PATCH v10 0/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
2026-10-03 12:40 ` [RFC PATCH v10 1/4] " Tao Cui
2026-10-03 13:21   ` bot+bpf-ci
2026-10-03 20:45   ` Alexei Starovoitov
2026-10-03 12:40 ` [RFC PATCH v10 2/4] selftests/bpf: add iocost cost model test Tao Cui
2026-10-03 12:40 ` [RFC PATCH v10 3/4] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary Tao Cui
2026-10-03 12:40 ` [RFC PATCH v10 4/4] docs: cgroup-v2: document the iocost BPF cost model attachment Tao Cui

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®