mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Tao Cui <cui.tao@linux.dev>
To: tj@kernel.org, josef@toxicopanda.com, axboe@kernel.dk, hch@infradead.org
Cc: cgroups@vger.kernel.org, linux-block@vger.kernel.org,
	linux-kernel@vger.kernel.org, cui.tao@linux.dev,
	cuitao@kylinos.cn
Subject: [PATCH v2 1/4] blk-iocost: charge flushes as pageless random writes
Date: Wed, 16 Sep 2026 16:53:01 +0800	[thread overview]
Message-ID: <20260916085304.1080271-2-cui.tao@linux.dev> (raw)
In-Reply-To: <20260916085304.1080271-1-cui.tao@linux.dev>

From: Tao Cui <cuitao@kylinos.cn>

Standalone flushes issued by blkdev_issue_flush() are represented as
dataless REQ_OP_WRITE | REQ_PREFLUSH bios, which
calc_vtime_cost_builtin() prices at zero.  The flush component of
flush-heavy workloads such as database commits, journal flushes, and
metadata sync is thus neither charged nor throttled: a cgroup at 1% weight
can issue ~510k flushes per 12s, monopolizing the device while iocost
reports zero usage.

The same is true for the flush component of data-bearing
REQ_OP_WRITE | REQ_PREFLUSH bios, e.g. journal commit writes: they are
charged for their data only, and the cache flush the flush machine runs
ahead of it is free.

Charge the flush component of any REQ_PREFLUSH bio on top of its data
cost, priced as a pageless random write (LCOEF_WRANDIO), which provides
an approximation of the device time consumed by a flush.  For profiles
where WRANDIO clamps to zero (ssd_dfl / ssd_fast), use a one-page floor
(LCOEF_WPAGE).  A dataless flush bio falls out of the switch with zero
data cost and picks up the same surcharge, so standalone and pre-flush
forms are priced the same way.  After this patch, the same 1%-weight
cgroup is limited to 24 flushes per 12s; on ext4, write+fsync workloads
are correctly accounted through the journal layer (~2.2us per flush on
the ssd_fast profile).

A standalone flush must also not update iocg->cursor: its bi_sector
(usually 0) is not a data position, so setting the cursor from it
would misclassify the following READ/WRITE bios, and a zero cursor
defeats the !iocg->cursor sentinel in calc_vtime_cost_builtin().
Skip the cursor update for dataless bios.

Fixes: 7caa47151ab2 ("blkcg: implement blk-iocost")
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
Changes in v2:
- Skip the iocg->cursor update for dataless flush bios, which would
  otherwise corrupt the seq/rand classification of the following IOs
  (reported in review of v1).
- Charge the flush component of data-bearing REQ_PREFLUSH bios too;
  v1 only priced standalone flushes (reported in review of v1).
---
 block/blk-iocost.c | 20 +++++++++++++++++---
 1 file changed, 17 insertions(+), 3 deletions(-)

diff --git a/block/blk-iocost.c b/block/blk-iocost.c
index 2745bffcd5eef..082f26d6e27b6 100644
--- a/block/blk-iocost.c
+++ b/block/blk-iocost.c
@@ -2532,8 +2532,20 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg,
 	u64 pages = max_t(u64, bio_sectors(bio) >> IOC_SECT_TO_PAGE_SHIFT, 1);
 	u64 seek_pages = 0;
 	u64 cost = 0;
+	u64 flush_cost = 0;
 
-	/* Can't calculate cost for empty bio */
+	/*
+	 * A WRITE|REQ_PREFLUSH bio carries a flush component: the flush
+	 * machine runs a cache flush for it, either standalone (dataless)
+	 * or ahead of the data.  Charge the flush on top of the data cost,
+	 * priced as a pageless random write with a one-page floor so fast
+	 * profiles still charge something.  Flush bios are never merged.
+	 */
+	if (!is_merge && (bio->bi_opf & REQ_PREFLUSH))
+		flush_cost = max(ioc->params.lcoefs[LCOEF_WRANDIO],
+				 ioc->params.lcoefs[LCOEF_WPAGE]);
+
+	/* Can't calculate data cost for empty bio */
 	if (!bio->bi_iter.bi_size)
 		goto out;
 
@@ -2566,7 +2578,7 @@ static void calc_vtime_cost_builtin(struct bio *bio, struct ioc_gq *iocg,
 	}
 	cost += pages * coef_page;
 out:
-	*costp = cost;
+	*costp = cost + flush_cost;
 }
 
 static u64 calc_vtime_cost(struct bio *bio, struct ioc_gq *iocg, bool is_merge)
@@ -2708,7 +2720,9 @@ static void ioc_rqos_throttle(struct rq_qos *rqos, struct bio *bio)
 	if (!iocg_activate(iocg, &now))
 		return;
 
-	iocg->cursor = bio_end_sector(bio);
+	/* dataless bios have no meaningful position for seq/rand detection */
+	if (bio->bi_iter.bi_size)
+		iocg->cursor = bio_end_sector(bio);
 	vtime = atomic64_read(&iocg->vtime);
 	cost = adjust_inuse_and_calc_cost(iocg, vtime, abs_cost, &now);
 
-- 
2.43.0


  reply	other threads:[~2026-09-16  8:53 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  8:53 [PATCH v2 0/4] blk-iocost: charge flushes and zone appends Tao Cui
2026-09-16  8:53 ` Tao Cui [this message]
2026-09-16  8:53 ` [PATCH v2 2/4] blk-iocost: charge zone appends as page-counted sequential writes Tao Cui
2026-09-16  8:53 ` [PATCH v2 3/4] blk-iocost: account zone append completions in latency stats Tao Cui
2026-09-16  8:53 ` [PATCH v2 4/4] blk-iocost: fix stale comment in ioc_rqos_throttle() Tao Cui

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260916085304.1080271-2-cui.tao@linux.dev \
    --to=cui.tao@linux.dev \
    --cc=axboe@kernel.dk \
    --cc=cgroups@vger.kernel.org \
    --cc=cuitao@kylinos.cn \
    --cc=hch@infradead.org \
    --cc=josef@toxicopanda.com \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®