mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [v2 for-next 0/3] block: delay support for error injection
@ 2026-08-30  1:19 Md Haris Iqbal
  2026-08-30  1:20 ` [v2 for-next 1/3] block: reject unknown status tags in error injection rules Md Haris Iqbal
                   ` (2 more replies)
  0 siblings, 3 replies; 9+ messages in thread
From: Md Haris Iqbal @ 2026-08-30  1:19 UTC (permalink / raw)
  To: Jens Axboe, linux-block
  Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet,
	linux-doc, Md Haris Iqbal

Error injection can only fail a bio today.  This adds a delay_us option so
that a rule can hold a bio back first, to model a slow device.

The delay happens above the driver, so it is invisible to the I/O
statistics and never reaches the blk-mq timeout handler or SCSI error
handling.  What it does exercise is the code waiting above the block
layer: io_uring cancellation, hung task detection, and filesystem or
userspace timeouts.

Two things are worth a look.  A delayed bio is marked with a new
BIO_ERROR_INJECTED bio flag, so the rules are not applied to it again and
it can never pick up a status from another rule.  Resubmitting it below the
injection hook is not enough on its own, because a bio is split below the
hook and the remainder is resubmitted above it.  And holding a bio back
reorders it against bios submitted later, which breaks sequential write
ordering on zoned devices.  Both are documented in patch 3.

Patch 1 is a prep cleanup.  It moves the rejection of an unknown status
tag into the parser, because patch 2 makes a rule without a status valid.

v1: https://lore.kernel.org/linux-block/20260827000115.128093-1-haris.iqbal@linux.dev/

Changes since v1:

 - Add the BIO_ERROR_INJECTED bio flag.  v1 only resubmitted a delayed bio
   below the injection hook, which left the remainder of a split going
   through the hook and matching the same rule again, so a bio was held
   once per split instead of once.

 - Patch 1: tag_to_blk_status() returns an error and passes the status back
   through a pointer, instead of returning BLK_STS_OK for both the "OK" tag
   and an unknown one.  A delay-only rule has no status, so the two have to
   be told apart.

 - Documentation: a delayed bio is held once rather than once per split.

Tested in a VM.

Md Haris Iqbal (3):
  block: reject unknown status tags in error injection rules
  block: allow error injection rules to delay bios
  Documentation: block: document error injection delays

 Documentation/block/error-injection.rst |  58 ++++++++-
 block/blk-core.c                        |  27 ++--
 block/blk.h                             |   3 +-
 block/error-injection.c                 | 165 +++++++++++++++++++++---
 block/error-injection.h                 |   1 +
 include/linux/blk_types.h               |   1 +
 6 files changed, 222 insertions(+), 33 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [v2 for-next 1/3] block: reject unknown status tags in error injection rules
  2026-08-30  1:19 [v2 for-next 0/3] block: delay support for error injection Md Haris Iqbal
@ 2026-08-30  1:20 ` Md Haris Iqbal
  2026-09-02 13:58   ` Christoph Hellwig
  2026-08-30  1:20 ` [v2 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal
  2026-08-30  1:20 ` [v2 for-next 3/3] Documentation: block: document error injection delays Md Haris Iqbal
  2 siblings, 1 reply; 9+ messages in thread
From: Md Haris Iqbal @ 2026-08-30  1:20 UTC (permalink / raw)
  To: Jens Axboe, linux-block
  Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet,
	linux-doc, Md Haris Iqbal

tag_to_blk_status() returns BLK_STS_OK both for the "OK" tag and for a
tag it does not recognise, so a caller cannot tell the two apart.
match_status() leaves *status at BLK_STS_OK for an unknown tag and relies
on error_inject_add() rejecting BLK_STS_OK.

That holds only while a rule without a status is meaningless.  The delay
option added next makes such a rule valid, and "OK" is what
blk_error_injection_show() prints for one, so the two cases have to be
told apart.  Return the status through a pointer and report an unknown
tag as -EINVAL.  There is no spare blk_status_t to encode "not found" in:
every value in the table has a tag that can be typed, and anything
outside the table trips the WARN_ON_ONCE() in blk_status_to_str(),
blk_status_to_tag() and blk_status_to_errno().

For a single status= this does not change behaviour: an unknown tag still
fails the write with -EINVAL.  A repeated status= where an invalid tag
comes first is now rejected instead of being overridden by the later one.

Cc: Christoph Hellwig <hch@lst.de>
Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev>
---
 block/blk-core.c        | 14 ++++++--------
 block/blk.h             |  2 +-
 block/error-injection.c |  7 ++++---
 3 files changed, 11 insertions(+), 12 deletions(-)

diff --git a/block/blk-core.c b/block/blk-core.c
index 196bccf27f58..29c86addb8f2 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -225,21 +225,19 @@ const char *blk_status_to_tag(blk_status_t status)
 	return blk_errors[idx].tag;
 }
 
-blk_status_t tag_to_blk_status(const char *tag)
+int tag_to_blk_status(const char *tag, blk_status_t *status)
 {
 	int i;
 
 	for (i = 0; i < ARRAY_SIZE(blk_errors); i++) {
 		if (blk_errors[i].tag &&
-		    !strcmp(blk_errors[i].tag, tag))
-			return (__force blk_status_t)i;
+		    !strcmp(blk_errors[i].tag, tag)) {
+			*status = (__force blk_status_t)i;
+			return 0;
+		}
 	}
 
-	/*
-	 * Return BLK_STS_OK for mismatches as this function is intended to
-	 * parse error status values.
-	 */
-	return BLK_STS_OK;
+	return -EINVAL;
 }
 
 /**
diff --git a/block/blk.h b/block/blk.h
index 50abfd932886..8a8ab961528d 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -52,7 +52,7 @@ void blk_free_flush_queue(struct blk_flush_queue *q);
 
 const char *blk_status_to_str(blk_status_t status);
 const char *blk_status_to_tag(blk_status_t status);
-blk_status_t tag_to_blk_status(const char *tag);
+int tag_to_blk_status(const char *tag, blk_status_t *status);
 enum req_op str_to_blk_op(const char *op);
 
 bool __blk_mq_unfreeze_queue(struct request_queue *q, bool force_atomic);
diff --git a/block/error-injection.c b/block/error-injection.c
index e14bc4b723ef..41ee8f788bb5 100644
--- a/block/error-injection.c
+++ b/block/error-injection.c
@@ -171,15 +171,16 @@ static int match_op(substring_t *args, enum req_op *op)
 static int match_status(substring_t *args, blk_status_t *status)
 {
 	const char *tag;
+	int ret;
 
 	tag = match_strdup(args);
 	if (!tag)
 		return -ENOMEM;
-	*status = tag_to_blk_status(tag);
-	if (!*status)
+	ret = tag_to_blk_status(tag, status);
+	if (ret)
 		pr_warn("invalid status '%s'\n", tag);
 	kfree(tag);
-	return 0;
+	return ret;
 }
 
 static ssize_t blk_error_injection_parse_options(struct gendisk *disk,
-- 
2.53.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [v2 for-next 2/3] block: allow error injection rules to delay bios
  2026-08-30  1:19 [v2 for-next 0/3] block: delay support for error injection Md Haris Iqbal
  2026-08-30  1:20 ` [v2 for-next 1/3] block: reject unknown status tags in error injection rules Md Haris Iqbal
@ 2026-08-30  1:20 ` Md Haris Iqbal
  2026-09-15  9:11   ` Christoph Hellwig
  2026-08-30  1:20 ` [v2 for-next 3/3] Documentation: block: document error injection delays Md Haris Iqbal
  2 siblings, 1 reply; 9+ messages in thread
From: Md Haris Iqbal @ 2026-08-30  1:20 UTC (permalink / raw)
  To: Jens Axboe, linux-block
  Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet,
	linux-doc, Md Haris Iqbal

Error injection can only fail a bio today.  Add a delay_us option so that
a rule can hold a bio back first, to model a slow device.

If a matching rule has delay_us set, the bio is held for that long.  It is
then failed if the rule also has a status, or resubmitted below the
injection hook so that the rules are not applied to it again.

Submitting below the hook is not enough on its own.  A bio is split below
the hook, and bio_submit_split_bioset() resubmits the remainder through
submit_bio_noacct_nocheck(), which is above it.  bio_split() advances the
original bio and returns a clone of the front piece, so the remainder is
the same bio on a range the rule still covers and would be delayed once
per split.  Mark a delayed bio with BIO_ERROR_INJECTED instead and skip
the hook for a bio that has it.  __bio_clone() does not propagate the
flag, so the front pieces and the clones a stacking driver aims at a
lower device are still evaluated.

The bio is submitted from a workqueue rather than from the timer, because
submitting a bio can sleep.  Bios with REQ_NOWAIT are never delayed, and
values above 600 seconds are rejected.

Cc: Christoph Hellwig <hch@lst.de>
Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev>
---
 block/blk-core.c          |  13 +++-
 block/blk.h               |   1 +
 block/error-injection.c   | 158 ++++++++++++++++++++++++++++++++++----
 block/error-injection.h   |   1 +
 include/linux/blk_types.h |   1 +
 5 files changed, 155 insertions(+), 19 deletions(-)

diff --git a/block/blk-core.c b/block/blk-core.c
index 29c86addb8f2..6ba21fd37b6c 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -757,11 +757,8 @@ static void __submit_bio_noacct_mq(struct bio *bio)
 	current->bio_list = NULL;
 }
 
-void submit_bio_noacct_nocheck(struct bio *bio, bool split)
+void __submit_bio_noacct_nocheck(struct bio *bio, bool split)
 {
-	if (unlikely(blk_error_inject(bio)))
-		return;
-
 	blk_cgroup_bio_start(bio);
 
 	if (!bio_flagged(bio, BIO_TRACE_COMPLETION)) {
@@ -791,6 +788,14 @@ void submit_bio_noacct_nocheck(struct bio *bio, bool split)
 	}
 }
 
+void submit_bio_noacct_nocheck(struct bio *bio, bool split)
+{
+	if (unlikely(blk_error_inject(bio)))
+		return;
+
+	__submit_bio_noacct_nocheck(bio, split);
+}
+
 static blk_status_t blk_validate_atomic_write_op_size(struct request_queue *q,
 						 struct bio *bio)
 {
diff --git a/block/blk.h b/block/blk.h
index 8a8ab961528d..389c9a487067 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -61,6 +61,7 @@ bool __blk_freeze_queue_start(struct request_queue *q,
 			      struct task_struct *owner);
 int __bio_queue_enter(struct request_queue *q, struct bio *bio);
 void submit_bio_noacct_nocheck(struct bio *bio, bool split);
+void __submit_bio_noacct_nocheck(struct bio *bio, bool split);
 int bio_submit_or_kill(struct bio *bio, unsigned int flags);
 
 static inline bool blk_try_enter_queue(struct request_queue *q, bool pm)
diff --git a/block/error-injection.c b/block/error-injection.c
index 41ee8f788bb5..db8384c7c6f9 100644
--- a/block/error-injection.c
+++ b/block/error-injection.c
@@ -6,9 +6,17 @@
 #include <linux/blkdev.h>
 #include <linux/parser.h>
 #include <linux/seq_file.h>
+#include <linux/workqueue.h>
 #include "blk.h"
 #include "error-injection.h"
 
+/*
+ * Cap the delay so that a typo can't wedge a device for good.  This is still
+ * well beyond the default hung task timeout, which is one of the things a
+ * delay is useful for triggering.
+ */
+#define BLK_ERROR_INJECT_MAX_DELAY_US	(600 * USEC_PER_SEC)
+
 struct blk_error_inject {
 	struct list_head		entry;
 	sector_t			start;
@@ -18,14 +26,102 @@ struct blk_error_inject {
 
 	/* only inject every 1 / chance times */
 	unsigned int			chance;
+
+	/* hold the bio for this long before submitting or failing it */
+	unsigned int			delay_us;
 };
 
+/*
+ * A bio held by a delay rule.  This is self-contained on purpose: it does not
+ * point back at the rule, so rules can be removed while delayed bios are
+ * outstanding, and it does not point at the gendisk, so nothing has to be
+ * cleaned up when the disk goes away.  A delayed bio holds no queue usage
+ * counter reference either, so one that outlives its disk is failed by the
+ * GD_DEAD check in __bio_queue_enter() once it is finally submitted.
+ */
+struct blk_error_inject_delay {
+	struct delayed_work		dwork;
+	struct bio			*bio;
+	blk_status_t			status;
+};
+
+static struct workqueue_struct *blk_error_inject_wq;
+
 DEFINE_STATIC_KEY_FALSE(blk_error_injection_enabled);
 
+static void blk_error_inject_delay_work(struct work_struct *work)
+{
+	struct blk_error_inject_delay *d = container_of(to_delayed_work(work),
+			struct blk_error_inject_delay, dwork);
+	struct bio *bio = d->bio;
+	blk_status_t status = d->status;
+
+	kfree(d);
+
+	if (status != BLK_STS_OK) {
+		bio->bi_status = status;
+		bio_endio(bio);
+	} else {
+		/*
+		 * Submit below the injection hook.  Together with
+		 * BIO_ERROR_INJECTED, which also covers the resubmission of
+		 * the remainder of a split, this means a bio that was delayed
+		 * once skips error injection entirely from here on, including
+		 * any other rule that covers it.
+		 */
+		__submit_bio_noacct_nocheck(bio, false);
+	}
+}
+
+/*
+ * Hand the bio to a workqueue that submits or fails it once the delay has
+ * expired.  Both blk_mq_submit_bio() and ->submit_bio can sleep, so this can't
+ * be completed from the timer itself.
+ *
+ * Returns false if the bio can't be delayed, in which case the caller handles
+ * it immediately instead.
+ */
+static bool blk_error_inject_delay(struct gendisk *disk, struct bio *bio,
+		blk_status_t status, unsigned int delay_us)
+{
+	struct blk_error_inject_delay *d;
+
+	/* never block a bio that asked not to be blocked */
+	if (bio->bi_opf & REQ_NOWAIT)
+		return false;
+
+	d = kmalloc_obj(*d, GFP_NOIO);
+	if (!d)
+		return false;
+
+	pr_info_ratelimited("%pg: delaying %s at sector %llu:%u by %uus\n",
+			disk->part0, blk_op_str(bio_op(bio)),
+			bio->bi_iter.bi_sector, bio_sectors(bio), delay_us);
+
+	d->bio = bio;
+	d->status = status;
+	INIT_DELAYED_WORK(&d->dwork, blk_error_inject_delay_work);
+
+	/*
+	 * Mark the bio before queueing the work, which can complete it as soon
+	 * as it is queued.  Splitting happens below the injection hook, but
+	 * bio_submit_split_bioset() resubmits the remainder through the hook
+	 * again, and as bio_split() only advances the original bio that
+	 * remainder still matches the same rule.  Without this a bio would be
+	 * delayed once per split.
+	 */
+	bio_set_flag(bio, BIO_ERROR_INJECTED);
+	queue_delayed_work(blk_error_inject_wq, &d->dwork,
+			usecs_to_jiffies(delay_us));
+	return true;
+}
+
 bool __blk_error_inject(struct bio *bio)
 {
 	struct gendisk *disk = bio->bi_bdev->bd_disk;
 	struct blk_error_inject *inj;
+	blk_status_t status = BLK_STS_OK;
+	unsigned int delay_us = 0;
 
 	rcu_read_lock();
 	list_for_each_entry_rcu(inj, &disk->error_injection_list, entry) {
@@ -45,29 +141,38 @@ bool __blk_error_inject(struct bio *bio)
 		if (inj->chance > 1 && (get_random_u32() % inj->chance) != 0)
 			continue;
 
-		pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n",
-				disk->part0, blk_status_to_str(inj->status),
-				blk_op_str(inj->op), bio->bi_iter.bi_sector,
-				bio_sectors(bio));
-		bio->bi_status = inj->status;
-		rcu_read_unlock();
-		bio_endio(bio);
-		return true;
+		status = inj->status;
+		delay_us = inj->delay_us;
+		break;
 	}
 	rcu_read_unlock();
-	return false;
+
+	if (delay_us && blk_error_inject_delay(disk, bio, status, delay_us))
+		return true;
+	if (status == BLK_STS_OK)
+		return false;
+
+	pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n",
+			disk->part0, blk_status_to_str(status),
+			blk_op_str(bio_op(bio)), bio->bi_iter.bi_sector,
+			bio_sectors(bio));
+	bio->bi_status = status;
+	bio_endio(bio);
+	return true;
 }
 
 static int error_inject_add(struct gendisk *disk, enum req_op op,
 		sector_t start, u64 nr_sectors, blk_status_t status,
-		unsigned int chance)
+		unsigned int chance, unsigned int delay_us)
 {
 	struct blk_error_inject *inj;
 	int error = -EINVAL;
 
 	if (op == REQ_OP_LAST)
 		return -EINVAL;
-	if (status == BLK_STS_OK)
+	if (status == BLK_STS_OK && !delay_us)
+		return -EINVAL;
+	if (delay_us > BLK_ERROR_INJECT_MAX_DELAY_US)
 		return -EINVAL;
 
 	inj = kzalloc_obj(*inj);
@@ -86,6 +191,7 @@ static int error_inject_add(struct gendisk *disk, enum req_op op,
 	inj->start = start;
 	inj->status = status;
 	inj->chance = chance;
+	inj->delay_us = delay_us;
 
 	pr_debug_ratelimited("%pg: adding %s injection for %s at sector %llu:%llu\n",
 			disk->part0, blk_status_to_str(status),
@@ -139,6 +245,7 @@ enum options {
 	Opt_nr_sectors		= (1u << 18),
 	Opt_status		= (1u << 19),
 	Opt_chance		= (1u << 20),
+	Opt_delay_us		= (1u << 21),
 
 	Opt_invalid,
 };
@@ -151,6 +258,7 @@ static const match_table_t opt_tokens = {
 	{ Opt_nr_sectors,		"nr_sectors=%u"		},
 	{ Opt_status,			"status=%s"		},
 	{ Opt_chance,			"chance=%u"		},
+	{ Opt_delay_us,			"delay_us=%u"		},
 	{ Opt_invalid,			NULL,			},
 };
 
@@ -187,7 +295,7 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk,
 		char *options)
 {
 	enum { Unset, Add, Removeall } action = Unset;
-	unsigned int option_mask = 0, chance = 1;
+	unsigned int option_mask = 0, chance = 1, delay_us = 0;
 	enum req_op op = REQ_OP_LAST;
 	u64 start = 0, nr_sectors = 0;
 	blk_status_t status = BLK_STS_OK;
@@ -230,6 +338,9 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk,
 			if (!error && chance == 0)
 				error = -EINVAL;
 			break;
+		case Opt_delay_us:
+			error = match_uint(args, &delay_us);
+			break;
 		default:
 			pr_warn("unknown parameter or missing value '%s'\n", p);
 			error = -EINVAL;
@@ -241,7 +352,7 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk,
 	switch (action) {
 	case Add:
 		return error_inject_add(disk, op, start, nr_sectors, status,
-				chance);
+				chance, delay_us);
 	case Removeall:
 		if (option_mask & ~Opt_removeall)
 			return -EINVAL;
@@ -277,10 +388,11 @@ static int blk_error_injection_show(struct seq_file *s, void *private)
 
 	rcu_read_lock();
 	list_for_each_entry_rcu(inj, &disk->error_injection_list, entry) {
-		seq_printf(s, "%llu:%llu op=%s,status=%s,chance=%u",
+		seq_printf(s, "%llu:%llu op=%s,status=%s,chance=%u,delay_us=%u",
 			   inj->start, inj->end,
 			   blk_op_str(inj->op),
-			   blk_status_to_tag(inj->status), inj->chance);
+			   blk_status_to_tag(inj->status), inj->chance,
+			   inj->delay_us);
 		seq_putc(s, '\n');
 	}
 	rcu_read_unlock();
@@ -315,3 +427,19 @@ void blk_error_injection_exit(struct gendisk *disk)
 {
 	error_inject_removeall(disk);
 }
+
+static int __init blk_error_injection_init_wq(void)
+{
+	/*
+	 * WQ_MEM_RECLAIM so that a delayed bio on the reclaim path can still
+	 * find a worker under memory pressure.  Note that this only guarantees
+	 * a worker exists, not that it is free: submitting a bio can block on
+	 * a queue freeze or on tag allocation, so a delayed bio can still be
+	 * held up behind another one.
+	 */
+	blk_error_inject_wq = alloc_workqueue("blk_error_inject", WQ_MEM_RECLAIM | WQ_UNBOUND, 0);
+	if (!blk_error_inject_wq)
+		panic("Failed to create blk_error_inject wq\n");
+	return 0;
+}
+subsys_initcall(blk_error_injection_init_wq);
diff --git a/block/error-injection.h b/block/error-injection.h
index 9821d773abab..8b3809d85e83 100644
--- a/block/error-injection.h
+++ b/block/error-injection.h
@@ -13,6 +13,7 @@ static inline bool blk_error_inject(struct bio *bio)
 {
 	if (IS_ENABLED(CONFIG_BLK_ERROR_INJECTION) &&
 	    static_branch_unlikely(&blk_error_injection_enabled) &&
+	    !bio_flagged(bio, BIO_ERROR_INJECTED) &&
 	    test_bit(GD_ERROR_INJECT, &bio->bi_bdev->bd_disk->state))
 		return __blk_error_inject(bio);
 	return false;
diff --git a/include/linux/blk_types.h b/include/linux/blk_types.h
index 98e21b4cbf32..50e28dc0f1f9 100644
--- a/include/linux/blk_types.h
+++ b/include/linux/blk_types.h
@@ -323,6 +323,7 @@ enum {
 	BIO_ZONE_WRITE_PLUGGING, /* bio handled through zone write plugging */
 	BIO_EMULATES_ZONE_APPEND, /* bio emulates a zone append operation */
 	BIO_COMPLETE_IN_TASK, /* complete bi_end_io() in task context */
+	BIO_ERROR_INJECTED, /* error injection rules already applied */
 	BIO_FLAG_LAST
 };
 
-- 
2.53.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [v2 for-next 3/3] Documentation: block: document error injection delays
  2026-08-30  1:19 [v2 for-next 0/3] block: delay support for error injection Md Haris Iqbal
  2026-08-30  1:20 ` [v2 for-next 1/3] block: reject unknown status tags in error injection rules Md Haris Iqbal
  2026-08-30  1:20 ` [v2 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal
@ 2026-08-30  1:20 ` Md Haris Iqbal
  2 siblings, 0 replies; 9+ messages in thread
From: Md Haris Iqbal @ 2026-08-30  1:20 UTC (permalink / raw)
  To: Jens Axboe, linux-block
  Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet,
	linux-doc, Md Haris Iqbal

Document the delay_us option: the delay happens above the driver, a
delayed bio is not run through the rules again, and holding a bio back
reorders it against bios submitted later.

Cc: Christoph Hellwig <hch@lst.de>
Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev>
---
 Documentation/block/error-injection.rst | 58 ++++++++++++++++++++++++-
 1 file changed, 56 insertions(+), 2 deletions(-)

diff --git a/Documentation/block/error-injection.rst b/Documentation/block/error-injection.rst
index 81f31af82e65..a1f50d41211d 100644
--- a/Documentation/block/error-injection.rst
+++ b/Documentation/block/error-injection.rst
@@ -9,7 +9,8 @@ Overview
 
 Configurable error injection allows injecting specific block layer status codes
 for sector ranges of a block device.  Errors can be injected unconditionally, or
-with a given probability.
+with a given probability.  Instead of, or before, failing a bio it can also be
+held back for a while to model a slow device.
 
 To use configurable error injection, CONFIG_BLK_ERROR_INJECTION must be enabled.
 
@@ -34,15 +35,60 @@ op=<string>		block layer operation this rule applies to.  This uses
 			the XYZ for each REQ_OP_XYZ operation, e.g. READ, WRITE
 			or DISCARD. Mandatory.
 status=<string>		Status to return.  This uses XYZ for each BLK_STS_XYZ
-			code, e.g. IOERR or MEDIUM. Mandatory.
+			code, e.g. IOERR or MEDIUM. Mandatory unless delay_us
+			is given.
 start=<number>		First block layer sector the rule applies to.
 			Optional, defaults to 0.
 nr_sectors=<number>	Number of sectors this rule applies.
 			Optional, defaults to the remainder of the device.
 chance=<number>		Only return a failure with a likelihood of 1/chance.
 			Optional, defaults to 1 (always).
+delay_us=<number>	Hold the bio back for this many microseconds.  Without
+			status the bio is then submitted to the device as
+			usual, with status it is failed once the delay has
+			expired.  Optional, defaults to 0 (no delay).
+			Values above 600 seconds are rejected.
 ===================	=======================================================
 
+Delays
+------
+
+A delayed bio is held before it is submitted, so the device itself never sees a
+slow I/O: the delay is not visible to the driver, to the I/O statistics, or to
+anything else below submission such as writeback throttling.  Throttling by
+blk-throttle happens before a bio can be delayed, so it is not affected either.
+What it does exercise is everything waiting above the block layer, for instance
+io_uring cancellation, hung task detection, and filesystem or userspace
+timeouts.  Because the low level driver is not involved, a delay does not reach
+the blk-mq timeout handler or SCSI error handling.
+
+Once a bio has been delayed no rule is evaluated for it a second time, not when
+its delay expires and it is submitted below the injection hook, and not when
+the block layer splits it and resubmits the remainder above the hook.  A bio
+that matched a delay rule therefore never gets an error from another rule, even
+one covering the same sectors, and is held for the delay once rather than once
+per split.  Put the delay and the status in a single rule to fail a bio after
+holding it back.
+
+A delayed bio is issued after bios submitted while it was held, which reorders
+the I/O stream.  On zoned devices this breaks sequential write ordering: zone
+write plugging happens below the injection hook, so the writes issued while a
+write is held reach the zone out of order and are failed as misaligned.  Only
+delay reads there.
+
+Bios that must not block are never delayed.  A bio with REQ_NOWAIT set is
+submitted, or failed with the rule's status, immediately.
+
+The delay is a lower bound for anything longer than a timer tick, and the timer
+wheel adds further slack as the delay grows.  Values shorter than a tick are of
+little use: they expire on the next tick, which is anywhere between now and one
+tick away.
+
+Removing rules does not release bios that are already being delayed by them;
+those run out on their own.  A delayed bio whose disk is removed in the meantime
+is not submitted until its delay expires, by which point the queue no longer
+accepts I/O, so it fails with EIO.
+
 Example
 -------
 
@@ -54,6 +100,14 @@ Return BLK_STS_MEDIUM for every write to /dev/nvme0n1:
 
 	$ echo 'add,op=WRITE,start=0,status=MEDIUM' > /sys/kernel/debug/block/nvme0n1/error_injection
 
+Delay every read of /dev/nvme0n1 by 10 milliseconds, then issue it normally:
+
+	$ echo 'add,op=READ,delay_us=10000' > /sys/kernel/debug/block/nvme0n1/error_injection
+
+Fail one in 100 writes with BLK_STS_TIMEOUT, but only after 30 seconds:
+
+	$ echo 'add,op=WRITE,status=TIMEOUT,chance=100,delay_us=30000000' > /sys/kernel/debug/block/nvme0n1/error_injection
+
 Remove all rules for /dev/nvme0n1:
 
 	$ echo 'removeall' > /sys/kernel/debug/block/nvme0n1/error_injection
-- 
2.53.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [v2 for-next 1/3] block: reject unknown status tags in error injection rules
  2026-08-30  1:20 ` [v2 for-next 1/3] block: reject unknown status tags in error injection rules Md Haris Iqbal
@ 2026-09-02 13:58   ` Christoph Hellwig
  2026-09-02 19:48     ` Haris Iqbal
  0 siblings, 1 reply; 9+ messages in thread
From: Christoph Hellwig @ 2026-09-02 13:58 UTC (permalink / raw)
  To: Md Haris Iqbal
  Cc: Jens Axboe, linux-block, linux-kernel, Christoph Hellwig,
	Keith Busch, Jonathan Corbet, linux-doc

On Sun, Aug 30, 2026 at 03:20:00AM +0200, Md Haris Iqbal wrote:
> tag_to_blk_status() returns BLK_STS_OK both for the "OK" tag and for a
> tag it does not recognise, so a caller cannot tell the two apart.
> match_status() leaves *status at BLK_STS_OK for an unknown tag and relies
> on error_inject_add() rejecting BLK_STS_OK.

There is no *status yet.

> That holds only while a rule without a status is meaningless.  The delay

This reads a lot like AI slop.  Can you please self-write a short and
descriptive commit message?

> -	 * Return BLK_STS_OK for mismatches as this function is intended to
> -	 * parse error status values.
> -	 */
> -	return BLK_STS_OK;
> +	return -EINVAL;

No need for an int return here, this can easily be done with a bool.



^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [v2 for-next 1/3] block: reject unknown status tags in error injection rules
  2026-09-02 13:58   ` Christoph Hellwig
@ 2026-09-02 19:48     ` Haris Iqbal
  0 siblings, 0 replies; 9+ messages in thread
From: Haris Iqbal @ 2026-09-02 19:48 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Jens Axboe, linux-block, linux-kernel, Keith Busch,
	Jonathan Corbet, linux-doc



On 9/2/26 15:58, Christoph Hellwig wrote:
> On Sun, Aug 30, 2026 at 03:20:00AM +0200, Md Haris Iqbal wrote:
>> tag_to_blk_status() returns BLK_STS_OK both for the "OK" tag and for a
>> tag it does not recognise, so a caller cannot tell the two apart.
>> match_status() leaves *status at BLK_STS_OK for an unknown tag and relies
>> on error_inject_add() rejecting BLK_STS_OK.
> 
> There is no *status yet.

Ah yes. I'll correct it.

> 
>> That holds only while a rule without a status is meaningless.  The delay
> 
> This reads a lot like AI slop.  Can you please self-write a short and
> descriptive commit message?

My bad. I'll correct it.

> 
>> -	 * Return BLK_STS_OK for mismatches as this function is intended to
>> -	 * parse error status values.
>> -	 */
>> -	return BLK_STS_OK;
>> +	return -EINVAL;
> 
> No need for an int return here, this can easily be done with a bool.

True. I'll change it.

I'll wait for your comments for the other patches before sending a v3.


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [v2 for-next 2/3] block: allow error injection rules to delay bios
  2026-08-30  1:20 ` [v2 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal
@ 2026-09-15  9:11   ` Christoph Hellwig
  2026-09-15 21:24     ` Haris Iqbal
  0 siblings, 1 reply; 9+ messages in thread
From: Christoph Hellwig @ 2026-09-15  9:11 UTC (permalink / raw)
  To: Md Haris Iqbal
  Cc: Jens Axboe, linux-block, linux-kernel, Christoph Hellwig,
	Keith Busch, Jonathan Corbet, linux-doc

The delay does look fine, but I'd really like to not use up one of the
scare bio flags for it.  Can we insist on the delay only working for
an error, so that we never have to reinsert?


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [v2 for-next 2/3] block: allow error injection rules to delay bios
  2026-09-15  9:11   ` Christoph Hellwig
@ 2026-09-15 21:24     ` Haris Iqbal
  2026-09-15 21:29       ` Haris Iqbal
  0 siblings, 1 reply; 9+ messages in thread
From: Haris Iqbal @ 2026-09-15 21:24 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Jens Axboe, linux-block, linux-kernel, Keith Busch,
	Jonathan Corbet, linux-doc



On 9/15/26 11:11, Christoph Hellwig wrote:
> The delay does look fine, but I'd really like to not use up one of the
> scare bio flags for it.  Can we insist on the delay only working for
> an error, so that we never have to reinsert?

We can do that, but that's only half of the feature, and the weaker half 
IMHO. Only delay with a successful completion covers a more useful 
scenario for testing in the upper layers.

I am thinking about how else we can do this, but did not find a nice 
approach.

Maybe we can #ifdef the new flag on CONFIG_BLK_ERROR_INJECTION?

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [v2 for-next 2/3] block: allow error injection rules to delay bios
  2026-09-15 21:24     ` Haris Iqbal
@ 2026-09-15 21:29       ` Haris Iqbal
  0 siblings, 0 replies; 9+ messages in thread
From: Haris Iqbal @ 2026-09-15 21:29 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Jens Axboe, linux-block, linux-kernel, Keith Busch,
	Jonathan Corbet, linux-doc



On 9/15/26 23:24, Haris Iqbal wrote:
> 
> 
> On 9/15/26 11:11, Christoph Hellwig wrote:
>> The delay does look fine, but I'd really like to not use up one of the
>> scare bio flags for it.  Can we insist on the delay only working for
>> an error, so that we never have to reinsert?
> 
> We can do that, but that's only half of the feature, and the weaker half 
> IMHO. Only delay with a successful completion covers a more useful 
> scenario for testing in the upper layers.
> 
> I am thinking about how else we can do this, but did not find a nice 
> approach.
> 
> Maybe we can #ifdef the new flag on CONFIG_BLK_ERROR_INJECTION?

I forgot one small point in favour of the flag.

We can remove addition of __submit_bio_noacct_nocheck() function, since 
with the flag the re-entry after the delay would fail on 
blk_error_inject(). This would remove confine the change to 
error_injection files.

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-15 21:29 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-08-30  1:19 [v2 for-next 0/3] block: delay support for error injection Md Haris Iqbal
2026-08-30  1:20 ` [v2 for-next 1/3] block: reject unknown status tags in error injection rules Md Haris Iqbal
2026-09-02 13:58   ` Christoph Hellwig
2026-09-02 19:48     ` Haris Iqbal
2026-08-30  1:20 ` [v2 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal
2026-09-15  9:11   ` Christoph Hellwig
2026-09-15 21:24     ` Haris Iqbal
2026-09-15 21:29       ` Haris Iqbal
2026-08-30  1:20 ` [v2 for-next 3/3] Documentation: block: document error injection delays Md Haris Iqbal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®