* [v3 for-next 0/3] block: delay support for error injection
@ 2026-09-28 22:16 Md Haris Iqbal
2026-09-28 22:16 ` [v3 for-next 1/3] block: Reject unknown status tags in error injection rules Md Haris Iqbal
` (2 more replies)
0 siblings, 3 replies; 7+ messages in thread
From: Md Haris Iqbal @ 2026-09-28 22:16 UTC (permalink / raw)
To: Jens Axboe, linux-block
Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet,
linux-doc, Md Haris Iqbal
Error injection can only fail a bio today. This adds a delay_us option so
that a rule can hold a bio back first, to model a slow device. The delay
happens above the driver, so it tests the code waiting above the block
layer, not the blk-mq timeout handler.
A delayed bio is marked with a new BIO_ERROR_INJECTED flag so the rules
are not applied to it again, even after a split. Christoph asked on v2
to avoid the flag by allowing a delay only together with an error. I kept
delay-only rules, as a slow device that succeeds is the more useful case,
so the flag question is still open.
Holding a bio back reorders it against bios submitted later, which breaks
sequential write ordering on zoned devices.
Patch 1 is a prep cleanup. It moves the rejection of an unknown status
tag into the parser, because patch 2 makes a rule without a status valid.
Tested in a VM.
v1: https://lore.kernel.org/linux-block/20260827000115.128093-1-haris.iqbal@linux.dev/
v2: https://lore.kernel.org/linux-block/20260830012002.80275-1-haris.iqbal@linux.dev/
Changes since v2:
- Patch 1: return bool, rewrite the commit message.
- Patch 2: drop __submit_bio_noacct_nocheck(); the flag alone stops the
rules being applied again. Log the error when a delayed bio is failed.
- Patch 3: trim the documentation.
Changes since v1:
- Add the BIO_ERROR_INJECTED bio flag. v1 only resubmitted a delayed bio
below the injection hook, which left the remainder of a split going
through the hook and matching the same rule again, so a bio was held
once per split instead of once.
- Patch 1: tag_to_blk_status() returns an error and passes the status back
through a pointer, instead of returning BLK_STS_OK for both the "OK" tag
and an unknown one. A delay-only rule has no status, so the two have to
be told apart.
- Documentation: a delayed bio is held once rather than once per split.
Md Haris Iqbal (3):
block: Reject unknown status tags in error injection rules
block: allow error injection rules to delay bios
Documentation: block: document error injection delay feature
Documentation/block/error-injection.rst | 39 +++++-
block/blk-core.c | 14 +-
block/blk.h | 2 +-
block/error-injection.c | 167 +++++++++++++++++++++---
block/error-injection.h | 1 +
include/linux/blk_types.h | 1 +
6 files changed, 192 insertions(+), 32 deletions(-)
--
2.53.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [v3 for-next 1/3] block: Reject unknown status tags in error injection rules 2026-09-28 22:16 [v3 for-next 0/3] block: delay support for error injection Md Haris Iqbal @ 2026-09-28 22:16 ` Md Haris Iqbal 2026-10-05 8:49 ` Christoph Hellwig 2026-09-28 22:16 ` [v3 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal 2026-09-28 22:16 ` [v3 for-next 3/3] Documentation: block: document error injection delay feature Md Haris Iqbal 2 siblings, 1 reply; 7+ messages in thread From: Md Haris Iqbal @ 2026-09-28 22:16 UTC (permalink / raw) To: Jens Axboe, linux-block Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc, Md Haris Iqbal tag_to_blk_status() returns BLK_STS_OK both for the "OK" tag and for a tag it does not recognize, so a caller cannot tell the two apart. The block error injection code (which is the sole user of this function currently) relies on error_inject_add() later to reject adding a rule with status=BLK_STS_OK. This is okay for the current state of error injection. The delay option added in the next commit makes the rule with status=OK valid, meaning the function tag_to_blk_status() now needs to explicitly match BLK_STS_OK for it, and fail for a tag it does not recognize. Hence make tag_to_blk_status() take a second param to update the matched status, and return true in case of a successful match. If a match is not found, the function tag_to_blk_status() returns false and *status remains unchanged. A repeated status= where an invalid tag comes first is now rejected instead of being overridden by the later one. Cc: Christoph Hellwig <hch@lst.de> Assisted-by: Claude:claude-opus-5 Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev> --- block/blk-core.c | 14 ++++++-------- block/blk.h | 2 +- block/error-injection.c | 7 ++++--- 3 files changed, 11 insertions(+), 12 deletions(-) diff --git a/block/blk-core.c b/block/blk-core.c index 13dc70e8f55d..c45846b9c50d 100644 --- a/block/blk-core.c +++ b/block/blk-core.c @@ -225,21 +225,19 @@ const char *blk_status_to_tag(blk_status_t status) return blk_errors[idx].tag; } -blk_status_t tag_to_blk_status(const char *tag) +bool tag_to_blk_status(const char *tag, blk_status_t *status) { int i; for (i = 0; i < ARRAY_SIZE(blk_errors); i++) { if (blk_errors[i].tag && - !strcmp(blk_errors[i].tag, tag)) - return (__force blk_status_t)i; + !strcmp(blk_errors[i].tag, tag)) { + *status = (__force blk_status_t)i; + return true; + } } - /* - * Return BLK_STS_OK for mismatches as this function is intended to - * parse error status values. - */ - return BLK_STS_OK; + return false; } /** diff --git a/block/blk.h b/block/blk.h index 2cc03aa54c53..5fa54162c686 100644 --- a/block/blk.h +++ b/block/blk.h @@ -52,7 +52,7 @@ void blk_free_flush_queue(struct blk_flush_queue *q); const char *blk_status_to_str(blk_status_t status); const char *blk_status_to_tag(blk_status_t status); -blk_status_t tag_to_blk_status(const char *tag); +bool tag_to_blk_status(const char *tag, blk_status_t *status); enum req_op str_to_blk_op(const char *op); bool __blk_mq_unfreeze_queue(struct request_queue *q, bool force_atomic); diff --git a/block/error-injection.c b/block/error-injection.c index e14bc4b723ef..db543fe27630 100644 --- a/block/error-injection.c +++ b/block/error-injection.c @@ -171,15 +171,16 @@ static int match_op(substring_t *args, enum req_op *op) static int match_status(substring_t *args, blk_status_t *status) { const char *tag; + bool found; tag = match_strdup(args); if (!tag) return -ENOMEM; - *status = tag_to_blk_status(tag); - if (!*status) + found = tag_to_blk_status(tag, status); + if (!found) pr_warn("invalid status '%s'\n", tag); kfree(tag); - return 0; + return found ? 0 : -EINVAL; } static ssize_t blk_error_injection_parse_options(struct gendisk *disk, -- 2.53.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [v3 for-next 1/3] block: Reject unknown status tags in error injection rules 2026-09-28 22:16 ` [v3 for-next 1/3] block: Reject unknown status tags in error injection rules Md Haris Iqbal @ 2026-10-05 8:49 ` Christoph Hellwig 0 siblings, 0 replies; 7+ messages in thread From: Christoph Hellwig @ 2026-10-05 8:49 UTC (permalink / raw) To: Md Haris Iqbal Cc: Jens Axboe, linux-block, linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc On Tue, Sep 29, 2026 at 12:16:32AM +0200, Md Haris Iqbal wrote: > tag_to_blk_status() returns BLK_STS_OK both for the "OK" tag and for a > tag it does not recognize, so a caller cannot tell the two apart. The > block error injection code (which is the sole user of this function > currently) relies on error_inject_add() later to reject adding a rule with > status=BLK_STS_OK. This is okay for the current state of error injection. > > The delay option added in the next commit makes the rule with status=OK > valid, meaning the function tag_to_blk_status() now needs to explicitly > match BLK_STS_OK for it, and fail for a tag it does not recognize. Hence > make tag_to_blk_status() take a second param to update the matched status, > and return true in case of a successful match. If a match is not found, > the function tag_to_blk_status() returns false and *status remains > unchanged. > > A repeated status= where an invalid tag comes first is now rejected > instead of being overridden by the later one. > > Cc: Christoph Hellwig <hch@lst.de> > Assisted-by: Claude:claude-opus-5 > Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev> > --- > block/blk-core.c | 14 ++++++-------- > block/blk.h | 2 +- > block/error-injection.c | 7 ++++--- > 3 files changed, 11 insertions(+), 12 deletions(-) > > diff --git a/block/blk-core.c b/block/blk-core.c > index 13dc70e8f55d..c45846b9c50d 100644 > --- a/block/blk-core.c > +++ b/block/blk-core.c > @@ -225,21 +225,19 @@ const char *blk_status_to_tag(blk_status_t status) > return blk_errors[idx].tag; > } > > -blk_status_t tag_to_blk_status(const char *tag) > +bool tag_to_blk_status(const char *tag, blk_status_t *status) > { > int i; > > for (i = 0; i < ARRAY_SIZE(blk_errors); i++) { > if (blk_errors[i].tag && > - !strcmp(blk_errors[i].tag, tag)) > - return (__force blk_status_t)i; > + !strcmp(blk_errors[i].tag, tag)) { > + *status = (__force blk_status_t)i; > + return true; > + } > } > > - /* > - * Return BLK_STS_OK for mismatches as this function is intended to > - * parse error status values. > - */ > - return BLK_STS_OK; > + return false; > } > > /** > diff --git a/block/blk.h b/block/blk.h > index 2cc03aa54c53..5fa54162c686 100644 > --- a/block/blk.h > +++ b/block/blk.h > @@ -52,7 +52,7 @@ void blk_free_flush_queue(struct blk_flush_queue *q); > > const char *blk_status_to_str(blk_status_t status); > const char *blk_status_to_tag(blk_status_t status); > -blk_status_t tag_to_blk_status(const char *tag); > +bool tag_to_blk_status(const char *tag, blk_status_t *status); > enum req_op str_to_blk_op(const char *op); > > bool __blk_mq_unfreeze_queue(struct request_queue *q, bool force_atomic); > diff --git a/block/error-injection.c b/block/error-injection.c > index e14bc4b723ef..db543fe27630 100644 > --- a/block/error-injection.c > +++ b/block/error-injection.c > @@ -171,15 +171,16 @@ static int match_op(substring_t *args, enum req_op *op) > static int match_status(substring_t *args, blk_status_t *status) > { > const char *tag; > + bool found; > > tag = match_strdup(args); > if (!tag) > return -ENOMEM; > - *status = tag_to_blk_status(tag); > - if (!*status) > + found = tag_to_blk_status(tag, status); > + if (!found) > pr_warn("invalid status '%s'\n", tag); > kfree(tag); > - return 0; > + return found ? 0 : -EINVAL; if (!found) { pr_warn("invalid status '%s'\n", tag); return -EINVAL; } return 0; to keep the code a bit more readable. Otherwise looks good: Reviewed-by: Christoph Hellwig <hch@lst.de> ^ permalink raw reply [flat|nested] 7+ messages in thread
* [v3 for-next 2/3] block: allow error injection rules to delay bios 2026-09-28 22:16 [v3 for-next 0/3] block: delay support for error injection Md Haris Iqbal 2026-09-28 22:16 ` [v3 for-next 1/3] block: Reject unknown status tags in error injection rules Md Haris Iqbal @ 2026-09-28 22:16 ` Md Haris Iqbal 2026-10-05 8:56 ` Christoph Hellwig 2026-09-28 22:16 ` [v3 for-next 3/3] Documentation: block: document error injection delay feature Md Haris Iqbal 2 siblings, 1 reply; 7+ messages in thread From: Md Haris Iqbal @ 2026-09-28 22:16 UTC (permalink / raw) To: Jens Axboe, linux-block Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc, Md Haris Iqbal Error injection can only fail a bio today. Add a delay_us option so that a rule can hold a bio back first, to model a slow device. If a matching rule has delay_us set, the bio is held for that long. It is then failed if the rule also has a status, or resubmitted otherwise. The resubmission is done through submit_bio_noacct_nocheck(), which is where the injection hook sits. To avoid the same delay rule being applied to it again, mark the bio with the BIO_ERROR_INJECTED flag before queueing the work in blk_error_inject_delay(). This also prevents bios that are split later and resubmitted from being evaluated for delay (or error injection) again. __bio_clone() does not propagate the flag, so a clone a stacking driver aims at a lower device is still evaluated. The bio is submitted from a workqueue rather than from the timer, because submitting a bio can sleep. Bios with REQ_NOWAIT are never delayed, and values above 600 seconds are rejected. Cc: Christoph Hellwig <hch@lst.de> Assisted-by: Claude:claude-opus-5 Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev> --- block/error-injection.c | 160 ++++++++++++++++++++++++++++++++++---- block/error-injection.h | 1 + include/linux/blk_types.h | 1 + 3 files changed, 147 insertions(+), 15 deletions(-) diff --git a/block/error-injection.c b/block/error-injection.c index db543fe27630..a35fe41a43e4 100644 --- a/block/error-injection.c +++ b/block/error-injection.c @@ -6,9 +6,17 @@ #include <linux/blkdev.h> #include <linux/parser.h> #include <linux/seq_file.h> +#include <linux/workqueue.h> #include "blk.h" #include "error-injection.h" +/* + * Cap the delay so that a typo can't wedge a device for good. This is still + * well beyond the default hung task timeout, which is one of the things a + * delay is useful for triggering. + */ +#define BLK_ERROR_INJECT_MAX_DELAY_US (600 * USEC_PER_SEC) + struct blk_error_inject { struct list_head entry; sector_t start; @@ -18,14 +26,104 @@ struct blk_error_inject { /* only inject every 1 / chance times */ unsigned int chance; + + /* hold the bio for this long before submitting or failing it */ + unsigned int delay_us; +}; + +/* + * A bio held by a delay rule. It does not point back at the rule, so rules + * can be removed while delayed bios are outstanding. The submitter holds the + * bdev open until the bio completes, so the disk and queue stay allocated. + * A delayed bio holds no queue usage counter reference, so it does not hold + * up a queue freeze either. + */ +struct blk_error_inject_delay { + struct delayed_work dwork; + struct bio *bio; + blk_status_t status; }; +static struct workqueue_struct *blk_error_inject_wq; + DEFINE_STATIC_KEY_FALSE(blk_error_injection_enabled); +static void blk_error_inject_delay_work(struct work_struct *work) +{ + struct blk_error_inject_delay *d = container_of(to_delayed_work(work), + struct blk_error_inject_delay, dwork); + struct bio *bio = d->bio; + struct gendisk *disk = bio->bi_bdev->bd_disk; + blk_status_t status = d->status; + + kfree(d); + + if (status != BLK_STS_OK) { + pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n", + disk->part0, blk_status_to_str(status), + blk_op_str(bio_op(bio)), bio->bi_iter.bi_sector, + bio_sectors(bio)); + bio->bi_status = status; + bio_endio(bio); + } else { + /* + * Resubmit. The BIO_ERROR_INJECTED flag ensures that a bio + * that was delayed once skips error injection entirely from + * here on, including any other rule that covers it. + */ + submit_bio_noacct_nocheck(bio, false); + } +} + +/* + * Hand the bio to a workqueue that submits or fails it once the delay has + * expired. Both blk_mq_submit_bio() and ->submit_bio can sleep, so this can't + * be completed from the timer itself. + * + * Returns false if the bio can't be delayed, in which case the caller handles + * it immediately instead. + */ +static bool blk_error_inject_delay(struct gendisk *disk, struct bio *bio, + blk_status_t status, unsigned int delay_us) +{ + struct blk_error_inject_delay *d; + + /* never block a bio that asked not to be blocked */ + if (bio->bi_opf & REQ_NOWAIT) + return false; + + d = kmalloc_obj(*d, GFP_NOIO); + if (!d) + return false; + + pr_info_ratelimited("%pg: delaying %s at sector %llu:%u by %uus\n", + disk->part0, blk_op_str(bio_op(bio)), + bio->bi_iter.bi_sector, bio_sectors(bio), delay_us); + + d->bio = bio; + d->status = status; + INIT_DELAYED_WORK(&d->dwork, blk_error_inject_delay_work); + + /* + * Mark the bio before queueing the work, which can complete it as soon + * as it is queued. Splitting happens below the injection hook, but + * bio_submit_split_bioset() resubmits the remainder through the hook + * again, and as bio_split() only advances the original bio that + * remainder still matches the same rule. Without this a bio would be + * delayed once per split. + */ + bio_set_flag(bio, BIO_ERROR_INJECTED); + queue_delayed_work(blk_error_inject_wq, &d->dwork, + usecs_to_jiffies(delay_us)); + return true; +} + bool __blk_error_inject(struct bio *bio) { struct gendisk *disk = bio->bi_bdev->bd_disk; struct blk_error_inject *inj; + blk_status_t status = BLK_STS_OK; + unsigned int delay_us = 0; rcu_read_lock(); list_for_each_entry_rcu(inj, &disk->error_injection_list, entry) { @@ -45,29 +143,38 @@ bool __blk_error_inject(struct bio *bio) if (inj->chance > 1 && (get_random_u32() % inj->chance) != 0) continue; - pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n", - disk->part0, blk_status_to_str(inj->status), - blk_op_str(inj->op), bio->bi_iter.bi_sector, - bio_sectors(bio)); - bio->bi_status = inj->status; - rcu_read_unlock(); - bio_endio(bio); - return true; + status = inj->status; + delay_us = inj->delay_us; + break; } rcu_read_unlock(); - return false; + + if (delay_us && blk_error_inject_delay(disk, bio, status, delay_us)) + return true; + if (status == BLK_STS_OK) + return false; + + pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n", + disk->part0, blk_status_to_str(status), + blk_op_str(bio_op(bio)), bio->bi_iter.bi_sector, + bio_sectors(bio)); + bio->bi_status = status; + bio_endio(bio); + return true; } static int error_inject_add(struct gendisk *disk, enum req_op op, sector_t start, u64 nr_sectors, blk_status_t status, - unsigned int chance) + unsigned int chance, unsigned int delay_us) { struct blk_error_inject *inj; int error = -EINVAL; if (op == REQ_OP_LAST) return -EINVAL; - if (status == BLK_STS_OK) + if (status == BLK_STS_OK && !delay_us) + return -EINVAL; + if (delay_us > BLK_ERROR_INJECT_MAX_DELAY_US) return -EINVAL; inj = kzalloc_obj(*inj); @@ -86,6 +193,7 @@ static int error_inject_add(struct gendisk *disk, enum req_op op, inj->start = start; inj->status = status; inj->chance = chance; + inj->delay_us = delay_us; pr_debug_ratelimited("%pg: adding %s injection for %s at sector %llu:%llu\n", disk->part0, blk_status_to_str(status), @@ -139,6 +247,7 @@ enum options { Opt_nr_sectors = (1u << 18), Opt_status = (1u << 19), Opt_chance = (1u << 20), + Opt_delay_us = (1u << 21), Opt_invalid, }; @@ -151,6 +260,7 @@ static const match_table_t opt_tokens = { { Opt_nr_sectors, "nr_sectors=%u" }, { Opt_status, "status=%s" }, { Opt_chance, "chance=%u" }, + { Opt_delay_us, "delay_us=%u" }, { Opt_invalid, NULL, }, }; @@ -187,7 +297,7 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk, char *options) { enum { Unset, Add, Removeall } action = Unset; - unsigned int option_mask = 0, chance = 1; + unsigned int option_mask = 0, chance = 1, delay_us = 0; enum req_op op = REQ_OP_LAST; u64 start = 0, nr_sectors = 0; blk_status_t status = BLK_STS_OK; @@ -230,6 +340,9 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk, if (!error && chance == 0) error = -EINVAL; break; + case Opt_delay_us: + error = match_uint(args, &delay_us); + break; default: pr_warn("unknown parameter or missing value '%s'\n", p); error = -EINVAL; @@ -241,7 +354,7 @@ static ssize_t blk_error_injection_parse_options(struct gendisk *disk, switch (action) { case Add: return error_inject_add(disk, op, start, nr_sectors, status, - chance); + chance, delay_us); case Removeall: if (option_mask & ~Opt_removeall) return -EINVAL; @@ -277,10 +390,11 @@ static int blk_error_injection_show(struct seq_file *s, void *private) rcu_read_lock(); list_for_each_entry_rcu(inj, &disk->error_injection_list, entry) { - seq_printf(s, "%llu:%llu op=%s,status=%s,chance=%u", + seq_printf(s, "%llu:%llu op=%s,status=%s,chance=%u,delay_us=%u", inj->start, inj->end, blk_op_str(inj->op), - blk_status_to_tag(inj->status), inj->chance); + blk_status_to_tag(inj->status), inj->chance, + inj->delay_us); seq_putc(s, '\n'); } rcu_read_unlock(); @@ -315,3 +429,19 @@ void blk_error_injection_exit(struct gendisk *disk) { error_inject_removeall(disk); } + +static int __init blk_error_injection_init_wq(void) +{ + /* + * WQ_MEM_RECLAIM so that a delayed bio on the reclaim path can still + * find a worker under memory pressure. Note that this only guarantees + * a worker exists, not that it is free: submitting a bio can block on + * a queue freeze or on tag allocation, so a delayed bio can still be + * held up behind another one. + */ + blk_error_inject_wq = alloc_workqueue("blk_error_inject", WQ_MEM_RECLAIM | WQ_UNBOUND, 0); + if (!blk_error_inject_wq) + panic("Failed to create blk_error_inject wq\n"); + return 0; +} +subsys_initcall(blk_error_injection_init_wq); diff --git a/block/error-injection.h b/block/error-injection.h index 9821d773abab..8b3809d85e83 100644 --- a/block/error-injection.h +++ b/block/error-injection.h @@ -13,6 +13,7 @@ static inline bool blk_error_inject(struct bio *bio) { if (IS_ENABLED(CONFIG_BLK_ERROR_INJECTION) && static_branch_unlikely(&blk_error_injection_enabled) && + !bio_flagged(bio, BIO_ERROR_INJECTED) && test_bit(GD_ERROR_INJECT, &bio->bi_bdev->bd_disk->state)) return __blk_error_inject(bio); return false; diff --git a/include/linux/blk_types.h b/include/linux/blk_types.h index 98e21b4cbf32..50e28dc0f1f9 100644 --- a/include/linux/blk_types.h +++ b/include/linux/blk_types.h @@ -323,6 +323,7 @@ enum { BIO_ZONE_WRITE_PLUGGING, /* bio handled through zone write plugging */ BIO_EMULATES_ZONE_APPEND, /* bio emulates a zone append operation */ BIO_COMPLETE_IN_TASK, /* complete bi_end_io() in task context */ + BIO_ERROR_INJECTED, /* error injection rules already applied */ BIO_FLAG_LAST }; -- 2.53.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [v3 for-next 2/3] block: allow error injection rules to delay bios 2026-09-28 22:16 ` [v3 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal @ 2026-10-05 8:56 ` Christoph Hellwig 0 siblings, 0 replies; 7+ messages in thread From: Christoph Hellwig @ 2026-10-05 8:56 UTC (permalink / raw) To: Md Haris Iqbal Cc: Jens Axboe, linux-block, linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc On Tue, Sep 29, 2026 at 12:16:33AM +0200, Md Haris Iqbal wrote: > +static void blk_error_inject_delay_work(struct work_struct *work) > +{ > + struct blk_error_inject_delay *d = container_of(to_delayed_work(work), > + struct blk_error_inject_delay, dwork); > + struct bio *bio = d->bio; > + struct gendisk *disk = bio->bi_bdev->bd_disk; > + blk_status_t status = d->status; > + > + kfree(d); > + > + if (status != BLK_STS_OK) { > + pr_info_ratelimited("%pg: injecting %s error for %s at sector %llu:%u\n", > + disk->part0, blk_status_to_str(status), > + blk_op_str(bio_op(bio)), bio->bi_iter.bi_sector, > + bio_sectors(bio)); > + bio->bi_status = status; > + bio_endio(bio); This duplicates the tail of __blk_error_inject, please share the code. > + } else { And return after this to keep the other path straight line. > +/* > + * Hand the bio to a workqueue that submits or fails it once the delay has > + * expired. Both blk_mq_submit_bio() and ->submit_bio can sleep, so this can't > + * be completed from the timer itself. > + * > + * Returns false if the bio can't be delayed, in which case the caller handles > + * it immediately instead. > + */ I'd drop the comment, it just states what is obvious from the code below. > +static bool blk_error_inject_delay(struct gendisk *disk, struct bio *bio, > + blk_status_t status, unsigned int delay_us) > +{ > + struct blk_error_inject_delay *d; > + > + /* never block a bio that asked not to be blocked */ > + if (bio->bi_opf & REQ_NOWAIT) > + return false; We should probably complete it with BLK_STS_AGAIN as we would do for a real delay? > + > + d = kmalloc_obj(*d, GFP_NOIO); > + if (!d) > + return false; Should we log a warning that the intended delay did not happen? > + /* > + * Mark the bio before queueing the work, which can complete it as soon > + * as it is queued. Splitting happens below the injection hook, but > + * bio_submit_split_bioset() resubmits the remainder through the hook > + * again, and as bio_split() only advances the original bio that > + * remainder still matches the same rule. Without this a bio would be > + * delayed once per split. > + */ Over-eager comment. But we should probably also set the flag for non-delayed bios anyway and do it as soon as any rule matches? > +static int __init blk_error_injection_init_wq(void) > +{ > + /* > + * WQ_MEM_RECLAIM so that a delayed bio on the reclaim path can still > + * find a worker under memory pressure. Note that this only guarantees > + * a worker exists, not that it is free: submitting a bio can block on > + * a queue freeze or on tag allocation, so a delayed bio can still be > + * held up behind another one. > + */ Very long comment explaining the obvious, everything doing I/O needs WQ_MEM_RECLAIM. Please drop it. > + blk_error_inject_wq = alloc_workqueue("blk_error_inject", WQ_MEM_RECLAIM | WQ_UNBOUND, 0); Overly long line. > diff --git a/include/linux/blk_types.h b/include/linux/blk_types.h > index 98e21b4cbf32..50e28dc0f1f9 100644 > --- a/include/linux/blk_types.h > +++ b/include/linux/blk_types.h > @@ -323,6 +323,7 @@ enum { > BIO_ZONE_WRITE_PLUGGING, /* bio handled through zone write plugging */ > BIO_EMULATES_ZONE_APPEND, /* bio emulates a zone append operation */ > BIO_COMPLETE_IN_TASK, /* complete bi_end_io() in task context */ > + BIO_ERROR_INJECTED, /* error injection rules already applied */ > BIO_FLAG_LAST > }; With this we've used up the last BIO_ flag. I hope this won't cause problems in the future.. ^ permalink raw reply [flat|nested] 7+ messages in thread
* [v3 for-next 3/3] Documentation: block: document error injection delay feature 2026-09-28 22:16 [v3 for-next 0/3] block: delay support for error injection Md Haris Iqbal 2026-09-28 22:16 ` [v3 for-next 1/3] block: Reject unknown status tags in error injection rules Md Haris Iqbal 2026-09-28 22:16 ` [v3 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal @ 2026-09-28 22:16 ` Md Haris Iqbal 2026-10-05 8:56 ` Christoph Hellwig 2 siblings, 1 reply; 7+ messages in thread From: Md Haris Iqbal @ 2026-09-28 22:16 UTC (permalink / raw) To: Jens Axboe, linux-block Cc: linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc, Md Haris Iqbal Document the delay_us option: the delay happens above the driver, a delayed bio is not run through the rules again, and holding a bio back reorders it against bios submitted later. Cc: Christoph Hellwig <hch@lst.de> Assisted-by: Claude:claude-opus-5 Signed-off-by: Md Haris Iqbal <haris.iqbal@linux.dev> --- Documentation/block/error-injection.rst | 39 +++++++++++++++++++++---- 1 file changed, 34 insertions(+), 5 deletions(-) diff --git a/Documentation/block/error-injection.rst b/Documentation/block/error-injection.rst index 81f31af82e65..6f7609b91d24 100644 --- a/Documentation/block/error-injection.rst +++ b/Documentation/block/error-injection.rst @@ -7,9 +7,9 @@ Configurable Error Injection Overview -------- -Configurable error injection allows injecting specific block layer status codes -for sector ranges of a block device. Errors can be injected unconditionally, or -with a given probability. +Configurable error injection allows injecting delays and/or specific block +layer status codes for sector ranges of a block device by adding rules. The +rules can be configured to trigger unconditionally, or with a given probability. To use configurable error injection, CONFIG_BLK_ERROR_INJECTION must be enabled. @@ -34,15 +34,36 @@ op=<string> block layer operation this rule applies to. This uses the XYZ for each REQ_OP_XYZ operation, e.g. READ, WRITE or DISCARD. Mandatory. status=<string> Status to return. This uses XYZ for each BLK_STS_XYZ - code, e.g. IOERR or MEDIUM. Mandatory. + code, e.g. IOERR or MEDIUM. Mandatory unless delay_us + is given. start=<number> First block layer sector the rule applies to. Optional, defaults to 0. nr_sectors=<number> Number of sectors this rule applies. Optional, defaults to the remainder of the device. -chance=<number> Only return a failure with a likelihood of 1/chance. +chance=<number> Only apply the rule with a likelihood of 1/chance. Optional, defaults to 1 (always). +delay_us=<number> Hold the bio back for this many microseconds. Without + status the bio is then submitted to the device as + usual. With status it is failed once the delay has + expired. Bios with REQ_NOWAIT set are never delayed. + Optional, defaults to 0 (no delay). Values above 600 + seconds are rejected. =================== ======================================================= +Delays +------ + +A delayed bio is held above the driver, so the device never sees a slow I/O. +A delay does not reach the blk-mq timeout handler or SCSI error handling. + +A bio that matched a delay rule is not evaluated against the rules again, even +after it is split and resubmitted internally. Removing a rule does not release +bios it is already delaying. + +Holding a bio back reorders it against bios submitted later. On zoned devices +this breaks sequential write ordering and the block layer fails the +out-of-order writes, so only delay reads there. + Example ------- @@ -54,6 +75,14 @@ Return BLK_STS_MEDIUM for every write to /dev/nvme0n1: $ echo 'add,op=WRITE,start=0,status=MEDIUM' > /sys/kernel/debug/block/nvme0n1/error_injection +Delay every read of /dev/nvme0n1 by 10 milliseconds, then issue it normally: + + $ echo 'add,op=READ,delay_us=10000' > /sys/kernel/debug/block/nvme0n1/error_injection + +Fail one in 100 writes with BLK_STS_TIMEOUT, but only after 30 seconds: + + $ echo 'add,op=WRITE,status=TIMEOUT,chance=100,delay_us=30000000' > /sys/kernel/debug/block/nvme0n1/error_injection + Remove all rules for /dev/nvme0n1: $ echo 'removeall' > /sys/kernel/debug/block/nvme0n1/error_injection -- 2.53.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [v3 for-next 3/3] Documentation: block: document error injection delay feature 2026-09-28 22:16 ` [v3 for-next 3/3] Documentation: block: document error injection delay feature Md Haris Iqbal @ 2026-10-05 8:56 ` Christoph Hellwig 0 siblings, 0 replies; 7+ messages in thread From: Christoph Hellwig @ 2026-10-05 8:56 UTC (permalink / raw) To: Md Haris Iqbal Cc: Jens Axboe, linux-block, linux-kernel, Christoph Hellwig, Keith Busch, Jonathan Corbet, linux-doc Looks good: Reviewed-by: Christoph Hellwig <hch@lst.de> ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-05 8:56 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-09-28 22:16 [v3 for-next 0/3] block: delay support for error injection Md Haris Iqbal 2026-09-28 22:16 ` [v3 for-next 1/3] block: Reject unknown status tags in error injection rules Md Haris Iqbal 2026-10-05 8:49 ` Christoph Hellwig 2026-09-28 22:16 ` [v3 for-next 2/3] block: allow error injection rules to delay bios Md Haris Iqbal 2026-10-05 8:56 ` Christoph Hellwig 2026-09-28 22:16 ` [v3 for-next 3/3] Documentation: block: document error injection delay feature Md Haris Iqbal 2026-10-05 8:56 ` Christoph Hellwig
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®