* [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration
@ 2026-03-04 7:38 Yu Kuai
2026-03-04 7:38 ` [PATCH v3 1/7] blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with blkcg_mutex Yu Kuai
` (7 more replies)
0 siblings, 8 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
This series fixes several race conditions related to q->blkg_list iteration
and improves the locking around blkcg policy activation/deactivation.
Patch 1-2: Protect q->blkg_list iteration with blkcg_mutex in blkg_destroy_all()
and bfq_end_wr_async() to prevent races with blkg_free_workfn().
Patch 3-4: Fix use-after-free and memory leak issues in blkcg_activate_policy()
by extending blkcg_mutex coverage and skipping dying blkgs.
Patch 5: Refactor policy pd teardown into a helper function.
Patch 6: Restructure blkcg_activate_policy() to allocate pds before freezing
the queue, avoiding potential deadlocks from percpu allocation. Also fix
locking order in blkcg_deactivate_policy() to be consistent with
blkcg_activate_policy() (mutex -> freeze).
Patch 7: Move rq_qos_mutex handling inside rq_qos_add()/rq_qos_del() to
simplify the locking and eliminate potential deadlocks.
Note: queue_lock is still used in many places to protect queue blkg.
Future work is to convert it to blkcg_mutex entirely.
Changes v2 -> v3:
- Patch 2: Wrap mutex_lock/unlock with #ifdef CONFIG_BFQ_GROUP_IOSCHED to
fix compile error when CONFIG_BLK_CGROUP is disabled.
- Patch 6: Fix locking order in blkcg_deactivate_policy() to match
blkcg_activate_policy() (mutex -> freeze instead of freeze -> mutex).
- Patch 7: Remove stale lockdep_assert_held() in iolatency_set_limit().
Changes v1 -> v2:
- Link: https://lore.kernel.org/all/20260108014416.3656493-1-zhengqixing@huaweicloud.com/
Yu Kuai (4):
blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with
blkcg_mutex
bfq: protect q->blkg_list iteration in bfq_end_wr_async() with
blkcg_mutex
blk-cgroup: allocate pds before freezing queue in
blkcg_activate_policy()
blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del
Zheng Qixing (3):
blk-cgroup: fix race between policy activation and blkg destruction
blk-cgroup: skip dying blkg in blkcg_activate_policy()
blk-cgroup: factor policy pd teardown loop into helper
block/bfq-cgroup.c | 3 +-
block/bfq-iosched.c | 6 ++
block/blk-cgroup.c | 205 ++++++++++++++----------------------------
block/blk-cgroup.h | 2 -
block/blk-iocost.c | 11 +--
block/blk-iolatency.c | 5 --
block/blk-rq-qos.c | 31 ++++---
block/blk-wbt.c | 2 -
8 files changed, 97 insertions(+), 168 deletions(-)
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 1/7] blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with blkcg_mutex
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 2/7] bfq: protect q->blkg_list iteration in bfq_end_wr_async() " Yu Kuai
` (6 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
blkg_destroy_all() iterates q->blkg_list without holding blkcg_mutex,
which can race with blkg_free_workfn() that removes blkgs from the list
while holding blkcg_mutex.
Add blkcg_mutex protection around the q->blkg_list iteration to prevent
potential list corruption or use-after-free issues.
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index 3cffb68ba5d8..0bc7b19399b6 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -574,6 +574,7 @@ static void blkg_destroy_all(struct gendisk *disk)
int i;
restart:
+ mutex_lock(&q->blkcg_mutex);
spin_lock_irq(&q->queue_lock);
list_for_each_entry(blkg, &q->blkg_list, q_node) {
struct blkcg *blkcg = blkg->blkcg;
@@ -592,6 +593,7 @@ static void blkg_destroy_all(struct gendisk *disk)
if (!(--count)) {
count = BLKG_DESTROY_BATCH_SIZE;
spin_unlock_irq(&q->queue_lock);
+ mutex_unlock(&q->blkcg_mutex);
cond_resched();
goto restart;
}
@@ -611,6 +613,7 @@ static void blkg_destroy_all(struct gendisk *disk)
q->root_blkg = NULL;
spin_unlock_irq(&q->queue_lock);
+ mutex_unlock(&q->blkcg_mutex);
}
static void blkg_iostat_set(struct blkg_iostat *dst, struct blkg_iostat *src)
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 2/7] bfq: protect q->blkg_list iteration in bfq_end_wr_async() with blkcg_mutex
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
2026-03-04 7:38 ` [PATCH v3 1/7] blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with blkcg_mutex Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 3/7] blk-cgroup: fix race between policy activation and blkg destruction Yu Kuai
` (5 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
bfq_end_wr_async() iterates q->blkg_list while only holding bfqd->lock,
but not blkcg_mutex. This can race with blkg_free_workfn() that removes
blkgs from the list while holding blkcg_mutex.
Add blkcg_mutex protection in bfq_end_wr() before taking bfqd->lock to
ensure proper synchronization when iterating q->blkg_list.
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/bfq-cgroup.c | 3 ++-
block/bfq-iosched.c | 6 ++++++
2 files changed, 8 insertions(+), 1 deletion(-)
diff --git a/block/bfq-cgroup.c b/block/bfq-cgroup.c
index 6a75fe1c7a5c..839d266a6aa6 100644
--- a/block/bfq-cgroup.c
+++ b/block/bfq-cgroup.c
@@ -940,7 +940,8 @@ void bfq_end_wr_async(struct bfq_data *bfqd)
list_for_each_entry(blkg, &bfqd->queue->blkg_list, q_node) {
struct bfq_group *bfqg = blkg_to_bfqg(blkg);
- bfq_end_wr_async_queues(bfqd, bfqg);
+ if (bfqg)
+ bfq_end_wr_async_queues(bfqd, bfqg);
}
bfq_end_wr_async_queues(bfqd, bfqd->root_group);
}
diff --git a/block/bfq-iosched.c b/block/bfq-iosched.c
index b180ce583951..14ca67a616b7 100644
--- a/block/bfq-iosched.c
+++ b/block/bfq-iosched.c
@@ -2645,6 +2645,9 @@ static void bfq_end_wr(struct bfq_data *bfqd)
struct bfq_queue *bfqq;
int i;
+#ifdef CONFIG_BFQ_GROUP_IOSCHED
+ mutex_lock(&bfqd->queue->blkcg_mutex);
+#endif
spin_lock_irq(&bfqd->lock);
for (i = 0; i < bfqd->num_actuators; i++) {
@@ -2656,6 +2659,9 @@ static void bfq_end_wr(struct bfq_data *bfqd)
bfq_end_wr_async(bfqd);
spin_unlock_irq(&bfqd->lock);
+#ifdef CONFIG_BFQ_GROUP_IOSCHED
+ mutex_unlock(&bfqd->queue->blkcg_mutex);
+#endif
}
static sector_t bfq_io_struct_pos(void *io_struct, bool request)
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 3/7] blk-cgroup: fix race between policy activation and blkg destruction
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
2026-03-04 7:38 ` [PATCH v3 1/7] blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with blkcg_mutex Yu Kuai
2026-03-04 7:38 ` [PATCH v3 2/7] bfq: protect q->blkg_list iteration in bfq_end_wr_async() " Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 4/7] blk-cgroup: skip dying blkg in blkcg_activate_policy() Yu Kuai
` (4 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
From: Zheng Qixing <zhengqixing@huawei.com>
When switching an IO scheduler on a block device, blkcg_activate_policy()
allocates blkg_policy_data (pd) for all blkgs attached to the queue.
However, blkcg_activate_policy() may race with concurrent blkcg deletion,
leading to use-after-free and memory leak issues.
The use-after-free occurs in the following race:
T1 (blkcg_activate_policy):
- Successfully allocates pd for blkg1 (loop0->queue, blkcgA)
- Fails to allocate pd for blkg2 (loop0->queue, blkcgB)
- Enters the enomem rollback path to release blkg1 resources
T2 (blkcg deletion):
- blkcgA is deleted concurrently
- blkg1 is freed via blkg_free_workfn()
- blkg1->pd is freed
T1 (continued):
- Rollback path accesses blkg1->pd->online after pd is freed
- Triggers use-after-free
In addition, blkg_free_workfn() frees pd before removing the blkg from
q->blkg_list. This allows blkcg_activate_policy() to allocate a new pd
for a blkg that is being destroyed, leaving the newly allocated pd
unreachable when the blkg is finally freed.
Fix these races by extending blkcg_mutex coverage to serialize
blkcg_activate_policy() rollback and blkg destruction, ensuring pd
lifecycle is synchronized with blkg list visibility.
Link: https://lore.kernel.org/all/20260108014416.3656493-3-zhengqixing@huaweicloud.com/
Fixes: f1c006f1c685 ("blk-cgroup: synchronize pd_free_fn() from blkg_free_workfn() and blkcg_deactivate_policy()")
Signed-off-by: Zheng Qixing <zhengqixing@huawei.com>
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index 0bc7b19399b6..a6ac6ba9430d 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -1599,6 +1599,8 @@ int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
if (queue_is_mq(q))
memflags = blk_mq_freeze_queue(q);
+
+ mutex_lock(&q->blkcg_mutex);
retry:
spin_lock_irq(&q->queue_lock);
@@ -1661,6 +1663,7 @@ int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
spin_unlock_irq(&q->queue_lock);
out:
+ mutex_unlock(&q->blkcg_mutex);
if (queue_is_mq(q))
blk_mq_unfreeze_queue(q, memflags);
if (pinned_blkg)
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 4/7] blk-cgroup: skip dying blkg in blkcg_activate_policy()
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
` (2 preceding siblings ...)
2026-03-04 7:38 ` [PATCH v3 3/7] blk-cgroup: fix race between policy activation and blkg destruction Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 5/7] blk-cgroup: factor policy pd teardown loop into helper Yu Kuai
` (3 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
From: Zheng Qixing <zhengqixing@huawei.com>
When switching IO schedulers on a block device, blkcg_activate_policy()
can race with concurrent blkcg deletion, leading to a use-after-free in
rcu_accelerate_cbs.
T1: T2:
blkg_destroy
kill(&blkg->refcnt) // blkg->refcnt=1->0
blkg_release // call_rcu(__blkg_release)
...
blkg_free_workfn
->pd_free_fn(pd)
elv_iosched_store
elevator_switch
...
iterate blkg list
blkg_get(blkg) // blkg->refcnt=0->1
list_del_init(&blkg->q_node)
blkg_put(pinned_blkg) // blkg->refcnt=1->0
blkg_release // call_rcu again
rcu_accelerate_cbs // uaf
Fix this by checking hlist_unhashed(&blkg->blkcg_node) before getting
a reference to the blkg. This is the same check used in blkg_destroy()
to detect if a blkg has already been destroyed. If the blkg is already
unhashed, skip processing it since it's being destroyed.
Link: https://lore.kernel.org/all/20260108014416.3656493-4-zhengqixing@huaweicloud.com/
Fixes: f1c006f1c685 ("blk-cgroup: synchronize pd_free_fn() from blkg_free_workfn() and blkcg_deactivate_policy()")
Signed-off-by: Zheng Qixing <zhengqixing@huawei.com>
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index a6ac6ba9430d..f5b14a1d6973 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -1610,6 +1610,8 @@ int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
if (blkg->pd[pol->plid])
continue;
+ if (hlist_unhashed(&blkg->blkcg_node))
+ continue;
/* If prealloc matches, use it; otherwise try GFP_NOWAIT */
if (blkg == pinned_blkg) {
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 5/7] blk-cgroup: factor policy pd teardown loop into helper
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
` (3 preceding siblings ...)
2026-03-04 7:38 ` [PATCH v3 4/7] blk-cgroup: skip dying blkg in blkcg_activate_policy() Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 6/7] blk-cgroup: allocate pds before freezing queue in blkcg_activate_policy() Yu Kuai
` (2 subsequent siblings)
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
From: Zheng Qixing <zhengqixing@huawei.com>
Move the teardown sequence which offlines and frees per-policy
blkg_policy_data (pd) into a helper for readability.
No functional change intended.
Signed-off-by: Zheng Qixing <zhengqixing@huawei.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Yu Kuai <yukuai@fnnas.com>
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 58 +++++++++++++++++++++-------------------------
1 file changed, 27 insertions(+), 31 deletions(-)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index f5b14a1d6973..0206050f81ea 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -1562,6 +1562,31 @@ struct cgroup_subsys io_cgrp_subsys = {
};
EXPORT_SYMBOL_GPL(io_cgrp_subsys);
+/*
+ * Tear down per-blkg policy data for @pol on @q.
+ */
+static void blkcg_policy_teardown_pds(struct request_queue *q,
+ const struct blkcg_policy *pol)
+{
+ struct blkcg_gq *blkg;
+
+ list_for_each_entry(blkg, &q->blkg_list, q_node) {
+ struct blkcg *blkcg = blkg->blkcg;
+ struct blkg_policy_data *pd;
+
+ spin_lock(&blkcg->lock);
+ pd = blkg->pd[pol->plid];
+ if (pd) {
+ if (pd->online && pol->pd_offline_fn)
+ pol->pd_offline_fn(pd);
+ pd->online = false;
+ pol->pd_free_fn(pd);
+ blkg->pd[pol->plid] = NULL;
+ }
+ spin_unlock(&blkcg->lock);
+ }
+}
+
/**
* blkcg_activate_policy - activate a blkcg policy on a gendisk
* @disk: gendisk of interest
@@ -1677,21 +1702,7 @@ int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
enomem:
/* alloc failed, take down everything */
spin_lock_irq(&q->queue_lock);
- list_for_each_entry(blkg, &q->blkg_list, q_node) {
- struct blkcg *blkcg = blkg->blkcg;
- struct blkg_policy_data *pd;
-
- spin_lock(&blkcg->lock);
- pd = blkg->pd[pol->plid];
- if (pd) {
- if (pd->online && pol->pd_offline_fn)
- pol->pd_offline_fn(pd);
- pd->online = false;
- pol->pd_free_fn(pd);
- blkg->pd[pol->plid] = NULL;
- }
- spin_unlock(&blkcg->lock);
- }
+ blkcg_policy_teardown_pds(q, pol);
spin_unlock_irq(&q->queue_lock);
ret = -ENOMEM;
goto out;
@@ -1710,7 +1721,6 @@ void blkcg_deactivate_policy(struct gendisk *disk,
const struct blkcg_policy *pol)
{
struct request_queue *q = disk->queue;
- struct blkcg_gq *blkg;
unsigned int memflags;
if (!blkcg_policy_enabled(q, pol))
@@ -1721,22 +1731,8 @@ void blkcg_deactivate_policy(struct gendisk *disk,
mutex_lock(&q->blkcg_mutex);
spin_lock_irq(&q->queue_lock);
-
__clear_bit(pol->plid, q->blkcg_pols);
-
- list_for_each_entry(blkg, &q->blkg_list, q_node) {
- struct blkcg *blkcg = blkg->blkcg;
-
- spin_lock(&blkcg->lock);
- if (blkg->pd[pol->plid]) {
- if (blkg->pd[pol->plid]->online && pol->pd_offline_fn)
- pol->pd_offline_fn(blkg->pd[pol->plid]);
- pol->pd_free_fn(blkg->pd[pol->plid]);
- blkg->pd[pol->plid] = NULL;
- }
- spin_unlock(&blkcg->lock);
- }
-
+ blkcg_policy_teardown_pds(q, pol);
spin_unlock_irq(&q->queue_lock);
mutex_unlock(&q->blkcg_mutex);
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 6/7] blk-cgroup: allocate pds before freezing queue in blkcg_activate_policy()
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
` (4 preceding siblings ...)
2026-03-04 7:38 ` [PATCH v3 5/7] blk-cgroup: factor policy pd teardown loop into helper Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-04 7:38 ` [PATCH v3 7/7] blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del Yu Kuai
2026-03-23 8:11 ` [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
Some policies like iocost and iolatency perform percpu allocation in
pd_alloc_fn(). Percpu allocation with queue frozen can cause deadlock
because percpu memory reclaim may issue IO.
Now that q->blkg_list is protected by blkcg_mutex, restructure
blkcg_activate_policy() to allocate all pds before freezing the queue:
1. Allocate all pds with GFP_KERNEL before freezing the queue
2. Freeze the queue
3. Initialize and online all pds
Note: Future work is to remove all queue freezing before
blkcg_activate_policy() to fix the deadlocks thoroughly.
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 95 +++++++++++++++++-----------------------------
1 file changed, 35 insertions(+), 60 deletions(-)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index 0206050f81ea..1620be75f124 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -1606,8 +1606,7 @@ static void blkcg_policy_teardown_pds(struct request_queue *q,
int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
{
struct request_queue *q = disk->queue;
- struct blkg_policy_data *pd_prealloc = NULL;
- struct blkcg_gq *blkg, *pinned_blkg = NULL;
+ struct blkcg_gq *blkg;
unsigned int memflags;
int ret;
@@ -1622,90 +1621,65 @@ int blkcg_activate_policy(struct gendisk *disk, const struct blkcg_policy *pol)
if (WARN_ON_ONCE(!pol->pd_alloc_fn || !pol->pd_free_fn))
return -EINVAL;
- if (queue_is_mq(q))
- memflags = blk_mq_freeze_queue(q);
-
+ /*
+ * Allocate all pds before freezing queue. Some policies like iocost
+ * and iolatency do percpu allocation in pd_alloc_fn(), which can
+ * deadlock with queue frozen because percpu memory reclaim may issue
+ * IO. blkcg_mutex protects q->blkg_list iteration.
+ */
mutex_lock(&q->blkcg_mutex);
-retry:
- spin_lock_irq(&q->queue_lock);
-
- /* blkg_list is pushed at the head, reverse walk to initialize parents first */
list_for_each_entry_reverse(blkg, &q->blkg_list, q_node) {
struct blkg_policy_data *pd;
- if (blkg->pd[pol->plid])
- continue;
+ /* Skip dying blkg */
if (hlist_unhashed(&blkg->blkcg_node))
continue;
- /* If prealloc matches, use it; otherwise try GFP_NOWAIT */
- if (blkg == pinned_blkg) {
- pd = pd_prealloc;
- pd_prealloc = NULL;
- } else {
- pd = pol->pd_alloc_fn(disk, blkg->blkcg,
- GFP_NOWAIT);
- }
-
+ pd = pol->pd_alloc_fn(disk, blkg->blkcg, GFP_KERNEL);
if (!pd) {
- /*
- * GFP_NOWAIT failed. Free the existing one and
- * prealloc for @blkg w/ GFP_KERNEL.
- */
- if (pinned_blkg)
- blkg_put(pinned_blkg);
- blkg_get(blkg);
- pinned_blkg = blkg;
-
- spin_unlock_irq(&q->queue_lock);
-
- if (pd_prealloc)
- pol->pd_free_fn(pd_prealloc);
- pd_prealloc = pol->pd_alloc_fn(disk, blkg->blkcg,
- GFP_KERNEL);
- if (pd_prealloc)
- goto retry;
- else
- goto enomem;
+ ret = -ENOMEM;
+ goto err_teardown;
}
- spin_lock(&blkg->blkcg->lock);
-
pd->blkg = blkg;
pd->plid = pol->plid;
+ pd->online = false;
blkg->pd[pol->plid] = pd;
+ }
+
+ /* Now freeze queue and initialize/online all pds */
+ if (queue_is_mq(q))
+ memflags = blk_mq_freeze_queue(q);
+ spin_lock_irq(&q->queue_lock);
+ list_for_each_entry_reverse(blkg, &q->blkg_list, q_node) {
+ struct blkg_policy_data *pd = blkg->pd[pol->plid];
+
+ /* Skip dying blkg */
+ if (hlist_unhashed(&blkg->blkcg_node))
+ continue;
+
+ spin_lock(&blkg->blkcg->lock);
if (pol->pd_init_fn)
pol->pd_init_fn(pd);
-
if (pol->pd_online_fn)
pol->pd_online_fn(pd);
pd->online = true;
-
spin_unlock(&blkg->blkcg->lock);
}
__set_bit(pol->plid, q->blkcg_pols);
- ret = 0;
-
spin_unlock_irq(&q->queue_lock);
-out:
- mutex_unlock(&q->blkcg_mutex);
+
if (queue_is_mq(q))
blk_mq_unfreeze_queue(q, memflags);
- if (pinned_blkg)
- blkg_put(pinned_blkg);
- if (pd_prealloc)
- pol->pd_free_fn(pd_prealloc);
- return ret;
+ mutex_unlock(&q->blkcg_mutex);
+ return 0;
-enomem:
- /* alloc failed, take down everything */
- spin_lock_irq(&q->queue_lock);
+err_teardown:
blkcg_policy_teardown_pds(q, pol);
- spin_unlock_irq(&q->queue_lock);
- ret = -ENOMEM;
- goto out;
+ mutex_unlock(&q->blkcg_mutex);
+ return ret;
}
EXPORT_SYMBOL_GPL(blkcg_activate_policy);
@@ -1726,18 +1700,19 @@ void blkcg_deactivate_policy(struct gendisk *disk,
if (!blkcg_policy_enabled(q, pol))
return;
+ /* Same locking order as blkcg_activate_policy(): mutex -> freeze */
+ mutex_lock(&q->blkcg_mutex);
if (queue_is_mq(q))
memflags = blk_mq_freeze_queue(q);
- mutex_lock(&q->blkcg_mutex);
spin_lock_irq(&q->queue_lock);
__clear_bit(pol->plid, q->blkcg_pols);
blkcg_policy_teardown_pds(q, pol);
spin_unlock_irq(&q->queue_lock);
- mutex_unlock(&q->blkcg_mutex);
if (queue_is_mq(q))
blk_mq_unfreeze_queue(q, memflags);
+ mutex_unlock(&q->blkcg_mutex);
}
EXPORT_SYMBOL_GPL(blkcg_deactivate_policy);
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 7/7] blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
` (5 preceding siblings ...)
2026-03-04 7:38 ` [PATCH v3 6/7] blk-cgroup: allocate pds before freezing queue in blkcg_activate_policy() Yu Kuai
@ 2026-03-04 7:38 ` Yu Kuai
2026-03-23 8:11 ` [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-04 7:38 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei, Nilay Shroff
The current rq_qos_mutex handling has an awkward pattern where callers
must acquire the mutex before calling rq_qos_add()/rq_qos_del(), and
blkg_conf_open_bdev_frozen() had to release and re-acquire the mutex
around queue freezing to maintain proper locking order (freeze queue
before mutex).
On the other hand, with rq_qos_mutex held after blkg_conf_prep(), there
are many possible deadlocks:
- allocating memory with GFP_KERNEL, like blk_throtl_init();
- allocating percpu memory, like pd_alloc_fn() for iocost/iolatency;
This patch refactors the locking by:
1. Moving queue freeze and rq_qos_mutex acquisition inside
rq_qos_add()/rq_qos_del(), with the correct order: freeze first,
then acquire mutex.
2. Removing external mutex handling from wbt_init() since rq_qos_add()
now handles it internally.
3. Removing rq_qos_mutex handling from blkg_conf_open_bdev() entirely,
making it only responsible for parsing MAJ:MIN and opening the bdev.
4. Removing blkg_conf_open_bdev_frozen() and blkg_conf_exit_frozen()
functions which are no longer needed.
5. Updating ioc_qos_write() to use the simpler blkg_conf_open_bdev()
and blkg_conf_exit() functions.
This eliminates the release-and-reacquire pattern and makes
rq_qos_add()/rq_qos_del() self-contained, which is cleaner and reduces
complexity. Each function now properly manages its own locking with
the correct order: queue freeze → mutex acquire → modify → mutex
release → queue unfreeze.
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
block/blk-cgroup.c | 50 -------------------------------------------
block/blk-cgroup.h | 2 --
block/blk-iocost.c | 11 ++++------
block/blk-iolatency.c | 5 -----
block/blk-rq-qos.c | 31 ++++++++++++++++-----------
block/blk-wbt.c | 2 --
6 files changed, 22 insertions(+), 79 deletions(-)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index 1620be75f124..02ef8f60f759 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -802,10 +802,8 @@ int blkg_conf_open_bdev(struct blkg_conf_ctx *ctx)
return -ENODEV;
}
- mutex_lock(&bdev->bd_queue->rq_qos_mutex);
if (!disk_live(bdev->bd_disk)) {
blkdev_put_no_open(bdev);
- mutex_unlock(&bdev->bd_queue->rq_qos_mutex);
return -ENODEV;
}
@@ -813,38 +811,6 @@ int blkg_conf_open_bdev(struct blkg_conf_ctx *ctx)
ctx->bdev = bdev;
return 0;
}
-/*
- * Similar to blkg_conf_open_bdev, but additionally freezes the queue,
- * ensures the correct locking order between freeze queue and q->rq_qos_mutex.
- *
- * This function returns negative error on failure. On success it returns
- * memflags which must be saved and later passed to blkg_conf_exit_frozen
- * for restoring the memalloc scope.
- */
-unsigned long __must_check blkg_conf_open_bdev_frozen(struct blkg_conf_ctx *ctx)
-{
- int ret;
- unsigned long memflags;
-
- if (ctx->bdev)
- return -EINVAL;
-
- ret = blkg_conf_open_bdev(ctx);
- if (ret < 0)
- return ret;
- /*
- * At this point, we haven’t started protecting anything related to QoS,
- * so we release q->rq_qos_mutex here, which was first acquired in blkg_
- * conf_open_bdev. Later, we re-acquire q->rq_qos_mutex after freezing
- * the queue to maintain the correct locking order.
- */
- mutex_unlock(&ctx->bdev->bd_queue->rq_qos_mutex);
-
- memflags = blk_mq_freeze_queue(ctx->bdev->bd_queue);
- mutex_lock(&ctx->bdev->bd_queue->rq_qos_mutex);
-
- return memflags;
-}
/**
* blkg_conf_prep - parse and prepare for per-blkg config update
@@ -978,7 +944,6 @@ EXPORT_SYMBOL_GPL(blkg_conf_prep);
*/
void blkg_conf_exit(struct blkg_conf_ctx *ctx)
__releases(&ctx->bdev->bd_queue->queue_lock)
- __releases(&ctx->bdev->bd_queue->rq_qos_mutex)
{
if (ctx->blkg) {
spin_unlock_irq(&bdev_get_queue(ctx->bdev)->queue_lock);
@@ -986,7 +951,6 @@ void blkg_conf_exit(struct blkg_conf_ctx *ctx)
}
if (ctx->bdev) {
- mutex_unlock(&ctx->bdev->bd_queue->rq_qos_mutex);
blkdev_put_no_open(ctx->bdev);
ctx->body = NULL;
ctx->bdev = NULL;
@@ -994,20 +958,6 @@ void blkg_conf_exit(struct blkg_conf_ctx *ctx)
}
EXPORT_SYMBOL_GPL(blkg_conf_exit);
-/*
- * Similar to blkg_conf_exit, but also unfreezes the queue. Should be used
- * when blkg_conf_open_bdev_frozen is used to open the bdev.
- */
-void blkg_conf_exit_frozen(struct blkg_conf_ctx *ctx, unsigned long memflags)
-{
- if (ctx->bdev) {
- struct request_queue *q = ctx->bdev->bd_queue;
-
- blkg_conf_exit(ctx);
- blk_mq_unfreeze_queue(q, memflags);
- }
-}
-
static void blkg_iostat_add(struct blkg_iostat *dst, struct blkg_iostat *src)
{
int i;
diff --git a/block/blk-cgroup.h b/block/blk-cgroup.h
index 1cce3294634d..d4e7f78ba545 100644
--- a/block/blk-cgroup.h
+++ b/block/blk-cgroup.h
@@ -219,11 +219,9 @@ struct blkg_conf_ctx {
void blkg_conf_init(struct blkg_conf_ctx *ctx, char *input);
int blkg_conf_open_bdev(struct blkg_conf_ctx *ctx);
-unsigned long blkg_conf_open_bdev_frozen(struct blkg_conf_ctx *ctx);
int blkg_conf_prep(struct blkcg *blkcg, const struct blkcg_policy *pol,
struct blkg_conf_ctx *ctx);
void blkg_conf_exit(struct blkg_conf_ctx *ctx);
-void blkg_conf_exit_frozen(struct blkg_conf_ctx *ctx, unsigned long memflags);
/**
* bio_issue_as_root_blkg - see if this bio needs to be issued as root blkg
diff --git a/block/blk-iocost.c b/block/blk-iocost.c
index ef543d163d46..104a9a9f563f 100644
--- a/block/blk-iocost.c
+++ b/block/blk-iocost.c
@@ -3220,16 +3220,13 @@ static ssize_t ioc_qos_write(struct kernfs_open_file *of, char *input,
u32 qos[NR_QOS_PARAMS];
bool enable, user;
char *body, *p;
- unsigned long memflags;
int ret;
blkg_conf_init(&ctx, input);
- memflags = blkg_conf_open_bdev_frozen(&ctx);
- if (IS_ERR_VALUE(memflags)) {
- ret = memflags;
+ ret = blkg_conf_open_bdev(&ctx);
+ if (ret)
goto err;
- }
body = ctx.body;
disk = ctx.bdev->bd_disk;
@@ -3346,14 +3343,14 @@ static ssize_t ioc_qos_write(struct kernfs_open_file *of, char *input,
blk_mq_unquiesce_queue(disk->queue);
- blkg_conf_exit_frozen(&ctx, memflags);
+ blkg_conf_exit(&ctx);
return nbytes;
einval:
spin_unlock_irq(&ioc->lock);
blk_mq_unquiesce_queue(disk->queue);
ret = -EINVAL;
err:
- blkg_conf_exit_frozen(&ctx, memflags);
+ blkg_conf_exit(&ctx);
return ret;
}
diff --git a/block/blk-iolatency.c b/block/blk-iolatency.c
index f7434278cd29..3f454fb3ff51 100644
--- a/block/blk-iolatency.c
+++ b/block/blk-iolatency.c
@@ -842,11 +842,6 @@ static ssize_t iolatency_set_limit(struct kernfs_open_file *of, char *buf,
if (ret)
goto out;
- /*
- * blk_iolatency_init() may fail after rq_qos_add() succeeds which can
- * confuse iolat_rq_qos() test. Make the test and init atomic.
- */
- lockdep_assert_held(&ctx.bdev->bd_queue->rq_qos_mutex);
if (!iolat_rq_qos(ctx.bdev->bd_queue))
ret = blk_iolatency_init(ctx.bdev->bd_disk);
if (ret)
diff --git a/block/blk-rq-qos.c b/block/blk-rq-qos.c
index 85cf74402a09..fe96183bcc75 100644
--- a/block/blk-rq-qos.c
+++ b/block/blk-rq-qos.c
@@ -327,8 +327,7 @@ int rq_qos_add(struct rq_qos *rqos, struct gendisk *disk, enum rq_qos_id id,
{
struct request_queue *q = disk->queue;
unsigned int memflags;
-
- lockdep_assert_held(&q->rq_qos_mutex);
+ int ret = 0;
rqos->disk = disk;
rqos->id = id;
@@ -337,20 +336,24 @@ int rq_qos_add(struct rq_qos *rqos, struct gendisk *disk, enum rq_qos_id id,
/*
* No IO can be in-flight when adding rqos, so freeze queue, which
* is fine since we only support rq_qos for blk-mq queue.
+ *
+ * Acquire rq_qos_mutex after freezing the queue to ensure proper
+ * locking order.
*/
memflags = blk_mq_freeze_queue(q);
+ mutex_lock(&q->rq_qos_mutex);
- if (rq_qos_id(q, rqos->id))
- goto ebusy;
- rqos->next = q->rq_qos;
- q->rq_qos = rqos;
- blk_queue_flag_set(QUEUE_FLAG_QOS_ENABLED, q);
+ if (rq_qos_id(q, rqos->id)) {
+ ret = -EBUSY;
+ } else {
+ rqos->next = q->rq_qos;
+ q->rq_qos = rqos;
+ blk_queue_flag_set(QUEUE_FLAG_QOS_ENABLED, q);
+ }
+ mutex_unlock(&q->rq_qos_mutex);
blk_mq_unfreeze_queue(q, memflags);
- return 0;
-ebusy:
- blk_mq_unfreeze_queue(q, memflags);
- return -EBUSY;
+ return ret;
}
void rq_qos_del(struct rq_qos *rqos)
@@ -359,9 +362,9 @@ void rq_qos_del(struct rq_qos *rqos)
struct rq_qos **cur;
unsigned int memflags;
- lockdep_assert_held(&q->rq_qos_mutex);
-
memflags = blk_mq_freeze_queue(q);
+ mutex_lock(&q->rq_qos_mutex);
+
for (cur = &q->rq_qos; *cur; cur = &(*cur)->next) {
if (*cur == rqos) {
*cur = rqos->next;
@@ -370,5 +373,7 @@ void rq_qos_del(struct rq_qos *rqos)
}
if (!q->rq_qos)
blk_queue_flag_clear(QUEUE_FLAG_QOS_ENABLED, q);
+
+ mutex_unlock(&q->rq_qos_mutex);
blk_mq_unfreeze_queue(q, memflags);
}
diff --git a/block/blk-wbt.c b/block/blk-wbt.c
index 6dba71e87387..dde03b9ea074 100644
--- a/block/blk-wbt.c
+++ b/block/blk-wbt.c
@@ -961,9 +961,7 @@ static int wbt_init(struct gendisk *disk, struct rq_wb *rwb)
/*
* Assign rwb and add the stats callback.
*/
- mutex_lock(&q->rq_qos_mutex);
ret = rq_qos_add(&rwb->rqos, disk, RQ_QOS_WBT, &wbt_rqos_ops);
- mutex_unlock(&q->rq_qos_mutex);
if (ret)
return ret;
--
2.51.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
` (6 preceding siblings ...)
2026-03-04 7:38 ` [PATCH v3 7/7] blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del Yu Kuai
@ 2026-03-23 8:11 ` Yu Kuai
7 siblings, 0 replies; 9+ messages in thread
From: Yu Kuai @ 2026-03-23 8:11 UTC (permalink / raw)
To: Tejun Heo, Josef Bacik, Jens Axboe
Cc: cgroups, linux-block, linux-kernel, Zheng Qixing, Ming Lei,
Nilay Shroff, yukuai
Friendly ping ...
Hope we can consider this for 7.1-rc1 merge window.
在 2026/3/4 15:38, Yu Kuai 写道:
> This series fixes several race conditions related to q->blkg_list iteration
> and improves the locking around blkcg policy activation/deactivation.
>
> Patch 1-2: Protect q->blkg_list iteration with blkcg_mutex in blkg_destroy_all()
> and bfq_end_wr_async() to prevent races with blkg_free_workfn().
>
> Patch 3-4: Fix use-after-free and memory leak issues in blkcg_activate_policy()
> by extending blkcg_mutex coverage and skipping dying blkgs.
>
> Patch 5: Refactor policy pd teardown into a helper function.
>
> Patch 6: Restructure blkcg_activate_policy() to allocate pds before freezing
> the queue, avoiding potential deadlocks from percpu allocation. Also fix
> locking order in blkcg_deactivate_policy() to be consistent with
> blkcg_activate_policy() (mutex -> freeze).
>
> Patch 7: Move rq_qos_mutex handling inside rq_qos_add()/rq_qos_del() to
> simplify the locking and eliminate potential deadlocks.
>
> Note: queue_lock is still used in many places to protect queue blkg.
> Future work is to convert it to blkcg_mutex entirely.
>
> Changes v2 -> v3:
> - Patch 2: Wrap mutex_lock/unlock with #ifdef CONFIG_BFQ_GROUP_IOSCHED to
> fix compile error when CONFIG_BLK_CGROUP is disabled.
> - Patch 6: Fix locking order in blkcg_deactivate_policy() to match
> blkcg_activate_policy() (mutex -> freeze instead of freeze -> mutex).
> - Patch 7: Remove stale lockdep_assert_held() in iolatency_set_limit().
>
> Changes v1 -> v2:
> - Link: https://lore.kernel.org/all/20260108014416.3656493-1-zhengqixing@huaweicloud.com/
>
> Yu Kuai (4):
> blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with
> blkcg_mutex
> bfq: protect q->blkg_list iteration in bfq_end_wr_async() with
> blkcg_mutex
> blk-cgroup: allocate pds before freezing queue in
> blkcg_activate_policy()
> blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del
>
> Zheng Qixing (3):
> blk-cgroup: fix race between policy activation and blkg destruction
> blk-cgroup: skip dying blkg in blkcg_activate_policy()
> blk-cgroup: factor policy pd teardown loop into helper
>
> block/bfq-cgroup.c | 3 +-
> block/bfq-iosched.c | 6 ++
> block/blk-cgroup.c | 205 ++++++++++++++----------------------------
> block/blk-cgroup.h | 2 -
> block/blk-iocost.c | 11 +--
> block/blk-iolatency.c | 5 --
> block/blk-rq-qos.c | 31 ++++---
> block/blk-wbt.c | 2 -
> 8 files changed, 97 insertions(+), 168 deletions(-)
>
--
Thansk,
Kuai
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-03-23 8:23 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-03-04 7:38 [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
2026-03-04 7:38 ` [PATCH v3 1/7] blk-cgroup: protect q->blkg_list iteration in blkg_destroy_all() with blkcg_mutex Yu Kuai
2026-03-04 7:38 ` [PATCH v3 2/7] bfq: protect q->blkg_list iteration in bfq_end_wr_async() " Yu Kuai
2026-03-04 7:38 ` [PATCH v3 3/7] blk-cgroup: fix race between policy activation and blkg destruction Yu Kuai
2026-03-04 7:38 ` [PATCH v3 4/7] blk-cgroup: skip dying blkg in blkcg_activate_policy() Yu Kuai
2026-03-04 7:38 ` [PATCH v3 5/7] blk-cgroup: factor policy pd teardown loop into helper Yu Kuai
2026-03-04 7:38 ` [PATCH v3 6/7] blk-cgroup: allocate pds before freezing queue in blkcg_activate_policy() Yu Kuai
2026-03-04 7:38 ` [PATCH v3 7/7] blk-rq-qos: move rq_qos_mutex acquisition inside rq_qos_add/del Yu Kuai
2026-03-23 8:11 ` [PATCH v3 0/7] blk-cgroup: fix races related to blkg_list iteration Yu Kuai
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®