From: Tao Cui <cui.tao@linux.dev>
To: Tejun Heo <tj@kernel.org>
Cc: cui.tao@linux.dev, josef@toxicopanda.com, axboe@kernel.dk,
cgroups@vger.kernel.org, linux-block@vger.kernel.org,
linux-kernel@vger.kernel.org, bpf@vger.kernel.org,
andrii@kernel.org, eddyz87@gmail.com, ast@kernel.org,
daniel@iogearbox.net, linux-kselftest@vger.kernel.org,
Tao Cui <cuitao@kylinos.cn>,
ameryhung@gmail.com, alexei.starovoitov@gmail.com
Subject: Re: [RFC PATCH v7 1/4] blk-iocost: add BPF struct_ops cost model support
Date: Wed, 30 Sep 2026 12:14:57 +0800 [thread overview]
Message-ID: <4f0e79fc-9413-4017-93a2-414b4370781f@linux.dev> (raw)
In-Reply-To: <aeb50cff4d3014a43ded85247a5786da@kernel.org>
Hello, Tejun.
在 2026/9/30 00:27, Tejun Heo 写道:
> Hello, Tao.
>
> On Tue, 29 Sep 2026 21:42:51 +0800, Tao Cui wrote:
>> That leaves one race: .unreg may observe a non-NULL ops->q before
>> ioc_rqos_exit() clears it, then block on rq_qos_mutex while the
>> ejection drops the last bdev reference. This looks analogous to
>> hid_bpf's .unreg vs. destroy_device synchronization.
>>
>> Does that seem acceptable here too, or would you rather have .unreg
>> own the final reference unconditionally?
>
> No, there can't be a crash window like that. The underlying problem is
> that .unreg sleeps on a mutex inside a queue it holds no reference on. The
> model should still be ejected when the device goes away, but exactly when
> doesn't matter much as long as it happens in a reasonable amount of time.
>
> request_queues are RCU-freed, so maybe .unreg can rcu_dereference() ops->q
> and try to get a queue reference under rcu_read_lock()? blk_get_queue()
> fails on a dying queue, so this would need a tryget wrapper around
> q->refs. Holding its own reference, .unreg can then take rq_qos_mutex and
> test ops->q for NULL to tell whether removal already ejected the model.
>
.unreg now takes a queue reference before entering the queue: it
rcu_dereferences ops->q under rcu_read_lock() and uses a queue
reference tryget helper to keep the queue valid across the mutex wait. With
the reference held, .unreg freezes and quiesces the queue, takes
rq_qos_mutex and re-checks ops->q for NULL, so it can handle both the
normal detach path and the case where removal has already ejected
the model.
> Also, when .unreg detaches, the device is still live and the BPF code may
> be running. Freezing and quiescing the queue covers the IO paths but not
> iocg_init() and iocg_free() called from ioc_pd_init() and ioc_pd_free(),
> so the iocg_free() walk and clearing the model need to be synchronized
> against those too.
>
Clearing the model and the iocg_free() walk are now synchronized
under q->blkcg_mutex. This is the same mutex used when
blkg_create() and blkg_free_workfn() invoke pd_init_fn() and
pd_free_fn(), so the detach path is synchronized with the normal pd
callback paths.
While reworking the attach path, I also fixed two related issues I
noticed: the attach path no longer uses the ioc after releasing
rq_qos_mutex, and blkdev_get_no_open() failures are handled
correctly for unknown dev_t values.
v8 will follow shortly.
Thanks.
Tao
> Thanks.
>
> --
> tejun
next prev parent reply other threads:[~2026-09-30 4:16 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 5:45 [RFC PATCH v7 0/4] blk-iocost: BPF struct_ops cost model Tao Cui
2026-09-24 5:45 ` [RFC PATCH v7 1/4] blk-iocost: add BPF struct_ops cost model support Tao Cui
2026-09-24 6:30 ` bot+bpf-ci
2026-09-29 0:42 ` Tejun Heo
2026-09-29 13:42 ` Tao Cui
2026-09-29 16:27 ` Tejun Heo
2026-09-30 4:14 ` Tao Cui [this message]
2026-09-24 5:45 ` [RFC PATCH v7 2/4] selftests/bpf: add iocost cost model test Tao Cui
2026-09-24 6:30 ` bot+bpf-ci
2026-09-24 5:45 ` [RFC PATCH v7 3/4] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary Tao Cui
2026-09-24 6:17 ` bot+bpf-ci
2026-09-29 0:42 ` Tejun Heo
2026-09-29 13:46 ` Tao Cui
2026-09-24 5:45 ` [RFC PATCH v7 4/4] docs: cgroup-v2: document the iocost BPF cost model attachment Tao Cui
2026-09-29 0:42 ` [RFC PATCH v7 0/4] blk-iocost: BPF struct_ops cost model Tejun Heo
2026-09-29 13:39 ` Tao Cui
2026-09-29 16:27 ` Tejun Heo
2026-09-30 4:09 ` Tao Cui
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4f0e79fc-9413-4017-93a2-414b4370781f@linux.dev \
--to=cui.tao@linux.dev \
--cc=alexei.starovoitov@gmail.com \
--cc=ameryhung@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=axboe@kernel.dk \
--cc=bpf@vger.kernel.org \
--cc=cgroups@vger.kernel.org \
--cc=cuitao@kylinos.cn \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=josef@toxicopanda.com \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®