Hello, Tao.

On Tue, 29 Sep 2026 21:42:51 +0800, Tao Cui wrote:
> That leaves one race: .unreg may observe a non-NULL ops->q before
> ioc_rqos_exit() clears it, then block on rq_qos_mutex while the
> ejection drops the last bdev reference.  This looks analogous to
> hid_bpf's .unreg vs. destroy_device synchronization.
>
> Does that seem acceptable here too, or would you rather have .unreg
> own the final reference unconditionally?

No, there can't be a crash window like that. The underlying problem is
that .unreg sleeps on a mutex inside a queue it holds no reference on. The
model should still be ejected when the device goes away, but exactly when
doesn't matter much as long as it happens in a reasonable amount of time.

request_queues are RCU-freed, so maybe .unreg can rcu_dereference() ops->q
and try to get a queue reference under rcu_read_lock()? blk_get_queue()
fails on a dying queue, so this would need a tryget wrapper around
q->refs. Holding its own reference, .unreg can then take rq_qos_mutex and
test ops->q for NULL to tell whether removal already ejected the model.

Also, when .unreg detaches, the device is still live and the BPF code may
be running. Freezing and quiescing the queue covers the IO paths but not
iocg_init() and iocg_free() called from ioc_pd_init() and ioc_pd_free(),
so the iocg_free() walk and clearing the model need to be synchronized
against those too.

Thanks.

--
tejun

Reply via email to