On Tue, Sep 22, 2026 at 09:35:29AM +0200, Christian KKKnig wrote: > On 9/22/26 08:08, Zhu, Lingshan wrote: > > On 9/1/2026 4:50 PM, Zhu, Lingshan wrote: > > > >> On 8/31/2026 4:21 PM, Christian König wrote: > >> > >>> On 8/28/26 17:59, Zhu, Lingshan wrote: > >>>> On 8/28/2026 9:08 PM, Christian König wrote: > >>>> > >>>>> On 8/28/26 11:53, Zhu Lingshan wrote: > >>>>>> This commit introduces a new helper > >>>>>> amdgpu_lookup_queue_by_doorbell which helps > >>>>>> look up a user queue with the given doorbell id > >>>>>> in a xarray. > >>>>>> > >>>>>> This function takes a kref of the user space queue. > >>>> Hello Christian > >>>> > >>>> Thanks for your comments. > >>>> > >>>>> Well absolutely clear NAK to the whole approach. > >>>>> > >>>>> This is the nonsense Sunil and I have worked quite hard to remove and > >>>>> we certainly shouldn't repeat such mistakes. > >>>>> > >>>>> When the userq needs to be used from interrupt context we need to hold > >>>>> the xa_lock_irqsave() or otherwise we don't have any guarantee that the > >>>>> userq, userq_mgr or associated fpriv went out of scope. > >>>> Holding the spin lock by xa_lock_irqsave() can surely avoid racing with > >>>> the destruction process, however, it does not apply to all scenarios, > >>>> for example, you can not hold spin lock in mes_userq_reset_queue(), > >>>> because it calls either amdgpu_mes_reset_queue_mmio or > >>>> amdgpu_mes_reset_queue_mmio, both of them acquire the MES mutex through > >>>> amdgpu_mes_lock. > >>> Yeah which is exactly the reason why mes_userq_reset_queue() should *NOT* > >>> be called from non IOCTL context. > >> I think it is not about whether called from IOCTL, it is a common racing > >> we should fix, and holding a kref is a low haning fruit. > >> > >>>> Another thing, out of the topic is, holding xa_lock does not guarantee > >>>> fpriv/userq_mgr alive, for example, when drm_device->unplugged is true, > >>>> all amdgpu teardown paths in amdgpu_drm_release are skipped, > >>>> and the fpriv/userq_mgr is freed, no matter whether holding the xa spin > >>>> lock. > >>> That would clearly be a massive bug. Those objects still need to be > >>> cleaned up independent of device hot plug. > >> I agree, when unplugged == true, means can not access any HW registers, so > >> this bug deserve another series to fix. > >> > >>>> So IMHO since we have userq->kref, lets use it to maintain the lifecycle > >>>> of the queues. > >>>> > >>>>> Grabbing references from this side would obviously result in circle > >>>>> dependencies. > >>>> I am not sure, we should use the lock/unlock and kref_put/get in pairs > >>>> in sequence, can you name some circle dependencies or AB-BA lockings as > >>>> examples? > >>> That is not AB-BA locking, but circle dependencies. E.g. A reference B, B > >>> referencing C, C referencing A again. > >> A kref is an atomic counter, not a lock, that means we can try hold more > >> than one kref in a thread, and a kref not > >> depend on another. As far as I can see, current amdgpu driver does not > >> have such problems, do you see any occurrences? > > > > Hello Christian, > > > > I have not heard back from you for three weeks, I wonder whether you have > > any further comments for this series, > > or shall I make any improvements? > > Well just completely drop that series. As I said this approach is a > fundamental no-go from my side. > > As far as I can see we have solved the problems at hand and the rules how to > handle the user queues should be pretty clear by now. >
Assuming every kref_get() has a matching kref_put(), and the queue object itself does not participate in a circular ownership chain, I don't immediately see how taking a temporary queue reference would cause correctness issues. Could you please point us to the historical issue or the corresponding commits that led to removing this approach previously? Thanks, Ray > Regards, > Christian. > > > > > Thanks > > Lingshan > > > >> Thanks > >> Lingshan > >> > >>> Regards, > >>> Christian. > >>> > >>>> Thanks > >>>> Lingshan > >>>> > >>>>> Regards, > >>>>> Christian. > >>>>> > >>>>>> Signed-off-by: Zhu Lingshan <[email protected]> > >>>>>> --- > >>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++ > >>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h | 2 ++ > >>>>>> 2 files changed, 32 insertions(+) > >>>>>> > >>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > >>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > >>>>>> index 0a816b3c5ff9..e0639f844a8e 100644 > >>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > >>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c > >>>>>> @@ -609,6 +609,36 @@ struct amdgpu_usermode_queue > >>>>>> *amdgpu_userq_get(struct amdgpu_userq_mgr *uq_mgr, > >>>>>> return queue; > >>>>>> } > >>>>>> > >>>>>> +/** > >>>>>> + * amdgpu_lookup_queue_by_doorbell - look up a user queue by doorbell > >>>>>> + * @xa: user queue XArray indexed by doorbell > >>>>>> + * @doorbell: doorbell index > >>>>>> + * > >>>>>> + * Return: A queue with the doorbell indexed, or NULL if no such a > >>>>>> queue found. > >>>>>> + * > >>>>>> + * This function increases kref of the queue, the caller > >>>>>> + * must release the reference with amdgpu_userq_put(). > >>>>>> + */ > >>>>>> +struct amdgpu_usermode_queue * > >>>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell) > >>>>>> +{ > >>>>>> + struct amdgpu_usermode_queue *queue; > >>>>>> + unsigned long flags; > >>>>>> + > >>>>>> + xa_lock_irqsave(xa, flags); > >>>>>> + queue = xa_load(xa, doorbell); > >>>>>> + if (!queue) > >>>>>> + goto out_unlock; > >>>>>> + > >>>>>> + if (!kref_get_unless_zero(&queue->refcount)) > >>>>>> + queue = NULL; > >>>>>> + > >>>>>> +out_unlock: > >>>>>> + xa_unlock_irqrestore(xa, flags); > >>>>>> + > >>>>>> + return queue; > >>>>>> +} > >>>>>> + > >>>>>> void amdgpu_userq_put(struct amdgpu_usermode_queue *queue) > >>>>>> { > >>>>>> if (queue) > >>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h > >>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h > >>>>>> index 6412a7f7b6ef..8fc73862f64e 100644 > >>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h > >>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h > >>>>>> @@ -151,6 +151,8 @@ struct amdgpu_db_info { > >>>>>> }; > >>>>>> > >>>>>> struct amdgpu_usermode_queue *amdgpu_userq_get(struct > >>>>>> amdgpu_userq_mgr *uq_mgr, u32 qid); > >>>>>> +struct amdgpu_usermode_queue * > >>>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell); > >>>>>> void amdgpu_userq_put(struct amdgpu_usermode_queue *queue); > >>>>>> > >>>>>> int amdgpu_userq_ioctl(struct drm_device *dev, void *data, struct > >>>>>> drm_file *filp); >
