On 8/31/2026 4:21 PM, Christian König wrote:

> On 8/28/26 17:59, Zhu, Lingshan wrote:
>> On 8/28/2026 9:08 PM, Christian König wrote:
>>
>>> On 8/28/26 11:53, Zhu Lingshan wrote:
>>>> This commit introduces a new helper
>>>> amdgpu_lookup_queue_by_doorbell which helps
>>>> look up a user queue with the given doorbell id
>>>> in a xarray.
>>>>
>>>> This function takes a kref of the user space queue.
>> Hello Christian
>>
>> Thanks for your comments.
>>
>>> Well absolutely clear NAK to the whole approach.
>>>
>>> This is the nonsense Sunil and I have worked quite hard to remove and we 
>>> certainly shouldn't repeat such mistakes.
>>>
>>> When the userq needs to be used from interrupt context we need to hold the 
>>> xa_lock_irqsave() or otherwise we don't have any guarantee that the userq, 
>>> userq_mgr or associated fpriv went out of scope.
>> Holding the spin lock by xa_lock_irqsave() can surely avoid racing with the 
>> destruction process, however, it does not apply to all scenarios, for 
>> example, you can not hold spin lock in mes_userq_reset_queue(),
>> because it calls either amdgpu_mes_reset_queue_mmio or 
>> amdgpu_mes_reset_queue_mmio, both of them acquire the MES mutex through 
>> amdgpu_mes_lock.
> Yeah which is exactly the reason why mes_userq_reset_queue() should *NOT* be 
> called from non IOCTL context.

I think it is not about whether called from IOCTL, it is a common racing we 
should fix, and holding a kref is a low haning fruit.

>
>> Another thing, out of the topic is, holding xa_lock does not guarantee 
>> fpriv/userq_mgr alive, for example, when drm_device->unplugged is true, all 
>> amdgpu teardown paths in amdgpu_drm_release are skipped,
>> and the fpriv/userq_mgr is freed, no matter whether holding the xa spin lock.
> That would clearly be a massive bug. Those objects still need to be cleaned 
> up independent of device hot plug.

I agree, when unplugged == true, means can not access any HW registers, so this 
bug deserve another series to fix.

>
>> So IMHO since we have userq->kref, lets use it to maintain the lifecycle of 
>> the queues. 
>>
>>> Grabbing references from this side would obviously result in circle 
>>> dependencies.
>> I am not sure, we should use the lock/unlock and kref_put/get in pairs in 
>> sequence, can you name some circle dependencies or AB-BA lockings as 
>> examples?
> That is not AB-BA locking, but circle dependencies. E.g. A reference B, B 
> referencing C, C referencing A again.

A kref is an atomic counter, not a lock, that means we can try hold more than 
one kref in a thread, and a kref not
depend on another. As far as I can see, current amdgpu driver does not have 
such problems, do you see any occurrences?

Thanks
Lingshan

>
> Regards,
> Christian.
>
>> Thanks
>> Lingshan
>>
>>> Regards,
>>> Christian.
>>>
>>>> Signed-off-by: Zhu Lingshan <[email protected]>
>>>> ---
>>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++
>>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h |  2 ++
>>>>  2 files changed, 32 insertions(+)
>>>>
>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c 
>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>>> index 0a816b3c5ff9..e0639f844a8e 100644
>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
>>>> @@ -609,6 +609,36 @@ struct amdgpu_usermode_queue *amdgpu_userq_get(struct 
>>>> amdgpu_userq_mgr *uq_mgr,
>>>>    return queue;
>>>>  }
>>>>  
>>>> +/**
>>>> + * amdgpu_lookup_queue_by_doorbell - look up a user queue by doorbell
>>>> + * @xa: user queue XArray indexed by doorbell
>>>> + * @doorbell: doorbell index
>>>> + *
>>>> + * Return: A queue with the doorbell indexed, or NULL if no such a queue 
>>>> found.
>>>> + *
>>>> + * This function increases kref of the queue, the caller
>>>> + * must release the reference with amdgpu_userq_put().
>>>> + */
>>>> +struct amdgpu_usermode_queue *
>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell)
>>>> +{
>>>> +  struct amdgpu_usermode_queue *queue;
>>>> +  unsigned long flags;
>>>> +
>>>> +  xa_lock_irqsave(xa, flags);
>>>> +  queue = xa_load(xa, doorbell);
>>>> +  if (!queue)
>>>> +          goto out_unlock;
>>>> +
>>>> +  if (!kref_get_unless_zero(&queue->refcount))
>>>> +          queue = NULL;
>>>> +
>>>> +out_unlock:
>>>> +  xa_unlock_irqrestore(xa, flags);
>>>> +
>>>> +  return queue;
>>>> +}
>>>> +
>>>>  void amdgpu_userq_put(struct amdgpu_usermode_queue *queue)
>>>>  {
>>>>    if (queue)
>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h 
>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>>> index 6412a7f7b6ef..8fc73862f64e 100644
>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
>>>> @@ -151,6 +151,8 @@ struct amdgpu_db_info {
>>>>  };
>>>>  
>>>>  struct amdgpu_usermode_queue *amdgpu_userq_get(struct amdgpu_userq_mgr 
>>>> *uq_mgr, u32 qid);
>>>> +struct amdgpu_usermode_queue *
>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell);
>>>>  void amdgpu_userq_put(struct amdgpu_usermode_queue *queue);
>>>>  
>>>>  int amdgpu_userq_ioctl(struct drm_device *dev, void *data, struct 
>>>> drm_file *filp);

Reply via email to