On 07/09/2026 12:53, Christian König wrote:
On 9/7/26 11:54, Tvrtko Ursulin wrote:

On 03/09/2026 12:36, Christian König wrote:
i915_gem_busy_ioctl uses dma_resv_for_each_fence_unlocked() to iterate
over the fences in an GEM object without holding a reference but only
the RCU read side lock.

What can happen here is that the GEM object is destroyed concurrently
while i915_gem_busy_ioctl is still running. This won't free the GEM
objects memory, but still drops all the dma_fence references.

Now when dma_resv_for_each_fence_unlocked() sees a destroyed dma_fence it
assumes that a new fence list was installed and re-starts the loop.

But in the case of a destroyed GEM object a new fence list is never
installed, only the old one freed and therefore the iteration never
finishes resulting in an endless loop.

Only i915_busy can get into this failure mode? None of the other users of the 
iterator?

Yes, at least as far as I can see.

The problem is completely i915 specific because it is the only driver (I could 
find) which protects GEM objects by RCU.

Also, the reference counting series makes the fix irrelevant?

No, that series just helped uncover the issue.

Sashiko-bot correctly complained that i915 is dropping the new dma-resv 
reference to early resulting in potential use after free. And I was thinking 
wait a second when the dma_resv_fini() is called to early in the existing code 
then the dma_fence references are dropped to early as well... so that is an 
pre-existing bug.

Before the commit mentioned in the fixes tag the i915_gem_busy_ioctl() could 
just return nonsense, but after that change it could result in an endless loop 
and that is problematic.

Where is this sashiko report, associated with which patch I mean?

Is the dma_fence_get_rcu() inside dma_resv_iter_walk_unlocked() what triggers the endless restarts? It's been some time since I looked at the dma-resv walks.. but fences on the list have reference held so that can trigger either via dma_resv_fini() or dma_resv_replace_fences(), right? If second is true then how does i915 having the dma-resv containing object RCU freed cause the problem?

Regards,

Tvrtko

The solution is to drop the fence references only after the RCU grace
period.

The fixes tag is not necessary the patch introducing the problem, but the
one making it so worse that we need to address it.

This problem was pointed out by Sashiko-bot.

Signed-off-by: Christian König <[email protected]>
Fixes: 912ff2ebd695 ("drm/i915: use the new iterator in i915_gem_busy_ioctl v2")
CC: [email protected]
---
   drivers/gpu/drm/i915/gem/i915_gem_object.c | 2 +-
   1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/i915/gem/i915_gem_object.c 
b/drivers/gpu/drm/i915/gem/i915_gem_object.c
index 5172d3982654..9e01f8b2079a 100644
--- a/drivers/gpu/drm/i915/gem/i915_gem_object.c
+++ b/drivers/gpu/drm/i915/gem/i915_gem_object.c
@@ -89,6 +89,7 @@ struct drm_i915_gem_object *i915_gem_object_alloc(void)
     void i915_gem_object_free(struct drm_i915_gem_object *obj)
   {
+    dma_resv_fini(&obj->base._resv);
       return kmem_cache_free(slab_objects, obj);
   }
   @@ -144,7 +145,6 @@ void __i915_gem_object_fini(struct drm_i915_gem_object 
*obj)
   {
       mutex_destroy(&obj->mm.get_page.lock);
       mutex_destroy(&obj->mm.get_dma_page.lock);
-    dma_resv_fini(&obj->base._resv);
   }
     /**



Reply via email to