On Fri Sep 25, 2026 at 3:33 PM CEST, Thomas Hellström wrote: > Driver and shared DRM helper code is increasingly relying on bare > drm_device references (drm_dev_get()/drm_dev_put()) to keep a device's > software state around, without also pairing that with a reference on > the owning kernel module. Xe itself does this in several places, and > so does drm_gpuvm for the lifetime of a GPU VM. None of these > references currently prevent the owning module from being unloaded > while they, or the teardown work they can still trigger, are > outstanding, meaning driver code can end up executing after its own > module's text has already been freed.
Since you mention DRM GPUVM in a couple of places, how can this ever happen? It wouldn't make sense to keep a VM alive beyond driver unbind. I.e. it can't make its drm_device reference count reach module unload in the first place. Besides that, can you please remind me whether there are any other reasons than the release() callback why a DRM device must not outlive module unload? I don't think the correct solution is to constrain module unload. The release() callback shouldn't really do anything other than free the memory of the drm_device allocation. All other resources a driver may have should be released on driver unbind. There may be shared resources, such as e.g. a common workqueue, but those are module level things that have nothing to do with the DRM device. I am aware that a few drivers implement release(), but TBH it looks pretty broken. qxl_drm_release() is a great example, and it already calls it out itself. > This series closes that gap for xe: > > - Patch 1 adds core DRM infrastructure allowing a driver to wait for > its outstanding device-release callbacks to finish before > proceeding with module unload. Please see above; I also wonder why Xe cares in the first place. Xe doesn't implement release(), no? > - Patch 2 makes xe use this infrastructure to hold up module unload > until every xe_device instance has actually been released, rather > than only until the module's own refcount happens to reach zero, > with a diagnostic if this ends up taking an unexpectedly long time. > > - Patch 3 fixes a related, previously unprotected case where the > teardown of a GPU VM or its address space mappings can be deferred > to run at an arbitrary later time, including after module unload has > already completed. Huh? GPUVM tracks the GPU's VA space mappings, but after driver unbind there's no access to the GPU to manage anything anymore. How can this even work? Thanks, Danilo
