On Mon Sep 28, 2026 at 4:20 AM CEST, Matthew Brost wrote:
> To me, this looks like a lifetime issue that should be fixed in AMDGPU.
> Either that, or we need to rework DRM to have proper lifetime management,
> or, of course, just deprecate drm_sched.

Agreed. The lifetime and ownership model in drm_sched is not fundamentally
wrong, but the implementation is not consistently keeping it up and drivers also
don't always honor it.

> In other words, we need refcounting so that drm_sched_fini() cannot be
> called while jobs are still in flight, nor can jobs be submitted before
> drm_sched_init(). Alternatively, drm_dep could serve as a replacement.

I'm not a huge fan of refcounting for those kind of things because it
fundamentally incentivises the wrong lifetime model in the context of the driver
model. The driver model requires a bounded lifetime scope, but refcounting
incentivises an unbounded lifetime model, which leads to other problems.

So, especially for the sake of deferring drm_sched_fini() from running jobs,
drivers still have to make sure that all jobs are torn down and drm_sched_fini()
is called *before* driver unbind completes. IOW, drivers should tear down the
hardware and hence all jobs latest in remove() and then call drm_sched_fini()
subsequently.

Refcounting does not provide a lot of value in this regard, because we have to
somehow guarantee that the hardware and all jobs are torn down at this specific
boundary anyway.

This is also my biggest concern about drm_dep, it seems to be designed with
exactly the idea of an unbounded lifetime model, which is not the correct design
for anything that represents a device resource (e.g. a GPU job).

Now, to be fair, refcounting is really the only mechanism that we have in C to
manage lifetimes, everything else is more or less just a convention. That said,
I'm not all against refcounting in general, but we have to be careful about how
it influences the design in terms of a bounded and unbounded lifetime model.

Reply via email to