Applied.  Thanks!

On Mon, Sep 14, 2026 at 6:51 AM Christian König
<[email protected]> wrote:
>
> On 9/11/26 20:38, Mike Lothian wrote:
> > amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
> > attachment.  GTT is only a fallback placement when VRAM is preferred, so
> > ttm_bo_validate() migrates the buffer from GTT into VRAM.  While the
> > exporting device is runtime suspended its SDMA rings are down and the
> > move fails:
> >
> >   amdgpu: Move buffer fallback to memcpy unavailable
> >
> > An importer on a second GPU reaches this holding no runtime PM
> > reference on the exporter, e.g. a compositor on the APU submitting a
> > frame that references a buffer exported by an idle dGPU:
> >
> >   amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
> >     -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
> >       -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move
> >
> > Pinning a dma-buf into VRAM has the same requirement, which
> > commit 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference
> > when we attach a buffer" v2") called out as the one case that would
> > need the reference back.
> >
> > Take it in attach and drop it in detach.  pm_runtime_get_if_active()
> > never resumes the device, so it cannot deadlock against the reservation
> > taken during resume, which is why the old pm_runtime_get_sync() had to
> > go.  If the device is not active, clear peer2peer instead: the buffer
> > then stays in GTT, which remains accessible while the GPU is powered
> > down.
> >
> > Fixes: 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference when we 
> > attach a buffer" v2")
> > Cc: [email protected]
> > Suggested-by: Christian König <[email protected]>
> > Signed-off-by: Mike Lothian <[email protected]>
> > Assisted-by: Claude:Opus-5 [Claude Code]
>
> Reviewed-by: Christian König <[email protected]>
>
> > ---
> >
> > v3: take the reference in attach and drop it in detach, clearing
> >     peer2peer when the device is not active, as suggested by Christian.
> >     v2 only covered amdgpu_dma_buf_map() and left VRAM pinning exposed.
> > v2: use pm_runtime_get_if_active() instead of testing
> >     adev->mman.buffer_funcs_enabled, which was racy against a
> >     concurrent suspend.  Reported by Sashiko AI review.
> >
> > Reproduced on a HawkPoint APU [1002:1900] driving the display with a
> > Navi 48 [Radeon AI PRO R9700] [1002:7551] on oculink for render
> > offload.  Without the patch kwin_wayland hits the call chain above
> > within a minute of the dGPU autosuspending and the desktop stops
> > repainting until it resumes.
> >
> > Tested with v3: Chromium rendering on the dGPU and composited by kwin
> > 6.7.5 for five minutes, then closed.  The dGPU stayed active while the
> > window was on screen and suspended six seconds after Chromium exited,
> > with no fallback errors or runtime PM usage count underflows.  After
> > an hour of yuzu render offload the dGPU also suspended once yuzu
> > exited.
> >
> >  drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c | 39 ++++++++++++++++++++-
> >  1 file changed, 38 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c 
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
> > index b33c300e26e2..fae695c3e531 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
> > @@ -43,6 +43,7 @@
> >  #include <linux/dma-buf.h>
> >  #include <linux/dma-fence-array.h>
> >  #include <linux/pci-p2pdma.h>
> > +#include <linux/pm_runtime.h>
> >
> >  static const struct dma_buf_attach_ops amdgpu_dma_buf_attach_ops;
> >
> > @@ -100,15 +101,50 @@ static int amdgpu_dma_buf_attach(struct dma_buf 
> > *dmabuf,
> >           pci_p2pdma_distance(adev->pdev, attach->dev, false) < 0)
> >               attach->peer2peer = false;
> >
> > +     /*
> > +      * P2P access needs the exporter awake for the lifetime of the
> > +      * attachment.  pm_runtime_get_if_active() never resumes the device,
> > +      * so it cannot deadlock against the reservation taken during resume.
> > +      * A negative return means runtime PM is disabled and the device
> > +      * cannot suspend, in which case the put in detach is a no-op.
> > +      */
> > +     if (attach->peer2peer &&
> > +         !pm_runtime_get_if_active(adev_to_drm(adev)->dev))
> > +             attach->peer2peer = false;
> > +
> >       r = dma_resv_lock(bo->tbo.base.resv, NULL);
> >       if (r)
> > -             return r;
> > +             goto err_pm_put;
> >
> >       amdgpu_vm_bo_update_shared(bo);
> >
> >       dma_resv_unlock(bo->tbo.base.resv);
> >
> >       return 0;
> > +
> > +err_pm_put:
> > +     if (attach->peer2peer)
> > +             pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
> > +     return r;
> > +}
> > +
> > +/**
> > + * amdgpu_dma_buf_detach - &dma_buf_ops.detach implementation
> > + *
> > + * @dmabuf: DMA-buf where we remove the attachment from
> > + * @attach: the attachment to remove
> > + *
> > + * Drop the runtime PM reference taken in amdgpu_dma_buf_attach().
> > + */
> > +static void amdgpu_dma_buf_detach(struct dma_buf *dmabuf,
> > +                               struct dma_buf_attachment *attach)
> > +{
> > +     struct drm_gem_object *obj = dmabuf->priv;
> > +     struct amdgpu_bo *bo = gem_to_amdgpu_bo(obj);
> > +     struct amdgpu_device *adev = amdgpu_ttm_adev(bo->tbo.bdev);
> > +
> > +     if (attach->peer2peer)
> > +             pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
> >  }
> >
> >  /**
> > @@ -350,6 +386,7 @@ static void amdgpu_dma_buf_vunmap(struct dma_buf 
> > *dma_buf, struct iosys_map *map
> >
> >  const struct dma_buf_ops amdgpu_dmabuf_ops = {
> >       .attach = amdgpu_dma_buf_attach,
> > +     .detach = amdgpu_dma_buf_detach,
> >       .pin = amdgpu_dma_buf_pin,
> >       .unpin = amdgpu_dma_buf_unpin,
> >       .map_dma_buf = amdgpu_dma_buf_map,
>

Reply via email to