amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
attachment.  GTT is only a fallback placement when VRAM is preferred, so
ttm_bo_validate() migrates the buffer from GTT into VRAM.  While the
exporting device is runtime suspended its SDMA rings are down and the
move fails:

  amdgpu: Move buffer fallback to memcpy unavailable

An importer on a second GPU reaches this holding no runtime PM
reference on the exporter, e.g. a compositor on the APU submitting a
frame that references a buffer exported by an idle dGPU:

  amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
    -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
      -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move

Pinning a dma-buf into VRAM has the same requirement, which
commit 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference
when we attach a buffer" v2") called out as the one case that would
need the reference back.

Take it in attach and drop it in detach.  pm_runtime_get_if_active()
never resumes the device, so it cannot deadlock against the reservation
taken during resume, which is why the old pm_runtime_get_sync() had to
go.  If the device is not active, clear peer2peer instead: the buffer
then stays in GTT, which remains accessible while the GPU is powered
down.

Fixes: 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference when we 
attach a buffer" v2")
Cc: [email protected]
Suggested-by: Christian König <[email protected]>
Signed-off-by: Mike Lothian <[email protected]>
Assisted-by: Claude:Opus-5 [Claude Code]
---

v3: take the reference in attach and drop it in detach, clearing
    peer2peer when the device is not active, as suggested by Christian.
    v2 only covered amdgpu_dma_buf_map() and left VRAM pinning exposed.
v2: use pm_runtime_get_if_active() instead of testing
    adev->mman.buffer_funcs_enabled, which was racy against a
    concurrent suspend.  Reported by Sashiko AI review.

Reproduced on a HawkPoint APU [1002:1900] driving the display with a
Navi 48 [Radeon AI PRO R9700] [1002:7551] on oculink for render
offload.  Without the patch kwin_wayland hits the call chain above
within a minute of the dGPU autosuspending and the desktop stops
repainting until it resumes.

Tested with v3: Chromium rendering on the dGPU and composited by kwin
6.7.5 for five minutes, then closed.  The dGPU stayed active while the
window was on screen and suspended six seconds after Chromium exited,
with no fallback errors or runtime PM usage count underflows.  After
an hour of yuzu render offload the dGPU also suspended once yuzu
exited.

 drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c | 39 ++++++++++++++++++++-
 1 file changed, 38 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c 
b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
index b33c300e26e2..fae695c3e531 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
@@ -43,6 +43,7 @@
 #include <linux/dma-buf.h>
 #include <linux/dma-fence-array.h>
 #include <linux/pci-p2pdma.h>
+#include <linux/pm_runtime.h>
 
 static const struct dma_buf_attach_ops amdgpu_dma_buf_attach_ops;
 
@@ -100,15 +101,50 @@ static int amdgpu_dma_buf_attach(struct dma_buf *dmabuf,
            pci_p2pdma_distance(adev->pdev, attach->dev, false) < 0)
                attach->peer2peer = false;
 
+       /*
+        * P2P access needs the exporter awake for the lifetime of the
+        * attachment.  pm_runtime_get_if_active() never resumes the device,
+        * so it cannot deadlock against the reservation taken during resume.
+        * A negative return means runtime PM is disabled and the device
+        * cannot suspend, in which case the put in detach is a no-op.
+        */
+       if (attach->peer2peer &&
+           !pm_runtime_get_if_active(adev_to_drm(adev)->dev))
+               attach->peer2peer = false;
+
        r = dma_resv_lock(bo->tbo.base.resv, NULL);
        if (r)
-               return r;
+               goto err_pm_put;
 
        amdgpu_vm_bo_update_shared(bo);
 
        dma_resv_unlock(bo->tbo.base.resv);
 
        return 0;
+
+err_pm_put:
+       if (attach->peer2peer)
+               pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
+       return r;
+}
+
+/**
+ * amdgpu_dma_buf_detach - &dma_buf_ops.detach implementation
+ *
+ * @dmabuf: DMA-buf where we remove the attachment from
+ * @attach: the attachment to remove
+ *
+ * Drop the runtime PM reference taken in amdgpu_dma_buf_attach().
+ */
+static void amdgpu_dma_buf_detach(struct dma_buf *dmabuf,
+                                 struct dma_buf_attachment *attach)
+{
+       struct drm_gem_object *obj = dmabuf->priv;
+       struct amdgpu_bo *bo = gem_to_amdgpu_bo(obj);
+       struct amdgpu_device *adev = amdgpu_ttm_adev(bo->tbo.bdev);
+
+       if (attach->peer2peer)
+               pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
 }
 
 /**
@@ -350,6 +386,7 @@ static void amdgpu_dma_buf_vunmap(struct dma_buf *dma_buf, 
struct iosys_map *map
 
 const struct dma_buf_ops amdgpu_dmabuf_ops = {
        .attach = amdgpu_dma_buf_attach,
+       .detach = amdgpu_dma_buf_detach,
        .pin = amdgpu_dma_buf_pin,
        .unpin = amdgpu_dma_buf_unpin,
        .map_dma_buf = amdgpu_dma_buf_map,
-- 
2.55.0

Reply via email to