Applied. Thanks!
On Mon, Sep 14, 2026 at 6:51 AM Christian König <[email protected]> wrote: > > On 9/11/26 20:38, Mike Lothian wrote: > > amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer > > attachment. GTT is only a fallback placement when VRAM is preferred, so > > ttm_bo_validate() migrates the buffer from GTT into VRAM. While the > > exporting device is runtime suspended its SDMA rings are down and the > > move fails: > > > > amdgpu: Move buffer fallback to memcpy unavailable > > > > An importer on a second GPU reaches this holding no runtime PM > > reference on the exporter, e.g. a compositor on the APU submitting a > > frame that references a buffer exported by an idle dGPU: > > > > amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate > > -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment > > -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move > > > > Pinning a dma-buf into VRAM has the same requirement, which > > commit 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference > > when we attach a buffer" v2") called out as the one case that would > > need the reference back. > > > > Take it in attach and drop it in detach. pm_runtime_get_if_active() > > never resumes the device, so it cannot deadlock against the reservation > > taken during resume, which is why the old pm_runtime_get_sync() had to > > go. If the device is not active, clear peer2peer instead: the buffer > > then stays in GTT, which remains accessible while the GPU is powered > > down. > > > > Fixes: 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference when we > > attach a buffer" v2") > > Cc: [email protected] > > Suggested-by: Christian König <[email protected]> > > Signed-off-by: Mike Lothian <[email protected]> > > Assisted-by: Claude:Opus-5 [Claude Code] > > Reviewed-by: Christian König <[email protected]> > > > --- > > > > v3: take the reference in attach and drop it in detach, clearing > > peer2peer when the device is not active, as suggested by Christian. > > v2 only covered amdgpu_dma_buf_map() and left VRAM pinning exposed. > > v2: use pm_runtime_get_if_active() instead of testing > > adev->mman.buffer_funcs_enabled, which was racy against a > > concurrent suspend. Reported by Sashiko AI review. > > > > Reproduced on a HawkPoint APU [1002:1900] driving the display with a > > Navi 48 [Radeon AI PRO R9700] [1002:7551] on oculink for render > > offload. Without the patch kwin_wayland hits the call chain above > > within a minute of the dGPU autosuspending and the desktop stops > > repainting until it resumes. > > > > Tested with v3: Chromium rendering on the dGPU and composited by kwin > > 6.7.5 for five minutes, then closed. The dGPU stayed active while the > > window was on screen and suspended six seconds after Chromium exited, > > with no fallback errors or runtime PM usage count underflows. After > > an hour of yuzu render offload the dGPU also suspended once yuzu > > exited. > > > > drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c | 39 ++++++++++++++++++++- > > 1 file changed, 38 insertions(+), 1 deletion(-) > > > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c > > index b33c300e26e2..fae695c3e531 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c > > @@ -43,6 +43,7 @@ > > #include <linux/dma-buf.h> > > #include <linux/dma-fence-array.h> > > #include <linux/pci-p2pdma.h> > > +#include <linux/pm_runtime.h> > > > > static const struct dma_buf_attach_ops amdgpu_dma_buf_attach_ops; > > > > @@ -100,15 +101,50 @@ static int amdgpu_dma_buf_attach(struct dma_buf > > *dmabuf, > > pci_p2pdma_distance(adev->pdev, attach->dev, false) < 0) > > attach->peer2peer = false; > > > > + /* > > + * P2P access needs the exporter awake for the lifetime of the > > + * attachment. pm_runtime_get_if_active() never resumes the device, > > + * so it cannot deadlock against the reservation taken during resume. > > + * A negative return means runtime PM is disabled and the device > > + * cannot suspend, in which case the put in detach is a no-op. > > + */ > > + if (attach->peer2peer && > > + !pm_runtime_get_if_active(adev_to_drm(adev)->dev)) > > + attach->peer2peer = false; > > + > > r = dma_resv_lock(bo->tbo.base.resv, NULL); > > if (r) > > - return r; > > + goto err_pm_put; > > > > amdgpu_vm_bo_update_shared(bo); > > > > dma_resv_unlock(bo->tbo.base.resv); > > > > return 0; > > + > > +err_pm_put: > > + if (attach->peer2peer) > > + pm_runtime_put_autosuspend(adev_to_drm(adev)->dev); > > + return r; > > +} > > + > > +/** > > + * amdgpu_dma_buf_detach - &dma_buf_ops.detach implementation > > + * > > + * @dmabuf: DMA-buf where we remove the attachment from > > + * @attach: the attachment to remove > > + * > > + * Drop the runtime PM reference taken in amdgpu_dma_buf_attach(). > > + */ > > +static void amdgpu_dma_buf_detach(struct dma_buf *dmabuf, > > + struct dma_buf_attachment *attach) > > +{ > > + struct drm_gem_object *obj = dmabuf->priv; > > + struct amdgpu_bo *bo = gem_to_amdgpu_bo(obj); > > + struct amdgpu_device *adev = amdgpu_ttm_adev(bo->tbo.bdev); > > + > > + if (attach->peer2peer) > > + pm_runtime_put_autosuspend(adev_to_drm(adev)->dev); > > } > > > > /** > > @@ -350,6 +386,7 @@ static void amdgpu_dma_buf_vunmap(struct dma_buf > > *dma_buf, struct iosys_map *map > > > > const struct dma_buf_ops amdgpu_dmabuf_ops = { > > .attach = amdgpu_dma_buf_attach, > > + .detach = amdgpu_dma_buf_detach, > > .pin = amdgpu_dma_buf_pin, > > .unpin = amdgpu_dma_buf_unpin, > > .map_dma_buf = amdgpu_dma_buf_map, >
