On Thu, Sep 24, 2026 at 8:44 AM Pierre-Eric Pelloux-Prayer
<[email protected]> wrote:
>
> For unknown reasons, on gfx12 using multiple entities can causes
> random corruption of BOs with DCC.
> This workaround seems to prevent the issue until the root cause
> is understood and fixed.
>
> Link: https://gitlab.freedesktop.org/drm/amd/-/work_items/5663
> Fixes: 3a6f6eeb3db5 ("drm/amdgpu: give ttm entities access to all the sdma 
> scheds")
> Signed-off-by: Pierre-Eric Pelloux-Prayer <[email protected]>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c | 5 +++++
>  1 file changed, 5 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c 
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
> index 3d620ec2937f..a826e091ea58 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
> @@ -2392,6 +2392,11 @@ void amdgpu_ttm_enable_buffer_funcs(struct 
> amdgpu_device *adev)
>
>         num_clear_entities = MIN(adev->mman.num_buffer_funcs_scheds, 
> TTM_NUM_MOVE_FENCES);
>         num_move_entities = MIN(adev->mman.num_buffer_funcs_scheds, 
> TTM_NUM_MOVE_FENCES);
> +        /* TODO: workaround for DCC corruption when moving BOs from multiple 
> queues at
> +         * the same time: use a single queue until the root cause is 
> identified and fixed.
> +         */
> +       if (IP_VERSION_MAJ(amdgpu_ip_version(adev, SDMA0_HWIP, 0)) == 7)

Probably just limit this to 7.0.0 and 7.0.1.  sdma 7.1 doesn't have
DCC.  WIth that fixed:
Reviewed-by: Alex Deucher <[email protected]>

Alex

> +               num_move_entities = 1;
>
>         adev->mman.clear_entities = kcalloc(num_clear_entities,
>                                                 sizeof(struct 
> amdgpu_ttm_buffer_entity),
> --
> 2.43.0
>

Reply via email to