[PATCH v2] drm/amdgpu: implement workaround for sdma dcc corruption

From: Pierre-Eric Pelloux-Prayer

Date: Thu Sep 24 2026 - 10:15:07 EST


For unknown reasons, on gfx12 using multiple entities can causes
random corruption of BOs with DCC.
This workaround seems to prevent the issue until the root cause
is understood and fixed.

Link: https://gitlab.freedesktop.org/drm/amd/-/work_items/5663
Fixes: 3a6f6eeb3db5 ("drm/amdgpu: give ttm entities access to all the sdma scheds")
Signed-off-by: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@xxxxxxx>
Reviewed-by: Alex Deucher <alexander.deucher@xxxxxxx>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
index 3d620ec2937f..e9de5c4b7f63 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
@@ -2367,7 +2367,7 @@ void amdgpu_ttm_fini(struct amdgpu_device *adev)
void amdgpu_ttm_enable_buffer_funcs(struct amdgpu_device *adev)
{
struct ttm_resource_manager *man = ttm_manager_type(&adev->mman.bdev, TTM_PL_VRAM);
- u32 num_clear_entities, num_move_entities;
+ u32 num_clear_entities, num_move_entities, sdma_ip_version;
int r, i, j;

if (!adev->mman.initialized || amdgpu_in_reset(adev) ||
@@ -2392,6 +2392,12 @@ void amdgpu_ttm_enable_buffer_funcs(struct amdgpu_device *adev)

num_clear_entities = MIN(adev->mman.num_buffer_funcs_scheds, TTM_NUM_MOVE_FENCES);
num_move_entities = MIN(adev->mman.num_buffer_funcs_scheds, TTM_NUM_MOVE_FENCES);
+ /* TODO: workaround for DCC corruption when moving BOs from multiple queues at
+ * the same time: use a single queue until the root cause is identified and fixed.
+ */
+ sdma_ip_version = amdgpu_ip_version(adev, SDMA0_HWIP, 0);
+ if (sdma_ip_version == IP_VERSION(7, 0, 0) || sdma_ip_version == IP_VERSION(7, 0, 1))
+ num_move_entities = 1;

adev->mman.clear_entities = kcalloc(num_clear_entities,
sizeof(struct amdgpu_ttm_buffer_entity),
--
2.43.0