[PATCH] drm/amdgpu: fix NULL RAS context dereference in UMC page retirement
From: Haotian Zhang
Date: Thu Oct 08 2026 - 12:11:58 EST
amdgpu_umc_handle_bad_pages() and amdgpu_umc_do_page_retirement() take the
RAS context from amdgpu_ras_get_context() and dereference it without
checking for NULL. That context is NULL whenever RAS is disabled: during
amdgpu_ras_init() the release_con path calls amdgpu_ras_set_context(adev,
NULL) and frees it when !adev->ras_enabled on non-VEGA20 hardware. The
GFX11 KFD poison consumption interrupt (kfd_int_process_v11.c) is not
gated on RAS and reaches amdgpu_umc_pasid_poison_handler() ->
amdgpu_umc_do_page_retirement() -> amdgpu_umc_handle_bad_pages(), which
then crashes on mutex_lock(&con->page_retirement_lock) and
con->eeprom_control.
Return early from amdgpu_umc_handle_bad_pages() and
amdgpu_umc_do_page_retirement() when the RAS context is NULL.
Fixes: 513befa63446 ("drm/amdgpu: message smu to update hbm bad page number")
Assisted-by: DeepSeek-V4.1-Flash
Signed-off-by: Haotian Zhang <vulab@xxxxxxxxxxx>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_umc.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_umc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_umc.c
index 71d34fd09385..832e9f64421d 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_umc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_umc.c
@@ -101,6 +101,10 @@ void amdgpu_umc_handle_bad_pages(struct amdgpu_device *adev,
int ret = 0;
unsigned long err_count;
+ /* RAS is disabled, nothing to retire. */
+ if (!con)
+ return;
+
amdgpu_ras_get_error_query_mode(adev, &error_query_mode);
err_data->err_addr =
@@ -216,6 +220,9 @@ static int amdgpu_umc_do_page_retirement(struct amdgpu_device *adev,
kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
amdgpu_umc_handle_bad_pages(adev, ras_error_status);
+ if (!con)
+ return AMDGPU_RAS_SUCCESS;
+
if ((err_data->ue_count || err_data->de_count) &&
(reset || amdgpu_ras_is_rma(adev))) {
con->gpu_reset_flags |= reset;
--
2.25.1