Re: [PATCH] amdkfd: fix doorbell allocation race creating duplicate allocations

From: Kuehling, Felix

Date: Fri Oct 09 2026 - 11:34:27 EST


On 2026-10-09 11:00, Kuehling, Felix wrote:
On 2026-10-09 06:18, Henry Martin wrote:
The check-and-allocate sequence on pdd->qpd.proc_doorbells is
lockless at every call site of kfd_alloc_process_doorbells():
kfd_get_process_doorbells() (mmap of /dev/kfd), kfd_ioctl_create_queue()
and the CRIU restore path each test the pointer under different
locking (or none), so two racing threads can both observe NULL and
each allocate a doorbell BO and bitmap, orphaning one allocation and
handing the two mappings different backing pages.

Move the serialization into kfd_alloc_process_doorbells() itself:
take kfd->doorbell_mutex and re-check under it, so all three call
sites are covered regardless of their own locking.  The existing
doorbell_mutex users (kfd_get_kernel_doorbell() /
kfd_release_kernel_doorbell()) only guard the device-global
kfd->doorbell_bitmap in short non-sleeping sections and never call
into this path, so there is no nesting.

This vulnerability was discovered by Tencent CodeBuddy Security.

Fixes: 428542d91772 ("drm/amd/display: Setup for mmhubbub3_warmup_mcif with big buffer")
Signed-off-by: Henry Martin <bsdhenrymartin@xxxxxxxxx>

Reviewed-by: Felix Kuehling <felix.kuehling@xxxxxxx>

I'm applying the patch to amd-staging-drm-next.

An automatic Claude review pointed out two problems with the patch. I agree with both:

The Fixes: tag points at an unrelated drm/amd/display commit (428542d91772). It must be corrected — 2105a15a2046 ("drm/amdgpu: use doorbell mgr for kfd process doorbells") is the commit that introduced the allocation being fixed.

The patch holds the device-global kfd->doorbell_mutex across bitmap_zalloc(GFP_KERNEL) and amdgpu_bo_create_kernel(), while that same mutex is already acquired under dqm_lock via start_cpsch -> pm_init -> kq_initialize -> kfd_get_kernel_doorbell(). That is a plausible ABBA inversion once reclaim/TTM eviction runs the KFD eviction fence back into dqm_lock. Please either post lockdep (CONFIG_PROVE_LOCKING) results for a boot + queue-create cycle, or switch to a per-process-device lock, which also avoids serialising unrelated processes behind a device-wide mutex for per-process state.

Regards,
  Felix



Thanks,
  Felix


---
  drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c | 15 ++++++++++++++-
  1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c b/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
index fdcf7f2d1b5b4..07f881573837e 100644
--- a/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
+++ b/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
@@ -257,12 +257,21 @@ int kfd_alloc_process_doorbells(struct kfd_dev *kfd, struct kfd_process_device *
      int r;
      struct qcm_process_device *qpd = &pdd->qpd;

+    mutex_lock(&kfd->doorbell_mutex);
+
+    /* Another thread may have won the check-and-allocate race */
+    if (qpd->proc_doorbells) {
+        mutex_unlock(&kfd->doorbell_mutex);
+        return 0;
+    }
+
      /* Allocate bitmap for dynamic doorbell allocation */
      qpd->doorbell_bitmap = bitmap_zalloc(KFD_MAX_NUM_OF_QUEUES_PER_PROCESS,
                           GFP_KERNEL);
      if (!qpd->doorbell_bitmap) {
          DRM_ERROR("Failed to allocate process doorbell bitmap\n");
-        return -ENOMEM;
+        r = -ENOMEM;
+        goto unlock;
      }

      r = init_doorbell_bitmap(&pdd->qpd, kfd);
@@ -284,11 +293,15 @@ int kfd_alloc_process_doorbells(struct kfd_dev *kfd, struct kfd_process_device *
          DRM_ERROR("Failed to allocate process doorbells\n");
          goto err;
      }
+
+    mutex_unlock(&kfd->doorbell_mutex);
      return 0;

  err:
      bitmap_free(qpd->doorbell_bitmap);
      qpd->doorbell_bitmap = NULL;
+unlock:
+    mutex_unlock(&kfd->doorbell_mutex);
      return r;
  }

--
2.43.7