Re: [PATCH] amdkfd: fix doorbell allocation race creating duplicate allocations
From: Kuehling, Felix
Date: Fri Oct 09 2026 - 11:34:27 EST
On 2026-10-09 11:00, Kuehling, Felix wrote:
On 2026-10-09 06:18, Henry Martin wrote:
The check-and-allocate sequence on pdd->qpd.proc_doorbells is
lockless at every call site of kfd_alloc_process_doorbells():
kfd_get_process_doorbells() (mmap of /dev/kfd), kfd_ioctl_create_queue()
and the CRIU restore path each test the pointer under different
locking (or none), so two racing threads can both observe NULL and
each allocate a doorbell BO and bitmap, orphaning one allocation and
handing the two mappings different backing pages.
Move the serialization into kfd_alloc_process_doorbells() itself:
take kfd->doorbell_mutex and re-check under it, so all three call
sites are covered regardless of their own locking. The existing
doorbell_mutex users (kfd_get_kernel_doorbell() /
kfd_release_kernel_doorbell()) only guard the device-global
kfd->doorbell_bitmap in short non-sleeping sections and never call
into this path, so there is no nesting.
This vulnerability was discovered by Tencent CodeBuddy Security.
Fixes: 428542d91772 ("drm/amd/display: Setup for mmhubbub3_warmup_mcif with big buffer")
Signed-off-by: Henry Martin <bsdhenrymartin@xxxxxxxxx>
Reviewed-by: Felix Kuehling <felix.kuehling@xxxxxxx>
I'm applying the patch to amd-staging-drm-next.
An automatic Claude review pointed out two problems with the patch. I agree with both:
The Fixes: tag points at an unrelated drm/amd/display commit (428542d91772). It must be corrected — 2105a15a2046 ("drm/amdgpu: use doorbell mgr for kfd process doorbells") is the commit that introduced the allocation being fixed.
The patch holds the device-global kfd->doorbell_mutex across bitmap_zalloc(GFP_KERNEL) and amdgpu_bo_create_kernel(), while that same mutex is already acquired under dqm_lock via start_cpsch -> pm_init -> kq_initialize -> kfd_get_kernel_doorbell(). That is a plausible ABBA inversion once reclaim/TTM eviction runs the KFD eviction fence back into dqm_lock. Please either post lockdep (CONFIG_PROVE_LOCKING) results for a boot + queue-create cycle, or switch to a per-process-device lock, which also avoids serialising unrelated processes behind a device-wide mutex for per-process state.
Regards,
Felix
Thanks,
Felix
---
drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c | 15 ++++++++++++++-
1 file changed, 14 insertions(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c b/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
index fdcf7f2d1b5b4..07f881573837e 100644
--- a/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
+++ b/drivers/gpu/drm/amd/amdkfd/kfd_doorbell.c
@@ -257,12 +257,21 @@ int kfd_alloc_process_doorbells(struct kfd_dev *kfd, struct kfd_process_device *
int r;
struct qcm_process_device *qpd = &pdd->qpd;
+ mutex_lock(&kfd->doorbell_mutex);
+
+ /* Another thread may have won the check-and-allocate race */
+ if (qpd->proc_doorbells) {
+ mutex_unlock(&kfd->doorbell_mutex);
+ return 0;
+ }
+
/* Allocate bitmap for dynamic doorbell allocation */
qpd->doorbell_bitmap = bitmap_zalloc(KFD_MAX_NUM_OF_QUEUES_PER_PROCESS,
GFP_KERNEL);
if (!qpd->doorbell_bitmap) {
DRM_ERROR("Failed to allocate process doorbell bitmap\n");
- return -ENOMEM;
+ r = -ENOMEM;
+ goto unlock;
}
r = init_doorbell_bitmap(&pdd->qpd, kfd);
@@ -284,11 +293,15 @@ int kfd_alloc_process_doorbells(struct kfd_dev *kfd, struct kfd_process_device *
DRM_ERROR("Failed to allocate process doorbells\n");
goto err;
}
+
+ mutex_unlock(&kfd->doorbell_mutex);
return 0;
err:
bitmap_free(qpd->doorbell_bitmap);
qpd->doorbell_bitmap = NULL;
+unlock:
+ mutex_unlock(&kfd->doorbell_mutex);
return r;
}
--
2.43.7