Re: [PATCH v3] misc: mei: fix race condition between client teardown and read completion
From: nirbhayykumarr
Date: Mon Aug 24 2026 - 07:00:37 EST
Hi Greg,
To answer your first question directly: Yes, I used an LLM to help polish my changelog and commit description, and learn the strict kernel mailing list etiquette. This is my very first kernel contribution, and I used it to help navigate the patching process. However, the core technical work, finding the bug, root causing it, writing the C patch, and testing it is entirely my own.
Regarding how this was found: This came out of a broader research project centered around systematic fuzzing and interface testing of /dev/mei0. As part of testing various subsystem boundaries, I wrote a custom multi-threaded C stress utility that rapidly hammers the driver, more specifically, one thread streaming asynchronous MKHI requests while a second thread concurrently drives rapid open/close/reconnect cycles.
During this stress test, dmesg caught the WARNING at drivers/misc/mei/client.c:698 (mei_cl_unlink) because cl->rd_completed was not empty. Then I traced the driver source code in drivers/misc/mei/ to understand why rd_completed was non-empty at unlink time: mei_release() called mei_cl_flush_queues() before mei_cl_disconnect(), but mei_cl_disconnect() temporarily drops dev->device_lock while waiting for firmware response. During that lock-drop window, deferred read completions on the interrupt thread re-populated cl->rd_completed before mei_cl_unlink() ran. Moving the queue flush inside mei_cl_unlink() under the link state lock resolved the race and silenced the WARN_ON.
And my colleague simply helped me sanity check the trace to ensure it was a legitimate kernel driver bug and not a misunderstanding on my end before reporting it.
Thanks,
Nirbhay Kumar
On Monday, August 24th, 2026 at 3:10 PM, gregkh@xxxxxxxxxxxxxxxxxxx <gregkh@xxxxxxxxxxxxxxxxxxx> wrote:
> On Mon, Aug 24, 2026 at 09:32:27AM +0000, nirbhayykumarr@xxxxxxxxx wrote:
> > In mei_release(), a host client is torn down upon close(). During this
> > teardown sequence, mei_cl_disconnect() is invoked, which releases
> > dev->device_lock while waiting for the firmware response.
> >
> > If an in-flight read request was previously submitted, an incoming
> > completion interrupt processed concurrently by the MEI interrupt
> > handler can add a completed callback into cl->rd_completed via
> > mei_cl_add_rd_completed().
> >
> > Because mei_cl_flush_queues(cl, NULL) was invoked before mei_cl_unlink(cl),
> > an incoming completion callback can slip into cl->rd_completed after the
> > flush has completed but before the client is unlinked from dev->file_list.
> > When mei_cl_unlink() is subsequently called, the invariant check at
> > drivers/misc/mei/client.c:698 triggers:
> >
> > WARN_ON(!list_empty(&cl->rd_completed) ||
> > !list_empty(&cl->rd_pending) ||
> > !list_empty(&cl->link));
> >
> > Call trace:
> > WARNING: CPU: 2 PID: 5056 at drivers/misc/mei/client.c:698 mei_cl_unlink+0xaa/0x140 [mei]
> > RIP: 0010:mei_cl_unlink+0xaa/0x140 [mei]
> > Call Trace:
> > <TASK>
> > mei_release+0x202/0x270 [mei]
> > __fput+0x105/0x2e0
> > __x64_sys_close+0x90/0x140
> > do_syscall_64+0xaa/0x660
> > entry_SYSCALL_64_after_hwframe+0x77/0x7f
> > </TASK>
> >
> > Immediately following mei_cl_unlink(), mei_release() calls kfree(cl).
> > If any remaining or deferred callback references the freed client, a
> > use-after-free occurs.
> >
> > Fix this by flushing queues after unlinking the client from dev->file_list
> > inside mei_cl_unlink(), preventing concurrent IRQ completions from
> > populating the client's completed queue during teardown.
> >
> > Fixes: f35fe5f47ed0 ("mei: add a vtag map for each client")
> > Signed-off-by: Nirbhay Kumar <nirbhayykumarr@xxxxxxxxx>
> > Cc: stable@xxxxxxxxxxxxxxx
> > ---
> > v3:
> > - Removed non-standard Helped-by tags.
> > - Omitted Assisted-by: No AI or co-authors were used. The "we" in my original report referred to a colleague who merely verified and confirmed the bug.
>
> Just to confirm, the original bug report, and test program, and this
> hugely long changelog were not generated by any sort of LLM at all?
>
> So how was this issue found? What tool was used to poke around on the
> mei codebase?
>
> thanks,
>
> greg k-h
>