Re: [PATCH 6.6.y] scsi: lpfc: Handle mailbox timeouts in lpfc_get_sfp_info
From: Artem Dinaburg
Date: Fri Oct 02 2026 - 13:53:17 EST
Hi Sasha,
Thanks for catching this.
I'm working on, and will send, a v2 addressing the issues.
While working on a fix, I also found what appears to be a separate
mailbox timeout/completion ownership race in both 6.6.y and current
mainline.
I do not have proper hardware to validate certainty, and can't emulate
it in QEMU, but I do have a hardware-independent KUnit test that
reproduces one of the failing interleavings. I'll send that separately
as an RFC to the LPFC/SCSI maintainers.
To be clear, v2 will not attempt to fix this newly identified
potential race. It will remain scoped to the existing backport and the
6.6-specific problems. If the RFC is accepted upstream, I can submit
its stable backport separately.
Thanks,
Artem Dinaburg
On Thu, Oct 1, 2026 at 9:52 AM Sasha Levin <sashal@xxxxxxxxxx> wrote:
>
> > [ Backport to 6.6.y: v6.6 predates ext_buf and uses ctx_buf for the SLI3
> > raw payload. Restore ctx_buf to the saved struct lpfc_dmabuf before
> > testing LPFC_MBX_WAKE so a timed-out mailbox's late default completion
> > sees the DMA descriptor rather than payload bytes. ]
>
> This isn't safe on SLI3 HBAs. When the new 60s wait times out, ctx_buf
> is pointed back at mpsave, a 40-byte struct lpfc_dmabuf, while
> out_ext_byte_len is still 256. If the firmware completes between 60s and
> the 300s mailbox timeout, the SLI3 interrupt handler copies those 256
> bytes into the dmabuf. That is a slab overflow, and lpfc_mbuf_free() then
> runs on the corrupted virt/phys.
>
> --
> Thanks,
> Sasha