Re: [PATCH] PCI: quirks: Fix out-of-bounds MMIO read in nvme_disable_and_flr()
From: Mohamad Raizudeen
Date: Thu Sep 03 2026 - 02:07:35 EST
On Wed, Sep 02, 2026 at 10:21:56PM -0600, Alex Williamson wrote:
> On Thu, 3 Sep 2026 08:54:56 +0530
> Mohamad Raizudeen <raizudeen.kerneldev@xxxxxxxxx> wrote:
>
> > On Wed, Sep 02, 2026 at 11:51:17AM -0600, Alex Williamson wrote:
> > > On Mon, 17 Aug 2026 14:54:47 +0530
> > > Mohamad Raizudeen <raizudeen.kerneldev@xxxxxxxxx> wrote:
> > >
> > > > In nvme_disable_and_flr(), the PCI bar is mapped using
> > > > NVME_REG_CC + sizeof(cfg) which is (0x14 + 4 = 0x18 bytes)
> > > >
> > > > However, the function later reads the controller status from
> > > > NVME_REG_CSTS - offset 0x1C, which is outside the mapped 0x18 byte
> > > > boundary and it can cause a page fault or kernel panic on architectures
> > > > that enforce strict MMIO boundaries.
> > >
> > > What are those architectures? The bug and fix look correct, but the
> > > risk seems overstated. Thanks,
> > >
> > > Alex
> > >
> > Hi Alex,
> >
> > Arm64, risc-v do panic on out-of-bounds mmio. But you are right that the
> > risk is overstated, since most standard servers return all ones instead
> > of crashing.
>
> Are you building with some sort of sanity or debug checking enabled?
>
> While we're clearly violating the API accessing beyond the requested
> length, it's my understanding that we're generally working with
> PAGE_SIZE mappings at the MMU, so unless we're crossing a page, the
> access should work regardless. Even for the cited archs.
>
> The -1 return is typically related to how the platform handles
> master-abort when accessing unimplemented or disabled MMIO space. This
> out-of-bounds access shouldn't be triggering that, it's still within
> the enabled BAR range of the device.
>
> Anyway, I'd be curious to see the backtrace and whether there's a
> config option to enable such sanity checking. Thanks,
>
> Alex
>
No, I am not runnning with any debug enabled, So I don't have a
bactrace. You are completely right about the PAGE_SIZE mappings. I agree
that I was mistaken about the panic.
Also, I already sent a v2 that drops that speculation and just focuses
on the undefined behavior of the out-of-bounds read.
Thanks,
Mohamad Raizudeen
> > > > Fix this by increasing the mapping size to include NVME_REG_CSTS.
> > > >
> > > > Fixes: ffb0863426eb9 ("PCI: Disable Samsung SM961/PM961 NVMe before FLR")
> > > > Signed-off-by: Mohamad Raizudeen <raizudeen.kerneldev@xxxxxxxxx>
> > > > ---
> > > > drivers/pci/quirks.c | 2 +-
> > > > 1 file changed, 1 insertion(+), 1 deletion(-)
> > > >
> > > > diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c
> > > > index b09f27f7846f..ed03892cc960 100644
> > > > --- a/drivers/pci/quirks.c
> > > > +++ b/drivers/pci/quirks.c
> > > > @@ -4090,7 +4090,7 @@ static int nvme_disable_and_flr(struct pci_dev *dev, bool probe)
> > > > if (probe)
> > > > return 0;
> > > >
> > > > - bar = pci_iomap(dev, 0, NVME_REG_CC + sizeof(cfg));
> > > > + bar = pci_iomap(dev, 0, NVME_REG_CSTS + sizeof(cfg));
> > > > if (!bar)
> > > > return -ENOTTY;
> > > >
> > >
>