Re: [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay
From: Daniel Machon
Date: Tue Sep 15 2026 - 06:54:01 EST
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider:
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> - [Medium] The FDMA interrupt is hooked up (and thereby unmasked in the
> OIC) long before any FDMA state exists.
>
> --
>
> misc: lan966x-pci: dts: add fdma interrupt to overlay
>
> This adds OIC interrupt 14 as the "fdma" named interrupt to the lan966x
> PCI device tree overlay, so FDMA-based injection/extraction can be used
> when the switch is reached over PCIe.
>
> > diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> > index 7bb726550caf..5bb12dbc0843 100644
> > --- a/drivers/misc/lan966x_pci.dtso
> > +++ b/drivers/misc/lan966x_pci.dtso
> > @@ -141,8 +141,9 @@ switch: switch@e0000000 {
> >
> > interrupt-parent = <&oic>;
> > interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
> > + <14 IRQ_TYPE_LEVEL_HIGH>,
> > <9 IRQ_TYPE_LEVEL_HIGH>;
> > - interrupt-names = "xtr", "ana";
> > + interrupt-names = "xtr", "fdma", "ana";
>
> [Severity: Medium]
> With source 14 now named "fdma", platform_get_irq_byname(pdev, "fdma")
> starts succeeding on the PCIe path, and lan966x_probe() requests the
> handler right there:
>
> drivers/net/ethernet/microchip/lan966x/lan966x_main.c:lan966x_probe() {
> lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
> if (lan966x->fdma_irq > 0) {
> err = devm_request_irq(&pdev->dev, lan966x->fdma_irq,
> lan966x_fdma_irq_handler, 0,
> "fdma irq", lan966x);
> ...
> lan966x->fdma = true;
> }
> }
>
> Is this ordering safe? The OIC maps and unmasks source 14 at
> devm_request_irq() time, but at that point no FDMA state exists yet:
>
> - netif_napi_add()/napi_enable() only run later, from
> lan966x_port_init() -> lan966x_fdma_netdev_init() in the per-port
> probe loop, so lan966x->napi is still the zeroed devm_kzalloc()
> memory and napi.poll is NULL.
>
> - the DCB rings, the LLP/ATU windows and the FDMA_CTRL NRESET pulse are
> all done even later, from lan966x->ops->fdma_init() ->
> lan966x_fdma_pci_init().
>
> If the OIC delivers an FDMA interrupt in that window, the handler
> schedules NAPI unconditionally:
>
> drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c:lan966x_fdma_irq_handler() {
> db = lan_rd(lan966x, FDMA_INTR_DB);
> err = lan_rd(lan966x, FDMA_INTR_ERR);
>
> if (db) {
> lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
> lan_wr(db, lan966x, FDMA_INTR_DB);
>
> napi_schedule(&lan966x->napi);
> }
> }
>
> Can this then oops in __napi_poll(), either calling n->poll() while it
> is still NULL, or running lan966x_fdma_pci_napi_poll() against
> rx->fdma.dcbs == NULL and n_dcbs == 0, where fdma_has_frames() ->
> fdma_db_get() dereferences the dcbs array?
>
> How reachable is that window in practice? The series itself describes
> the card arriving with FDMA state from a previous boot. The commit
> "net: lan966x: clear FDMA interrupt stickies after switch reset" says
> the stickies "sit in FDMA_INTR_ERR until the FDMA IRQ is requested later
> in probe, at which point the handler fires immediately and WARNs", and
> "net: lan966x: add shutdown callback to stop FDMA on reboot" says that
> on the endpoint "channels stay active and interrupt sources stay armed
> across the reset" and that "FDMA_INTR_ENA persists on the card across a
> warm reboot".
>
> The sticky clear in lan966x_reset_switch() clears the latched bits at
> one instant, but does not mask FDMA_INTR_ENA/FDMA_INTR_DB_ENA and does
> not stop the channels. After a reboot that skips device_shutdown()
> (panic reboot, kexec/kdump, watchdog warm reset with the card still
> powered), could a still-active channel latch FDMA_INTR_DB again after
> the clear and have it delivered as soon as source 14 is unmasked?
>
> Would it be more robust to reset and mask the FDMA block (the
> FDMA_CTRL NRESET pulse that already exists in lan966x_fdma_pci_init(),
> plus clearing FDMA_INTR_ENA/FDMA_INTR_DB_ENA) before
> devm_request_irq(), or to request the FDMA IRQ only after fdma_init()
> has built the rings and added the NAPI instance?
>
I think this is pre-existing. It can oops, both variants, and yes - a channel the
previous kernel left active can re-latch FDMA_INTR_DB before devm_request_irq()
unmasks the source. Your second suggestion is the right fix; the reset half of the
first is what fdma_init()'s NRESET pulse already does.
None of this originates here: the request has sat before netif_napi_add() and
ops->fdma_init() since c8349639324a ("net: lan966x: Add FDMA functionality"), and
the SoC device tree has named "fdma" since 2952d32c0611 ("ARM: dts: lan966x: add
switch node") - this patch adds one entry to an interrupt-names list. Nothing to
change in this patch; it goes to net with the same reorder as the MTU-change problem
earlier in the series.
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com