Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot

From: Alex Deucher

Date: Tue Oct 06 2026 - 12:40:04 EST


On Mon, Oct 5, 2026 at 7:19 PM Bjorn Helgaas <helgaas@xxxxxxxxxx> wrote:
>
> [+cc Alex]
>
> On Mon, Oct 05, 2026 at 01:53:55PM -0400, Perlow, Jason wrote:
> > Hi Lukas, Bjorn,
> >
> > Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
> > Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
> > power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
> > not. Bisection lands on that commit, and v7.3-rc6 with only that
> > commit reverted no longer powers off. aer.c has not changed since, as
> > of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
> > existing report.
>
> Alex reported something similar at
> https://bugzilla.kernel.org/show_bug.cgi?id=222095.
>
> Can you try the experiment mentioned there? The patches Lukas
> mentioned don't apply cleanly on v7.3-rc1, but the conflict looks
> trivial?

Lukas' patches fixes some boards, but unfortunately, I'm still seeing
failures on others. Reverting the patch fixes those failures. I've
updated the ticket.

Alex

>
> > Hardware
> > --------
> >
> > MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
> > machine has a Core i7-9750H; a second unit of the same model, used for
> > cross-checks, has a Core i9-9980HK. I have not tested any other T2
> > model. The internal SSD, the T2 and the audio device are functions of
> > one Apple PCIe device behind root port 00:1b.0:
> >
> > 04:00.0 Apple ANS2 NVMe [106b:2005]
> > 04:00.1 Apple T2 Bridge [106b:1801]
> > 04:00.2 Apple T2 Secure Enclave [106b:1802]
> > 04:00.3 Apple Audio Device [106b:1803]
> >
> > Symptom
> > -------
> >
> > The machine powers off abruptly 17 to 24 s after boot. There is no
> > shutdown sequence; the previous boot's journal simply ends. Nothing is
> > logged beforehand: no AER message, no oops, no lockup, no thermal
> > event. The T2 controls power on these machines, so I assume (but
> > cannot show) that the T2/SMC removes power.
> >
> > Bisect
> > ------
> >
> > v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
> > patches in the kernel image, CONFIG_PCIEAER=y,
> > CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
> > 19 steps, 4 skipped because those trees oops early for an unrelated
> > reason (the ones I looked at were in thunderbolt icm_probe at about
> > 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
> > logging between 17.9 and 24.1 s. The complete history, every commit
> > tested and every result, is in Appendix B, and the exact method in
> > Appendix A. The result:
> >
> > # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
> > # Non-Fatal Errors
> >
> > Revert test, v7.3-rc6, same config:
> >
> > v7.3-rc6 unmodified: powers off at 16.7 s
> > only eddba19b8b5f reverted: survives the full 64 s capture
> >
> > With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
> > series (one kernel image, built once), a normal desktop runs on both
> > MacBookPro16,1 units: the second one has been up for more than 45
> > minutes, and the bisect machine ran sessions of 32 and 22 minutes.
> > For completeness: the bisect machine once lost power after about 10
> > minutes while I was manually switching the display mux and powering
> > off the AMD GPU, which I do not expect the firmware to support. I
> > have not established the cause and have no evidence either way on
> > whether it is related.
> >
> > Which devices are affected
> > --------------------------
> >
> > On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
> > Correctable Error Status register has AdvNonFatalErr latched on
> > exactly these functions, with every Uncorrectable Error Status
> > register clear:
> >
> > 04:00.0 04:00.1 04:00.2 04:00.3 (the Apple device above)
> > 01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
> >
> > The pattern is identical on both MacBookPro16,1 units (same model, so
> > this says nothing about other T2 models). The Titan Ridge 4C bridges
> > and NHI on the same machine do not have the bit set.
> > These devices report the advisory bit without any matching
> > Uncorrectable Error status, which looks like the "non-compliant
> > products" case the commit message mentions.
> >
> > Only one of the two machines was used for the bisect and the
> > power-off tests above. I have not yet booted an unreverted 7.3 kernel
> > on the second one, so I cannot yet say that the power-off reproduces
> > there.
> >
> > Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
> > T2) has the same bit latched on its NVIDIA TU106 functions
> > [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
> > 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
> > applied (plus the t2linux series) for more than 15 hours without a
> > problem, and the same reverted kernel image as above also runs on it
> > normally. So unmasking the bit is not harmful in general; something
> > specific to the T2 platform is.
> >
> > What I do not know
> > ------------------
> >
> > Which device triggers the power-off, and why. My guess, and it is
> > only a guess: treating a possible Advisory Non-Fatal Error as
> > non-Advisory and recovering through the uncorrectable path resets or
> > disturbs a T2 function, and the T2 then powers the machine down. I
> > intend to build a diagnostic kernel that can leave the bit masked per
> > device to find out which one.
> >
> > Possibly related, different symptom: "PCI/portdev: Disable AER for
> > Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
> > warnings on T2 iMacs.
> >
> > What I am asking
> > ----------------
> >
> > Which direction would you prefer: a revert, or a quirk that keeps
> > Advisory Non-Fatal Errors masked on the affected Apple functions (and
> > possibly the AMD switch port)? The t2linux project carries a revert
> > for now: https://github.com/t2linux/linux-t2-patches/pull/70
> >
> > I can test patches on real hardware and can provide full lspci -vvv
> > output, the complete bisect log and the captured boot logs.
> >
> > #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
> >
> > Thanks,
> > Jason Perlow
> >
> >
> > APPENDIX A - METHOD
> >
> > Machine: one MacBookPro16,1 for every boot below, running a Debian
> > trixie userland from its internal SSD. The distribution's 7.2.9 kernel
> > was the default boot entry between tests.
> >
> > Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
> > candidate was built on a separate x86-64 build host (16 threads, gcc
> > 15.2.0, binutils 2.46) with the same recipe:
> >
> > cp bisect-trimmed.config .config
> > make olddefconfig
> > make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
> > KDEB_PKGVERSION=<release>-1
> >
> > bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
> > quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
> > CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
> > CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
> > for both v7.3-rc6 tests. No out-of-tree patches were applied to the
> > kernel. The release strings come from each tree's Makefile, so commits
> > on 7.2-based topic branches show as 7.2.0.
> >
> > Boot: the .deb packages were installed on the laptop and booted through
> > a one-shot rEFInd entry; the default entry stayed the stable kernel.
> > Kernel command line for every test boot:
> >
> > console=tty0 ignore_loglevel keep_bootcon initcall_debug
> > printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
> > systemd.show_status=1 systemd.unit=multi-user.target
> > modprobe.blacklist=sbs,sbshc
> > systemd.mask=ncz-usb2-rescan.service
> > systemd.wants=ncz-netlog.service
> >
> > (plus the root= options). So there was no graphical session. sbs and
> > sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
> > unrelated distribution boot workaround. Steps 13 to 19 and the two
> > v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
> > added after the early oopses (in thunderbolt icm_probe, at about 5 s)
> > had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
> > enabled.
> >
> > Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
> > and the journal over TCP to a second machine from early multi-user
> > boot, so captures begin at roughly 10 s of uptime. Each capture file
> > records the uptime of its last line.
> >
> > Verdicts:
> > good: the machine kept running and logging past the point where bad
> > kernels die (the cut always came before 25 s).
> > bad: the capture stops before 45 s, the machine stays unreachable
> > for at least 60 s, and the next boot's journal shows the
> > previous boot ending with no clean shutdown.
> > skip: a kernel oops or panic in the capture or the previous boot's
> > journal.
> >
> > After a cut the laptop does not restart by itself, so it was powered
> > on by hand and booted the stable kernel; the verdict was then confirmed
> > from the previous boot's journal.
> >
> > Caveats:
> > - One machine was used for the bisect and for the power-off tests.
> > - No desktop session was running.
> > - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
> > on.
> > - The laptop was powered from, and networked through, a Thunderbolt
> > dock during the bisect. An earlier test of a T2-patched
> > 7.3.0-rc5 kernel with the dock unplugged also lost power, but I
> > did not repeat the bisect undocked.
> > - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
> > t2smp, t2thunderbolt) were installed for these kernels and may
> > have been loaded, which would taint them. An earlier test with
> > those modules blocked still lost power.
> > - Step 12 ran 816 s before an unrelated oops, so it did not show
> > the cut and was treated as a skip rather than as good; this does
> > not change the result.
> > - The device that triggers the cut is not identified.
> >
> > APPENDIX B - FULL BISECT HISTORY
> >
> > Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
> > commit tested is the merge base, "Linux 7.2". Observed = uptime of the
> > last line captured. Results: 8 good, 7 bad, 4 skip.
> >
> > # commit result observed subject
> > 1 8d3ae59288f1 good 55.8 s Linux 7.2
> > 2 56ea4e86832d good 67.7 s nstree: check listing permission
> > before taking a namespace ref
> > 3 93e4b3076b5f bad 24.1 s Merge tag 'char-misc-7.3-rc1'
> > 4 21bd0802cd3f good 60.1 s Merge tag 'for-linus' (rdma)
> > 5 0b0e645ed2c8 bad 22.9 s Merge tag 'auxdisplay-v7.3-1'
> > 6 e5f92606156a good 69.5 s Merge tag 'mm-nonmm-stable-2026-08-
> > 22-16-57'
> > 7 b6b019a1d9b9 good 69.2 s Merge tag 'parisc-for-7.3-rc1'
> > 8 b130a2caf5d3 skip oops 5.3 s Merge branch
> > 'pci/controller/tegra264'
> > 9 455b454c87bb good 69.0 s i3c: mipi-i3c-hci: Add support for
> > AMD_PT I3C controller
> > 10 9bb52aa1972d skip oops 5.5 s Merge branch
> > 'pci/controller/dwc-meson'
> > 11 b4b07fb82b9e skip oops 5.3 s Merge branch 'pci/wake'
> > 12 651fb94aaf24 skip oops 816 s alpha/PCI: Fix I/O port accessor
> > argument order in
> > pci_legacy_write()
> > 13 625ae0ff41e5 bad 19.9 s Merge branch 'pci/dt-binding'
> > 14 9f91b2b716a0 bad 19.1 s Merge branch 'pci/procfs'
> > 15 ea55835bc538 bad 23.3 s Merge branch 'pci/dpc'
> > 16 d358e9ad15c2 bad 22.3 s Merge branch 'pci/aer'
> > 17 8446e1147f65 good 69.7 s PCI/AER: Deduplicate logging of
> > Error Source Identification
> > 18 f141f74c45c6 good 67.6 s PCI/AER: Move retrieval of FEP and
> > TLP Log into helper
> > 19 eddba19b8b5f bad 17.9 s PCI/AER: Support Advisory
> > Non-Fatal Errors
> >
> > Full hashes of the decisive steps:
> > eddba19b8b5f76d57424ee328a68fd495c5db857 (first bad)
> > f141f74c45c6f774eebdb7e45bd609be5122bfa8 (last good before it)
> > 8446e1147f65563d374ffa54dc3ba81adb1342c5
> >
> > The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
> > unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
> > survived the full 64 s capture.
> >
> > --
> > Jason