From: Bjorn Helgaas <helgaas@kernel.org>
To: "Perlow, Jason" <jperlow@gmail.com>
Cc: Lukas Wunner <lukas@wunner.de>,
Bjorn Helgaas <bhelgaas@google.com>,
linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org,
regressions@lists.linux.dev, Aditya Garg <aditya.garg@linux.dev>,
Alex Deucher <alexdeucher@gmail.com>
Subject: Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
Date: Mon, 5 Oct 2026 18:19:14 -0500 [thread overview]
Message-ID: <20261005231914.GA639967@bhelgaas> (raw)
In-Reply-To: <CABZrw2HcaOh3DHTJtxU5dXrb3Vg19scr6+qR-xe_+R7rME35_Q@mail.gmail.com>
[+cc Alex]
On Mon, Oct 05, 2026 at 01:53:55PM -0400, Perlow, Jason wrote:
> Hi Lukas, Bjorn,
>
> Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
> Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
> power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
> not. Bisection lands on that commit, and v7.3-rc6 with only that
> commit reverted no longer powers off. aer.c has not changed since, as
> of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
> existing report.
Alex reported something similar at
https://bugzilla.kernel.org/show_bug.cgi?id=222095.
Can you try the experiment mentioned there? The patches Lukas
mentioned don't apply cleanly on v7.3-rc1, but the conflict looks
trivial?
> Hardware
> --------
>
> MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
> machine has a Core i7-9750H; a second unit of the same model, used for
> cross-checks, has a Core i9-9980HK. I have not tested any other T2
> model. The internal SSD, the T2 and the audio device are functions of
> one Apple PCIe device behind root port 00:1b.0:
>
> 04:00.0 Apple ANS2 NVMe [106b:2005]
> 04:00.1 Apple T2 Bridge [106b:1801]
> 04:00.2 Apple T2 Secure Enclave [106b:1802]
> 04:00.3 Apple Audio Device [106b:1803]
>
> Symptom
> -------
>
> The machine powers off abruptly 17 to 24 s after boot. There is no
> shutdown sequence; the previous boot's journal simply ends. Nothing is
> logged beforehand: no AER message, no oops, no lockup, no thermal
> event. The T2 controls power on these machines, so I assume (but
> cannot show) that the T2/SMC removes power.
>
> Bisect
> ------
>
> v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
> patches in the kernel image, CONFIG_PCIEAER=y,
> CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
> 19 steps, 4 skipped because those trees oops early for an unrelated
> reason (the ones I looked at were in thunderbolt icm_probe at about
> 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
> logging between 17.9 and 24.1 s. The complete history, every commit
> tested and every result, is in Appendix B, and the exact method in
> Appendix A. The result:
>
> # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
> # Non-Fatal Errors
>
> Revert test, v7.3-rc6, same config:
>
> v7.3-rc6 unmodified: powers off at 16.7 s
> only eddba19b8b5f reverted: survives the full 64 s capture
>
> With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
> series (one kernel image, built once), a normal desktop runs on both
> MacBookPro16,1 units: the second one has been up for more than 45
> minutes, and the bisect machine ran sessions of 32 and 22 minutes.
> For completeness: the bisect machine once lost power after about 10
> minutes while I was manually switching the display mux and powering
> off the AMD GPU, which I do not expect the firmware to support. I
> have not established the cause and have no evidence either way on
> whether it is related.
>
> Which devices are affected
> --------------------------
>
> On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
> Correctable Error Status register has AdvNonFatalErr latched on
> exactly these functions, with every Uncorrectable Error Status
> register clear:
>
> 04:00.0 04:00.1 04:00.2 04:00.3 (the Apple device above)
> 01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
>
> The pattern is identical on both MacBookPro16,1 units (same model, so
> this says nothing about other T2 models). The Titan Ridge 4C bridges
> and NHI on the same machine do not have the bit set.
> These devices report the advisory bit without any matching
> Uncorrectable Error status, which looks like the "non-compliant
> products" case the commit message mentions.
>
> Only one of the two machines was used for the bisect and the
> power-off tests above. I have not yet booted an unreverted 7.3 kernel
> on the second one, so I cannot yet say that the power-off reproduces
> there.
>
> Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
> T2) has the same bit latched on its NVIDIA TU106 functions
> [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
> 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
> applied (plus the t2linux series) for more than 15 hours without a
> problem, and the same reverted kernel image as above also runs on it
> normally. So unmasking the bit is not harmful in general; something
> specific to the T2 platform is.
>
> What I do not know
> ------------------
>
> Which device triggers the power-off, and why. My guess, and it is
> only a guess: treating a possible Advisory Non-Fatal Error as
> non-Advisory and recovering through the uncorrectable path resets or
> disturbs a T2 function, and the T2 then powers the machine down. I
> intend to build a diagnostic kernel that can leave the bit masked per
> device to find out which one.
>
> Possibly related, different symptom: "PCI/portdev: Disable AER for
> Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
> warnings on T2 iMacs.
>
> What I am asking
> ----------------
>
> Which direction would you prefer: a revert, or a quirk that keeps
> Advisory Non-Fatal Errors masked on the affected Apple functions (and
> possibly the AMD switch port)? The t2linux project carries a revert
> for now: https://github.com/t2linux/linux-t2-patches/pull/70
>
> I can test patches on real hardware and can provide full lspci -vvv
> output, the complete bisect log and the captured boot logs.
>
> #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
>
> Thanks,
> Jason Perlow
>
>
> APPENDIX A - METHOD
>
> Machine: one MacBookPro16,1 for every boot below, running a Debian
> trixie userland from its internal SSD. The distribution's 7.2.9 kernel
> was the default boot entry between tests.
>
> Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
> candidate was built on a separate x86-64 build host (16 threads, gcc
> 15.2.0, binutils 2.46) with the same recipe:
>
> cp bisect-trimmed.config .config
> make olddefconfig
> make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
> KDEB_PKGVERSION=<release>-1
>
> bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
> quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
> CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
> CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
> for both v7.3-rc6 tests. No out-of-tree patches were applied to the
> kernel. The release strings come from each tree's Makefile, so commits
> on 7.2-based topic branches show as 7.2.0.
>
> Boot: the .deb packages were installed on the laptop and booted through
> a one-shot rEFInd entry; the default entry stayed the stable kernel.
> Kernel command line for every test boot:
>
> console=tty0 ignore_loglevel keep_bootcon initcall_debug
> printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
> systemd.show_status=1 systemd.unit=multi-user.target
> modprobe.blacklist=sbs,sbshc
> systemd.mask=ncz-usb2-rescan.service
> systemd.wants=ncz-netlog.service
>
> (plus the root= options). So there was no graphical session. sbs and
> sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
> unrelated distribution boot workaround. Steps 13 to 19 and the two
> v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
> added after the early oopses (in thunderbolt icm_probe, at about 5 s)
> had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
> enabled.
>
> Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
> and the journal over TCP to a second machine from early multi-user
> boot, so captures begin at roughly 10 s of uptime. Each capture file
> records the uptime of its last line.
>
> Verdicts:
> good: the machine kept running and logging past the point where bad
> kernels die (the cut always came before 25 s).
> bad: the capture stops before 45 s, the machine stays unreachable
> for at least 60 s, and the next boot's journal shows the
> previous boot ending with no clean shutdown.
> skip: a kernel oops or panic in the capture or the previous boot's
> journal.
>
> After a cut the laptop does not restart by itself, so it was powered
> on by hand and booted the stable kernel; the verdict was then confirmed
> from the previous boot's journal.
>
> Caveats:
> - One machine was used for the bisect and for the power-off tests.
> - No desktop session was running.
> - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
> on.
> - The laptop was powered from, and networked through, a Thunderbolt
> dock during the bisect. An earlier test of a T2-patched
> 7.3.0-rc5 kernel with the dock unplugged also lost power, but I
> did not repeat the bisect undocked.
> - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
> t2smp, t2thunderbolt) were installed for these kernels and may
> have been loaded, which would taint them. An earlier test with
> those modules blocked still lost power.
> - Step 12 ran 816 s before an unrelated oops, so it did not show
> the cut and was treated as a skip rather than as good; this does
> not change the result.
> - The device that triggers the cut is not identified.
>
> APPENDIX B - FULL BISECT HISTORY
>
> Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
> commit tested is the merge base, "Linux 7.2". Observed = uptime of the
> last line captured. Results: 8 good, 7 bad, 4 skip.
>
> # commit result observed subject
> 1 8d3ae59288f1 good 55.8 s Linux 7.2
> 2 56ea4e86832d good 67.7 s nstree: check listing permission
> before taking a namespace ref
> 3 93e4b3076b5f bad 24.1 s Merge tag 'char-misc-7.3-rc1'
> 4 21bd0802cd3f good 60.1 s Merge tag 'for-linus' (rdma)
> 5 0b0e645ed2c8 bad 22.9 s Merge tag 'auxdisplay-v7.3-1'
> 6 e5f92606156a good 69.5 s Merge tag 'mm-nonmm-stable-2026-08-
> 22-16-57'
> 7 b6b019a1d9b9 good 69.2 s Merge tag 'parisc-for-7.3-rc1'
> 8 b130a2caf5d3 skip oops 5.3 s Merge branch
> 'pci/controller/tegra264'
> 9 455b454c87bb good 69.0 s i3c: mipi-i3c-hci: Add support for
> AMD_PT I3C controller
> 10 9bb52aa1972d skip oops 5.5 s Merge branch
> 'pci/controller/dwc-meson'
> 11 b4b07fb82b9e skip oops 5.3 s Merge branch 'pci/wake'
> 12 651fb94aaf24 skip oops 816 s alpha/PCI: Fix I/O port accessor
> argument order in
> pci_legacy_write()
> 13 625ae0ff41e5 bad 19.9 s Merge branch 'pci/dt-binding'
> 14 9f91b2b716a0 bad 19.1 s Merge branch 'pci/procfs'
> 15 ea55835bc538 bad 23.3 s Merge branch 'pci/dpc'
> 16 d358e9ad15c2 bad 22.3 s Merge branch 'pci/aer'
> 17 8446e1147f65 good 69.7 s PCI/AER: Deduplicate logging of
> Error Source Identification
> 18 f141f74c45c6 good 67.6 s PCI/AER: Move retrieval of FEP and
> TLP Log into helper
> 19 eddba19b8b5f bad 17.9 s PCI/AER: Support Advisory
> Non-Fatal Errors
>
> Full hashes of the decisive steps:
> eddba19b8b5f76d57424ee328a68fd495c5db857 (first bad)
> f141f74c45c6f774eebdb7e45bd609be5122bfa8 (last good before it)
> 8446e1147f65563d374ffa54dc3ba81adb1342c5
>
> The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
> unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
> survived the full 64 s capture.
>
> --
> Jason
next prev parent reply other threads:[~2026-10-05 23:19 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 17:53 Perlow, Jason
2026-10-05 18:37 ` Perlow, Jason
2026-10-05 23:19 ` Bjorn Helgaas [this message]
2026-10-06 0:17 ` Perlow, Jason
2026-10-06 2:56 ` Lukas Wunner
2026-10-06 8:31 ` Aditya Garg
2026-10-06 9:44 ` Lukas Wunner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005231914.GA639967@bhelgaas \
--to=helgaas@kernel.org \
--cc=aditya.garg@linux.dev \
--cc=alexdeucher@gmail.com \
--cc=bhelgaas@google.com \
--cc=jperlow@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=lukas@wunner.de \
--cc=regressions@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®