mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Aditya Garg <aditya.garg@linux.dev>
To: Alex Deucher <alexdeucher@gmail.com>, Bjorn Helgaas <helgaas@kernel.org>
Cc: "Perlow, Jason" <jperlow@gmail.com>,
	Lukas Wunner <lukas@wunner.de>,
	Bjorn Helgaas <bhelgaas@google.com>,
	linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org,
	regressions@lists.linux.dev
Subject: Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
Date: Wed, 07 Oct 2026 00:15:44 +0530	[thread overview]
Message-ID: <02FB87FF-9CFD-4CD6-BC94-1D2F7B9C4C94@linux.dev> (raw)
In-Reply-To: <CADnq5_O2TWeWEE1NsuCVvtWLKWyOiJHChPGP9CN4Zu=hvs+jmA@mail.gmail.com>



On 6 October 2026 10:07:18 pm IST, Alex Deucher <alexdeucher@gmail.com> wrote:
>On Mon, Oct 5, 2026 at 7:19 PM Bjorn Helgaas <helgaas@kernel.org> wrote:
>>
>> [+cc Alex]
>>
>> On Mon, Oct 05, 2026 at 01:53:55PM -0400, Perlow, Jason wrote:
>> > Hi Lukas, Bjorn,
>> >
>> > Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
>> > Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
>> > power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
>> > not. Bisection lands on that commit, and v7.3-rc6 with only that
>> > commit reverted no longer powers off. aer.c has not changed since, as
>> > of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
>> > existing report.
>>
>> Alex reported something similar at
>> https://bugzilla.kernel.org/show_bug.cgi?id=222095.
>>
>> Can you try the experiment mentioned there?  The patches Lukas
>> mentioned don't apply cleanly on v7.3-rc1, but the conflict looks
>> trivial?
>
>Lukas' patches fixes some boards, but unfortunately, I'm still seeing
>failures on others.  Reverting the patch fixes those failures.  I've
>updated the ticket.

Considering the number of devices breaking has started increasing inspite of being just in rc, that commit should be reverted.
>
>Alex
>
>>
>> > Hardware
>> > --------
>> >
>> > MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
>> > machine has a Core i7-9750H; a second unit of the same model, used for
>> > cross-checks, has a Core i9-9980HK. I have not tested any other T2
>> > model. The internal SSD, the T2 and the audio device are functions of
>> > one Apple PCIe device behind root port 00:1b.0:
>> >
>> >   04:00.0 Apple ANS2 NVMe          [106b:2005]
>> >   04:00.1 Apple T2 Bridge          [106b:1801]
>> >   04:00.2 Apple T2 Secure Enclave  [106b:1802]
>> >   04:00.3 Apple Audio Device       [106b:1803]
>> >
>> > Symptom
>> > -------
>> >
>> > The machine powers off abruptly 17 to 24 s after boot. There is no
>> > shutdown sequence; the previous boot's journal simply ends. Nothing is
>> > logged beforehand: no AER message, no oops, no lockup, no thermal
>> > event. The T2 controls power on these machines, so I assume (but
>> > cannot show) that the T2/SMC removes power.
>> >
>> > Bisect
>> > ------
>> >
>> > v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
>> > patches in the kernel image, CONFIG_PCIEAER=y,
>> > CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
>> > 19 steps, 4 skipped because those trees oops early for an unrelated
>> > reason (the ones I looked at were in thunderbolt icm_probe at about
>> > 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
>> > logging between 17.9 and 24.1 s. The complete history, every commit
>> > tested and every result, is in Appendix B, and the exact method in
>> > Appendix A. The result:
>> >
>> >   # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
>> >   # Non-Fatal Errors
>> >
>> > Revert test, v7.3-rc6, same config:
>> >
>> >   v7.3-rc6 unmodified:            powers off at 16.7 s
>> >   only eddba19b8b5f reverted:     survives the full 64 s capture
>> >
>> > With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
>> > series (one kernel image, built once), a normal desktop runs on both
>> > MacBookPro16,1 units: the second one has been up for more than 45
>> > minutes, and the bisect machine ran sessions of 32 and 22 minutes.
>> > For completeness: the bisect machine once lost power after about 10
>> > minutes while I was manually switching the display mux and powering
>> > off the AMD GPU, which I do not expect the firmware to support. I
>> > have not established the cause and have no evidence either way on
>> > whether it is related.
>> >
>> > Which devices are affected
>> > --------------------------
>> >
>> > On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
>> > Correctable Error Status register has AdvNonFatalErr latched on
>> > exactly these functions, with every Uncorrectable Error Status
>> > register clear:
>> >
>> >   04:00.0 04:00.1 04:00.2 04:00.3  (the Apple device above)
>> >   01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
>> >
>> > The pattern is identical on both MacBookPro16,1 units (same model, so
>> > this says nothing about other T2 models). The Titan Ridge 4C bridges
>> > and NHI on the same machine do not have the bit set.
>> > These devices report the advisory bit without any matching
>> > Uncorrectable Error status, which looks like the "non-compliant
>> > products" case the commit message mentions.
>> >
>> > Only one of the two machines was used for the bisect and the
>> > power-off tests above. I have not yet booted an unreverted 7.3 kernel
>> > on the second one, so I cannot yet say that the power-off reproduces
>> > there.
>> >
>> > Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
>> > T2) has the same bit latched on its NVIDIA TU106 functions
>> > [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
>> > 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
>> > applied (plus the t2linux series) for more than 15 hours without a
>> > problem, and the same reverted kernel image as above also runs on it
>> > normally. So unmasking the bit is not harmful in general; something
>> > specific to the T2 platform is.
>> >
>> > What I do not know
>> > ------------------
>> >
>> > Which device triggers the power-off, and why. My guess, and it is
>> > only a guess: treating a possible Advisory Non-Fatal Error as
>> > non-Advisory and recovering through the uncorrectable path resets or
>> > disturbs a T2 function, and the T2 then powers the machine down. I
>> > intend to build a diagnostic kernel that can leave the bit masked per
>> > device to find out which one.
>> >
>> > Possibly related, different symptom: "PCI/portdev: Disable AER for
>> > Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
>> > warnings on T2 iMacs.
>> >
>> > What I am asking
>> > ----------------
>> >
>> > Which direction would you prefer: a revert, or a quirk that keeps
>> > Advisory Non-Fatal Errors masked on the affected Apple functions (and
>> > possibly the AMD switch port)? The t2linux project carries a revert
>> > for now: https://github.com/t2linux/linux-t2-patches/pull/70
>> >
>> > I can test patches on real hardware and can provide full lspci -vvv
>> > output, the complete bisect log and the captured boot logs.
>> >
>> > #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
>> >
>> > Thanks,
>> > Jason Perlow
>> >
>> >
>> > APPENDIX A - METHOD
>> >
>> > Machine: one MacBookPro16,1 for every boot below, running a Debian
>> > trixie userland from its internal SSD. The distribution's 7.2.9 kernel
>> > was the default boot entry between tests.
>> >
>> > Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
>> > candidate was built on a separate x86-64 build host (16 threads, gcc
>> > 15.2.0, binutils 2.46) with the same recipe:
>> >
>> >   cp bisect-trimmed.config .config
>> >   make olddefconfig
>> >   make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
>> >        KDEB_PKGVERSION=<release>-1
>> >
>> > bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
>> > quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
>> > CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
>> > CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
>> > for both v7.3-rc6 tests. No out-of-tree patches were applied to the
>> > kernel. The release strings come from each tree's Makefile, so commits
>> > on 7.2-based topic branches show as 7.2.0.
>> >
>> > Boot: the .deb packages were installed on the laptop and booted through
>> > a one-shot rEFInd entry; the default entry stayed the stable kernel.
>> > Kernel command line for every test boot:
>> >
>> >   console=tty0 ignore_loglevel keep_bootcon initcall_debug
>> >   printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
>> >   systemd.show_status=1 systemd.unit=multi-user.target
>> >   modprobe.blacklist=sbs,sbshc
>> >   systemd.mask=ncz-usb2-rescan.service
>> >   systemd.wants=ncz-netlog.service
>> >
>> > (plus the root= options). So there was no graphical session. sbs and
>> > sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
>> > unrelated distribution boot workaround. Steps 13 to 19 and the two
>> > v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
>> > added after the early oopses (in thunderbolt icm_probe, at about 5 s)
>> > had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
>> > enabled.
>> >
>> > Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
>> > and the journal over TCP to a second machine from early multi-user
>> > boot, so captures begin at roughly 10 s of uptime. Each capture file
>> > records the uptime of its last line.
>> >
>> > Verdicts:
>> >   good: the machine kept running and logging past the point where bad
>> >         kernels die (the cut always came before 25 s).
>> >   bad:  the capture stops before 45 s, the machine stays unreachable
>> >         for at least 60 s, and the next boot's journal shows the
>> >         previous boot ending with no clean shutdown.
>> >   skip: a kernel oops or panic in the capture or the previous boot's
>> >         journal.
>> >
>> > After a cut the laptop does not restart by itself, so it was powered
>> > on by hand and booted the stable kernel; the verdict was then confirmed
>> > from the previous boot's journal.
>> >
>> > Caveats:
>> >   - One machine was used for the bisect and for the power-off tests.
>> >   - No desktop session was running.
>> >   - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
>> >     on.
>> >   - The laptop was powered from, and networked through, a Thunderbolt
>> >     dock during the bisect. An earlier test of a T2-patched
>> >     7.3.0-rc5 kernel with the dock unplugged also lost power, but I
>> >     did not repeat the bisect undocked.
>> >   - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
>> >     t2smp, t2thunderbolt) were installed for these kernels and may
>> >     have been loaded, which would taint them. An earlier test with
>> >     those modules blocked still lost power.
>> >   - Step 12 ran 816 s before an unrelated oops, so it did not show
>> >     the cut and was treated as a skip rather than as good; this does
>> >     not change the result.
>> >   - The device that triggers the cut is not identified.
>> >
>> > APPENDIX B - FULL BISECT HISTORY
>> >
>> > Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
>> > commit tested is the merge base, "Linux 7.2". Observed = uptime of the
>> > last line captured. Results: 8 good, 7 bad, 4 skip.
>> >
>> >  #  commit        result observed   subject
>> >  1  8d3ae59288f1  good   55.8 s     Linux 7.2
>> >  2  56ea4e86832d  good   67.7 s     nstree: check listing permission
>> >                                     before taking a namespace ref
>> >  3  93e4b3076b5f  bad    24.1 s     Merge tag 'char-misc-7.3-rc1'
>> >  4  21bd0802cd3f  good   60.1 s     Merge tag 'for-linus' (rdma)
>> >  5  0b0e645ed2c8  bad    22.9 s     Merge tag 'auxdisplay-v7.3-1'
>> >  6  e5f92606156a  good   69.5 s     Merge tag 'mm-nonmm-stable-2026-08-
>> >                                     22-16-57'
>> >  7  b6b019a1d9b9  good   69.2 s     Merge tag 'parisc-for-7.3-rc1'
>> >  8  b130a2caf5d3  skip   oops 5.3 s Merge branch
>> >                                     'pci/controller/tegra264'
>> >  9  455b454c87bb  good   69.0 s     i3c: mipi-i3c-hci: Add support for
>> >                                     AMD_PT I3C controller
>> > 10  9bb52aa1972d  skip   oops 5.5 s Merge branch
>> >                                     'pci/controller/dwc-meson'
>> > 11  b4b07fb82b9e  skip   oops 5.3 s Merge branch 'pci/wake'
>> > 12  651fb94aaf24  skip   oops 816 s alpha/PCI: Fix I/O port accessor
>> >                                     argument order in
>> >                                     pci_legacy_write()
>> > 13  625ae0ff41e5  bad    19.9 s     Merge branch 'pci/dt-binding'
>> > 14  9f91b2b716a0  bad    19.1 s     Merge branch 'pci/procfs'
>> > 15  ea55835bc538  bad    23.3 s     Merge branch 'pci/dpc'
>> > 16  d358e9ad15c2  bad    22.3 s     Merge branch 'pci/aer'
>> > 17  8446e1147f65  good   69.7 s     PCI/AER: Deduplicate logging of
>> >                                     Error Source Identification
>> > 18  f141f74c45c6  good   67.6 s     PCI/AER: Move retrieval of FEP and
>> >                                     TLP Log into helper
>> > 19  eddba19b8b5f  bad    17.9 s     PCI/AER: Support Advisory
>> >                                     Non-Fatal Errors
>> >
>> > Full hashes of the decisive steps:
>> >   eddba19b8b5f76d57424ee328a68fd495c5db857  (first bad)
>> >   f141f74c45c6f774eebdb7e45bd609be5122bfa8  (last good before it)
>> >   8446e1147f65563d374ffa54dc3ba81adb1342c5
>> >
>> > The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
>> > unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
>> > survived the full 64 s capture.
>> >
>> > --
>> > Jason

  reply	other threads:[~2026-10-06 18:46 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05 17:53 Perlow, Jason
2026-10-05 18:37 ` Perlow, Jason
2026-10-05 23:19 ` Bjorn Helgaas
2026-10-06  0:17   ` Perlow, Jason
2026-10-06  2:56     ` Lukas Wunner
2026-10-06  8:31       ` Aditya Garg
2026-10-06  9:44         ` Lukas Wunner
2026-10-06 15:00           ` Perlow, Jason
2026-10-06 16:37   ` Alex Deucher
2026-10-06 18:45     ` Aditya Garg [this message]
2026-10-06 21:19     ` Lukas Wunner

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=02FB87FF-9CFD-4CD6-BC94-1D2F7B9C4C94@linux.dev \
    --to=aditya.garg@linux.dev \
    --cc=alexdeucher@gmail.com \
    --cc=bhelgaas@google.com \
    --cc=helgaas@kernel.org \
    --cc=jperlow@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=lukas@wunner.de \
    --cc=regressions@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®