mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
@ 2026-10-05 17:53 Perlow, Jason
  2026-10-05 18:37 ` Perlow, Jason
  2026-10-05 23:19 ` Bjorn Helgaas
  0 siblings, 2 replies; 6+ messages in thread
From: Perlow, Jason @ 2026-10-05 17:53 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas
  Cc: linux-pci, linux-kernel, regressions, Aditya Garg

Hi Lukas, Bjorn,

Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
not. Bisection lands on that commit, and v7.3-rc6 with only that
commit reverted no longer powers off. aer.c has not changed since, as
of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
existing report.

Hardware
--------

MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
machine has a Core i7-9750H; a second unit of the same model, used for
cross-checks, has a Core i9-9980HK. I have not tested any other T2
model. The internal SSD, the T2 and the audio device are functions of
one Apple PCIe device behind root port 00:1b.0:

  04:00.0 Apple ANS2 NVMe          [106b:2005]
  04:00.1 Apple T2 Bridge          [106b:1801]
  04:00.2 Apple T2 Secure Enclave  [106b:1802]
  04:00.3 Apple Audio Device       [106b:1803]

Symptom
-------

The machine powers off abruptly 17 to 24 s after boot. There is no
shutdown sequence; the previous boot's journal simply ends. Nothing is
logged beforehand: no AER message, no oops, no lockup, no thermal
event. The T2 controls power on these machines, so I assume (but
cannot show) that the T2/SMC removes power.

Bisect
------

v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
patches in the kernel image, CONFIG_PCIEAER=y,
CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
19 steps, 4 skipped because those trees oops early for an unrelated
reason (the ones I looked at were in thunderbolt icm_probe at about
5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
logging between 17.9 and 24.1 s. The complete history, every commit
tested and every result, is in Appendix B, and the exact method in
Appendix A. The result:

  # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
  # Non-Fatal Errors

Revert test, v7.3-rc6, same config:

  v7.3-rc6 unmodified:            powers off at 16.7 s
  only eddba19b8b5f reverted:     survives the full 64 s capture

With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
series (one kernel image, built once), a normal desktop runs on both
MacBookPro16,1 units: the second one has been up for more than 45
minutes, and the bisect machine ran sessions of 32 and 22 minutes.
For completeness: the bisect machine once lost power after about 10
minutes while I was manually switching the display mux and powering
off the AMD GPU, which I do not expect the firmware to support. I
have not established the cause and have no evidence either way on
whether it is related.

Which devices are affected
--------------------------

On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
Correctable Error Status register has AdvNonFatalErr latched on
exactly these functions, with every Uncorrectable Error Status
register clear:

  04:00.0 04:00.1 04:00.2 04:00.3  (the Apple device above)
  01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]

The pattern is identical on both MacBookPro16,1 units (same model, so
this says nothing about other T2 models). The Titan Ridge 4C bridges
and NHI on the same machine do not have the bit set.
These devices report the advisory bit without any matching
Uncorrectable Error status, which looks like the "non-compliant
products" case the commit message mentions.

Only one of the two machines was used for the bisect and the
power-off tests above. I have not yet booted an unreverted 7.3 kernel
on the second one, so I cannot yet say that the power-off reproduces
there.

Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
T2) has the same bit latched on its NVIDIA TU106 functions
[10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
applied (plus the t2linux series) for more than 15 hours without a
problem, and the same reverted kernel image as above also runs on it
normally. So unmasking the bit is not harmful in general; something
specific to the T2 platform is.

What I do not know
------------------

Which device triggers the power-off, and why. My guess, and it is
only a guess: treating a possible Advisory Non-Fatal Error as
non-Advisory and recovering through the uncorrectable path resets or
disturbs a T2 function, and the T2 then powers the machine down. I
intend to build a diagnostic kernel that can leave the bit masked per
device to find out which one.

Possibly related, different symptom: "PCI/portdev: Disable AER for
Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
warnings on T2 iMacs.

What I am asking
----------------

Which direction would you prefer: a revert, or a quirk that keeps
Advisory Non-Fatal Errors masked on the affected Apple functions (and
possibly the AMD switch port)? The t2linux project carries a revert
for now: https://github.com/t2linux/linux-t2-patches/pull/70

I can test patches on real hardware and can provide full lspci -vvv
output, the complete bisect log and the captured boot logs.

#regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857

Thanks,
Jason Perlow


APPENDIX A - METHOD

Machine: one MacBookPro16,1 for every boot below, running a Debian
trixie userland from its internal SSD. The distribution's 7.2.9 kernel
was the default boot entry between tests.

Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
candidate was built on a separate x86-64 build host (16 threads, gcc
15.2.0, binutils 2.46) with the same recipe:

  cp bisect-trimmed.config .config
  make olddefconfig
  make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
       KDEB_PKGVERSION=<release>-1

bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
for both v7.3-rc6 tests. No out-of-tree patches were applied to the
kernel. The release strings come from each tree's Makefile, so commits
on 7.2-based topic branches show as 7.2.0.

Boot: the .deb packages were installed on the laptop and booted through
a one-shot rEFInd entry; the default entry stayed the stable kernel.
Kernel command line for every test boot:

  console=tty0 ignore_loglevel keep_bootcon initcall_debug
  printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
  systemd.show_status=1 systemd.unit=multi-user.target
  modprobe.blacklist=sbs,sbshc
  systemd.mask=ncz-usb2-rescan.service
  systemd.wants=ncz-netlog.service

(plus the root= options). So there was no graphical session. sbs and
sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
unrelated distribution boot workaround. Steps 13 to 19 and the two
v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
added after the early oopses (in thunderbolt icm_probe, at about 5 s)
had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
enabled.

Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
and the journal over TCP to a second machine from early multi-user
boot, so captures begin at roughly 10 s of uptime. Each capture file
records the uptime of its last line.

Verdicts:
  good: the machine kept running and logging past the point where bad
        kernels die (the cut always came before 25 s).
  bad:  the capture stops before 45 s, the machine stays unreachable
        for at least 60 s, and the next boot's journal shows the
        previous boot ending with no clean shutdown.
  skip: a kernel oops or panic in the capture or the previous boot's
        journal.

After a cut the laptop does not restart by itself, so it was powered
on by hand and booted the stable kernel; the verdict was then confirmed
from the previous boot's journal.

Caveats:
  - One machine was used for the bisect and for the power-off tests.
  - No desktop session was running.
  - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
    on.
  - The laptop was powered from, and networked through, a Thunderbolt
    dock during the bisect. An earlier test of a T2-patched
    7.3.0-rc5 kernel with the dock unplugged also lost power, but I
    did not repeat the bisect undocked.
  - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
    t2smp, t2thunderbolt) were installed for these kernels and may
    have been loaded, which would taint them. An earlier test with
    those modules blocked still lost power.
  - Step 12 ran 816 s before an unrelated oops, so it did not show
    the cut and was treated as a skip rather than as good; this does
    not change the result.
  - The device that triggers the cut is not identified.

APPENDIX B - FULL BISECT HISTORY

Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
commit tested is the merge base, "Linux 7.2". Observed = uptime of the
last line captured. Results: 8 good, 7 bad, 4 skip.

 #  commit        result observed   subject
 1  8d3ae59288f1  good   55.8 s     Linux 7.2
 2  56ea4e86832d  good   67.7 s     nstree: check listing permission
                                    before taking a namespace ref
 3  93e4b3076b5f  bad    24.1 s     Merge tag 'char-misc-7.3-rc1'
 4  21bd0802cd3f  good   60.1 s     Merge tag 'for-linus' (rdma)
 5  0b0e645ed2c8  bad    22.9 s     Merge tag 'auxdisplay-v7.3-1'
 6  e5f92606156a  good   69.5 s     Merge tag 'mm-nonmm-stable-2026-08-
                                    22-16-57'
 7  b6b019a1d9b9  good   69.2 s     Merge tag 'parisc-for-7.3-rc1'
 8  b130a2caf5d3  skip   oops 5.3 s Merge branch
                                    'pci/controller/tegra264'
 9  455b454c87bb  good   69.0 s     i3c: mipi-i3c-hci: Add support for
                                    AMD_PT I3C controller
10  9bb52aa1972d  skip   oops 5.5 s Merge branch
                                    'pci/controller/dwc-meson'
11  b4b07fb82b9e  skip   oops 5.3 s Merge branch 'pci/wake'
12  651fb94aaf24  skip   oops 816 s alpha/PCI: Fix I/O port accessor
                                    argument order in
                                    pci_legacy_write()
13  625ae0ff41e5  bad    19.9 s     Merge branch 'pci/dt-binding'
14  9f91b2b716a0  bad    19.1 s     Merge branch 'pci/procfs'
15  ea55835bc538  bad    23.3 s     Merge branch 'pci/dpc'
16  d358e9ad15c2  bad    22.3 s     Merge branch 'pci/aer'
17  8446e1147f65  good   69.7 s     PCI/AER: Deduplicate logging of
                                    Error Source Identification
18  f141f74c45c6  good   67.6 s     PCI/AER: Move retrieval of FEP and
                                    TLP Log into helper
19  eddba19b8b5f  bad    17.9 s     PCI/AER: Support Advisory
                                    Non-Fatal Errors

Full hashes of the decisive steps:
  eddba19b8b5f76d57424ee328a68fd495c5db857  (first bad)
  f141f74c45c6f774eebdb7e45bd609be5122bfa8  (last good before it)
  8446e1147f65563d374ffa54dc3ba81adb1342c5

The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
survived the full 64 s capture.

-- 
Jason

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
  2026-10-05 17:53 [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot Perlow, Jason
@ 2026-10-05 18:37 ` Perlow, Jason
  2026-10-05 23:19 ` Bjorn Helgaas
  1 sibling, 0 replies; 6+ messages in thread
From: Perlow, Jason @ 2026-10-05 18:37 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas
  Cc: linux-pci, linux-kernel, regressions, Aditya Garg

Follow-up: the trigger is one device, the T2 Secure Enclave function
(106b:1802, 04:00.2).

To find which device matters I built a diagnostic kernel on v7.3-rc6
(same config and recipe as the bisect, no revert). It adds a boot
parameter, aer_anfe_skip=<list>, that leaves the Advisory Non-Fatal
bit masked on the listed devices (a BDF, a vendor:device pair, or
"all") and logs one line per device at unmask time. Without the
parameter it behaves as v7.3-rc6 does.

Each boot ran on the same MacBookPro16,1, with kernel messages
streamed over the network. A cut means the machine powered off with
no clean shutdown; every cut so far has come before 27 s.

 #  masked (skipped)                      result
 1  all five functions that latch it      survived
 2  04:00.0 04:00.1 04:00.2 04:00.3       survived, 343 s
    (the four Apple functions only;
    the AMD switch port 01:00.0 unmasked)
 3  04:00.0 04:00.1 only                  powered off at 14 s
    (NVMe and T2 Bridge; Enclave and
    Audio unmasked)
 4  04:00.2 only (106b:1802)              survived, 120 s

So the AMD switch port is not involved. Test 3 cuts with the Enclave
and Audio unmasked, and test 4 survives with Audio still unmasked, so
unmasking the bit on 04:00.2 alone is what triggers the power-off.
Each case was run once; I can repeat them if that would help.

I still do not know why the T2 reacts. No AER message is logged before
the cut, and the Uncorrectable Error Status of that function is clear,
so this is only about the Correctable path that the commit enables.

Since the bit is only wrong on this one function, a quirk may be
better than a revert: keep Advisory Non-Fatal Errors masked on Apple
106b:1802, for example with a DECLARE_PCI_FIXUP_EARLY entry that makes
pci_aer_init() leave the mask alone. I am happy to write and test
that patch on both of my T2 machines if you tell me which form you
prefer (a quirk flag in the PCI core, or a device-specific fixup).

The diagnostic patch, the lspci -vvv output and the captured boot
logs are available on request.

Thanks,
Jason Perlow

On Mon, Oct 5, 2026 at 1:53 PM Perlow, Jason <jperlow@gmail.com> wrote:
>
> Hi Lukas, Bjorn,
>
> Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
> Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
> power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
> not. Bisection lands on that commit, and v7.3-rc6 with only that
> commit reverted no longer powers off. aer.c has not changed since, as
> of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
> existing report.
>
> Hardware
> --------
>
> MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
> machine has a Core i7-9750H; a second unit of the same model, used for
> cross-checks, has a Core i9-9980HK. I have not tested any other T2
> model. The internal SSD, the T2 and the audio device are functions of
> one Apple PCIe device behind root port 00:1b.0:
>
>   04:00.0 Apple ANS2 NVMe          [106b:2005]
>   04:00.1 Apple T2 Bridge          [106b:1801]
>   04:00.2 Apple T2 Secure Enclave  [106b:1802]
>   04:00.3 Apple Audio Device       [106b:1803]
>
> Symptom
> -------
>
> The machine powers off abruptly 17 to 24 s after boot. There is no
> shutdown sequence; the previous boot's journal simply ends. Nothing is
> logged beforehand: no AER message, no oops, no lockup, no thermal
> event. The T2 controls power on these machines, so I assume (but
> cannot show) that the T2/SMC removes power.
>
> Bisect
> ------
>
> v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
> patches in the kernel image, CONFIG_PCIEAER=y,
> CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
> 19 steps, 4 skipped because those trees oops early for an unrelated
> reason (the ones I looked at were in thunderbolt icm_probe at about
> 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
> logging between 17.9 and 24.1 s. The complete history, every commit
> tested and every result, is in Appendix B, and the exact method in
> Appendix A. The result:
>
>   # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
>   # Non-Fatal Errors
>
> Revert test, v7.3-rc6, same config:
>
>   v7.3-rc6 unmodified:            powers off at 16.7 s
>   only eddba19b8b5f reverted:     survives the full 64 s capture
>
> With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
> series (one kernel image, built once), a normal desktop runs on both
> MacBookPro16,1 units: the second one has been up for more than 45
> minutes, and the bisect machine ran sessions of 32 and 22 minutes.
> For completeness: the bisect machine once lost power after about 10
> minutes while I was manually switching the display mux and powering
> off the AMD GPU, which I do not expect the firmware to support. I
> have not established the cause and have no evidence either way on
> whether it is related.
>
> Which devices are affected
> --------------------------
>
> On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
> Correctable Error Status register has AdvNonFatalErr latched on
> exactly these functions, with every Uncorrectable Error Status
> register clear:
>
>   04:00.0 04:00.1 04:00.2 04:00.3  (the Apple device above)
>   01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
>
> The pattern is identical on both MacBookPro16,1 units (same model, so
> this says nothing about other T2 models). The Titan Ridge 4C bridges
> and NHI on the same machine do not have the bit set.
> These devices report the advisory bit without any matching
> Uncorrectable Error status, which looks like the "non-compliant
> products" case the commit message mentions.
>
> Only one of the two machines was used for the bisect and the
> power-off tests above. I have not yet booted an unreverted 7.3 kernel
> on the second one, so I cannot yet say that the power-off reproduces
> there.
>
> Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
> T2) has the same bit latched on its NVIDIA TU106 functions
> [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
> 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
> applied (plus the t2linux series) for more than 15 hours without a
> problem, and the same reverted kernel image as above also runs on it
> normally. So unmasking the bit is not harmful in general; something
> specific to the T2 platform is.
>
> What I do not know
> ------------------
>
> Which device triggers the power-off, and why. My guess, and it is
> only a guess: treating a possible Advisory Non-Fatal Error as
> non-Advisory and recovering through the uncorrectable path resets or
> disturbs a T2 function, and the T2 then powers the machine down. I
> intend to build a diagnostic kernel that can leave the bit masked per
> device to find out which one.
>
> Possibly related, different symptom: "PCI/portdev: Disable AER for
> Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
> warnings on T2 iMacs.
>
> What I am asking
> ----------------
>
> Which direction would you prefer: a revert, or a quirk that keeps
> Advisory Non-Fatal Errors masked on the affected Apple functions (and
> possibly the AMD switch port)? The t2linux project carries a revert
> for now: https://github.com/t2linux/linux-t2-patches/pull/70
>
> I can test patches on real hardware and can provide full lspci -vvv
> output, the complete bisect log and the captured boot logs.
>
> #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
>
> Thanks,
> Jason Perlow
>
>
> APPENDIX A - METHOD
>
> Machine: one MacBookPro16,1 for every boot below, running a Debian
> trixie userland from its internal SSD. The distribution's 7.2.9 kernel
> was the default boot entry between tests.
>
> Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
> candidate was built on a separate x86-64 build host (16 threads, gcc
> 15.2.0, binutils 2.46) with the same recipe:
>
>   cp bisect-trimmed.config .config
>   make olddefconfig
>   make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
>        KDEB_PKGVERSION=<release>-1
>
> bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
> quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
> CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
> CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
> for both v7.3-rc6 tests. No out-of-tree patches were applied to the
> kernel. The release strings come from each tree's Makefile, so commits
> on 7.2-based topic branches show as 7.2.0.
>
> Boot: the .deb packages were installed on the laptop and booted through
> a one-shot rEFInd entry; the default entry stayed the stable kernel.
> Kernel command line for every test boot:
>
>   console=tty0 ignore_loglevel keep_bootcon initcall_debug
>   printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
>   systemd.show_status=1 systemd.unit=multi-user.target
>   modprobe.blacklist=sbs,sbshc
>   systemd.mask=ncz-usb2-rescan.service
>   systemd.wants=ncz-netlog.service
>
> (plus the root= options). So there was no graphical session. sbs and
> sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
> unrelated distribution boot workaround. Steps 13 to 19 and the two
> v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
> added after the early oopses (in thunderbolt icm_probe, at about 5 s)
> had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
> enabled.
>
> Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
> and the journal over TCP to a second machine from early multi-user
> boot, so captures begin at roughly 10 s of uptime. Each capture file
> records the uptime of its last line.
>
> Verdicts:
>   good: the machine kept running and logging past the point where bad
>         kernels die (the cut always came before 25 s).
>   bad:  the capture stops before 45 s, the machine stays unreachable
>         for at least 60 s, and the next boot's journal shows the
>         previous boot ending with no clean shutdown.
>   skip: a kernel oops or panic in the capture or the previous boot's
>         journal.
>
> After a cut the laptop does not restart by itself, so it was powered
> on by hand and booted the stable kernel; the verdict was then confirmed
> from the previous boot's journal.
>
> Caveats:
>   - One machine was used for the bisect and for the power-off tests.
>   - No desktop session was running.
>   - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
>     on.
>   - The laptop was powered from, and networked through, a Thunderbolt
>     dock during the bisect. An earlier test of a T2-patched
>     7.3.0-rc5 kernel with the dock unplugged also lost power, but I
>     did not repeat the bisect undocked.
>   - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
>     t2smp, t2thunderbolt) were installed for these kernels and may
>     have been loaded, which would taint them. An earlier test with
>     those modules blocked still lost power.
>   - Step 12 ran 816 s before an unrelated oops, so it did not show
>     the cut and was treated as a skip rather than as good; this does
>     not change the result.
>   - The device that triggers the cut is not identified.
>
> APPENDIX B - FULL BISECT HISTORY
>
> Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
> commit tested is the merge base, "Linux 7.2". Observed = uptime of the
> last line captured. Results: 8 good, 7 bad, 4 skip.
>
>  #  commit        result observed   subject
>  1  8d3ae59288f1  good   55.8 s     Linux 7.2
>  2  56ea4e86832d  good   67.7 s     nstree: check listing permission
>                                     before taking a namespace ref
>  3  93e4b3076b5f  bad    24.1 s     Merge tag 'char-misc-7.3-rc1'
>  4  21bd0802cd3f  good   60.1 s     Merge tag 'for-linus' (rdma)
>  5  0b0e645ed2c8  bad    22.9 s     Merge tag 'auxdisplay-v7.3-1'
>  6  e5f92606156a  good   69.5 s     Merge tag 'mm-nonmm-stable-2026-08-
>                                     22-16-57'
>  7  b6b019a1d9b9  good   69.2 s     Merge tag 'parisc-for-7.3-rc1'
>  8  b130a2caf5d3  skip   oops 5.3 s Merge branch
>                                     'pci/controller/tegra264'
>  9  455b454c87bb  good   69.0 s     i3c: mipi-i3c-hci: Add support for
>                                     AMD_PT I3C controller
> 10  9bb52aa1972d  skip   oops 5.5 s Merge branch
>                                     'pci/controller/dwc-meson'
> 11  b4b07fb82b9e  skip   oops 5.3 s Merge branch 'pci/wake'
> 12  651fb94aaf24  skip   oops 816 s alpha/PCI: Fix I/O port accessor
>                                     argument order in
>                                     pci_legacy_write()
> 13  625ae0ff41e5  bad    19.9 s     Merge branch 'pci/dt-binding'
> 14  9f91b2b716a0  bad    19.1 s     Merge branch 'pci/procfs'
> 15  ea55835bc538  bad    23.3 s     Merge branch 'pci/dpc'
> 16  d358e9ad15c2  bad    22.3 s     Merge branch 'pci/aer'
> 17  8446e1147f65  good   69.7 s     PCI/AER: Deduplicate logging of
>                                     Error Source Identification
> 18  f141f74c45c6  good   67.6 s     PCI/AER: Move retrieval of FEP and
>                                     TLP Log into helper
> 19  eddba19b8b5f  bad    17.9 s     PCI/AER: Support Advisory
>                                     Non-Fatal Errors
>
> Full hashes of the decisive steps:
>   eddba19b8b5f76d57424ee328a68fd495c5db857  (first bad)
>   f141f74c45c6f774eebdb7e45bd609be5122bfa8  (last good before it)
>   8446e1147f65563d374ffa54dc3ba81adb1342c5
>
> The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
> unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
> survived the full 64 s capture.
>
> --
> Jason



-- 
Jason Perlow | Argonaut Media Communications

Voice/Text  (954) 242-3484 | jperlow@gmail.com | Blog: techbroiler.net
Read my Tech and Food Industry articles: https://linktr.ee/jperlow
Bluesky: https://bsky.app/profile/jperlow.bsky.social

Need to schedule a meeting with me? https://bit.ly/3y8P3Gp

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
  2026-10-05 17:53 [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot Perlow, Jason
  2026-10-05 18:37 ` Perlow, Jason
@ 2026-10-05 23:19 ` Bjorn Helgaas
  2026-10-06  0:17   ` Perlow, Jason
  1 sibling, 1 reply; 6+ messages in thread
From: Bjorn Helgaas @ 2026-10-05 23:19 UTC (permalink / raw)
  To: Perlow, Jason
  Cc: Lukas Wunner, Bjorn Helgaas, linux-pci, linux-kernel,
	regressions, Aditya Garg, Alex Deucher

[+cc Alex]

On Mon, Oct 05, 2026 at 01:53:55PM -0400, Perlow, Jason wrote:
> Hi Lukas, Bjorn,
> 
> Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
> Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
> power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
> not. Bisection lands on that commit, and v7.3-rc6 with only that
> commit reverted no longer powers off. aer.c has not changed since, as
> of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
> existing report.

Alex reported something similar at
https://bugzilla.kernel.org/show_bug.cgi?id=222095.

Can you try the experiment mentioned there?  The patches Lukas
mentioned don't apply cleanly on v7.3-rc1, but the conflict looks
trivial?

> Hardware
> --------
> 
> MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
> machine has a Core i7-9750H; a second unit of the same model, used for
> cross-checks, has a Core i9-9980HK. I have not tested any other T2
> model. The internal SSD, the T2 and the audio device are functions of
> one Apple PCIe device behind root port 00:1b.0:
> 
>   04:00.0 Apple ANS2 NVMe          [106b:2005]
>   04:00.1 Apple T2 Bridge          [106b:1801]
>   04:00.2 Apple T2 Secure Enclave  [106b:1802]
>   04:00.3 Apple Audio Device       [106b:1803]
> 
> Symptom
> -------
> 
> The machine powers off abruptly 17 to 24 s after boot. There is no
> shutdown sequence; the previous boot's journal simply ends. Nothing is
> logged beforehand: no AER message, no oops, no lockup, no thermal
> event. The T2 controls power on these machines, so I assume (but
> cannot show) that the T2/SMC removes power.
> 
> Bisect
> ------
> 
> v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
> patches in the kernel image, CONFIG_PCIEAER=y,
> CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
> 19 steps, 4 skipped because those trees oops early for an unrelated
> reason (the ones I looked at were in thunderbolt icm_probe at about
> 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
> logging between 17.9 and 24.1 s. The complete history, every commit
> tested and every result, is in Appendix B, and the exact method in
> Appendix A. The result:
> 
>   # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
>   # Non-Fatal Errors
> 
> Revert test, v7.3-rc6, same config:
> 
>   v7.3-rc6 unmodified:            powers off at 16.7 s
>   only eddba19b8b5f reverted:     survives the full 64 s capture
> 
> With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
> series (one kernel image, built once), a normal desktop runs on both
> MacBookPro16,1 units: the second one has been up for more than 45
> minutes, and the bisect machine ran sessions of 32 and 22 minutes.
> For completeness: the bisect machine once lost power after about 10
> minutes while I was manually switching the display mux and powering
> off the AMD GPU, which I do not expect the firmware to support. I
> have not established the cause and have no evidence either way on
> whether it is related.
> 
> Which devices are affected
> --------------------------
> 
> On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
> Correctable Error Status register has AdvNonFatalErr latched on
> exactly these functions, with every Uncorrectable Error Status
> register clear:
> 
>   04:00.0 04:00.1 04:00.2 04:00.3  (the Apple device above)
>   01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
> 
> The pattern is identical on both MacBookPro16,1 units (same model, so
> this says nothing about other T2 models). The Titan Ridge 4C bridges
> and NHI on the same machine do not have the bit set.
> These devices report the advisory bit without any matching
> Uncorrectable Error status, which looks like the "non-compliant
> products" case the commit message mentions.
> 
> Only one of the two machines was used for the bisect and the
> power-off tests above. I have not yet booted an unreverted 7.3 kernel
> on the second one, so I cannot yet say that the power-off reproduces
> there.
> 
> Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
> T2) has the same bit latched on its NVIDIA TU106 functions
> [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
> 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
> applied (plus the t2linux series) for more than 15 hours without a
> problem, and the same reverted kernel image as above also runs on it
> normally. So unmasking the bit is not harmful in general; something
> specific to the T2 platform is.
> 
> What I do not know
> ------------------
> 
> Which device triggers the power-off, and why. My guess, and it is
> only a guess: treating a possible Advisory Non-Fatal Error as
> non-Advisory and recovering through the uncorrectable path resets or
> disturbs a T2 function, and the T2 then powers the machine down. I
> intend to build a diagnostic kernel that can leave the bit masked per
> device to find out which one.
> 
> Possibly related, different symptom: "PCI/portdev: Disable AER for
> Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
> warnings on T2 iMacs.
> 
> What I am asking
> ----------------
> 
> Which direction would you prefer: a revert, or a quirk that keeps
> Advisory Non-Fatal Errors masked on the affected Apple functions (and
> possibly the AMD switch port)? The t2linux project carries a revert
> for now: https://github.com/t2linux/linux-t2-patches/pull/70
> 
> I can test patches on real hardware and can provide full lspci -vvv
> output, the complete bisect log and the captured boot logs.
> 
> #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
> 
> Thanks,
> Jason Perlow
> 
> 
> APPENDIX A - METHOD
> 
> Machine: one MacBookPro16,1 for every boot below, running a Debian
> trixie userland from its internal SSD. The distribution's 7.2.9 kernel
> was the default boot entry between tests.
> 
> Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
> candidate was built on a separate x86-64 build host (16 threads, gcc
> 15.2.0, binutils 2.46) with the same recipe:
> 
>   cp bisect-trimmed.config .config
>   make olddefconfig
>   make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
>        KDEB_PKGVERSION=<release>-1
> 
> bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
> quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
> CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
> CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
> for both v7.3-rc6 tests. No out-of-tree patches were applied to the
> kernel. The release strings come from each tree's Makefile, so commits
> on 7.2-based topic branches show as 7.2.0.
> 
> Boot: the .deb packages were installed on the laptop and booted through
> a one-shot rEFInd entry; the default entry stayed the stable kernel.
> Kernel command line for every test boot:
> 
>   console=tty0 ignore_loglevel keep_bootcon initcall_debug
>   printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
>   systemd.show_status=1 systemd.unit=multi-user.target
>   modprobe.blacklist=sbs,sbshc
>   systemd.mask=ncz-usb2-rescan.service
>   systemd.wants=ncz-netlog.service
> 
> (plus the root= options). So there was no graphical session. sbs and
> sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
> unrelated distribution boot workaround. Steps 13 to 19 and the two
> v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
> added after the early oopses (in thunderbolt icm_probe, at about 5 s)
> had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
> enabled.
> 
> Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
> and the journal over TCP to a second machine from early multi-user
> boot, so captures begin at roughly 10 s of uptime. Each capture file
> records the uptime of its last line.
> 
> Verdicts:
>   good: the machine kept running and logging past the point where bad
>         kernels die (the cut always came before 25 s).
>   bad:  the capture stops before 45 s, the machine stays unreachable
>         for at least 60 s, and the next boot's journal shows the
>         previous boot ending with no clean shutdown.
>   skip: a kernel oops or panic in the capture or the previous boot's
>         journal.
> 
> After a cut the laptop does not restart by itself, so it was powered
> on by hand and booted the stable kernel; the verdict was then confirmed
> from the previous boot's journal.
> 
> Caveats:
>   - One machine was used for the bisect and for the power-off tests.
>   - No desktop session was running.
>   - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
>     on.
>   - The laptop was powered from, and networked through, a Thunderbolt
>     dock during the bisect. An earlier test of a T2-patched
>     7.3.0-rc5 kernel with the dock unplugged also lost power, but I
>     did not repeat the bisect undocked.
>   - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
>     t2smp, t2thunderbolt) were installed for these kernels and may
>     have been loaded, which would taint them. An earlier test with
>     those modules blocked still lost power.
>   - Step 12 ran 816 s before an unrelated oops, so it did not show
>     the cut and was treated as a skip rather than as good; this does
>     not change the result.
>   - The device that triggers the cut is not identified.
> 
> APPENDIX B - FULL BISECT HISTORY
> 
> Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
> commit tested is the merge base, "Linux 7.2". Observed = uptime of the
> last line captured. Results: 8 good, 7 bad, 4 skip.
> 
>  #  commit        result observed   subject
>  1  8d3ae59288f1  good   55.8 s     Linux 7.2
>  2  56ea4e86832d  good   67.7 s     nstree: check listing permission
>                                     before taking a namespace ref
>  3  93e4b3076b5f  bad    24.1 s     Merge tag 'char-misc-7.3-rc1'
>  4  21bd0802cd3f  good   60.1 s     Merge tag 'for-linus' (rdma)
>  5  0b0e645ed2c8  bad    22.9 s     Merge tag 'auxdisplay-v7.3-1'
>  6  e5f92606156a  good   69.5 s     Merge tag 'mm-nonmm-stable-2026-08-
>                                     22-16-57'
>  7  b6b019a1d9b9  good   69.2 s     Merge tag 'parisc-for-7.3-rc1'
>  8  b130a2caf5d3  skip   oops 5.3 s Merge branch
>                                     'pci/controller/tegra264'
>  9  455b454c87bb  good   69.0 s     i3c: mipi-i3c-hci: Add support for
>                                     AMD_PT I3C controller
> 10  9bb52aa1972d  skip   oops 5.5 s Merge branch
>                                     'pci/controller/dwc-meson'
> 11  b4b07fb82b9e  skip   oops 5.3 s Merge branch 'pci/wake'
> 12  651fb94aaf24  skip   oops 816 s alpha/PCI: Fix I/O port accessor
>                                     argument order in
>                                     pci_legacy_write()
> 13  625ae0ff41e5  bad    19.9 s     Merge branch 'pci/dt-binding'
> 14  9f91b2b716a0  bad    19.1 s     Merge branch 'pci/procfs'
> 15  ea55835bc538  bad    23.3 s     Merge branch 'pci/dpc'
> 16  d358e9ad15c2  bad    22.3 s     Merge branch 'pci/aer'
> 17  8446e1147f65  good   69.7 s     PCI/AER: Deduplicate logging of
>                                     Error Source Identification
> 18  f141f74c45c6  good   67.6 s     PCI/AER: Move retrieval of FEP and
>                                     TLP Log into helper
> 19  eddba19b8b5f  bad    17.9 s     PCI/AER: Support Advisory
>                                     Non-Fatal Errors
> 
> Full hashes of the decisive steps:
>   eddba19b8b5f76d57424ee328a68fd495c5db857  (first bad)
>   f141f74c45c6f774eebdb7e45bd609be5122bfa8  (last good before it)
>   8446e1147f65563d374ffa54dc3ba81adb1342c5
> 
> The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
> unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
> survived the full 64 s capture.
> 
> -- 
> Jason

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
  2026-10-05 23:19 ` Bjorn Helgaas
@ 2026-10-06  0:17   ` Perlow, Jason
  2026-10-06  2:56     ` Lukas Wunner
  0 siblings, 1 reply; 6+ messages in thread
From: Perlow, Jason @ 2026-10-06  0:17 UTC (permalink / raw)
  To: Bjorn Helgaas
  Cc: Lukas Wunner, Bjorn Helgaas, linux-pci, linux-kernel,
	regressions, Aditya Garg, Alex Deucher

Hi Bjorn, Lukas,

Thanks for the pointer. I ran that experiment on the same
MacBookPro16,1.

Kernel: v7.3-rc6 (a90ee4305c4a) plus exactly the two top commits of
l1k/linux aer_unbound, and nothing else (no T2 patches, no revert, no
diagnostic parameter):

  9ce840949ee9 PCI/ERR: Always notify drivers of slot reset
  db16ca536fd0 PCI/ERR: Allow unbound devices to recover from
               Uncorrectable Errors

The first one conflicts trivially in drivers/pci/pcie/err.c on rc6
(the TODO comment in its context is already gone); I resolved it by
making the same removal. The second applied cleanly. Same trimmed
config and boot recipe as my earlier bisect (console only, no
graphical session).

Result: the machine still loses power. It answered over ssh at about
16 s of uptime and the journal of that boot ends at 16.9 s with no
shutdown sequence. Unmodified v7.3-rc6 powered off at 16.7 s in my
earlier test, so these two commits make no difference here.

So on this machine the failure is not the "link reset before a
driver is bound" case from Alex's report. As in my earlier tests no
AER error message is logged before the power loss (only the usual
"OS assumes control of AER" line) and the Uncorrectable Error Status
registers are clear.

For reference, the earlier result with a diagnostic boot parameter
that leaves the Advisory Non-Fatal bit masked per device:

  masked only on 04:00.2 (Apple T2 Secure Enclave, 106b:1802):
      stays up (120 s)
  masked only on 04:00.0 and 04:00.1 (NVMe, T2 Bridge):
      power off at 14 s
  masked on all four Apple functions 04:00.0-.3: stays up (343 s)

So here the trigger is unmasking Advisory Non-Fatal Errors on that
one function. It looks like a second, separate way for
eddba19b8b5f to hurt a platform.

Would you accept a quirk that keeps the bit masked on Apple
106b:1802? I am happy to write and test that, or any other patch you
would like tried on this machine, and I can send lspci -vvv and the
journals.

Thanks,
Jason


On Mon, Oct 5, 2026 at 7:19 PM Bjorn Helgaas <helgaas@kernel.org> wrote:
>
> [+cc Alex]
>
> On Mon, Oct 05, 2026 at 01:53:55PM -0400, Perlow, Jason wrote:
> > Hi Lukas, Bjorn,
> >
> > Since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal
> > Errors"), first in v7.3-rc1, an Apple MacBookPro16,1 (T2 chip) loses
> > power 17 to 24 s after boot. v7.2.9 is fine; v7.3-rc5 and v7.3-rc6 are
> > not. Bisection lands on that commit, and v7.3-rc6 with only that
> > commit reverted no longer powers off. aer.c has not changed since, as
> > of v7.3-rc6-3, so I expect it is still unfixed; I did not find an
> > existing report.
>
> Alex reported something similar at
> https://bugzilla.kernel.org/show_bug.cgi?id=222095.
>
> Can you try the experiment mentioned there?  The patches Lukas
> mentioned don't apply cleanly on v7.3-rc1, but the conflict looks
> trivial?
>
> > Hardware
> > --------
> >
> > MacBookPro16,1 with an AMD Navi 14 dGPU and the Apple T2. The bisect
> > machine has a Core i7-9750H; a second unit of the same model, used for
> > cross-checks, has a Core i9-9980HK. I have not tested any other T2
> > model. The internal SSD, the T2 and the audio device are functions of
> > one Apple PCIe device behind root port 00:1b.0:
> >
> >   04:00.0 Apple ANS2 NVMe          [106b:2005]
> >   04:00.1 Apple T2 Bridge          [106b:1801]
> >   04:00.2 Apple T2 Secure Enclave  [106b:1802]
> >   04:00.3 Apple Audio Device       [106b:1803]
> >
> > Symptom
> > -------
> >
> > The machine powers off abruptly 17 to 24 s after boot. There is no
> > shutdown sequence; the previous boot's journal simply ends. Nothing is
> > logged beforehand: no AER message, no oops, no lockup, no thermal
> > event. The T2 controls power on these machines, so I assume (but
> > cannot show) that the T2/SMC removes power.
> >
> > Bisect
> > ------
> >
> > v7.2.9 good, v7.3-rc5 bad. Vanilla mainline trees, no out-of-tree
> > patches in the kernel image, CONFIG_PCIEAER=y,
> > CONFIG_ACPI_APEI_GHES=y, kernel messages captured over the network.
> > 19 steps, 4 skipped because those trees oops early for an unrelated
> > reason (the ones I looked at were in thunderbolt icm_probe at about
> > 5 s). Good boots were watched for 55.8 to 69.7 s; bad boots stopped
> > logging between 17.9 and 24.1 s. The complete history, every commit
> > tested and every result, is in Appendix B, and the exact method in
> > Appendix A. The result:
> >
> >   # first bad commit: [eddba19b8b5f] PCI/AER: Support Advisory
> >   # Non-Fatal Errors
> >
> > Revert test, v7.3-rc6, same config:
> >
> >   v7.3-rc6 unmodified:            powers off at 16.7 s
> >   only eddba19b8b5f reverted:     survives the full 64 s capture
> >
> > With that revert on top of 7.3.0-rc6 plus the out-of-tree t2linux
> > series (one kernel image, built once), a normal desktop runs on both
> > MacBookPro16,1 units: the second one has been up for more than 45
> > minutes, and the bisect machine ran sessions of 32 and 22 minutes.
> > For completeness: the bisect machine once lost power after about 10
> > minutes while I was manually switching the display mux and powering
> > off the AMD GPU, which I do not expect the firmware to support. I
> > have not established the cause and have no evidence either way on
> > whether it is related.
> >
> > Which devices are affected
> > --------------------------
> >
> > On v7.2.9, where the Advisory Non-Fatal Error bit is still masked, the
> > Correctable Error Status register has AdvNonFatalErr latched on
> > exactly these functions, with every Uncorrectable Error Status
> > register clear:
> >
> >   04:00.0 04:00.1 04:00.2 04:00.3  (the Apple device above)
> >   01:00.0 AMD Navi 10 XL PCIe switch upstream port [1002:1478]
> >
> > The pattern is identical on both MacBookPro16,1 units (same model, so
> > this says nothing about other T2 models). The Titan Ridge 4C bridges
> > and NHI on the same machine do not have the bit set.
> > These devices report the advisory bit without any matching
> > Uncorrectable Error status, which looks like the "non-compliant
> > products" case the commit message mentions.
> >
> > Only one of the two machines was used for the bisect and the
> > power-off tests above. I have not yet booted an unreverted 7.3 kernel
> > on the second one, so I cannot yet say that the power-off reproduces
> > there.
> >
> > Control: an ASUS ROG Zephyrus M15 GU502LV (i7-10750H, RTX 2060, no
> > T2) has the same bit latched on its NVIDIA TU106 functions
> > [10de:10f9, 10de:1ada, 10de:1adb] and on Titan Ridge 2C [8086:15e7,
> > 8086:15e8, 8086:15e9]. It ran a 7.3.0-rc6 build with the commit
> > applied (plus the t2linux series) for more than 15 hours without a
> > problem, and the same reverted kernel image as above also runs on it
> > normally. So unmasking the bit is not harmful in general; something
> > specific to the T2 platform is.
> >
> > What I do not know
> > ------------------
> >
> > Which device triggers the power-off, and why. My guess, and it is
> > only a guess: treating a possible Advisory Non-Fatal Error as
> > non-Advisory and recovering through the uncorrectable path resets or
> > disturbs a T2 function, and the T2 then powers the machine down. I
> > intend to build a diagnostic kernel that can leave the bit masked per
> > device to find out which one.
> >
> > Possibly related, different symptom: "PCI/portdev: Disable AER for
> > Titan Ridge 4C 2018" (Atharva Tiwari, January 2026) concerned AER
> > warnings on T2 iMacs.
> >
> > What I am asking
> > ----------------
> >
> > Which direction would you prefer: a revert, or a quirk that keeps
> > Advisory Non-Fatal Errors masked on the affected Apple functions (and
> > possibly the AMD switch port)? The t2linux project carries a revert
> > for now: https://github.com/t2linux/linux-t2-patches/pull/70
> >
> > I can test patches on real hardware and can provide full lspci -vvv
> > output, the complete bisect log and the captured boot logs.
> >
> > #regzbot introduced: eddba19b8b5f76d57424ee328a68fd495c5db857
> >
> > Thanks,
> > Jason Perlow
> >
> >
> > APPENDIX A - METHOD
> >
> > Machine: one MacBookPro16,1 for every boot below, running a Debian
> > trixie userland from its internal SSD. The distribution's 7.2.9 kernel
> > was the default boot entry between tests.
> >
> > Build: git bisect in a clone of torvalds/linux (git.kernel.org). Each
> > candidate was built on a separate x86-64 build host (16 threads, gcc
> > 15.2.0, binutils 2.46) with the same recipe:
> >
> >   cp bisect-trimmed.config .config
> >   make olddefconfig
> >   make -j10 bindeb-pkg LOCALVERSION=-bisN-vanilla-rc0 \
> >        KDEB_PKGVERSION=<release>-1
> >
> > bisect-trimmed.config is a 7.3.0-rc5 configuration trimmed to build
> > quickly (2068 options built in, 204 modules). It has CONFIG_PCIEAER=y,
> > CONFIG_PCIE_DPC=y, CONFIG_PCIEASPM=y, CONFIG_ACPI_APEI=y and
> > CONFIG_ACPI_APEI_GHES=y. The same file was used for all 19 steps and
> > for both v7.3-rc6 tests. No out-of-tree patches were applied to the
> > kernel. The release strings come from each tree's Makefile, so commits
> > on 7.2-based topic branches show as 7.2.0.
> >
> > Boot: the .deb packages were installed on the laptop and booted through
> > a one-shot rEFInd entry; the default entry stayed the stable kernel.
> > Kernel command line for every test boot:
> >
> >   console=tty0 ignore_loglevel keep_bootcon initcall_debug
> >   printk.time=1 log_buf_len=16M panic=0 fbcon=font:TER16x32
> >   systemd.show_status=1 systemd.unit=multi-user.target
> >   modprobe.blacklist=sbs,sbshc
> >   systemd.mask=ncz-usb2-rescan.service
> >   systemd.wants=ncz-netlog.service
> >
> > (plus the root= options). So there was no graphical session. sbs and
> > sbshc are the ACPI smart battery drivers; ncz-usb2-rescan is an
> > unrelated distribution boot workaround. Steps 13 to 19 and the two
> > v7.3-rc6 tests also had module_blacklist=thunderbolt,t2thunderbolt,
> > added after the early oopses (in thunderbolt icm_probe, at about 5 s)
> > had cost several skipped steps. Steps 1 to 12 ran with Thunderbolt
> > enabled.
> >
> > Capture: a small userspace unit (ncz-netlog.service) streams /dev/kmsg
> > and the journal over TCP to a second machine from early multi-user
> > boot, so captures begin at roughly 10 s of uptime. Each capture file
> > records the uptime of its last line.
> >
> > Verdicts:
> >   good: the machine kept running and logging past the point where bad
> >         kernels die (the cut always came before 25 s).
> >   bad:  the capture stops before 45 s, the machine stays unreachable
> >         for at least 60 s, and the next boot's journal shows the
> >         previous boot ending with no clean shutdown.
> >   skip: a kernel oops or panic in the capture or the previous boot's
> >         journal.
> >
> > After a cut the laptop does not restart by itself, so it was powered
> > on by hand and booted the stable kernel; the verdict was then confirmed
> > from the previous boot's journal.
> >
> > Caveats:
> >   - One machine was used for the bisect and for the power-off tests.
> >   - No desktop session was running.
> >   - sbs/sbshc were blocked on every boot, and Thunderbolt from step 13
> >     on.
> >   - The laptop was powered from, and networked through, a Thunderbolt
> >     dock during the bisect. An earlier test of a T2-patched
> >     7.3.0-rc5 kernel with the dock unplugged also lost power, but I
> >     did not repeat the bisect undocked.
> >   - Out-of-tree DKMS modules for the T2 hardware (t2smc, t2gmux,
> >     t2smp, t2thunderbolt) were installed for these kernels and may
> >     have been loaded, which would taint them. An earlier test with
> >     those modules blocked still lost power.
> >   - Step 12 ran 816 s before an unrelated oops, so it did not show
> >     the cut and was treated as a skip rather than as good; this does
> >     not change the result.
> >   - The device that triggers the cut is not identified.
> >
> > APPENDIX B - FULL BISECT HISTORY
> >
> > Good: v7.2.9 (5fce161649b4). Bad: v7.3-rc5 (72d3fcf802c4). The first
> > commit tested is the merge base, "Linux 7.2". Observed = uptime of the
> > last line captured. Results: 8 good, 7 bad, 4 skip.
> >
> >  #  commit        result observed   subject
> >  1  8d3ae59288f1  good   55.8 s     Linux 7.2
> >  2  56ea4e86832d  good   67.7 s     nstree: check listing permission
> >                                     before taking a namespace ref
> >  3  93e4b3076b5f  bad    24.1 s     Merge tag 'char-misc-7.3-rc1'
> >  4  21bd0802cd3f  good   60.1 s     Merge tag 'for-linus' (rdma)
> >  5  0b0e645ed2c8  bad    22.9 s     Merge tag 'auxdisplay-v7.3-1'
> >  6  e5f92606156a  good   69.5 s     Merge tag 'mm-nonmm-stable-2026-08-
> >                                     22-16-57'
> >  7  b6b019a1d9b9  good   69.2 s     Merge tag 'parisc-for-7.3-rc1'
> >  8  b130a2caf5d3  skip   oops 5.3 s Merge branch
> >                                     'pci/controller/tegra264'
> >  9  455b454c87bb  good   69.0 s     i3c: mipi-i3c-hci: Add support for
> >                                     AMD_PT I3C controller
> > 10  9bb52aa1972d  skip   oops 5.5 s Merge branch
> >                                     'pci/controller/dwc-meson'
> > 11  b4b07fb82b9e  skip   oops 5.3 s Merge branch 'pci/wake'
> > 12  651fb94aaf24  skip   oops 816 s alpha/PCI: Fix I/O port accessor
> >                                     argument order in
> >                                     pci_legacy_write()
> > 13  625ae0ff41e5  bad    19.9 s     Merge branch 'pci/dt-binding'
> > 14  9f91b2b716a0  bad    19.1 s     Merge branch 'pci/procfs'
> > 15  ea55835bc538  bad    23.3 s     Merge branch 'pci/dpc'
> > 16  d358e9ad15c2  bad    22.3 s     Merge branch 'pci/aer'
> > 17  8446e1147f65  good   69.7 s     PCI/AER: Deduplicate logging of
> >                                     Error Source Identification
> > 18  f141f74c45c6  good   67.6 s     PCI/AER: Move retrieval of FEP and
> >                                     TLP Log into helper
> > 19  eddba19b8b5f  bad    17.9 s     PCI/AER: Support Advisory
> >                                     Non-Fatal Errors
> >
> > Full hashes of the decisive steps:
> >   eddba19b8b5f76d57424ee328a68fd495c5db857  (first bad)
> >   f141f74c45c6f774eebdb7e45bd609be5122bfa8  (last good before it)
> >   8446e1147f65563d374ffa54dc3ba81adb1342c5
> >
> > The v7.3-rc6 tests (a90ee4305c4a) used the same recipe and command line:
> > unmodified, power off at 16.7 s; with only eddba19b8b5f reverted,
> > survived the full 64 s capture.
> >
> > --
> > Jason



-- 
Jason Perlow | Argonaut Media Communications

Voice/Text  (954) 242-3484 | jperlow@gmail.com | Blog: techbroiler.net
Read my Tech and Food Industry articles: https://linktr.ee/jperlow
Bluesky: https://bsky.app/profile/jperlow.bsky.social

Need to schedule a meeting with me? https://bit.ly/3y8P3Gp

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
  2026-10-06  0:17   ` Perlow, Jason
@ 2026-10-06  2:56     ` Lukas Wunner
  2026-10-06  8:31       ` Aditya Garg
  0 siblings, 1 reply; 6+ messages in thread
From: Lukas Wunner @ 2026-10-06  2:56 UTC (permalink / raw)
  To: Perlow, Jason
  Cc: Bjorn Helgaas, Bjorn Helgaas, linux-pci, linux-kernel,
	regressions, Aditya Garg, Alex Deucher

On Mon, Oct 05, 2026 at 08:17:12PM -0400, Perlow, Jason wrote:
> > > My guess, and it is
> > > only a guess: treating a possible Advisory Non-Fatal Error as
> > > non-Advisory and recovering through the uncorrectable path resets or
> > > disturbs a T2 function, and the T2 then powers the machine down.

Advisory Non-Fatal Errors aren't handled through the Uncorrectable
Error code path.  The kernel just reports and clears the errors.

And if the device has a driver and that driver implements the
->cor_error_detected callback() in struct pci_error_handlers,
that callback is invoked.  (Currently only the CXL driver
implements the callback.)

>   masked only on 04:00.2 (Apple T2 Secure Enclave, 106b:1802):
>       stays up (120 s)
[...]
> So here the trigger is unmasking Advisory Non-Fatal Errors on that
> one function.

That's quite unexpected.  Maybe there's a bug in the PCIe IP which
Apple used for the T2 Secure Enclave device (which I believe is a
fully-fledged arm64 CPU with a PCIe port in Endpoint mode).

Another explanation might be that there's a watchdog on the T2 device
which monitors its Config Space and triggers a poweroff whenever
some suspicious register mutation is detected.  I'm just guessing
here because this is really weird.

macOS likely never unmasks the bit and so they never saw this during
validation testing.  Is the T2 device visible on Windows?  If it is,
it seems Windows doesn't support and unmask Advisory Non-Fatal Errors
either.

> Would you accept a quirk that keeps the bit masked on Apple
> 106b:1802?
[...]
> I am happy to write and test that, or any other patch you
> would like tried on this machine, and I can send lspci -vvv and the
> journals.

A quirk does sound like the only possible solution.  However I'd
really like to see lspci and dmesg output with the bit unmasked
shortly before poweroff to get a full understanding.

Thanks!

Lukas

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot
  2026-10-06  2:56     ` Lukas Wunner
@ 2026-10-06  8:31       ` Aditya Garg
  0 siblings, 0 replies; 6+ messages in thread
From: Aditya Garg @ 2026-10-06  8:31 UTC (permalink / raw)
  To: Lukas Wunner, Perlow, Jason
  Cc: Bjorn Helgaas, Bjorn Helgaas, linux-pci, linux-kernel,
	regressions, Alex Deucher



On 06-10-2026 08:26 am, Lukas Wunner wrote:
> On Mon, Oct 05, 2026 at 08:17:12PM -0400, Perlow, Jason wrote:
>>>> My guess, and it is
>>>> only a guess: treating a possible Advisory Non-Fatal Error as
>>>> non-Advisory and recovering through the uncorrectable path resets or
>>>> disturbs a T2 function, and the T2 then powers the machine down.
> 
> Advisory Non-Fatal Errors aren't handled through the Uncorrectable
> Error code path.  The kernel just reports and clears the errors.
> 
> And if the device has a driver and that driver implements the
> ->cor_error_detected callback() in struct pci_error_handlers,
> that callback is invoked.  (Currently only the CXL driver
> implements the callback.)
> 
>>   masked only on 04:00.2 (Apple T2 Secure Enclave, 106b:1802):
>>       stays up (120 s)
> [...]
>> So here the trigger is unmasking Advisory Non-Fatal Errors on that
>> one function.
> 
> That's quite unexpected.  Maybe there's a bug in the PCIe IP which
> Apple used for the T2 Secure Enclave device (which I believe is a
> fully-fledged arm64 CPU with a PCIe port in Endpoint mode).
> 
> Another explanation might be that there's a watchdog on the T2 device
> which monitors its Config Space and triggers a poweroff whenever
> some suspicious register mutation is detected.  I'm just guessing
> here because this is really weird.
> 
> macOS likely never unmasks the bit and so they never saw this during
> validation testing.  Is the T2 device visible on Windows?  If it is,
> it seems Windows doesn't support and unmask Advisory Non-Fatal Errors
> either.

T2 device is visible on Windows as well as "System Device"

PS C:\Users\Aditya> Get-PnpDevice -Class 'System' -PresentOnly | Where-Object {$_.InstanceId -like "*PCI*"} | Format-Table -AutoSize FriendlyName, InstanceId

FriendlyName                                                                              InstanceId
------------                                                                              ----------
Intel(R) Xeon(R) E3 - 1200/1500 v5/6th Gen Intel(R) Core(TM) PCIe Controller (x4) - 1909  PCI\VEN_8086&DEV_1909&SUBSYS_72708086&REV_07\3&11583659&0&0A
Intel(R) PCI Express Root Port #1 - A338                                                  PCI\VEN_8086&DEV_A338&SUBSYS_72708086&REV_F0\3&11583659&0&E0
Intel(R) Serial IO UART Host Controller - A328                                            PCI\VEN_8086&DEV_A328&SUBSYS_72708086&REV_10\3&11583659&0&F0
AMD PCI Express Upstream Switch Port                                                      PCI\VEN_1002&DEV_1478&SUBSYS_00000000&REV_43\4&D48CA2C&0&0008
Intel(R) PCI Express Root Port #17 - A340                                                 PCI\VEN_8086&DEV_A340&SUBSYS_72708086&REV_F0\3&11583659&0&D8
Intel(R) Xeon(R) E3 - 1200/1500 v5/6th Gen Intel(R) Core(TM) PCIe Controller (x8) - 1905  PCI\VEN_8086&DEV_1905&SUBSYS_72708086&REV_07\3&11583659&0&09
Intel(R) LPC Controller/eSPI Controller - A313                                            PCI\VEN_8086&DEV_A313&SUBSYS_72708086&REV_10\3&11583659&0&F8
High Definition Audio Bus                                                                 PCI\VEN_1002&DEV_AB38&SUBSYS_AB381002&REV_00\6&16883A51&0&01000008
Intel(R) SPI (flash) Controller - A324                                                    PCI\VEN_8086&DEV_A324&SUBSYS_72708086&REV_10\3&11583659&0&FD
System Device                                                                             PCI\VEN_106B&DEV_1802&SUBSYS_1802106B&REV_01\4&3AC8FC3&0&02D8
Intel(R) Host Bridge/DRAM Registers - 3EC4                                                PCI\VEN_8086&DEV_3EC4&SUBSYS_72708086&REV_07\3&11583659&0&00
Intel(R) Management Engine Interface #1                                                   PCI\VEN_8086&DEV_A360&SUBSYS_72708086&REV_10\3&11583659&0&B0
AMD PCI Express Downstream Switch Port                                                    PCI\VEN_1002&DEV_1479&SUBSYS_14791002&REV_00\5&130BE3BB&0&000008
Intel(R) Thermal Subsystem - A379                                                         PCI\VEN_8086&DEV_A379&SUBSYS_72708086&REV_10\3&11583659&0&90
Apple USB Virtual Host Controller                                                         PCI\VEN_106B&DEV_1801&SUBSYS_1801106B&REV_01\4&3AC8FC3&0&01D8
PCI standard RAM Controller                                                               PCI\VEN_8086&DEV_A36F&SUBSYS_00000000&REV_10\3&11583659&0&A2
Intel(R) SMBus - A323                                                                     PCI\VEN_8086&DEV_A323&SUBSYS_72708086&REV_10\3&11583659&0&FC
Intel(R) Xeon(R) E3 - 1200/1500 v5/6th Gen Intel(R) Core(TM) PCIe Controller (x16) - 1901 PCI\VEN_8086&DEV_1901&SUBSYS_72708086&REV_07\3&11583659&0&08

> 
>> Would you accept a quirk that keeps the bit masked on Apple
>> 106b:1802?
> [...]
>> I am happy to write and test that, or any other patch you
>> would like tried on this machine, and I can send lspci -vvv and the
>> journals.
> 
> A quirk does sound like the only possible solution.  However I'd
> really like to see lspci and dmesg output with the bit unmasked
> shortly before poweroff to get a full understanding.
> 
> Thanks!
> 
> Lukas


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-10-06  8:32 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05 17:53 [REGRESSION] PCI/AER: MacBookPro16,1 powers off ~20 s after boot Perlow, Jason
2026-10-05 18:37 ` Perlow, Jason
2026-10-05 23:19 ` Bjorn Helgaas
2026-10-06  0:17   ` Perlow, Jason
2026-10-06  2:56     ` Lukas Wunner
2026-10-06  8:31       ` Aditya Garg

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®