Commit 79c168e2aa95 for kernel
commit 79c168e2aa957ed0054c93db17caba653b95e377
Author: Lukas Wunner <lukas@wunner.de>
Date: Thu Oct 8 14:26:00 2026 +0200
PCI/AER: Skip error recovery on false alarms
Alex is seeing a probe failure of the amdgpu driver after the Root Port
above an AMD Navi10 GPU has been reset. The reset was performed to recover
from a Firmware First reported Fatal Error.
However all status registers in the Root Port's AER Extended Capability are
blank, so apparently the platform firmware raised a false alarm.
The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support
Advisory Non-Fatal Errors"). It looks like enabling Advisory Non-Fatal
Errors causes code paths to be exercised in platform firmware which were
never validated before.
Skip error recovery on false alarms, i.e. if no unmasked errors were
actually signaled.
Note that this will also skip recovery if both the Status and Mask
registers are "all ones", as would be the case for inaccessible devices.
However that seems justified because it would imply either a hot-unplug
event or a Surprise Down Error further up in the hierarchy. Interfering
with recovery from that seems uncalled for.
Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors")
Reported-by: Alex Deucher <alexander.deucher@amd.com>
Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Tested-by: Alex Deucher <alexander.deucher@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de
diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index d8dcd238fda1..922a726a52a5 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -1360,8 +1360,10 @@ static DEFINE_KFIFO(aer_recover_ring, struct aer_recover_entry,
static void aer_recover_work_func(struct work_struct *work)
{
+ struct aer_capability_regs *regs;
struct aer_recover_entry entry;
struct pci_dev *pdev;
+ u32 err;
while (kfifo_get(&aer_recover_ring, &entry)) {
pdev = pci_get_domain_bus_and_slot(entry.domain, entry.bus,
@@ -1375,6 +1377,12 @@ static void aer_recover_work_func(struct work_struct *work)
}
pci_print_aer(pdev, entry.severity, entry.regs);
+ regs = entry.regs;
+ if (entry.severity == AER_CORRECTABLE)
+ err = regs->cor_status & ~regs->cor_mask;
+ else
+ err = regs->uncor_status & ~regs->uncor_mask;
+
/*
* Memory for aer_capability_regs(entry.regs) is being
* allocated from the ghes_estatus_pool to protect it from
@@ -1385,12 +1393,14 @@ static void aer_recover_work_func(struct work_struct *work)
ghes_estatus_pool_region_free((unsigned long)entry.regs,
sizeof(struct aer_capability_regs));
- if (entry.severity == AER_NONFATAL)
- pcie_do_recovery(pdev, pci_channel_io_normal,
- aer_root_reset);
- else if (entry.severity == AER_FATAL)
- pcie_do_recovery(pdev, pci_channel_io_frozen,
- aer_root_reset);
+ if (err) {
+ if (entry.severity == AER_NONFATAL)
+ pcie_do_recovery(pdev, pci_channel_io_normal,
+ aer_root_reset);
+ else if (entry.severity == AER_FATAL)
+ pcie_do_recovery(pdev, pci_channel_io_frozen,
+ aer_root_reset);
+ }
pci_dev_put(pdev);
}
}