Symptoms & scope
- Kernel repeatedly prints PCIe Bus Error messages.
- An endpoint may reset or disappear, while corrected errors can occur without visible interruption.
Relevant environment
PCI Express systems where firmware grants Linux AER control; physical devices, bridges and root ports can all appear in reports.
Recognizable messages (synthetic examples)
pcieport 0000:00:03.1: PCIe Bus Error: severity=Corrected, type=Physical Layer, (Receiver ID)Monitor rate and symptoms; corrected does not mean harmless forever, and a single report does not identify the failing component.
pcieport 0000:00:03.1: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Transaction Layer, (Requester ID)Inspect Fatal versus Non-Fatal and the recovery result; preserve full requester/status fields before any reset.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- Signal integrity, risers, connectors, endpoint firmware or platform firmware may contribute.
- Driver transaction behavior can also trigger reports; the root port that reports the event need not be the defective part.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read AER severity, source IDs and recovery events; administrator journal access may be required.
journalctl -b -k --no-pager --grep='AER|PCIe Bus Error|DPC:'Interpret the result: Corrected means the protocol recovered that error; Uncorrectable Fatal is more serious. Compare event rate and device symptoms instead of treating all AER messages alike.
Check 2
Read the PCI bridge topology without changing devices. Match bus addresses with the detailed AER record.
lspci -tInterpret the result: Locate the reporting upstream port and its endpoints. A shared bridge can explain several affected devices without proving every endpoint is defective.
Evidence-guided next steps
Inspect the identified link path while powered off
If repeated link errors map to one physical path, shut down normally and inspect seating, supported cables and risers using manufacturer instructions. Compare a supported direct connection with the same workload if possible.
Precautions: Disconnect mains power before reseating. A lower negotiated link speed alone does not prove faulty wiring; power-saving and device capabilities also matter.
Recovery / rollback: Return to the recorded supported layout once inspected; keep the stable direct path if an intermediary is implicated.
Did this solution help you?
Apply a targeted firmware or kernel fix
If the errors started after an update or match a vendor erratum, compare the relevant supported kernel or device/platform firmware fix during maintenance. Include topology and full status fields in a vendor report.
Precautions: Do not use pci=noaer as a repair; it hides reporting. Do not force unsupported PCIe generations or disable protection across all devices without device-specific evidence.
Recovery / rollback: Boot the earlier supported kernel or use the vendor’s documented firmware recovery procedure if the update worsens access.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- Kernel PCIe AER severity and recovery (project or distribution documentation)
- Kernel PCI error recovery model (project or distribution documentation)