Kernel & System Stability

PCIe AER reports corrected or uncorrected errors

PCIe AER records describe transaction or link errors with different severities; map the reporting port to its devices before blaming an endpoint.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Kernel repeatedly prints PCIe Bus Error messages.
  • An endpoint may reset or disappear, while corrected errors can occur without visible interruption.

Relevant environment

PCI Express systems where firmware grants Linux AER control; physical devices, bridges and root ports can all appear in reports.

Recognizable messages (synthetic examples)
pcieport 0000:00:03.1: PCIe Bus Error: severity=Corrected, type=Physical Layer, (Receiver ID)

Monitor rate and symptoms; corrected does not mean harmless forever, and a single report does not identify the failing component.

pcieport 0000:00:03.1: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Transaction Layer, (Requester ID)

Inspect Fatal versus Non-Fatal and the recovery result; preserve full requester/status fields before any reset.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • Signal integrity, risers, connectors, endpoint firmware or platform firmware may contribute.
  • Driver transaction behavior can also trigger reports; the root port that reports the event need not be the defective part.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read AER severity, source IDs and recovery events; administrator journal access may be required.

journalctl -b -k --no-pager --grep='AER|PCIe Bus Error|DPC:'

Interpret the result: Corrected means the protocol recovered that error; Uncorrectable Fatal is more serious. Compare event rate and device symptoms instead of treating all AER messages alike.

Check 2

Read the PCI bridge topology without changing devices. Match bus addresses with the detailed AER record.

lspci -t

Interpret the result: Locate the reporting upstream port and its endpoints. A shared bridge can explain several affected devices without proving every endpoint is defective.

Evidence-guided next steps

Apply a targeted firmware or kernel fix

If the errors started after an update or match a vendor erratum, compare the relevant supported kernel or device/platform firmware fix during maintenance. Include topology and full status fields in a vendor report.

Precautions: Do not use pci=noaer as a repair; it hides reporting. Do not force unsupported PCIe generations or disable protection across all devices without device-specific evidence.

Recovery / rollback: Boot the earlier supported kernel or use the vendor’s documented firmware recovery procedure if the update worsens access.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.