Symptoms & scope
- Display output or GPU work stops and nvidia-smi cannot use the device.
- NVRM logs Xid 79 or explicitly reports that the GPU has fallen off the bus.
Relevant environment
NVIDIA GPUs driven by the NVIDIA kernel module; desktop, compute and passthrough systems can all lose PCIe accessibility.
Recognizable messages (synthetic examples)
NVRM: Xid (PCI:0000:01:00): 79, GPU has fallen off the bus.Investigate PCIe, power and platform evidence; a restart may recover access but cannot identify the cause.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- PCIe link instability, device power loss or platform firmware interactions are possible.
- Driver and virtualization defects are also possible; Xid 79 identifies lost access rather than a unique failed component.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read kernel events, with administrator access if needed; inspect the boot containing the failure.
journalctl -b -k --no-pager --grep='NVRM|AER|PCIe Bus Error'Interpret the result: Preceding uncorrected AER errors strengthen a PCIe-path investigation. Their absence does not prove a healthy link because logs may be incomplete.
Check 2
Replace 01:00.0 with the GPU address found in lspci. This reads PCI configuration; some detail needs administrator rights.
lspci -nnk -vv -s 01:00.0Interpret the result: Compare link capability and current status. A device missing after the incident is stronger evidence than a slow negotiated link alone.
Evidence-guided next steps
Inspect the powered-down PCIe path
If logs suggest lost PCIe access, power down normally and inspect GPU seating, risers and manufacturer-specified power connectors. A supported direct-slot comparison can isolate an intermediary path.
Precautions: Disconnect mains power and follow hardware instructions; never reseat a live GPU or mix modular PSU cables from different supplies.
Recovery / rollback: Return to the documented original layout only after inspecting it; retain the safer stable path if a riser is implicated.
Did this solution help you?
Collect evidence before another full test
If physical checks do not explain the loss, record BIOS version, driver version, topology and the first Xid/AER messages for NVIDIA or the system vendor. Compare one supported driver or firmware change during maintenance.
Precautions: Do not use an in-session GPU reset on the display GPU or an active compute workload. Firmware flashing requires the vendor’s recovery procedure.
Recovery / rollback: Keep the previous driver package and supported boot entry; restore them if the comparison worsens behavior.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- NVIDIA Xid catalog: Xid 79 (project or distribution documentation)
- NVIDIA: working with Xid errors (project or distribution documentation)