Storage & Filesystems

NVMe I/O timeouts and controller resets

Trace NVMe resets to controller health, PCIe evidence and kernel history instead of interpreting an I/O timeout as proof of one defective SSD.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • File access stops while nvme logs reset or abort messages.
  • An NVMe namespace disappears until reboot.

Relevant environment

PCIe NVMe storage using the Linux nvme driver; NVMe over Fabrics requires additional transport checks.

Recognizable messages (synthetic examples)
nvme nvme0: I/O 42 QID 2 timeout, reset controller

Correlate health and PCIe events; this does not by itself prove a media fault.

nvme nvme0: Device not ready; aborting reset, CSTS=0x1

Stop unnecessary writes and preserve data; firmware, power and PCIe availability need checking.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • A stuck controller, firmware defect, thermal condition, PCIe fault or driver regression can prevent completion.
  • A reset following idle or resume may need a separate power-management comparison; heavy-load resets alone do not identify APST.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read kernel events with journal privileges if necessary; retain NVMe, PCIe/AER and filesystem lines around the incident.

journalctl -k -b --no-pager -n 400

Interpret the result: A reset is recovery after a command problem. Controller not-ready or all-ones status is a stronger availability fault; absence of AER does not exclude PCIe issues.

Check 2

Replace /dev/nvme0 with the controller corresponding to the affected namespace. Read health and error records; no test is started.

sudo smartctl -x /dev/nvme0

Interpret the result: Media errors, critical-warning bits or high recorded temperature provide specific leads. An error-log count can include unsupported admin commands and is not automatically media damage.

Evidence-guided next steps

Compare a supported kernel and SSD firmware

When the first failure follows a kernel update, boot an installed known-good kernel with the same workload and no other tuning. If resets occur across kernels, check the exact SSD model against vendor firmware advisories before considering an update.

Precautions: Secure data first and keep a bootable alternative. Do not reset a mounted root controller manually as a diagnostic.

Recovery / rollback: Select the previous boot entry if the comparison fails. Firmware updates may not be reversible, so backup-based recovery remains necessary.

Did this solution help you?

Share this solution#

Check the physical NVMe path

If failures coincide with AER reports or heat, power down and check SSD seating, heatsink contact and adapter placement. Compare one supported slot or adapter at a time and measure whether the same errors recur.

Precautions: Avoid benchmarks on storage already returning uncorrectable errors. Slot changes can affect boot order and PCIe lane sharing.

Recovery / rollback: Restore the recorded hardware arrangement and firmware boot order while powered off if detection worsens.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.