Symptoms & scope
- A game stops drawing while sound or SSH may continue.
- The kernel reports a ring timeout followed by recovery attempts.
Relevant environment
AMD GPUs using the amdgpu kernel driver; OpenGL, Vulkan or compute workloads. A black screen without a retained log needs separate investigation.
Recognizable messages (synthetic examples)
amdgpu 0000:03:00.0: ring gfx_0.0.0 timeout, signaled seq=42, emitted seq=43Investigate the first failure and recovery outcome; neither the queue name nor this line identifies the root cause.
amdgpu 0000:03:00.0: GPU Recovery Failed: -110Save evidence from another console if available; restarting affected applications may be insufficient and a controlled reboot may be needed.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- An application, Mesa or kernel regression may submit work that stops completing.
- Firmware, unstable clocks, power delivery or PCIe errors can produce similar symptoms; the timeout does not distinguish them.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read this boot’s kernel events; journal access may require an administrator. Use -b -1 after a reboot if that boot was retained.
journalctl -b -k --no-pager --grep='amdgpu|AER|PCIe Bus Error'Interpret the result: The first ring error and any preceding PCIe event matter more than the final reset. A successful reset does not prove the underlying fault is fixed.
Check 2
List PCI devices and their active drivers without changing binding; locate the AMD display controller.
lspci -nnkInterpret the result: Kernel driver in use should identify amdgpu for a supported device. Record the PCI address for the report; installed modules alone do not establish binding.
Evidence-guided next steps
Compare a retained kernel
If the fault began after an update, boot an already installed earlier supported kernel and repeat only the shortest reliable trigger. Keep the workload and Mesa version constant to isolate the kernel change.
Precautions: Save work before testing; keep both boot entries and avoid repeated hard locks. A single stable run is weak evidence.
Recovery / rollback: Choose the original kernel entry at the next boot; remove a temporary default change if one was made.
Did this solution help you?
Return GPU tuning to defaults
If overclocking, undervolting or custom driver parameters were added, restore the recorded defaults one change at a time. If a minimal reproducer still fails, report complete pre-reset logs and versions upstream.
Precautions: Do not extend scheduler timeouts merely to hide the error or disable recovery on a daily workstation. Never alter voltage without a documented device-specific procedure.
Recovery / rollback: Reapply only a previously recorded setting that is necessary and demonstrably stable; retain a boot entry without experimental parameters.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- AMDGPU module parameters: scheduler timeout and recovery (project or distribution documentation)
- AMDGPU scheduler timeout implementation (upstream implementation; behavior can vary by version)