Symptoms & scope
- One application crashes or reports device lost.
- A GPU address fault may appear before a scheduler timeout.
Relevant environment
AMDGPU workloads where kernel logs name gfxhub or mmhub page faults, often alongside a process, VMID or PASID.
Recognizable messages (synthetic examples)
amdgpu 0000:03:00.0: [gfxhub0] no-retry page fault (src_id:0 ring:24 vmid:3 pasid:32777)This is GPU virtual-memory evidence; preserve adjacent status fields rather than conclude that physical memory is defective.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- An application or userspace driver may reference an invalid GPU mapping.
- Kernel mapping bugs or hardware instability remain possible; a GPU page fault is not a CPU RAM test result.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read the complete fault context with administrator journal access if required. Avoid collecting only the last device-lost line.
journalctl -b -k --no-pager --grep='amdgpu|gfxhub|mmhub'Interpret the result: Record fault address, read/write status and named process. Nearby process attribution narrows the triggering workload but does not prove that process is defective.
Check 2
Query the installed Vulkan runtime as the desktop user; requires the distribution’s vulkaninfo utility.
vulkaninfo --summaryInterpret the result: Record the selected device and driver version. If only a CPU renderer is listed, first investigate driver discovery rather than interpret this as the same hardware fault.
Evidence-guided next steps
Isolate the triggering application path
If only one application triggers faults, compare its supported stable release without optional overlays, mods or experimental rendering settings. Change one setting and keep the same scene or input.
Precautions: Back up application settings and saves. Reducing texture use may change the trigger but does not establish faulty VRAM.
Recovery / rollback: Restore the backed-up application profile after comparison; keep only changes with a reproducible benefit.
Did this solution help you?
Compare a supported graphics stack
If the failure appeared with a Mesa or kernel update, compare a retained supported stack or distribution snapshot and include the original fault fields in an upstream report. A matching reproducer is more useful than a list of random flags.
Precautions: Keep kernel, firmware and 32/64-bit userspace packages coherent. Do not mix libraries copied from unrelated releases.
Recovery / rollback: Return to the previous generation or package snapshot if the comparison introduces new failures.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- AMDGPU GPU debugging: page faults (project or distribution documentation)
- RADV driver documentation (project or distribution documentation)