Graphics, GPU & Display

AMDGPU reports a GPU virtual-memory page fault

A GPU page fault identifies an invalid GPU address access; preserve the process and fault details before blaming physical RAM or video memory.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • One application crashes or reports device lost.
  • A GPU address fault may appear before a scheduler timeout.

Relevant environment

AMDGPU workloads where kernel logs name gfxhub or mmhub page faults, often alongside a process, VMID or PASID.

Recognizable messages (synthetic examples)
amdgpu 0000:03:00.0: [gfxhub0] no-retry page fault (src_id:0 ring:24 vmid:3 pasid:32777)

This is GPU virtual-memory evidence; preserve adjacent status fields rather than conclude that physical memory is defective.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • An application or userspace driver may reference an invalid GPU mapping.
  • Kernel mapping bugs or hardware instability remain possible; a GPU page fault is not a CPU RAM test result.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read the complete fault context with administrator journal access if required. Avoid collecting only the last device-lost line.

journalctl -b -k --no-pager --grep='amdgpu|gfxhub|mmhub'

Interpret the result: Record fault address, read/write status and named process. Nearby process attribution narrows the triggering workload but does not prove that process is defective.

Check 2

Query the installed Vulkan runtime as the desktop user; requires the distribution’s vulkaninfo utility.

vulkaninfo --summary

Interpret the result: Record the selected device and driver version. If only a CPU renderer is listed, first investigate driver discovery rather than interpret this as the same hardware fault.

Evidence-guided next steps

Isolate the triggering application path

If only one application triggers faults, compare its supported stable release without optional overlays, mods or experimental rendering settings. Change one setting and keep the same scene or input.

Precautions: Back up application settings and saves. Reducing texture use may change the trigger but does not establish faulty VRAM.

Recovery / rollback: Restore the backed-up application profile after comparison; keep only changes with a reproducible benefit.

Did this solution help you?

Share this solution#

Compare a supported graphics stack

If the failure appeared with a Mesa or kernel update, compare a retained supported stack or distribution snapshot and include the original fault fields in an upstream report. A matching reproducer is more useful than a list of random flags.

Precautions: Keep kernel, firmware and 32/64-bit userspace packages coherent. Do not mix libraries copied from unrelated releases.

Recovery / rollback: Return to the previous generation or package snapshot if the comparison introduces new failures.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.