Symptoms & scope
- A PCI device stops responding while IOMMU faults accumulate.
- Kernel logs DMA Read/Write faults or AMD IO_PAGE_FAULT events.
Relevant environment
Intel VT-d/DMAR or AMD-Vi systems using translated DMA, including virtual-machine device assignment.
Recognizable messages (synthetic examples)
DMAR: [DMA Read NO_PASID] Request device [03:00.0] fault addr 0x1000 [fault reason 0x06] PTE Read access is not setUse device address, DMA direction and reason to investigate the host mapping; do not classify this as an application CPU page fault.
AMD-Vi: Event logged [IO_PAGE_FAULT device=03:00.0 domain=0x000a address=0x1000 flags=0x0000]Check the device’s active driver and domain; the record demonstrates isolation enforcement, not that disabling isolation is a fix.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- A driver may submit DMA outside a valid mapping or use a stale mapping.
- Firmware reservations, reset behavior or incorrect VM assignment can conflict with DMA translation; a GPU virtual-memory fault is a separate mechanism.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read device address, fault reason and DMA direction; journal access may require administrator rights.
journalctl -b -k --no-pager --grep='DMAR:|AMD-Vi:|IO_PAGE_FAULT'Interpret the result: Repeated faults from one device narrow the affected driver/domain. Early firmware messages before driver binding need separate interpretation from runtime DMA faults.
Check 2
Replace 03:00.0 with the address in the fault record. Read binding without changing it.
lspci -nnk -s 03:00.0Interpret the result: Confirm whether the host driver or vfio-pci owns the device. Binding unexpected for the intended VM setup warrants checking its configuration before launching another guest.
Evidence-guided next steps
Compare a fix for the identified device driver
If a runtime fault repeatedly identifies the same normally host-driven device, use a supported kernel or device firmware fix matching that device and fault reason. Temporarily stop its optional workload while collecting evidence.
Precautions: Do not disable the IOMMU globally to silence faults; it provides DMA isolation and may be essential to the VM configuration.
Recovery / rollback: Restore the earlier supported kernel and stopped workload configuration if the comparison introduces new faults.
Did this solution help you?
Correct a documented passthrough assignment
If faults begin on guest start or stop, review the hypervisor’s supported device group, host binding and reset requirements. Correct the VM definition with guests shut down, then compare a single controlled lifecycle.
Precautions: Do not force ACS overrides, unsafe interrupt settings or assignment of a device needed by the host. Preserve the original VM definition.
Recovery / rollback: Restore the saved VM definition and host driver configuration with guests stopped if the corrected assignment cannot initialize safely.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- Kernel x86 IOMMU support (project or distribution documentation)
- AMD IOMMU fault reporting implementation (upstream implementation; behavior can vary by version)
- Intel DMAR implementation (upstream implementation; behavior can vary by version)