← FIELD NOTES & ARTIFACTS

FIELD NOTE #008 · DOCUMENTED 5 OCT 2026 · INVESTIGATION STILL OPEN

THE RX 9060 XT HARD-LOCK
THAT SURVIVED EVERYTHING

Kernel switches, Mesa rollbacks, RADV flags, power limits and reset experiments kept failing. Then the PCIe path changed.

RX 9060 XT 16 GBNAVI44 / RDNA4NIXOSAMDGPUMESA 26.2.4PCIE 4.0 x16

Short version

Cthulhu's PowerColor Radeon RX 9060 XT could run light games for hours, yet heavy titles repeatedly produced stalls, black screens and sometimes a full GPU hard-lock with the fans at 100 percent.

After a long software investigation, Mesa 26.2.4 alone still crashed on the old PCIe setup within seconds. I then simplified the hardware path in one maintenance window: the GPU went directly into the motherboard, the Corsair riser came out, the Intel 10 GbE card came out, and the BIOS link setting returned from forced Gen3 to Gen4/Auto.

The external GPU link changed from 8 GT/s x8 to 16 GT/s x16. In that state, S.T.A.L.K.E.R. 2 ran for about 25 minutes with no traversal stalls or black screen, followed by roughly 20 minutes of Arma 3 with no problem at all.

Important:

This is the strongest lead so far, not a final root-cause proof. Several hardware variables changed together. The current evidence says the PCIe/topology change matters; it does not yet tell us whether the decisive part was the riser, lane allocation, Gen3 versus Gen4, the removed 10 GbE card, or an interaction between them. Mesa 26.2.4 may also contribute, but it was not sufficient by itself on the old topology.

The machine and the symptom

The workstation is Cthulhu: Ryzen 7 5800X3D, PowerColor Radeon RX 9060 XT 16 GiB, Gigabyte B550 VISION D, NixOS, Plasma/Wayland and two displays. The GPU identifies as Navi44 / RDNA4 / GFX1200, PCI ID 1002:7590, subsystem 148c:2437.

The failure was never one clean bug. That distinction turned out to matter.

Traversal stalls

S.T.A.L.K.E.R. 2 could stand still, shoot and reload normally, but moving through the world produced one- to several-second freezes. During captured stalls there was often no amdgpu timeout in the kernel log.

Black screen, audio alive

Some runs lost the display while audio continued and the GPU fans stayed normal. That looked more like a display/presentation or render wedge than the catastrophic failure.

Black screen, fans at 100%

The nasty one. Heavy games could make both monitors go black, the GPU fans jump to full speed and the entire machine require a hard reboot.

Arma 3 as the brutal trigger

Arma became the most reproducible test: joining a game could trigger the hard-lock almost immediately.

The kernel finally left a fingerprint

Some crashes died too hard to leave anything useful. Others produced a much better signature. The interesting part was not just that the graphics ring timed out; recovery itself also failed:

amdgpu: ring gfx_0.0.0 timeout
MES(1) failed to respond to msg=RESET
failed to reset legacy queue
reset via MES failed and try pipe reset -110
The CPFW hasn't support pipe reset yet.
GPU reset begin!
amdgpu: device lost from bus!
GPU Recovery Failed: -19

That changed the working model. The MES failure could be a recovery failure after the initiating hang, not necessarily the original cause.

Firmware inventory instead of guessing

Before blaming “old firmware” in the abstract, I read what the card had actually loaded:

ME   feature 29  firmware 0x00000c12
PFP  feature 29  firmware 0x00000c76
MEC  feature 29  firmware 0x00000d7a
SMC  firmware 0x00664700 (102.71.0)
DMCUB firmware 0x0a004b00
MES  feature 1   firmware 0x00000093

That was useful because it stopped the investigation from turning into “install a newer blob and hope.” The loaded SMU was not obviously obsolete, while newer upstream display firmware existed and remained relevant mainly to the normal-fan black-screen class.

The terminal became the lab notebook

The most useful part of the process was keeping the commands boring and repeatable. One read-only check at a time made it much easier to tell whether a new theory had actually changed anything.

journalctl -b -1 -k --no-pager | grep -Ei 'amdgpu|ring|timeout|reset|mes|device lost|gpu recovery|gpuvm|fault|pcie|aer' | tail -n 100

sudo cat /sys/kernel/debug/dri/0000:55:00.0/amdgpu_firmware_info

nix-shell -p pciutils --run 'sudo lspci -s 00:03.1 -vv'

vulkaninfo --summary | grep -E 'driverName|driverInfo'

Those four checks kept answering four different questions: what failed, what firmware was loaded, how PCIe was actually negotiated, and which Mesa driver a fresh process really saw. That separation prevented several false conclusions.

What we tested — and what did not fix it

The rule was simple: change one software variable where possible, reproduce the same workload, read the actual error, and do not declare victory because one run happened to last longer.

Kernels

Linux 7.2.8 remained the preferred baseline. Linux 6.18.54 was worse: S.T.A.L.K.E.R. could black-screen with 100% fans immediately. Linux 7.3-rc6, despite relevant GFX12 fixes, also black-screened almost immediately in the first S.T.A.L.K.E.R. movement test.

Older Mesa

Mesa 25.1.7 was not a known-good fallback. At a location where 26.1.8 usually stalled, 25.1.7 produced a persistent black screen with audio continuing.

Scheduler and power knobs

amdgpu.uni_mes=0, amdgpu.async_gfx_ring=0, GFXOFF inhibition, RADV synchronous shaders and GPU power caps from 112 W through the normal 160 W did not remove the problem.

Rendering paths

Arma also crashed with PROTON_USE_WINED3D=1, so DXVK/Vulkan was not required for the catastrophic class. RADV_DEBUG=nodcc did not save S.T.A.L.K.E.R. either.

Storage

The game was moved from the SATA game SSDs to NVMe. The traversal stalls remained. The Samsung SATA drives do have real transport errors of their own, but they are a separate maintenance problem, not a complete explanation for the GPU hard-lock.

Specific upstream workarounds

A targeted GFX12 SDMA/DCC kernel workaround was backported and tested. It did not change the S.T.A.L.K.E.R. stalls. Steam shader pre-caching was also tested and did not prevent another hard-lock.

ASPM looked promising — until we measured it

A very similar PowerColor RX 9060 XT case had reportedly become stable with global PCIe ASPM disabled. The subsystem ID even matched ours, so this deserved a proper check.

Instead of adding another boot parameter blindly, every relevant hop was inspected. Root port, upstream switch, downstream switch and GPU all reported:

LnkCtl: ASPM Disabled

The L1 substates were also disabled. That made pcie_aspm=off a poor candidate for Cthulhu: the relevant path was already running with ASPM off.

Then Mesa 26.2.4 fought back in a very NixOS way

The first attempt to pull Mesa 26.2.4 directly from a newer nixpkgs snapshot produced a clean example of why mixing package sets can be dangerous:

libm.so.6: version 'GLIBC_2.44' not found
required by ... libLLVM.so.21.1
loader_icd_scan: Failed loading ... libvulkan_radeon.so

No GPU mystery there — just an ABI mismatch. That experiment was rolled back immediately.

Instead, Mesa 26.2.4 was rebuilt from its upstream source using Cthulhu's current NixOS package set and toolchain, for both 64-bit and 32-bit graphics. The result loaded correctly:

driverName = radv
driverInfo = Mesa 26.2.4

driverName = llvmpipe
driverInfo = Mesa 26.2.4 (LLVM 21.1.8)

Then S.T.A.L.K.E.R. entered the game and still hard-locked after only a few seconds on the old PCIe arrangement. That was an important negative result: Mesa 26.2.4 by itself was not the fix.

The hardware simplification

At that point the software tree had been shaken hard enough. Instead of doing another dozen reboots, the next test deliberately simplified the physical path as much as possible in one shutdown:

GPU: Corsair riser  → removed
GPU: indirect mount → direct motherboard slot
Intel dual-port 10 GbE → removed for the test
BIOS PCIe mode: forced Gen3 → Gen4 / Auto
software: Linux 7.2.8 + Mesa 26.2.4 → unchanged

This was intentionally not a one-variable experiment. Its purpose was to answer a different question: can the machine become stable when the PCIe environment is stripped down?

Before the change, the outer GPU link had been:

LnkSta: Speed 8GT/s (downgraded), Width x8 (downgraded)

After the GPU went directly into the board:

00:03.1 root port
LnkCap: Speed 16GT/s, Width x16
LnkSta: Speed 16GT/s, Width x16

53:00.0 GPU upstream port
LnkSta: Speed 16GT/s (downgraded), Width x16

The card still exposes its own AMD switch functions internally, but the motherboard-facing path had changed from PCIe 3.0 x8 to PCIe 4.0 x16.

The first result that actually felt different

S.T.A.L.K.E.R. 2 had previously been able to fail within seconds. In the simplified PCIe 4.0 x16 configuration it ran for about 25 minutes with no traversal stalls, no black screen and no 100%-fan hard-lock.

Arma 3 had been our nastiest reproducer. A quick test turned into roughly 20 minutes of normal play with no problem at all.

This was the first genuinely different result:

Not “the crash took a little longer.” Not “the power limit changed the symptom.” The two workloads that had repeatedly exposed the problem suddenly behaved normally.

One more PCIe detail: sticky lane errors

After booting the direct Gen4 x16 setup, the root port initially showed sticky lane-error bits on lanes 3, 7, 11 and 15:

LnkSta: Speed 16GT/s, Width x16
LaneErrStat: LaneErr at lane: 3 7 11 15

Those bits can be left behind by link training, so they were cleared to establish a clean baseline. Read-back then showed:

00000000

The useful test is not whether training ever set a bit. It is whether new lane errors appear after sustained gaming load. That check is still pending.

What I think now

The strongest evidence currently points toward the PCIe path / lane topology / riser environment being part of the real problem. There is historical support too: on the previous Arch installation, removing the riser had already stopped a separate WoW Classic black-screen behavior.

I do not think the Intel 10 GbE card is the leading suspect. The NAS at the other end was not even active during these tests, so network traffic was irrelevant. Still, the card can affect PCIe lane allocation simply by being installed, so it is not scientifically eliminated yet.

Windows remaining stable on the old physical setup is also important. It argues against the simplistic conclusion “the riser is broken.” A marginal link, lane-layout interaction or reset/power-state edge case can be tolerated differently by Windows and Linux. The Linux GFX12 stack may simply be much less forgiving when the GPU stops responding.

Mesa 26.2.4 may be helping as well, but the evidence is asymmetric: it failed quickly on the old topology, while the new topology has not yet been tested with Mesa 26.1.8. So the honest statement is PCIe/topology change: strongly implicated; Mesa 26.2.4: plausible contributor; exact root cause: still open.

What comes next

The next step is deliberately boring: leave the stable state alone and play longer sessions. After that, read LaneErrStat again. If the system remains clean, the Intel 10 GbE card can be reinstalled first while keeping the GPU direct at Gen4 x16. If that stays stable, the NIC drops further down the list. The riser and x8/bifurcation path can then be isolated separately.

No more “change five kernel parameters and reboot” roulette. We finally have a useful reference state.

Lesson

The biggest mistake would have been treating every black screen as the same bug.

One class looked like a display/render wedge. Another produced a graphics-ring timeout followed by a failed MES recovery and device loss. The traversal stalls could happen without either.

Once those were separated, the investigation became much more useful: measure the layer that failed, preserve the known-good variables, and only then change the next one.

And after what felt like reboot number 224, sometimes the most valuable command is not another kernel parameter.

Sometimes it is a screwdriver.

← RETURN TO FIELD NOTES & ARTIFACTS