Kernel & System Stability

CPU hard lockup detected by a watchdog

A hard-lockup report means CPU interrupt heartbeats stopped; preserve watchdog traces and hardware context without assuming it is a GPU hang.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • The system stops responding to normal interrupts or restarts by panic policy.
  • Kernel reports Watchdog detected hard LOCKUP on a CPU.

Relevant environment

Kernels with NMI/perf or another supported hard-lockup detector. A silent freeze without a watchdog report is not classified by this entry.

Recognizable messages (synthetic examples)
NMI watchdog: Watchdog detected hard LOCKUP on cpu 1

Inspect the NMI trace and earlier hardware context; absence of a detector report cannot exclude a silent system lockup.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • A kernel path may loop with interrupts disabled or stall CPU progress.
  • Hardware instability or platform problems can also stop progress; the watchdog reports detection, not the uniquely faulty part.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read the previous retained boot after recovery, with administrator access if needed; use -b when the same boot is still accessible.

journalctl -b -1 -k --no-pager --grep='hard LOCKUP|watchdog|Hardware Error|Call Trace|RIP:'

Interpret the result: Preserve the first watchdog trace and any preceding hardware error. Empty prior logs can mean retention or write-out failed during the lockup.

Check 2

Read available detector/panic policy without setting it; kernel build and architecture can omit these keys.

sysctl kernel.nmi_watchdog kernel.hardlockup_panic

Interpret the result: A disabled or unavailable detector explains missing watchdog evidence, not a healthy machine. Panic configuration can explain an automatic reboot after detection.

Evidence-guided next steps

Compare at supported platform defaults

If tuning preceded lockups, restore documented CPU/memory defaults and compare only a short non-destructive trigger. If the same stack appears after a kernel update, compare an already installed supported earlier kernel.

Precautions: Back up critical data and avoid repeated hard power cycles. Changing several firmware and kernel options together destroys useful comparison evidence.

Recovery / rollback: Boot the retained original supported entry; restore only recorded stable hardware settings rather than known-triggering tuning.

Did this solution help you?

Share this solution#

Plan a supported crash capture

If repeated hard lockups lose all logs, prepare the distribution’s supported kdump or external-console capture during maintenance and report the earliest trace. Verify the capture configuration without deliberately crashing a production session.

Precautions: Crash dumps can contain sensitive memory and reserve RAM. Kdump is not guaranteed to run when hardware or every CPU has lost progress.

Recovery / rollback: Restore the previous crash-capture configuration and boot memory reservation if the setup prevents normal operation.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.