Kernel & System Stability

CPU soft lockup detected by the kernel watchdog

A soft lockup means the watchdog task did not run in time; inspect its CPU trace and distinguish a guest pause from a repeatable kernel stall.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • The desktop or service becomes unresponsive for an extended interval.
  • The kernel names a stuck CPU and task in a soft lockup warning.

Relevant environment

Linux with the soft-lockup detector enabled, on bare metal or virtual machines. The threshold can be configured and is not a benchmark.

Recognizable messages (synthetic examples)
watchdog: BUG: soft lockup - CPU#2 stuck for 26s! [worker:440]

Inspect execution traces and virtualization context; a soft lockup is different from a GPU ring timeout or an interrupt hard lockup.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • A long non-preemptible kernel path, driver loop or scheduling starvation may delay the watchdog.
  • A paused or severely descheduled virtual machine and altered watchdog thresholds can produce misleading timing evidence.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read watchdog and stack lines; administrator journal access may be required. Also retain the surrounding unfiltered log when reporting.

journalctl -b -k --no-pager --grep='soft lockup|watchdog|Call Trace|RIP:'

Interpret the result: Repeated stalls with the same execution path are more actionable than isolated warnings after a VM pause. The named task may be interrupted rather than responsible.

Check 2

Read threshold and panic policy without assigning values. Availability depends on the kernel configuration.

sysctl kernel.watchdog_thresh kernel.softlockup_panic

Interpret the result: An unusually low threshold can explain earlier reporting, but raising it would hide rather than fix a recurring stall. A panic policy explains automatic restarts.

Evidence-guided next steps

Reduce identified tracing or console overhead

If the incident occurs only with heavy tracing or very verbose slow-console logging, disable that optional instrumentation through its documented configuration and repeat the minimal workload. Compare the stack and elapsed pause.

Precautions: Keep normal kernel error logging and watchdog detection active. Do not stop essential monitoring solely because the warning is inconvenient.

Recovery / rollback: Restore the recorded tracing or console profile if it is needed for a targeted follow-up capture.

Did this solution help you?

Share this solution#

Compare the kernel or VM scheduling context

For a repeated bare-metal trace, compare a supported earlier kernel and a matching packaged fix. For warnings only after guest pause or migration, investigate the host timeline and scheduling before changing guest kernel policy.

Precautions: Keep workloads and hardware settings constant; avoid repeated destructive stress when the machine cannot recover cleanly.

Recovery / rollback: Return to the original supported boot entry or restore the prior VM scheduling profile if the comparison worsens behavior.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.