Kernel & System Stability

RCU grace-period progress stalls

RCU stall warnings concern delayed kernel synchronization; compare CPU traces and timer progress before attributing the failure to RCU itself.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Kernel prints detected stalls on CPUs/tasks while responsiveness degrades.
  • Repeated reports may show the same CPU stack or a starved RCU thread.

Relevant environment

Linux using RCU stall detection. Stalls can arise in other code, under real-time scheduling or with heavy instrumentation.

Recognizable messages (synthetic examples)
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:

Compare the subsequent CPU/task traces; a detecting CPU is not necessarily the stalled CPU.

rcu: rcu_preempt kthread starved for 24000 jiffies! g700 f0x0 RCU_GP_WAIT_FQS(3)

Inspect real-time scheduling and timer context; this is more specific than treating every RCU warning as a broken synchronization implementation.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • Code may fail to reach a quiescent state, disable interrupts too long or starve an RCU thread.
  • Timer faults, verbose console output and tracing overhead can resemble an RCU implementation problem.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read RCU and CPU trace context; administrator journal access may be required. Preserve repeated traces, not only a single headline.

journalctl -b -k --no-pager --grep='rcu|RCU|Call Trace|RIP:'

Interpret the result: Compare stable stack frames, softirq counts and timer messages between reports. The CPU detecting the stall can differ from the CPU causing it.

Check 2

Read this kernel’s RCU stall threshold if the parameter exists; do not change it during evidence collection.

cat /sys/module/rcupdate/parameters/rcu_cpu_stall_timeout

Interpret the result: A custom very low value can report short delays. A normal value plus repeated unchanged stack frames supports investigating stalled progress, not just tuning the timeout.

Evidence-guided next steps

Remove an identified scheduling or tracing trigger

If RCU warnings begin only with an optional real-time workload or heavy tracer, compare its documented default scheduling or reduced tracing profile. Verify that normal RCU progress returns while keeping error detection enabled.

Precautions: Do not change safety-critical real-time policy casually. Record the previous profile and change only the optional triggering component.

Recovery / rollback: Restore the recorded profile if needed for an isolated reproducer; avoid returning a known-stalling workload to daily operation.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.