Services & systemd

A service misses its userspace watchdog heartbeat

A running service aborts after missing WATCHDOG=1 heartbeats. Check notification support, sender identity and workload stalls before changing its deadline.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • The service journal reports Watchdog timeout.
  • An abort may be followed by Restart-policy recovery.

Relevant environment

systemd per-service WatchdogSec; separate from kernel lockup detectors and hardware watchdog devices.

Recognizable messages (synthetic examples)
systemd[1]: example.service: Watchdog timeout (limit 30s)!

A service heartbeat was missed; this is not a kernel watchdog diagnosis.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • Watchdog support may be absent or notifications sent by a process not accepted by NotifyAccess.
  • An event-loop blockage, scheduling delay or I/O stall can prevent otherwise implemented heartbeats.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Replace the service name; reads watchdog state without changing it.

systemctl show example.service -p WatchdogUSec -p NotifyAccess -p WatchdogTimestampMonotonic -p Result

Interpret the result: A nonzero interval requires periodic application support. Result=watchdog differs from an ordinary exit-code failure.

Check 2

Use the affected unit; restricted journals require authorized read access.

journalctl -b -u example.service -o short-monotonic --no-pager -n 120

Interpret the result: Compare expiration with last progress. Repeated stalls at one stage justify profiling that stage, but do not identify its blocking resource.

Evidence-guided next steps

Repair the watchdog notification path

If support exists, configure the documented notification mode and accepted sender. If the program lacks support, remove a locally added WatchdogSec setting for that service.

Precautions: Removing it removes hang detection. A heartbeat unrelated to application health can hide deadlocks.

Recovery / rollback: Restore notification mode and prior WatchdogSec/NotifyAccess values, then observe normal operation.

Did this solution help you?

Share this solution#

Address a measured stall or deadline mismatch

If progress stops, investigate blocked work and fix its resource or concurrency issue. If a documented healthy operation exceeds the interval, choose a finite service-specific watchdog budget covering it.

Precautions: A larger interval delays genuine hang recovery. Preserve timing and core-dump evidence before repeated restarts.

Recovery / rollback: Restore the previous interval or application change and compare progress plus heartbeats under the same workload.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.