Performance & Virtualization

A service hits its cgroup memory boundary

A process is killed although host RAM remains available. Correlate cgroup limits, memory.events and the kill record before increasing a service budget.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • A kernel message names Memory cgroup out of memory.
  • The service fails with oom-kill while other applications remain active.

Relevant environment

Linux cgroup v2 memory controller; systemd MemoryHigh, MemoryMax and ancestor slices.

Recognizable messages (synthetic examples)
kernel: Memory cgroup out of memory: Killed process 932 (example) total-vm:3000000kB, anon-rss:1900000kB, file-rss:0kB

Inspect local and parent limits plus the surrounding OOM constraints.

systemd[1]: example.service: Failed with result 'oom-kill'.

This reports an OOM-related death, not necessarily a local MemoryMax violation.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • The cgroup or ancestor may reach memory.max and be unable to reclaim enough for an allocation.
  • Excessive concurrency or a growing cache can breach an otherwise intentional memory boundary.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Replace the unit; reads policy and failure result without restarting it.

systemctl show example.service -p ControlGroup -p MemoryHigh -p MemoryMax -p MemorySwapMax -p Result

Interpret the result: MemoryHigh drives pressure and reclaim; MemoryMax is a hard boundary. Parent limits can dominate the local values.

Check 2

Use ControlGroup under /sys/fs/cgroup; requires cgroup v2 and a still-existing cgroup.

cat /sys/fs/cgroup/system.slice/example.service/memory.events /sys/fs/cgroup/system.slice/example.service/memory.current /sys/fs/cgroup/system.slice/example.service/memory.max

Interpret the result: Correlate max/oom events with the kill record. oom_kill counts members killed by any OOM type, so that counter alone does not prove a local limit hit.

Evidence-guided next steps

Right-size a confirmed memory limit

If legitimate peak usage exceeds an accidentally low budget, increase this service's finite MemoryMax and review MemoryHigh and its parent's allocation. Check host headroom and repeat the intended workload.

Precautions: A larger budget can move OOM pressure to the host. Increasing a child limit cannot override its parent.

Recovery / rollback: Restore recorded budgets and restart affected work through its supported recovery path; a killed process's unsaved state is not restored.

Did this solution help you?

Share this solution#

Keep the budget and reduce demand

If isolation is deliberate, reduce worker count, batch size or application cache to fit it and investigate leaks. Make the workload resumable or checkpointed using the application's supported mechanism.

Precautions: Do not disable OOM protection or set every service unlimited. A cgroup kill record does not identify which allocation caused the peak.

Recovery / rollback: Restore prior application sizing only if it fits the restored budget; recover from valid application checkpoints or backups.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.