Performance & Virtualization

Memory reclaim stalls interactive work

Applications pause during memory pressure before any OOM kill. Use memory PSI, available memory and reclaim deltas to separate pressure from a CPU bottleneck.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Interactive latency rises while the machine reclaims memory.
  • Memory PSI rises even without a killed process.

Relevant environment

Linux with PSI and proc memory accounting; host-wide pressure and cgroup-local pressure can differ.

Recognizable messages (synthetic examples)
kernel: Out of memory: Killed process 824 (example) total-vm:8000000kB, anon-rss:6200000kB, file-rss:0kB, shmem-rss:0kB

An allocation reached OOM handling; inspect surrounding constraints before calling this simple total-RAM exhaustion.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • The active working set may exceed accessible memory and force direct reclaim.
  • A constrained cgroup may be reclaiming locally although the host has spare memory.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Reads host memory PSI; missing output can mean PSI support is disabled or unavailable.

cat /proc/pressure/memory

Interpret the result: some measures time with at least some tasks stalled; full means all non-idle tasks stalled together. Percentages are time, not RAM utilization.

Check 2

Reads selected kernel accounting fields; normally no elevation required.

rg '^(MemAvailable|MemFree|Cached|SwapTotal|SwapFree):' /proc/meminfo

Interpret the result: Low MemFree alone is normal with caching. Low MemAvailable combined with PSI supports real pressure; swap occupancy alone does not prove thrashing.

Check 3

Reads cumulative reclaim counters; repeat before and after the same workload, without changing VM settings.

rg '^(pgscan_direct|pgsteal_direct|allocstall)' /proc/vmstat

Interpret the result: Rising direct-reclaim and allocation-stall counters correlate with reclaim work. A historical total without a workload-time delta is weak evidence.

Evidence-guided next steps

Reduce the measured working set

If pressure tracks a specific application's memory growth, reduce its documented cache, batch size or worker count and investigate unbounded growth. Compare pressure and latency under the same workload.

Precautions: Dropping system caches is not a durable fix and can increase I/O. Save workload results before stopping heavy jobs.

Recovery / rollback: Restore recorded application settings if the throughput tradeoff is unsuitable; retain the measured baseline.

Did this solution help you?

Share this solution#

Correct the applicable memory budget

If pressure is local, review the service's MemoryHigh and parent budget rather than assuming host RAM is exhausted. If the host itself is constrained, schedule fewer concurrent jobs or add suitable capacity with a planned swap policy.

Precautions: Raising one budget can pressure neighboring services. PSI does not identify a single leaking application.

Recovery / rollback: Restore the prior service budgets or scheduling plan and compare host plus local pressure again.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.