Performance & Virtualization

A service is slow while the host still has idle CPU

CPU quota can suspend a busy cgroup while other CPUs remain idle. Compare quota settings and throttling deltas before adding threads or changing governors.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Work slows in bursts although the machine has spare CPU capacity.
  • The affected cgroup's throttling counters increase during the workload.

Relevant environment

Linux cgroup v2 CPU controller and systemd CPUQuota; controller files depend on the active hierarchy.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • A finite cpu.max or systemd CPUQuota can cap aggregate execution time across the service's threads.
  • An ancestor slice may impose the effective cap; CPUWeight only distributes contention and is not a hard quota.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Replace example.service; no runtime settings are changed.

systemctl show example.service -p ControlGroup -p CPUQuotaPerSecUSec -p CPUQuotaPeriodUSec -p CPUWeight

Interpret the result: A quota equivalent to 100% caps about one CPU's aggregate time, not one fixed core. Inspect parent slices if the local quota is unlimited.

Check 2

Replace the path with ControlGroup under /sys/fs/cgroup on cgroup v2; repeat this read around the same workload.

cat /sys/fs/cgroup/system.slice/example.service/cpu.max /sys/fs/cgroup/system.slice/example.service/cpu.stat

Interpret the result: cpu.max gives budget and period. Rising nr_throttled or throttled_usec shows local quota enforcement; inspect ancestors because their throttling may not appear in the child's counters.

Evidence-guided next steps

Right-size the measured CPU quota

If throttling coincides with latency and more CPU time is intended, raise the finite quota for this service or its constraining slice through the owner configuration. Recheck latency and neighboring workloads.

Precautions: A child cannot escape a tighter parent cap. Record old settings and avoid removing isolation for all services.

Recovery / rollback: Restore previous quota and period at the same hierarchy level and compare the same workload.

Did this solution help you?

Share this solution#

Reduce the application's burst demand

If the quota is intentional, reduce worker concurrency or batch size so the application fits its CPU budget. For best-effort prioritization rather than a hard cap, assess CPUWeight with measured competing workloads.

Precautions: Less concurrency may reduce throughput; changing quota to weight also changes resource-isolation semantics.

Recovery / rollback: Restore application concurrency and the original quota/weight settings if measured results worsen.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.