Performance & Virtualization

Background writeback stalls foreground work

Foreground latency rises during bulk writes without device errors. Compare I/O PSI, queue latency and dirty pages before changing writeback or I/O budgets.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Interactive reads pause while a backup or large copy runs.
  • I/O pressure rises without a corresponding hardware-error record.

Relevant environment

Linux PSI, dirty-page writeback and cgroup v2 I/O control; sysstat tools must already be installed.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • A bulk writer may saturate the device queue or cause foreground dirty-page throttling.
  • An io.max or parent resource budget can limit throughput; high utilization alone does not prove a failing disk.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Reads host I/O stall time; PSI must be supported and enabled.

cat /proc/pressure/io

Interpret the result: Rising some or full identifies time lost to I/O stalls, not a specific device. Compare changes during and after the bulk writer.

Check 2

Read-only sysstat sampling; first report is usually since boot, so use later intervals.

iostat -xz 1 6

Interpret the result: Compare read/write await, queue depth and throughput for the actual backing device. %util alone is not a universal saturation threshold on parallel devices.

Check 3

Reads relevant memory accounting; repeat around the pause.

rg '^(Dirty|Writeback|MemAvailable):' /proc/meminfo

Interpret the result: Dirty pages awaiting writeback and active Writeback support a buffered-writer correlation. One snapshot does not establish queue service time.

Evidence-guided next steps

Limit the identified background writer

If bulk writes correlate with foreground pauses, reduce that job's application-level concurrency or use a finite device-specific IO bandwidth budget for its cgroup. Verify the device mapping and latency impact.

Precautions: The backup or copy will take longer. Device-mapper stacks and filesystem support can affect where cgroup accounting applies.

Recovery / rollback: Restore the recorded job settings or I/O limit on the same device and cgroup.

Did this solution help you?

Share this solution#

Adjust a measured writeback budget

If buffered dirty-page bursts are the verified trigger, trial documented dirty_background_bytes and dirty_bytes thresholds appropriate to device speed and RAM. Alternatively reschedule bulk writers away from interactive work.

Precautions: Byte and ratio forms supersede their counterpart; record all original values. Smaller bursts can improve latency while reducing peak throughput.

Recovery / rollback: Restore original VM thresholds or schedule and compare identical I/O work; do not disable writeback safety mechanisms.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.