Kernel & System Stability

Processor reports a machine-check hardware error

Machine-check records need CPU-family and bank-specific decoding; separate corrected reports from fatal events before replacing a component.

On this page
  1. Symptoms & scope
  2. Possible causes
  3. Diagnose safely
  4. Evidence-guided next steps
  5. References & review
  6. Related problems

Symptoms & scope

  • Kernel logs mce or Hardware Error records.
  • A workload may crash, the machine may reboot, or only corrected-error counters rise.

Relevant environment

x86 systems with Machine Check Architecture and optional rasdaemon/EDAC collection; reporting coverage depends on hardware and firmware.

Recognizable messages (synthetic examples)
mce: [Hardware Error]: Machine check events logged

Obtain the detailed bank/status records and corrected severity; the summary cannot identify a replaceable component.

Possible causes

These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.

  • CPU, memory, interconnect or platform errors may be reported through MCA.
  • Overclocking, undervolting and firmware errata can contribute; an MCA bank number is not universally a DIMM or core number.

Diagnose safely

Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.

Check 1

Read this boot’s hardware-error reports, or -b -1 for a retained previous boot. Access may require an administrator.

journalctl -b -k --no-pager --grep='mce:|Hardware Error|EDAC'

Interpret the result: Preserve status, bank, CPU family and corrected/uncorrected designation. A single summary saying events logged requires the detailed records, not a part replacement.

Check 2

If rasdaemon recording is already configured, this reads its saved events; database access may need administrator rights. This syntax is documented by distribution packages such as Debian trixie; newer upstream releases use ras-mc-ctl db --errors.

ras-mc-ctl --errors

Interpret the result: Repeated decoded errors on the same component merit vendor review. An empty database does not exclude errors if recording was absent or unsupported.

Evidence-guided next steps

Restore recorded CPU and memory defaults

If custom clocks, memory profiles or voltage offsets were applied, return those settings to manufacturer-supported defaults one at a time and observe normal workloads. Preserve the original error records for comparison.

Precautions: Record firmware settings before changing them and retain encryption recovery information. Do not raise voltage or disable machine-check reporting to hide an event.

Recovery / rollback: Use the saved firmware profile only if it is supported and not implicated; keep the stable default when tuning reproduces errors.

Did this solution help you?

Share this solution#

Act on the decoded hardware event

If uncorrected events recur at supported defaults, back up important data and present decoded records to the hardware vendor. Use a CPU/board-specific firmware fix or targeted component diagnosis when the records and vendor guidance justify it.

Precautions: Corrected events need trend and context; an uncorrected event can threaten reliable computation. Do not force exhaustive stress on a repeatedly failing production machine.

Recovery / rollback: Retain vendor-supported recovery firmware and the original configuration; hardware replacement should follow the verified diagnosis.

Did this solution help you?

Share this solution#

References & review

This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.