Symptoms & scope
- Kernel logs mce or Hardware Error records.
- A workload may crash, the machine may reboot, or only corrected-error counters rise.
Relevant environment
x86 systems with Machine Check Architecture and optional rasdaemon/EDAC collection; reporting coverage depends on hardware and firmware.
Recognizable messages (synthetic examples)
mce: [Hardware Error]: Machine check events loggedObtain the detailed bank/status records and corrected severity; the summary cannot identify a replaceable component.
Possible causes
These are possible explanations, not a confirmed diagnosis. Several independent faults can coexist.
- CPU, memory, interconnect or platform errors may be reported through MCA.
- Overclocking, undervolting and firmware errata can contribute; an MCA bank number is not universally a DIMM or core number.
Diagnose safely
Run one command at a time in the relevant session. Read the explanation first. Uppercase placeholders need your own values; tools and privileges vary by distribution. These commands are displayed here and never executed by the website.
Check 1
Read this boot’s hardware-error reports, or -b -1 for a retained previous boot. Access may require an administrator.
journalctl -b -k --no-pager --grep='mce:|Hardware Error|EDAC'Interpret the result: Preserve status, bank, CPU family and corrected/uncorrected designation. A single summary saying events logged requires the detailed records, not a part replacement.
Check 2
If rasdaemon recording is already configured, this reads its saved events; database access may need administrator rights. This syntax is documented by distribution packages such as Debian trixie; newer upstream releases use ras-mc-ctl db --errors.
ras-mc-ctl --errorsInterpret the result: Repeated decoded errors on the same component merit vendor review. An empty database does not exclude errors if recording was absent or unsupported.
Evidence-guided next steps
Restore recorded CPU and memory defaults
If custom clocks, memory profiles or voltage offsets were applied, return those settings to manufacturer-supported defaults one at a time and observe normal workloads. Preserve the original error records for comparison.
Precautions: Record firmware settings before changing them and retain encryption recovery information. Do not raise voltage or disable machine-check reporting to hide an event.
Recovery / rollback: Use the saved firmware profile only if it is supported and not implicated; keep the stable default when tuning reproduces errors.
Did this solution help you?
Act on the decoded hardware event
If uncorrected events recur at supported defaults, back up important data and present decoded records to the hardware vendor. Use a CPU/board-specific firmware fix or targeted component diagnosis when the records and vendor guidance justify it.
Precautions: Corrected events need trend and context; an uncorrected event can threaten reliable computation. Do not force exhaustive stress on a repeatedly failing production machine.
Recovery / rollback: Retain vendor-supported recovery firmware and the original configuration; hardware replacement should follow the verified diagnosis.
Did this solution help you?
References & review
This guide was prepared from primary project or distribution sources and reviewed on the date shown. This is an editorial source check, not evidence that a fix was reproduced on your hardware. Diagnostic log examples are synthetic fixtures. Version-dependent details must be checked against your installed release.
- Kernel RAS and MCA error handling (project or distribution documentation)
- x86 machine-check reporting parameters (project or distribution documentation)
- Debian ras-mc-ctl: reading the recorded error database (project or distribution documentation)