What SEL is and why you should read it
SEL (System Event Log) is a non-volatile event log that the server's BMC keeps independently of the operating system. It records voltage drops, ECC events, overheating, fan stoppage, and other hardware events — even if the server was powered off or completely hung at that moment.
When a server reboots for no visible reason and the OS logs stay silent, the first place to check is the SEL, not syslog: the OS often does not have time to write anything before an emergency reboot.
How to read the log through IPMI
ipmitool sel list
ipmitool sel elist
The sel list command prints a short summary: record number, date, event type. The sel elist command decodes the event in more detail — usually that is enough to understand the nature of the failure without the documentation for a specific board.
If SSH access to the OS is unavailable because it hung, you can still read the log remotely through IPMI and KVM-over-IP — the BMC works even when the OS itself does not respond.
Common entries: ECC, power, overheating
| Event type | What it means | What to check |
|---|---|---|
| Correctable ECC | A single memory error was corrected | Repeat frequency, the DIMM module |
| Uncorrectable ECC | A memory error was not corrected | Module replacement, memtest |
| Power Supply Failure | One power supply unit went down | The backup PSU, PDU load |
| Temperature Upper Critical | A component is overheating | Fans, dust, thermal paste |
Uncorrectable ECC is the most serious entry in the table: it means data in memory was corrupted and not recovered on the fly. After such an entry, the module should be checked with memtest86+, not just fixed by rebooting the server and forgetting about it.
How to decode an event code
Each SEL line contains a sensor type, sensor number, and event data — three fields that, together with the IPMI specification version, point to the exact meaning of the event. For a specific vendor's board (Supermicro, Dell iDRAC, HPE iLO) there is a code mapping table in the server documentation — the general IPMI format is the same, but sensor numbering differs per vendor.
If an event code is unclear, it helps to check the sensor list output — it shows the current reading of the same sensor that triggered the event.
ipmitool sdr list
ipmitool sensor list
When to clear the SEL and what happens then
The SEL has a limited size — usually a few thousand entries — and once it fills up, old entries either get overwritten by new ones or the log gets locked. Some platforms stop writing new critical events once the SEL is completely full, which is more dangerous than it sounds.
ipmitool sel clear
Clear the log only after you have investigated and recorded the cause of the failure — for example, in a ticket about replacing a drive or a memory module. This complements continuous hardware monitoring through IPMI, which should catch new events immediately, not through a manual review once a month.
Checklist for diagnosing a failure with SEL
- Check the SEL before syslog if the server rebooted for no visible reason.
- Record the time and type of the event before clearing the log.
- For Uncorrectable ECC, test the memory module separately instead of relying on a reboot.
- Set up automatic SEL export to your monitoring system instead of relying on the administrator's memory.