The platform itself does not reboot without a reason: almost always something external triggers it — power glitches, overheating, faulty memory, or a cluster watchdog that resets the node when it loses contact with the rest. Checks should follow the order from common and cheap to rare.
Reboot or hang — how to tell them apart
First it helps to understand what actually happened: a reboot is a full power-off and power-on cycle, and the node reappears on the network after a known delay. A hang is when the node stops responding while power itself was never cut, and it will not return without intervention.
The difference shows up in the log: after a real reboot, the new entry starts with clean kernel initialization and hardware enumeration. After a hang followed by a hardware reset, the record before the new boot simply breaks off — the last messages never made it to disk.
Where to look first
A handful of sources are enough for the first pass, and each answers a different question.
- the log from the previous boot — the command
journalctl -b -1shows what happened before the stop - current uptime — the command
uptimeshows how long the system has been running without interruption right now - kernel messages about hardware — the command
journalctl -kshows memory, disk and controller errors - component temperature — the command
sensorsshows CPU and board heat, if the sensor package is installed
Table of causes: from common to rare
Reboot causes are not spread evenly: some show up in nearly every case, others are extremely rare. The order in the table reflects that.
| Cause | How it shows up | What to check |
|---|---|---|
| Power glitches or voltage dips | the reboot is sudden, often under load, with no warnings | the log around the failure, the power supply and cable, the UPS log |
| CPU or board overheating | the node reboots specifically under load | sensors, kernel messages about thermal protection |
| Faulty memory | reboots and hangs are random, with no pattern | a memory test, correction-error messages in the kernel log |
| Filesystem error | the reboot happens during a disk check at startup | filesystem messages in the log, a disk check |
| Cluster watchdog on quorum loss | only clustered nodes reboot, and only during network trouble | the cluster service log, quorum status |
| Update that pulls in a new kernel | the reboot is predictable and follows a package update | the package manager log, the list of installed kernels |
| Disk full | the node hangs and drops services one by one instead of rebooting cleanly | free space on the partitions |
Cluster watchdog and quorum loss
In a cluster a separate mechanism — the watchdog — tracks the link between nodes and the quorum state. When a node loses contact with the rest and quorum is lost, the watchdog deliberately reboots it so it does not keep running in isolation and drift out of sync with the rest in its data.
This happens most often in a cluster of two nodes: when the link between them breaks, neither one can gather a majority of votes, and the reboot fires on almost any network glitch. In a cluster with a larger number of nodes, the minority is left isolated while the majority keeps running without rebooting.
Leaving yourself evidence for next time
When the cause was not found on the first try, it helps to prepare the node for the next failure in advance — otherwise there will be no evidence again.
- enable saving the log across boots instead of keeping it only in memory
- forward the log to another node or server so the record is not lost along with the disk
- set up monitoring for temperature and power supply load instead of checking them only after a failure
- connect the node through a UPS that reports loss of external power
Why guest machines feel slow
That is a neighboring but different question: the node runs stably and still hands guests too few resources.
- too many virtual machines share the same processor cores
- storage disks cannot keep up with simultaneous requests from several guests
- a guest lacks enough allocated memory and starts swapping inside itself
A detailed look at this topic matters most for compact nodes — it is covered by the article on a home lab on a mini PC.
Frequently asked questions
Why does a Proxmox server reboot on its own?
The platform itself contains no code that reboots a healthy node without a reason. Almost always something external triggers it: power glitches, overheating, faulty memory, or a cluster watchdog that resets the node when it loses contact with the rest. Checks run from common and cheap to rare and complex.
How do I view the log from the last boot?
The command journalctl with the flag -b -1 shows the log from the previous boot, if saving the log across restarts is enabled. It shows the last messages before the stop: kernel errors, memory or disk warnings, or a sharp break in the record if what happened was not a clean shutdown but a hardware reset.
Can a cluster reboot a node?
Yes: a node in a cluster watches the quorum state, and when it loses contact with the rest, a dedicated watchdog deliberately reboots it so the node does not keep running in isolation. On two-node clusters this happens especially often, because when the link breaks neither side can gather a majority of votes.
How do I check memory?
Faulty memory is found with a test that runs the node's free memory through cycles of reads and writes and looks for mismatches; it is worth running whenever reboots look random and unexplained. Memory correction-error messages are also visible in the kernel log, and they often appear long before the first hang.
Conclusion
Diagnosis is best done in order: first the cheap and common causes — power, heat, memory and disk, and only then the rare ones — a kernel update and watchdog behavior. A log saved across boots saves most of the time on the next investigation. The article on Proxmox VE covers the broader picture of the platform and its role in the infrastructure.