A Proxmox VE cluster without quorum switches into a protective mode: the /etc/pve filesystem becomes read-only, and configuration changes, starting new VMs, and HA failover are blocked. This protects against split-brain, but to an administrator it just looks like "the panel won't let me do anything" at the worst possible moment.
What quorum is and why it breaks
Quorum is a simple majority vote of cluster nodes: if only two out of five nodes are reachable, the cluster cannot be sure the remaining three are not running separately and making conflicting changes, so it blocks writes. Quorum breaks due to a network outage between nodes, several nodes going down at once, or an incorrect vote count left over after removing a node from the cluster.
How to quickly check quorum status
The first command whenever you suspect a cluster problem is checking corosync status on a live node:
pvecm status
corosync-quorumtool -s
Two lines in the output matter: Quorate must read Yes, and Expected votes must match the actual number of nodes. If Quorate: No, the cluster runs in read-only mode until a majority is restored.
Diagnosing the network between nodes
Quorum most often breaks because of problems on the corosync network, not because the servers themselves went down. Check connectivity and latency between nodes at the addresses listed in /etc/pve/corosync.conf:
| Check | Command | Normal value |
|---|---|---|
| Ping between nodes | ping -c 20 10.10.10.2 | 0% loss, latency under 1 ms |
| Corosync link status | corosync-cfgtool -s | every link shows OK |
| Service log | journalctl -u corosync -n 100 | no token lost entries |
How to restore quorum when some nodes are unreachable
If nodes have been down for a long time and a majority cannot be reached, an administrator can temporarily lower the quorum threshold on the surviving nodes:
pvecm expected 1
The command changes the expected vote count on the fly, without touching the configuration file, and unblocks writes to /etc/pve on the remaining nodes. This is a temporary measure until the network is restored or the failed servers are repaired — do not leave it in place permanently, or the cluster loses its protection against split-brain.
What to do after every node is back online
As soon as connectivity is restored, set the vote count back to normal and check that the nodes see each other again without manual commands:
pvecm expected 3
pvecm status
If the cluster does not reassemble on its own after the network is back, restart the corosync service one node at a time, starting with the one that was unreachable the longest. Restarting it on every node at once can cause another brief split for a few seconds.
How to lower the risk of another quorum loss
A cluster of three or more nodes survives losing one node without issues — quorum holds with a majority of the rest. For a two-node cluster, add a QDevice as a third vote on an external server, and it helps to put quorum status on a separate monitoring dashboard so you notice a problem before your users do. Also remember that the Proxmox VE firewall must not block corosync ports between nodes — that is another common cause of a false quorum loss after changing network rules.
Checklist for quorum recovery
- The root cause is identified: network, node failure, or an incorrect expected vote count.
pvecm statusshowsQuorate: Yeson every live node.- The temporary
expected votesvalue is set back to normal after the repair. - An alert for quorum loss is configured so you do not find out from customer complaints.