What Ceph Is and Why It Matters in Proxmox VE
Ceph is a distributed software-defined storage system that combines the local disks of several servers into a single fault-tolerant pool. Proxmox VE integrates Ceph directly into the web interface: there is no need to run a separate storage cluster, disks are configured right from the node panel.
The main advantage is shared storage without a dedicated SAN. VM disks are automatically replicated between nodes, and if one server fails, machines keep running from other data copies. This makes Ceph the foundation for live migration and high availability in a Proxmox VE cluster.
What Requirements the Cluster Nodes Must Meet
Ceph is designed for a cluster of at least three nodes — fewer nodes leave the monitors without quorum and cause the pool to stop if one server is lost. Production setups need a separate network for Ceph traffic: replication and heartbeat generate heavy load, and a shared link with VM traffic quickly becomes a bottleneck.
| Parameter | Minimum | Recommended |
|---|---|---|
| Nodes in the cluster | 3 | 5 or more |
| Network for Ceph | 1 Gbps | 10 Gbps, dedicated VLAN |
| Disks for OSD | SSD | NVMe for journal and data |
| RAM | 4 GB per OSD | 8 GB per OSD |
How to Install Ceph on Every Cluster Node
Ceph packages are installed on each Proxmox VE cluster node separately, through the web interface or the console. The steps are the same for every server.
pveceph install --repository no-subscription
pveceph init --network 10.10.10.0/24
The pveceph init command sets the network over which nodes exchange service traffic. After initialization on the first node, the other cluster servers immediately see the Ceph configuration.
How to Create Monitors, Managers, and OSDs
A monitor (MON) stores the cluster map and handles quorum, a manager (MGR) provides metrics and statistics. You need at least three of each on different nodes:
- Create a monitor on each of the three nodes through the Ceph → Monitor section of the web interface.
- Add a manager on the same nodes through Ceph → Manager.
- For each physical data disk, create an OSD with the command
pveceph osd create /dev/sdb.
An OSD (Object Storage Daemon) is a process that serves one physical disk. The more OSDs are spread across nodes, the faster and more reliable the pool becomes.
How to Create an RBD Pool and Attach It as Storage
A pool is a logical space with a set replication level. VM disks use the RBD (RADOS Block Device) type.
pveceph pool create vm-pool --size 3 --min_size 2
The size 3 parameter means three copies of every data block on different OSDs, and min_size 2 means at least two available copies are required for the pool to keep accepting writes. After creation, the pool automatically appears under Datacenter → Storage and is available for VM and LXC disks.
How to Check Cluster Health and What to Watch
The main diagnostic command is ceph -s. A HEALTH_OK status means the cluster is running normally. HEALTH_WARN usually points to degradation from a disabled OSD or low free space, while HEALTH_ERR requires immediate action.
In the web interface, the Ceph → Log tab shows monitor events, and the Ceph → OSD section shows the state of each disk and how much space is used. If one OSD fails repeatedly, check the disk with smartctl and remove it from the pool instead of waiting for a complete failure.
Checklist Before Putting Ceph Into Production
- The cluster has at least three nodes and three monitors.
- A separate network is dedicated to Ceph traffic.
- Enough OSDs are created for the required replication level.
- The pool is created with
size 3andmin_size 2for production. - The
ceph -scommand showsHEALTH_OK.
After that, the Ceph pool can serve as the main storage for VM disks, and for further data protection set up backups on a separate backup server.