The Proxmox VE web interface keeps load graphs for only 30 days and refreshes them once a minute. That is not enough for analyzing a week-old incident or for SMS alerts. The fix is a standard metric export to an external time-series database and dashboards built on top of it.
Why move Proxmox VE monitoring outside the panel
The built-in CPU and RAM graph on a cluster node is fine for a quick glance, but it does not store history longer than a month and cannot send a notification to Telegram or email when a threshold is crossed. A separate Proxmox VE monitoring stack solves three tasks: long-term metric storage, flexible alerting, and a shared dashboard for every node in the cluster on one screen.
Below is a working combination of the built-in metric exporter, InfluxDB as storage, and Grafana for visualization. This setup assumes that a Proxmox VE cluster is already built and every node can see the others over the network.
How to enable metric export to InfluxDB
Proxmox VE can send metrics to an external server without extra agents — configure it through Datacenter → Metric Server or directly in /etc/pve/status.cfg.
influxdb: influxdb-main
server 10.0.0.50
port 8089
protocol udp
timeout 2
mtu 1400
UDP port 8089 is the standard InfluxDB receiver using Line Protocol. If the network between the nodes and the metrics server is unstable, choose TCP on port 8086 instead and enable token authentication.
Installing InfluxDB and Grafana: minimum requirements
Both services can run on a single Debian 12 virtual machine, separate from the cluster hypervisors. For a cluster of 3-5 nodes, 2 vCPU, 4 GB RAM, and 40 GB of disk for a year of metric data is enough.
| Component | Port | Purpose |
|---|---|---|
| InfluxDB | 8086/8089 | receiving and storing metrics |
| Grafana | 3000 | dashboard web interface |
| Telegraf | — | optional metrics collection for the monitoring VM itself |
After installing the packages, create a database for Proxmox VE metrics:
influx -execute 'CREATE DATABASE proxmox'
influx -execute 'CREATE USER pve WITH PASSWORD 'changeme''
How to connect Grafana and import a dashboard
Open Grafana at the server address on port 3000, add a data source of type InfluxDB, and point it to the proxmox database. Then follow these steps:
- Add Data Source → InfluxDB, InfluxQL protocol, server address and port.
- Import a ready community dashboard for Proxmox VE by ID through Dashboards → Import.
- Select the created data source as the dashboard variable.
- Check that the panels show cluster node names and graphs for CPU, RAM, disk, and network.
If the panels are empty, metrics have not reached InfluxDB yet — check the next section of this article.
How to set up alerts for overload and failures
Alerting in Grafana is built on top of a regular data query: a rule fires when a value crosses a threshold over a set interval. Typical rules for a Proxmox VE node are CPU load above 90% for longer than 5 minutes and free storage space below 10%.
Create a contact point with an email address or a webhook to a Telegram bot, then link it to an alert rule under Alerting → Alert rules. In the condition, use the last() aggregation function and the threshold, set the evaluation interval to 1 minute, and set the For field to 5 minutes so short load spikes do not trigger it. Set up the same kind of alert for a lost quorum — the node will then send a signal before you even need the quorum recovery procedure.
Checklist after setting up monitoring
- Metrics from every cluster node appear in Grafana with no gaps longer than 5 minutes.
- The InfluxDB retention policy is set to at least 90 days for retrospective incident analysis.
- Alerting was tested by disabling one node — the notification arrived within 5 minutes.
- Access to Grafana is protected by a password and is not exposed to the open internet without a TLS reverse proxy.
Once every checklist item passes, the monitoring stack is ready. It is convenient to test alert delivery and metric collection with the same calls used by the Proxmox REST API — after that, just keep the stack up to date as new nodes join the cluster.