What DNS Failover is
DNS Failover is the automatic replacement of an IP address in an A or AAAA record when the primary server stops responding. A dedicated monitor keeps polling the primary host, and once it detects a failure, the system rewrites the zone record to point to a backup server without human involvement. It is a cheap alternative to a load balancer for sites that do not need two servers running at the same time, only insurance for the case when the primary one fails.
The technology does not replace backups and does not protect against data loss — it only solves the problem of domain availability while the main infrastructure is being restored.
How the availability check works
DNS Failover relies on a health check: an external agent contacts the server over HTTP, TCP or ping every 10 to 60 seconds and waits for a correct response. A typical configuration looks like this:
check_type: http
url: https://example.com/health
interval: 30s
timeout: 5s
failures_before_switch: 3A threshold of several consecutive failed checks (3 in the example) is needed so the system does not switch to the backup over a single random network hiccup. Once the threshold is hit, the system replaces the A record with the backup server's IP and waits for the primary to come back to life before switching back.
Setting a low TTL for a fast switch
How fast visitors see the new IP is determined by the record's TTL value, not by the speed of the monitoring itself. As long as the old value lives in the cache of the user's provider resolver, that resolver keeps knocking on the server that is down. For domains with Failover, TTL is usually lowered to 60 to 300 seconds:
example.com. 60 IN A 203.0.113.10For a detailed look at choosing a TTL value and its effect on update speed, see the article on configuring TTL. The trade-off to remember here: too low a TTL increases the number of requests to DNS servers and slightly raises resolution latency, too high a TTL slows down the switch during a failure.
DNS Failover versus Anycast
DNS Failover and Anycast solve a similar problem — service availability — but in different ways. Failover changes the record content and depends on TTL and the speed at which resolver caches around the world refresh, so switching takes minutes. Anycast announces the same IP from several points of presence at once, and BGP-level routing itself steers traffic away from a failed node — the switch happens at the network level, not DNS, and takes seconds.
Anycast is more complex and more expensive to deploy, so for small projects DNS Failover remains a reasonable trade-off between cost and downtime.
Limitations and pitfalls
| Problem | Why it happens |
|---|---|
| Some users still see the old IP | The provider's resolver keeps the record longer than the TTL, ignoring it |
| False switch | Health was checked from a single point and the network blipped locally |
| The switch happened but the site is still unreachable | Data or configuration on the backup server is not in sync |
| The switch back never happens | The primary server passes the health check but the application on it is still broken |
DNS Failover deployment checklist
- Set up the health check over the same protocol real traffic uses, not just ping.
- Lower the record's TTL in advance, a day before enabling Failover, so old values drop out of caches.
- Keep the backup server with current data and configuration, not just the software installed.
- Verify the switch with dig from several resolvers, not just one device.
- Set a threshold of several consecutive failed checks to avoid false triggers.