Why bond server network ports
Port bonding (or, at the switch level, link aggregation) solves two problems: it increases channel throughput beyond the speed of a single port and adds fault tolerance — if one cable or switch port goes down, traffic keeps flowing through the remaining interfaces without dropping connections.
Most servers have 2-4 network ports at 1 or 10 Gbps. Individually each one is limited to its own speed, while bonding two 10 Gbps ports in LACP mode theoretically gives 20 Gbps of total throughput, although a single TCP connection is still limited to the speed of one physical link.
Bonding modes: which one to choose
In Linux, the bonding driver supports several modes, but in practice three are used:
- active-backup (mode 1) — one port is active, the other is standby. Requires no switch configuration but does not increase speed.
- balance-rr (mode 0) — packets go out round-robin across ports. Gains speed for UDP but often breaks TCP packet order.
- 802.3ad, also known as LACP (mode 4) — the standard aggregation protocol, negotiated with the switch. It distributes connections by a hash function of MAC/IP/port and requires LACP configuration on both sides.
Production servers in a data center use LACP: it is the only one of the three officially negotiated with the switch side and it correctly handles a link failure.
Configuring LACP on the server and switch
On Linux with NetworkManager, a bond is created like this:
nmcli connection add type bond ifname bond0 mode 802.3ad
nmcli connection modify bond0 +bond.options "lacp_rate=fast,miimon=100"
nmcli connection add type ethernet ifname eth0 master bond0
nmcli connection add type ethernet ifname eth1 master bond0
nmcli connection up bond0
With netplan the configuration looks like this:
network:
version: 2
bonds:
bond0:
interfaces: [eth0, eth1]
parameters:
mode: 802.3ad
lacp-rate: fast
mii-monitor-interval: 100
addresses: [203.0.113.10/24]
On the switch side, the ports the server is connected to are grouped into a port-channel with LACP active enabled. If the switch is set to static mode (without LACP), bonding in 802.3ad mode will not work — both ends must negotiate the protocol.
How to check that bonding is working
Interface status and LACP negotiation are checked like this:
cat /proc/net/bonding/bond0
ip -s link show bond0
In the output, the important line is MII Status: up for each slave interface, and the Aggregator ID — it must match on both ports if they are actually aggregated. After configuration, throughput should be measured with iperf3 on both sides at once, running several parallel streams (-P 4) to engage hashing across different connections.
Common mistakes: why LACP "does not come up"
The problem is most often a mismatch between server and switch settings. Common causes:
| Symptom | Cause |
|---|---|
| Bond is up, but speed is like a single port | Switch ports are not in the same port-channel |
| Aggregator ID differs between ports | Different VLANs or speed on switch ports |
| Link keeps flapping | miimon is too small or the cable/SFP is faulty |
| A single TCP connection does not speed up | The LACP hash algorithm distributes by flow, not by splitting one flow |
Another common mistake is enabling LACP on the server but leaving the switch ports in separate VLANs without a port-channel. Then bond0 comes up locally, but there is no real aggregation: traffic simply goes through one physical port.
Checklist before going into production
- LACP mode (802.3ad) is configured the same way on the server and the switch, and the ports are grouped into one port-channel.
- In
/proc/net/bonding/bond0, all slave ports showMII Status: upand the same Aggregator ID. - Throughput has been verified with a load test using several parallel streams.
- Failover has been tested manually: unplugging one cable must not drop active connections.
- If the server needs several public addresses on top of bond0, IP aliases or additional subnets are configured on the bond0 interface itself, not on eth0/eth1.
Network checks should be part of the general list during server acceptance — a failed link or a misconfigured LACP setup is easier to catch in the first hour of operation than a month later under load.