What NUMA is and why it matters on two processors
NUMA (Non-Uniform Memory Access) is the memory architecture of a two-socket server, where each CPU has "local" memory with fast access and "remote" memory on the other socket, reached through the inter-processor bus. The topic matters most for servers with two or more sockets — how the processor was chosen directly determines the number of NUMA nodes and their size.
If an application is not NUMA-aware, the OS scheduler can place a process on one socket while its memory sits on another. The result is a performance drop with no obvious cause in CPU monitoring.
How to check NUMA topology
lscpu | grep -i numa
numactl --hardware
The numactl --hardware command shows the number of nodes, the memory size on each, and the node distances. A value of 10 means access to the local node, 20-21 means access to a neighboring one: the higher the number, the more expensive a miss is.
The problem: memory on the "wrong" node
A typical scenario: a two-processor server with 256 GB of memory, an application started without any binding. The Linux scheduler moves the process between cores on different sockets while its allocated memory stays on the first node. Every access to the second socket goes through the inter-processor interconnect, and latency grows.
This is especially visible on database servers, where cache and buffers need to live next to the cores that read them thousands of times per second.
Binding processes: numactl and cpuset
Explicit binding of a process to a node solves the problem directly:
numactl --cpunodebind=0 --membind=0 /usr/bin/mysqld
numactl --interleave=all /usr/bin/myapp
The --cpunodebind flag pins the process to node 0's cores, --membind pins it to that same node's memory. The --interleave=all flag spreads memory evenly across all nodes — useful for applications that are not NUMA-aware and gain nothing from locality.
| Parameter | What it does | When to use it |
|---|---|---|
| --cpunodebind | Pins the node's cores | Databases, JVM with a large heap |
| --membind | Pins the node's memory | Together with cpunodebind |
| --interleave=all | Interleaves memory across nodes | NUMA-unaware applications |
| --preferred | Prefers a node but does not forbid others | Flexible workloads |
NUMA in databases and virtualization
PostgreSQL and MySQL recommend running one instance per NUMA node for large memory sizes, or explicitly binding the process with numactl if there is a single instance for the whole server. In virtualization (KVM, Proxmox) you should allocate a virtual machine's vCPUs and memory within a single physical node — otherwise the hypervisor runs into the same cross-node traffic problem.
Before changing any binding, measure the effect: run a test before and after with sysbench under a load close to the real one, not a synthetic single-threaded run.
NUMA tuning checklist
- Check the topology with numactl --hardware before tuning, not after.
- Bind the process to CPU and memory together, not separately.
- For NUMA-unaware applications use --interleave=all instead of manual binding.
- Measure latency and throughput before and after — the gain is not always obvious under light load.