Skip to main content

Dedicated GPU Server: How to Choose a Card for ML

Dedicated Servers · 29.09.2026

Which GPU do you need: training, inference, or rendering

Choosing a GPU for a dedicated server depends on the task, not on "which card is more powerful." Training large models requires a lot of VRAM and high memory bandwidth. Inference (running an already trained model) can run on a card with less memory, but with a focus on low latency. Rendering and video encoding depend more on the number of CUDA/Tensor cores and support for the codecs you need.

A common beginner mistake is buying the most powerful card "just in case" without checking whether the model or scene even fits in its memory. If it does not fit, thousands of extra cores do not help much.

VRAM matters more than core count

The amount of video memory determines which model you can load at all and what batch size you can train with. Rough task classes:

Task classMinimum VRAMComment
Inference of small models, scene rendering16 GBEnough for most applied tasks
Fine-tuning medium models24-48 GBRequires a workstation-class card
Training large models from scratch80 GB and upUsually several cards with NVLink

If a model does not fit in one card's memory, you either have to cut the batch size or split the model across several GPUs — and that means a different server architecture and different requirements for the processor, which has to keep feeding data to the cards.

Power and cooling for a GPU server

A top-tier ML card draws 300-700 W under load. A server with 4-8 such cards easily reaches 4-6 kW of total consumption — that is no longer a standard 16 A rack circuit, but a separate calculation for power redundancy with two independent feeds.

Cooling should not be underestimated either: a GPU server packed with cards needs directed front-to-back airflow and may not fit into a compact chassis without hurting the thermal envelope.

CPU, PCIe, and form factor for multiple cards

Each data-center-class GPU occupies a PCIe x16 slot. With 4 or more cards the processor must provide 64+ PCIe lanes without sharing them with network cards and NVMe drives — one of the reasons GPU servers are almost always built on platforms with a high PCIe lane count rather than consumer chipsets.

The form factor matters too: full-height cards with active cooling physically do not fit into just any chassis, and it is worth checking the article on 1U and 2U form factors — most multi-card GPU servers are built in 3U-4U chassis exactly because of cooling and card-length requirements.

NVLink, passthrough, and GPU virtualization

If the cards need to exchange data directly, without going through the host's system memory, you need NVLink or a similar bus — it speeds up distributed training across several GPUs several times over compared to exchanging data through PCIe.

Virtualization uses either PCIe passthrough (the whole card is given to one virtual machine) or vGPU (the card is split among several VMs in software, if the vendor supports it). Passthrough is simpler to configure and delivers full performance, but rules out sharing the card.

nvidia-smi
nvidia-smi -q -d POWER

GPU server selection checklist

  • Calculate how much VRAM your model actually needs instead of buying extra just in case.
  • Check the total power draw of all cards and the power supply headroom.
  • Confirm the number of PCIe lanes on the processor when using 4 or more cards.
  • Decide in advance: passthrough or vGPU, if the server will be virtualized.
← Back to Knowledge Base Ask Support