A 7-billion-parameter model takes 14 GB in fp16 or 4 GB with 4-bit quantization — a 3.5x difference for the same set of weights. Buying a GPU server without calculating VRAM first is a common mistake: the card either sits half-idle, or the model does not fit and crashes with CUDA out of memory. Let's work out how to size it precisely instead of guessing.
Why calculate VRAM in advance
Video memory is the one GPU resource you cannot extend on the fly, unlike system RAM or disk. If a model weighs more than the available VRAM, inference either fails to start, or the driver starts offloading layers to system RAM through --gpu-memory-utilization or a similar flag, and speed drops several times over. Sizing it before buying a server saves both money and the time spent rebuilding the setup.
The rule is simple: you need to count not just the model weights but also the context, the batch, and CUDA's own buffers. Sometimes the context alone eats half the card on long conversations.
The formula for sizing VRAM for inference
A basic estimate of memory needed for the weights:
VRAM_GB = (params_billion * bytes_per_param) * 1.2
The 1.2 factor is a margin for activations, the KV cache, and memory fragmentation. bytes_per_param depends on precision: 2 bytes for fp16/bf16, 1 byte for int8, 0.5 bytes for int4. For a 13-billion-parameter model in fp16 you get 13 * 2 * 1.2 ≈ 31.2 GB — that already will not fit on a single 24 GB card.
Add the KV cache separately: it grows linearly with context length and the number of concurrent requests. For a 7B model with an 8192-token context and one conversation, the cache takes roughly 1-2 GB, but with ten parallel sessions it climbs to 10-20 GB.
How quantization changes memory requirements
Quantization lowers weight precision but sharply cuts VRAM usage and barely hurts answer quality at 8 and 4 bits. Below is a reference point for a 7-billion-parameter model.
| Precision | Bytes per param | VRAM for weights | Quality loss |
|---|---|---|---|
| fp16 / bf16 | 2 | ~14 GB | none |
| int8 | 1 | ~7 GB | minimal |
| int4 (GGUF Q4) | 0.5 | ~3.5 GB | noticeable on hard tasks |
| int2 | 0.25 | ~1.8 GB | significant |
For running on a CPU with no GPU at all, check out GGUF quantization in llama.cpp — the same formula applies there, but to system RAM instead of VRAM.
VRAM for fine-tuning differs from inference
Fine-tuning needs 3-4 times more memory than plain inference of the same model, because you also have to store gradients and optimizer state. Full fine-tuning of a 7B model in fp16 takes about 60-70 GB of VRAM — already several GPUs. The LoRA method removes this problem: only small adapter matrices get trained, while the base weights stay frozen at 4-bit quantization.
With LoRA and QLoRA quantization, fine-tuning that same 7B model fits into 10-12 GB of VRAM — doable on a single mid-range card instead of a cluster.
How much VRAM popular models need
| Model | fp16, inference | int4, inference |
|---|---|---|
| 7B (Mistral, Llama 3 8B) | 16 GB | 5 GB |
| 13-14B | 28 GB | 8 GB |
| 34B | 68 GB | 20 GB |
| 70B | 140 GB | 40 GB |
These figures include a margin for a 4096-token context with a single user; as load grows, budget an extra 10-20% for each parallel request stream.
Which GPU to pick on a budget
For a 7B model at 4-bit quantization, a card with 8 GB of VRAM is enough — a fitting task for Ollama on a VDS with a single consumer GPU. For 13B in fp16, or a chat service with several concurrent users, you need a card with 24 GB of VRAM. For 70B without quantization there is no reasonable alternative to a multi-GPU cluster — take the int4 version and a single 48 GB card, or split layers across two cards using vLLM with tensor parallelism.
- Up to 8 GB VRAM — 7B models in int4, personal and small projects.
- 16-24 GB VRAM — 7-13B in fp16 or 34B in int4.
- 48 GB and above — 70B in int4 or fine-tuning mid-sized models.
- Multiple GPUs — 70B in fp16 or training from scratch.
Summary: how to avoid sizing mistakes
Before ordering a server, calculate memory with the formula, add a KV cache for the real context length and the number of parallel requests, then add a 15-20% margin for fragmentation. It is cheaper to buy a card with extra VRAM once than to migrate a model later and lose time moving the weights.
- Determine the model's parameter count and target precision (fp16, int8, int4).
- Calculate the base amount with the formula and add a KV cache for the context.
- For fine-tuning, budget 3-4 times more memory than for inference, or use LoRA.
- Pick a GPU with a 15-20% margin above the calculation — for fragmentation and load growth.