What Is vLLM and Why Run It on a Server
vLLM is a library for fast inference of large language models on GPU. It uses PagedAttention technology, which saves video memory and multiplies throughput compared to a plain run through Transformers. vLLM starts an OpenAI-compatible HTTP server, so existing ChatGPT API clients switch to your own model without rewriting code.
Teams pick this library when a server handles many parallel requests: a chat widget on a website, an internal company assistant, or an API for a mobile app. For single experiments on CPU, llama.cpp fits better, while vLLM shows its strength on a GPU server with one or several video cards.
Server Requirements for Running vLLM
Check server parameters before installing. vLLM needs an NVIDIA GPU with CUDA support and enough video memory — the exact amount depends on model size and quantization.
| Parameter | Minimum | Recommended |
|---|---|---|
| GPU | NVIDIA, 8 GB VRAM | NVIDIA, 24 GB VRAM |
| NVIDIA driver | 525 | 550 |
| CUDA | 12.1 | 12.4 |
| Python | 3.10 | 3.11 |
| RAM | 16 GB | 32 GB |
Calculate the exact VRAM amount for a given model in advance — see VRAM calculation for a model.
Installing vLLM on a VDS
Install the NVIDIA driver and CUDA Toolkit, then create a Python virtual environment and install the package through pip.
sudo apt update
sudo apt install -y nvidia-driver-550 python3.11-venv
python3.11 -m venv /opt/vllm-env
source /opt/vllm-env/bin/activate
pip install --upgrade pip
pip install vllm
Check that the system sees the GPU with the nvidia-smi command. If the card does not show up, reboot the server and make sure the nouveau module is disabled in the kernel configuration.
How to Launch an OpenAI-Compatible Server
vLLM ships with a ready-made server that mirrors the OpenAI API request format. Launch it with a single command, specifying a model from Hugging Face.
python -m vllm.entrypoints.openai.api_server --model mistralai/Mistral-7B-Instruct-v0.2 --host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.9
Once launched, the server accepts requests on port 8000 at the path /v1/chat/completions — the format matches OpenAI API, so the openai libraries for Python and JavaScript connect without changes: just swap the base_url.
How to Speed Things Up and Save Memory
Speed and memory usage are controlled by launch parameters. Main tricks:
- Quantizing a model to AWQ or GPTQ cuts VRAM usage almost in half without a noticeable quality loss.
- The
--max-model-lenparameter limits context length and frees memory for the request queue. - Continuous batching in vLLM merges parallel requests automatically, no separate configuration needed.
- Multiple video cards connect through the
--tensor-parallel-sizeparameter, letting you run models larger than a single GPU.
When video memory is still short even with quantization, an option is running on CPU through llama.cpp and the GGUF format: slower, but no GPU required.
What to Do About Out-of-Memory Errors
The CUDA out of memory error is a common problem on the first launch. Steps to take:
- Lower the
--gpu-memory-utilizationvalue to 0.8 and restart the server. - Reduce
--max-model-lenwhen the model's full context is excessive for the task. - Use a quantized version of the model instead of full fp16 precision.
- Check for stuck processes on the GPU with
nvidia-smiand stop them withkill. - If the model still does not fit, recalculate the required amount and pick a GPU with more VRAM headroom.
After setup, send a test request through curl and compare the response speed with a plain model run. The basic vLLM installation is complete at this point — next, the server gets connected to a web interface or embedded into a product.