Skip to main content

Setting Up vLLM on a Server for Fast LLM Inference

AI Agents on VPS · 29.09.2026

What Is vLLM and Why Run It on a Server

vLLM is a library for fast inference of large language models on GPU. It uses PagedAttention technology, which saves video memory and multiplies throughput compared to a plain run through Transformers. vLLM starts an OpenAI-compatible HTTP server, so existing ChatGPT API clients switch to your own model without rewriting code.

Teams pick this library when a server handles many parallel requests: a chat widget on a website, an internal company assistant, or an API for a mobile app. For single experiments on CPU, llama.cpp fits better, while vLLM shows its strength on a GPU server with one or several video cards.

Server Requirements for Running vLLM

Check server parameters before installing. vLLM needs an NVIDIA GPU with CUDA support and enough video memory — the exact amount depends on model size and quantization.

ParameterMinimumRecommended
GPUNVIDIA, 8 GB VRAMNVIDIA, 24 GB VRAM
NVIDIA driver525550
CUDA12.112.4
Python3.103.11
RAM16 GB32 GB

Calculate the exact VRAM amount for a given model in advance — see VRAM calculation for a model.

Installing vLLM on a VDS

Install the NVIDIA driver and CUDA Toolkit, then create a Python virtual environment and install the package through pip.

sudo apt update
sudo apt install -y nvidia-driver-550 python3.11-venv
python3.11 -m venv /opt/vllm-env
source /opt/vllm-env/bin/activate
pip install --upgrade pip
pip install vllm

Check that the system sees the GPU with the nvidia-smi command. If the card does not show up, reboot the server and make sure the nouveau module is disabled in the kernel configuration.

How to Launch an OpenAI-Compatible Server

vLLM ships with a ready-made server that mirrors the OpenAI API request format. Launch it with a single command, specifying a model from Hugging Face.

python -m vllm.entrypoints.openai.api_server --model mistralai/Mistral-7B-Instruct-v0.2 --host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.9

Once launched, the server accepts requests on port 8000 at the path /v1/chat/completions — the format matches OpenAI API, so the openai libraries for Python and JavaScript connect without changes: just swap the base_url.

How to Speed Things Up and Save Memory

Speed and memory usage are controlled by launch parameters. Main tricks:

  • Quantizing a model to AWQ or GPTQ cuts VRAM usage almost in half without a noticeable quality loss.
  • The --max-model-len parameter limits context length and frees memory for the request queue.
  • Continuous batching in vLLM merges parallel requests automatically, no separate configuration needed.
  • Multiple video cards connect through the --tensor-parallel-size parameter, letting you run models larger than a single GPU.

When video memory is still short even with quantization, an option is running on CPU through llama.cpp and the GGUF format: slower, but no GPU required.

What to Do About Out-of-Memory Errors

The CUDA out of memory error is a common problem on the first launch. Steps to take:

  • Lower the --gpu-memory-utilization value to 0.8 and restart the server.
  • Reduce --max-model-len when the model's full context is excessive for the task.
  • Use a quantized version of the model instead of full fp16 precision.
  • Check for stuck processes on the GPU with nvidia-smi and stop them with kill.
  • If the model still does not fit, recalculate the required amount and pick a GPU with more VRAM headroom.

After setup, send a test request through curl and compare the response speed with a plain model run. The basic vLLM installation is complete at this point — next, the server gets connected to a web interface or embedded into a product.

← Back to Knowledge Base Ask Support