Skip to main content

llama.cpp and GGUF: Running Language Models on CPU

AI Agents on VPS · 29.09.2026

What Is llama.cpp and the GGUF Format

llama.cpp is a C++ implementation of language model inference that works without a video card, using only the CPU and RAM. Models are stored in the GGUF format: a single file holding all weights, metadata, and the tokenizer, easy to download and run with one command. The format replaced the older GGML and became the standard for CPU inference.

This setup fits a VDS without a GPU, tests on modest hardware, and cases where a model is needed rarely and keeping an expensive video card is not worth it. If a server already has a GPU, vLLM runs faster, and llama.cpp stays as a backup option or the main choice on a budget plan.

How GGUF Quantization Levels Differ

Quantization reduces the precision of model weights to shrink the file and speed up computation on CPU. The lower the bit count, the smaller the size and memory needs, but the higher the loss in answer quality.

QuantizationBits per weight7B model sizeQuality
Q8_08~7.2 GBalmost like fp16
Q6_K6~5.5 GBvery close to the original
Q5_K_M5~4.8 GBgood balance
Q4_K_M4~4.1 GBnoticeable but small loss
Q2_K2~2.8 GBstrong degradation

For most tasks Q4_K_M or Q5_K_M is optimal — they save memory and barely damage the model's answers.

Installing llama.cpp on a Server

Build the project from source: this gives you the current version and optimization for the specific server's CPU.

sudo apt update
sudo apt install -y build-essential cmake git
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j

After the build, binaries appear in the build/bin directory. Check the number of CPU cores with the nproc command — this value is passed at launch to fully load the CPU.

How to Download a Model and Run Inference

Ready-made GGUF files are published on Hugging Face — just download one file with the needed quantization level and pass its path to the binary.

wget https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF/resolve/main/llama-2-7b-chat.Q4_K_M.gguf
./build/bin/llama-cli -m llama-2-7b-chat.Q4_K_M.gguf -p "Describe a cloud server briefly" -n 200 -t 8

The -t parameter sets the number of CPU threads, and -n sets the maximum answer length. For continuous use it is more convenient to run a server with an HTTP API — see also the ready-made interface Text Generation WebUI, which can run on top of llama.cpp.

How to Pick Quantization for the Server's Memory

Base your choice on free RAM, not just the model file size — running inference needs headroom for context and the system.

  • 8 GB RAM — 7B models in Q4_K_M, with no parallel tasks on the server.
  • 16 GB RAM — 7B models in Q6_K or Q8_0, or 13B models in Q4_K_M.
  • 32 GB RAM — 13B models in Q6_K or 30B models in Q4_K_M.
  • Free memory should always stay 20-30% above the model file size to cover context and the KV cache.

For an exact video memory calculation in a GPU scenario, see the article how much VRAM a model needs — the calculation principle is similar, only RAM is counted instead of VRAM.

Common Errors on the First Launch

Most startup problems can be fixed without reinstalling anything:

  • The process exits immediately — there is not enough RAM for the chosen quantization level, pick a smaller file.
  • The answer generates very slowly — check the -t parameter and make sure it matches the number of physical cores on the server.
  • The model answers incoherently — the quantization level is likely too aggressive, such as Q2_K, raise it to Q4_K_M.
  • The build fails at the cmake step — update the compiler and make sure the build-essential package is installed.

After a successful CLI test, llama.cpp is ready to embed into scripts or run through a web interface. This setup needs no GPU and fits most VDS plans with enough RAM.

← Back to Knowledge Base Ask Support