Model choice mostly comes down to how much video memory you have and the task at hand: a model that does not fit in memory at once starts running several times slower, because part of its weights keeps getting reloaded.
What a self-hosted LLM is and how it differs from a third-party API
A self-hosted LLM is a language model that runs on your own server instead of being called through a third-party API. You take on the hardware, but your data never leaves your infrastructure.
The price for that is usually lower answer quality compared with the largest closed models available only through an API, and the cost shifts from paying per token to paying for a server that runs constantly, whether or not you are calling the model right now.
The main limit is memory
Before picking a model by name, it helps to know how much memory it takes at different levels of compression. The figures below are approximate: exact numbers depend on the specific build, but the order of magnitude holds.
| Model class (parameters) | Full precision | 8-bit | 4-bit |
|---|---|---|---|
| 7-8 billion | about 15 GB | about 8 GB | about 4-5 GB |
| 13-14 billion | about 28 GB | about 14 GB | about 7-8 GB |
| 30-34 billion | about 65 GB | about 33 GB | about 17 GB |
| 70 billion and up | about 140 GB | about 70 GB | about 35-40 GB |
The practical takeaway: a model should fit into available memory with room to spare for context and internal buffers, not exactly at the edge.
What quantization is and what it costs you
Quantization compresses a model's weights to a smaller bit width: instead of 16-bit numbers, it uses 8-bit or 4-bit ones. This cuts memory use and speeds up inference on the same hardware.
The cost is a gradual loss of precision: the heavier the compression, the higher the risk that the model starts losing track of details and long reasoning. For most practical tasks, 8-bit and moderate 4-bit compression work acceptably, but it is worth testing on your own data rather than trusting general descriptions.
Choosing a model for the task
Different tasks call for different model classes, and there is no point paying for memory a task does not need.
| Task | Model class | What to watch |
|---|---|---|
| Chat and messaging | 7-8 billion parameters is usually enough | how well it holds conversation context |
| Coding help | 13-14 billion parameters and up | a separate version trained on code |
| Text translation | 7-8 billion parameters | support for the language pair you need |
| Extracting data from documents | 7-8 billion parameters with light compression | robustness to formatting and typos in the source |
| Vector embeddings for search | a separate, more compact model class | vector size and speed, not conversational skill |
CPU or GPU
A model also runs without a GPU if it fits in system memory, but the answer forms noticeably slower, word by word, and each request can take seconds or longer to finish.
A GPU becomes necessary when answers need to come fast and often, for example in a service with live users. For rare background tasks with no hard speed requirement, a processor is sometimes enough.
How to check a model fits before building a service on it
Three signs show up in the very first requests, and they are worth checking before an entire service gets built around the model.
- The model holds exactly the context length the task needs, instead of losing the start of a conversation
- The answer comes back within an acceptable time on real hardware, not on paper
- Quality on your own real examples, not demo ones, stays stable
What to do next
Once the model class is clear, the next step is setup and the tooling around it. Installation is covered in the article on setting up Ollama on a VDS, and running local models without external APIs is covered in the article on local LLMs on a VPS.