Skip to main content

How to Choose a Local LLM Model for Your Task

AI Agents on VPS · 09.10.2026 · 4 min read
Illustration for “How to Choose a Local LLM Model for Your Task”

Model choice mostly comes down to how much video memory you have and the task at hand: a model that does not fit in memory at once starts running several times slower, because part of its weights keeps getting reloaded.

What a self-hosted LLM is and how it differs from a third-party API

A self-hosted LLM is a language model that runs on your own server instead of being called through a third-party API. You take on the hardware, but your data never leaves your infrastructure.

The price for that is usually lower answer quality compared with the largest closed models available only through an API, and the cost shifts from paying per token to paying for a server that runs constantly, whether or not you are calling the model right now.

The main limit is memory

Before picking a model by name, it helps to know how much memory it takes at different levels of compression. The figures below are approximate: exact numbers depend on the specific build, but the order of magnitude holds.

Model class (parameters)Full precision8-bit4-bit
7-8 billionabout 15 GBabout 8 GBabout 4-5 GB
13-14 billionabout 28 GBabout 14 GBabout 7-8 GB
30-34 billionabout 65 GBabout 33 GBabout 17 GB
70 billion and upabout 140 GBabout 70 GBabout 35-40 GB

The practical takeaway: a model should fit into available memory with room to spare for context and internal buffers, not exactly at the edge.

What quantization is and what it costs you

Quantization compresses a model's weights to a smaller bit width: instead of 16-bit numbers, it uses 8-bit or 4-bit ones. This cuts memory use and speeds up inference on the same hardware.

The cost is a gradual loss of precision: the heavier the compression, the higher the risk that the model starts losing track of details and long reasoning. For most practical tasks, 8-bit and moderate 4-bit compression work acceptably, but it is worth testing on your own data rather than trusting general descriptions.

Choosing a model for the task

Different tasks call for different model classes, and there is no point paying for memory a task does not need.

TaskModel classWhat to watch
Chat and messaging7-8 billion parameters is usually enoughhow well it holds conversation context
Coding help13-14 billion parameters and upa separate version trained on code
Text translation7-8 billion parameterssupport for the language pair you need
Extracting data from documents7-8 billion parameters with light compressionrobustness to formatting and typos in the source
Vector embeddings for searcha separate, more compact model classvector size and speed, not conversational skill

CPU or GPU

A model also runs without a GPU if it fits in system memory, but the answer forms noticeably slower, word by word, and each request can take seconds or longer to finish.

A GPU becomes necessary when answers need to come fast and often, for example in a service with live users. For rare background tasks with no hard speed requirement, a processor is sometimes enough.

How to check a model fits before building a service on it

Three signs show up in the very first requests, and they are worth checking before an entire service gets built around the model.

  • The model holds exactly the context length the task needs, instead of losing the start of a conversation
  • The answer comes back within an acceptable time on real hardware, not on paper
  • Quality on your own real examples, not demo ones, stays stable

What to do next

Once the model class is clear, the next step is setup and the tooling around it. Installation is covered in the article on setting up Ollama on a VDS, and running local models without external APIs is covered in the article on local LLMs on a VPS.

Was this article helpful?
← Back to Knowledge Base Ask Support