How Semantic Search Differs from Full-Text Search
Regular website search looks for matching words: a query like "how to get a refund" will not find an article titled "refund policy" if the text has none of the exact query words. Semantic search compares meaning, not letters — it turns the query and article texts into numeric vectors (embeddings) and finds the closest matches by meaning, even when the words differ.
This kind of search is useful for a knowledge base, a product catalog, and support desks where users phrase questions in their own words. The result is almost always more accurate than full-text keyword search, and implementation takes one working day on an average VDS.
How Embeddings Work and Which Model to Choose
An embedding model turns text into a fixed-length vector — usually from 384 to 1536 numbers. Texts with similar meaning get close vectors, and the distance between vectors (cosine similarity) shows how alike they are.
| Model | Dimensions | Languages | Where to run |
|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | mostly English | CPU, easy |
| multilingual-e5-base | 768 | many, including Ukrainian | CPU or GPU |
| text-embedding-3-small | 1536 | many | API only |
For a site with Russian and Ukrainian content, multilingual-e5-base is a good fit — it runs locally on a VDS without API keys and without the cost of an external service.
Setting Up the Environment and Generating Embeddings
Install the sentence-transformers library and calculate vectors for every article in the knowledge base.
python3.11 -m venv /opt/search-env
source /opt/search-env/bin/activate
pip install sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-base")
texts = ["How to set up backups", "What to do about a 500 error"]
vectors = model.encode(texts, normalize_embeddings=True)
print(vectors.shape)
Each call to encode returns an array of numbers for every text. Normalizing vectors simplifies later comparison through a dot product instead of a full cosine calculation.
Where to Store Vectors and Search for the Closest Matches
Ready-made vectors need somewhere to live and a fast way to search the closest ones to a user's query. There are two main paths: add a vector type right into an existing relational database, or run a separate specialized database.
- For a project already running on PostgreSQL, the simplest option is an extension — see pgvector in PostgreSQL: one SQL query finds the nearest articles without a separate service.
- For large catalogs with millions of documents, it is worth comparing specialized vector databases — an overview is given in the article comparing vector databases.
- For a ready-made combination of search and a language model answer, the scheme from the article RAG on a VDS comes in handy.
How to Add Search to a Website
Adding semantic search to an existing website usually follows the same steps regardless of the chosen database:
- Calculate embeddings for all current articles and store them together with the material's id.
- Add a background job that recalculates the vector whenever an article is published or edited.
- On the backend, accept the search query, turn it into a vector with the same model, and search for the closest records.
- Return not just the title to the user but also a relevance score — this helps drop weak matches below a threshold.
The threshold is usually tuned by trial: for multilingual-e5-base, a cosine similarity above 0.75 is a good starting point.
Common Mistakes During Implementation
Most problems come not from the model itself but from how the pipeline around it is organized:
- Using different models for indexing and for search — vectors from different models are not comparable, and the result is always meaningless.
- Skipping the embedding recalculation after editing an article — search returns the outdated version of the text.
- Texts that are too long without paragraph breaks — the model averages the meaning of the whole article and loses detail.
- Ignoring the relevance threshold — without a cutoff, search always returns something, even when there is no good answer.
Once the pipeline is set up, semantic search can be extended further: add a hybrid mode with regular full-text search, or connect a separate model for reranking results.