Skip to main content

Semantic Search for a Website Using Embeddings

AI Agents on VPS · 29.09.2026

How Semantic Search Differs from Full-Text Search

Regular website search looks for matching words: a query like "how to get a refund" will not find an article titled "refund policy" if the text has none of the exact query words. Semantic search compares meaning, not letters — it turns the query and article texts into numeric vectors (embeddings) and finds the closest matches by meaning, even when the words differ.

This kind of search is useful for a knowledge base, a product catalog, and support desks where users phrase questions in their own words. The result is almost always more accurate than full-text keyword search, and implementation takes one working day on an average VDS.

How Embeddings Work and Which Model to Choose

An embedding model turns text into a fixed-length vector — usually from 384 to 1536 numbers. Texts with similar meaning get close vectors, and the distance between vectors (cosine similarity) shows how alike they are.

ModelDimensionsLanguagesWhere to run
all-MiniLM-L6-v2384mostly EnglishCPU, easy
multilingual-e5-base768many, including UkrainianCPU or GPU
text-embedding-3-small1536manyAPI only

For a site with Russian and Ukrainian content, multilingual-e5-base is a good fit — it runs locally on a VDS without API keys and without the cost of an external service.

Setting Up the Environment and Generating Embeddings

Install the sentence-transformers library and calculate vectors for every article in the knowledge base.

python3.11 -m venv /opt/search-env
source /opt/search-env/bin/activate
pip install sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("intfloat/multilingual-e5-base")
texts = ["How to set up backups", "What to do about a 500 error"]
vectors = model.encode(texts, normalize_embeddings=True)
print(vectors.shape)

Each call to encode returns an array of numbers for every text. Normalizing vectors simplifies later comparison through a dot product instead of a full cosine calculation.

Where to Store Vectors and Search for the Closest Matches

Ready-made vectors need somewhere to live and a fast way to search the closest ones to a user's query. There are two main paths: add a vector type right into an existing relational database, or run a separate specialized database.

  • For a project already running on PostgreSQL, the simplest option is an extension — see pgvector in PostgreSQL: one SQL query finds the nearest articles without a separate service.
  • For large catalogs with millions of documents, it is worth comparing specialized vector databases — an overview is given in the article comparing vector databases.
  • For a ready-made combination of search and a language model answer, the scheme from the article RAG on a VDS comes in handy.

How to Add Search to a Website

Adding semantic search to an existing website usually follows the same steps regardless of the chosen database:

  • Calculate embeddings for all current articles and store them together with the material's id.
  • Add a background job that recalculates the vector whenever an article is published or edited.
  • On the backend, accept the search query, turn it into a vector with the same model, and search for the closest records.
  • Return not just the title to the user but also a relevance score — this helps drop weak matches below a threshold.

The threshold is usually tuned by trial: for multilingual-e5-base, a cosine similarity above 0.75 is a good starting point.

Common Mistakes During Implementation

Most problems come not from the model itself but from how the pipeline around it is organized:

  • Using different models for indexing and for search — vectors from different models are not comparable, and the result is always meaningless.
  • Skipping the embedding recalculation after editing an article — search returns the outdated version of the text.
  • Texts that are too long without paragraph breaks — the model averages the meaning of the whole article and loses detail.
  • Ignoring the relevance threshold — without a cutoff, search always returns something, even when there is no good answer.

Once the pipeline is set up, semantic search can be extended further: add a hybrid mode with regular full-text search, or connect a separate model for reranking results.

← Back to Knowledge Base Ask Support