An embedding is just a list of numbers that places a piece of text — or an image, or a snippet of audio — at a point in a high-dimensional space, arranged so that things with similar meaning land near each other. Once you accept that, "similarity search" stops being mysterious: it's literally geometry. Find the points closest to your query's point, and you've found the most similar items.
Why semantic search works
That's the whole trick behind semantic search. Keyword search matches characters; embedding search matches meaning — so a query like "how do I reset my password" can surface a document titled "account recovery steps" that shares not one keyword with it. The model has learned, from vast amounts of text, that those two phrases live in the same neighbourhood.
The vector database
A vector database is the specialised tool for finding those nearest neighbours fast. Comparing a query against millions of vectors one at a time is too slow, so these systems use approximate nearest-neighbour indexes — trading a sliver of exactness for a large speedup. That word approximate matters: occasionally the true closest match is missed. For recommendations that's usually fine; for some retrieval tasks you need to know it can happen.
The distance metric is a decision
The distance metric is a real decision, not a checkbox. Cosine similarity (the angle between two vectors) and dot product answer subtly different questions, and — this one bites people — mixing embeddings from different models in the same index is meaningless, because their coordinate systems don't line up. One embedding model per index, and stay consistent.
import numpy as np
def cosine(a, b):
# 1.0 = same direction, 0 = unrelated, -1 = opposite.
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
# "nearest neighbours" is just: sort every item by distance to the query.
ranked = sorted(items, key=lambda x: cosine(query_vec, x.vec), reverse=True)The mental model
The mental model I keep is this: embeddings turn "is this similar?" — a fuzzy human judgement — into "how far apart are these points?", a number a computer can sort in milliseconds. Almost everything interesting in modern search, recommendation, and RAG is built on that single translation.
Sources & further reading
- Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013) — word2vec, where "meaning as geometry" went mainstream.
- Reimers & Gurevych, Sentence-BERT (2019) — embeddings for whole sentences, not just words.
- Malkov & Yashunin, Approximate nearest neighbor search using HNSW graphs (2018) — how the search stays fast at scale.
Editorial note — A conceptual explainer. The techniques (embeddings, approximate nearest-neighbour search, cosine vs. dot-product) are well established; no specific database, model, or benchmark is endorsed.
