Retrieval & data
Embedding
Definition
An embedding is a list of numbers representing the meaning of a piece of text, such that semantically similar texts have mathematically similar vectors. Embeddings make meaning-based search possible.
An embedding model converts text into a fixed-length vector, typically a few hundred to a few thousand dimensions. The training objective places related concepts near each other in that space, so "car" sits closer to "automobile" than to "banana" despite sharing no letters.
This is what enables semantic search. Comparing embedding vectors finds documents about the same topic even when they use entirely different vocabulary — something keyword search fundamentally cannot do.
Embeddings from different models are not comparable. If you change embedding models, you must re-embed your entire corpus.
Related terms
RAG (Retrieval-Augmented Generation)
RAG retrieves relevant passages from your own documents and inserts them into the prompt before the model answers. It grounds responses in your data, cuts hallucination, and needs no retraining.
Vector database
A vector database stores embeddings and finds the most similar ones to a query vector quickly. It is the retrieval layer of most RAG systems. Examples include Pinecone, Weaviate, Qdrant and pgvector.
Semantic search
Semantic search finds results by meaning rather than exact keywords, using embeddings to compare concepts. It matches "how do I get my money back" to a document titled "Refund Policy".
Cosine similarity
Cosine similarity measures how closely two vectors point in the same direction, on a scale from -1 to 1. It is the standard way to compare embeddings, because it captures semantic similarity while ignoring text length.
Chunking
Chunking splits documents into smaller passages before embedding them for retrieval. Chunk size is a key quality lever: too large dilutes relevance, too small loses the context needed to make sense.
Reranking
Reranking takes an initial set of retrieved candidates and reorders them with a more accurate but slower model. It is one of the cheapest ways to materially improve RAG quality.
Put this into practice
Understanding the term is step one. Our free courses and tools let you actually use it.