Learning path · Embeddings & Representation · 35
Embeddings
Dense vector representations of text (or other modalities) where semantic similarity approximates geometric closeness.
Why it matters
- Foundation of semantic search, RAG retrieval, and clustering.
- Choice of embedding model affects recall on domain jargon.
- Separate from generative LLM—you often use both.
Key ideas
- Vector representations
- Similarity search
- Domain adaptation
Top resources
- 01DocsOpenAI
Embeddings guide
Why this resource. How to obtain vectors and what similarity is allowed to mean.
Covers in this concept
- embedding models
- vector space
- 02DocsPinecone
Understanding embeddings
Why this resource. Geometric intuition for why nearest neighbors retrieve meaning.
Covers in this concept
- similarity
- dimensions
Embeddings turn a sentence into a vector so similar meanings sit nearby. That is how you find a passage without sharing keywords. Use an embedder trained on text like yours; legal and code often need their own. When you change the model, re-embed the corpus. Mixing two embedding versions in one index produces similarity scores that look precise and are nonsense.
Updated 2026-08-09 · Full learning path