I once built a semantic search system where "dog food" returned results about "hot dogs." The…
embeddings were technically correct — both contain "dog" — but practically useless.
Embedding models capture similarity in ways that don't always align with what humans mean by "similar." This is the gap that trips up most engineers building search and RAG systems.
Lessons from debugging hundreds of embedding-related issues:
The embedding model matters enormously. OpenAI's ada-002 and text-embedding-3-large behave very differently on the same data. Always benchmark on YOUR specific use case.
Domain-specific fine-tuning of embedding models is underrated. A general embedding model treats "bank" (river) and "bank" (financial) identically. A finance-domain model doesn't.
Chunk size affects embedding quality more than most people realize. Too small and you lose context. Too large and the embedding becomes a blurry average of multiple concepts.
Hybrid search (combining semantic embeddings with keyword matching) beats pure semantic search in almost every production system I've built. Keywords catch the exact matches that embeddings miss, and embeddings catch the semantic matches that keywords miss.
Embeddings aren't magic. They're useful, imperfect tools that need engineering judgment to work well.