Your RAG system is only as good as your evaluation. And most RAG evaluations are terrible.
The standard approach: throw 10 questions at it, eyeball the answers, say "looks good." That's not evaluation. That's vibes.
Here's the evaluation framework I use for every RAG system:
Retrieval quality: for each question, are the RIGHT documents being retrieved? Measure this separately from answer quality. If retrieval fails, no amount of LLM magic will save you.
Answer faithfulness: is the answer actually supported by the retrieved documents? Or is the LLM ignoring the context and hallucinating?
Answer relevance: does the answer actually address the question? (LLMs love to give technically correct but irrelevant responses.)
Context precision: of the documents retrieved, how many were actually useful? Retrieving 10 documents when only 2 are relevant wastes context window and dilutes quality.
Build a test set of at least 50 question-answer pairs with annotated ground truth. Automate the evaluation with RAGAS or DeepEval. Run it every time you change anything — the model, the chunking, the retrieval, the prompt.
The teams that systematically evaluate their RAG systems iterate 5x faster than teams that eyeball it. Because they know exactly WHAT is broken, not just THAT something is broken.