We shipped an LLM-powered feature last year with no evaluation framework.
"It looks good in testing" was our quality bar.
Within two weeks, users found that it hallucinated medical advice, gave different answers to the same question depending on phrasing, and sometimes just... made up citations.
Never again.
LLM evaluation is now the first thing I set up, not the last. Before writing application logic, I build the evaluation pipeline.
What to measure: factuality (is it making stuff up?), relevance (does it answer the actual question?), coherence (does the response make sense?), safety (will it say something harmful?), and instruction following (does it do what was asked?).
The tools have gotten quite good. DeepEval for comprehensive evaluation. RAGAS specifically for RAG systems. Promptfoo for systematic prompt testing. LangSmith for monitoring in production.
Three approaches, each with tradeoffs: automated metrics (fast but shallow), LLM-as-Judge where another LLM evaluates outputs (decent balance), and human evaluation (gold standard but expensive and slow).
You wouldn't ship code without tests. Shipping LLMs without evaluation is the same mistake — just with more creative failure modes.