Our LLM application was working fine. Until it wasn't. And we had no idea what changed.
The problem: LLM behavior is non-deterministic. The same input can produce different outputs. So when quality degrades gradually, it's incredibly hard to pinpoint when and why.
This is why observability for LLM applications is different from traditional application monitoring.
What you need to track: every prompt sent to the model (not just the user input — the full prompt with system message and context). Every response received. Latency per request. Token usage and cost. Retrieval quality for RAG (what was retrieved, was it relevant). User feedback signals.
Tools that work: LangSmith for LangChain-based apps. Helicone for API proxy-level logging. Weights & Biases Prompts. Or a custom solution with structured logging to a database.
The pattern that saves you: tag every request with metadata (user segment, query type, model version). When quality drops, you can filter to find exactly which queries are affected and correlate with what changed.
We found our issue: a prompt template update that improved quality for 95% of queries but catastrophically degraded a specific query type. Without granular observability, we would have been debugging for days. With it, the filter took 5 minutes.