We changed one word in our system prompt and broke the entire application.
The word "must" was changed to "should." The LLM started treating previously strict format requirements as optional. JSON outputs became inconsistent. Downstream parsers failed silently.
One word. Millions of predictions affected.
This is why prompt versioning is critical and why "just updating the prompt in the code" is dangerous.
How I manage prompts now:
Every prompt is stored in version control with a unique version identifier. Changes go through review, just like code changes.
Prompt templates are separate from application code. Changing a prompt doesn't require a code deployment.
A/B testing for prompt changes. The new prompt runs alongside the old one on a subset of traffic before full rollout.
Evaluation runs automatically when a prompt changes. If quality drops below threshold, the change is blocked.
Rollback capability. If a prompt change causes issues in production, reverting takes seconds because we know exactly which version was running before.
The tooling: LangSmith, Promptfoo, or even a simple Git repository dedicated to prompts. The specific tool matters less than the practice.
Prompts are code. They affect system behavior. They should be versioned, reviewed, tested, and deployed with the same rigor.
Treat your prompts as production artifacts, not as casual text files.