We deployed a model update on a Thursday afternoon. Monday morning, customer complaints tripled.
The new model was objectively better on our test set. But it was worse on a specific segment of high-value customers that our test set underrepresented.
Rollback took 4 hours because we didn't have proper model versioning. The previous model weights were on someone's laptop. The configuration was in a Slack thread. The preprocessing code had been updated in-place.
This is why model versioning isn't a nice-to-have:
Every model gets a unique version ID. Weights, config, preprocessing code, and training data reference are ALL stored together as a single artifact. The model registry (MLflow is my go-to) tracks lineage — which data version, which code version, which training run.
Deployment uses model versions, not "latest." Rolling back means switching a version pointer, not reconstructing an artifact.
And the rule that saved me countless times: never deploy on a Friday. Actually, never deploy without the ability to roll back in under 5 minutes.
The cost of proper versioning is a few hours of setup. The cost of not having it is measured in customer trust and engineer sleep.