The scariest production bug I ever encountered: the model was right 99% of the time and…
catastrophically wrong the other 1%.
It was a medical triage system. Worked beautifully on routine cases. But for rare conditions — the ones where getting it right matters most — it confidently misclassified them as low-priority.
Why? The training data had 50,000 routine cases and 47 rare ones. The model learned to be a very good "it's probably fine" machine.
Took me a week to diagnose because the aggregate metrics looked perfect. The fix wasn't technical — it was spending three days with domain experts labeling 500 synthetic rare cases and rebalancing the training set.
Lesson that stuck: always, ALWAYS check your model's performance on the tails of the distribution. The average hides the disasters.