We deployed a model with a bias we didn't catch. A user caught it. They were upset. Rightfully so.
The model performed 18% worse for users with non-English names in our system. Not because of the ML model itself — because the preprocessing pipeline handled Unicode characters poorly, corrupting non-ASCII names before they reached the model.
A technical bug with discriminatory impact. The hardest kind to catch because it doesn't look like bias — it looks like a bug.
What we changed after that incident:
Pre-deployment bias testing became mandatory. Not optional. Not "when we have time." Every model, every deployment.
We test across demographic segments explicitly: age, gender, geography, language, and any other relevant dimension.
We set performance floors: the model must achieve at least X% accuracy on EVERY segment, not just on average. If any segment falls below the floor, deployment is blocked.
We created a rapid response plan for bias reports. When a user reports unfair behavior, it triggers an immediate investigation, not a ticket in the backlog.
Catching bias before deployment is ideal. Having a system to respond quickly when you miss something is essential.
Perfect fairness might be impossible. But treating bias as a bug — something to detect, fix, and prevent — is the minimum standard.