All posts
// / Blog

A model I deployed for a client started giving increasingly weird predictions.

Nobody noticed for six weeks.

By the time they flagged it, the model had been confidently making bad decisions for 10,000+ transactions. The input data had shifted gradually — a supplier changed their data format — and the model's performance degraded slowly enough that no one caught it.

That experience made me obsessive about monitoring.

ML models degrade silently. Unlike traditional software that crashes visibly, a deteriorating model keeps running, keeps serving predictions, and keeps looking fine on the surface. The only way to catch problems is active monitoring.

What I track now: prediction distribution shift (are outputs changing?), feature drift (are inputs changing?), performance metrics with real labels when available, latency and throughput, error rates on edge cases.

Evidently AI is excellent open source tooling for this. Whylabs for data profiling. Grafana + Prometheus for infrastructure. Custom Streamlit dashboards for business stakeholders.

The rule I follow: if you can't see it degrading, you can't fix it before users suffer. Deploy monitoring with the same urgency as the model itself.

#MLMonitoring#ModelDrift#MLOps#Observability#MachineLearning#Production