We set up 43 monitoring alerts for our ML system. Within a week, we turned off 38 of them.
Alert fatigue is real. When everything is alerting, nothing is alerting.
The mistake: alerting on every metric deviation. A 2% change in prediction distribution at 3 AM is not an emergency. But a 15% drop in key business metric during peak hours is.
What I alert on now — the minimal set that catches real problems:
Business metric degradation (the metric stakeholders care about).
Significant distribution shift in model outputs (more than 2 standard deviations from baseline, sustained for more than 1 hour).
Data pipeline failures (upstream data not arriving on schedule).
Latency SLA breaches (p99 above threshold).
Error rate spikes.
That's five alerts. Not forty-three. Each one is actionable — when it fires, there's a clear runbook for what to investigate and what to do.
Everything else gets logged to a dashboard for periodic review but doesn't page anyone.
The principle: alert on things that require immediate human action. Dashboard everything else. And review your alerts quarterly — if an alert hasn't fired in 3 months, consider if it's still relevant.
Good monitoring is about signal-to-noise ratio. Maximize signal. Minimize noise. Protect your team's attention.