All posts
// / Blog

We set up 43 monitoring alerts for our ML system. Within a week, we turned off 38 of them.

Alert fatigue is real. When everything is alerting, nothing is alerting.

The mistake: alerting on every metric deviation. A 2% change in prediction distribution at 3 AM is not an emergency. But a 15% drop in key business metric during peak hours is.

What I alert on now — the minimal set that catches real problems:

Business metric degradation (the metric stakeholders care about).

Significant distribution shift in model outputs (more than 2 standard deviations from baseline, sustained for more than 1 hour).

Data pipeline failures (upstream data not arriving on schedule).

Latency SLA breaches (p99 above threshold).

Error rate spikes.

That's five alerts. Not forty-three. Each one is actionable — when it fires, there's a clear runbook for what to investigate and what to do.

Everything else gets logged to a dashboard for periodic review but doesn't page anyone.

The principle: alert on things that require immediate human action. Dashboard everything else. And review your alerts quarterly — if an alert hasn't fired in 3 months, consider if it's still relevant.

Good monitoring is about signal-to-noise ratio. Maximize signal. Minimize noise. Protect your team's attention.

#Monitoring#MLOps#Observability#AlertFatigue#MachineLearning#SRE