Dashboards show metrics, but someone has to notice when something's wrong. Anomaly detection automates that.
Approaches
- Static thresholds: alert when a metric crosses a fixed value. Simple but ignores seasonality.
- Statistical bands: alert when a value falls outside expected ranges based on history.
- Seasonal models: compare against the same hour, day or week, accounting for trends.
- Machine learning: models forecasting expected values across many metrics.
Designing Useful Alerts
- Focus on metrics that trigger action.
- Account for seasonality, holidays and known events.
- Require anomalies to persist before alerting, to reduce noise.
- Include context: magnitude, affected segments, likely causes.
Root Cause Exploration
When a metric moves, break it down by dimensions — region, device, channel — to find where the change comes from.
Avoid Alert Fatigue
Too many false alarms and people stop paying attention. Tune thresholds and review alert quality regularly.
Data Issues First
Many "anomalies" are data pipeline failures. Check data freshness and completeness first.